* [Ollama v0.33.1](https://github.com/ollama/ollama) – Tool for running and managing large language models. * [HelixML 2.12.7](https://github.com/helixml/helix) – Private GenAI stack for deploying AI agents with support for RAG, API calls, vision, and efficient GPU scheduling. * [go-llama v0.2.3](https://github.com/goccy/go-llama) – Pure Go inference engine for GGUF models, built from llama.cpp compiled to WebAssembly and translated to Go without wasm runtime.