* [NVIDIA GPU Operator v26.7.0](https://github.com/NVIDIA/gpu-operator) – Automates installation and lifecycle management of GPU drivers, container runtimes, and monitoring on Kubernetes nodes. * [Cog v0.22.0](https://github.com/replicate/cog) – Tool for packaging machine learning models in production-ready containers. * [ggrun v3.2.8](https://github.com/raketenkater/ggrun) – Auto-tuned launcher that measures multi-GPU hardware for GGUF models, picks an optimal llama.cpp/ik\_llama.cpp backend, and serves an OpenAI-compatible API. * [pg\_onnx v1.29.0](https://github.com/kibae/pg_onnx) – ONNX Runtime integration enabling machine learning inference within PostgreSQL databases. * [Cumo v0.5.10](https://github.com/sonots/cumo) – CUDA-aware GPU-optimized numerical library compatible with Ruby Numo for enhanced performance. * [Atlas Inference Engine b236](https://github.com/Avarok-Cybersecurity/atlas) – Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs. * [Beta9 gateway-0.1.757](https://github.com/beam-cloud/beta9) – Fast serverless runtime for GPU inference, isolated sandboxes, and scalable background jobs.