* [llm-d Router v0.10.0](https://github.com/llm-d/llm-d-router) – llm-d Router provides load and prefix-cache aware inference routing with prioritization and flow control, supporting standalone and Gateway API deployment via an endpoint picker. * [RunAnywhere v0.20.19](https://github.com/RunanywhereAI/runanywhere-sdks) – On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally. * [NVIDIA Cloud Functions (NVCF) deploy/helm/containe...](https://github.com/NVIDIA/nvcf) – Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters. * [pg\_onnx v1.29.0](https://github.com/kibae/pg_onnx) – ONNX Runtime integration enabling machine learning inference within PostgreSQL databases. * [AIKit v0.22.1](https://github.com/kaito-project/aikit) – Platform to host, deploy, build, and fine-tune large language models with OpenAI API compatibility. * [LLMKube llmkube-0.9.19](https://github.com/defilantech/LLMKube) – Kubernetes operator managing self-hosted LLM inference on NVIDIA GPUs and Apple Silicon, with autoscaling, model routing, and OpenAI-compatible API. * [InferCrane v1.0.0-rc.1](https://github.com/infercrane/infercrane) – InferCrane provides an evidence-gated release lifecycle for self-hosted models behind a stable OpenAI-compatible endpoint, covering deploy, observe, scale, optimize, and safe promotion.