GitHub Release Tracker
All JS React Ruby Go Postgres Frontend Node

inference

Past 14d, sorted by best first
5 results Markdown version
08/17 7
llm-d/llm-d Router v0.10.0
llm-d Router provides load and prefix-cache aware inference routing with prioritization and flow control, supporting standalone and Gateway API deployment via an endpoint picker.
Go 304☆ 476d old #golang #ai #kubernetes #go #networking
08/27 6
RunanywhereAI/RunAnywhere v0.20.31
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10280☆ 403d old #swift #c++ #llm #ios #kotlin
08/24 6
NVIDIA/NVIDIA Cloud Functions (NVCF) deploy/helm/containe...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/20 4
defilantech/LLMKube llmkube-0.9.19
Kubernetes operator managing self-hosted LLM inference on NVIDIA GPUs and Apple Silicon, with autoscaling, model routing, and OpenAI-compatible API.
Go 196☆ 284d old #golang #ai #kubernetes #go #gpu
08/27 3
infercrane/InferCrane v1.0.0-rc.1
InferCrane provides an evidence-gated release lifecycle for self-hosted models behind a stable OpenAI-compatible endpoint, covering deploy, observe, scale, optimize, and safe promotion.
Go 64☆ 20d old #golang #aws #go #gcp #inference