| 08/17 | 7 |
llm-d Router provides load and prefix-cache aware inference routing with prioritization and flow control, supporting standalone and Gateway API deployment via an endpoint picker.
|
| 08/14 | 7 |
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
|
| 08/24 | 6 |
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
|
| 08/12 | 6 |
ONNX Runtime integration enabling machine learning inference within PostgreSQL databases.
|
| 08/08 | 6 |
Platform to host, deploy, build, and fine-tune large language models with OpenAI API compatibility.
|
| 08/20 | 4 |
Kubernetes operator managing self-hosted LLM inference on NVIDIA GPUs and Apple Silicon, with autoscaling, model routing, and OpenAI-compatible API.
|
| 08/27 | 3 |
InferCrane provides an evidence-gated release lifecycle for self-hosted models behind a stable OpenAI-compatible endpoint, covering deploy, observe, scale, optimize, and safe promotion.
|