GitHub Release Tracker
All JS React Ruby Go Postgres Frontend Node

inference

Past 30d, sorted by best first, all versions
38 results Markdown version
08/17 7
llm-d/llm-d Router v0.10.0
llm-d Router provides load and prefix-cache aware inference routing with prioritization and flow control, supporting standalone and Gateway API deployment via an endpoint picker.
Go 304☆ 476d old #golang #ai #kubernetes #go #networking
08/14 7
RunanywhereAI/RunAnywhere v0.20.19
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/13 7
RunanywhereAI/RunAnywhere v0.20.18
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/11 7
RunanywhereAI/RunAnywhere v0.20.17
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/11 7
RunanywhereAI/RunAnywhere v0.20.16
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/11 7
RunanywhereAI/RunAnywhere v0.20.15
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/09 7
RunanywhereAI/RunAnywhere v0.20.14
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/08 7
RunanywhereAI/RunAnywhere v0.20.13
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/27 6
RunanywhereAI/RunAnywhere v0.20.31
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/27 6
RunanywhereAI/RunAnywhere v0.20.30
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/26 6
RunanywhereAI/RunAnywhere v0.20.29
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/25 6
RunanywhereAI/RunAnywhere v0.20.28
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/24 6
RunanywhereAI/RunAnywhere v0.20.27
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/24 6
NVIDIA/NVIDIA Cloud Functions (NVCF) deploy/helm/containe...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/22 6
RunanywhereAI/RunAnywhere v0.20.25
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/17 6
RunanywhereAI/RunAnywhere v0.20.24
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/12 6
kibae/pg_onnx v1.29.0
ONNX Runtime integration enabling machine learning inference within PostgreSQL databases.
08/08 6
kaito-project/AIKit v0.22.1
Platform to host, deploy, build, and fine-tune large language models with OpenAI API compatibility.
08/24 4
NVIDIA/NVIDIA Cloud Functions (NVCF) deploy/helm/nvca-ope...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/23 4
RunanywhereAI/RunAnywhere qairt-runtime-v2.47....
On-device mobile SDKs for running LLMs, speech-to-text, and text-to-speech locally.
C++ 10279☆ 403d old #swift #c++ #llm #ios #kotlin
08/20 4
defilantech/LLMKube llmkube-0.9.19
Kubernetes operator managing self-hosted LLM inference on NVIDIA GPUs and Apple Silicon, with autoscaling, model routing, and OpenAI-compatible API.
Go 196☆ 284d old #golang #ai #kubernetes #go #gpu
08/15 4
NVIDIA/NVIDIA Cloud Functions (NVCF) src/control-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/13 4
defilantech/LLMKube llmkube-0.9.17
Kubernetes operator managing self-hosted LLM inference on NVIDIA GPUs and Apple Silicon, with autoscaling, model routing, and OpenAI-compatible API.
Go 196☆ 284d old #golang #ai #kubernetes #go #gpu
08/10 4
defilantech/LLMKube llmkube-0.9.16
Kubernetes operator managing self-hosted LLM inference on NVIDIA GPUs and Apple Silicon, with autoscaling, model routing, and OpenAI-compatible API.
Go 196☆ 284d old #golang #ai #kubernetes #go #gpu
08/08 4
defilantech/LLMKube llmkube-0.9.15
Kubernetes operator managing self-hosted LLM inference on NVIDIA GPUs and Apple Silicon, with autoscaling, model routing, and OpenAI-compatible API.
Go 196☆ 284d old #golang #ai #kubernetes #go #gpu
08/27 3
infercrane/InferCrane v1.0.0-rc.1
InferCrane provides an evidence-gated release lifecycle for self-hosted models behind a stable OpenAI-compatible endpoint, covering deploy, observe, scale, optimize, and safe promotion.
Go 64☆ 20d old #golang #aws #go #gcp #inference
08/24 3
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/22 3
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/24 2
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/23 2
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/23 2
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/23 2
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/23 2
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/09 2
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/09 2
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/03 2
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/03 2
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.
08/01 2
NVIDIA/NVIDIA Cloud Functions (NVCF) src/compute-plane-se...
Platform for deploying, managing, and running GPU-accelerated inference, streaming, and batch workloads across worker clusters.