| 08/21 | 8 |
Automates installation and lifecycle management of GPU drivers, container runtimes, and monitoring on Kubernetes nodes.
|
| 08/20 | 6 |
Auto-tuned launcher that measures multi-GPU hardware for GGUF models, picks an optimal llama.cpp/ik_llama.cpp backend, and serves an OpenAI-compatible API.
|
| 08/17 | 6 |
Auto-tuned launcher that measures multi-GPU hardware for GGUF models, picks an optimal llama.cpp/ik_llama.cpp backend, and serves an OpenAI-compatible API.
|
| 08/23 | 5 |
CUDA-aware GPU-optimized numerical library compatible with Ruby Numo for enhanced performance.
|
| 08/20 | 5 |
CUDA-aware GPU-optimized numerical library compatible with Ruby Numo for enhanced performance.
|
| 08/17 | 5 |
CUDA-aware GPU-optimized numerical library compatible with Ruby Numo for enhanced performance.
|
| 08/17 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|
| 08/17 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|
| 08/16 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|
| 08/16 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|
| 08/16 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|
| 08/16 | 3 |
Fast serverless runtime for GPU inference, isolated sandboxes, and scalable background jobs.
|
| 08/16 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|