| 08/21 | 8 |
Automates installation and lifecycle management of GPU drivers, container runtimes, and monitoring on Kubernetes nodes.
|
| 08/13 | 8 |
Tool for packaging machine learning models in production-ready containers.
|
| 08/20 | 6 |
Auto-tuned launcher that measures multi-GPU hardware for GGUF models, picks an optimal llama.cpp/ik_llama.cpp backend, and serves an OpenAI-compatible API.
|
| 08/17 | 6 |
Auto-tuned launcher that measures multi-GPU hardware for GGUF models, picks an optimal llama.cpp/ik_llama.cpp backend, and serves an OpenAI-compatible API.
|
| 08/12 | 6 |
ONNX Runtime integration enabling machine learning inference within PostgreSQL databases.
|
| 08/23 | 5 |
CUDA-aware GPU-optimized numerical library compatible with Ruby Numo for enhanced performance.
|
| 08/20 | 5 |
CUDA-aware GPU-optimized numerical library compatible with Ruby Numo for enhanced performance.
|
| 08/17 | 5 |
CUDA-aware GPU-optimized numerical library compatible with Ruby Numo for enhanced performance.
|
| 08/13 | 5 |
CUDA-aware GPU-optimized numerical library compatible with Ruby Numo for enhanced performance.
|
| 08/09 | 5 |
CUDA-aware GPU-optimized numerical library compatible with Ruby Numo for enhanced performance.
|
| 08/17 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|
| 08/17 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|
| 08/16 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|
| 08/16 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|
| 08/16 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|
| 08/16 | 3 |
Fast serverless runtime for GPU inference, isolated sandboxes, and scalable background jobs.
|
| 08/16 | 3 |
Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.
|
| 08/10 | 3 |
Fast serverless runtime for GPU inference, isolated sandboxes, and scalable background jobs.
|
| 08/09 | 3 |
Fast serverless runtime for GPU inference, isolated sandboxes, and scalable background jobs.
|