Tag
Nereus is a cost-aware runtime that adapts reinforcement learning post-training jobs for large language models to efficient execution plans, reducing latency by 27.7% and improving throughput up to 7.27 times over existing tools.
This article provides a comparative analysis of self-hosted inference orchestrators for AI model deployment as of September 2026, covering features like multi-machine support, cache-aware routing, and platform compatibility.
Lablup has joined the PyTorch Foundation as a Silver Member to advance open-source AI. They develop Backend.AI, a software tool for AI infrastructure that enables GPU orchestration and virtualization.
ENOVA is an open-source service that simplifies the deployment, monitoring, and auto-scaling of AI language models on GPU clusters, providing configuration recommendations and performance observability to improve efficiency.
This article presents a constraint-aware GPU allocator that improves GPU utilization by up to 33 percentage points compared to FIFO scheduling, demonstrating the critical role of priority-based allocation in enterprise AI systems.
The article argues that GPU utilization is becoming the key constraint in enterprise AI, analogous to aircraft utilization in aviation, and that idle GPUs represent wasted capacity that determines competitive advantage.
llama.cpp Console is a Windows desktop app that provides a GUI for managing llama.cpp in WSL/Ubuntu, handling installation, building, model downloading, and serving.
swm is an open-source tool that simplifies cloud GPU usage by installing frameworks like ComfyUI and Ollama in one command, and automatically saves your entire workspace between sessions, enabling seamless migration across providers.