Tag
Red Hat's AI blog series part 3 details how llm-d routes model inference traffic on Amazon EKS within their managed Kubernetes platform, using Kubernetes resources and Envoy integration for real-time decisions.
Red Hat AI team released quantized checkpoints for GLM-5.2 using NVFP4 and FP8 quantization, reducing model size by over 70% while maintaining high accuracy on GPQA. The quantized model, paired with the DSpark speculator, enables efficient deployment with vLLM.
Gemma 4 12B has been released under Apache 2.0, supporting multimodal inputs (text, image, audio, video), 256K context, built-in thinking, and native tool calling, running on Red Hat OpenShift AI.
Dozens of Red Hat packages were backdoored through the company's official NPM channel using the Shai-Hulud worm, which compromised Red Hat's CI/CD pipeline via GitHub Actions OIDC. Red Hat has removed the malicious packages and stated they were internal only, but the attack underscores escalating supply-chain risks.
A README for the RedHatInsights/javascript-clients monorepo that auto-generates Javascript API clients for Swagger/OpenAPI specs, using NX for monorepo management and GitHub Actions for CI/CD and NPM publishing.
IBM and Red Hat announce a $5 billion investment into Project Lightwell, a security clearinghouse that uses AI to identify and fix vulnerabilities in open source software, offering commercial subscriptions for enterprise use.
A technical discussion validates TurboQuant performance data on NVIDIA H100 GPUs with FP8 Tensor Cores and promises further insights from non-H100 testing.
Red Hat AI released a DFlash speculator model for Qwen3-8B, achieving 82.2% first-token acceptance on math reasoning tasks. The model was trained using the Speculators library and vLLM to optimize inference speed.