Tag
This paper introduces CRAWO, a framework for adaptive workload orchestration of AI pipelines across heterogeneous edge infrastructures. It uses a control-loop model and Kubernetes-based implementation to improve workload distribution and reduce reliance on centralized cloud processing, demonstrated in a vehicle surveillance scenario.
A comprehensive guide on building an internal developer platform using Backstage, ArgoCD, and Crossplane, enabling teams to self-serve infrastructure and deployments without manual ticket-based processes.
Devin Outposts enables running Devin on any machine, including Mac mini, GPU boxes, VMs, private networks, or Kubernetes clusters, bringing the AI coding agent to local or private infrastructure.
An educational thread that explains Kubernetes' core concepts using relatable analogies like an Amazon warehouse and cruise control, breaking down components and the desired state mechanism.
A tweet lists key projects to build in inference engineering for understanding production LLM systems, including inference servers, paged KV cache, speculative decoding, quantization libraries, and guardrails.
A tutorial explaining how to build a bridge network for containers from scratch using standard Linux tools like network namespaces, veth pairs, bridges, and NAT, to demystify Docker and Kubernetes networking.
Discusses the operational challenges of deploying AI agents at scale, drawing a parallel to how Kubernetes solved container orchestration. Suggests the agent ecosystem needs a similar infrastructure breakthrough.
Explored NVIDIA Dynamo, a tool for deploying LLMs across multiple GPU cluster nodes with features like model caching, autoscaling, multinode deployments, and Kubernetes integration.
A solo developer shares the pros and cons of building and maintaining Luxury Yacht, a desktop app for Kubernetes cluster management, emphasizing the freedom and responsibility of solo development.
A developer ported Kubernetes to run entirely in the browser by rewriting core components in TypeScript, aided by AI for code conversion. The tool requires no installation and is designed for learning, teaching, and interview prep.
Tigera launches Lynx, a Kubernetes-native unified control plane for managing AI agents at scale, providing identity, policy enforcement, and real-time visibility across multi-cluster environments.
The article explores how platform engineering teams must adapt to AI workloads by managing costs and risks without overhauling existing infrastructure, highlighting the shift from traditional DevOps to a new paradigm.
PyTorchCon China is co-located with KubeCon, CloudNativeCon, and OpenInfra Summit in Shanghai from September 7-9, 2026, featuring technical sessions and community collaboration. Registration discounts are available until July 28.
OpenAI's voice AI architecture uses WebRTC with a split relay-transceiver design to handle low-latency audio for 900 million weekly active users.
Zalando's engineering team describes how they implemented client-side load balancing to replace the shared edge load balancer for their high-traffic Product Read API, reducing latency and improving observability by eliminating the fan-out bottleneck through Skipper.
A detailed account of tracking down and fixing a memory leak in the Kubernetes kubelet component in version 1.36, caused by a context leak in startPodSync. The author used Go pprof profiling to identify the issue and implemented a fix.
Ray 2.56 has been released with improvements to Ray Data, Ray Serve for LLMs, GPU-domain-aware placement groups, and Kubernetes integration.
Sam Rose from ngrok demonstrates how they ported Kubernetes to run entirely in the browser using WebAssembly, allowing developers to simulate a local cluster without needing a remote server or complex setup.
A developer criticizes Kubernetes' deployment method as too complex, advocating for using tmux sessions and while true loops to simply manage service processes, believing this approach has sufficient performance and is easy to maintain.
Netflix describes migrating their batch compute workloads from a custom solution (CMB) to Kueue, a Kubernetes-native job queueing system, to simplify management and leverage the Kubernetes ecosystem.