Tag
The article presents research on distributed LLM inference for Intel PC fleets, focusing on pipeline-parallel sharded inference using OpenVINO with performance optimizations for heterogeneous hardware.