Tag
GORGO introduces a proxy architecture for LLM inference that jointly optimizes network latency, prefill cost, and queueing delay using evolutionary strategy tuning on a new synthetic dataset, improving p95 TTFT by 6.9-15.5% and end-to-end latency by 14.3-30.9%.
NERVE proposes a network-aware bilinear tokenization method for self-supervised learning on brain functional connectivity matrices using masked autoencoders, improving representation learning across developmental cohorts.