Tag
Norway's National Library is building a sovereign Norwegian LLM using 2 PB of Huawei OceanStor Dorado flash storage for its AI training data pipeline, addressing the need for a local language model.
Snowflake now supports job-based batch inference powered by Ray, enabling distributed GPU execution for scaling model inference over millions of unstructured datapoints with a single API call.
The article discusses the importance of quality control for reinforcement learning data, outlining the shortcomings of current data vendors and the evaluation criteria used by frontier AI labs for RL data.
Anyscale is hosting a hands-on virtual lab session teaching developers how to build and scale data pipelines with Ray, covering video data curation, distributed GPU inference, and CPU/GPU streaming pipelines.
Prefect is an open-source workflow orchestration framework for building and automating data pipelines in Python.