Tag
Release of SmolDataEnvs, a collection of 5,000 verifiable RL environment tasks for training small models in code and data science, fully open source.
This article details the implementation of a Parquet file writer in Haskell for the DataHaskell/Dataframe library, enabling efficient data serialization and interoperability with the data science ecosystem.
This article promotes a free live workshop on building AI with Python, offering a roadmap for data scientists to transition to AI/DS System Builder roles with practical demos.
The author shares their past efforts to achieve efficient single-expert loading for high-quality token-level classification at scale, echoing Maxime Rivest's thoughts on the utility of domain-specific expert models.
The Khipu Field Guide is a web-based resource that provides detailed symbolic renderings and interactive analyses of over 600 Incan khipus, aiming to democratize khipu research through computational methods and data science.
TabPFN-3.5 is released, claiming state-of-the-art performance for tabular data beyond IID small data settings, with features like handling grouped and temporal data, uncertainty calibration, and faster inference.
MIT CSAIL shares a beginner's guide to statistics, providing foundational educational content for data science and AI.
The article questions why software engineering jobs remain high compared to data science roles despite AI advancements, referencing US and EU labor statistics and growth projections for ICT professionals.
This paper introduces Calendar-SPCA, a calendar-structured sparse principal component method for interpretable representation learning in multi-periodic electricity consumption profiles, demonstrating high explained variance with increased sparsity and coherence.
A user shares insights from analyzing a decade's worth of lunch data using a state-of-the-art LLM, highlighting interesting findings about expenses and common restaurants.
The paper introduces StatFormBench, a benchmark for evaluating LLMs on statistical problem formulation, and finds that current models have significant limitations in classifying problems and identifying variables.
DS-Lighting is a unified harness toolkit that makes harness design explicit for data-science automation, improving reproducibility, comparability, and reliability in end-to-end workflows.
This paper presents the first application of data science to evaluate the UK Honours system using natural language processing, introducing a novel sentiment analysis algorithm called Minos to assess public opinion on honours recipients.
A Penn State-led study published in PLOS One finds that over half of U.S. adults report lacking basic statistical knowledge, yet most would rely on statistics more if they understood them better.
Rupert Young, chief product officer at MaxMind, discusses his career in cybersecurity and data science, emphasizing the company's GeoIP tool used for fraud prevention across streaming, security, and merchant industries.
Introduces a list on GitHub named Awesome Public Datasets, which compiles high-quality public datasets from around the world, helping data analysts save time searching for data.
This article discusses the parallels between Soviet economic planning and modern data science practices, focusing on issues like resource allocation and simplifying assumptions based on historical books.
A data scientist has created a benchmark for reinforcement learning and large language models by transforming Pokelike.xyz into a game environment where AI bots can be trained and tested. The project is open-source and invites contributions to improve bot performance.
This study uses connected vehicle telemetry data from Sydney to predict high-risk driving hotspots, benchmarking machine learning and time-series models for proactive road safety interventions.
This tweet announces the second edition of 'Introduction to Modern Statistics,' an open-source textbook available for free online or at a name-your-own-price PDF.