data-science

Tag

Cards List
#data-science

@ClementDelangue: Super happy to release SmolDataEnvs: 5,000 verifiable RL environment tasks for hill-climbing small models in code and d…

X AI KOLs Timeline ↗ · yesterday Cached

Release of SmolDataEnvs, a collection of 5,000 verifiable RL environment tasks for training small models in code and data science, fully open source.

0 favorites 0 likes
#data-science

Writing Parquet files using Haskell

Hacker News Top ↗ · 4d ago Cached

This article details the implementation of a Parquet file writer in Haskell for the DataHaskell/Dataframe library, enabling efficient data serialization and interoperability with the data science ecosystem.

0 favorites 0 likes
#data-science

@mdancho84: RIP traditional data scientists The role is splitting: 1. ML modelers ($65k/yr) 2. AI/DS System Builders ($200k/yr) Wan…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

This article promotes a free live workshop on building AI with Python, offering a roadmap for data scientists to transition to AI/DS System Builder roles with practical demos.

0 favorites 0 likes
#data-science

@MaximeRivest: this was essentially me looking for ways to have a jev at home before they made it.

X AI KOLs Following ↗ · 2026-09-17 Cached

The author shares their past efforts to achieve efficient single-expert loading for high-quality token-level classification at scale, echoing Maxime Rivest's thoughts on the utility of domain-specific expert models.

0 favorites 0 likes
#data-science

Khipu (Quipu) Field Guide

Hacker News Top ↗ · 2026-09-16 Cached

The Khipu Field Guide is a web-based resource that provides detailed symbolic renderings and interactive analyses of over 600 Incan khipus, aiming to democratize khipu research through computational methods and data science.

0 favorites 0 likes
#data-science

@FrankRHutter: The data science revolution continues. TabPFN-3.5 is live, claiming SOTA beyond the vanilla IID small data setting and …

X AI KOLs Timeline ↗ · 2026-09-15 Cached

TabPFN-3.5 is released, claiming state-of-the-art performance for tabular data beyond IID small data settings, with features like handling grouped and temporal data, uncertainty calibration, and faster inference.

0 favorites 0 likes
#data-science

@MIT_CSAIL: A guide to statistics for beginners: https://tinyurl.com/y258cupb v/@OACerebro

X AI KOLs Timeline ↗ · 2026-09-12 Cached

MIT CSAIL shares a beginner's guide to statistics, providing foundational educational content for data science and AI.

0 favorites 0 likes
#data-science

if software engineering is mostly replaced by frontier models why is the volume of jobs on software eng still so high compared to data science related ones?

Reddit r/singularity ↗ · 2026-09-12

The article questions why software engineering jobs remain high compared to data science roles despite AI advancements, referencing US and EU labor statistics and growth projections for ICT professionals.

0 favorites 0 likes
#data-science

Calendar-SPCA: Interpretable Representation Learning for Multi-Periodic Electricity Consumption Profiles

arXiv cs.LG ↗ · 2026-09-10 Cached

This paper introduces Calendar-SPCA, a calendar-structured sparse principal component method for interpretable representation learning in multi-periodic electricity consumption profiles, demonstrating high explained variance with increased sparsity and coherence.

0 favorites 0 likes
#data-science

@VirtualElena: some takeaways from running a sota llm on a decade worth of lunches: - noted X anon @litcapital had the 4th most expens…

X AI KOLs Following ↗ · 2026-09-08 Cached

A user shares insights from analyzing a decade's worth of lunch data using a state-of-the-art LLM, highlighting interesting findings about expenses and common restaurants.

0 favorites 0 likes
#data-science

Benchmarking Language Models for Statistical Problem Formulation

arXiv cs.AI ↗ · 2026-09-03 Cached

The paper introduces StatFormBench, a benchmark for evaluating LLMs on statistical problem formulation, and finds that current models have significant limitations in classifying problems and identifying variables.

0 favorites 0 likes
#data-science

DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation

arXiv cs.AI ↗ · 2026-09-01 Cached

DS-Lighting is a unified harness toolkit that makes harness design explicit for data-science automation, improving reproducibility, comparability, and reliability in end-to-end workflows.

0 favorites 0 likes
#data-science

Data Science Approaches to Evaluating Honours Candidates

arXiv cs.CL ↗ · 2026-08-28 Cached

This paper presents the first application of data science to evaluate the UK Honours system using natural language processing, introducing a novel sentiment analysis algorithm called Minos to assess public opinion on honours recipients.

0 favorites 0 likes
#data-science

More than half of adults in U.S. say they lack basic statistical understanding

Hacker News Top ↗ · 2026-08-26 Cached

A Penn State-led study published in PLOS One finds that over half of U.S. adults report lacking basic statistical knowledge, yet most would rely on statistics more if they understood them better.

0 favorites 0 likes
#data-science

A new stamp on cyberfraud prevention

MIT Technology Review ↗ · 2026-08-25 Cached

Rupert Young, chief product officer at MaxMind, discusses his career in cybersecurity and data science, emphasizing the company's GeoIP tool used for fraud prevention across streaming, security, and merchant industries.

0 favorites 0 likes
#data-science

@XAMTO_AI: In the years I've been doing data analysis, the most frustrating part has never been writing code, but finding data. I'd search and search for hours, only to find that either it's not publicly available or the format is so messy it's a total turnoff. Yesterday, while browsing GitHub, I came across a list called Awesome Public Datasets, wow, 78k+ Stars,...

X AI KOLs Timeline ↗ · 2026-08-21 Cached

Introduces a list on GitHub named Awesome Public Datasets, which compiles high-quality public datasets from around the world, helping data analysts save time searching for data.

0 favorites 0 likes
#data-science

Optimizing things in the USSR (2016)

Hacker News Top ↗ · 2026-08-20 Cached

This article discusses the parallels between Soviet economic planning and modern data science practices, focusing on issues like resource allocation and simplifying assumptions based on historical books.

0 favorites 0 likes
#data-science

I transformed Pokelike.xyz into a LLM and RL benchmark!

Reddit r/LocalLLaMA ↗ · 2026-08-19

A data scientist has created a benchmark for reinforcement learning and large language models by transforming Pokelike.xyz into a game environment where AI bots can be trained and tested. The project is open-source and invites contributions to improve bot performance.

0 favorites 0 likes
#data-science

Proactive Road Safety Intervention in Australia: Predicting Risky Driving Hotspots from Connected Vehicle Data

arXiv cs.LG ↗ · 2026-08-19 Cached

This study uses connected vehicle telemetry data from Sydney to predict high-risk driving hotspots, benchmarking machine learning and time-series models for proactive road safety interventions.

0 favorites 0 likes
#data-science

@KirkDBorne: Introduction to Modern Statistics: https://openintro-ims.netlify.app Choose: Read it FREE online. Or name-your-own-pric…

X AI KOLs Timeline ↗ · 2026-08-18 Cached

This tweet announces the second edition of 'Introduction to Modern Statistics,' an open-source textbook available for free online or at a name-your-own-price PDF.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback