Tag
This tweet promotes the book 'Mathematical Methods in Data Science' by Ren and Wang, which covers mathematical methods like differential equations for data science applications, suitable for advanced students.
Databricks raised $5 billion at a $190 billion valuation after investor demand hit $15 billion, despite originally seeking only $1 billion. CEO Ali Ghodsi cited $7 billion annualized revenue growing 80% and heavy AI investment costs as reasons for the raise.
A developer shows that Apache DataFusion can perform billion-scale graph analytics like PageRank and weakly connected components on a laptop with 5–10GB RAM by offloading data to disk, challenging the need for Spark/GraphFrames.
Perplexity AI announces the Bumblebee Pipeline, a tool for big data and AI workflows, shared by prominent AI influencer gp_pulipaka.
Databricks Delta Sharing is an open protocol that enables secure, live data sharing across cloud providers without replication, reducing egress costs and simplifying multi-cloud data access.
Proposes Big-means++, a simple algorithm that achieves global optimization quality for big data K-means clustering by systematically curating inputs and using sample-induced surrogate landscapes.
The Multimodal Universe (MMU), an 80TB+ collection of astronomical survey data, has been converted to the HATS parquet format, enabling crossmatching on a laptop via LSDB and Hugging Face ecosystems without bulk downloads.
There is a meticulously compiled data engineering interview question bank on GitHub called data-engineering-interview-questions, containing over 2,000 questions covering databases, big data frameworks, cloud platforms, data visualization, and other core areas.
The Vera C. Rubin Observatory has begun initial observations, already discovering rapidly spinning asteroids, supernovas, and a potential interstellar visitor, heralding a new era of big-data astronomy.
This paper introduces a knowledge-based approach using knowledge graph embeddings to automatically assess big data quality by predicting missing edges between context representations and quality rules, outperforming traditional matching methods.