big-data

Tag

Cards List
#big-data

@gp_pulipaka: Mathematical Methods in Data Science! #BigData #Analytics #DataScience #AI #MachineLearning #IoT #IIoT #PyTorch #Python…

X AI KOLs Timeline · 2026-08-18 Cached

This tweet promotes the book 'Mathematical Methods in Data Science' by Ren and Wang, which covers mathematical methods like differential equations for data science applications, suitable for advanced students.

0 favorites 0 likes
#big-data

Databricks wanted to raise $1B, investors wanted $15B. It settled on $5B at a $190B valuation.

TechCrunch AI · 2026-08-13 Cached

Databricks raised $5 billion at a $190 billion valuation after investor demand hit $15 billion, despite originally seeking only $1 billion. CEO Ali Ghodsi cited $7 billion annualized revenue growing 80% and heavy AI investment costs as reasons for the raise.

0 favorites 0 likes
#big-data

Algorithms on billion-scale graph using 10GB RAM: I love DataFusion

Hacker News Top · 2026-07-31 Cached

A developer shows that Apache DataFusion can perform billion-scale graph analytics like PageRank and weakly connected components on a laptop with 5–10GB RAM by offloading data to disk, challenging the need for Spark/GraphFrames.

0 favorites 0 likes
#big-data

@gp_pulipaka: Perplexity’s Bumblebee Pipeline! @perplexity_ai #BigData #Analytics #DataScience #AI #MachineLearning #NLProc #LLM #IoT…

X AI KOLs Timeline · 2026-07-27 Cached

Perplexity AI announces the Bumblebee Pipeline, a tool for big data and AI workflows, shared by prominent AI influencer gp_pulipaka.

0 favorites 0 likes
#big-data

@gp_pulipaka: Databricks Delta Sharing! #BigData #Analytics #DataScience #AI #MachineLearning #NLProc #LLM #IoT #IIoT #PyTorch #Pytho…

X AI KOLs Timeline · 2026-07-26 Cached

Databricks Delta Sharing is an open protocol that enables secure, live data sharing across cloud providers without replication, reducing egress costs and simplifying multi-cloud data access.

0 favorites 0 likes
#big-data

Data-Native Global Optimization for Big Data K-means Clustering

arXiv cs.LG · 2026-07-20 Cached

Proposes Big-means++, a simple algorithm that achieves global optimization quality for big data K-means clustering by systematically curating inputs and using sample-induced surrogate landscapes.

0 favorites 0 likes
#big-data

@Thom_Wolf: people are sleeping on the mega-release happening every week in AI x Science on Hugging Face this one is 80TB of astrop…

X AI KOLs Following · 2026-06-30 Cached

The Multimodal Universe (MMU), an 80TB+ collection of astronomical survey data, has been converted to the HATS parquet format, enabling crossmatching on a laptop via LSDB and Hugging Face ecosystems without bulk downloads.

0 favorites 0 likes
#big-data

@GitHub_Daily: A carefully curated data engineering interview question bank on GitHub: data-engineering-interview-questions, featuring over 2,000 questions covering core areas such as databases and data warehouses, big data processing frameworks, cloud platform services, data formats, data visualization, and more.

X AI KOLs Timeline · 2026-06-10 Cached

There is a meticulously compiled data engineering interview question bank on GitHub called data-engineering-interview-questions, containing over 2,000 questions covering databases, big data frameworks, cloud platforms, data visualization, and other core areas.

0 favorites 0 likes
#big-data

Rubin Tracks Skyscraper-Size Asteroids and Failed Supernovas

Hacker News Top · 2026-06-01 Cached

The Vera C. Rubin Observatory has begun initial observations, already discovering rapidly spinning asteroids, supernovas, and a potential interstellar visitor, heralding a new era of big-data astronomy.

0 favorites 0 likes
#big-data

Automated Big Data Quality Assessment using Knowledge Graph Embeddings

arXiv cs.LG · 2026-05-20

This paper introduces a knowledge-based approach using knowledge graph embeddings to automatically assess big data quality by predicting missing edges between context representations and quality rules, outperforming traditional matching methods.

0 favorites 0 likes
← Back to home

Submit Feedback