data-engineering

Tag

Cards List
#data-engineering

@JohnKutay: A year ago we had to hire Databricks specialists to build spark jobs. Now we have GTM Engineers building data pipelines…

X AI KOLs Following · 2d ago Cached

A tweet discusses how Databricks has simplified Spark, enabling non-specialists like GTM Engineers to build data pipelines without deep technical knowledge.

0 favorites 0 likes
#data-engineering

Every Millisecond Counts

Lobsters Hottest · 5d ago Cached

This blog post details incremental optimizations to reduce ClickHouse query latency from over 85 seconds to sub-second, including changes to partition keys, join elimination, and aggressive merges.

0 favorites 0 likes
#data-engineering

I've operated petabyte-scale ClickHouse clusters for 5 years

Hacker News Top · 2026-09-07 Cached

The article shares lessons learned from operating petabyte-scale ClickHouse clusters for five years, discussing best practices and promoting Tinybird's managed data services.

0 favorites 0 likes
#data-engineering

Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning

arXiv cs.CL · 2026-09-02 Cached

This paper proposes a neurosymbolic layer for LLMs that improves logical reasoning accuracy by 8.5% without finetuning and reduces token usage by over 50% on long context tasks.

0 favorites 0 likes
#data-engineering

@no_stp_on_snek: Really considering getting that second spark. Convince me! (note it won't be hard)

X AI KOLs Following · 2026-08-23 Cached

A Twitter user contemplates obtaining a second Spark instance, likely referring to Apache Spark, and asks for persuasion on the decision.

0 favorites 0 likes
#data-engineering

How We Pushed CDC into Postgres

Hacker News Top · 2026-08-10 Cached

Snowflake's engineering blog details the new Postgres data mirroring feature that pushes changes via a snowflake_cdc extension into Iceberg tables, aiming for resilient, low-lag transactional replication.

0 favorites 0 likes
#data-engineering

DuckDB – Data power tools for your laptop, now in Clojure (2023)

Hacker News Top · 2026-08-04 Cached

TechAscent illustrates how DuckDB's high-performance vectorized SQL engine can now be leveraged from Clojure via tech.ml.dataset (TMD), enabling large out-of-memory joins and 50GB CSV ingestion that compresses to 18GB in about two minutes.

0 favorites 0 likes
#data-engineering

@PythonHub: Guide to data tools landscape for developers Found yourself on a data project and have no idea what they all are talkin…

X AI KOLs Timeline · 2026-07-31 Cached

A comprehensive guide for software developers entering the data field, explaining the data tools landscape, key concepts, and workflows from ingestion to visualization.

0 favorites 0 likes
#data-engineering

@freeCodeCamp: Production ETL pipelines need to keep running reliably, even when data is messy or APIs fail. In this handbook, Brookly…

X AI KOLs Timeline · 2026-07-31 Cached

This freeCodeCamp handbook teaches readers how to build a production-grade ETL pipeline in Python using real flood data, covering incremental loading, type coercion, deduplication, and idempotency.

0 favorites 0 likes
#data-engineering

@gudanglifehack: SQL Project Series #19 Logistics & Shipment Analytics Analyze shipments, warehouses, delivery performance, customers, a…

X AI KOLs Timeline · 2026-07-29 Cached

A SQL tutorial for building a logistics and shipment analytics database, covering schema design, sample data, and key business KPIs for supply chain optimization.

0 favorites 0 likes
#data-engineering

@LangChain: Our data agent now handles roughly 40x the request volume our 3 person data team could manage directly. Now, our data t…

X AI KOLs Timeline · 2026-07-28 Cached

LangChain rebuilt its data stack around an AI agent that handles ~40x the request volume of its 3-person data team, enabling self-serve analysis and shifting the team's focus to models, context, and guardrails.

0 favorites 0 likes
#data-engineering

Guide to data tools landscape for developers

Hacker News Top · 2026-07-16 Cached

A comprehensive guide for software developers unfamiliar with data tools, covering the data lifecycle, data professions, and the tool landscape, with insights from the author's experience at Deepnote and Metabase.

0 favorites 0 likes
#data-engineering

@floydophone: Big news! Thrilled for the Dagster family to be joining Prefect.

X AI KOLs Following · 2026-07-13 Cached

Prefect announces acquisition of Dagster, combining two leading open-source data orchestration platforms.

0 favorites 0 likes
#data-engineering

Coding agents for data eng/analytics/ds

Reddit r/AI_Agents · 2026-07-03

Introduces a new coding agent designed to assist with data engineering, analytics, and data science tasks.

0 favorites 0 likes
#data-engineering

@EcZachly: Senior data engineers at Netflix can make $500k per year. The interview process is intense though: - SQL Make sure you …

X AI KOLs Timeline · 2026-06-24 Cached

A tweet thread from Zach outlines the skills and interview process for senior data engineers at Netflix, including SQL, data pipelines, and system design, while promoting DataExpert.io's mock interview service.

0 favorites 0 likes
#data-engineering

@sspaeti: Data engineering spent fifteen years building orchestrators — Airflow, Dagster, Prefect, Kestra. We are now doing the s…

X AI KOLs Timeline · 2026-06-16 Cached

Data engineering has spent years building orchestrators like Airflow and Dagster, and now the same pattern is emerging for AI agents with projects like Agor, Agent Teams, and Omnigent from major companies.

0 favorites 0 likes
#data-engineering

@GitHub_Daily: A carefully curated data engineering interview question bank on GitHub: data-engineering-interview-questions, featuring over 2,000 questions covering core areas such as databases and data warehouses, big data processing frameworks, cloud platform services, data formats, data visualization, and more.

X AI KOLs Timeline · 2026-06-10 Cached

There is a meticulously compiled data engineering interview question bank on GitHub called data-engineering-interview-questions, containing over 2,000 questions covering databases, big data frameworks, cloud platforms, data visualization, and other core areas.

0 favorites 0 likes
#data-engineering

@riba2534: https://x.com/riba2534/status/2062495991421616319

X AI KOLs Timeline · 2026-06-04 Cached

Anthropic shared best practices for implementing self-service data analysis with Claude, achieving 95% automation of business analysis queries with an overall accuracy of about 95%, and detailed the agent analysis tech stack, three main failure modes, and corresponding countermeasures.

0 favorites 0 likes
#data-engineering

databow: a Rust CLI to query any database with an ADBC driver

Hacker News Top · 2026-06-02 Cached

databow is a new open-source Rust CLI tool that provides a unified interface for querying any database with an ADBC driver, supporting over 30 databases including PostgreSQL, DuckDB, and Snowflake.

0 favorites 0 likes
#data-engineering

@svpino: Super-fast CLI tool for moving data around: $ ingestr ingest --source-uri --dest-uri Ingestr is open-source, requires a…

X AI KOLs Following · 2026-06-01

Ingestr is an open-source CLI tool for high-speed data movement between any source and destination, supporting numerous databases, data warehouses, and SaaS applications.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback