data-engineering

Tag

Cards List
#data-engineering

@freeCodeCamp: Production ETL pipelines need to keep running reliably, even when data is messy or APIs fail. In this handbook, Brookly…

X AI KOLs Timeline · 4h ago Cached

This freeCodeCamp handbook teaches readers how to build a production-grade ETL pipeline in Python using real flood data, covering incremental loading, type coercion, deduplication, and idempotency.

0 favorites 0 likes
#data-engineering

@gudanglifehack: SQL Project Series #19 Logistics & Shipment Analytics Analyze shipments, warehouses, delivery performance, customers, a…

X AI KOLs Timeline · 2d ago Cached

A SQL tutorial for building a logistics and shipment analytics database, covering schema design, sample data, and key business KPIs for supply chain optimization.

0 favorites 0 likes
#data-engineering

@LangChain: Our data agent now handles roughly 40x the request volume our 3 person data team could manage directly. Now, our data t…

X AI KOLs Timeline · 2d ago Cached

LangChain rebuilt its data stack around an AI agent that handles ~40x the request volume of its 3-person data team, enabling self-serve analysis and shifting the team's focus to models, context, and guardrails.

0 favorites 0 likes
#data-engineering

Guide to data tools landscape for developers

Hacker News Top · 2026-07-16 Cached

A comprehensive guide for software developers unfamiliar with data tools, covering the data lifecycle, data professions, and the tool landscape, with insights from the author's experience at Deepnote and Metabase.

0 favorites 0 likes
#data-engineering

@floydophone: Big news! Thrilled for the Dagster family to be joining Prefect.

X AI KOLs Following · 2026-07-13 Cached

Prefect announces acquisition of Dagster, combining two leading open-source data orchestration platforms.

0 favorites 0 likes
#data-engineering

Coding agents for data eng/analytics/ds

Reddit r/AI_Agents · 2026-07-03

Introduces a new coding agent designed to assist with data engineering, analytics, and data science tasks.

0 favorites 0 likes
#data-engineering

@EcZachly: Senior data engineers at Netflix can make $500k per year. The interview process is intense though: - SQL Make sure you …

X AI KOLs Timeline · 2026-06-24 Cached

A tweet thread from Zach outlines the skills and interview process for senior data engineers at Netflix, including SQL, data pipelines, and system design, while promoting DataExpert.io's mock interview service.

0 favorites 0 likes
#data-engineering

@sspaeti: Data engineering spent fifteen years building orchestrators — Airflow, Dagster, Prefect, Kestra. We are now doing the s…

X AI KOLs Timeline · 2026-06-16 Cached

Data engineering has spent years building orchestrators like Airflow and Dagster, and now the same pattern is emerging for AI agents with projects like Agor, Agent Teams, and Omnigent from major companies.

0 favorites 0 likes
#data-engineering

@GitHub_Daily: A carefully curated data engineering interview question bank on GitHub: data-engineering-interview-questions, featuring over 2,000 questions covering core areas such as databases and data warehouses, big data processing frameworks, cloud platform services, data formats, data visualization, and more.

X AI KOLs Timeline · 2026-06-10 Cached

There is a meticulously compiled data engineering interview question bank on GitHub called data-engineering-interview-questions, containing over 2,000 questions covering databases, big data frameworks, cloud platforms, data visualization, and other core areas.

0 favorites 0 likes
#data-engineering

@riba2534: https://x.com/riba2534/status/2062495991421616319

X AI KOLs Timeline · 2026-06-04 Cached

Anthropic shared best practices for implementing self-service data analysis with Claude, achieving 95% automation of business analysis queries with an overall accuracy of about 95%, and detailed the agent analysis tech stack, three main failure modes, and corresponding countermeasures.

0 favorites 0 likes
#data-engineering

databow: a Rust CLI to query any database with an ADBC driver

Hacker News Top · 2026-06-02 Cached

databow is a new open-source Rust CLI tool that provides a unified interface for querying any database with an ADBC driver, supporting over 30 databases including PostgreSQL, DuckDB, and Snowflake.

0 favorites 0 likes
#data-engineering

@svpino: Super-fast CLI tool for moving data around: $ ingestr ingest --source-uri --dest-uri Ingestr is open-source, requires a…

X AI KOLs Following · 2026-06-01

Ingestr is an open-source CLI tool for high-speed data movement between any source and destination, supporting numerous databases, data warehouses, and SaaS applications.

0 favorites 0 likes
#data-engineering

Exploring Autonomous Agentic Data Engineering for Model Specialization

arXiv cs.CL · 2026-06-01 Cached

This paper formalizes Autonomous Agentic Data Engineering, where LLMs act as autonomous data engineers to curate and optimize training data for specialized domains, showing a 57.29% improvement in student model performance using GPT-5.2.

0 favorites 0 likes
#data-engineering

Show HN: Streambed – Stream Postgres to Iceberg on S3, Supports Postgres Wire

Hacker News Top · 2026-05-31 Cached

Streambed is an open-source CDC engine that streams Postgres WAL changes to Iceberg tables on S3, with a built-in query server using DuckDB that speaks the Postgres wire protocol.

0 favorites 0 likes
#data-engineering

@github: Build data pipelines without the complexity. Tomorrow on Open Source Friday, dev advocate Elvis Kahoro explains how @dl…

X AI KOLs Following · 2026-05-28 Cached

GitHub's Open Source Friday event features Elvis Kahoro from dltHub discussing dlt, an open-source Python library for building data pipelines without complexity.

0 favorites 0 likes
#data-engineering

Exploring Autonomous Agentic Data Engineering for Model Specialization

Hugging Face Daily Papers · 2026-05-28 Cached

This paper introduces Autonomous Agentic Data Engineering, a task where LLMs autonomously execute end-to-end data curation pipelines for model specialization, showing significant performance gains (e.g., GPT-5.2 improves a student model by 57.29%).

0 favorites 0 likes
#data-engineering

@shabnam_774: https://x.com/shabnam_774/status/2058517919760355729

X AI KOLs Timeline · 2026-05-24 Cached

This article provides a comprehensive step-by-step breakdown of how modern Large Language Models like ChatGPT and Claude are built from scratch, covering data collection, tokenization, transformer architectures, training, alignment, and deployment.

0 favorites 0 likes
#data-engineering

@Teknium: Our database and data engineering expert @yoniebans made some major improvements to the way sessions are stored and acc…

X AI KOLs Following · 2026-05-21 Cached

Major improvements to session storage and access for Hermes Agent, saving 20-40% disk space and improving speed.

0 favorites 0 likes
#data-engineering

Quack: The DuckDB Client-Server Protocol

Hacker News Top · 2026-05-12 Cached

DuckDB introduces 'Quack', a new client-server protocol that enables DuckDB instances to communicate via HTTP, supporting concurrent writers and remote access while maintaining simplicity and performance.

0 favorites 0 likes
#data-engineering

DataTalksClub/data-engineering-zoomcamp

GitHub Trending (daily) · 2026-05-29 Cached

DataTalksClub is offering a free 9-week data engineering zoomcamp course covering containers, orchestration, data warehousing, analytics, batch and streaming processing.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback