DataTalksClub/data-engineering-zoomcamp
Summary
DataTalksClub is offering a free 9-week data engineering zoomcamp course covering containers, orchestration, data warehousing, analytics, batch and streaming processing.
View Cached Full Text
Cached at: 05/29/26, 12:42 PM
DataTalksClub/data-engineering-zoomcamp
Source: https://github.com/DataTalksClub/data-engineering-zoomcamp
Data Engineering Zoomcamp: A Free 9-Week Course on Data Engineering Fundamentals
Master the fundamentals of data engineering by building an end-to-end data pipeline from scratch. Gain hands-on experience with industry-standard tools and best practices.
Join Slack β’ #course-data-engineering Channel β’ Telegram Announcements β’ Course Playlist β’ FAQ
How to Enroll
2026 Cohort
- Start Date: 12 January 2026
- Register Here: Sign up
Self-Paced Learning
All course materials are freely available for independent study. Follow these steps:
- Watch the course videos.
- Join the Slack community.
- Refer to the FAQ document for guidance.
Syllabus Overview
The course consists of structured modules, hands-on workshops, and a final project to reinforce your learning.
Prerequisites
To get the most out of this course, you should have:
- Basic coding experience
- Familiarity with SQL
- Experience with Python (helpful but not required)
No prior data engineering experience is necessary.
Modules
Module 1: Containerization and Infrastructure as Code
- Introduction to GCP
- Docker and Docker Compose
- Running PostgreSQL with Docker
- Infrastructure setup with Terraform
- Homework
Module 2: Workflow Orchestration
- Data Lakes and Workflow Orchestration
- Workflow orchestration with Kestra
- Homework
Workshop 1: Data Ingestion
- API reading and pipeline scalability
- Data normalization and incremental loading
- Homework
Module 3: Data Warehousing
- Introduction to BigQuery
- Partitioning, clustering, and best practices
- Machine learning in BigQuery
Module 4: Analytics Engineering
- Analytics Engineering and Data Modeling
- dbt (data build tool) with DuckDB & BigQuery
- Testing, documentation, and deployment
Module 5: Data Platforms
- Building end-to-end data pipelines with Bruin
- Data ingestion, transformation, and quality
- Deployment to cloud (BigQuery)
Module 6: Batch Processing
- Introduction to Apache Spark
- DataFrames and SQL
- Internals of GroupBy and Joins
Module 7: Streaming
- Introduction to Kafka
- Kafka Streams and KSQL
- Schema management with Avro
Final Project
- Apply all concepts learned in a real-world scenario
- Peer review and feedback process
Testimonials
Thank you for what you do! The Data Engineering Zoomcamp gave me skills that helped me land my first tech job.
β Tim Claytor (Source)
Three months might seem like a long time, but the growth and learning during this period are truly remarkable. It was a great experience with a lot of learning, connecting with like-minded people from all around the world, and having fun. I must admit, this was really hard. But the feeling of accomplishment and learning made it all worthwhile. And I would do it again!
β Nevenka Lukic (Source)
One of the significant things I inferred from the Zoomcamp is to prioritize fundamentals and principles over ever-evolving tools and tech stacks. Hugely grateful to Alexey Grigorev for putting together this incredible course and offering it for free.
β Siddhartha Gogoi (Source)
Such a fun deep dive into data engineering, cloud automation, and orchestration. I learned so much along the way. Big shoutout to Alexey Grigorev and the DataTalksClub team for the opportunity and guidance throughout the 3 months of the free course.
β Assitan NIARE (Source)
If youβre serious about breaking into data engineering, start here. The repoβs structure, community, and hands-on focus make it unparalleled.
β Wady Osama (Source)
Community & Support
Getting Help on Slack
Join the #course-data-engineering channel on DataTalks.Club Slack for discussions, troubleshooting, and networking.
To keep discussions organized:
- Follow our guidelines when posting questions.
- Review the community guidelines.
Meet the Instructors
Past instructors:
Sponsors & Supporters
A special thanks to our course sponsors for making this initiative possible!
Interested in supporting our community? Reach out to [email protected].
About DataTalks.Club
DataTalks.Club is a global online community of data enthusiasts. It's a place to discuss data, learn, share knowledge, ask and answer questions, and support each other.
Website β’ Join Slack Community β’ Newsletter β’ Upcoming Events β’ YouTube β’ GitHub β’ LinkedIn β’ Twitter
All the activity at DataTalks.Club mainly happens on Slack. We post updates there and discuss different aspects of data, career questions, and more.
At DataTalksClub, we organize online events, community activities, and free courses. You can learn more about what we do at DataTalksClub Community Navigation.
Similar Articles
@anyscalecompute: In this session, you'll learn: - Build and scale data pipelines with Ray - What is video data curation - Stream large dβ¦
Anyscale is hosting a hands-on virtual lab session teaching developers how to build and scale data pipelines with Ray, covering video data curation, distributed GPU inference, and CPU/GPU streaming pipelines.
@DataScienceDojo: Five days. 40 hours. One πππ ππ©π©π₯π’ππππ’π¨π§ you'll build yourself. That's the shape of our next LLM Bootcampβ¦
DataScienceDojo announces their next LLM Bootcamp (August 24-28, 2026), covering transformers, attention mechanisms, vector databases, fine-tuning, RAG, and agent building. Over 12,000 alumni have completed the program.
@DataScienceDojo: Your LLM agent can call tools, write files, and hit APIs on its own β which means it needs somewhere safe to do that. Jβ¦
DataScienceDojo hosts a Docker-sponsored webinar on running LLM agents safely with Docker sandboxes, featuring a hands-on session by Docker Developer Success Advocate Dan Ndombe.
Enabling a data-driven workforce
OpenAI hosted a webinar on August 8, 2024 demonstrating how employees can use ChatGPT Enterprise for data analysis and business insights.
DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
DataFlow is an LLM-driven framework for automated data preparation and workflow engineering, featuring nearly 200 reusable operators and six domain-general pipelines that improve LLM performance across tasks like math, code, and Text-to-SQL.
