@GitHub_Daily: GitHub 上一份精心收集的数据工程师面试题库:data-engineering-interview-questions ,收录了超过 2000 道题。 覆盖数据库与数据仓库、大数据处理框架、云平台服务、数据格式、数据可视化等核心方向,…

X AI KOLs Timeline 工具

摘要

GitHub 上有一个精心收集的数据工程师面试题库 data-engineering-interview-questions,收录了超过 2000 道题,覆盖数据库、大数据框架、云平台、数据可视化等核心方向。

GitHub 上一份精心收集的数据工程师面试题库:data-engineering-interview-questions ,收录了超过 2000 道题。 覆盖数据库与数据仓库、大数据处理框架、云平台服务、数据格式、数据可视化等核心方向,可以按主题逐个击破,也能用完整题单做全面模拟。 GitHub:http://github.com/OBenner/data-engineering-interview-questions… 数据库部分涵盖 Cassandra、MongoDB、HBase、Hive 以及 Redshift、BigQuery 等主流选型。 理论部分也没落下,数据建模、数据质量、系统设计、SQL 和 Python 都有专门的题目集。 每个主题还附带了官方文档和 Awesome 资源列表的链接,方便深入学习。 如果你正在准备数据工程相关的面试,这份题库值得收藏,系统刷一遍心里会踏实很多。
查看原文
查看缓存全文

缓存时间: 2026/06/10 09:48

GitHub 上一份精心收集的数据工程师面试题库:data-engineering-interview-questions ,收录了超过 2000 道题。

覆盖数据库与数据仓库、大数据处理框架、云平台服务、数据格式、数据可视化等核心方向,可以按主题逐个击破,也能用完整题单做全面模拟。

GitHub:http://github.com/OBenner/data-engineering-interview-questions…

数据库部分涵盖 Cassandra、MongoDB、HBase、Hive 以及 Redshift、BigQuery 等主流选型。

理论部分也没落下,数据建模、数据质量、系统设计、SQL 和 Python 都有专门的题目集。

每个主题还附带了官方文档和 Awesome 资源列表的链接,方便深入学习。

如果你正在准备数据工程相关的面试,这份题库值得收藏,系统刷一遍心里会踏实很多。


OBenner/data-engineering-interview-questions

Source: https://github.com/OBenner/data-engineering-interview-questions

More than 2000+ questions for preparing a Data Engineer interview.

Full list of questions

Pick a topic below or use the full list to practice end-to-end.

Interview questions for Data Engineer

Databases and Data Warehouses
GitHub Repo Official page Questions Description Useful links
Cassandra Cassandra Apache Cassandra Cassandra is a distributed, wide-column store, NoSQL database management system. Awesome Cassandra
Greenplum Greenplum Greenplum Greenplum is a big data technology based on MPP architecture and the Postgres open source database technology. Awesome Greenplum
MongoDB MongoDB MongoDB MongoDB is a document-oriented database. Awesome MongoDB
Hbase Hbase Apache Hbase HBase is an open-source non-relational distributed database. Awesome HBase
Hive Hive Apache Hive Apache Hive is a data warehouse software project built on top of Apache Hadoop for providing data query and analysis. Awesome Hive
Amazon DynamoDB Amazon DynamoDB Amazon DynamoDB is a fully managed proprietary NoSQL database service. Awesome DynamoDB Awesome AWS
Amazon Redshift Amazon Redshift Amazon Redshift is a data warehouse product. Amazon Redshift Utilities Awesome AWS
BigQuery BigQuery GCP BigQuery is a fully-managed, serverless data warehouse. Awesome BigQuery
Bigtable Bigtable GCP Bigtable is a fully managed wide-column and key-value NoSQL database service. Awesome Bigtable
Data Formats
Avro Avro Apache Avro Avro is a row-oriented remote procedure call and data serialization framework. Awesome Avro
Parquet Parquet Apache Parquet Apache Parquet is a column-oriented data file format designed for efficient data storage and retrieval. Parquet format · Docs
Delta Delta Delta Delta Lake is a storage framework that enables building a Lakehouse architecture with compute engines Delta examples
Iceberg Iceberg Apache Iceberg Apache Iceberg is an open table format for huge analytic datasets. Iceberg docs
Hudi Hudi Apache Hudi Apache Hudi brings upserts, deletes, and incremental processing to data lakes. Hudi docs
Big Data Frameworks
Airflow Airflow Apache Airflow Apache Airflow is a workflow management platform for data engineering pipelines. Awesome Airflow
Flume Flume Apache Flume Apache Flume is a distributed, reliable, and available software for efficiently collecting, aggregating, and moving large amounts of log data. Flume User Guide
Hadoop Hadoop Apache Hadoop Apache Hadoop is a collection of software utilities that facilitates using a network of many computers to solve problems involving massive amounts of data and computation. Awesome Hadoop
Impala Impala Apache Impala Apache Impala is a parallel processing SQL query engine for data stored in a computer cluster running Apache Hadoop. Impala docs
Kafka Kafka Apache Kafka Apache Kafka is a distributed event store and stream-processing platform. Awesome Kafka
NiFi NiFi Apache NiFi Apache NiFi is a software project designed to automate the flow of data between software systems. Awesome NiFi
Spark Spark Apache Spark Apache Spark is unified analytics engine for large-scale data processing. Awesome Spark
Flink Flink Apache Flink Apache Flink is unified stream-processing and batch-processing framework. Awesome Flink
Kubernetes Kubernetes Kubernetes Kubernetes is a system for managing containerized applications across multiple hosts. Awesome Kubernetes
Cloud providers
AWS AWS Amazon Web Services Amazon web service is an online platform that provides scalable and cost-effective cloud computing solutions. Awesome AWS
Azure Azure Microsoft Azure Microsoft Azure is Microsoft's public cloud computing platform. Awesome Azure
GCP GCP Google Cloud Platform Google Cloud Platform is a suite of cloud computing services. Awesome GCP
Modern Data Stack
dbt dbt dbt dbt is a transformation framework for building tested and documented SQL models. dbt tests
Theory
DWHA DWH Architectures A data warehouse architecture is a method of defining the overall architecture of data communication processing and presentation that exist for end-clients computing within the enterprise. Awesome databases
CDC Change Data Capture (CDC) CDC captures inserts/updates/deletes from source systems for low-latency ingestion. Debezium docs
Data Modeling Data Modeling Dimensional modeling concepts used to build reliable analytics datasets. Kimball Group
Data Quality Data Quality Tests, monitoring, and practices to ensure datasets are trusted and correct. Great Expectations docs
Data Observability Data Observability Monitoring and incident response practices for pipeline and dataset health. OpenLineage
Data Governance Data Governance Ownership, policies, privacy, and access controls for data platforms. DataHub
Cost Optimization Cost Optimization Practical techniques to reduce compute and storage costs while meeting SLAs. Spark tuning
Python Python for Data Engineering Python fundamentals for reliable, scalable data pipelines and tooling. PyArrow docs
System Design Data System Design System design interview questions for batch/streaming data platforms. Data mesh overview
Airflow Data Structures A data structure is a specialized format for organizing, processing, retrieving and storing data. Awesome Algorithms
SQL SQL SQL is a domain-specific language used in programming and designed for managing data held in a relational database management system (RDBMS). Awesome SQL
Data visualization tools/BI
Tableau Tableau Tableau is a powerful data visualization tool used in the Business Intelligence. Tableau Desktop docs
Looker Looker Looker is an enterprise platform for BI, data applications, and embedded analytics that helps you explore and share insights in real time. Looker docs
Kafka Apache Superset Apache Superset Superset is a modern data exploration and data visualization platform Superset docs

Contribution

Please contribute to this repository to help it make better. Any change like new question, code improvement, doc improvement etc is very welcome.

See CONTRIBUTING.md for quick checks and guidelines.

相似文章

@GitHub_Daily: 准备运维和 DevOps 面试,网上搜到的大多是那种拼凑的题库,跟真实面试差别挺大。 DevOps-Interview-Guide 收录了 151 份真实面试记录,来自 85 家公司,问的问题都是候选人原样记下来的,没有二手转述。 覆盖 …

X AI KOLs Timeline

推荐一个收录了151份真实DevOps/SRE面试记录的GitHub仓库,涵盖85家公司,问题由候选人原样记录,覆盖Kubernetes、Docker、Terraform、AWS、CI/CD等方向,方便求职者按公司或主题准备面试。

@vintcessun: 一早翻到一个有意思的项目,改变了我对面试准备的认知。一直以为大厂面试刷题就够了,但本质上它考察的是完整的计算机科学知识体系。这个项目把离散的知识点串成了一个系统计划,从 Big-O、数据结构、算法到系统设计、面试技巧全覆盖,甚至包含如何写…

X AI KOLs Timeline

A popular GitHub project providing a comprehensive multi-month study plan for software engineering interviews, covering CS fundamentals, algorithms, system design, and resume tips.

litu54/DevOps-Interview-Guide

GitHub Trending (daily)

一个 GitHub 仓库,收集了 2025-2026 年真实的 DevOps、SRE 和云工程面试题,按公司整理,涵盖 85 家公司的 151 篇文章。