Skip to content
#

data-lakehouse

Here are 279 public repositories matching this topic...

End-to-end Data Lakehouse project built on Databricks, following the Medallion Architecture (Bronze, Silver, Gold). Covers real-world data engineering and analytics workflows using Spark, PySpark, SQL, Delta Lake, and Unity Catalog. Designed for learning, portfolio building, and job interviews.

  • Updated Jan 19, 2026
  • Jupyter Notebook

ETL / ELT / Reverse ETL Framework powered by DuckDB, designed to seamlessly integrate and process data from diverse sources. It leverages Markdown as a configuration medium, where YAML blocks define metadata for each data source, and embedded SQL blocks specify the extraction, transformation, and loading logic.

  • Updated Oct 4, 2026
  • Go

This repo provides a step-by-step approach to building a modern data warehouse using PostgreSQL. It covers the ETL (Extract, Transform, Load) process, data modeling, exploratory data analysis (EDA), and advanced data analysis techniques.

  • Updated Mar 7, 2025
  • PLpgSQL

Production-ready Apache Superset with DuckLake integration. Stateless analytics architecture using DuckDB for compute, PostgreSQL for metadata, and S3/GCS/MinIO for data lake storage. Includes Docker Compose, Kubernetes Helm charts, BigQuery Integration, and CI/CD workflows. Supports MotherDuck cloud integration.

  • Updated Feb 6, 2026
  • Python

Add this topic to your repo

To associate your repository with the data-lakehouse topic, visit your repo's landing page and select "manage topics."

Learn more