Skip to content

Repository files navigation

DATA-KING

Data Engineering. ETL. Big Data. Cloud Pipelines.

Welcome to DATA-KING — the Data Engineering version of “Backend-King”.
This repository is not about APIs, microservices, or backend frameworks.
This is about data pipelines, transformations, analytics platforms, and cloud-scale ETL systems.

Think of this as a hands-on lab for mastering modern Data Engineering using Python, Spark, Snowflake, Databricks, Alteryx, Talend, and all major cloud providers.


🔥 What is DATA-KING?

DATA-KING is a collection of real-world examples for:

  • Building batch & streaming pipelines
  • Performing large-scale transformations
  • Working with data warehouses and lakehouses
  • Using cloud-native ETL tools
  • Designing production-grade data architectures

It’s designed for:

  • Data Engineers
  • Analytics Engineers
  • Cloud Data Engineers
  • ML Engineers who want strong data foundations

🧠 Core Stack

Area Tools
Programming Python, PySpark
Big Data Apache Spark
Data Warehousing Snowflake
Lakehouse Databricks + Delta Lake
Cloud ETL AWS Glue, Azure Data Factory, GCP Dataflow
Low-Code ETL Alteryx, Talend
Orchestration Airflow / Prefect
Storage S3, GCS, Azure Blob
Compute EMR, Dataproc, Synapse

🏗 Architecture Overview

Typical pipeline flow:

About

Data-King is a curated collection of sample data pipelines, transformations, and cloud architectures using top-tier tools. Think of it as your data engineering lab where you learn by doing ETL, ELT, streaming, batch jobs, orchestration, and analytics across AWS, GCP, Azure, Databricks, Snowflake, and open-source engines.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages