Seasoned Data Engineer with extensive expertise in building end-to-end data pipelines and real-time streaming solutions. Specialized in cloud platforms (GCP, AWS, Azure) and modern data technologies. Passionate about architecting scalable, reliable, and cost-efficient data solutions that drive business value.
- Google Cloud Platform (GCP) - Databricks, BigQuery, Dataproc, Pub/Sub, Cloud Run, Cloud Composer, and GCS
- Amazon Web Services (AWS) - S3, DynamoDB, Redshift, Kinesis, Lambda, Airflow, EC2, and Glue
- Microsoft Azure - ADLS, Synapse, ADF, CosmosDB, MS-Fabric, Event Hub, Stream Analytics
- Orchestration: Apache Airflow, Cloud Composer, Azure Data Factory, and SSIS
- Streaming: Apache Kafka, Confluent Cloud, Azure Event Hub, Apache Flink, Redpanda
- Data Warehousing: Snowflake, BigQuery, Redshift, Azure Synapse, Databricks
- Data Processing: PySpark, Apache Spark, dbt, Microsoft Fabric, Delta Lake
- Transformation: dbt, Apache Beam, Dataflow, Azure Data Factory
- Storage: GCS, S3, ADLS, Apache Iceberg, Delta Lake
- Languages: Python, SQL (T-SQL), and Pyspark
- Tools: dbt, SSIS, SSRS, PowerBI, ThoughtSpot, Looker, Streamlit, n8n
- Infrastructure: Docker, GitHub Actions, Azure DevOps CI/CD, ARM Templates
- Version Control: Git, GitHub, TFS
- Kafka Stock Market Pipeline - Real-time stock market data processing
- Azure Event Hub BookMyShow Pipeline - Real-time event booking and payment streams with Stream Analytics
- GCP Uber Car Idle Alerts - Pub/Sub to BigQuery real-time alerts using Dataflow
- Databricks UPI Transactions CDC - Real-time transaction change data capture with PySpark Streaming
- Confluent Kafka MongoDB Streaming - Kafka to MongoDB real-time pipelines
- AWS Kinesis Spark Streaming - Real-time food delivery pipeline with Spark Streaming
- Flink Real-Time Processing - Redpanda, Flink, and Postgres integration for real-time analytics
- Flight Booking Pipeline - Airflow β GCS β BigQuery β Looker
- Credit Card Processing - End-to-end credit card data pipeline with Looker dashboards
- YouTube Trends Pipeline - GCS Iceberg tables to BigQuery analytics
- Weather Map API Integration - API data extraction and PySpark transformation
- Bigtable CRUD Operations - NoSQL database operations with Python
- News API GCS to Snowflake - Multi-warehouse pipeline with Airflow
- Airflow Dataproc PySpark - Ephemeral Dataproc cluster orchestration
- Pyspark Backfill Pipeline - Backfill DAGs with PySpark jobs
- Real-time Gaming Leaderboard - Low-latency analytics with Dataflow
- Snowflake Car Rental - GCP Airflow to Snowflake integration
- SCD Type 2 CI/CD - GitHub Actions automation for data processing
- Airline Data Ingestion - S3 to Redshift pipeline with data quality checks
- Movie Quality Pipeline - Data quality framework and monitoring
- S3 to Redshift Pipeline - Airflow orchestration with Redshift
- DBT Snowflake Preset - dbt transformations with Preset dashboarding
- Airbnb CosmosDB Pipeline - Near real-time pipeline with CosmosDB and n8n workflows
- Airport ADF with CI/CD - Data Factory with DevOps automation and ARM templates
- Fintech SQL Pipeline - Multi-warehouse architecture with Synapse
- Analysis Services Model - OLAP cube development from SQL Server
- Olympic Data Engineering - End-to-end analytics project
- Snowflake Movies Streamlit App - Dynamic tables with interactive Streamlit dashboards
- Databricks Travel Booking SCD2 - Slowly Changing Dimensions Type 2 implementation
- Databricks dbt Project - dbt frameworks on Databricks
- dbt Snowflake - dbt with Snowflake integration
- Microsoft Fabric End-to-End - Bronze-Silver-Gold architecture on Fabric
- Microsoft Fabric Uber Analytics - Comprehensive analytics platform
- Microsoft Fabric API PowerBI - API ingestion with PowerBI reporting
- GCP PySpark Streaming Iceberg - Open table format implementation
- Databricks Ecommerce Event-Driven - Event-driven architecture on Databricks
- Databricks Healthcare DLT - Delta Live Tables for healthcare data
- Airflow GCP Complete Stack - Complete GCP visualization stack
- Yahoo Finance API - Financial data extraction with Airflow
- Data Engineer Handbook - Comprehensive data engineering resources and best practices
- Awesome Data Engineering - Curated list of tools, frameworks, and resources
- System Design Academy - System design principles for scalability
- Data Structures & Algorithms (Python) - DSA templates, solutions, and practical projects
- Python Web Scraping Projects - Web data extraction techniques and examples
- PySpark Project Files - Spark learning resources and implementations
- Linux Project Files - Linux system administration knowledge base
- βοΈ Azure Cloud Professional Certifications
- βοΈ GCP Data Engineering Certifications
- βοΈ AWS Solutions Architect Certifications
- π Data Engineering Specialized Certifications
- π Technology Stack Specific Training
- π Advanced Tools & Frameworks Certifications
- π Udacity Data Engineering Nanodegree
- π Multiple Udemy courses on data platforms
- π Official cloud provider training programs
- ποΈ Awards and Recognitions in IT career
- ποΈ Extensive hands-on project experience
- ποΈ Community contributions and knowledge sharing
| Category | Count | Focus |
|---|---|---|
| GCP Projects | 15+ | Cloud data pipelines, BigQuery, Dataproc, Airflow |
| AWS Projects | 10+ | Redshift, S3, Kinesis, Lambda pipelines |
| Azure Projects | 12+ | Synapse, Data Factory, ADLS, CosmosDB |
| Databricks Projects | 5+ | dbt, Delta, Streaming, MLOps |
| Streaming Projects | 8+ | Kafka, Kinesis, Event Hub, Flink, Pub/Sub |
| Microsoft Fabric | 3+ | End-to-end lakehouse solutions |
| Learning Resources | 10+ | DSA, Python, Linux, System Design |
| Total Repositories | 75+ | Production-grade implementations |
β
Multi-Cloud Architecture - Design and implement solutions across GCP, AWS, Azure
β
Real-Time Streaming - Kafka, Kinesis, Event Hub, Flink, Pub/Sub expertise
β
Data Warehousing - Snowflake, BigQuery, Redshift, Synapse optimization
β
Data Transformation - dbt, PySpark, Apache Spark, SQL optimization
β
Orchestration - Apache Airflow, Cloud Composer, Data Factory automation
β
CI/CD & DevOps - GitHub Actions, Azure DevOps, Infrastructure as Code
β
Data Modeling - SCD, Dimensional modeling, Star Schema, Lakehouse architecture
β
Analytics & BI - PowerBI, Looker, Streamlit dashboard development
β
Full-Stack Pipelines - End-to-end implementations from ingestion to visualization
β
Documentation & Knowledge Sharing - Best practices, tutorials, mentoring
I believe in building production-grade data solutions that embody:
- Scalability - Handle data at any volume with optimal performance
- Reliability - Robust error handling, monitoring, and alerting
- Maintainability - Clean, well-documented, and testable code
- Cost-Efficiency - Optimized resource utilization and cloud spending
- Future-Proof - Aligned with industry best practices and emerging technologies
- GitHub Profile: github.com/ViinayKumaarMamidi
- All Repositories: 75+ Active Projects
I'm passionate about:
- π Designing scalable data architectures
- π Building reliable data pipelines
- βοΈ Leveraging cloud technologies effectively
- π Transforming data into actionable insights
- π€ Sharing knowledge and mentoring others
- βοΈ Exploring AI tools and leveraging to build data pipelines
Interested in collaborating? Feel free to explore my repositories, fork projects, or reach out with ideas!

