Skip to content
View adrianc7's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report adrianc7

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
adrianc7/README.md

Data network animation

Data Analytics Engineer

I consider myself an analytics engineer because I thrive on connecting data engineering with analytics. I want to be involved in the entire data journey, from the initial source systems all the way to the final insights and dashboards that drive business decisions.

I enjoy building scalable data systems that transform raw, messy data into reliable, analytics-ready datasets using ETL/ELT processes, data warehousing, and solid modeling practices.

I'm fascinated by how data moves through organizations and becomes actionable intelligence. I've become comfortable working with various tools and technologies including Python, SQL, Tableau, Power BI, and PySpark. I'm always exploring new techniques and tools for different business needs, currently diving into DuckDB, Apache Airflow, and streaming technologies like Kafka and Flink.

This field constantly evolves and I'm committed to continuous learning, currently focusing on pipeline automation, data quality, and building cloud data stacks that make analytics seamless and decision making more effective.


About Me

  • Graduate student in Management Information Systems (M.S.) at Utah State University (4.0 GPA) and International Commerce at Seoul National University
  • Fluent in English, Korean, and Spanish
  • Based in Korea, originally from Salt Lake City, Utah

Certifications

  • AWS Certified Data Engineer – Associate
  • Snowflake SnowPro Core Certification
  • Tableau Desktop Certified Professional
  • Professional Scrum Master I (PSM I)
  • TOPIK Level 6 (Advanced Korean Proficiency)

Featured Projects

Data Engineering & ETL Pipelines

Repository Description Tech Stack
Airbnb Analytics Engineering Project (Snowflake + dbt + Dagster) End-to-end ELT project built on Inside Airbnb open data for Berlin (2015–2021). Integrated Snowflake as the data warehouse, dbt for transformations and testing, and Dagster for orchestration and lineage tracking. Visualized insights in Preset (Apache Superset) showing host performance, pricing, and review trends. Snowflake · dbt · Dagster · Preset (Superset) · SQL · Jinja · ELT Orchestration · Dimensional Modeling
AWS Data ETL Health & Mortality Analysis Project Serverless AWS ETL pipeline analyzing PAHO/WHO health indicators across the Americas. Built a multi-stage workflow integrating data ingestion, transformation, and visualization using serverless compute. AWS S3 · AWS Glue · Athena · PySpark · Tableau · Data Lake Architecture · ETL Orchestration
Sam's Subs Data Warehouse (dbt + Snowflake) Designed a modern ELT architecture using Airbyte for ingestion from SQL Server into Snowflake. Applied dbt transformations and built Tableau dashboards to analyze sales, orders, and customer engagement. Snowflake · dbt · Airbyte · SQL Server · Tableau · Dimensional Modeling · ELT Orchestration
Spotify AWS Scalable Data Pipeline Scalable data pipeline for Spotify API data built on AWS. Automated ingestion, schema discovery, transformation, and querying to extract insights on music listening trends. AWS Lambda · S3 · Glue Crawlers · DataBrew · Athena · REST API Integration · Serverless ETL
Jetflix Snowflake Data Mart Built a complete ELT data pipeline and dimensional model for Jetflix, a mock streaming service. Designed OLTP-to-OLAP transformation using dbt in Snowflake and integrated AWS S3 as external stage storage. Snowflake · dbt · SQL · AWS S3 · Lucidchart · Star Schema Design · Incremental Models
Iron Will Sportsbook Data Mart Developed a PostgreSQL data mart combining SQL Server betting logs and NFL datasets. Built a Python ETL pipeline to extract, transform, and load structured data for profitability analytics, testing, and regression modeling on customer value and personality insights. PostgreSQL · Python · pandas · psycopg2 · AWS RDS · SQL Server · Data Normalization · Regression Analysis · Feature Engineering

Data Analysis & Machine Learning

Repository Description Tech Stack
World Happiness Indicators Analysis Comprehensive analysis of global happiness trends using R statistical analysis and Tableau visualization. Examined economic, social, and health factors influencing happiness across 150+ countries from 2015–2019. R · Tableau · Statistical Modeling · Correlation Analysis · Data Visualization
Predicting Initiating Events in Narrative Texts Machine learning analysis identifying story-initiating events using linguistic features. Random Forest model achieved 58.6% accuracy in predicting narrative structure from writing samples. Python · Random Forest · Feature Importance · Cross-validation · Natural Language Analysis
Near-Earth Objects NASA Hazard Classification Predictive modeling of asteroid hazard potential using NASA’s Near-Earth Object dataset. Logistic regression models achieved 89% accuracy in classifying hazardous asteroids. Python · Logistic Regression · LASSO · PCA · Feature Selection · Cross-validation
Key Indicators: Life Expectancy Analysis Regression-based study of WHO/UN global health data (2000–2015). Modeled the effects of socioeconomic, health, and education factors on life expectancy using multivariate and polynomial regression. Python · pandas · statsmodels · seaborn · Linear & Quadratic Regression · Feature Engineering
Advanced Python Projects & Exercises Collection of advanced Python projects including stock trading automation, cryptocurrency arbitrage detection, API ingestion, and algorithmic problem-solving with OOP design and analytical focus. Python · REST APIs · OOP · pandas · NumPy · Matplotlib · JSON Data Processing
Korean Lottery (Lotto 6/45) Data Analysis Exploratory data analysis of Korean Lotto 6/45 draws. Investigated statistical distributions, number frequency, and randomness using Python-based analytics and data visualization techniques. Python · pandas · matplotlib · seaborn · EDA · Probability Analysis
NBA Player Stats Pipeline (AWS RDS + S3) Built a lightweight cloud data workflow to process NBA player statistics, load data into AWS RDS (MySQL), and publish an HTML leaderboard to Amazon S3. Used Python for ETL operations and Boto3 for automated file upload and hosting setup. Python · pandas · PyMySQL · Boto3 · AWS RDS · Amazon S3 · Static Website Hosting

Interactive Dashboards & Visualizations

Project Description Dashboard Link Repository
Shopify Sales Performance Dashboard Analytical Power BI dashboard built on anonymized Shopify e-commerce data, visualizing sales KPIs, product performance, customer behavior, and payment method distribution. View Power BI Dashboard GitHub Repo
World Happiness Analysis Dashboard Interactive exploration of global happiness trends, regional comparisons, and factor importance across 150+ countries from 2015–2019. View Tableau Dashboard GitHub Repo
HR Attrition & Satisfaction Dashboard Comprehensive analysis of employee attrition patterns, satisfaction scores, and workforce demographics to identify retention risks and improvement opportunities. View Tableau Dashboard GitHub Repo
Formula 1 Races Analysis Dashboard Interactive exploration of F1 race statistics, team performance, and historical trends across multiple seasons and circuits. View Power BI Dashboard GitHub Repo
COVID-19 School Closures Analysis Analysis of global school closure patterns during the COVID-19 pandemic, examining regional differences and educational impact. View Power BI Dashboard GitHub Repo
Inter-University Summit Project SQL + Power BI case study demonstrating how SQL Views streamline analytics integration. Showcased performance tuning, pre-aggregation, and dashboard development for collaborative reporting. Documentation in Repository SQL Server · T-SQL · Power BI · Data Visualization · Query Optimization

Technical Skills

Programming & Machine Learning: Python (Pandas, NumPy, Scikit-learn, TensorFlow, PyTorch), R (dplyr, tidyr), regex, Pyarrow, XGBoost
Data Engineering & Cloud: AWS (S3, Glue, Lambda, EC2, Athena), dbt (with Jinja), Airflow, PySpark, Docker
Databases & Warehousing: PostgreSQL, Snowflake, BigQuery, Redshift, SQL Server, DuckDB
Data Modeling & Analytics: Feature Engineering, Regression, Forecasting, Hypothesis Testing, Data Pipelines, Kimball Methodology
Visualization & BI Tools: Tableau, Power BI, Preset, Seaborn, Matplotlib, ggplot2, Plotly
Development & Collaboration: Git, JupyterLab, VS Code/Cursor, Excel, PowerPoint, Bash/Shell


Connect with Me

Popular repositories Loading

  1. Deep_forecasting-USU Deep_forecasting-USU Public

    Forked from PJalgotrader/Deep_forecasting-USU

    GitHub repository for deep forecasting courses owned and maintained by prof. Jahangiry

    Jupyter Notebook

  2. Deep_Learning-USU Deep_Learning-USU Public

    Forked from PJalgotrader/Deep_Learning-USU

    GitHub repository for deep learning courses

    Jupyter Notebook

  3. Data-Science--Cheat-Sheet Data-Science--Cheat-Sheet Public

    Forked from georgearun/Data-Science--Cheat-Sheet

    Cheat Sheets

  4. advanced-python-projects-and-exercises advanced-python-projects-and-exercises Public

    This repository showcases advanced Python projects and coding challenges that explore key programming concepts. It demonstrates practical applications of Python for solving real-world problems and …

    Jupyter Notebook

  5. sams-subs-datawarehouse sams-subs-datawarehouse Public

    Jupyter Notebook

  6. inter-university-summit-project-presentation inter-university-summit-project-presentation Public