I consider myself an analytics engineer because I thrive on connecting data engineering with analytics. I want to be involved in the entire data journey, from the initial source systems all the way to the final insights and dashboards that drive business decisions.
I enjoy building scalable data systems that transform raw, messy data into reliable, analytics-ready datasets using ETL/ELT processes, data warehousing, and solid modeling practices.
I'm fascinated by how data moves through organizations and becomes actionable intelligence. I've become comfortable working with various tools and technologies including Python, SQL, Tableau, Power BI, and PySpark. I'm always exploring new techniques and tools for different business needs, currently diving into DuckDB, Apache Airflow, and streaming technologies like Kafka and Flink.
This field constantly evolves and I'm committed to continuous learning, currently focusing on pipeline automation, data quality, and building cloud data stacks that make analytics seamless and decision making more effective.
- Graduate student in Management Information Systems (M.S.) at Utah State University (4.0 GPA) and International Commerce at Seoul National University
- Fluent in English, Korean, and Spanish
- Based in Korea, originally from Salt Lake City, Utah
- AWS Certified Data Engineer – Associate
- Snowflake SnowPro Core Certification
- Tableau Desktop Certified Professional
- Professional Scrum Master I (PSM I)
- TOPIK Level 6 (Advanced Korean Proficiency)
| Repository | Description | Tech Stack |
|---|---|---|
| Airbnb Analytics Engineering Project (Snowflake + dbt + Dagster) | End-to-end ELT project built on Inside Airbnb open data for Berlin (2015–2021). Integrated Snowflake as the data warehouse, dbt for transformations and testing, and Dagster for orchestration and lineage tracking. Visualized insights in Preset (Apache Superset) showing host performance, pricing, and review trends. | Snowflake · dbt · Dagster · Preset (Superset) · SQL · Jinja · ELT Orchestration · Dimensional Modeling |
| AWS Data ETL Health & Mortality Analysis Project | Serverless AWS ETL pipeline analyzing PAHO/WHO health indicators across the Americas. Built a multi-stage workflow integrating data ingestion, transformation, and visualization using serverless compute. | AWS S3 · AWS Glue · Athena · PySpark · Tableau · Data Lake Architecture · ETL Orchestration |
| Sam's Subs Data Warehouse (dbt + Snowflake) | Designed a modern ELT architecture using Airbyte for ingestion from SQL Server into Snowflake. Applied dbt transformations and built Tableau dashboards to analyze sales, orders, and customer engagement. | Snowflake · dbt · Airbyte · SQL Server · Tableau · Dimensional Modeling · ELT Orchestration |
| Spotify AWS Scalable Data Pipeline | Scalable data pipeline for Spotify API data built on AWS. Automated ingestion, schema discovery, transformation, and querying to extract insights on music listening trends. | AWS Lambda · S3 · Glue Crawlers · DataBrew · Athena · REST API Integration · Serverless ETL |
| Jetflix Snowflake Data Mart | Built a complete ELT data pipeline and dimensional model for Jetflix, a mock streaming service. Designed OLTP-to-OLAP transformation using dbt in Snowflake and integrated AWS S3 as external stage storage. | Snowflake · dbt · SQL · AWS S3 · Lucidchart · Star Schema Design · Incremental Models |
| Iron Will Sportsbook Data Mart | Developed a PostgreSQL data mart combining SQL Server betting logs and NFL datasets. Built a Python ETL pipeline to extract, transform, and load structured data for profitability analytics, testing, and regression modeling on customer value and personality insights. | PostgreSQL · Python · pandas · psycopg2 · AWS RDS · SQL Server · Data Normalization · Regression Analysis · Feature Engineering |
| Repository | Description | Tech Stack |
|---|---|---|
| World Happiness Indicators Analysis | Comprehensive analysis of global happiness trends using R statistical analysis and Tableau visualization. Examined economic, social, and health factors influencing happiness across 150+ countries from 2015–2019. | R · Tableau · Statistical Modeling · Correlation Analysis · Data Visualization |
| Predicting Initiating Events in Narrative Texts | Machine learning analysis identifying story-initiating events using linguistic features. Random Forest model achieved 58.6% accuracy in predicting narrative structure from writing samples. | Python · Random Forest · Feature Importance · Cross-validation · Natural Language Analysis |
| Near-Earth Objects NASA Hazard Classification | Predictive modeling of asteroid hazard potential using NASA’s Near-Earth Object dataset. Logistic regression models achieved 89% accuracy in classifying hazardous asteroids. | Python · Logistic Regression · LASSO · PCA · Feature Selection · Cross-validation |
| Key Indicators: Life Expectancy Analysis | Regression-based study of WHO/UN global health data (2000–2015). Modeled the effects of socioeconomic, health, and education factors on life expectancy using multivariate and polynomial regression. | Python · pandas · statsmodels · seaborn · Linear & Quadratic Regression · Feature Engineering |
| Advanced Python Projects & Exercises | Collection of advanced Python projects including stock trading automation, cryptocurrency arbitrage detection, API ingestion, and algorithmic problem-solving with OOP design and analytical focus. | Python · REST APIs · OOP · pandas · NumPy · Matplotlib · JSON Data Processing |
| Korean Lottery (Lotto 6/45) Data Analysis | Exploratory data analysis of Korean Lotto 6/45 draws. Investigated statistical distributions, number frequency, and randomness using Python-based analytics and data visualization techniques. | Python · pandas · matplotlib · seaborn · EDA · Probability Analysis |
| NBA Player Stats Pipeline (AWS RDS + S3) | Built a lightweight cloud data workflow to process NBA player statistics, load data into AWS RDS (MySQL), and publish an HTML leaderboard to Amazon S3. Used Python for ETL operations and Boto3 for automated file upload and hosting setup. | Python · pandas · PyMySQL · Boto3 · AWS RDS · Amazon S3 · Static Website Hosting |
| Project | Description | Dashboard Link | Repository |
|---|---|---|---|
| Shopify Sales Performance Dashboard | Analytical Power BI dashboard built on anonymized Shopify e-commerce data, visualizing sales KPIs, product performance, customer behavior, and payment method distribution. | View Power BI Dashboard | GitHub Repo |
| World Happiness Analysis Dashboard | Interactive exploration of global happiness trends, regional comparisons, and factor importance across 150+ countries from 2015–2019. | View Tableau Dashboard | GitHub Repo |
| HR Attrition & Satisfaction Dashboard | Comprehensive analysis of employee attrition patterns, satisfaction scores, and workforce demographics to identify retention risks and improvement opportunities. | View Tableau Dashboard | GitHub Repo |
| Formula 1 Races Analysis Dashboard | Interactive exploration of F1 race statistics, team performance, and historical trends across multiple seasons and circuits. | View Power BI Dashboard | GitHub Repo |
| COVID-19 School Closures Analysis | Analysis of global school closure patterns during the COVID-19 pandemic, examining regional differences and educational impact. | View Power BI Dashboard | GitHub Repo |
| Inter-University Summit Project | SQL + Power BI case study demonstrating how SQL Views streamline analytics integration. Showcased performance tuning, pre-aggregation, and dashboard development for collaborative reporting. | Documentation in Repository | SQL Server · T-SQL · Power BI · Data Visualization · Query Optimization |
Programming & Machine Learning: Python (Pandas, NumPy, Scikit-learn, TensorFlow, PyTorch), R (dplyr, tidyr), regex, Pyarrow, XGBoost
Data Engineering & Cloud: AWS (S3, Glue, Lambda, EC2, Athena), dbt (with Jinja), Airflow, PySpark, Docker
Databases & Warehousing: PostgreSQL, Snowflake, BigQuery, Redshift, SQL Server, DuckDB
Data Modeling & Analytics: Feature Engineering, Regression, Forecasting, Hypothesis Testing, Data Pipelines, Kimball Methodology
Visualization & BI Tools: Tableau, Power BI, Preset, Seaborn, Matplotlib, ggplot2, Plotly
Development & Collaboration: Git, JupyterLab, VS Code/Cursor, Excel, PowerPoint, Bash/Shell
