I specialize in optimizing AI model inference on edge devices and analyzing low-level kernels for large-scale model serving. I enjoy solving complex engineering problems with a focus on system efficiency and latency reduction.
- 🎓 M.S. Candidate in Computer Software at Hanyang University (Feb. 2025)
- 🔭 Research Focus: On-device AI Optimization, Model Quantization (INT8/FP16), LLM Serving
- 💼 Experience: Former Front-End Developer at SmartDoctor
- 🏆 Awards: HCPC (Hanyang Collegiate Programming Contest) Beginner Division - 2nd Place
🔒 Note: Most research codes (Thesis, vLLM analysis) are private due to security regulations. Please refer to my portfolio for details.
- Real-time Object Detection Optimization on Embedded Systems
- Optimized inference speed by 75% on NVIDIA Jetson Orin Nano via INT8 Quantization & TensorRT.
- vLLM Attention Kernel Analysis
- Analyzed PagedAttention CUDA kernels and memory hierarchy for Low-bit quantization implementation.
- Algorithm Problem Solving
- Solving various algorithmic problems (Graph, DP, Greedy) with a focus on time complexity.

