Building fault-tolerant, high-throughput AI data systems β from raw ingestion to semantic search.
I'm an AI Infrastructure Engineer, specialising in building resilient pipelines, LLM-powered classification systems, and vector embedding infrastructure at scale.
- π Currently @ SkAI
π Nexus
Because Urban traffic data is noisy, fast-moving, and difficult to use directly for analytics. Raw streaming events are not enough for decision-making because they need validation, structure, and business-friendly models before they can support monitoring or dashboards.
End-to-end traffic analytics pipeline built with Kafka, Spark Structured Streaming, Delta Lake, and a Streamlit dashboard
Python Apache Kafka Docker Delta Lake Spark
βοΈ DE-Boto3 Project
Turning raw S3 data into actionable sales intelligence β fast.
Engineered a Python-based ETL pipeline wiring AWS S3 + PySpark to a MySQL warehouse, achieving a 30% boost in processing speed and 20% reduction in data latency, with a custom PySpark aggregation framework for sales performance analytics.
Python AWS S3 PySpark MySQL Boto3 Git
| πΌ LinkedIn | linkedin.com/in/nikrrai | | π GitHub | github.com/NikrrGit |