Lokesh S

ML Engineer | Backend Developer

I build machine learning systems, backend services, and data pipelines that transform complex data into scalable applications.

About

I'm a B.Tech Information Technology student focused on machine learning, backend engineering, and data-intensive systems. I enjoy building reliable software that transforms complex real-world data into scalable applications, from data pipelines and ML-powered services to evaluation systems and APIs. My work emphasizes clean architecture, automation, and practical problem-solving rather than isolated model development.

Projects

Built with product impact in mind.

Hands-on work across data engineering, ML systems, and backend architecture.

FloatChat – Oceanographic Data Engineering Pipeline

FloatChat – Oceanographic Data Engineering Pipeline

Designed and implemented a scalable data engineering pipeline for processing ARGO oceanographic datasets stored in NetCDF format. Built a multi-stage ETL workflow that automatically downloads, validates, transforms, and stores float data as optimized Parquet datasets while generating metadata for PostgreSQL and ChromaDB indexing. Implemented parallel processing and hash-based incremental updates to eliminate redundant computation and handle inconsistencies across large scientific datasets. The system enables efficient querying and downstream analytics on previously difficult-to-use oceanographic data.

Impact: Automated ingestion and transformation of large-scale scientific datasets, significantly reducing manual preprocessing effort while improving processing throughput through concurrent execution and incremental updates.

  • Python
  • PostgreSQL
  • ChromaDB
  • NetCDF
  • Pandas
  • Joblib
  • ThreadPoolExecutor
View on GitHub
InterviewPilot – AI-Powered Interview Simulation Platform

InterviewPilot – AI-Powered Interview Simulation Platform

Built an AI-powered interview simulation platform that conducts adaptive mock interviews and provides automated candidate evaluation. Designed a multi-stage backend pipeline that manages question generation, response processing, transcription analysis, and feedback generation while maintaining interview state across sessions. Implemented asynchronous workflows to decouple candidate interaction from computationally intensive evaluation tasks, enabling seamless interview progression without blocking user responses. Structured the system using modular agents and reusable services to support extensibility across different interview domains and skill levels.

Impact: Automates key stages of the interview process by combining adaptive questioning, response evaluation, and feedback generation into a unified workflow, reducing manual assessment effort and enabling scalable mock interview experiences.

  • Python
  • FastAPI
  • PostgreSQL
  • SQLAlchemy
  • AsyncIO
  • LLMs
  • REST APIs
  • Agent Architecture
View on GitHub
MindEase – Mental Health Risk Assessment System

MindEase – Mental Health Risk Assessment System

Developed a machine learning-based mental health assessment system using PHQ-9 questionnaire data and classical ML techniques. Built a complete training pipeline including data preprocessing, feature engineering, model evaluation, and deployment of a Random Forest classifier through a Flask API. Containerized the application using Docker and implemented automated testing workflows with GitHub Actions to support reliable deployment and reproducible development.

Impact: Automated mental health risk assessment through a deployable ML service, transforming questionnaire responses into real-time predictions while maintaining a reproducible development and deployment workflow.

  • Python
  • Flask
  • Scikit-learn
  • Random Forest
  • Docker
  • GitHub Actions
View on GitHub
Semiconductor Manufacturing Yield Analysis

Semiconductor Manufacturing Yield Analysis

Contributed to a data science project focused on semiconductor manufacturing process data. Worked with a high-dimensional industrial dataset containing noisy sensor measurements and process variables, performing data cleaning, exploratory analysis, feature selection, and predictive modeling to identify factors associated with manufacturing yield outcomes. Collaborated on preprocessing strategies and model evaluation to improve understanding of complex process behavior.

Impact: Explored methods for extracting insights from complex manufacturing data and evaluated machine learning approaches for yield-related prediction tasks.

  • Python
  • Pandas
  • NumPy
  • Scikit-learn
  • Matplotlib
  • Data Analysis
View on GitHub
Smart Parking Safety and Surveillance System

Smart Parking Safety and Surveillance System

Developed a research-oriented computer vision prototype for smart parking and vehicle surveillance. Implemented a multi-stage vision pipeline using YOLOv8 for vehicle detection, a YOLOv11-based license plate detector for plate localization, and EasyOCR for license plate recognition. The system processes road and parking-lot images, isolates individual vehicles and license plates, performs OCR, and displays recognized plate information with confidence scores.

Impact: Built and validated an end-to-end ALPR pipeline as the Computer Vision component of a published Smart Parking research project, demonstrating vehicle detection, license plate localization, and OCR-based identification.

  • Python
  • YOLOv8
  • YOLOv11
  • OpenCV
  • EasyOCR
  • PyTorch
  • Computer Vision
  • Object Detection
View on GitHub

Skills

Core capabilities

A practical stack optimized for ML experimentation and production backend delivery.

Programming visual

Programming

  • Python
  • Java
  • SQL
Machine Learning visual

Machine Learning

  • Scikit-learn
  • Feature Engineering
  • Model Evaluation
  • Classical ML Algorithms
  • ML Pipelines
  • OpenCV
Backend visual

Backend

  • Flask
  • Django
  • FastAPI
  • REST API Design
  • Async Programming
Data Engineering visual

Data Engineering

  • ETL Pipelines
  • Data Preprocessing
  • Parquet
  • Pandas
  • NetCDF Data Handling
DevOps & Tooling visual

DevOps & Tooling

  • Git
  • Docker
  • GitHub Actions
  • CI/CD Pipelines
Databases visual

Databases

  • PostgreSQL
  • MySQL
  • ChromaDB
  • Database Design

Publication

Smart Parking Safety and Surveillance System Using Computer Vision

Intelligent Transportation and Smart Systems, IGI Global (2026)

Proposed a multi-view smart parking pipeline integrating computer vision, IoT, and blockchain, with YOLO and OpenCV for real-time pedestrian and vehicle monitoring.

Read publication DOI

Contact

Let us build something meaningful.

Open to internships, backend and ML collaborations, and product engineering opportunities.

Reach out directly for project discussions, internships, or engineering opportunities.

Fill in your details and send via Gmail. A compose window opens in a new tab with your message pre-filled.