- Built a multi-vehicle tracking REST API using FastAPI and YOLOv8n, maintaining persistent object identities across occlusions (up to 5 missed frames) by implementing SORT (Simple Online and Realtime Tracking).
- Improved real-time cloud processing speed on Hugging Face Spaces by bypassing YOLO detection on 50% of frames via Kalman predictions, utilizing FP16 inference, and encoding WebM streams.
- Conducted workshops and classes on Machine and Deep Learning for 100+ BE/BCA/Diploma students
Experience
- Jan 26 - May 26Tech Fortune Technologies, BengaluruData Science Intern
- May 25 - Sep 25Dept. of Neurosurgery, AIIMS DelhiResearch Intern
- Built a multi-agent AI system using Llama vision language models and Visual RAG to automate surgical suturing evaluation. Curated a 514-image dataset, achieving an RMSE of 2.2/10 against surgeon annotations.
- Deployed a Python/PyQt5 computer vision application with multi-threaded OpenCV live feeds for real-time craniotomy scoring. Integrated a Grad-CAM pipeline to visually highlight spatial drilling performance.
- Configured 3 ESP32S3 controllers for a wearable device to study hand-eye coordination during endoscopy.
- May 24 - Jun 24IIT DelhiGraphics & Vision Summer School
Participated in a 8-week program at IIT Delhi covering image processing, computer vision, and physically-based rendering.
- Implemented CNNs from scratch for visual recognition and constructed feature extraction pipelines with PyTorch.
- Built a multi-image house price prediction model that fuses visual features (facade appearance, neighborhood density) with structured regression — bridging visual perception and quantitative inference.
- Explored photorealistic 3D rendering through ray-tracing algorithms and light transport simulation techniques.