Projects

Engineering projects from NVIDIA, CMU, and personal exploration in ML/Robotics.

2D/3D Pose Estimation for Mesh Recovery

NVIDIA • 2021–present

Working on 2D and 3D human pose estimation and mesh recovery, and optimizing models for large-scale inference.

Eye Contact — Real-Time AI Gaze Correction

NVIDIA • shipped in NVIDIA Broadcast and Maxine

Worked end to end on Eye Contact, a live AI feature that redirects your gaze to the camera during a call. It runs in NVIDIA Broadcast and in enterprise video conferencing today.

Single-Camera Sports Biomechanics

NVIDIA × UNC Blue Sky Innovations • shown at NAB Show

Human pose models applied to sports: real-time biomechanics and performance stats for athletes from a single camera angle, for broadcast analytics and sports medicine. I work on the pose and evaluation side — multi-person tracking in crowded scenes is a very different problem from one person at a webcam.

NVIDIA Maxine — AI Video Features

NVIDIA • 2021–present

Built and shipped AI features for video conferencing in the Maxine SDK, and owned the release and testing infrastructure behind them.

Realistic Scenario Generation in CARLA (Capstone)

CMU with Ridecell • Jan–Aug 2020

Implemented goal-conditioned imitation learning for modeling traffic agents using expert trajectories in bird's eye view. Used DeepLabV3+ with MobileNet for vehicle motion prediction. Developed 3D to 2D homography projection for generating ground truth bounding boxes and built multi-object tracker for BEV vehicle tracking.

Python Deep Learning CARLA Detectron2 Autonomous Driving

SLAM Implementation with AMCL & Visual Odometry

Carnegie Mellon University • Aug–Dec 2020

Implemented Adaptive Monte Carlo Localization (AMCL) in C++ with OpenMP, achieving 10Hz performance with 5000 particles. Built monocular camera tracking framework integrating classic visual odometry with DL-based image-scan alignment using GTSAM factor graphs, showing improved performance over pure VO.

C++ Python SLAM GTSAM OpenMP

Mobile Robot Localization & Obstacle Detection

Carnegie Robotics • Summer 2020
Advised by David LaRose, then Chief Scientist at Carnegie Robotics, now at Sooth Labs

Formulated bundle adjustment problem in Ceres Solver for offline robot localization using AprilTags. Developed software pipeline for obstacle detection using UV disparity maps in C++ and ROS2, achieving real-time performance for mobile robot navigation.

C++ ROS2 Ceres Solver Computer Vision AprilTags

Image Synthesis with GANs and NeRF

CMU Course Project • Jan 2021–May 2021

Implemented various generative models including Poisson blending, DCGAN, CycleGAN, and StyleGAN using PyTorch. Explored Neural Style Transfer and used Neural Radiance Fields (NeRF) for super-resolution. Gained deep understanding of modern generative AI techniques.

PyTorch GANs NeRF Computer Vision Deep Learning

Culinary Skills Evaluation — Knife Motion as an HMM

CMU Biorobotics Lab × Culinary Institute of America • 2018–2019
Advised by Matt Travers and Howie Choset

Can you tell an expert from a novice by how they hold a knife? We instrumented knives with IMUs and recorded chefs julienning potatoes, then modelled the gyroscope trace as a 6-state 2D Gaussian HMM — time and angular velocity as the observed variables, transition probabilities initialised to force a sequential structure.

Cuts were normalised and temporally aligned with dynamic time warping before fitting, then by peak alignment when DTW smeared the fast strokes. The result is visible without statistics: an expert's cuts collapse into a tight, repeatable trajectory; a novice's wander. The HMM states are essentially the phases of a single cut — lift, drive, contact, recover.

Python HMM DTW IMU ROS Time Series

DARPA SubT Challenge & Point Cloud Registration

CMU Biorobotics Lab • Jan 2018–Aug 2019

Evaluated point cloud registration methods (ICP, GICP, GO-ICP) against proposed algorithms. Built multi-robot Gazebo simulation for DARPA Subterranean Challenge.

Python C++ ROS PCL Gazebo

Computational Photography Suite

CMU Course • March–May 2021

Implemented uncalibrated photometric stereo, depth from focus, and structured light triangulation. Developed white balancing, color correction, and tone mapping algorithms. Created depth estimation from light field images and used GDP for image reconstruction.

Python Computer Vision Image Processing 3D Reconstruction

Autonomous Mobile Robot Navigation Stack

Addverb Technologies • June 2017–Jan 2018

Started and led the autonomous robotics division. Developed complete navigation pipeline for Autonomous Mobile Robots (AMRs) using ROS Navigation Stack. Implemented lower-level motor controls in C++ and recruited team members for the new division. Reliance later acquired a 54% stake for $132 million; customers during my time included Patanjali, ITC and Coca-Cola.

C++ Python ROS Navigation Robotics

Computer Vision Fundamentals (Lead TA)

CMU Teaching • Fall 2020

Implemented variants of Lucas-Kanade tracking algorithms, homography estimation, and multiview 3D reconstruction from scratch. Served as Teaching Assistant and led the 3D reconstruction assignment module, helping students understand fundamental CV concepts.

Python Computer Vision 3D Reconstruction Teaching

6D Object Pose Estimation with RGBD

CMU Visual Learning • March–May 2021

Reviewed and implemented various 6D object pose estimation refinement techniques using RGBD input. Compared different approaches for accuracy and computational efficiency. Explored applications in robotic manipulation and augmented reality.

PyTorch 3D Vision RGBD Deep Learning