Harsh Sharma

Harsh Sharma

sharmaharsh2308 [at] gmail [dot] com

I'm a Computer Vision and ML Engineer at NVIDIA, in AI for Media, working on human motion models and agents.

My job is essentially the distance between a research checkpoint and something that survives production. That means C++ and CUDA integration, TensorRT inference pipelines, multi-stream batching — and the evaluation harnesses that tell you when a model is quietly getting worse. Evaluation is the part nobody puts on a slide and the part that decides whether a model ships.

At NVIDIA, I work on:

(1) Generative 3D human pose and motion estimation, from research model to shipped SDK
(2) Large-scale evaluation across real and synthetic benchmarks
(3) LLM and vision-language post-training, LoRA fine-tuning, and adversarial evaluation of agents
(4) CUDA-accelerated inference for real-time AI video, at scale

Outside of shipping, I review for the HuMoGen workshop at CVPR 2026 and judge robotics and AI hackathons around the Bay Area — most recently on the physical-AI panel at the Open World Hackathon, alongside folks from Google DeepMind, NASA and Meta. More on that here.

I completed my Master of Science in Robotics from Carnegie Mellon University's School of Computer Science (2021), where I focused on Computer Vision and SLAM. Before grad school, I worked at CMU (2018-19) on the DARPA SubT Challenge and collaborated with the Culinary Institute of America on the future of culinary education using AI.

Prior to CMU, I was one of the early engineers at Addverb Technologies, a robotics startup that was acquired by Reliance for $132 million. At Addverb, I worked on the full stack for warehouse autonomous robots—from perception to planning prototypes—and helped recruit the founding engineering team. The company quickly scaled to serve major clients like Patanjali, ITC, and Coca-Cola. I graduated from IIT Indore (2017) with a BTech in Mechanical Engineering, focused on ML, Computer Vision, and Mechatronics.

The best way to reach me is via email. Connect with me:

Email GitHub LinkedIn Twitter Substack

Featured Projects

1.
NVIDIA · AI for Media
Generative 3D Human Motion — Research Model to Shipped SDK
2025 - Present
Taking generative 3D body pose estimation from a research model into a production SDK: C++/CUDA integration, TensorRT inference, multi-stream batching, and cross-skeleton conversion across COCO-17, SMPL-24 and proprietary joint formats. Built the golden-eval suite that gates every release across real (EMDB, 3DPW, H36M) and synthetic benchmarks.
2.
NVIDIA
Eye Contact — Real-Time AI Gaze Correction
Shipped in NVIDIA Broadcast & Maxine
Worked end to end on Eye Contact, a live AI feature in the Maxine AR/VFX SDK — model work through SDK productization and CUDA acceleration for real-time deployment. It runs in NVIDIA Broadcast and in enterprise video conferencing today.
3.
NVIDIA × UNC Blue Sky Innovations
Single-Camera Sports Biomechanics
Shown at NAB Show
Human pose models applied to sports: real-time biomechanics and performance stats for athletes from a single camera angle, for broadcast analytics and sports medicine. I work on the pose and evaluation side — multi-person tracking in crowded scenes is a very different problem from one person at a webcam.
4.
Addverb Technologies
Founding Robotics Engineer — Warehouse Autonomy
2017 - 2018
Founded and staffed the autonomous robotics division: full navigation stack for autonomous mobile robots in ROS, plus low-level motor control in C++. Reliance later acquired a 54% stake for $132M; customers during my time included Patanjali, ITC and Coca-Cola.
5.
CMU Biorobotics Lab
DARPA Subterranean Challenge & Robot Perception
Jan 2018 - Aug 2019
Multi-robot Gazebo simulation for the DARPA SubT Challenge — autonomous navigation in GPS-denied underground environments. Also evaluated point cloud registration methods (ICP, GICP, GO-ICP) and built an HMM-based skill evaluation pipeline in ROS.