The Story So Far...
Download the boring PDF version • Or keep reading for the fun one
Hi! I'm Harsh, and I've been on quite the journey from mechanical engineering to making computers see things. Currently, I'm helping NVIDIA's Maxine team make your video calls look less like potato quality and more like Hollywood productions. But let me back up a bit...
Fun fact: A decade ago, I was an NTSE scholar (yes, I was that kid who actually enjoyed standardized tests). These days, I channel that same nerdy enthusiasm into optimizing CUDA kernels and arguing about whether tabs or spaces are better (spaces, obviously).
The Academic Adventures
Carnegie Mellon University (2019-2021) was where I truly fell in love with making robots see. I pursued an MS in Robotic Systems Development at the School of Computer Science, which basically meant I spent two years teaching machines to understand the world around them. My coursework read like a computer vision enthusiast's wishlist: Computer Vision (obviously), Computational Photography (making pictures pretty with math), SLAM (helping robots not bump into walls), and Machine Learning (teaching computers to learn from their mistakes, unlike me with my coffee consumption).
I dove deep into Visual Learning and Recognition, and Image Synthesis - essentially learning how to make computers hallucinate in productive ways. My research focused on perception systems for autonomous robots, because someone needs to help these metal friends navigate our chaotic world.
Before CMU, I spent four years at IIT Indore (2013-2017) pretending to be a Mechanical Engineer while secretly taking every computer vision and ML course I could find. Got my B.Tech in Mechanical Engineering, but my heart was already with the pixels and matrices. The mechanical background actually helps though - I understand why robots break, not just how to program them!
The Professional Journey (or: How I Learned to Stop Worrying and Love the GPU)
NVIDIA (June 2021 - Present): This is where the magic happens. These days I'm in AI for Media, working on human motion models and agents — teaching computers to understand how bodies actually move, from a single camera, in real time. Before that I spent years on Maxine making pandemic-era video calls look less terrible, including shipping Eye Contact, the feature that quietly fixes your gaze so you look like you're paying attention. (You're welcome.)
The through-line of my career is the unglamorous gap between "the research model works" and "the research model ships." That means C++/CUDA integration, TensorRT inference pipelines, multi-stream batching, and building the evaluation harnesses that catch a model quietly getting worse before a customer does. Nobody puts eval infrastructure on a conference slide. It's also the thing that decides whether anything ships.
Lately I've been dragged happily into LLM territory — post-training, LoRA fine-tuning, and adversarial evaluation of agents. Coming from computer vision, the punchline is that it's the same problem in a new outfit: you still can't improve what you can't measure.
Carnegie Robotics (Summer 2020): Spent a summer in Pittsburgh teaching robots where they are (harder than it sounds). I worked on localization using AprilTags - think QR codes but for robots. Used fancy math (bundle adjustment in Ceres Solver) to help robots figure out where they were even when their sensors were lying to them. Also built obstacle detection pipelines because robots running into walls is only funny the first time.
CMU Biorobotics Lab (2018-2019): This was wild. Under Prof. Howie Choset and Matt Travers, I worked on everything from evaluating culinary skills with Hidden Markov Models (yes, we taught computers to judge cooking) to preparing robots for the DARPA SubT Challenge. Imagine sending robot teams into underground mines and tunnels - it's as cool and terrifying as it sounds. I spent a lot of time making point clouds play nice with each other and building simulations where virtual robots could practice being brave in dangerous places.
Addverb Technologies (2017-2018): Fresh out of undergrad, I had the audacity to found and lead the autonomous robotics division. Built the team from scratch, developed navigation pipelines for warehouse robots, and wrote a lot of C++ to make motors do exactly what we wanted. It was like being a robot whisperer, except with more debugging and less whispering.
Teaching at CMU (Fall 2020): Was the lead TA for Computer Vision, specifically the 3D reconstruction module. Taught students how to track things (Lucas-Kanade), transform perspectives (homography), and build 3D models from 2D images. The best part? Watching students have that "aha!" moment when the math suddenly makes sense and their code actually works.
Judging Other People's Homework
Program Committee Reviewer, HuMoGen Workshop @ CVPR 2026: Reviewing human motion generation papers for the biggest venue in computer vision. Reading other people's work carefully is the cheapest way I know to stay current — you see where the field is going about a year before it shows up in a product.
Judge & Panelist, Open World Hackathon (VLGE AI, SF, 2026): A day on physical AI with 100+ builders and $10K across three tracks, judging alongside people from Google DeepMind, NASA, Meta and the Harvard Wyss Institute. Also sat on the closing panel arguing about the sim-to-real gap, which is my favorite thing to argue about.
Judge, RoboHacks @ Y Combinator (2025): SF's largest general-purpose robotics hackathon. 25 teams, 120+ builders, robots doing manipulation and navigation in a building full of people who very much wanted their robot to work.
Judge, Intelligence at the Frontier (DevSpot): Evaluated AI projects from teams pushing on frontier model capabilities.
Lead TA, 16-720A Computer Vision @ CMU (Fall 2020): 100+ students, and I owned the 3D reconstruction and multiview geometry assignment. Best part was watching the "aha" moment land when the math finally clicked. Full details here.
The Trophy Cabinet (Both Physical and Virtual)
Verdictron - Triple crown winner at the Weights & Biases Multimodal AI Hackathon. We built an AI that could actually understand legal documents (a miracle, really).
NTSE Scholar (2011) - A decade ago, I was that nerdy kid who got a national scholarship for being good at standardized tests. My parents were thrilled. I used the money to buy more programming books.
Various Hackathon Wins - Let's just say I've consumed unhealthy amounts of Red Bull in the pursuit of glory and free t-shirts. The projects ranged from "surprisingly useful" to "what were we thinking at 3 AM?"
The Technical Arsenal (or: Things I Can Do Without Stack Overflow)
Languages I Speak Fluently: C++ (my first love), Python (my daily driver), and CUDA (for when Python is too slow and I need to summon the GPU gods). I can also debug JavaScript, but I prefer not to talk about it.
Deep Learning & Vision Magic: I'm fluent in PyTorch (sorry TensorFlow), can make TensorRT optimizations that would make your GPU weep tears of joy, and OpenCV is basically my Swiss Army knife. I've done everything from making neural networks see in 3D to teaching them to track objects like a particularly obsessive cat.
Robot Wrangling: ROS and ROS2 are my go-to for making robots do things. I've spent quality time with Gazebo simulations (where robots can crash without expensive consequences), Point Cloud Library for 3D data wrangling, and various SLAM implementations because robots need to know where they are.
The Boring But Essential Stuff: Git (I've resolved merge conflicts that would make grown developers cry), CMake (because someone has to understand those build files), Docker (for when "it works on my machine" isn't good enough), and I actually write tests (I know, shocking).
Special Superpowers: Making things run in real-time that have no business running in real-time, optimizing code until it's unrecognizable but 100x faster, and explaining complex computer vision concepts using food analogies.
When I'm Not Teaching Computers to See
Fitness Enthusiast: I lift weights with the same precision I apply to optimizing algorithms. My PRs in the gym are as carefully tracked as my git commits. Currently exploring the intersection of longevity research and not dying early from too much sitting.
Reading Addiction: I consume books on everything from quantum computing to philosophy. Currently deep in the rabbit hole of reasoning models in LLMs and why consciousness is weird. My Kindle library looks like it belongs to three different people having an identity crisis.
Investment Tinkering: I analyze stock charts with the same intensity I debug segfaults. Particularly interested in the solar energy revolution and why my NVIDIA stock options make me very happy. Also trying to understand crypto without losing my shirt.
Cooking Experiments: I approach cooking like I approach coding - lots of iteration, occasional spectacular failures, and a git-like version control system in my head for recipes. My kitchen has more sensors than necessary because why not?
Travel & Photography: I take photos with the dedication of someone who's implemented computational photography algorithms. Every vacation becomes a dataset collection opportunity. Yes, I've calculated the optimal HDR settings for sunset photos.
Building Things: Whether it's a new ML model, a ridiculous home automation system, or furniture from IKEA, I love creating things. The IKEA furniture usually requires less debugging than the ML models.
Let's Connect!
If you want to talk about computer vision, debate whether we'll achieve AGI before 2030, discuss why C++ is still relevant, or just share memes about segmentation faults, I'm always up for a chat!
📧 Email • 🐙 GitHub • 💼 LinkedIn • 🐦 Twitter • 📝 Substack