Skip to Content

Hi, my name is

Mayank Deshpande.

I love to build Software

I'm a Perception Engineer at Contoro Robotics, where I own the stack that lets our robots actually see what they're picking up—cameras, LiDAR, segmentation models, and a whole lot of 3D geometry.

Outside of work my head is mostly in Multimodal and Embodied AI, and I build RAG and AI engineering projects on weekends just to see how far I can push them. I also read about markets and finance almost every day.

Always looking to work on exciting projects—feel free to reach out if you'd like to build something amazing together!

About Me

I absolutely love talking Football, Finance, and Startups—and honestly, I think the life would've been way more exciting if I was already working in one of these (but hey, I'm getting there!).

Something else I find really cool? Robots. I did my Master's in Robotics at Maryland, where I spent countless hours diving into Controls, Perception, and way too much C++. These days I'm a Perception Engineer at Contoro Robotics, owning the perception stack for a fleet of autonomous unloading robots—RGB-D cameras, LiDAR, segmentation models, calibration, and the 3D geometry that ties it all together.

What actually got me here was a course on Multimodal Foundation Models. I got completely hooked and couldn't stop exploring them, and that's still where my head goes outside of work: Embodied AI, plus RAG and AI engineering projects I build on weekends just to see how far I can push them. One thing I've realized along the way: it's incredibly valuable to master one area deeply before branching out—being able to claim mastery over a domain has so much cross-applicability (Thank You, Kyle!)

I'm excited for all the great conversations and adventures ahead as I keep exploring this fascinating journey!

  • Languages: Python, C++ (11/14/17), CUDA, MATLAB, Bash
  • 3D Perception: Open3D, PCL, RANSAC & plane fitting, ICP, camera calibration, TF2, occupancy mapping, 6-DoF pose estimation
  • Deep Learning: PyTorch, Detectron2, Mask R-CNN, PointRend, DETR, SAM, DINOv2, ViT, ONNX, TensorRT
  • Robotics & Tooling: ROS1/ROS2, RGB-D cameras, 2D/3D LiDAR, MoveIt, Gazebo, CARLA, Docker, Git, pytest, GoogleTest
Mayank Deshpande - Robotics Software Engineer

Where I’ve Worked

Robotics Perception Engineer @ Contoro Robotics

Oct 2025 - Present

  • Own the perception stack for the entire fleet of autonomous box unloading robots. RGB-D cameras and 2D/3D LiDAR feeding segmentation, point-cloud reconstruction, 6-DoF grasp poses, and obstacle detection.
  • Raised segmentation recall 15 points on unseen sites and cut merged/split boxes 40%. Rebuilt the dataset pipeline around visual clustering and negative mining based on field heuristics.
  • Train and ship the segmentation and detection models the robots run: PointRend, DETR, SAM and ViT backbones in PyTorch and Detectron2. Redesigned the complete ML development architecture to be scalable and maintainable, allowing developers to systematically evaluate and improve model performance.
  • Cut inference time on the robot by 35% and got two models sharing one GPU, by fixing how frames are preprocessed and batched and when weights are held in memory. Exported through ONNX/TensorRT and quantized where accuracy allowed.
  • Developed a per-site adaptive learning loop that learns from operator corrections while the robot keeps unloading, then swaps the new weights asynchronously.
  • Took box dimension error from about 4 cm down to under 5 mm by measuring faces from mask pixels projected onto the fitted plane, improving the major pipeline algorithms and point cloud manipulation.
  • Own calibration for the sensor rig: intrinsics, extrinsics, hand-eye. Solved as rigid-body transforms in SE(3), checked against fiducial targets, and developed an auto calibration pipeline.
  • Wrote the point-cloud and 3D LiDAR layer: plane fitting, clustering, outlier rejection, occupancy mapping to construct a high accuracy World Model of the scene at a tight latency budget.
  • Developed strategies to tackle the common depth issues like dark surfaces drop out, depth edges throw flying pixels, and multipath interference. That work drives sensor choice and mounting geometry for the future hardware decisions.

Some Things I’ve Built

Other Noteworthy Projects

view the archive
  • Folder

    AutoPano

    Developed an automatic panorama stitching solution using traditional techniques and deep learning models (HomographyNet), achieving high-quality results with supervised and unsupervised learning, validated on synthetic and real-world image sets.

    • Python
    • OpenCV
    • git
  • Folder

    Human detection and Tracking

    A C++ module for detecting and tracking human obstacles using Monocular Camera with ResNet & OpenCV integration within the robot's reference frame.

    • C++
    • OpenCV
    • MiDAS Resnet
    • GoogleTest
    • CMake
  • Folder

    Stereo Vision

    A Python-based stereo vision pipeline to accurately generate disparity and depth maps from stereo image pairs with OpenCV and NumPy.

    • Python
    • OpenCV
    • NumPy
    • Matplotlib
  • Folder

    Pb-Lite and DL for Edge Detection

    An edge detection implementation that integrates Probability-based methods (Pb-Lite) with deep learning models like DenseNet, ResNext and a cutom architecture to improve accuracy and robustness.

    • Python
    • OpenCV
    • PyTorch
    • DenseNet
    • ResNext
  • Folder

    AutoCalib

    A Python implementation of Zhang's camera calibration method, automating checkerboard detection and accurately estimating camera parameters using OpenCV and NumPy.

    • Python
    • OpenCV
    • NumPy
    • Matplotlib
    • CMake
  • LQR and LQG for stabilizing Two Pendulum cart

    Modeled and controlled a two-pendulum crane system using cutom LQR and LQG controllers implementation in MATLAB.

    • MATLAB
    • LQR
    • LQG
    • Leunberger Observer

Publications

What’s Next?

Get In Touch

I'd love to connect! Feel free to email me, check out my resume above, or visit my LinkedIn page to learn more.