Deepak Kumar Singh

I'm a Computer Vision and Deep Learning Engineer at Mercedes-Benz Research and Development India, Bengaluru, part of the Interior-Sensing (IS) Algorithm team. I build production-grade computer vision systems for in-cabin automotive applications — including face recognition, depth estimation, and IR-based occupant analysis — owning the pipeline from proof-of-concept to deployment.

I completed my MS by Research at CVIT, IIIT Hyderabad, advised by Prof. C.V. Jawahar and co-advised by Prof. Vineeth N Balasubramanian, Prof. Chetan Arora, and Dr. Anbumani Subramanian. My current areas of interest are Face Recognition, Vision-Language Models, Contrastive Learning, Knowledge Distillation, and Open-World / Incremental Learning.

I like photography, majorly urban photography and landscapes; check out my Flickr. I've recently started learning Bansuri(Indian Flute). To know what I'm currently reading, checkout my Goodreads.

Email: deepaksinghcv@gmail.com

|    CV  /  LinkedIn  /  Github     | 

profile photo
Research

My current work spans in-cabin face recognition, cross-domain (IR↔RGB) learning, and occupant analysis for automotive systems. Prior research focused on Open-World Object Detection and benchmarking detection/segmentation models on road-scene datasets. I'm broadly interested in translating research into scalable, deployable systems. Authors marked with a * have an equal contribution in the research.

News
  • [Mar 2026] : 🎉 Received Kaizen (Operational Excellence) Recognition at the department level for innovation leading to significant cost and efficiency gains.
  • [Nov 2025] : Presented benefits, optimizations, and achievements of our solutions at World Quality Week.
  • [Sep 2025] : Received Customer Appreciation Recognition from the Mercedes-Benz client team.
  • [Feb 2025] : Received the Impact Award (Silver) for delivering high-impact ML solutions.
  • [Oct 2024] : Received the Team Excellence Award (Silver) for collaborative delivery of production-grade systems.
  • [2023] : Received the Impact Award (Silver).
  • [Oct 2022] : Presented our work, "New Vehicles on the Road? No Problem - We'll Learn Them Too" at IROS 2022, Kyoto, Japan.
  • [Apr 2022] : Joined Mercedes-Benz Research and Development India as a Computer Vision and Deep Learning Engineer.
  • [Dec 2021] : Completed MS by Research at CVIT, IIIT Hyderabad.
  • [Dec 2021] : Presented our work, "ORDER: Open World Object Detection on Road Scenes" at NeurIPS Workshop on Machine Learning for Autonomous Driving 2021.
  • [Dec 2021] : Presented our work, "Evaluation of Detection and Segmentation Tasks on Driving Datasets" at CVIP 2021.
  • [Oct 2021] : 🎉 One paper accepted in NeurIPS Workshop on Machine Learning for Autonomous Driving 2021.
  • [Oct 2021] : 🎉 One paper accepted in CVIP 2021 for Oral presentation.
  • [Aug 2021] : In the organizing team for IIIT-H's annual Summer School Program
  • [Aug 2020] : Joined Vision for Mobility and Safety Team in CVIT.
  • [Nov 2019] : Attended NCVPRIPG at Hubbali, Karnataka.
  • [Jan 2019] : Joined CVIT as a Research Fellow.
  • [Dec 2018] : Got the opportunity to attend ICVGIP 2018 and talk to various speakers and fellow researchers.
  • [Aug 2018] : Joined IIIT-H as an MS by Research student.
Experience
Computer Vision and Deep Learning Engineer, Mercedes-Benz Research and Development India — Apr 2022 – Present
Bengaluru, India

Member of the Interior-Sensing (IS) Algorithm team, building computer vision solutions for in-cabin automotive systems across face recognition, depth estimation, and weight prediction — owning the pipeline from proof-of-concept to production.

  • Led end-to-end development of production-grade in-cabin face recognition systems, from data strategy and model development to deployment, achieving stringent accuracy and robustness targets.
  • Architected a dual-level authentication framework enabling both low-latency personalization and high-assurance security for privacy- and payment-sensitive use cases.
  • Developed cross-domain face recognition and IR-to-RGB colorization solutions to bridge RGB/NIR modality gaps and reduce dependency on large-scale RGB data collection.
  • Built deep learning models for occupant analysis, including IR-based weight estimation, replacing hardware sensors with software-driven solutions.
  • Collaborated across software, quantization, validation, and platform teams; drove GCP-based ML infrastructure initiatives and contributed to patentable innovations.
Software Engineer, Celstream Systems Pvt. Ltd. — Sep 2014 – Sep 2016
Bangalore, India
  • Built the product's main JavaScript UI console and dynamic data-visualization modules (IgniteUI), and led migration of the in-house application from Adobe Flash to a JavaScript environment.
  • Developed REST APIs in Java and data adapters powering live data-visualization modules.
Publications
New Vehicles on the Road? No Problem - We'll Learn Them Too
Deepak Kumar Singh, Shyam Nandan Rai, K J Joseph, Rohit Saluja,
Vineeth N Balasubramanian, Chetan Arora, Anbumani Subramanian, C.V. Jawahar
IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2022, Kyoto, Japan
Paper

Extends our open-world object detection work to incrementally adapt to new vehicle categories without catastrophic forgetting, improving robustness for real-world autonomous driving deployments.

ORDER: Open World Object Detection on Road Scenes
Deepak Kumar Singh*, Shyam Nandan Rai*, K J Joseph, Rohit Saluja,
Vineeth N Balasubramanian, Chetan Arora, Anbumani Subramanian, C.V. Jawahar
NeurIPS 2021 Workshop on Machine Learning for Autonomous Driving
Poster / Paper / Video(Soon) /

This work formulates Open-World Object Detection by addressing the inherent issues present in road scene datasets India Driving Dataset(IDD) and Berkeley Deep Drive(BDD).

Evaluation of Detection and Segmentation Tasks on Driving Datasets
Deepak Kumar Singh, Ameet Rahane, Ajoy Mondal, Anbumani Subramanian, C.V. Jawahar
CVIP, 2021 (Oral)
Paper / Project Page(Soon) / Slides(Soon) / Video(Soon) /

Benchmark latest models of object detection, semantic segmentation, and instance segmentation on road scene datasets. We perform a detailed study on their behaviour on both constrained and unconstrained datasets; Cityscapes, India Driving Dataset(IDD) and Berkeley Deep Drive(BDD).

Patents
  • In-Cabin Face Authentication SystemFiled in Germany and India
    Implemented Patent Award · Inventor lead. Enables secure multi-level authentication in automotive systems.

  • Intelligent In-Vehicle Hyper-PersonalizationFiled in Germany and India
    Co-Inventor. Enhances user experience by assimilating signals, memory, and preferences.
Personal Projects
DocLens — Local Multimodal RAG for Technical Documents
Python · sentence-transformers · ChromaDB · Ollama (Qwen3 / Qwen-VL) · PyMuPDF · RAG · cross-encoder reranking
GitHub

Built a fully-local multimodal RAG assistant for technical PDFs with two-stage retrieval (bi-encoder recall + cross-encoder reranking) and page-level citations, running entirely on-device with no external APIs. Figures are extracted, captioned by a local VLM, embedded into a unified index, and passed back to the vision model for image-grounded answers. A Recall@K / MRR evaluation harness over a 100-question gold set drove evidence-based tuning, reaching Recall@5 ~ 0.93 (reranked) and MRR ~ 0.74.

Nutrition Coach — Vision-Language Model for Calorie Estimation
PyTorch · Hugging Face Transformers · PEFT/LoRA · TRL · bitsandbytes · Qwen3-VL · Kaggle T4 GPU
GitHub

Fine-tuned Qwen3-VL-4B with QLoRA (4-bit, LoRA) to estimate calories and macros from food images, outputting structured JSON at 100% schema compliance on held-out data. A controlled experiment comparing two label sources cut calorie error to 81.4 kcal MAE (~34% lower than the base model) and reduced carbohydrate error by ~45% on a clean lab-measured holdout. Demonstrated that fine-tuning on synthetic class-average labels (Food101) degraded accuracy below the un-tuned baseline, while real measured data (Nutrition5k) produced genuine estimation — a data-quality finding with direct modeling implications.

Awards & Recognition
  • Kaizen (Operational Excellence) RecognitionMar 2026
    Selected at department level for innovation leading to significant cost and efficiency gains.

  • World Quality Week RecognitionNov 2025
    Presented the benefits, optimizations, and achievements of our solutions.

  • Customer Appreciation RecognitionSep 2025
    Recognized by the Mercedes-Benz client team.

  • Impact Award (Silver)Feb 2025, 2023
    Recognized for delivering high-impact ML solutions.

  • Team Excellence Award (Silver)Oct 2024
    Awarded for collaborative delivery of production-grade systems.

Template from here and here. Thank you :)

Last Updated: 15th June 2026