|
Deepak Kumar Singh
|
|
I'm a Computer Vision and Deep Learning Engineer at Mercedes-Benz Research and Development India, Bengaluru, part of the Interior-Sensing (IS) Algorithm team. I build production-grade computer vision systems for in-cabin automotive applications — including face recognition, depth estimation, and IR-based occupant analysis — owning the pipeline from proof-of-concept to deployment.
I completed my MS by Research at CVIT, IIIT Hyderabad, advised by Prof. C.V. Jawahar and co-advised by Prof. Vineeth N Balasubramanian, Prof. Chetan Arora, and Dr. Anbumani Subramanian. My current areas of interest are Face Recognition, Vision-Language Models, Contrastive Learning, Knowledge Distillation, and Open-World / Incremental Learning.
I like photography, majorly urban photography and landscapes; check out my Flickr. I've recently started learning Bansuri(Indian Flute).
To know what I'm currently reading, checkout my Goodreads.
Email: deepaksinghcv@gmail.com
|   
CV  / 
LinkedIn  / 
Github   
 | 
|
|
|
Research
My current work spans in-cabin face recognition, cross-domain (IR↔RGB) learning, and occupant analysis for automotive systems. Prior research focused on Open-World Object Detection and benchmarking detection/segmentation models on road-scene datasets. I'm broadly interested in translating research into scalable, deployable systems. Authors marked with a * have an equal contribution in the research.
|
News
- [Mar 2026] : 🎉 Received Kaizen (Operational Excellence) Recognition at the department level for innovation leading to significant cost and efficiency gains.
- [Nov 2025] : Presented benefits, optimizations, and achievements of our solutions at World Quality Week.
- [Sep 2025] : Received Customer Appreciation Recognition from the Mercedes-Benz client team.
- [Feb 2025] : Received the Impact Award (Silver) for delivering high-impact ML solutions.
- [Oct 2024] : Received the Team Excellence Award (Silver) for collaborative delivery of production-grade systems.
- [2023] : Received the Impact Award (Silver).
- [Oct 2022] : Presented our work, "New Vehicles on the Road? No Problem - We'll Learn Them Too" at IROS 2022, Kyoto, Japan.
- [Apr 2022] : Joined Mercedes-Benz Research and Development India as a Computer Vision and Deep Learning Engineer.
- [Dec 2021] : Completed MS by Research at CVIT, IIIT Hyderabad.
- [Dec 2021] : Presented our work, "ORDER: Open World Object Detection on Road Scenes" at NeurIPS Workshop on Machine Learning for Autonomous Driving 2021.
- [Dec 2021] : Presented our work, "Evaluation of Detection and Segmentation Tasks on Driving Datasets" at CVIP 2021.
- [Oct 2021] : 🎉 One paper accepted in NeurIPS Workshop on Machine Learning for Autonomous Driving 2021.
- [Oct 2021] : 🎉 One paper accepted in CVIP 2021 for Oral presentation.
- [Aug 2021] : In the organizing team for IIIT-H's annual Summer School Program
- [Aug 2020] : Joined Vision for Mobility and Safety Team in CVIT.
- [Nov 2019] : Attended NCVPRIPG at Hubbali, Karnataka.
- [Jan 2019] : Joined CVIT as a Research Fellow.
- [Dec 2018] : Got the opportunity to attend ICVGIP 2018 and talk to various speakers and fellow researchers.
- [Aug 2018] : Joined IIIT-H as an MS by Research student.
|
Computer Vision and Deep Learning Engineer, Mercedes-Benz Research and Development India — Apr 2022 – Present
Bengaluru, India
Member of the Interior-Sensing (IS) Algorithm team, building computer vision solutions for in-cabin automotive systems across face recognition, depth estimation, and weight prediction — owning the pipeline from proof-of-concept to production.
- Led end-to-end development of production-grade in-cabin face recognition systems, from data strategy and model development to deployment, achieving stringent accuracy and robustness targets.
- Architected a dual-level authentication framework enabling both low-latency personalization and high-assurance security for privacy- and payment-sensitive use cases.
- Developed cross-domain face recognition and IR-to-RGB colorization solutions to bridge RGB/NIR modality gaps and reduce dependency on large-scale RGB data collection.
- Built deep learning models for occupant analysis, including IR-based weight estimation, replacing hardware sensors with software-driven solutions.
- Collaborated across software, quantization, validation, and platform teams; drove GCP-based ML infrastructure initiatives and contributed to patentable innovations.
|
Software Engineer, Celstream Systems Pvt. Ltd. — Sep 2014 – Sep 2016
Bangalore, India
- Built the product's main JavaScript UI console and dynamic data-visualization modules (IgniteUI), and led migration of the in-house application from Adobe Flash to a JavaScript environment.
- Developed REST APIs in Java and data adapters powering live data-visualization modules.
|
|
|
New Vehicles on the Road? No Problem - We'll Learn Them Too
Deepak Kumar Singh,
Shyam Nandan Rai,
K J Joseph,
Rohit Saluja,
Vineeth N Balasubramanian,
Chetan Arora,
Anbumani Subramanian,
C.V. Jawahar
IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2022, Kyoto, Japan
Paper
Extends our open-world object detection work to incrementally adapt to new vehicle categories without catastrophic forgetting, improving robustness for real-world autonomous driving deployments.
|
|
|
ORDER: Open World Object Detection on Road Scenes
Deepak Kumar Singh*,
Shyam Nandan Rai*,
K J Joseph,
Rohit Saluja,
Vineeth N Balasubramanian,
Chetan Arora,
Anbumani Subramanian,
C.V. Jawahar
NeurIPS 2021 Workshop on Machine Learning for Autonomous Driving
Poster /
Paper /
Video(Soon) /
This work formulates Open-World Object Detection by addressing the inherent issues present in road scene datasets India Driving Dataset(IDD) and Berkeley Deep Drive(BDD).
|
|
|
Evaluation of Detection and Segmentation Tasks on Driving Datasets
Deepak Kumar Singh,
Ameet Rahane,
Ajoy Mondal,
Anbumani Subramanian,
C.V. Jawahar
CVIP, 2021 (Oral)
Paper /
Project Page(Soon) /
Slides(Soon) /
Video(Soon) /
Benchmark latest models of object detection, semantic segmentation, and instance segmentation on road scene datasets. We perform a detailed study on their behaviour on both constrained and unconstrained datasets; Cityscapes, India Driving Dataset(IDD) and Berkeley Deep Drive(BDD).
|
-
In-Cabin Face Authentication System — Filed in Germany and India
Implemented Patent Award · Inventor lead. Enables secure multi-level authentication in automotive systems.
-
Intelligent In-Vehicle Hyper-Personalization — Filed in Germany and India
Co-Inventor. Enhances user experience by assimilating signals, memory, and preferences.
|
|
|
DocLens — Local Multimodal RAG for Technical Documents
Python · sentence-transformers · ChromaDB · Ollama (Qwen3 / Qwen-VL) · PyMuPDF · RAG · cross-encoder reranking
GitHub
Built a fully-local multimodal RAG assistant for technical PDFs with two-stage retrieval (bi-encoder recall + cross-encoder reranking) and page-level citations, running entirely on-device with no external APIs. Figures are extracted, captioned by a local VLM, embedded into a unified index, and passed back to the vision model for image-grounded answers. A Recall@K / MRR evaluation harness over a 100-question gold set drove evidence-based tuning, reaching Recall@5 ~ 0.93 (reranked) and MRR ~ 0.74.
|
|
|
Nutrition Coach — Vision-Language Model for Calorie Estimation
PyTorch · Hugging Face Transformers · PEFT/LoRA · TRL · bitsandbytes · Qwen3-VL · Kaggle T4 GPU
GitHub
Fine-tuned Qwen3-VL-4B with QLoRA (4-bit, LoRA) to estimate calories and macros from food images, outputting structured JSON at 100% schema compliance on held-out data. A controlled experiment comparing two label sources cut calorie error to 81.4 kcal MAE (~34% lower than the base model) and reduced carbohydrate error by ~45% on a clean lab-measured holdout. Demonstrated that fine-tuning on synthetic class-average labels (Food101) degraded accuracy below the un-tuned baseline, while real measured data (Nutrition5k) produced genuine estimation — a data-quality finding with direct modeling implications.
|
- Kaizen (Operational Excellence) Recognition — Mar 2026
Selected at department level for innovation leading to significant cost and efficiency gains.
- World Quality Week Recognition — Nov 2025
Presented the benefits, optimizations, and achievements of our solutions.
- Customer Appreciation Recognition — Sep 2025
Recognized by the Mercedes-Benz client team.
- Impact Award (Silver) — Feb 2025, 2023
Recognized for delivering high-impact ML solutions.
- Team Excellence Award (Silver) — Oct 2024
Awarded for collaborative delivery of production-grade systems.
|
Template from here and here. Thank you :)
|
|
Last Updated: 15th June 2026
|
|