Abstract
One of the most crucial tasks of computer vision systems in autonomous navigation, robotics, obstacle avoidance, and real-time situational awareness is depth estimation. However, using depth estimation algorithms on embedded platforms can be limited by computational complexity, hardware resource usage, latency, and power consumption. This study provides a comparative analysis of relative and absolute depth estimation techniques from the perspective of hardware efficiency for embedded vision systems. The proposed relative depth estimation framework is based on a lightweight run-based connected component labeling (CCL) architecture to identify regions (objects) and infer ordinal depth relationships using area and centroid features. In contrast, the absolute depth estimation pipeline combines object localization with disparity-based metric depth estimation using camera geometry. Both approaches are tested on KITTI, NYU V2 and self-generated datasets, under FPGA and System-on-Chip (SoC) deployment environments. The experimental results show that the relative depth estimation framework consumes significantly fewer computational resources, with 3460 LUTs, 4 BRAMs, a power consumption of 0.78 W, and a processing latency of 0.30 ms/frame. The absolute depth estimation pipeline, on the other hand, consumes more hardware resources, such as approximately 18k LUTs, 28 BRAMs, 1.8 W power, and 3.80 ms/frame latency from the disparity estimation and CNN-assisted processing pipeline. The comparative analysis shows that relative depth estimation offers a better latency–power–hardware trade-off for resource constrained embedded platforms, while absolute depth estimation has better metric depth accuracy at the cost of higher computational complexity. The proposed framework demonstrates how lightweight relative depth inference can be used in real-time embedded perception applications.References
- Eigen, David, Christian Puhrsch, and Rob Fergus. ”Depth Map Prediction from a Single Image Using a Multi-Scale Deep Network.” Advances in neural information processing systems 27 (2014).
- Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision Meets Robotics: The KITTI Dataset,” The International Journal of Robotics Research (IJRR), vol. 32, no. 11, 2013, 1231–1237.
- Endsley, Mica R. ”Design and Evaluation for Situation Awareness Enhancement.” In Proceedings of the Human Factors Society annual meeting, vol. 32, no. 2, Sage CA: Los Angeles, CA: Sage Publications, 1988, 97-101.
- Godard, O. Mac Aodha, and G. J. Brostow, “Unsupervised Monocular Depth Estimation with Left-Right Consistency,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
- Ranftl, René, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. ”Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer.” IEEE transactions on pattern analysis and machine intelligence 44, no. 3. 2020: 1623-1637.
- Wu, Haifeng, Shuhang Gu, Lixin Duan, and Wen Li. ”Geodepth: From Point-to-Depth to Plane-to-Depth Modeling for Self-Supervised Monocular Depth Estimation.” In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2025, 11525-11535.
- Liu, Fayao, Chunhua Shen, and Guosheng Lin. ”Deep Convolutional Neural Fields for Depth Estimation from a Single Image.” In Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, 5162-5170.
- Godard, Clément, Oisin Mac Aodha, Michael Firman, and Gabriel J. Brostow. ”Digging into self-supervised monocular depth estimation.” In Proceedings of the IEEE/CVF international conference on computer vision, 2019, 3828-3838.
- Bhat, Shariq Farooq, Ibraheem Alhashim, and Peter Wonka. ”Adabins: Depth estimation using adaptive bins.” In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, 4009-4018.
- Zoran, D., Isola, P., Krishnan, D. and Freeman, W.T., 2015. Learning Ordinal Relationships for Mid-Level Vision. In Proceedings of the IEEE international conference on computer vision (388-396).
- Fu, Huan, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao. ”Deep Ordinal Regression Network for Monocular Depth Estimation.” In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE, 2018, 2002-2011.
- Zhou, Hang, David Greenwood, and Sarah Taylor. ”Self-Supervised Monocular Depth Estimation with Internal Feature Fusion.” arXiv preprint arXiv:2110.09482 (2021).
- Dosovitskiy, Alexey, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani et al. ”An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.” arXiv preprint arXiv:2010.11929 (2020).
- Ranftl, René, Alexey Bochkovskiy, and Vladlen Koltun. ”Vision Transformers for Dense Prediction.” In Proceedings of the IEEE/CVF international conference on computer vision, 2021,12179-12188.
- Yang, Lihe, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. ”Depth Anything v2.” Advances in Neural Information Processing Systems 37 (2024): 21875-21911.
- Wofk, Diana, Fangchang Ma, Tien-Ju Yang, Sertac Karaman, and Vivienne Sze. ”Fastdepth: Fast Monocular Depth Estimation on Embedded Systems.” In 2019 International Conference on Robotics and Automation (ICRA), IEEE, 2019, 6101-6108.
- Zhang, Ning, Francesco Nex, George Vosselman, and Norman Kerle. ”Lite-mono: A lightweight CNN and Transformer Architecture for Self-Supervised Monocular Depth Estimation.” In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, 18537-18546.
- He, Lifeng, Yuyan Chao, and Kenji Suzuki. ”A Run-Based Two-Scan Labeling Algorithm.” IEEE transactions on image processing 17, no. 5 (2008): 749-756.
- Di Stefano, Luigi, and Andrea Bulgarelli. ”A Simple and Efficient Connected Components Labeling Algorithm.” In Proceedings 10th international conference on image analysis and processing, IEEE, 1999, 322-327.
- Grana, Costantino, Daniele Borghesani, and Rita Cucchiara. ”Optimized Block-Based Connected Components Labeled with Decision Trees.” IEEE Transactions on Image Processing 19, no. 6 (2010): 1596-1609.
- Kiran, Divya, Abdul Imran Rasheed, and Hariharan Ramasangu. 2023. ”Method, System and Apparatus for Object Detection.” Indian Patent 466134 filed November 24, 2014 (Application No. 3573/CHE/2014) and issued November 6, 2023.
- Kiran, Divya, Abdul Imran Rasheed, and Hariharan Ramasangu. ”FPGA Implementation Of Blob Detection Algorithm for Object Detection in Visual Navigation.” In 2013 International conference on Circuits, Controls and Communications (CCUBE), IEEE, 2013, 1-5.
- He, Kaiming, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. ”Mask r-cnn.” In Proceedings of the IEEE international conference on computer vision, 2017, 2961-2969.
- Lavreniuk, Mykola, and Alla Lavreniuk. ”Spidepth: Strengthened Pose Information for Self-Supervised Monocular Depth Estimation.” In Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, 874-884.
- ul Islam, Nida, Divya Kiran, and Hariharan Ramasangu. ”System-on-Chip Implementation of Situational Awareness System for Autonomous Cars.” In 2018 15th IEEE India Council International Conference (INDICON), IEEE, 2018, 1-6.
- Suda, Naveen, Vikas Chandra, Ganesh Dasika, Abinash Mohanty, Yufei Ma, Sarma Vrudhula, Jae-sun Seo, and Yu Cao. ”Throughput-optimized OpenCL-based FPGA Accelerator for Large-Scale Convolutional Neural Networks.” In Proceedings of the 2016 ACM/SIGDA international symposium on field-programmable gate arrays, 2016, 16-25.
- Wang, Han, Xiaoyue Liu, Yi Yang, Qin Liao, Shiyu Ke, and Yong Dong. ”FPGA-Based Acceleration of Deep Learning Networks: Techniques and Applications.” In Proceedings of the 2025 6th International Conference on Computer Information and Big Data Applications, 2025, 774-778.
- Obukhov, Anton, Matteo Poggi, Fabio Tosi, Ripudaman Singh Arora, Jaime Spencer, Chris Russel, Simon Hadfield et al. ”The Fourth Monocular Depth Estimation Challenge.” In Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, 6182-6195.
- Benkrid, Khaled, S. Sukhsawas, Danny Crookes, and Abdsamad Benkrid. ”An FPGA-based Image Connected Component Labeller.” In International Conference on Field Programmable Logic and Applications, Berlin, Heidelberg: Springer Berlin Heidelberg, 2003, 1012-1015.
- Spagnolo, Fanny, Fabio Frustaci, Stefania Perri, and Pasquale Corsonello. ”An Efficient Connected Component Labeling Architecture for Embedded Systems.” Journal of Low Power Electronics and Applications 8, no. 1 (2018): 7.
- Perri, Stefania, Fanny Spagnolo, and Pasquale Corsonello. ”A Parallel Connected Component Labeling Architecture for Heterogeneous Systems-on-Chip.” Electronics 9, no. 2 (2020): 292.
- Sada, Youki, Naoto Soga, Masayuki Shimoda, Akira Jinguji, Shimpei Sato, and Hiroki Nakahara. ”Fast Monocular Depth Estimation on an FPGA.” In 2020 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), IEEE, 2020, 143-146.
- Hashimoto, Nobuho, and Shinya Takamaeda-Yamazaki. ”FADEC: FPGA-based Acceleration of Video Depth Estimation by HW/SW Co-design.” In 2022 International Conference on Field-Programmable Technology (ICFPT), IEEE, 2022, 1-9.
- Li, Jie, Chuanlun Zhang, Heng Li, Shuangli Du, Wenxuan Yang, Xiaoyan Wang, and Yiguang Liu. ”RE-LFDE: A Resource-Efficient Hardware Accelerator for Low-Bit Light Field Image Depth Estimation.” ACM Transactions on Embedded Computing Systems 25, no. 3 (2026): 1-20.

Journal of Innovative Image Processing