Performance Analysis of an Offline Text Detection System Based on Edge AI A Case Study of DokuScan Pro
DOI:
https://doi.org/10.47709/cnahpc.v8i1.7854Keywords:
offline text detection edge AI mobile document scanning on-device inference text detection evaluation mobile vision system resource-constrained devices offline mobile application document image analysis optical character recognition OCR preprocessing edge computing mobile AI inference time resource utilization performance evaluation integrated system real-world document data computer visionAbstract
The growing use of mobile document scanning applications has increased the demand for text detection systems that can operate reliably in offline and on-device environments. Although Edge AI enables local inference without network dependency, system-level empirical evidence regarding its performance under real-world mobile usage conditions remains limited. This study presents a system-level evaluation of an offline Edge AI–based text detection system for mobile document scanning, using DokuScan Pro as a case study. The evaluation was conducted on 40 document images captured under varying lighting conditions, capture angles, and background characteristics. System performance was assessed using precision, recall, F1-score, and inference time to characterize on-device behavior rather than algorithmic novelty. Experimental results show that the system achieved a precision of 1.00, a recall of 0.975, and an F1-score of approximately 0.98, with an average inference time of 63.8 ms per image during fully offline execution on mobile devices. These results indicate stable system-level performance under real-world document scanning conditions with controlled computational overhead. This study provides empirical system-level insights into the feasibility and practical limitations of deploying Edge AI–based text detection in offline mobile document scanning applications, thereby complementing existing model-centric research with evidence from real-world, on-device evaluation.
Downloads
References
Baek, Y., Lee, B., Han, D., Yun, S., & Lee, H. (2019). Character Region Awareness for Text Detection. http://arxiv.org/abs/1904.01941
Epshtein, B., Ofek, E., & Wexler, Y. (2010). Detecting text in natural scenes with stroke width transform. IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2963–2970. https://doi.org/10.1109/CVPR.2010.5540041
Fang, Z., & Yin, B. (2025). A low functional redundancy-based network slimming method for accelerating deep neural networks. Alexandria Engineering Journal, 119, 437–450. https://doi.org/10.1016/j.aej.2024.12.118
Francis, L. M., & Sreenath, N. (2020). TEDLESS – Text detection using least-square SVM from natural scene. Journal of King Saud University - Computer and Information Sciences, 32(3), 287–299. https://doi.org/10.1016/j.jksuci.2017.09.001
Gunawardena, N., Ginige, J. A., Javadi, B., & Lui, G. (2022). Performance Analysis of CNN Models for Mobile Device Eye Tracking with Edge Computing. Procedia Computer Science, 207, 2291–2300. https://doi.org/10.1016/j.procs.2022.09.288
Liao, M., Wan, Z., Yao, C., Chen, K., & Bai, X. (2020). Real-Time Scene Text Detection with Differentiable Binarization. Proceedings of the AAAI Conference on Artificial Intelligence, 34(07), 11474–11481. https://doi.org/10.1609/aaai.v34i07.6812
Lins, R. D. , C. G. D. C. , & S. J. D. (2021). Image enhancement techniques for mobile document scanning systems. Pattern Recognition Letters.
MohiEldeen Alabbasy, F., Abohamama, A. S., & Alrahmawy, M. F. (2023). Compressing medical deep neural network models for edge devices using knowledge distillation. Journal of King Saud University - Computer and Information Sciences, 35(7). https://doi.org/10.1016/j.jksuci.2023.101616
N, K., Choudhury, A. R., G S, R., G, B., & Sanchi, C. (2025). On-Device Deep Learning for Retrieving System and User Timestamps from Noisy Chat Images. Procedia Computer Science, 258, 3760–3770. https://doi.org/10.1016/j.procs.2025.04.631
Peng, Z., Li, J., Hao, H., & Zhong, Y. (2024). Smart structural health monitoring using computer vision and edge computing. Engineering Structures, 319. https://doi.org/10.1016/j.engstruct.2024.118809
Pineau, J., Vincent-Lamarre, P., Sinha, K., Larivì, V., Beygelzimer, A., D’alché-Buc, F., Paris, T., Larochelle, H., Florence D’alché-Buc, E., & Fox, H. L. (2021). Improving Reproducibility in Machine Learning Research (A Report from the NeurIPS 2019 Reproducibility Program). In Journal of Machine Learning Research (Vol. 22).
Rajan, R., & Devasena, M. S. G. (2025). Deep learning based optimization model for document layout and text recognition. Ain Shams Engineering Journal, 16(10), 103587. https://doi.org/10.1016/j.asej.2025.103587
Wang, P., Ma, Z., Dong, B., Liu, X., Ding, J., Sun, K., & Chen, Y. (2024). Generative data augmentation by conditional inpainting for multi-class object detection in infrared images. Pattern Recognition, 153. https://doi.org/10.1016/j.patcog.2024.110501
Y?ld?z, S. (2024). Turkish scene text recognition: Introducing extensive real and synthetic datasets and a novel recognition model. Engineering Science and Technology, an International Journal, 60, 101881. https://doi.org/10.1016/j.jestch.2024.101881
Zhou, X. , Y. C. , W. H. , W. Y. , Z. S. , H. W. , & L. J. (2017). EAST: An efficient and accurate scene text detector.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Defi Pujianto, Kadarsih, Sri Hartati

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.











