About | Contact Us | Register | Login
ProceedingsSeriesJournalsSearchEAI
airo 25(1):

Research Article

Hardware-Aware INT8 Quantization and FPGA Deployment of MobileNetV2 for Real-Time Facial Landmark Detection

Download2 downloads
Cite
BibTeX Plain Text
  • @ARTICLE{10.4108/airo.10070,
        author={Van-Khoa Pham and Manh-Dung Do and Trung-Nghia Dang and Long Tran},
        title={Hardware-Aware INT8 Quantization and FPGA Deployment of MobileNetV2 for Real-Time Facial Landmark Detection},
        journal={EAI Endorsed Transactions on AI and Robotics},
        volume={5},
        number={1},
        publisher={EAI},
        journal_a={AIRO},
        year={2026},
        month={3},
        keywords={facial landmark detection, MobileNetV2, quantization-aware training, post-training quantization, edge computing, AMD/Xilinx Kria KV260},
        doi={10.4108/airo.10070}
    }
    
  • Van-Khoa Pham
    Manh-Dung Do
    Trung-Nghia Dang
    Long Tran
    Year: 2026
    Hardware-Aware INT8 Quantization and FPGA Deployment of MobileNetV2 for Real-Time Facial Landmark Detection
    AIRO
    EAI
    DOI: 10.4108/airo.10070
Van-Khoa Pham1,*, Manh-Dung Do1, Trung-Nghia Dang1, Long Tran1
  • 1: Ho Chi Minh City University of Technology and Engineering
*Contact email: khoapv@hcmute.edu.vn

Abstract

Facial landmark detection is a key component of always-on edge vision systems, but practical deployment requires balancing localization accuracy, model size, throughput, and power consumption. This study proposes a two-stage, hardware-aware framework for MobileNetV2-based facial landmark detection. In Stage I, a lightweight detector is developed in PyTorch and evaluated in FP32 and INT8 using post-training quantization (PTQ) and quantization-aware training (QAT). In Stage II, the quantized model is realized on the AMD/Xilinx Kria KV260 FPGA and assessed in terms of real-time throughput, power consumption, and hardware resource utilization. INT8 quantization reduces the model size from 6.59 MB to 1.65 MB, and QAT retains accuracy more effectively than PTQ (91.74% vs. 90.94%) relative to the FP32 baseline (92.42%). The hardware implementation achieves approximately 30 FPS at approximately 3 W while using 14.8% of LUTs, 8.1% of FFs, 16.3% of BRAM, and 4.5% of DSPs. Among the evaluated platforms, the KV260 delivers the highest measured energy efficiency, whereas the RTX 4060 delivers the highest throughput. Within the evaluated setup, these results support the practicality of explicitly separating software-stage quantization analysis from hardware-stage realization for real-time, low-power facial landmark detection on FPGA-based edge platforms.  

Keywords
facial landmark detection, MobileNetV2, quantization-aware training, post-training quantization, edge computing, AMD/Xilinx Kria KV260
Received
2025-08-25
Accepted
2026-03-20
Published
2026-03-25
Publisher
EAI
http://dx.doi.org/10.4108/airo.10070

Copyright © 2026 Van-Khoa Pham et al., licensed to EAI. This is an open access article distributed under the terms of the CC BY-NC-SA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.

EBSCOProQuestDBLPDOAJPortico
EAI Logo

About EAI

  • Who We Are
  • Leadership
  • Research Areas
  • Partners
  • Media Center
  • Cookie Preferences

Community

  • Membership
  • Conference
  • Recognition
  • Sponsor Us

Publish with EAI

  • Publishing
  • Journals
  • Proceedings
  • Books
  • EUDL