
Research Article
A Multi-Dimensional Evaluation Framework for Matrix-Based AI Accelerators: GPU, FPGA, and ASIC
@INPROCEEDINGS{10.4108/eai.24-4-2026.2364875, author={Kangzhe Peng}, title={A Multi-Dimensional Evaluation Framework for Matrix-Based AI Accelerators: GPU, FPGA, and ASIC}, proceedings={Proceedings of the 3rd International Conference on Mechanics, Electronics Engineering and Automation, ICMEEA 2026, April 24-26, 2026, Singapore, Singapore}, publisher={EAI}, proceedings_a={ICMEEA}, year={2026}, month={9}, keywords={AI Hardware Accelerators Matrix Multiplication (GEMM) Hardware Selection Framework Compute Gap Deep Learning}, doi={10.4108/eai.24-4-2026.2364875} }- Kangzhe Peng
Year: 2026
A Multi-Dimensional Evaluation Framework for Matrix-Based AI Accelerators: GPU, FPGA, and ASIC
ICMEEA
EAI
DOI: 10.4108/eai.24-4-2026.2364875
Abstract
The fast developing pace in deep learning is giving pressure on computing hardware continuously. In many practical cases, model size and training cost increase faster than the performance improvement of general-purpose processors. The huge different pace between them makes a gap called “compute gap”. Consequently, heterogeneous accelerators such as GPUs, FPGAs, and ASICs have become central to modern AI systems. It is not a easy work to select the right and appropriate accelerator because current studies focus on peak throughput while ignore the latency characteristics, energy efficiency and deployment constraints. This essay introduces a multi-dimensional evaluation framework for matrix-based AI accelerators which concentrated on convolution-dominated workload. Analyzing the represented hardware platform, including NVIDIA A100 GPUs, AMD Xilinx Versal AI Core FPGAs and Google TPU v4 ASICs, and comparing their throughput, determinism, scalability and power consumption. At the same time, in order to provide the guide in selection of hardware in different and detailed scenarios and distinguish compute-bound and memory-bound workloads, the Roofline Model is utilized. The analysis suggests that GPUs and TPUs are generally more suitable for large-scale cloud training, whereas FPGAs tend to be more advantageous in latency-sensitive and energy-constrained edge inference scenarios.

