Zhang, Huaizhi and Al-Hasan, Tamim M and Zhu, Xuqi and Zhu, Jiacheng and Si, Weiyong and McDonald-Maier, Klaus D and Zhai, Xiaojun (2025) Computationally Efficient FPGA-based Large Language Model Inference for Real-Time Decision-Making in Robotic Systems. In: 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025-10-19 - 2025-10-25, Hangzhou, China.
Zhang, Huaizhi and Al-Hasan, Tamim M and Zhu, Xuqi and Zhu, Jiacheng and Si, Weiyong and McDonald-Maier, Klaus D and Zhai, Xiaojun (2025) Computationally Efficient FPGA-based Large Language Model Inference for Real-Time Decision-Making in Robotic Systems. In: 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025-10-19 - 2025-10-25, Hangzhou, China.
Zhang, Huaizhi and Al-Hasan, Tamim M and Zhu, Xuqi and Zhu, Jiacheng and Si, Weiyong and McDonald-Maier, Klaus D and Zhai, Xiaojun (2025) Computationally Efficient FPGA-based Large Language Model Inference for Real-Time Decision-Making in Robotic Systems. In: 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2025-10-19 - 2025-10-25, Hangzhou, China.
Abstract
Integrating Large Language Models (LLMs) into modern robotic systems presents significant computational and energy constraint challenges, particularly for human-centered robotic applications. This paper presents a novel hardware optimization technique for deploying LLMs on resource-constrained embedded devices, achieving an up to 77% reduction in computational latency through an FPGA implementation in comparison to other popular embedded computing devices (e.g., CPU and GPUs). Additionally, we demonstrate our methodology by deploying a LLaMA 2-7B model on a Unitree Go2 robotic dog integrated with the proposed FPGA platform. The proposed optimization framework preserves real-time interaction capabilities while significantly reducing computational and energy overhead, facilitating efficient natural language processing for human-robot interaction in safety-critical and dynamic environments. Experimental results demonstrate that the FPGA-based LLaMA 2-7B implementation achieves up to 6.06-fold and 1.95-fold higher throughput compared to baseline CPU and GPU implementations while maintaining comparable inference accuracy. Furthermore, the proposed FPGA design surpasses existing state-of-the-art FPGA implementations, delivering a 30% improvement in computational efficiency.
| Item Type: | Conference or Workshop Item (Paper) |
|---|---|
| Uncontrolled Keywords: | Large language models, Human-robot interaction, Graphics processing units, Dogs, Throughput, Real-time systems, Computational efficiency, Robots, Field programmable gate arrays, Optimization |
| Subjects: | Z Bibliography. Library Science. Information Resources > ZR Rights Retention |
| Divisions: | Faculty of Science and Health Faculty of Science and Health > Computer Science and Electronic Engineering, School of |
| SWORD Depositor: | Unnamed user with email elements@essex.ac.uk |
| Depositing User: | Unnamed user with email elements@essex.ac.uk |
| Date Deposited: | 02 Sep 2026 10:44 |
| Last Modified: | 02 Sep 2026 10:47 |
| URI: | http://repository.essex.ac.uk/id/eprint/43780 |
Available files
Filename: IROS (1).pdf
Licence: Creative Commons: Attribution 4.0