Abstract
Deploying language models on resource-constrained mobile/wearable devices while maintaining output quality is challenging. To address such challenges, many floating-point (FP) and integer (INT) quantization methods have been explored. FP arithmetic provides high quality at the cost of heavy area and energy costs, while INT-quantized models deliver superior efficiency at the cost of accuracy/ perplexity loss. In addition to the design choices between FP and INT, we explore an alternative option based on a logarithmic number system (LNS), which delivers FP-like accuracy/perplexity at an efficiency close to INT. In addition to applying low-precision (8-bit) LNS, we adaptively assign bits for the INT and the fraction depending on data distribution, which enables near-FP16 accuracy/perplexity. We also co-design the LNS arithmetic and accelerator architecture, which leads to 33% less energy than the FP8 (E4M3) accelerator with similar area as an INT8 accelerator, while delivering 30% lower perplexity compared to FP8 (E4M3).
| Original language | English |
|---|---|
| Pages (from-to) | 43-56 |
| Number of pages | 14 |
| Journal | IEEE Micro |
| Volume | 46 |
| Issue number | 3 |
| DOIs | |
| Publication status | Published - 2026 May 1 |
Bibliographical note
Publisher Copyright:© 2026 IEEE. All rights reserved.
ASJC Scopus subject areas
- Software
- Hardware and Architecture
- Electrical and Electronic Engineering
Fingerprint
Dive into the research topics of 'LogFlex: Flexible-Bit Log Arithmetic Accelerator for Language Models on Edge'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS