NPUWattch: ML-based Power, Area, and Timing Modeling for Neural Accelerators
Pre-silicon modeling tools for characterizing power, area, and timing (PAT) have enabled numerous architectural studies, but traditional analytical and table-based models begin to exhibit limitations in their applicability as architectural design complexity increases and process technology scales below 5nm with the emergence of advanced transistors. Previous modeling techniques typically assumed static scaling factors across different designs and technology nodes, derived from small circuit design benchmarks using old processes. Consequently, they do not reflect complex design variability and nonlinear projection to advanced technology nodes. Moreover, reference logic and SRAM implementations serving as the baseline for design and technology scaling were often created using different technologies and design rules, leading to significant estimation inaccuracies that distort the relative contributions of individual components. To address these challenges, this paper introduces NPUWattch, a machine learning-based PAT modeling framework for neural accelerators. It leverages neural network regression models to learn complex nonlinear relationships in technology and design scaling based on diverse post-layout logic and SRAM design datasets formulated using unified technology libraries. To this end, we developed technology libraries from 65nm to 2nm, constructed and validated diverse logic and SRAM datasets, and trained neural network models using an adaptive loss function to reinforce underrepresented regions of the design space. NPUWattch is validated against the post-layout results of numerous open-source neural accelerators, and evaluation results demonstrate that NPUWattch outperforms existing tools with an average estimation error of 2.7%, offering reliable and accurate PAT estimation.
Tue 3 FebDisplayed time zone: Hobart change
15:50 - 17:10 | |||
15:50 20mTalk | NPUWattch: ML-based Power, Area, and Timing Modeling for Neural Accelerators Main Conference Sehyeon Kim Yonsei University, Minkwan Kim Yonsei University, Chanho Park Yonsei University, Hanmok Park Kyungpook National University, Seonghoon Kim Kyungpook National University, Taigon Song Kyungpook National University, William Song Yonsei University | ||
16:10 20mTalk | Area Bloating and the Future of Specialization Main Conference | ||
16:30 20mTalk | Advancing Full-stack Acceleration for Schrödinger-Style Quantum Simulation Main Conference Shuang Liang Imperial College London, Yuncheng Lu Imperial College London, Ce Guo Imperial College London, Paul H J Kelly Imperial College London, Wayne Luk Imperial College London, Hongxiang Fan Imperial College London | ||
16:50 20mTalk | COMET: Communication and Memory Co-Design for Fine-Grained AI Inference in MCM Accelerators Main Conference Taishu Sheng College of Computer Science and Technology, National University of Defense Technology, Guangyu Sun Peking University, Dezun Dong NUDT | ||