TY - GEN
T1 - A Reconfigurable IEEE-754 FP32 Fused Multiply Accumulate Architecture with Hybrid Multiplier and Prefix Adder Optimization
AU - Krishnananda, K. R.
AU - Nayak, Subramanya G.
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - The proposed work presents the design, verification, and full RTL-to-GDSII implementation of a single-precision fused multiply-accumulate (FMA) unit that can be changed or reconfigured and follows the IEEE-754 standard. The unit is defined by y = a × b + c. The design incorporates selectable multiplier and adder architectures, such as Booth + CSA, Array, and Vedic hybrid multipliers with Kogge-Stone and Han-Carlson adders. It also has an optional approximate arithmetic mode to minimize dynamic power. The proposed five-stage pipelined datapath delivers high throughput with a balanced latency profile while keeping IEEE-754 rounding accuracy and special-case handling. Using Cadence tools, we carried out functional and physical verification, which showed that the power and area were being used efficiently. Using Cadence Genus-Innovus flow, the design was tested in a 90 nm CMOS technology node. It worked at a maximum frequency of 175 MHz with positive slack and a total power of 33.55 mW in a layout footprint of 48,197 μm2. A comparative evaluation shows that the reconfigurable architecture has a better balance of performance, power, and flexibility than traditional fixed datapaths. The proposed work addresses Sustainable Developmental Goals 4 and 9.
AB - The proposed work presents the design, verification, and full RTL-to-GDSII implementation of a single-precision fused multiply-accumulate (FMA) unit that can be changed or reconfigured and follows the IEEE-754 standard. The unit is defined by y = a × b + c. The design incorporates selectable multiplier and adder architectures, such as Booth + CSA, Array, and Vedic hybrid multipliers with Kogge-Stone and Han-Carlson adders. It also has an optional approximate arithmetic mode to minimize dynamic power. The proposed five-stage pipelined datapath delivers high throughput with a balanced latency profile while keeping IEEE-754 rounding accuracy and special-case handling. Using Cadence tools, we carried out functional and physical verification, which showed that the power and area were being used efficiently. Using Cadence Genus-Innovus flow, the design was tested in a 90 nm CMOS technology node. It worked at a maximum frequency of 175 MHz with positive slack and a total power of 33.55 mW in a layout footprint of 48,197 μm2. A comparative evaluation shows that the reconfigurable architecture has a better balance of performance, power, and flexibility than traditional fixed datapaths. The proposed work addresses Sustainable Developmental Goals 4 and 9.
UR - https://www.scopus.com/pages/publications/105041612193
UR - https://www.scopus.com/pages/publications/105041612193#tab=citedBy
U2 - 10.1109/ICSFT66733.2026.11506967
DO - 10.1109/ICSFT66733.2026.11506967
M3 - Conference contribution
AN - SCOPUS:105041612193
T3 - 2026 International Conference on Smart Futuristic Technology, ICSFT 2026
BT - 2026 International Conference on Smart Futuristic Technology, ICSFT 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 International Conference on Smart Futuristic Technology, ICSFT 2026
Y2 - 2 January 2026 through 3 January 2026
ER -