TY - GEN
T1 - Algorithm and Hardware Co-Design Exploration for NTRU Equation Solving
AU - Liao, Jun Hao
AU - Shieh, Ming Der
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - With the rapid advancement of quantum computing technologies, traditional public-key algorithms face the threat of being broken. Falcon, now standardized for digital signatures in post-quantum cryptography, relies heavily on repeated FFT, NTT, RNS and CRT operations during its key-generation phase. Since each of these modules contributes similarly to the total computation, standalone optimization yields only marginal gains. This paper presents a dedicated hardware accelerator design that unifies arithmetic units at the microarchitectural level and explores efficient parallel structures to shorten the critical path and enhance throughput. We target the three core routines, Field Norm, Lift and Reduce, by tailoring dataflows and dynamic switching strategies to prime bit-width and algorithmic complexity. Additionally, we adopt a compact SRAM partitioning scheme to achieve conflict-free data arrangement and efficient coexistence of all intermediate values. Post-layout simulations of our design in a 40-nm process indicate a footprint of 1.987 mm2 at 333 MHz, achieving approximately a 6× speedup over a software reference and a 100× improvement compared to a high-level-synthesis based hardware prototype.
AB - With the rapid advancement of quantum computing technologies, traditional public-key algorithms face the threat of being broken. Falcon, now standardized for digital signatures in post-quantum cryptography, relies heavily on repeated FFT, NTT, RNS and CRT operations during its key-generation phase. Since each of these modules contributes similarly to the total computation, standalone optimization yields only marginal gains. This paper presents a dedicated hardware accelerator design that unifies arithmetic units at the microarchitectural level and explores efficient parallel structures to shorten the critical path and enhance throughput. We target the three core routines, Field Norm, Lift and Reduce, by tailoring dataflows and dynamic switching strategies to prime bit-width and algorithmic complexity. Additionally, we adopt a compact SRAM partitioning scheme to achieve conflict-free data arrangement and efficient coexistence of all intermediate values. Post-layout simulations of our design in a 40-nm process indicate a footprint of 1.987 mm2 at 333 MHz, achieving approximately a 6× speedup over a software reference and a 100× improvement compared to a high-level-synthesis based hardware prototype.
UR - https://www.scopus.com/pages/publications/105030477653
UR - https://www.scopus.com/pages/publications/105030477653#tab=citedBy
U2 - 10.1109/ICECS66544.2025.11270592
DO - 10.1109/ICECS66544.2025.11270592
M3 - Conference contribution
AN - SCOPUS:105030477653
T3 - 2025 32nd IEEE International Conference on Electronics, Circuits and Systems, ICECS 2025
BT - 2025 32nd IEEE International Conference on Electronics, Circuits and Systems, ICECS 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 32nd IEEE International Conference on Electronics, Circuits and Systems, ICECS 2025
Y2 - 17 November 2025 through 19 November 2025
ER -