High-speed and space-efficient arithmetic units are necessary for the hardware implementation of cryptographic algorithms to guarantee security and performance. Efficient multi-operand addition, especially three-operand binary addition, is crucial for modular operations like multiplication and exponentiation. A high-speed, low-area three-operand binary adder for cryptography and pseudorandom bit generator (PRBG) applications is presented in this study. Kogge-Stone, Han-Carlson, and Ladner-Fischer are examples of parallel prefix adders that are used to increase throughput and decrease propagation latency. For VLSI systems, heterogeneous delay-insensitive coding is used to further optimize power, area, and performance. A Carry Look-Ahead (CLA) adder, which lowers critical route delay through parallel carry generation, is added to the design to improve carry calculation. When compared to traditional designs, the suggested hybrid architecture provides increased speed and efficiency. Its usefulness for high-performance computing and digital signal processing applications is demonstrated via implementation using Xilinx Vivado.
Introduction
This study focuses on the design of high-speed, area-efficient, and low-power multi-operand adders (MOAs) for VLSI systems, particularly for applications such as multipliers, digital signal processors (DSPs), cryptography, and medical IoT devices. Since arithmetic units—especially multipliers—consume a significant portion of system power and area, optimizing multi-operand addition is essential for improving overall performance.
Background
Multi-operand adders are used to efficiently sum multiple partial products generated during multiplication. Different adder architectures offer various trade-offs:
Ripple Carry Adder (RCA): Small area and low power but high delay.
Carry Look-Ahead Adder (CLA): Faster operation but larger area and power consumption.
Carry Select Adder (CSLA): Balances speed and hardware complexity.
Carry Save Adder (CSA): Efficient for multi-operand addition by avoiding immediate carry propagation.
Parallel Prefix Adders (PPA) (e.g., Kogge-Stone, Han-Carlson): Very fast but require more hardware.
The main challenge is reducing carry propagation delay while maintaining low area and power consumption.
Literature Review
Previous studies have explored:
Wallace Tree Adders and compressor trees for reducing addition delay.
Optimized CSLA designs using Binary-to-Excess-1 Converter (BEC) and Square Root CSLA to reduce area and power.
Reversible logic for minimizing energy dissipation in low-power VLSI and quantum computing.
Parallel prefix architectures for achieving logarithmic delay.
However, existing approaches often sacrifice one design metric (speed, area, or power) to improve another.
Existing Limitations
Current adder architectures face several trade-offs:
RCA is compact but slow.
CLA and PPAs are fast but consume more hardware resources.
Wallace Tree Adders offer low delay but have irregular structures that complicate implementation.
Carry Save Adders still require a final carry-propagation stage, limiting overall speed.
An efficient three-operand adder that simultaneously minimizes delay, area, and power remains a key research need.
Proposed Methodology
The proposed design introduces a Variable Stage Structure (VSS) combined with a hybrid Parallel Prefix Adder (PPA) architecture.
Key features include:
Hybrid design replacing the most delay-critical stage with a Ling–Kogge-Stone Parallel Prefix Adder.
Three-stage PPA implementation:
Preprocessing
Prefix carry computation
Post-processing (sum generation)
Faster carry computation using Ling equations and Kogge-Stone architecture.
Reduced logic complexity and lower critical path delay.
Carry Select Adder Optimization
To further improve efficiency, the design incorporates:
Binary-to-Excess-1 Converter (BEC)-based CSLA, replacing duplicate RCAs to reduce hardware and power.
Common Boolean Logic (CBL)-based Square Root CSLA, which shares logic between carry-in cases to minimize transistor count and improve the power-delay product.
Reliability Enhancement
The proposed architecture also introduces concurrent error detection, featuring:
Predicted parity checking.
Dual carry outputs.
Modified hardware blocks for complete fault detection.
100% fault coverage with only 10 test patterns, regardless of adder bit width.
Improved testability (C-testability) suitable for safety-critical VLSI applications.
Conclusion
In this paper, a high-speed, area-efficient adder technique and its VLSI architecture are proposed to perform the three-operand binary addition for efficient computation of modular arithmetic used in cryptography and PRBG applications. The proposed three-operand adder technique is a parallel prefix adder that uses four-stage structures to compute the addition of three input operands. The novelty of this proposed architecture is the reduction of delay and area in the prefix computation stages in PG logic and bit-addition logic, which leads to an overall reduction in area and delay. For a fair comparison, the concept of a hybrid Han-Carlson two-operand adder is extended to develop a hybrid Han-Carlson three-operand adder (HHC3A) topology. The same coding style adopted in the proposed adder architecture is extended to implement the hybrid Han-Carlson three-operand adder, Kogge Stone, using Verilog HDL. Further, all these designs are synthesized using the commercially available Spartan 3 technology library to obtain the core area and delay timing. Application-wise, we implemented our adder in the adder part of an FIR filter for further efficiency in the field of area and delay. Concluding that our adder was comparatively better than other adders by considering both the area and delay.
References
[1] M. M. Islam, M. S. Hossain, M. K. Hasan, M. Shahjalal, and Y. M. Jang, “FPGA implementation of high-speed area-efficient processor for elliptic curve point multiplication over prime field,” IEEE Access, vol. 7, pp. 178811–178826, 2019.
[2] Z. Liu, J. GroBschadl, Z. Hu, K. Jarvinen, H.Wang, and I. Verbauwhede, “Elliptic curve cryptography with efficiently computable endomorphisms and its hardware implementations for the Internet of Things,” IEEE Trans. Comput., vol. 66, no. 5, pp. 773–785, May 2017.
[3] Z. Liu, D. Liu, and X. Zou, “An efficient and flexible hardware implementation of the dual-field elliptic curve cryptographic processor,” IEEE Trans. Ind. Electron., vol. 64, no. 3, pp. 2353–2362, Mar. 2017.
[4] B. Parhami, Computer Arithmetic: Algorithms and Hardware Design. New York, NY, USA: Oxford Univ. Press, 2000.
[5] P. L. Montgomery, “Modular multiplication without trial division,” Math. Comput., vol. 44, no. 170, pp. 519–521, Apr. 1985.
[6] S.-R. Kuang, K.-Y. Wu, and R.-Y. Lu, “Low-cost high-performance VLSI architecture for montgomery modular multiplication,” IEEE Trans. Very Large Scale Integr. (VLSI) Syst., vol. 24, no. 2, pp. 434–443, Feb. 2016.
[7] S.-R. Kuang, J.-P. Wang, K.-C. Chang, and H.-W. Hsu, “Energy-efficient high-throughput montgomery modular multipliers for RSA cryptosystems,” IEEE Trans. Very Large Scale Integr. (VLSI) Syst., vol. 21, no. 11, pp. 1999–2009, Nov. 2013.
[8] S. S. Erdem, T. Yanik, and A. Celebi, “A general digit-serial architecture for montgomery modular multiplication,” IEEE Trans. Very Large Scale Integr. (VLSI) Syst., vol. 25, no. 5, pp. 1658–1668, May 2017.
[9] R. S. Katti and S. K. Srinivasan, “Efficient hardware implementation of a new pseudo-random bit sequence generator,” in Proc. IEEE Int. Symp. Circuits Syst., Taipei, Taiwan, May 2009, pp. 1393–1396.
[10] A. K. Panda and K. C. Ray, “Modified dual-CLCG method and its VLSI architecture for pseudorandom bit generation,” IEEE Trans. Circuits Syst. I, Reg. Papers, vol. 66, no. 3, pp. 989–1002, Mar. 2019.
[11] A. Kumar Panda and K. Chandra Ray, “A coupled variable input LCG method and its VLSI architecture for pseudorandom bit generation,” IEEE Trans. Instrum. Meas., vol. 69, no. 4, pp. 1011–1019, Apr. 2020.
[12] N. Weste and K. Eshraghian, Principles of CMOS VLSI Design—A Systems Perspective. Reading, MA, USA: Addison-Wesley, 1985.
[13] T. Kim, W. Jao, and S. Tjiang, “Circuit optimization using carry-save adder cells,” IEEE Trans. Comput.-Aided Design Integr. Circuits Syst., vol. 17, no. 10, pp. 974–984, Oct. 1998.
[14] A. Rezai and P. Keshavarzi, “High-throughput modular multiplication and exponentiation algorithms using multibit-scan– multibit-shift technique,” IEEE Trans. Very Large Scale Integr. (VLSI) Syst., vol. 23, no. 9, pp. 1710–1719, Sep. 2015.
[15] A. K. Panda and K. C. Ray, “Design and FPGA prototype of 1024- bit Blum-Blum-Shub PRBG architecture,” in Proc. IEEE Int. Conf. Inf. Commun. Signal Process. (ICICSP), Singapore, Sep. 2018, pp. 38–43.
[16] T. Han and D. A. Carlson, “Fast area-efficient VLSI adders,” in Proc. IEEE 8th Symp. Comput. Arithmetic (ARITH), May 1987, pp. 49–56.
[17] D. L. Harris, “Parallel prefix networks that make tradeoffs between logic levels, fanout and wiring racks,” U.S. Patent 0 225 706 A1, Nov. 11, 2004.
[18] H. Ling, “High-speed binary adder,” IBM J. Res. Develop., vol. 25, no. 3, pp. 156–166, Mar. 1981.
[19] R. Jackson and S. Talwar, “High speed binary addition,” in Proc. Conf. Rec. 38th Asilomar Conf. Signals, Syst. Comput., vol. 2. Pacific Grove, CA, USA, Nov. 2004, pp. 1350–1353.
[20] K. S. Pandey, D. K. B. N. Goel, and H. Shrimali, “An ultra-fast parallel prefix adder,” in Proc. IEEE 26th Symp. Comput. Arithmetic (ARITH), Kyoto, Japan, Jun. 2019, pp. 125–134.
[21] F. Jafarzadehpour, A. S. Molahosseini, A. A. Emrani Zarandi, and L. Sousa, “New energy-efficient hybrid wide-operand adder architecture,” IET Circuits, Devices Syst., vol. 13, no. 8, pp. 1221–1231, Nov. 2019.
[22] S. Muthyala Sudhakar, K. P. Chidambaram, and E. E. Swartzlander, “Hybrid Han-Carlson adder,” in Proc. IEEE 55th Int. Midwest Symp. Circuits Syst. (MWSCAS), Boise, ID, USA, Aug. 2012, pp. 818–821.