Rapid Training of Neural Networks Using the Arctur Reconfigurable Computer System

Authors

  • Ilya I. Levin "Scientific Research Center of Supercomputers and Neurocomputers" Co Ltd ("SRC SC & NC" Co Ltd), Taganrog, Russian Federation
  • Dmitriy A. Sorokin "Scientific Research Center of Supercomputers and Neurocomputers" Co Ltd ("SRC SC & NC" Co Ltd), Taganrog, Russian Federation
  • Vasiliy B. Kovalenko "Scientific Research Center of Supercomputers and Neurocomputers" Co Ltd ("SRC SC & NC" Co Ltd), Taganrog, Russian Federation

DOI:

https://doi.org/10.14529/jsfi260204

Keywords:

reconfigurable computer systems, FPGA, neural network training, dataflow discontinuities

Abstract

 Linear scaling of hardware resources in widely used systems based on graphics processing units (GPUs) and central processing units (CPUs) for neural network training problems provides only logarithmic scaling of real performance. This is due to the presence of dataow discontinuities in the computational graphs of neural network training problems, which lead to a significant decrease in computational intensity and may even lead to a complete stop between execution stages. These limitations are unacceptable if neural network training must be performed within a limited time regardless of the amount of processed data. Reconfigurable computer systems (RCS) based on field-programmable gate arrays (FPGAs), unlike traditional systems, demonstrate the potential to address these limitations through the structural organization of calculations, as well as by adapting to the specifics of solving the problem. Methods for minimizing the impact of data discontinuities have been formulated and theoretically justified for RCS. Based on these methods, a methodology has been developed to construct efficient computing structures to solve neural network training problems, ensuring real RCS performance that exceeds 70% of peak performance. For GPU-based systems, achieving this level of computational efficiency is generally impossible. For the first time, neural network-based data processing problems have been implemented using a parallel-pipeline approach on the Arctur RCS, which contains 96 XCVU37P FPGAs. Experimental results have shown that for these problems, the energy efficiency of a single Arctur RCS block is 56% higher than the energy efficiency of 16 NVIDIA DGX A100 blocks. Increasing the number of Arctur RCS nodes enables near-linear scaling of real performance. For example, the RCS rack consisting of 16 Arctur nodes is capable of providing the real performance of about 7 PFLOPS. For neural network problems such as ResNet-50v1.5 training, the performance of this rack is comparable to that of 12 NVIDIA DGX A100 racks with 2.2× lower power consumption.

References

Lavrentyev, A.S., Solovyev, V.D.: Artificial intelligence: theory and practice. Nauka Publ., Moscow (2020).

NVIDIA DGX A100. https://www.nvidia.com/ru-ru/data-center/dgx-a100/, accessed: 2026-04-06

Cloud TPU system architecture. https://cloud.google.com/tpu/docs/system-architecture-tpu-vm, accessed: 2026-04-06.

Huawei Ascend 910 (512 TFLOPS). https://www.ixbt.com/news/2019/08/23/huawei-ascend-910-512-tflops-310.html, accessed: 2026-04-06

Xiong, R., Yang, Y., He, D., et al.: On layer normalization in the transformer architecture. Proceedings of the 37th International Conference on Machine Learning, 2020, vol. 119, pp. 10524–10533. https://proceedings.mlr.press/v119/xiong20b.html, accessed: 2026-04-08

He, K., Zhang, X., Ren, S., Sun, J.: Identity mappings in deep residual networks. Computer Vision – ECCV 2016, vol. 9908, pp. 630–645. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46493-0-8

Kasarkin, A.V., Levin, I.I., Sorokin, D.A.: New iteration parallel-based method for solving graph NP-complete problems with recon gurable computer systems. IOP Conference Series: Materials Science and Engineering 919(1), 052007 (2020). https://doi.org/10.1088/1757-899X/919/5/052007

Sorokin, D.A., Matrosov, A.Y., et al.: Real-time implementation of the problem of surface-related multiple prediction on RCS. Parallel Computational Technologies, PCT 2019, Kaliningrad, Russia, 2019, pp. 91–98.

Sorokin, D.A., Dordopulo, A.I., Levin, I.I.: Implementation of docking for molecular modeling on reconfigurable computing systems. Izvestiya SFedU. Engineering Sciences 7, 217–224. (2011).

APEgate project. https://apegate.roma1.infn.it/?page_id=656, accessed: 2026-04-08

Kalyaev, A.V., Levin, I.I.: Modular scalable multiprocessor systems with structural-procedural organization of computations. Janus-K Publ., Moscow (2003).

Kalyaev, A.V.: Training procedure for a multilayer neuroprocessor network using feedback. Information Technologies 6 (2002).

AMD. UltraScale Architecture GTH Transceivers User Guide (UG576). Version 1.7.1. San Jose, 2021. https://www.amd.com/content/dam/xilinx/support/documents/user_guides/ug576-ultrascale-gth-transceivers.pdf, accessed: 2026-04-08

Lee, J.-C., Kim, J., Kim, K.W., et al.: High bandwidth memory (HBM) with TSV technique. 2016 International SoC Design Conference, ISOCC, pp. 181–182 (2016). https://doi.org/10.1109/ISOCC.2016.7799847

Levin, I.I., Doronchenko, Yu.I.: Advanced reconfigurable supercomputer with immersion cooling. Proceedings of the XIV All-Russian Conference on Control Problems, VSPU-2024, pp. 2256–2260 (2024). https://doi.org/10.25699/vspu2024.2256

NVIDIA DGX A100 user guide. https://docs.nvidia.com/dgx/dgxa100-user-guide/introduction-to-dgxa100.html, accessed: 2026-04-08

Lohrmann, B., Warneke, D., Kao, O.: Nephele streaming: stream processing under QoS constraints at scale. https://doi.org/10.1007/s10586-013-0281-8 (2013), accessed: 2026-04-08

Intel Corporation. FPGA AI Suite – AI inference development tools for FPGAs. https://www.intel.com/content/www/us/en/products/details/fpga/development-tools/fpga-ai-suite.html, accessed: 2026-04-08

AMD. Vitis AI software for adaptive SoCs and FPGAs. https://www.amd.com/en/products/software/vitis-ai.html, accessed: 2026-04-08

AMD Research I& Advanced Development. FINN: dataflow compiler for QNN inference on FPGAs. https://xilinx.github.io/finn/, accessed: 2026-04-09

Zhao, W., Fu, H., Luk, W., et al.: F-CNN: an FPGA-based framework for training convolutional neural networks. Proc. 27th IEEE Int. Conf. Application-Specific Systems, Architectures and Processors, ASAP 2016, London, UK, 2016. pp. 107–114. IEEE (2016). https://doi.org/10.1109/ASAP.2016.7760779

Levin, I.I., Dordopulo, A.I.: Reconfigurable computing systems: resource-independent programming. Southern Federal University Publ., Rostov-on-Don (2025).

Chu, P.P.: RTL hardware design using VHDL: coding for efficiency, portability, and scalability. Wiley, Hoboken, NJ (2006).

NVIDIA DeepLearningExamples: ResNet-50 v1.5. https://github.com/NVIDIA/DeepLearningExamples/blob/master/PyTorch/Classification/ConvNets/resnet50v1.5/README.md, accessed: 2026-04-09

NVIDIA A100 datasheet. https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/a100/pdf/nvidia-a100-datasheet-nvidia-us-2188504-web.pdf, accessed: 2026-04-09

NVIDIA AI Enterprise sizing guide (archived). https://web.archive.org/web/20251110193817/https://docs.nvidia.com/ai-enterprise/sizing-guide/latest/appendix.html, accessed: 2026-04-09

MLCommons training benchmarks. https://mlcommons.org/benchmarks/training/, accessed: 2026-04-09

Downloads

Published

2026-07-30

How to Cite

Levin, I. I., Sorokin, D. A., & Kovalenko, V. B. (2026). Rapid Training of Neural Networks Using the Arctur Reconfigurable Computer System. Supercomputing Frontiers and Innovations, 13(2), 63–78. https://doi.org/10.14529/jsfi260204