Multi-Agent Task Flow Dispatching in a Heterogeneous Shared Supercomputer Center

Authors

DOI:

https://doi.org/10.14529/jsfi260203

Keywords:

high performance computing, hybrid computing systems, multi-agent scheduler, reconfigurable systems, DAG scheduling

Abstract

The slowdown in performance growth of general-purpose processors is driving the adoption of specialized accelerators such as GPUs and FPGAs in shared supercomputer centers (SSCs). FPGAs are attractive due to hardware-level parallelism and energy efficiency, but their integration significantly increases system heterogeneity: even with similar hardware, differences in bitstreams make nodes functionally nonequivalent. For streams of short jobs, frequent reconfiguration causes substantial overhead, while limited configuration-memory endurance constrains the number of configuration cycles. Classical HPC schedulers usually ignore current FPGA configurations and reconfiguration costs, leading to suboptimal task placement and reduced throughput. This paper presents a multi-agent task dispatcher for a heterogeneous SSC that explicitly accounts for FPGA configuration state and reconfiguration latency. Software agents on each FPGA node make local decisions on task admission and configuration changes, coordinating via bulletin-board queues. The model incorporates computation time, data transfer, and reconfiguration overhead into the scheduling objective. A prototype was implemented on a cluster with up to 18 Tertius2T reconfigurable units and two server nodes. Experiments with queues of up to 3500 tasks show that the multi-agent dispatcher reduces average task completion time and keeps placement and reconfiguration overheads within 40–1200 ms despite individual FPGA configuration times of at least 13 s, demonstrating resilience to heterogeneity and improved resource utilization.

References

Adimora, K., Gundla, S.R., Sun, H.: Machine learning approaches for optimizing high-performance computing scheduling: a comprehensive survey and analysis. Cluster Computing 28(13), 831 (2025). https://doi.org/10.1007/s10586-025-05521-8

Chen, S., Huang, J., Xu, X., et al.: Integrated optimization of partitioning, scheduling, and floorplanning for partially dynamically reconfigurable systems. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 39(1), 199–212 (2018). https://doi.org/10.1109/tcad.2018.2883982

Houssein, E.H., Gad, A.G., Wazery, Y.M., Suganthan, P.N.: Task scheduling in cloud computing based on meta-heuristics: Review, taxonomy, open challenges, and future trends. Swarm and Evolutionary Computation 62, 100841 (2021). https://doi.org/10.1016/j.swevo.2021.100841

Hsieh, F.S.: Analysis of contract net in multi-agent systems. Automatica 42(5), 733–740 (2006). https://doi.org/10.1016/j.automatica.2005.12.002

Huang, M., Narayana, V.K., Simmler, H., et al.: Reconfiguration and communication-aware task scheduling for high-performance reconfigurable computing. ACM Transactions on Reconfigurable Technology and Systems (TRETS) 3(4), 1–25 (2010). https://doi.org/10.1145/1862648.1862650

Jiang, Y., Yi, P., Zhang, S., Zhong, Y.: Constructing agents blackboard communication architecture based on graph theory. Computer Standards & Interfaces 27(3), 285–301 (2005). https://doi.org/10.1016/j.csi.2004.09.003

Kaliaev, A.: Multiagent approach for building distributed adaptive computing system. Procedia Computer Science 18, 2193–2202 (2013). https://doi.org/10.1016/j.procs.2013.05.390

Kostenetskiy, P., Shamsutdinov, A., Chulkevich, R., et al.: HPC TaskMaster – Task Efficiency Monitoring System for the Supercomputer Center. In: Sokolinsky, L., Zymbler, M. (eds.) Parallel Computational Technologies. pp. 17–29. Springer International Publishing, Cham (2022). https://doi.org/10.1007/978-3-031-11623-0_2

Macronix International Co., L.: Wear Leveling in NAND Flash Memory. https://www.mxic.com.tw/Lists/ApplicationNote/Attachments/1913/AN0289V1-Wear%20Leveling%20in%20NAND%20Flash%20Memory.pdf (2014), accessed: 2016-04-08

Parasumanna Gokulan, B., Srinivasan, D.: An Introduction to Multi-Agent Systems, vol. 310, pp. 1–27 (07 2010). https://doi.org/10.1007/978-3-642-14435-6_1

Rodriguez-Canal, G., Brown, N., Torres, Y., Gonzalez-Escribano, A.: Task-based preemptive scheduling on FPGAs leveraging partial reconfiguration. Concurrency and Computation: Practice and Experience 35(25), e7867 (2023). https://doi.org/10.1002/cpe.7867

Samayoa, W.F., Crespo, M.L., Cicuttin, A., Carrato, S.: A survey on FPGA-based heterogeneous clusters architectures. IEEE Access 11, 67679–67706 (2023). https://doi.org/10.1109/ACCESS.2023.3288431

Downloads

Published

2026-07-30

How to Cite

Kaliaev, I. A. ., Kaliaev, A. I., & Semenistyi, S. A. (2026). Multi-Agent Task Flow Dispatching in a Heterogeneous Shared Supercomputer Center. Supercomputing Frontiers and Innovations, 13(2), 47–62. https://doi.org/10.14529/jsfi260203