Multi-Agent Task Flow Dispatching in a Heterogeneous Shared Supercomputer Center
DOI:
https://doi.org/10.14529/jsfi260203Keywords:
high performance computing, hybrid computing systems, multi-agent scheduler, reconfigurable systems, DAG schedulingAbstract
The slowdown in performance growth of general-purpose processors is driving the adoption of specialized accelerators such as GPUs and FPGAs in shared supercomputer centers (SSCs). FPGAs are attractive due to hardware-level parallelism and energy efficiency, but their integration significantly increases system heterogeneity: even with similar hardware, differences in bitstreams make nodes functionally nonequivalent. For streams of short jobs, frequent reconfiguration causes substantial overhead, while limited configuration-memory endurance constrains the number of configuration cycles. Classical HPC schedulers usually ignore current FPGA configurations and reconfiguration costs, leading to suboptimal task placement and reduced throughput. This paper presents a multi-agent task dispatcher for a heterogeneous SSC that explicitly accounts for FPGA configuration state and reconfiguration latency. Software agents on each FPGA node make local decisions on task admission and configuration changes, coordinating via bulletin-board queues. The model incorporates computation time, data transfer, and reconfiguration overhead into the scheduling objective. A prototype was implemented on a cluster with up to 18 Tertius2T reconfigurable units and two server nodes. Experiments with queues of up to 3500 tasks show that the multi-agent dispatcher reduces average task completion time and keeps placement and reconfiguration overheads within 40–1200 ms despite individual FPGA configuration times of at least 13 s, demonstrating resilience to heterogeneity and improved resource utilization.
References
Adimora, K., Gundla, S.R., Sun, H.: Machine learning approaches for optimizing high-performance computing scheduling: a comprehensive survey and analysis. Cluster Computing 28(13), 831 (2025). https://doi.org/10.1007/s10586-025-05521-8
Chen, S., Huang, J., Xu, X., et al.: Integrated optimization of partitioning, scheduling, and floorplanning for partially dynamically reconfigurable systems. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 39(1), 199–212 (2018). https://doi.org/10.1109/tcad.2018.2883982
Houssein, E.H., Gad, A.G., Wazery, Y.M., Suganthan, P.N.: Task scheduling in cloud computing based on meta-heuristics: Review, taxonomy, open challenges, and future trends. Swarm and Evolutionary Computation 62, 100841 (2021). https://doi.org/10.1016/j.swevo.2021.100841
Hsieh, F.S.: Analysis of contract net in multi-agent systems. Automatica 42(5), 733–740 (2006). https://doi.org/10.1016/j.automatica.2005.12.002
Huang, M., Narayana, V.K., Simmler, H., et al.: Reconfiguration and communication-aware task scheduling for high-performance reconfigurable computing. ACM Transactions on Reconfigurable Technology and Systems (TRETS) 3(4), 1–25 (2010). https://doi.org/10.1145/1862648.1862650
Jiang, Y., Yi, P., Zhang, S., Zhong, Y.: Constructing agents blackboard communication architecture based on graph theory. Computer Standards & Interfaces 27(3), 285–301 (2005). https://doi.org/10.1016/j.csi.2004.09.003
Kaliaev, A.: Multiagent approach for building distributed adaptive computing system. Procedia Computer Science 18, 2193–2202 (2013). https://doi.org/10.1016/j.procs.2013.05.390
Kostenetskiy, P., Shamsutdinov, A., Chulkevich, R., et al.: HPC TaskMaster – Task Efficiency Monitoring System for the Supercomputer Center. In: Sokolinsky, L., Zymbler, M. (eds.) Parallel Computational Technologies. pp. 17–29. Springer International Publishing, Cham (2022). https://doi.org/10.1007/978-3-031-11623-0_2
Macronix International Co., L.: Wear Leveling in NAND Flash Memory. https://www.mxic.com.tw/Lists/ApplicationNote/Attachments/1913/AN0289V1-Wear%20Leveling%20in%20NAND%20Flash%20Memory.pdf (2014), accessed: 2016-04-08
Parasumanna Gokulan, B., Srinivasan, D.: An Introduction to Multi-Agent Systems, vol. 310, pp. 1–27 (07 2010). https://doi.org/10.1007/978-3-642-14435-6_1
Rodriguez-Canal, G., Brown, N., Torres, Y., Gonzalez-Escribano, A.: Task-based preemptive scheduling on FPGAs leveraging partial reconfiguration. Concurrency and Computation: Practice and Experience 35(25), e7867 (2023). https://doi.org/10.1002/cpe.7867
Samayoa, W.F., Crespo, M.L., Cicuttin, A., Carrato, S.: A survey on FPGA-based heterogeneous clusters architectures. IEEE Access 11, 67679–67706 (2023). https://doi.org/10.1109/ACCESS.2023.3288431
Downloads
Published
How to Cite
License
Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution-Non Commercial 3.0 License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.