Elasticity in Parallel Sparse Triangular Solve
2026-07-02 • Distributed, Parallel, and Cluster Computing
Distributed, Parallel, and Cluster Computing
AI summaryⓘ
The authors introduce a new way to speed up solving certain math problems on computers called stale synchronous parallel. They created a scheduler named ElasticDivide that smartly handles task order to overlap waiting times and computing times, making the process faster. Their method showed a 7% to 30% speed improvement on an ARM machine and 19% to 60% on an x86 machine with many cores, compared to existing schedulers. This means their approach can make complex calculations more efficient on modern processors.
stale synchronous parallelsparse triangular linear systemdirected acyclic graph (DAG) schedulerparallel computingsynchronizationElasticDivideGrowLocalSpMPARM architecturex86 architecture
Authors
Raphael S. Steiner, Christos K. Matzoros, Pál András Papp, Toni Böhnlein, A. N. Yzelman
Abstract
We introduce stale synchronous parallel as a mode of execution in parallel sparse triangular linear system solve and present a general directed-acyclic-graph scheduler capable of producing such schedules. Stale-synchronous-parallel schedules allow the overlap of synchronisation and compute which results in a geometric-mean speed-up of $7$-$30\%$ of our scheduler, ElasticDivide, over state-of-the-art synchronous scheduler GrowLocal on an ARM machine using 48 cores. On an x86 machine using 48 cores, we report geometric-mean speed-ups of $19$-$60\%$ over SpMP.