Journals & Magazines >IEEE Journal on Selected Area... >Volume: 39 Issue: 8

LOSP: Overlap Synchronization Parallel With Local Compensation for Fast Distributed Training

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

When running in Parameter Server (PS), the Distributed Stochastic Gradient Descent (D-SGD) incurs significant communication delays and huge communication overhead due to ...Show More

Metadata

Abstract:

When running in Parameter Server (PS), the Distributed Stochastic Gradient Descent (D-SGD) incurs significant communication delays and huge communication overhead due to the model synchronization. Moreover, considering the heterogeneity of computational capability among workers, traditional synchronization modes incur under-utilization of computational resources because fast workers have to wait for slow ones finishing the computation. Although our previous work OSP can effectively solve these problems by overlapping the computation and communication procedures and allowing adaptive multiple local updates in distributed training, it causes the staleness problem brought by the overlap, yielding a performance degradation. In this paper, we propose a new method named LOSP by introducing local compensation to our previous synchronization mechanism, which mitigates adverse effects caused by the overlapping synchronization. We theoretically prove that LOSP (1) preserves the same convergence rate as the sequential SGD for non-convex problems, and (2) exhibits good scalability due to the linear speedup property with respect to both the number of workers and the average number of local updates. Evaluations show that LOSP significantly improves performance over the state-of-the-art ones in terms of both convergence accuracy and communication cost.

Published in: IEEE Journal on Selected Areas in Communications ( Volume: 39, Issue: 8, August 2021)

Page(s): 2541 - 2557

Date of Publication: 07 June 2021

ISSN Information:

DOI: 10.1109/JSAC.2021.3087272

Funding Agency:

Contents

I. Introduction

Machine learning has demonstrated great promises in a wide range of application domains, e.g., self-driving, smart city, language processing, etc., which are fundamentally altering the way individuals and organizations live, work and interact [2]–[4]. With the rapid growth of training data and machine learning model size, how to efficiently train machine learning model in a distributed manner has received much attention since computation can be parallelized in multiple nodes. A widely used framework in distributed machine learning is data parallelism over Parameter Server (PS) architecture, i.e., data are distributed over multiple workers and a global model is cooperatively optimized with the coordination of servers [5], [6]. As to the algorithm running in PS for solving training problems, distributed Stochastic Gradient Descent (D-SGD) is usually adopted because it is applicable to various model optimizations with proven efficiency in terms of scalability [5], [7], [8].

References is not available for this document.

LOSP: Overlap Synchronization Parallel With Local Compensation for Fast Distributed Training

Abstract:

Metadata

Abstract:

ISSN Information:

Funding Agency:

I. Introduction

References

IEEE Account

Purchase Details

Profile Information

Need Help?

LOSP: Overlap Synchronization Parallel With Local Compensation for Fast Distributed Training

Alerts

Abstract:

Metadata

Abstract:

ISSN Information:

Funding Agency:

I. Introduction

References