Memory-Efficient Differentially Private Training with Gradient Random Projection

18 June 2025

Alex Mulrooney

Devansh Gupta

James Flemings

Huanyu Zhang

Murali Annavaram

Meisam Razaviyayn

Xinwei Zhang

Author Contacts:

ArXiv (abs)PDF HTML

Main:8 Pages

5 Figures

Bibliography:4 Pages

15 Tables

Appendix:15 Pages

Abstract

Differential privacy (DP) protects sensitive data during neural network training, but standard methods like DP-Adam suffer from high memory overhead due to per-sample gradient clipping, limiting scalability. We introduce DP-GRAPE (Gradient RAndom ProjEction), a DP training method that significantly reduces memory usage while maintaining utility on par with first-order DP approaches. Rather than directly applying DP to GaLore, DP-GRAPE introduces three key modifications: (1) gradients are privatized after projection, (2) random Gaussian matrices replace SVD-based subspaces, and (3) projection is applied during backpropagation. These contributions eliminate the need for costly SVD computations, enable substantial memory savings, and lead to improved utility. Despite operating in lower-dimensional subspaces, our theoretical analysis shows that DP-GRAPE achieves a privacy-utility trade-off comparable to DP-SGD. Our extensive empirical experiments show that DP-GRAPE can reduce the memory footprint of DP training without sacrificing accuracy or training time. In particular, DP-GRAPE reduces memory usage by over 63% when pre-training Vision Transformers and over 70% when fine-tuning RoBERTa-Large as compared to DP-Adam, while achieving similar performance. We further demonstrate that DP-GRAPE scales to fine-tuning large models such as OPT with up to 6.7 billion parameters.

View on arXiv

@article{mulrooney2025_2506.15588,
  title={ Memory-Efficient Differentially Private Training with Gradient Random Projection },
  author={ Alex Mulrooney and Devansh Gupta and James Flemings and Huanyu Zhang and Murali Annavaram and Meisam Razaviyayn and Xinwei Zhang },
  journal={arXiv preprint arXiv:2506.15588},
  year={ 2025 }
}

Comments on this paper