paperNovember 2022

Mean Estimation with User-level Privacy under Data Heterogeneity

AuthorsRachel Cummings*, Vitaly Feldman*, Audra McMillan*, Kunal Talwar*

A key challenge in many modern data analysis tasks is that user data is heterogeneous. Different users may possess vastly different numbers of data points. More importantly, it cannot be assumed that all users sample from the same underlying distribution. This is true, for example in language data, where different speech styles result in data heterogeneity. In this work we propose a simple model of heterogeneous user data that differs in both distribution and quantity of data, and we provide a method for estimating the population-level mean while preserving user-level differential privacy. We demonstrate asymptotic optimality of our estimator and also prove general lower bounds on the error achievable in our problem. In particular, while the optimal non-private estimator can be shown to be linear, we show that privacy constrains us to use a non-linear estimator.

*=Equal Contributors

Related readings and updates.

December 5, 2024research area Methods and Algorithms, research area Privacyconference NeurIPS

Motivated by the problem of next word prediction on user devices we introduce and study the problem of personalized frequency histogram estimation in a federated setting. In this problem, over some domain, each user observes a number of samples from a distribution which is specific to that user. The goal is to compute for all users a personalized estimate of the user's distribution with error measured in KL divergence. We focus on addressing two...

November 10, 2022research area Methods and Algorithms, research area Privacyconference NeurIPS

*= Equal Contributions

Recovering linear subspaces from data is a fundamental and important task in statistics and machine learning. Motivated by heterogeneity in Federated Learning settings, we study a basic formulation of this problem: the principal component analysis (PCA), with a focus on dealing with irregular noise. Our data come from $n$ users with user $i$ contributing data samples from a $d$ -dimensional distribution with mean $\mu_i$ ....

Mean Estimation with User-level Privacy under Data Heterogeneity

Related readings and updates.

Private and Personalized Frequency Estimation in a Federated Setting

Subspace Recovery from Heterogeneous Data with Non-isotropic Noise

Discover opportunities in Machine Learning.