Privacy-Preserving Synthetic Data Generation for Recommendation Systems

by   Fan Liu, et al.

Recommendation systems make predictions chiefly based on users' historical interaction data (e.g., items previously clicked or purchased). There is a risk of privacy leakage when collecting the users' behavior data for building the recommendation model. However, existing privacy-preserving solutions are designed for tackling the privacy issue only during the model training and results collection phases. The problem of privacy leakage still exists when directly sharing the private user interaction data with organizations or releasing them to the public. To address this problem, in this paper, we present a User Privacy Controllable Synthetic Data Generation model (short for UPC-SDG), which generates synthetic interaction data for users based on their privacy preferences. The generation model aims to provide certain privacy guarantees while maximizing the utility of the generated synthetic data at both data level and item level. Specifically, at the data level, we design a selection module that selects those items that contribute less to a user's preferences from the user's interaction data. At the item level, a synthetic data generation module is proposed to generate a synthetic item corresponding to the selected item based on the user's preferences. Furthermore, we also present a privacy-utility trade-off strategy to balance the privacy and utility of the synthetic data. Extensive experiments and ablation studies have been conducted on three publicly accessible datasets to justify our method, demonstrating its effectiveness in generating synthetic data under users' privacy preferences.


page 1

page 2

page 3

page 4


Synthetic Data and Simulators for Recommendation Systems: Current State and Future Directions

Synthetic data and simulators have the potential to markedly improve the...

FedGNN: Federated Graph Neural Network for Privacy-Preserving Recommendation

Graph neural network (GNN) is widely used for recommendation to model hi...

Active Preference Elicitation via Adjustable Robust Optimization

We consider the problem faced by a recommender system which seeks to off...

Privacy-Preserving Synthetic Educational Data Generation

Institutions collect massive learning traces but they may not disclose i...

Synthetic Data for Social Good

Data for good implies unfettered access to data. But data owners must be...

Generating synthetic transactional profiles

Financial institutions use clients' payment transactions in numerous ban...

GeoPointGAN: Synthetic Spatial Data with Local Label Differential Privacy

Synthetic data generation is a fundamental task for many data management...

Please sign up or login with your details

Forgot password? Click here to reset