一种基于随机掩码的低通信量Logistic回归外包训练方案
2021-02-11黄晓文王政杰崔硕硕张宇浩邓国强
黄晓文 王政杰 崔硕硕 张宇浩 邓国强




摘要:Logistic回归是一种典型的机器学习模型,因其在疾病诊断、金融预测等许多应用表现优越而受到广泛关注。logistic回归模型的建立不仅依赖于算法,更依赖于大量有效的训练数据。尽管构建高精度模型并提供预测服务有诸多优点,但用户的敏感信息数据造成隐私问题。因此,该文提出一个新的logistic回归外包训练方案。在该方案中,用户会预先对私有数据进行处理,并添加随机掩码的数据矩阵上传给聚合器,聚合器将聚合得到的全局训练矩阵上传给云服务器进行训练。该方案在满足数据隐私的安全性需求下具有较高的计算效率和较低的通信开销。
关键词:Logistic回归隐私保护随机掩码低通信量
中图分类号:TP309 文献标识码:A 文章编号:1672-3791(2021)12(a)-0000-00
An Outsourcing Training Scheme of Low-traffic Logistic Regression Based on Random Mask
HUANG XiaowenWANG ZhengjieCUI ShuoshuoZHANG YuhaoDENG Guoqiang*
(School of Mathematics and Computing Science, Guilin University of Electronic Technology, Guilin, Guangxi Zhuang Autonomous Region, 541004 China)
Abstract: Logistic regression is a typical machine learning model, and its superior performance in many applications such as disease diagnosis and financial forecasting is widely welcomed. Providing user data to the server for logistic regression is a new service mode. Although predictive services have many advantages, the user 's sensitive data itself has privacy problems. Therefore, a new outsourcing privacy protection logistic training framework is proposed. In our framework, the user processes the private data in advance, and uploads the data matrix with random mask to the aggregator. The aggregator uploads the aggregated global training matrix to the cloud server for training. The scheme meets the security requirements of data privacy and has high efficiency in computing and communication overhead.
Key Words: Logistic regression; Privacy-preserving; Random mask; Low-traffic
機器学习模型在各种应用领域取得了前所未有的发展[1-3]。然而,由于庞大的数据量,训练过程是一项计算和存储密集型任务。此外,通常针对敏感数据(如医疗记录、浏览历史记录或金融交易)进行训练时,会引发数据集的安全性和隐私问题。
一方面,由于其复杂性,训练过程往往需要外包给如云这样的更强大的计算平台。另一方面,训练数据集通常是敏感的,它可能包含一些敏感或私有信息,一旦披露,将导致灾难性后果。因此,对于参与云计算的数据需要进行隐藏得到密文数据。然而,机器学习算法不能直接访问密文,如果将解密密钥提供给诚实且好奇的云服务器又无法确保数据隐私。……
