Sketching Datasets for Large-Scale Learning (long version)

التفاصيل البيبلوغرافية
العنوان: Sketching Datasets for Large-Scale Learning (long version)
المؤلفون: Gribonval, Rémi, Chatalic, Antoine, Keriven, Nicolas, Schellekens, Vincent, Jacques, Laurent, Schniter, Philip
سنة النشر: 2020
المجموعة: Computer Science
Mathematics
Statistics
مصطلحات موضوعية: Statistics - Machine Learning, Computer Science - Information Theory, Computer Science - Machine Learning
الوصف: This article considers "compressive learning," an approach to large-scale machine learning where datasets are massively compressed before learning (e.g., clustering, classification, or regression) is performed. In particular, a "sketch" is first constructed by computing carefully chosen nonlinear random features (e.g., random Fourier features) and averaging them over the whole dataset. Parameters are then learned from the sketch, without access to the original dataset. This article surveys the current state-of-the-art in compressive learning, including the main concepts and algorithms, their connections with established signal-processing methods, existing theoretical guarantees -- on both information preservation and privacy preservation, and important open problems.
نوع الوثيقة: Working Paper
URL الوصول: http://arxiv.org/abs/2008.01839
رقم الأكسشن: edsarx.2008.01839
قاعدة البيانات: arXiv