PDSS: A Privacy-Preserving Framework for Step-by-Step Distillation of Large Language Models

التفاصيل البيبلوغرافية
العنوان: PDSS: A Privacy-Preserving Framework for Step-by-Step Distillation of Large Language Models
المؤلفون: Fan, Tao, Kang, Yan, Chen, Weijing, Gu, Hanlin, Song, Yuanfeng, Fan, Lixin, Chen, Kai, Yang, Qiang
سنة النشر: 2024
المجموعة: Computer Science
مصطلحات موضوعية: Computer Science - Computation and Language, Computer Science - Artificial Intelligence
الوصف: In the context of real-world applications, leveraging large language models (LLMs) for domain-specific tasks often faces two major challenges: domain-specific knowledge privacy and constrained resources. To address these issues, we propose PDSS, a privacy-preserving framework for step-by-step distillation of LLMs. PDSS works on a server-client architecture, wherein client transmits perturbed prompts to the server's LLM for rationale generation. The generated rationales are then decoded by the client and used to enrich the training of task-specific small language model(SLM) within a multi-task learning paradigm. PDSS introduces two privacy protection strategies: the Exponential Mechanism Strategy and the Encoder-Decoder Strategy, balancing prompt privacy and rationale usability. Experiments demonstrate the effectiveness of PDSS in various text generation tasks, enabling the training of task-specific SLM with enhanced performance while prioritizing data privacy protection.
نوع الوثيقة: Working Paper
URL الوصول: http://arxiv.org/abs/2406.12403
رقم الأكسشن: edsarx.2406.12403
قاعدة البيانات: arXiv