Using Spark Machine Learning Models to Perform Predictive Analysis on Flight Ticket Pricing Data

التفاصيل البيبلوغرافية
العنوان: Using Spark Machine Learning Models to Perform Predictive Analysis on Flight Ticket Pricing Data
المؤلفون: Wong, Philip, Thant, Phue, Yadav, Pratiksha, Antaliya, Ruta, Woo, Jongwook
سنة النشر: 2023
المجموعة: Computer Science
مصطلحات موضوعية: Computer Science - Machine Learning, Computer Science - Distributed, Parallel, and Cluster Computing
الوصف: This paper discusses predictive performance and processes undertaken on flight pricing data utilizing r2(r-square) and RMSE that leverages a large dataset, originally from Expedia.com, consisting of approximately 20 million records or 4.68 gigabytes. The project aims to determine the best models usable in the real world to predict airline ticket fares for non-stop flights across the US. Therefore, good generalization capability and optimized processing times are important measures for the model. We will discover key business insights utilizing feature importance and discuss the process and tools used for our analysis. Four regression machine learning algorithms were utilized: Random Forest, Gradient Boost Tree, Decision Tree, and Factorization Machines utilizing Cross Validator and Training Validator functions for assessing performance and generalization capability.
Comment: 4 pages, 13 figures, 1 table
نوع الوثيقة: Working Paper
URL الوصول: http://arxiv.org/abs/2310.07787
رقم الأكسشن: edsarx.2310.07787
قاعدة البيانات: arXiv