المجلة الدولية للعلوم والتقنية

International Science and Technology Journal

الرئيسية < البحوث والدراسات < تفاصيل بحث أو دراسة

Abstract Linear and Nonlinear Dimensionality Reduction and Their Impact on Nonparametric Regression: A Monte Carlo Comparison of PCA and KPCA

الملخص
يهدف هذا البحث إلى إجراء مقارنة تجريبية بين تقنية خطية لخفض الأبعاد، تُعرف باسم "تحليل المكونات الرئيسية" (PCA)، وتقنية غير خطية لخفض الأبعاد، وهي "تحليل المكونات الرئيسية المعتمد على النواة" (KPCA)، وذلك لاستخدامهما كخطوة معالجة أولية في الانحدار غير المعلمي (nonparametric regression). ويعود سبب هذه المقارنة إلى "لعنة الأبعاد" (curse of dimensionality)، التي من المرجح أن تؤدي إلى تدهور أداء الانحدار غير المعلمي مع زيادة عدد المتغيرات التفسيرية. ومع أن تقنية KPCA تُعد نظيراً غير خطي لتقنية PCA، فإن هذا لا يعني بالضرورة أنها ستقدم دائماً أداءً تنبؤياً أفضل من PCA، حتى في ظل وجود حدود غير خطية في دالة الانحدار. لذا، ندرس السيناريوهات التي تكون فيها PCA كافية، وتلك التي يُتوقع فيها أن تحقق KPCA تحسينات في الدقة. أُجريت الدراسة باستخدام محاكاة "مونت كارلو" (Monte Carlo) مع تغيير أحجام العينات (n = 20, 60, 150)، وعدد المتغيرات التفسيرية (p = 5, 10)، وعدد المكونات المُحتفظ بها (q = 2, 4) في بيئة البرمجة R. وقد تم النظر في ثلاث دوال انحدار ذات درجات متفاوتة من التعقيد في عدم الخطية، مع سحب المتغيرات التنبؤية والأخطاء العشوائية بشكل مستقل ومتماثل التوزيع (i.i.d.) من توزيعات طبيعية. وبعد عملية خفض الأبعاد، تم تقدير نموذج انحدار النواة "نادارايا-واتسون" (Nadaraya–Watson) باستخدام المكونات المستمدة من PCA وKPCA، مع استخدام نواة غاوسية (Gaussian kernel) في كلتا الخطوتين. وجرى تقييم الأداء التنبؤي باستخدام مقياسي "متوسط خطأ التنبؤ" (APE) و"متوسط مربع خطأ التنبؤ" (MSPE) عبر 500 تكرار لمحاكاة مونت كارلو لكل إعداد من إعدادات المحاكاة. وتُظهر النتائج أن الأداء التنبؤي يتحسن مع زيادة عدد المكونات المُحتفظ بها (من q = 2 إلى q = 4)، وأن أحجام العينات الأكبر تميل إلى خفض كل من APE وMSPE. وعند تثبيت عدد المكونات المُحتفظ بها، تؤدي زيادة عدد المتغيرات التفسيرية عموماً إلى ارتفاع أخطاء التنبؤ. كما أظهرت المقارنة بين PCA وKPCA عدم وجود تفوق ثابت لأي منهما على الآخر؛ إذ يعتمد الأداء النسبي لكل منهما على حجم العينة، وأبعاد المتغيرات التنبؤية، وعدد المكونات المُحتفظ بها، وبنية عدم الخطية في دالة الانحدار................ الكلمات المفتاحية:.............. اختزال الأبعاد، تحليل المكونات الرئيسية (PCA)، تحليل المكونات الرئيسية بالنواة (KPCA)، الانحدار اللامعلمي، لعنة الأبعاد، اللاخطية
Abstract
The objective of this paper is to provide an empirical comparison between a linear dimensionality reduction technique, referred to as principal component analysis (PCA), and a nonlinear dimensionality reduction technique, that is, kernel principal component analysis (KPCA), used as a preprocessing step for nonparametric regression. The reason for this contrast is the curse of dimensionality, which is likely to worsen the performance of nonparametric regression with an increasing number of explanatory variables. However, although KPCA constitutes a nonlinear analog of PCA, this does not imply that KPCA would always necessarily provide better predictive performance than PCA, given that there are nonlinear terms involved in the regression function. Thus, we study the scenarios where PCA is good enough and those where KPCA is expected to bring some improvements in accuracy. Using Monte Carlo simulation, the present study was performed with varying sample sizes (n = 20, 60, 150), numbers of explanatory variables (p = 5, 10), and numbers of retained components (q = 2, 4) in the R environment. Three regression functions with increasingly complex nonlinearity were considered, with predictors and random errors drawn i.i.d. from normal distributions. After the dimensionality reduction, the Nadaraya–Watson kernel regression estimator was estimated using the components derived from PCA and KPCA, where a Gaussian kernel was used in both steps. The predictive performance was evaluated using the Average Prediction Error (APE) and Mean Squared Prediction Error (MSPE) for 500 Monte Carlo replications for each simulation setting. The results show that the predictive performance improves with the number of retained components, from q = 2 to q = 4, and that larger sample sizes tend to decrease both APE and MSPE. For a fixed number of retained components, increasing the number of explanatory variables generally leads to higher prediction errors. The PCA vs. KPCA comparison does not show a consistent superiority of one over the other, but their relative performance depends on the sample size, predictor dimension, number of retained components, and the nonlinearity structure of the regression function....................... Keywords:.................. Dimensionality reduction, Principal Component Analysis, PCA, Kernel PCA, KPCA, Non-parametric regression, Curse of dimensionality, Nonlinearity.