INVESTIGATION OF A TWO-FACTOR FULLY CONNECTED LINEAR REGRESSION MODEL
Михаил Павлович Базилевский
Abstract
Open-access reader
Михаил Павлович Базилевский
Abstract
Open-access reader
Данная работа посвящена исследованию модели полносвязной линейной регрессии, представляющей собой синтез модели парной линейной регрессии и регрессии Деминга. Если множественная регрессия строится по принципу «независимые переменные влияют на зависимую», то принципом полносвязной регрессии является «все переменные влияют друг на друга». Полносвязная регрессия достаточно просто оценивается, лишена эффекта мультиколлинеарности, имеет гораздо более разнообразную интерпретацию, чем множественная регрессия, и пригодна для прогнозирования. Однако при построении полносвязной регрессии неизвестным остается соотношение дисперсий ошибок независимых переменных. В данной работе найдено такое соотношение дисперсий ошибок независимых переменных, которое обеспечивает наилучшие аппроксимационные качества вторичной модели полносвязной регрессии. Результаты исследования оформлены в виде теоремы. Из теоремы следует, что значение коэффициента детерминации вторичной модели полносвязной регрессии будет наибольшим либо когда она принимает вид двухфакторной линейной регрессии, либо вид наилучшей по коэффициенту детерминации однофакторной линейной регрессии. Таким образом, осуществляется отбор информативных регрессоров в регрессионной модели. Установлено, что в основе такого отбора лежит полная согласованность знаков коэффициентов при независимых переменных знакам соответствующих коэффициентов корреляции. This paper is devoted to the study of a fully connected linear regression model, which is a synthesis of the pairing linear regression model and the Deming regression model. If multiple regression is based on the principle “independent variables influence dependent”, then the principle of fully connected regression is “all variables influence each other”. A fully connected regression is fairly simply estimated, devoid of multicollinearity effect, has a much more diverse interpretation than multiple regression, and is suitable for prediction. However, when building a fully connected regression, the ratio of error variances of independent variables remains unknown. In this paper, we find the ratio of error variances of independent variables that provides the best approximation qualities of the secondary fully connected regression model. The research results are presented in the form of a theorem. It follows from the theorem that the value of the coefficient of determination of the secondary model of a fully connected regression will be greatest either when it takes the form of a two-factor linear regression or the best one in the coefficient of determination of a single-factor linear regression. Thus, the selection of informative regressors in the regression model is carried out. It is established that the basis of such a selection is the complete consistency of the signs of the coefficients with independent variable signs of the corresponding correlation coefficients.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Данная работа посвящена исследованию модели полносвязной линейной регрессии, представляющей собой синтез модели парной линейной регрессии и регрессии Деминга. Если множественная регрессия строится по принципу «независимые переменные влияют на зависимую», то принципом полносвязной регрессии является «все переменные влияют друг на друга». Полносвязная регрессия достаточно просто оценивается, лишена эффекта мультиколлинеарности, имеет гораздо более разнообразную интерпретацию, чем множественная регрессия, и пригодна для прогнозирования. Однако при построении полносвязной регрессии неизвестным остается соотношение дисперсий ошибок независимых переменных. В данной работе найдено такое соотношение дисперсий ошибок независимых переменных, которое обеспечивает наилучшие аппроксимационные качества вторичной модели полносвязной регрессии. Результаты исследования оформлены в виде теоремы. Из теоремы следует, что значение коэффициента детерминации вторичной модели полносвязной регрессии будет наибольшим либо когда она принимает вид двухфакторной линейной регрессии, либо вид наилучшей по коэффициенту детерминации однофакторной линейной регрессии. Таким образом, осуществляется отбор информативных регрессоров в регрессионной модели. Установлено, что в основе такого отбора лежит полная согласованность знаков коэффициентов при независимых переменных знакам соответствующих коэффициентов корреляции. This paper is devoted to the study of a fully connected linear regression model, which is a synthesis of the pairing linear regression model and the Deming regression model. If multiple regression is based on the principle “independent variables influence dependent”, then the principle of fully connected regression is “all variables influence each other”. A fully connected regression is fairly simply estimated, devoid of multicollinearity effect, has a much more diverse interpretation than multiple regression, and is suitable for prediction. However, when building a fully connected regression, the ratio of error variances of independent variables remains unknown. In this paper, we find the ratio of error variances of independent variables that provides the best approximation qualities of the secondary fully connected regression model. The research results are presented in the form of a theorem. It follows from the theorem that the value of the coefficient of determination of the secondary model of a fully connected regression will be greatest either when it takes the form of a two-factor linear regression or the best one in the coefficient of determination of a single-factor linear regression. Thus, the selection of informative regressors in the regression model is carried out. It is established that the basis of such a selection is the complete consistency of the signs of the coefficients with independent variable signs of the corresponding correlation coefficients.
Key concepts: Multicollinearity, Proper linear model, Linear predictor function, Regression diagnostic, Linear regression, Mathematics, Regression analysis, Statistics