Conflict of user accuracy requirements and regression model indicators
DOI:
https://doi.org/10.34121/1028-9763-2025-1-91-102Keywords:
regression analysis, accuracy of approximation, least squares, weighted least squares, nonparametric Theil regression, regression coefficients, heteroskedasticity, outliers, large dispersion, residual variance, multiple correlation coefficient, mean relative deviation of the model from experimental data, maximum relative deviation, mean absolute deviation, maximum absolute deviationAbstract
The paper investigates the relationship between different groups of quality indicators of the regression model. Three groups of indicators are considered: statistical indicators, indicators of approximation accuracy in accordance with the requirements of technical applications, and theoretical indicators in relation to regression coefficients. The statistical indicators are those that are the most common in applied research (although they are not enough for a reasonable assessment of the regression equation quality), namely the residual variance and the multiple correlation coefficient. Approximation accuracy estimates are those that are most often put forward as adequacy criteria in technical applications: average relative deviation of the model from experimental data, maximum relative deviation, average absolute deviation, and maximum absolute deviation. In addition, the difference between the values of the correlation coefficients and their estimates is considered. In order to cover the possible variants of the conditions under which the model is built, the research was conducted under the following conditions: a) compliance with all the prerequisites of the regression analysis; b) compliance with all the prerequisites of the regression analysis in the presence of large dispersion (large dispersion of reproducibility); c) the presence of heteroskedasticity; d) the presence of outliers (with different locations and different numbers in the training sample). It was established that when the prerequisites and assumptions of the regression analysis are met, the difference when using different methods is insignificant from an applied point of view for the indicators of all groups. With large dispersion and heteroskedasticity, methods with the aim of obtaining the best relative indicators give a significant improvement of «their» characteristics with a significant deterioration of all other characteristics. Recommendations for the selection of methods depending on the characteristics of the data and requirements for model characteristics have been developed.
References
1. Лапач С.М., Радченко С.Г. Основні проблеми побудови регресійних моделей. Математичні машини і системи. 2012. № 4. С. 125–133.
2. Лапач С.М. Регресійний аналіз. Процесний підхід. Математичні машини і системи. 2016. № 1. C. 129–138.
3. Kelleher J.D., Namee, B.M., D`Arcy A. Fundamentals of Machine Leaning for predictive data analysis. Algorithms, worked Examples, and cade studies. Cambridge: The MIT Press, 2015. 624 p.
4. Kuhn M., Lohnson K. Applied Predictive Modelling. New York: Springer, 2013. 612 p.
5. Hastie T., Tibshirani R., Friedman J. The elements of Statistical Leaning. Data Mining, Inference and Prediction. Second Edition. Springer, 2009. 764 p.
6. Brunton S.L., Kurtz J.N. Data-Driven Science and engineering. Machine Leaning, Dynamical Systems, and Control. Cambrige: Cambrige university Press, 2022. 574 p. 102 ISSN 1028-9763. Математичні машини і системи. 2025. № 1
7. Müller J.A., Ivachnenko A.G. Selbstorganisation von vorhersage Modellen. Berlin: Verlag Technik, 1984. 222 p.
8. URL: https://uk.wikipedia.org/wiki/Метод_групового_урахування_аргументів.
9. URL: https://en.wikipedia.org/wiki/Group_method_of_data_handling.
10. Лапач С.М. Теорія планування експериментів: Виконання розрахунково-графічної роботи: навч. посіб. для студ. спеціальності 131 «Прикладна механіка», спеціалізації «Технологія машинобудування». Київ: КПІ ім. Ігоря Сікорського, 2020. 86 с. URL: https://ela.kpi.ua/handle/123456789/38858.
11. Лапач С.Н., Чубенко А.В., Бабич П.Н. Статистические методы в медико-биологических исследованиях с использованием Excel. 2 изд. перераб. и доп. К.: Морион, 2001. 408 с.
12. Лапач С.М., Радченко С.Г. Проблеми визначення структури рівняння регресії в множинному регресійному аналізі. Наукові вісті НТУУ «КПІ». 2007. № 1 (51). С. 150–155.
13. Draper N.R., Smith H. Applied regression Analysis. Third edition. New York: Wiley-Interscience, 1998. 736 p.
14. Hettmansperger T.P. Statistical inference based on ranks. New York: Wiley, 1984. 323 p.
15. Лапач С.М. Застосування непараметричної регресії при наявності викидів. Теорія ймовірностей та математична статистика: Шістнадцята міжнар. конф. ім. акад. Михайла Кравчука (м. Київ, 14–15 травня 2015 р.). К.: НТУУ «КПІ», 2015. Т. 3. С. 43–45.
16. Лапач С.М. Проблеми побудови регресійних моделей процесів різання металів. Вісник НТУУ «КПІ». Машнобудування. 2014. № 3 (72). С. 40–47.
17. Лапач С.М. Визначення викидів у кореляційному і одновимірному регресійному аналізі. Математичні машини і системи. 2019. № 4. С. 126–138.

