Forecasting Daily Online Attention to K-Pop Artists: A Retrospective Comparison of Baselines, Ridge Regression and Random Forests
Keywords:
K-pop, online attention, Wikipedia pageviews, time-series forecasting, random forest, ridge regression, chronological evaluationAbstract
Daily public attention measures offer a low-cost basis for audience monitoring, but forecast improvements must be assessed against strong simple methods and with chronological information boundaries. This study compares persistence, seasonal naive and moving-average baselines with ridge regression and random forests using English Wikipedia user-classified pageviews for BTS, BLACKPINK, TWICE, Stray Kids, SEVENTEEN and aespa. The primary dataset contains 10,404 artist-day observations from January 2022 to September 2026. Parameters were selected using eight expanding quarterly validation blocks in 2024–2025. Retrospective evaluation used 1,638 artist-days in January–September 2026, with models fixed before the first forecast origin and lagged inputs updated as observations became available. Log-target ridge regression achieved a mean absolute error of 512.50 pageviews, compared with 603.18 for persistence, a 15.03% reduction (paired 14-day block-bootstrap 95% interval: 6.14%–21.51%). The development-selected log-target forest achieved 551.13 pageviews, an 8.63% reduction (−2.95%–18.20%). Both improved the point-estimate error for all six primary artists, although monthly gains were uneven. Forest error was stable across five random seeds; removing the latest count raised MAE by 33.50%. Only 26 large-traffic artist-days accounted for 38.33% of forest absolute error. Seven-day forecasts were less accurate, and transfer to two artists excluded from fitting did not establish a reliable advantage over persistence. These findings support transparent, lightweight forecasting for aggregate attention monitoring while showing that nonlinear models need not outperform a well-specified linear comparator. Pageviews measure article visits, not unique fans, music consumption or campaign effects.
References
Wikimedia Foundation. Page view analytics. Wikimedia Analytics API documentation. Accessed 2 October 2026. https://doc.wikimedia.org/generated-data-platform/aqs/analytics-api/reference/page-views.html
Wikimedia Foundation. Access policy: data licensing and downloads. Wikimedia Analytics API documentation. Accessed 2 October 2026. https://doc.wikimedia.org/generated-data-platform/aqs/analytics-api/documentation/access-policy.html
Mestyán M, Yasseri T, Kertész J. Early prediction of movie box office success based on Wikipedia activity big data. PLOS ONE. 2013;8(8):e71226. https://doi.org/10.1371/journal.pone.0071226
Hoerl AE, Kennard RW. Ridge regression: biased estimation for nonorthogonal problems. Technometrics. 1970;12(1):55–67. https://doi.org/10.1080/00401706.1970.10488634
Breiman L. Random forests. Machine Learning. 2001;45:5–32. https://doi.org/10.1023/A:1010933404324
Cerqueira V, Torgo L, Mozetič I. Evaluating time series forecasting models: an empirical study on performance estimation methods. Machine Learning. 2020;109:1997–2028. https://doi.org/10.1007/s10994-020-05910-7
Hewamalage H, Ackermann K, Bergmeir C. Forecast evaluation for data scientists: common pitfalls and best practices. Data Mining and Knowledge Discovery. 2023;37:788–832. https://doi.org/10.1007/s10618-022-00894-5
Gneiting T. Making and evaluating point forecasts. Journal of the American Statistical Association. 2011;106(494):746–762. https://doi.org/10.1198/jasa.2011.r10138
Hyndman RJ, Koehler AB. Another look at measures of forecast accuracy. International Journal of Forecasting. 2006;22(4):679–688. https://doi.org/10.1016/j.ijforecast.2006.03.001
Künsch HR. The jackknife and the bootstrap for general stationary observations. The Annals of Statistics. 1989;17(3):1217–1241. https://doi.org/10.1214/aos/1176347265
Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns. 2023;4(9):100804. https://doi.org/10.1016/j.patter.2023.100804
Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: machine learning in Python. Journal of Machine Learning Research. 2011;12:2825–2830. https://jmlr.org/papers/v12/pedregosa11a.html
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Frankie Gao

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Authors retain copyright and grant Perspective permission to publish their work as the original journal of publication. Accepted articles are made freely available under the Creative Commons Attribution–NonCommercial 4.0 International license (CC BY-NC 4.0).
The license permits non-commercial sharing and adaptation with appropriate attribution, a license link and identification of changes. Reuse beyond the license requires permission unless otherwise permitted by law. Follow any separate license or credit line for third-party material. Read the open access, copyright and fees policy for details.