Prediction of Investment Decision in Indonesian Conglomerate Stock Using Machine Learning with K-Means and XGBoost
Nanda Ammaa Tsuroyya Rahmah
*
Faculty of Economics and Business, University of Muhammadiyah Malang, Malang, Indonesia.
Novi Puji Lestari
Faculty of Economics and Business, University of Muhammadiyah Malang, Malang, Indonesia.
Mursidi
Faculty of Economics and Business, University of Muhammadiyah Malang, Malang, Indonesia.
*Author to whom correspondence should be addressed.
Abstract
The growing number of investors in the Indonesian capital market highlights the need for analytical methods that support accurate investment decision-making. Conglomerate stocks have complex characteristics due to business diversification across sectors, making their price movements more difficult to predict than stocks within a single sector. This complexity creates challenges for investors in determining accurate investment decisions because stock movement patterns are heterogeneous and may be difficult to model effectively. This study aims to develop and evaluate a model for predicting Buy, Hold, and Sell investment decisions in Indonesian conglomerate stocks through the integration of K-Means and XGBoost algorithms. The data consist of weekly historical data from 11 conglomerate companies during the 2020–2025 period. The research stages include data preprocessing, formation of technical indicators, clustering using K-Means, and classification of investment decisions using XGBoost. The dataset was divided into 80% training data and 20% testing data. Investment decision labels were determined using a predefined ±5% weekly return threshold, where returns above 5% were classified as Buy, returns between −5% and 5% were classified as Hold, and returns below −5% were classified as Sell. The results show that the addition of clustering features improves the performance of the classification model. The XGBoost model without clustering achieved an accuracy of 73.41%, while the combination of XGBoost and K-Means with four clusters achieved the highest accuracy of 81.69%, with a macro F1-score of 0.70, compared with a 69.8% majority-class baseline. These findings indicate that integrating clustering and machine learning methods can improve the prediction of investment decisions in Indonesian conglomerate stocks. This study contributes by demonstrating that using K-Means as a feature engineering stage before XGBoost classification improves prediction accuracy compared with an XGBoost model without clustering, providing a more effective approach to support investment decision-making in conglomerate stocks.
Keywords: Indonesian conglomerate stocks, investment decision prediction, machine learning, K-means clustering, XGBoost