Analysis of Direct Scoring and Similarity-Based Scoring Approaches in Automatic Short Answer Scoring (ASAS)
DOI:
https://doi.org/10.47709/brilliance.v5i1.6275Keywords:
Automatic Short Answer Scoring, Cross-Prompt, Direct Scoring, Outlier, Similarity-Based Scoring, Specific-PromptAbstract
In the era of digital education, the need for automated scoring systems for short text answers has been steadily increasing. Automatic Short Answer Scoring (ASAS) aims to automate this assessment process with efficient and consistent approaches. Two commonly used approaches in ASAS are direct scoring and similarity-based scoring. Although these two approaches have been widely used, previous research has mostly focused on metrics like RMSE and Pearson Correlation to assess model performance. This study aims to provide a more in-depth analysis by comparing both approaches in two evaluation scenarios, specific-prompt and cross-prompt, by evaluating the accuracy and stability of the models. The dataset used in this study is the Rahutomo dataset. The results of the analysis show that direct scoring outperforms similarity-based scoring in terms of lower RMSE, higher Pearson Correlation, and fewer outliers. In the specific-prompt scenario, an RMSE of 0.0817 and a Pearson Correlation of 0.9504 were obtained, while in the cross-prompt scenario, the RMSE was 0.0917 and the Pearson Correlation was 0.9286. This study provides a more comprehensive insight into model performance by not only relying on evaluation metrics but also examining the distribution of residuals and outliers, which offers a more complete picture of model stability. Based on these findings, direct scoring is recommended for implementation in ASAS systems and for future research that can extend the analysis to other datasets or languages.
References
Ahmed, A., Joorabchi, A., & Hayes, M. J. (2022). On the Application of Sentence Transformers to Automatic Short Answer Grading in Blended Assessment. 2022 33rd Irish Signals and Systems Conference, ISSC 2022, June 2022, 1–6. https://doi.org/10.1109/ISSC55427.2022.9826194
Amin, M., Afzal, S., Akram, M. N., Muse, A. H., Tolba, A. H., & Abushal, T. A. (2022). Outlier detection in gamma regression using Pearson residuals: Simulation and an application. AIMS Mathematics, 7(8), 15331–15347. https://doi.org/10.3934/math.2022840
Chamidah, N., Yulianti, E., & Budi, I. (2023). Evaluating the Impact of Sentence Tokenization on Indonesian Automated Essay Scoring Using Pretrained Sentence Embeddings. Revue d’Intelligence Artificielle, 37(5), 1101–1108. https://doi.org/10.18280/ria.370502
Devlin, J., Chang, M.-W., Lee, K., Google, K. T., & Language, A. I. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Naacl-Hlt 2019, Mlm, 4171–4186. https://aclanthology.org/N19-1423.pdf
Ferreira Mello, R., Pereira Junior, C., Rodrigues, L., Pereira, F. D., Cabral, L., Costa, N., Ramalho, G., & Gasevic, D. (2025). Automatic Short Answer Grading in the LLM Era: Does GPT-4 with Prompt Engineering beat Traditional Models? 15th International Conference on Learning Analytics and Knowledge, LAK 2025, 93–103. https://doi.org/10.1145/3706468.3706481
Haidir, M. H., & Purwarianti, A. (2020). Short Answer Grading Using Contextual Word Embedding and Linear Regression. Jurnal Linguistik Komputasional, 3(2), 54–61. https://inacl.id/journal/index.php/jlk/article/view/38
Kaya, M., & Cicekli, I. (2024). A Hybrid Approach for Automated Short Answer Grading. IEEE Access, 12(May), 96332–96341. https://doi.org/10.1109/ACCESS.2024.3420890
Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP. COLING 2020 - 28th International Conference on Computational Linguistics, Proceedings of the Conference, 757–770. https://doi.org/10.18653/v1/2020.coling-main.66
Mardini G, I. D., Quintero M, C. G., Viloria N, C. A., Percybrooks B, W. S., Robles N, H. S., & Villalba R, K. (2024). A deep-learning-based grading system (ASAG) for reading comprehension assessment by using aphorisms as open-answer-questions. Education and Information Technologies, 29(4), 4565–4590. https://doi.org/10.1007/s10639-023-11890-7
Rahutomo, F., Ari Roshinta, T., Rohadi, E., Siradjuddin, I., Ariyanto, R., Setiawan, A., & Adhisuwignjo, S. (2018). Open Problems in Indonesian Automatic Essay Scoring System. International Journal of Engineering & Technology, 7(4.44), 156. https://doi.org/10.14419/ijet.v7i4.44.26974
Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using siamese BERT-networks. EMNLP-IJCNLP 2019 - 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing, Proceedings of the Conference, 3982–3992. https://doi.org/10.18653/v1/d19-1410
Salim, H. R., De, C., Pratamaputra, N. D., & Suhartono, D. (2022). Indonesian automatic short answer grading system. Bulletin of Electrical Engineering and Informatics, 11(3), 1586–1603. https://doi.org/10.11591/eei.v11i3.3531
Santoso, R. R., Megasari, R., & Hambali, Y. A. (2020). Implementasi Metode Machinelearning. Jurnal Aplikasi Dan Teori Ilmu Komputer, 3(2), 85–97. https://ejournal.upi.edu/index.php/JATIKOM
Verma, V. (2025). A Comprehensive Framework for Residual Analysis in Regression and Machine Learning. January. https://doi.org/10.52783/jisem.v10i31s.4958
Wijanto, M. C., & Yong, H. S. (2024). Combining Balancing Dataset and SentenceTransformers to Improve Short Answer Grading Performance. Applied Sciences (Switzerland), 14(11). https://doi.org/10.3390/app14114532
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Bayu Wicaksono, Rasim, Yaya Wihardi

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.















