Enhancing test quality in Lesotho Basic Education through Classical Test and Item Response Theory Analyses
DOI:
https://doi.org/10.71291/pdrgvp19Keywords:
Educational measurement, Lesotho basic education, multiple-choice test development, Classical Test Theory, Item Response TheoryAbstract
This study developed and psychometrically evaluated a Grade 6 mathematics test for Lesotho’s Cambridge curriculum using Classical Test Theory (CTT) and Item Response Theory (IRT). A quantitative psychometric design involving 200 learners from Cambridge International Schools was employed. Data were analysed using jMetrik through CTT indices and 1PL and 2PL IRT models. Results revealed substantial psychometric weaknesses, including extreme item difficulty variation, weak and negative discrimination indices, low reliability coefficients (α = 0.1975), and several misfitting items. The 2PL model demonstrated a superior fit over the 1PL model based on deviance, AIC, and BIC statistics. While CTT identified global weaknesses in reliability and item discrimination, IRT provided richer diagnostic evidence regarding item functioning and conditional measurement precision across ability levels. The study concludes that integrating CTT and IRT strengthens assessment validity, reliability, fairness, and evidence-based test development in Lesotho basic education. The study recommends rigorous item piloting, routine use of CTT and IRT analyses, and enhanced psychometric training for teachers and assessment practitioners in Lesotho.
References
Abedalaziz, N., & Leng, C. H. (2013). The relationship between CTT and IRT approaches in analysing item characteristics. Malaysian Online Journal of Educational Sciences, 1(1), 64–70.
Ayanwale, M. A., Chere-Masopha, J., & Morena, M. C. (2022). The classical test or item response measurement theory: The status of the framework at the Examination Council of Lesotho. International Journal of Learning, Teaching and Educational Research, 21(8), 384–406. https://doi.org/10.26803/ijlter.21.8.22
Brzezińska, J. (2020). Computation item response theory models in the measurement theory. Communications in Statistics - Simulation and Computation, 49(12), 3299–3313. https://doi.org/10.1080/03610918.2018.1546399
Cappelleri, J. C., Lundy, J. J., & Hays, R. D. (2014). Overview of classical test theory and item response theory for quantitative assessment of items in developing patient-reported outcome measures. Science, 36(5), 648–662. https://doi.org/10.1016/j.clinthera.2014.04.006.Overview
Chang, C., Tseng, K., & Lou, S. (2012). A comparative analysis of the consistency and difference among in a web-based portfolio assessment environment for high school students. Computers and Education, 58(1), 303–320. https://doi.org/10.1016/j.compedu.2011.08.005
Chen, Y., Li, X., Liu, J., & Ying, Z. (2021). Item response theory – A statistical framework for educational and psychological measurement. Statistical Science, 40(2), 167–194.
Cook, L. L., & Pitoniak, M. J. (2025). Educational measurement (5th ed.). New York: Oxford University Press.
De Ayala, R. J. (2022). The theory and practice of item response theory. Change for the Better: Personal Development through Practical Psychotherapy, 77. https://doi.org/10.4135/9781526438447.n17
DeMars, C. (2010). Item response theory: Understandig statistics measurement. New York: Oxford University Press.
El-Hamamsy, L., Zapata-C´aceres, M., Martin-Barroso, E., Mondada, F., Zufferey, J. D., Bruno, B., & Roman-Gonzalez, M. (2023). The competent computational Thinking test (cCTt): A valid, reliable and gender-fair test for longitudinal CT studies in grades 3-6. Technology, Knowledge and Learning, 30(3), 1607–1661.
Embretson, S. E., & Reise, S. P. (2000). Item response theory for psychologists. New Jersey: Lawrence Erlbaum Associates.
Fu, J., Tan, X., & Kyllonen, P. C. (2024). Item and test characteristic curves of rank-2PL models for multidimensional forced-choice questionnaires. Applied Measurement in Education, 37(3), 272–288. https://doi.org/10.1080/08957347.2024.2386939.Item
Geleta, K. T., Fisseha, M., & Zenebe, N. (2024). The relationship between the psychometric and performance properties of teacher-made tests and students ’ academic performance in Ethiopian public universities: A baseline survey study. Cogent Education, 11(1), 2298049. https://doi.org/10.1080/2331186X.2023.2298049
Gordon, R. A. (2015). Measuring constructs in family science: How can item response theory improve precision and validity? Journal of Marriage and Family, 77(1), 147–176.
Hambleton, R. K., Swaminathan, H., Rogers, H. J., & Rogers, D. J. (1991). Fundamentals of item response theory. In Choice Reviews Online (Vol. 21). London: SAGE Publications. https://doi.org/10.2307/2075521
Helbling, L. A., Berger, S., & Verschoor, A. (2021). Flexibility at the price of volatility: Concurrent calibration in multistage tests in practice using a 2PL model. Frontiers in Education, 6, 679864. https://doi.org/10.3389/feduc.2021.679864
Hu, Z., Lin, L., Wang, Y., & Li, J. (2021). The integration of classical testing theory and item response theory. Psychology, 12(9), 1397–1409. https://doi.org/10.4236/psych.2021.129088
Huebner, A., & Skar, G. B. (2021). Conditional standard error of measurement: Classical test theory, generalizability theory and many-facet Rasch measurement with applications to writing assessment. Practical Assessment, Research, and Evaluation, 26(14), 1–21. https://doi.org/10.7275/vzmm-0z68
Idowu, O. I., Eluwa, A. N., & Abang, B. K. (2011). Evaluation of mathematics achievement test: A comparison between classical test theory( CTT) and item Response theory( IRT). Journal of Educational and Social Research, 1(4), 99–106.
Jumini, & Retnawati, H. (2022). Estimating item parameters and student abilities: An IRT 2PL analysis of mathematics examination. Al-Ishlah: Jurnal Pendidikan, 14(1), 385–398. https://doi.org/10.35445/alishlah.v14i1.926
Li, C., Lin, Y., Tosun, B., Wang, P., Ye Guo, H., Ling, C. R., … Zhang, L. (2025). Psychometric evaluation of the Chinese version of the BENEFITS-CCCSAT based on CTT and IRT : a cross-sectional design translation and validation study. Frontiers in Public Health, 13, 1532709. https://doi.org/10.3389/fpubh.2025.1532709
Magno, C. (2009). Demonstrating the difference between classical test theory and item response theory using derived test data. The International Journal of Educational and Psychological Assessment, 1(1), 1–11.
Meijer, R. R., Sijtsma, K., Smid, N. G., Philips, & Eindhoven. (1990). Theoretical and empirical comparison of the Mokken and the Rasch approach to IRT. Applied Psychological Measurement, 14(3), 283–298. https://doi.org/10.1177/014662169001400306
Mitee, T. L. (2019). Comparative study of classical test theory and item response theory using item analysis results of quantitative chemistry achievement test. AJB-SDR, 1(1), 26–36.
Mokken, R. J. (1971). A theory and procedure of scale analysis: With applications in political research. In Contemporary Sociology (1st ed.). Mouton: Mouton & Co. https://doi.org/10.2307/2062452
Reckase, M. D. (1979). Unifactor latent trait model applied to multifactor tests: Results and implications. Journal of Educational Statistics, 4(3), 207–230.
Reise, S. P., & Revicki, D. A. (2015). Handbook of item response theory modelling: Applications to typical performance assessment. In Handbook of Item Response Theory Modeling (1st ed.). New York: Routledge. https://doi.org/10.4324/9781315736013-27
Rezaee, R., Shafiayan, M., Jafari, P., & Zarifsanaiey, N. (2018). Invariance of item difficulty parameter estimates based on classical test theory and item response theory. Journal of Advance Pharmacy Education and Research, 8, 156-161., 8, 156–161.
Rusch, T., Lowry, P. B., Mair, P., & Treiblmaier, H. (2017). Breaking free from the limitations of classical test theory: Developing and measuring information systems scales using item response theory. Information and Management, 54(2), 189–203. https://doi.org/10.1016/j.im.2016.06.005
Santoso, P. H., Istiyono, E., Haryanto, & Retnawati, H. (2024). Validating light phenomena conceptual assessment through the lens of classical test theory and item response theory frameworks. Physics Education, 59(2), 1–19. https://doi.org/10.1088/1361-6552/ad183b
Schwarz, G. (1978). Estimating the dimension of model. The Annals of Statistics, 6(2), 461–464.
Setiawati, F. A., Amelia, R. N., Sumintono, B., & Purwanta, E. (2023). Study item parameters of classical and modern theory of differential aptitude test: Is it comparable. European Journal of Educational Research, 12(2), 1097–1107.
Siregar, H. N. I., & Panjaitan, A. (2021). A systematic literature review: The role of classical test theory and item response theory in item analysis to determine the quality of mathematics tests. Jurnal Fibonaci, 2(2), 29–48. Retrieved from https://doi.org/10.24114/jfi.v2i%0AJurnal
Tavakol, M., & Dennick, R. (2011). Making sense of Cronbach’s alpha. International Journal of Medical Education, 2, 53–55. https://doi.org/10.5116/ijme.4dfb.8dfd
Umobong, M., & Jacob, S. S. (2016). A comparison of classical and item response theory person/item parameters of physics achievement test for technical schools. African Journal of Theory and Practice of Educational Assessment, 4, 131.
Wongvorachan, T., & Bulut, O. (2025). Detecting construct-irrelevant variance: A comparison of network psychometrics and traditional psychometric methods using the HEXACO-PI dataset. Psychology International, 7(4), 88. Retrieved from https://doi.org/10.3390/ psycholint7040088
Yang, F. M., & Kao, S. T. (2014). Item Response Theory for measurement validity. Shanghai Archives of Psychiatry, 26(3), 171–177. https://doi.org/10.3969/j.issn.1002-0829.2014.03
Zakariya, Y. F. (2022). Cronbach’s alpha in mathematics education research: Its appropriateness, overuse, and alternatives in estimating scale reliability. Frontiers in Psychology, 13, 1–6. https://doi.org/10.3389/fpsyg.2022.1074430
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Lefa Clement Thamae (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.



