Albanese, M. A., & Sabers, D. L. (1988). Multiple true–false items: A study of interitem correlations, scoring alternatives, and reliability estimation. Journal of Educational Measurement, 25(2), 111–123.
Article
Google Scholar
Baldiga, K. (2013). Gender differences in willingness to guess. Management Science, 60(2), 434–448.
Article
Google Scholar
Bauer, D., Holzer, M., Kopp, V., & Fischer, M. R. (2011). Pick-N multiple choice-exams: A comparison of scoring algorithms. Advances in Health Sciences Education, 16(2), 211–221.
Article
Google Scholar
Case, S. M., & Swanson, D. B. (2002). Constructing written test questions for the basic and clinical sciences (3rd ed.). Philadelphia, PA: National Board of Medical Examiners.
Google Scholar
Cook, D. A., Brydges, R., Ginsburg, S., & Hatala, R. (2015). A contemporary approach to validity arguments: A practical guide to Kane’s framework. Medical Education, 49(6), 560–575. https://doi.org/10.1111/medu.12678.
Article
Google Scholar
Cronbach, L. (1939). Note on the multiple true–false test exercise. Journal of Educational Psychology, 30(8), 628.
Article
Google Scholar
Cronbach, L. (1941). An experimental comparison of the multiple true–false and multiple multiple-choice tests. Journal of Educational Psychology, 32(7), 533.
Article
Google Scholar
Downing, S. M., & Yudkowsky, R. (2009). Assessment in health professions education. New York: Routledge.
Google Scholar
Dudley, A. (2006). Multiple dichotomous-scored items in second language testing: Investigating the multiple true–false item type under norm-referenced conditions. Language Testing, 23(2), 198–228.
Article
Google Scholar
Dunham, M. L. (2006). An investigation of the multiple true–false item for nursing licensure and potential sources of construct-irrelevant difficulty. ProQuest.
Frisbie, D. A., & Sweeney, D. C. (1982). The relative merits of multiple true–false achievement tests. Journal of Educational Measurement, 19(1), 29–35. https://doi.org/10.2307/1434916.
Article
Google Scholar
Gross, L. J. (1982). Scoring multiple true/false tests some considerations. Evaluation and the Health Professions, 5(4), 459–468.
Article
Google Scholar
Guttormsen, S., Beyeler, C., Bonvin, R., Feller, S., Schirlo, C., Schnabel, K., et al. (2013). The new licencing examination for human medicine: From concept to implementation. Swiss Medical Weekly, 143, w13897. https://doi.org/10.4414/smw.2013.13897.
Google Scholar
Haladyna, T. M., Downing, S. M., & Rodriguez, M. C. (2002). A review of multiple-choice item-writing guidelines for classroom assessment. Applied Measurement in Education, 15(3), 309–333.
Article
Google Scholar
Itten, S., & Krebs, R. (1997). Messqualität der verschiedenen MC-Itemtypen in den beiden Vorprüfungen des Medizinstudiums an der Universität Bern 1997/2 (Forschungsbericht Institut für Aus-, Weiter-und Fortbildung (IAWF) der medizinischen Fakultät der Universität Bern). Bern: IAWF.
Google Scholar
Javid, L. (2014). The comparison between multiple-choice (MC) and multiple true–false (MTF) test formats in iranian intermediate EFL learners’ vocabulary learning. Procedia-Social and Behavioral Sciences, 98, 784–788.
Article
Google Scholar
Krebs, R. (1997). The swiss way to score multiple true–false items: theoretical and empirical evidence. In A. J. J. A. Scherpbier, C. P. M. van der Vleuten, J. J. Rethans, & A. F. W. van der Steeg (Eds.), Advances in medical education (pp. 158–161). Netherlands: Springer.
Chapter
Google Scholar
Krebs, R. (2004). Anleitung zur Herstellung von MC-Fragen und MC-Prüfungen für die ärztliche Ausbildung. Bern: Institut für Medizinische Lehre IML, Abteilung für Ausbildungs-und Examensforschung AAE.
Google Scholar
Kreiter, C. D., & Frisbie, D. A. (1989). Effectiveness of multiple true–false items. Applied Measurement in Education, 2(3), 207–216.
Article
Google Scholar
Mobalegh, A., & Barati, H. (2012). Multiple true–false (MTF) and multiple-choice (MC) test formats: A comparison between two versions of the same test paper of Iranian NUEE. Journal of Language Teaching and Research, 3(5), 1027–1037.
Article
Google Scholar
Muchinsky, P. M. (1996). The correction for attenuation. Educational and Psychological Measurement, 56(1), 63–75.
Article
Google Scholar
Norman, G. R., Swanson, D. B., & Case, S. M. (1996). Conceptual and methodological issues in studies comparing assessment formats. Teaching and Learning in Medicine: An International Journal, 8(4), 208–216.
Article
Google Scholar
R Core Team. (2013). R: A language and environment for statistical computing. Vienna, Austria. Retrieved from http://www.r-project.org/.
Ravesloot, C., Van der Schaaf, M., Muijtjens, A., Haaring, C., Kruitwagen, C., Beek, F., et al. (2015). The don’t know option in progress testing. Advances in Health Sciences Education, 20(5), 1325–1338.
Article
Google Scholar
Richardson, R. (1992). The multiple choice true/false question: What does it measure and what could it measure? Medical Teacher, 14(2–3), 201–204.
Article
Google Scholar
Romano, J. L., Kromrey, J. D., & Hibbard, S. T. (2010). A Monte Carlo study of eight confidence interval methods for coefficient alpha. Educational and Psychological Measurement, 70, 376–393.
Article
Google Scholar
Siddiqui, N. I., Bhavsar, V. H., Bhavsar, A. V., & Bose, S. (2016). Contemplation on marking scheme for Type X multiple choice questions, and an illustration of a practically applicable scheme. Indian Journal of Pharmacology, 48(2), 114.
Article
Google Scholar
Tarasowa, D., & Auer, S. (2013). Balanced scoring method for multiple-mark questions. Paper presented at the CSEDU.
Tsai, F.-J., & Suen, H. K. (1993). A brief report on a comparison of six scoring methods for multiple true–false items. Educational and Psychological Measurement, 53(2), 399–404.
Article
Google Scholar
Verbić, S. (2012). Information value of multiple response questions. Psihologija, 45(4), 467–485.
Article
Google Scholar
Wu, B. C. (2003). Scoring multiple true false items: A comparison of summed scores and response pattern scores at item and test levels. Retrieved from Eric: https://eric.ed.gov/?id=ED476148