Evaluating the quality of pediatric multiple-choice questions: a ten-year journey through docimology at Cadi Ayyad University
Mohamed-Amine Ben-Rrahilya, Meriem Elbaz
Corresponding author: Mohamed-Amine Ben-Rrahilya, Faculty of Medicine and Pharmacy in Marrakech, Cadi Ayyad University, Marrakech, Morocco 
Received: 15 Oct 2024 - Accepted: 13 Aug 2026 - Published: 14 Sep 2026
Domain: Health education
Keywords: Multiple-choice question, Bloom´s taxonomy, cognition, item-writing flaws, item analysis
Funding: This work received no specific grant from any funding agency in the public, commercial, or non-profit sectors.
©Mohamed-Amine Ben-Rrahilya et al. Pan African Medical Journal (ISSN: 1937-8688). This is an Open Access article distributed under the terms of the Creative Commons Attribution International 4.0 License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Cite this article: Mohamed-Amine Ben-Rrahilya et al. Evaluating the quality of pediatric multiple-choice questions: a ten-year journey through docimology at Cadi Ayyad University. Pan African Medical Journal. 2026;55:20. [doi: 10.11604/pamj.2026.55.20.45621]
Available online at: https://www.panafrican-med-journal.com//content/article/55/20/full
Research 
Evaluating the quality of pediatric multiple-choice questions: a ten-year journey through docimology at Cadi Ayyad University
Evaluating the quality of pediatric multiple-choice questions: a ten-year journey through docimology at Cadi Ayyad University
Mohamed-Amine Ben-Rrahilya1,&, Meriem Elbaz1
&Corresponding author
Introduction: multiple-choice questions (MCQs) are a widely used tool in medical education for their validity, reliability, and efficiency in assessing student knowledge. However, concerns remain regarding their ability to test higher-order cognitive functions and the presence of item-writing flaws (IWFs), which can compromise their quality and fairness. This study aimed to conduct a docimological evaluation of MCQs used in pediatric pathology exams at the Faculty of Medicine, Cadi Ayyad University, Marrakech, over a ten-year period, focusing on cognitive level distribution, prevalence of IWFs, and psychometric properties through item analysis.
Methods: retrospective cross-sectional study was conducted at the Faculty of Medicine, Cadi Ayyad University, Marrakech, Morocco. A total of 1,100 MCQs from 22 pediatric exams conducted between 2013 and 2023 were evaluated. MCQs were classified according to an adapted Bloom's taxonomy into recall/memorization, comprehension/interpretation, and problem-solving levels. IWFs were identified using a 17-criteria rubric. Due to data availability constraints, the difficulty index (DIF) and discrimination index (DI) were calculated for a subset of 500 MCQs from exams between 2019 and 2023.
Results: of the 1,100 MCQs analyzed, 66.36% targeted the recall/memorization level, 18.55% tested comprehension/interpretation, and 15.09% addressed problem-solving. The overall difficulty index averaged 0.57 ± 0.21, indicating moderate difficulty. However, 9.27% of MCQs exhibited IWFs, the most frequent being "multiple concepts within a single option". The discrimination index was excellent for 85.33% of the MCQs, indicating strong discriminative capacity.
Conclusion: despite limitations such as data constraints and potential observer bias, the study found that the majority of MCQs had acceptable difficulty and excellent discrimination, effectively distinguishing between varying levels of student performance. However, the over-reliance on low-cognitive-level items underscores the need for faculty development in constructing MCQs that assess higher-order thinking skills. Enhancing MCQ quality can improve assessment validity and better prepare students for clinical practice.
Multiple-choice questions (MCQs) are one of the most widely adopted assessment tools in medical and health disciplines since their incorporation into medical examinations in the 1950s [1]. Each MCQ, also known as an item, consists of a problem statement, or stem, and several answer options, typically including one correct answer and several distractors intended to appear plausible to less knowledgeable students [1,2].
Well-constructed MCQs are valued for their validity, reliability, and efficiency in assessing a broad range of knowledge in large groups of students [3,4]. They facilitate the evaluation of candidates with minimal human intervention and are noted for their standardization and ease of administration [3]. However, developing high-quality MCQs is time-consuming, often requiring about an hour per question [5]. Moreover, MCQs have been criticized for focusing on factual recall rather than assessing higher-order cognitive skills essential in the health sciences [6]. Assessments that test advanced cognitive functions are crucial for preparing professionals to handle complex information and make critical decisions in clinical settings [7].
Another significant concern is the prevalence of item-writing flaws (IWFs), defined as violations of established item-writing guidelines [6,8]. Studies indicate that 50% to 75% of MCQs in health professions education contain IWFs [9], which can negatively impact the validity of assessments and potentially alter student performance outcomes [6,10]. Eliminating flawed questions could reclassify 10-15% of students from failing to passing [10], highlighting the importance of quality control in MCQ construction.
To address these challenges, docimology, the study of assessment methods, has emerged as a methodological approach to enhance the quality of examinations [11,12]. Docimology originated in the 1920s with the work of Henri Piéron and Henri Laugier, as described by Leclercq et al. [11]; it focuses on analyzing exams to identify biases and improve assessment practices. Tools such as cognitive level classification, item analysis, and IWF identification are employed to ensure that exams measure intended skills effectively and discriminate between varying levels of student understanding [13,14].
In the context of our institution, the Faculty of Medicine at Cadi Ayyad University in Marrakech, there has been limited evaluation of the quality of MCQs used in the pediatrics module. Given the critical role of pediatrics in medical education and the potential impact of assessment quality on student learning and competency, it is essential to examine and enhance the quality of these assessments. Therefore, the objective of this study was to conduct a docimological evaluation of MCQs used in the Pediatrics module exams over a ten-year period (2013-2023) at our faculty. Specifically, we aimed to: analyze the distribution of cognitive levels assessed by the MCQs according to Bloom's taxonomy; determine the prevalence and types of item-writing flaws present in the MCQs; assess the psychometric properties of the MCQs through item analysis, including difficulty and discrimination indices.
Study design: a retrospective cross-sectional analysis of MCQs used in pediatrics module exams over a ten-year period.
Study setting: the study was conducted at the Faculty of Medicine, Cadi Ayyad University in Marrakech, Morocco. The pediatrics module is a core component of the undergraduate medical curriculum, undertaken by fourth-year medical students. The module includes theoretical courses assessed through MCQ exams administered at the end of each academic session. Inclusion criteria encompassed all MCQs administered in the pediatrics module exams between 2013 and 2023. Exclusion criteria included any assessment formats other than MCQs, such as essay questions or oral exams. A total of 22 exams were included, representing both first and second session exams over the ten-year period.
Data sources: data were obtained from the Department of Courses and Examinations at the Faculty of Medicine Marrakech (FMPM). The department provided copies of the pediatrics module exams and anonymized student performance data for the study period. The exams and results were securely transferred to the research team, ensuring the confidentiality and integrity of the data.
Variables and measurement
Cognitive level classification
Rationale for using Bloom's taxonomy: the cognitive level of each MCQ was classified using an adapted version of Bloom's taxonomy. Bloom's taxonomy is a widely recognized framework for categorizing educational objectives based on cognitive complexity [15]. However, determining the appropriate level within Bloom’s taxonomy can be challenging due to interpretative variability among evaluators [16]. This variability can lead to inconsistencies when "Bloomizing" exam questions, especially with multiple observers.
Strategies to enhance consistency
Normalization session between observers: before the full-scale analysis, a normalization session was conducted. During this session, the two independent observers reviewed and discussed a sample set of exam questions to align their interpretations of Bloom's levels. This process aimed to ensure consistent categorization based on a shared understanding of the taxonomy. This method has been successfully employed in previous research to improve intra-study reliability [17,18].
Condensed three-level cognitive framework: to further enhance reliability, we applied a condensed three-level cognitive framework rather than the full six-level Bloom's taxonomy. Specifically, we adopted the recall/interpretation/problem-solving taxonomy used in the analysis of in-training examinations by Frassica et al. [19] and subsequently by Murphy et al. [20]. In that literature, this cognitive taxonomy is applied independently of the content categories used to describe question topics (five categories in Frassica et al. [19] and six in Murphy et al. [20]); only the cognitive axis was retained here. The levels were defined as:
Recall/memorization questions (taxonomy I): these questions required the student to recall specific facts about the tested entity. In addition, they tested the student's ability to immediately recall knowledge on various treatment modalities;
Comprehension/interpretation questions (taxonomy II): these questions aimed to determine a diagnosis. They often provided a brief history, clinical data, simple X-rays, and blood test results. The student had to interpret the data and select a diagnosis from five choices;
Problem-solving questions (taxonomy III): this type of question required the highest level of cognitive performance and asked the student to choose a treatment method after interpreting the provided data (e.g., imaging studies). Similarly, evaluation/decision-making questions required the student to choose the next step in management, even when they could not establish a diagnosis based on the available data. This approach minimizes the risk of evaluator disagreement while maintaining meaningful distinctions between cognitive levels [16,21].
Classification process: each MCQ was independently classified by the two observers according to the adapted taxonomy. Any discrepancies were resolved through discussion until consensus was reached. Examples of MCQs at each taxonomy level are provided in Table 1.
Item-writing flaws identification
Rationale for assessing IWFs: IWFs are violations of established guidelines for constructing MCQs and can adversely affect the validity and reliability of assessments [6,8]. Identifying and categorizing IWFs is essential for understanding the quality of the exam items and for informing improvements.
Development of the IWFs rubric: we developed a comprehensive rubric consisting of 18 criteria to systematically identify and categorize IWFs in the MCQs (Table 2). This rubric was derived from established MCQ writing standards and common violations documented in previous studies [22,23]. Each criterion addressed specific aspects of item construction that could potentially compromise question validity or effectiveness.
Evaluation process: both observers independently evaluated each MCQ for the presence of IWFs using the rubric. Discrepancies were discussed, and consensus was reached for all items.
Item analysis
Rationale for item analysis: item analysis provides statistical information about the quality of test items, specifically their difficulty and ability to discriminate between high-performing and low-performing students [13,14]. This analysis is crucial for evaluating the effectiveness of MCQs and for identifying items that may need revision.
Data collection for item analysis: for each exam, we obtained paper copies of the tests, the corresponding answer keys, and an Excel file containing anonymized scores each student achieved on each question. Due to data availability constraints, detailed student performance data necessary for item analysis were only available for exams conducted between 2019 and 2023.
Calculation of indices
Difficulty index (DIF):
Definition: reflects the proportion of students who answered an item correctly.
Calculation: **DIF = C/T**. Where: C is the number of students who answered the question correctly; T is the total number of students who attempted the question.
Interpretation: items were classified based on DIF into: very easy: DIF ≥ 0.80; easy: 0.70 ≤ DIF < 0.80; acceptable: 0.30 ≤ DIF < 0.70; difficult: DIF < 0.30.
Discrimination index (DI):
Definition: measures an item's ability to distinguish between high-performing and low-performing students.
Calculation: DI=(H-L/N) *2. Where H: number of high-performing students (top 27%) who answered the item correctly; L: number of low-performing students (bottom 27%) who answered the item correctly; N: total number of students in both groups.
Interpretation: DI values were categorized as: poor discrimination: DI ≤ 0.15; acceptable discrimination: 0.15 < DI ≤ 0.25; excellent discrimination: DI > 0.25.
Rationale for using the 27% groups: using the upper and lower 27% of students is a standard psychometric practice that optimizes the discrimination power of the analysis [24]. This method balances the need for sufficient group sizes while maximizing the differences between high and low performers.
Limitations due to sample size: only 500 MCQs from 10 exams (2019-2023) were included in the item analysis due to limited availability of retrospective data. This constraint may affect the generalizability of the findings to the entire ten-year period.
Data sources: the Department of Courses and Examinations at the Faculty of Medicine Marrakech provided: copies of the pediatrics module exams; anonymized student performance data for the study period; data were securely transferred, ensuring confidentiality and integrity.
Measurement techniques: normalization sessions were conducted to standardize cognitive level classification; rubric-based evaluations were used for identifying IWFs; statistical formulas were applied for calculating DIF and DI.
Statistical analysis: data were entered and analyzed using Microsoft Excel 2016. Descriptive statistics, including frequencies, percentages, means, and standard deviations, were calculated. For DIF and DI, means and standard deviations were computed for each exam session; the overall values reported in Table 4 and Table 5 are the mean and standard deviation of all 500 individual item indices, the standard deviation therefore reflecting variation both within and between examination sessions. No inferential statistical tests were performed, as the study was primarily descriptive in nature. Results were presented in tables to enhance clarity.
Bias: to address potential sources of bias: observer bias: conducting a normalization session aimed to reduce variability in cognitive level classification between observers; selection bias: acknowledged the limitation of only having item analysis data for a subset of exams, which may affect representativeness; information bias: ensured data accuracy by cross-checking student performance data with official records.
Ethical considerations: the study was conducted on anonymized institutional examination records and did not involve any intervention on human subjects. Access to and examination of these data were authorized by the Department of Courses and Examinations of the Faculty of Medicine and Pharmacy, Cadi Ayyad University, Marrakech. All data collected were anonymized to protect student confidentiality. As the study involved retrospective analysis of existing data without direct involvement of human subjects, informed consent was not required. We adhered to the principles of privacy, transparency, and integrity to ensure ethical compliance and respect for data usage.
Study sample: a total of 22 pediatrics module exams conducted between 2013 and 2023 were analyzed, comprising 1,100 MCQs. All MCQs administered during this period met the inclusion criteria; no MCQs were excluded from the analysis.
Cognitive level classification: in terms of taxonomy, out of a total of 1,100 MCQs, 730 (66.36%) were in taxonomy level I, 204 (18.55%) were in taxonomy level II, and 166 (15.09%) were in taxonomy level III. The distribution of MCQs across taxonomy levels for each exam session is detailed in Table 3. Over the ten-year period, the proportion of higher-order cognitive level questions (taxonomy II and III) remained relatively consistent, with slight variations between sessions.
Item-writing flaws: of the 1,100 MCQs, there were 102 (9.27%) flawed questions. The most common violation of item-writing guidelines was the inclusion of "multiple concepts within a single option," noted in 68 cases. Other violations included "the longest option is the correct one" in 20 cases (1.8%), and the use of "all of the above" in 2 cases (0.18%) and "none of the above" in 2 cases (0.18%).
Item analysis: due to the availability of student performance data, the difficulty index (DIF) and discrimination index (DI) were calculated for a subset of 500 MCQs from 10 exams conducted between 2019 and 2023.
Difficulty index (DIF): the mean DIF was 0.57 with a standard deviation (SD) of ± 0.21, indicating acceptable difficulty overall. When classified based on DIF values, 104 MCQs (20.8%) were considered difficult (DIF < 0.30), 267 MCQs (53.4%) fell into the acceptable range (0.30 ≤ DIF ≤ 0.69), and 129 MCQs (25.8%) were categorized as easy (DIF ≥ 0.70). The session-wise DIF values are detailed in Table 4. Although there was variability in difficulty levels across different exam sessions, no consistent trend was observed over the years.
Discrimination index (DI): the mean DI was 0.27 with an SD of ± 0.21. Based on DI classification, 427 MCQs (85.4%) demonstrated excellent discrimination (DI > 0.25), 41 MCQs (8.2%) had acceptable discrimination (0.15 < DI ≤ 0.25), and 32 MCQs (6.4%) showed poor discrimination (DI ≤ 0.15). Detailed DI values for each exam session are presented in Table 5.
The ultimate goal of creating standard and high-quality peer-reviewed MCQ items requires not only training but consistent practice. These standards can be achieved through rigorous MCQ quality evaluation or docimology. Such evaluations can pinpoint areas for improvement, ensuring that assessments are both fair and effective [25]. In terms of taxonomy, the majority of MCQs were at taxonomy level I (66.36%), followed by level II (18.55%) and level III (15.09%). This finding aligns with Baig et al. [26] who reported that 76% of MCQs targeted recall of isolated facts, while 24% assessed data interpretation skills. Similar trends were observed by Marzieh Nojomi et al. [27], who found that 50.7% of MCQs required information recall.
The high percentage of MCQs in these studies that tested low cognitive levels could be due to the idea that these questions are simpler to create, less time-consuming, and require less specialized knowledge compared to questions that assess higher-order cognitive skills. In contrast, developing questions that require data synthesis and critical thinking demands more effort, time, and training, making them more challenging to produce. In the current study, the low cognitive levels of the MCQs can also be attributed to the fact that the sample for this study came from a test bank of MCQs designed for fourth-year medical school students. This group is relatively new to many medical concepts, so it is crucial for them to recall a significant amount of foundational information. The emphasis on the first cognitive level in Bloom's taxonomy, which focuses on recall and recognition, aligns with the educational needs of these students as they build their medical knowledge base. Although no consensus has been established regarding the optimal share of items that should target each cognitive level [7], our findings point to a clear need to strengthen the design of the assessment tools used in this module. When an examination is dominated by items rewarding memorisation, the inferences that can be drawn from the resulting scores are correspondingly weakened, and students are given an incentive to adopt superficial revision strategies that serve them poorly beyond the examination itself. In our study, the prevalence of IWFs was 102 out of 1,100 questions, representing 9.27%. The frequency of IWFs found in MCQs in this study is close to that reported by Khan et al. [28], who evaluated IWFs in a medical school examination in which 12% of MCQs contained IWFs. The proportion of flawed questions in this study is, however, well below the 21% to 67% range reported by Fayyaz et al. [29].
The issue of IWFs is not exclusive to medical schools. A similar pattern is observed in pharmacy education, where Dell et al. [9], in their guide for pharmacy instructors, summarise the available evidence across the health professions as placing the proportion of flawed items between 50% and 75%. In dentistry schools, Kowash et al. [30] reviewed 185 single-best-answer items from two postgraduate pediatric dentistry examinations and identified 92 item-writing flaws, a figure corresponding to 49.7% of the number of items analysed. As a single item may carry more than one flaw, this represents a count of flaws rather than a proportion of flawed questions, and is therefore not directly comparable with the item-level prevalence reported here. Nursing schools also face similar issues, as research by Hijji [31], focusing on teacher-constructed nursing examinations, revealed that 91.8% of the MCQs examined contained one or more IWFs. In medical education, Downing [10] reported that between 36% and 65% of MCQs contained item-writing flaws. This broad prevalence of IWFs across different healthcare disciplines underscores the need for enhanced quality assurance, peer review, and faculty training to improve the reliability and validity of MCQs in educational assessments.
In this study, the most frequent IWF was "multiple concepts within a single option," with 68 questions exhibiting this issue. Several of the flaws identified belong to a group that allows a student to reach the correct answer from features of the item itself rather than from knowledge of the subject. A correct option that is longer or more explicit than the distractors, repetition of stem wording in the key, mutually exclusive options, "all of the above" and "none of the above" formulations, and absolutist adverbs all narrow the field of plausible answers for a test-wise candidate, and in doing so weaken the inference that a correct response reflects mastery. Pham et al. [32] reported that only four of the ten types of IWFs they examined had the hypothesized effect on mean item scores, the remaining six having either the opposite effect or no significant effect, and that no impact of IWFs on the difficulty or discrimination indices was demonstrated. They concluded that the effect of IWFs is neither systematic nor predictable, and that this unpredictability itself poses a risk to test validity.
In terms of item analysis psychometrics, out of 500 MCQs, 53.4% had good or acceptable difficulty levels, while 20.8% were very difficult, and 25.9% were easy. Christian et al. [33] evaluated 200 MCQs administered to interns during compulsory rotating postings and found that 93 (46.5%) had an acceptable difficulty range (P = 30-70%), while 33 (16.5%) were too easy (P > 70%), and 74 (37%) were too difficult (P < 30%). In a study by Patil et al. [13], 16.7% of items were considered easy, 46.6% had good or acceptable difficulty levels, and 36.7% were highly difficult.. This supports our findings and emphasizes the need for continuous evaluation and revision of MCQs to ensure they adequately challenge students without being overly difficult or too simplistic. Similarly, Rao et al. [34] classified 85% of their 40 items as being of acceptable difficulty, 10% as difficult, and 5% as easy.
As for the discrimination index (DI), the average DI was 0.27± 21. Within this dataset, 427 MCQs (85.33%) had a high discrimination index (DI > 0.25), 41 MCQs (8.21%) had moderate discrimination (0.15 < DI ≤ 0.25), and 32 MCQs (6.45%) had low discrimination (DI ≤ 0.15). Regarding MCQs with high discrimination, our study's results were notably higher compared to other studies. For instance, Nojomi et al. [27] reported that only 38.02% of their MCQs had high discrimination, Islam ZU et al. [35] observed 54.3%, and Patil et al. [13] found 50% with high discrimination. In terms of MCQs with moderate discrimination, our study revealed that 8.21% of questions fell into this category, a rate considerably lower than those reported by other researchers. Nojomi et al. [27] found that 44.09% of their MCQs had moderate discrimination, Patil et al. [13] found 20%, and Islam ZU et al [35] found 24.3%. For MCQs with low discrimination, our study reported a 6.4% rate, while Nojomi et al. [27] found 13.47%, Patil et al. [13] noted 30%, and Islam ZU et al. [35] observed 15.7%.
The consistent use of MCQs with acceptable psychometric properties supports their continued use in assessments. However, there is a need to increase the proportion of questions that assess higher-order cognitive skills to better prepare students for clinical practice. Implementing structured training for faculty on MCQ construction and incorporating standardized guidelines could improve the cognitive level and quality of questions.
Limitations: the subjectivity of participants in evaluating the quality of MCQs can introduce bias into the results. This subjectivity may affect the consistency of question assessments, complicating the establishment of a precise standard for judging the quality of MCQs. The relationship between the quality of MCQs and student learning is not necessarily direct. Several external factors, such as motivation, learning methods, and pedagogical support, may influence student performance. These variables can make it challenging to interpret the study's results in terms of their broader impact on overall student learning. Inter-observer agreement prior to consensus was not formally quantified. Although a normalization session was held before the analysis and all discrepancies were resolved by discussion until consensus was reached, the absence of a reliability coefficient limits the extent to which the reproducibility of the cognitive-level and item-writing-flaw classifications can be assessed.
Generalizability: the findings of this study may be applicable to other medical schools with similar educational contexts, particularly in regions with comparable curricula and assessment practices. However, differences in institutional policies, faculty expertise, and student populations may limit the generalizability of the results. Collaborative studies across multiple institutions could provide more comprehensive insights and enhance the external validity of the findings.
The docimological evaluation of MCQs in this study provided valuable insights into the quality of assessments in the pediatrics module at the Faculty of Medicine, Cadi Ayyad University, Marrakech. The results indicated a significant proportion of questions with good or acceptable difficulty levels, and a high percentage of discriminative questions, demonstrating effective test design. However, the study also highlighted areas for improvement, with some questions exhibiting low cognitive level, low discrimination, and notable item IWFs. To enhance the reliability and validity of assessments, regular item analysis and peer review are recommended. The findings underscore the importance of continuous quality improvement in question design to ensure fair and valid student assessments in medical education.
What is known about this topic
- Challenges in constructing quality MCQs: it is recognized that constructing high-quality multiple-choice questions (MCQs) is challenging and time-consuming; many MCQs focus primarily on testing recall rather than higher-order cognitive skills, affecting the depth of assessment;
- Item writing flaws (IWFs) and their impact: research has documented that a substantial portion of MCQs used in medical assessments are flawed, often due to poor question construction, which can undermine the reliability and validity of the examination.
What this study adds
- Detailed docimological analysis over a decade: this study provides an unprecedented decade-long analysis of MCQs within the pediatric module, employing docimological methods to scrutinize the evolution and impact of item writing flaws over time;
- Cognitive complexity in pediatric MCQs: by examining the distribution of cognitive levels across a large dataset of MCQs, this study highlights how the depth of cognitive engagement in pediatric exams has shifted, offering insights into the balance of cognitive demands placed on learners in this discipline;
- Longitudinal analysis of item writing flaws: this study provides a comprehensive ten-year analysis of IWFs in pediatric MCQs, revealing trends and patterns in how these flaws have persisted or changed over time within a specific medical education discipline.
The authors declare no competing interests.
All the authors read and approved the final version of this manuscript.
We thank the Department of Courses and Examinations at the Faculty of Medicine Marrakech.
Table 1: examples of multiple-choice questions illustrating the three levels of cognitive learning of the adapted Bloom’s taxonomy (levels I, II and III), selected from the pediatrics and pediatric surgery examinations of the Faculty of Medicine, Cadi Ayyad University, Marrakech, Morocco, May 2022 (n=3 illustrative items)
Table 2: rubric of the 18 item-writing flaws (IWFs) applied by the two independent observers to the multiple-choice questions of the pediatrics module, Faculty of Medicine, Cadi Ayyad University, Marrakech, Morocco, 2013-2023 (n=18 criteria)
Table 3: number and percentage of multiple-choice questions at each level of the adapted Bloom’s taxonomy, by examination session, pediatrics module, Faculty of Medicine, Cadi Ayyad University, Marrakech, Morocco, 2013-2023 (N=1,100 items from 22 examinations)
Table 4: mean ± standard deviation (SD) of the difficulty index (DIF) of multiple-choice questions, by examination session, pediatrics module, Faculty of Medicine, Cadi Ayyad University, Marrakech, Morocco, 2019-2023 (n=500 items from 10 examinations)
Table 5: mean ± standard deviation (SD) of the discrimination index (DI) of multiple-choice questions, by examination session, pediatrics module, Faculty of Medicine, Cadi Ayyad University, Marrakech, Morocco, 2019-2023 (n=500 items from 10 examinations)
- Al-Rukban MO. Guidelines for the construction of multiple choice questions tests. J Family Community Med. 2006 Sep;13(3):125-33. PubMed | Google Scholar
- Cheung D, Bucat R. How can we construct good multiple-choice items. InScience and Technology Education Conference, Hong Kong. 2002. Google Scholar
- McCoubrie P. Improving the fairness of multiple-choice questions: a literature review. Med Teach. 2004 Dec;26(8):709-12. PubMed | Google Scholar
- Javaeed A. Assessment of Higher Ordered Thinking in Medical Education: Multiple Choice Questions and Modified Essay Questions. MedEdPublish (2016). 2018 Jun 12;7:128. PubMed | Google Scholar
- Farley JK. The multiple--choice test: developing the test blueprint. Nurse Educ. 1989 Sep-Oct;14(5):3-5. PubMed | Google Scholar
- Tarrant M, Ware J. Impact of item-writing flaws in multiple-choice questions on student achievement in high-stakes nursing assessments. Med Educ. 2008 Feb;42(2):198-206. PubMed | Google Scholar
- Masters JC, Hulsmeyer BS, Pike ME, Leichty K, Miller MT, Verst AL. Assessment of multiple-choice questions in selected test banks accompanying text books used in nursing education. J Nurs Educ. 2001 Jan;40(1):25-32. PubMed | Google Scholar
- Downing SM. Construct-irrelevant variance and flawed test questions: Do multiple-choice item-writing principles make any difference? Acad Med. 2002 Oct;77(10 Suppl):S103-4. PubMed | Google Scholar
- Dell KA, Wantuch GA. How-to-guide for writing multiple choice questions for the pharmacy instructor. Curr Pharm Teach Learn. 2017 Jan-Feb;9(1):137-144. PubMed | Google Scholar
- Downing SM. The effects of violating standard item writing principles on tests and students: the consequences of using flawed test items on achievement examinations in medical education. Adv Health Sci Educ Theory Pract. 2005;10(2):133-43. PubMed | Google Scholar
- Leclercq D, Nicaise J, Demeuse M. Docimologie critique: des difficultés de noter des copies et d’attribuer des notes aux élèves. 2004 Jan 1:273-92. Google Scholar
- Jozefowicz RF, Koeppen BM, Case S, Galbraith R, Swanson D, Glew RH. The quality of in-house medical school examinations. Acad Med. 2002 Feb;77(2):156-61. PubMed | Google Scholar
- Patil R, Palve SB, Vell K, Boratne AV. Evaluation of multiple choice questions by item analysis in a medical college at Pondicherry, India. International Journal of Community Medicine and Public Health. 2016 Jun;3(6):1612-6. Google Scholar
- Pais J, Silva A, Guimarães B, Povo A, Coelho E, Silva-Pereira F et al. Do item-writing flaws reduce examinations psychometric quality? BMC Res Notes. 2016 Aug 11;9(1):399. PubMed | Google Scholar
- Krathwohl DR. A Revision of Bloom’s Taxonomy: An Overview. Theory Pract. 2002;41(4):212-218. Google Scholar
- Thompson AR, Braun MW, O'Loughlin VD. A comparison of student performance on discipline-specific versus integrated exams in a medical school course. Adv Physiol Educ. 2013 Dec;37(4):370-6. PubMed | Google Scholar
- Freeman S, Haak D, Wenderoth MP. Increased course structure improves performance in introductory biology. CBE Life Sci Educ. 2011 Summer;10(2):175-86. PubMed | Google Scholar
- Zheng AY, Lawhorn JK, Lumley T, Freeman S. Assessment. Application of Bloom's taxonomy debunks the "MCAT myth". Science. 2008 Jan 25;319(5862):414-5. PubMed | Google Scholar
- Frassica FJ, Papp D, McCarthy E, Weber K. Analysis of the pathology section of the OITE will aid in trainee preparation. Clin Orthop Relat Res. 2008 Jun;466(6):1323-8. PubMed | Google Scholar
- Murphy RF, Nunez L, Barfield WR, Mooney JF 3rd. Evaluation of Pediatric Questions on the Orthopaedic In-Training Examination-An Update. J Pediatr Orthop. 2017 Sep;37(6):e394-e397. PubMed | Google Scholar
- Crowe A, Dirks C, Wenderoth MP. Biology in bloom: implementing Bloom's Taxonomy to enhance student learning in biology. CBE Life Sci Educ. 2008 Winter;7(4):368-81. PubMed | Google Scholar
- Haladyna TM. Developing and validating multiple-choice test items. 2004. Google Scholar
- Friedman SJ. Constructing Test Items: Multiple-Choice, Constructed-Response, Performance, and Other Formats. 1999. JSTOR.. PubMed | Google Scholar
- Wiersma W, Jurs SG. Educational measurement and testing. 1990.
- Jovanovska J. Designing Effective Multiple-Choice Questions for Assessing Learning Outcomes. Infotheca. 2018;18(1):25-42. Google Scholar
- Baig M, Ali SK, Ali S, Huda N. Evaluation of Multiple Choice and Short Essay Question items in Basic Medical Sciences. Pak J Med Sci. 2014 Jan;30(1):3-6. PubMed | Google Scholar
- Nojomi M, Mahmoudi M. Assessment of multiple-choice questions by item analysis for medical students’ examinations. Res Dev Med Educ. 2022;11:24. Google Scholar
- Khan MU, Aljarallah BM. Evaluation of Modified Essay Questions (MEQ) and Multiple Choice Questions (MCQ) as a tool for Assessing the Cognitive Skills of Undergraduate Medical Students. Int J Health Sci (Qassim). 2011 Jan;5(1):39-43. PubMed | Google Scholar
- Fayyaz Khan H, Farooq Danish K, Saeed Awan A, Anwar M. Identification of technical item flaws leads to improvement of the quality of single best Multiple Choice Questions. Pak J Med Sci. 2013 May;29(3):715-8. PubMed | Google Scholar
- Kowash M, Hussein I, Al Halabi M. Evaluating the Quality of Multiple Choice Question in Paediatric Dentistry Postgraduate Examinations. Sultan Qaboos Univ Med J. 2019 May;19(2):e135-e141. PubMed | Google Scholar
- Hijji BM. Flaws of Multiple Choice Questions in Teacher-Constructed Nursing Examinations: A Pilot Descriptive Study. J Nurs Educ. 2017 Aug 1;56(8):490-496. PubMed | Google Scholar
- Pham H, Besanko J, Devitt P. Examining the impact of specific types of item-writing flaws on student performance and psychometric properties of the multiple choice question. MedEdPublish (2016). 2018 Oct 2;7:225. PubMed | Google Scholar
- Christian DS, Prajapati AC, Rana BM, Dave VR. Evaluation of multiple choice questions using item analysis tool: a study from a medical institute of Ahmedabad, Gujarat. Int J Community Med Public Health. 2017;4(6):1876. Google Scholar
- Rao C, Kishan Prasad H, Sajitha K, Permi H, Shetty J. Item analysis of multiple choice questions: Assessing an assessment tool in medical students. Int J Educ Psychol Res. 2016;2(4):201. Google Scholar
- Islam ZU, Usmani A. Psychometric analysis of Anatomy MCQs in Modular examination. Pak J Med Sci. 2017 Sep-Oct;33(5):1138-1143. PubMed | Google Scholar



