Consequential Validity Evidence of Yes-No Angoff Standard Setting in a Pre-Clinical Medical School Curriculum | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Consequential Validity Evidence of Yes-No Angoff Standard Setting in a Pre-Clinical Medical School Curriculum Ketsia Dimanche, Edward Klatt, Marshall Angle This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4889026/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 14 Mar, 2025 Read the published version in BMC Medical Education → Version 1 posted 4 You are reading this latest preprint version Abstract Background Standard setting is consequential to student outcomes for defining sufficient mastery of knowledge and thus accurately classifying medical students as competent (or not) for advancing through a pre-clinical curriculum. In particular, for multiple-choice medical knowledge exams the Yes-No Angoff method for standard setting yields consequential cut-off scores based on faculty-experts’ item-level judgments of students possessing borderline but sufficient competence required for demonstrating sufficient mastery. Adapting a construct validity framework and defining consequential academic success in terms of unimpeded progress, we investigated consequential validity evidence of passing standards derived from the Yes-No Angoff method. Methods We analyzed academic success for four pre-clinical semesters across three student cohorts. First, we identified passing standards for pre-clinical courses using the Yes-No Angoff method. Then, we applied binary logistic regression and receiver-operator characteristic (ROC) analyses with area under the curve (AUC) to evaluate passing standards. For binary outcomes, we defined academic success in terms of unimpeded progress through the curriculum and students’ first-attempt passage of the United States Medical Licensing Examination (USMLE) Step 1. Model predictors included Yes-No Angoff passing standards, Medical College Admissions test scores, and grade point averages for math and science courses. Results ROC analyses showed a low but acceptable area under the curve for a single semester in one cohort and excellent or outstanding AUCs for the remaining 11 semesters. With rates ranging between 89% and 96%, overall classification accuracy for predicting academic success was adequate for all pre-clinical semesters across all cohorts. Conclusions Affirming that Yes-No Angoff standards accurately predicted academic success, study results indicated consequential validity evidence for our school’s passing standards in a pre-clinical medical curriculum. medical education standard setting Yes-No Angoff consequential validity Figures Figure 1 Background Undergraduate medical education typically requires a pre-clinical curriculum that focuses on the basic science knowledge and professional performance competencies students will need for clinical education and postgraduate residency training. For assessing knowledge in the pre-clinical curriculum, medical schools set passing standards for courses that must be achieved in order for students to advance. A school may employ any number of methods for deriving cut-off passing scores for exams and passing standards for courses. Selected standard setting methods include historic precedent, Angoff, Yes-No Angoff, Bookmarking, Contrasting Groups, Ebel, and Hofstee. 1 , 2 Adapted from Angoff standard setting, the Yes-No Angoff method is criterion-based and thus relies on the expert judgements of raters. With its relative ease of use and efficiency in administration, faculty experts rate exam questions based on subjective predictions of the item-level performance of “borderline” students performing at a level sufficient to pass and move on in the curriculum. 2 Construct validity defines how well a test measures the concept it was designed to evaluate. Construct validity becomes important when assessing a concept that cannot be directly measured, such as competence, and multiple measurable indicators, such as test items, are needed. 3 Ultimately, these item-level predictions of raters yield passing standards that define the borderline level of competence that is yet sufficient for progressing through a curriculum and adequate for attaining a terminal degree. Any given method, systematically and rigorously applied, can yield fair and defensible standards. 1 And yet, for any given method, standards are only as fair and defensible as the evidence for validity warrants. With this imperative in mind, we investigated the consequential validity of passing standards that were derived from the Yes-No Angoff method. 3 , 4 , 5 , 6 , 7 , 8 Methods This study was reviewed by the Mercer University School of Medicine Institutional Review Board and deemed not to involve health care intervention, and need for consent waived, and ethics approval granted, with reference number H2212295. Our school’s pre-clinical courses are organized in terms of four semester long blocks. For each block we assessed for students’ mastery of biomedical science with locally developed multiple choice question (MCQ) exams, setting standards per the Yes-No Angoff method. Our data included standard setting outcomes for three cohorts of medical students in 2017, 2018, and 2019, N = 126, 112, and 116 for each cohort. Over these years there were 24 exams in block 1 and 22 in block 2 with a total of 3025 items in the first pre-clinical year. There were 18 exams in block 3 and 24 in block 4 with 2775 items in the second pre-clinical year. The numbers of Yes-No Angoff raters typically ranged between 12 and 18 faculty members per exam. In our application of the Yes-No Angoff method, faculty members reviewed each MCQ exam item, and for that item answered the question, “Will the borderline student answer this question correctly?” For each item, faculty rated 1 for Yes or 0 for No. From these binary data, we summed the means of item ratings to determine cut-off scores for each exam. Then, we totaled the sum of cut-off scores to calculate pre-clinical passing standards for each block. This system for determining block passing standards was compensatory such that better performance on some exams compensated for poor performance on others. Table 1 illustrates an item × rater matrix for demonstrating how item means were calculated and then summed to determine the cut-off scores for exams. For example, the sum of item means and final cut-off score for this 10-item exam is 6.714. Table 1 Example of Yes-No Faculty Angoff Ratings by 7 faculty for 10 Exam Items Raters 1 2 3 4 5 6 7 Mean 1 1 0 1 1 1 0 1 0.714 2 0 1 0 1 1 0 1 0.571 3 0 1 1 0 1 0 1 0.571 4 0 0 1 1 0 1 1 0.571 5 1 1 1 1 1 1 1 1.000 6 1 1 1 1 1 1 1 1.000 7 1 0 1 1 0 1 1 0.714 8 0 0 0 1 1 1 0 0.429 9 1 0 0 0 1 1 1 0.571 10 0 1 1 0 1 0 1 0.571 Item Totals 5 5 7 7 8 6 9 6.714 We analyzed binary logistic regression results and receiver-operator characteristic curves (ROC) to evaluate consequential validity of passing standards. 9 , 10 Our independent variable was Academic Success , defined as students’ unimpeded progress through the pre-clinical curriculum plus their first-attempt passage of USMLE Step 1. The principle predictor in our models was Faculty Standard , which represented the dichotomized outcomes of students scoring at or above faculty passing standards versus those who scored below. Model predictors also included Medical College Admissions Test scores ( MCAT ) and pre-matriculating grade-point-averages for biology, chemistry, physics, and math courses ( BCPM-GPA ). 11 First, we calculated predicted probabilities and overall classification accuracy for Academic Success . Then, we plotted these predicted probabilities against observed outcomes for Academic Success in order to construct ROC-curves. Inferring validity evidence from ROC-curves, we applied conventional guidelines for evaluating area under the curve (AUC): AUC of 0.5 to 0.59 is poor, 0.6 to 0.69 is fair, 0.7 to 0.8 is good, 0.8 to 0.9 is excellent, and more than 0.9 is considered outstanding. 12 Results In plotting predicted probabilities against actual outcomes, ROC-curves illustrate simultaneously the true and false positive rates of students who should and should not have passed, showing the probability of success or failure across all probability thresholds between 0 and 1. For each pre-clinical block, Fig. 1 illustrates the ROC curves with the Yes-No Angoff method for predicting academic success, with all ROC-curves per cohort bending toward the upper-left corner. ROC curves closer to the upper-left corner indicate better evidence for validity. Thus, with MCAT , BCPM-GPA , and Faculty Standard as predictors, these curves visually demonstrate the goodness of fit of our models for predicting Academic Success . Complementing our study’s ROC-curves, Table 2 documents the AUC for all pre-clinical blocks and cohorts. The 2018 cohort’s Block 3 AUC was somewhat low but acceptable at 0.7941. All other models were excellent or outstanding with AUC results ranging between 0.8176 and 0.9238. Table 2 Area Under the Curve for Block Totals Block 1 Block 2 Block 3 Block 4 2017 Cohort 0.8697 0.9153 0.8962 0.9195 2018 Cohort 0.8176 0.8961 0.7941 0.9069 2019 Cohort 0.8992 0.9238 0.9096 0.9111 Note. AUC = 0.7 to 0.799 is acceptable; AUC = 0.8 to 0.899 is excellent; AUC = 0.9 + is outstanding. As referenced above, we defined Academic Success in terms of students’ unimpeded progress through the pre-clinical curriculum and their first-attempt passage of the USMLE Step 1. Across all cohorts, the actual rates of unimpeded progression through pre-clinical blocks were 90% for Blocks 1, 2, and 3, and 95% for Block 4. Our school’s first-time pass rate for USMLE Step 1 across all cohorts was 91%, compared with 94 to 96% for U.S. and Canadian medical school examinees in 2017-2019. 13 The rate of Academic Success , as defined for binary logistic regression, was 79%. As a measure of overall accuracy, the rates of correct classification in our models ranged between 89% and 96%. Discussion Assessment requires validity evidence that affirms academic standards, gives meaning to test scores, and informs high stakes decisions for advancing students through a curriculum. 3 , 5 , 6 , 7 , 8 In our study, “consequential validity” was defined in terms of the binary outcome, Academic Success . This outcome and our method for investigating validity are based on the work of Dunleavy et al who defined academic success in terms of “unimpeded progress” and investigated predictive validity with binary logistic regression and ROC-curve analyses. 4 In this way of thinking about validity: (1) the probabilities for achieving Academic Success were predicated, in large part, on faculty-set passing standards; and (2) AUC indicated goodness of fit and thus the quality of consequential validity evidence. Per our study’s method, AUC findings indicated adequate goodness of fit in our predictive models. Moreover, rates of classification accuracy affirmed the consequential outcomes of students progressing through the curriculum unimpeded and passing Step 1 on their first try. Significantly, validity ultimately concerns both the consequences of test use and the accuracy of inferences made about psychological constructs. 4 , 8 , 14 Thus, in addition to validating standard setting outcomes, our study’s results provided validity evidence for the foundational psychological construct that defines the Yes-No Angoff procedure, namely “borderline competence.” 1 , 2 Accordingly, our results pointed to consequential validity evidence that most students had developed (or not) the necessary competence in biomedical science. However, our models were not always helpful in predicting the likelihood that students who failed truly lacked the necessary competence for medical knowledge. Our school anticipated this limitation when it adopted the Yes-No Angoff procedure and therefore developed a protocol for adjusting faculty standards with an approach suggested by Camara et al. 15 As such, we addressed the possible classification error of unwarranted impeded progress by adjusting cut-off scores downward according to our calculations for standard error of measurement (SEM). 15 , 16 , 17 Concurrently, we adjusted faculty standards upward per SEM to address the other kind of classification error: that of passing incompetent students. For this kind of classification error, upward adjustments also helped us identify students who needed academic support. Ultimately, we see our application of SEM as sound protocol in standard setting that mitigates undesirable consequences while supporting the appropriate use of academic standards. 8 It should be noted that the borderline student is not an honors student or even an average student; but neither is the borderline student a failing student. The borderline competent student may struggle to pass at a given point in the curriculum, but having once passed, will have demonstrated the sufficient competence required for advancing through the pre-clinical curriculum. Faculty development aimed at honing expert raters’ shared understandings of borderline competence is foundational for supporting accurate inferences. 18 , 19 Moreover, we hold that validation of passing standards should be continuous in order to reflect the dynamic nature of testing for locally-developed assessments. Accordingly, we advocate for annual reviews of predictive models and classification statistics that affirm the competence of high achievers but also help schools of medicine identify students who need additional academic support. As regards our school’s dynamic application of the Yes-No Angoff method across multiple cohorts we observed a profoundly important value-added to teaching and learning: the systematic review of exams, item by item, was a necessary task fulfilled by multiple faculty colleagues. In reviewing items, faculty provided feedback on the accuracy of test items, their formatting and relevance to medical practice, and their alignment with curriculum goals and objectives. These reviews supported the quality improvement of new MCQ-items developed for new knowledge content in our pre-clinical curriculum. Certainly, our cohorts included a number of students who were not coded for Academic Success according to our study’s stringent rules for defining the binary outcome. And yet, most of our students who failed to achieve so-called Academic Success did in fact achieve ultimate success after remediation in the pre-clinical curriculum or after their passage of Step 1 on a second try. We would note that all of our matriculated students have gone through a rigorous evaluation process for selection and that our school’s curriculum is stepwise such that faculty do not expect students to have achieved full expertise all at once. It should also be noted that the typical medical school incorporates six competencies into a curriculum, not just medical knowledge. Thus, an MD graduate must also demonstrate competency in interpersonal and communication skills, professional behavior, patient care, systems-based practice, and practice-based learning and improvement. 20 With this broader view of the knowledge, skills, and attitudes students need for a successful professional career, the importance of consequential validity evidence comes into sharper focus as schools of medicine assess for more than just a single competency. Ultimately, validity studies for all competencies are required for affirming students’ probable success in graduate medical education and their career readiness for the entry-level responsibilities of first-year residents. Conclusions Out study showed that the Yes-No Angoff method was adequate for predicting academic success and therefore indicative of consequential validity. Moreover, the method provided a means for setting valid standards with substantial faculty input. This input supported test development, contributed to MCQ-item accuracy for inferring mastery of knowledge, and improved the clinical relevance of our exams. Significantly, our school’s systematic application of the Yes-No Angoff method aided in identifying borderline performing students in need of academic support. Future studies might consider the range of values in binary logistic cut-points for a more comprehensive discussion of positive and negative predictive values and the tradeoff associated with varying cut-point thresholds. Abbreviations ROC AUC USMLE MCQ MCAT BCPM-GPA SEM Declarations Ethics approval and consent to participate : This study was reviewed by the Mercer University School of Medicine Institutional Review Board and deemed not to involve health care intervention, and need for consent waived, and ethics approval granted, with reference number H2212295. Consent for publication : Not applicable. Availability of data and materials : The datasets used and analyzed during the current study are available from the corresponding author on reasonable request. Competing interests : The authors declare that they have no competing interests. Funding : None. Authors’ contributions : All authors contributed equally to the analysis of the data, the preparation of the manuscript, and read and approved the final manuscript. Acknowledgements : Not applicable. References Downing SM, Tekian A, Yudkowsky R. Procedures for establishing defensible absolute passing scores on performance examinations in health professions education. Teach Learn Med. 2006;18(1):50–7. https://doi.org/10.1207/s15328015tlm1801_11 . Yudkowsky R, Downing SM, Wirth S. Simpler standards for local performance examinations: The Yes/No Angoff and Whole Test Ebel. Teach Learn Med. 2008;20(3):212–7. https://doi.org/10.1080/10401330802199450 . Messick S. Validity of psychological assessment: validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. ETS Res Rep Ser. 1994;1994(2):i–28. https://doi.org/10.1002/j.2333-8504.1994.tb01618.x . Dunleavy DM, Kroopnick MH, Dowd KW, Searcy CA, Zhao X. The predictive validity of the MCAT exam in relation to academic performance through medical school: A national cohort study of 2001–2004 matriculants. Acad Med. 2013;88(5):666–71. https://doi.org/10.1097/ACM.0b013e3182864299 . Kane M. Validating the performance standards associated with passing scores. Rev Educ Res. 1994;64:425–61. https://www.jstor.org/stable/1170678 . Kane MT. Validating the Interpretations and Uses of Test Scores. J Educ Meas. 2013;50(1):1–73. https://doi.org/10.1111/jedm.12000 . Hambleton RK. Setting performance standards on educational assessments and criteria for evaluating the process. In: Cizek GJ, editor. Setting Performance Standards: Theory and Applications. New York: Routledge; 2001. Bachman LF. Building and supporting a case for test use. Lang Assess Q. 2005;2(1):1–34. https://doi.org/10.1207/s15434311laq0201_1 . Berwick V, Cheek L, Ball J. Statistics review 14: Logistic regression. Crit Care. 2005;9(1):112–8. https://doi.org/10.1186/cc3045 . Eng J. Receiver operating characteristic analysis: A primer. Acad Radiol. 2005;12(7):909–16. https://doi.org/10.1016/j.acra.2005.04.005 . Kleshinski J, Khuder SA, Shapiro JI, Gold JP. Impact of preadmission variables on USMLE step 1 and step 2 performance. Adv Health Sci Educ. 2009;14:69–78. https://doi.org/10.1007/s10459-007-9087-x . Mandrekar JN. Receiver operating characteristic curve in diagnostic test assessment. J Thoracic Oncology. 2010;5(9):1315–1316. https://doi.org10.1097/JTO.0b013e3181ec173d. USMLE Performance Data. https://www.usmle.org/performance-data . Accessed 9 August 2024. Loevinger J. Objective tests as instruments of psychological theory. Psychol Rep. 1957;3(3):635–94. https://doi.org/10.2466/pr0.1957.3.3.635 . Camara WJ, Allen JM, Moore JL. Empirically based college- and career-readiness cut scores and performance standards. In: McClarty KL, Mattern KD, Gaertner MN, editors. Preparing Students for College and Careers. 1st ed. New York, NY: Routledge; 2018. pp. 70–81. Hays R, Gupta TS, Veitch J. The practical value of the standard error of measurement in borderline pass/fail decisions. Med Educ. 2008;42(8):810–5. https://doi.org/10.1111/j.1365-2923.2008.03103.x . Biddle RE. How to set cutoff scores for knowledge tests used in promotion, training, certification, and licensing. Public Personnel Manage. 1993;229(1):63–79. https://doi.org/10.1177/009102609302200105 . Clauser JC, Clauser BE, Hambleton RK. Increasing the validity of Angoff standards through analysis of judge-level internal consistency. Appl Measur Educ. 2014;27:19–30. https://doi.org/10.1080/08957347.2013.853071 . Tannenbaum RJ, Kannan P. Consistency of Angoff-based standard setting judgments: Are item judgments and passing scores replicable across different panels of experts? Educational Assess. 2015;20:66–78. https://doi.org/10.1080/10627197.2015.997619 . Coalition for Physician Accountability. Consensus Statement on a Framework for Professional Competence by the Coalition for Physician Accountability. https://physicianaccountability.org/wp-content/uploads/2020/05/Coalition-Competencies-Consensus-Statement-FINAL.pdf . Updated. 2014. Accessed 22 May 2024. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 14 Mar, 2025 Read the published version in BMC Medical Education → Version 1 posted Editorial decision: Revision requested 20 Aug, 2024 Editor assigned by journal 19 Aug, 2024 Submission checks completed at journal 19 Aug, 2024 First submitted to journal 09 Aug, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4889026","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":342501953,"identity":"321a5652-75c7-4d04-ba19-32986d4771eb","order_by":0,"name":"Ketsia Dimanche","email":"","orcid":"","institution":"Mercer University School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ketsia","middleName":"","lastName":"Dimanche","suffix":""},{"id":342501958,"identity":"5f5e5727-8d0f-4047-ad44-8c17077909f6","order_by":1,"name":"Edward Klatt","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA20lEQVRIiWNgGAWjYBACCXYogx8hlkBACzMDY8MBIEOygWQtBgeI1SLZzPz88QcGG3nja4ePffhRcY+Bnz3HAK8WaWY2Q6AtaYbbbqclz+w5U8wg2fMGvxY5ZgaQlsMJZrdzjJkZ2xIYDG4QsEWOmf0jUMv/BOPZ+Z+ZGf8lMNgT0iLNzAOy5UCCgXQOMzNjA9AWCQJaJJt5CmecYUg2nHE7zZix51gCj8SZZwV4tUgcb9/woYLBTp5/dvJjhh81CXL87ckb8GoBA8Z/CDYPYeWjYBSMglEwCggCACokQSXXr39tAAAAAElFTkSuQmCC","orcid":"","institution":"Mercer University School of Medicine","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Edward","middleName":"","lastName":"Klatt","suffix":""},{"id":342501961,"identity":"2dbbe0b6-157c-4476-a589-39c70e38cbb6","order_by":2,"name":"Marshall Angle","email":"","orcid":"","institution":"Mercer University School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Marshall","middleName":"","lastName":"Angle","suffix":""}],"badges":[],"createdAt":"2024-08-09 20:37:52","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4889026/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4889026/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12909-025-06948-8","type":"published","date":"2025-03-14T15:57:45+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":64583613,"identity":"d6e95e5f-1a80-4293-a7f0-b095648175bc","added_by":"auto","created_at":"2024-09-16 07:12:16","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":720054,"visible":true,"origin":"","legend":"\u003cp\u003eSee image above for figure legend\u003c/p\u003e","description":"","filename":"AngoffValidityDimancheFig1.png","url":"https://assets-eu.researchsquare.com/files/rs-4889026/v1/ae7e247d142584568015ddf0.png"},{"id":78688986,"identity":"bda5b439-a925-455f-a261-5e98181df437","added_by":"auto","created_at":"2025-03-17 16:09:42","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1184384,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4889026/v1/5290c29b-837e-47a4-9db5-3ca8fe9c2bea.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Consequential Validity Evidence of Yes-No Angoff Standard Setting in a Pre-Clinical Medical School Curriculum","fulltext":[{"header":"Background","content":"\u003cp\u003eUndergraduate medical education typically requires a pre-clinical curriculum that focuses on the basic science knowledge and professional performance competencies students will need for clinical education and postgraduate residency training. For assessing knowledge in the pre-clinical curriculum, medical schools set passing standards for courses that must be achieved in order for students to advance. A school may employ any number of methods for deriving cut-off passing scores for exams and passing standards for courses. Selected standard setting methods include historic precedent, Angoff, Yes-No Angoff, Bookmarking, Contrasting Groups, Ebel, and Hofstee.\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eAdapted from Angoff standard setting, the Yes-No Angoff method is criterion-based and thus relies on the expert judgements of raters. With its relative ease of use and efficiency in administration, faculty experts rate exam questions based on subjective predictions of the item-level performance of \u0026ldquo;borderline\u0026rdquo; students performing at a level sufficient to pass and move on in the curriculum.\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e Construct validity defines how well a test measures the concept it was designed to evaluate. Construct validity becomes important when assessing a concept that cannot be directly measured, such as competence, and multiple measurable indicators, such as test items, are needed.\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eUltimately, these item-level predictions of raters yield passing standards that define the borderline level of competence that is yet sufficient for progressing through a curriculum and adequate for attaining a terminal degree. Any given method, systematically and rigorously applied, can yield fair and defensible standards.\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e And yet, for any given method, standards are only as fair and defensible as the evidence for validity warrants. With this imperative in mind, we investigated the consequential validity of passing standards that were derived from the Yes-No Angoff method.\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e,\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e,\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e This study was reviewed by the Mercer University School of Medicine Institutional Review Board and deemed not to involve health care intervention, and need for consent waived, and ethics approval granted, with reference number H2212295.\u003c/p\u003e \u003cp\u003eOur school\u0026rsquo;s pre-clinical courses are organized in terms of four semester long blocks. For each block we assessed for students\u0026rsquo; mastery of biomedical science with locally developed multiple choice question (MCQ) exams, setting standards per the Yes-No Angoff method. Our data included standard setting outcomes for three cohorts of medical students in 2017, 2018, and 2019, N\u0026thinsp;=\u0026thinsp;126, 112, and 116 for each cohort. Over these years there were 24 exams in block 1 and 22 in block 2 with a total of 3025 items in the first pre-clinical year. There were 18 exams in block 3 and 24 in block 4 with 2775 items in the second pre-clinical year. The numbers of Yes-No Angoff raters typically ranged between 12 and 18 faculty members per exam.\u003c/p\u003e \u003cp\u003eIn our application of the Yes-No Angoff method, faculty members reviewed each MCQ exam item, and for that item answered the question, \u0026ldquo;Will the borderline student answer this question correctly?\u0026rdquo; For each item, faculty rated 1 for Yes or 0 for No. From these binary data, we summed the means of item ratings to determine cut-off scores for each exam. Then, we totaled the sum of cut-off scores to calculate pre-clinical passing standards for each block. This system for determining block passing standards was compensatory such that better performance on some exams compensated for poor performance on others. Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e illustrates an item \u0026times; rater matrix for demonstrating how item means were calculated and then summed to determine the cut-off scores for exams. For example, the sum of item means and final cut-off score for this 10-item exam is 6.714.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eExample of Yes-No Faculty Angoff Ratings by 7 faculty for 10 Exam Items\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"9\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRaters\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eMean\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.714\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.571\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.571\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.571\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.714\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.429\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.571\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.571\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eItem Totals\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e6.714\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eWe analyzed binary logistic regression results and receiver-operator characteristic curves (ROC) to evaluate consequential validity of passing standards.\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e Our independent variable was \u003cem\u003eAcademic Success\u003c/em\u003e, defined as students\u0026rsquo; unimpeded progress through the pre-clinical curriculum plus their first-attempt passage of USMLE Step 1. The principle predictor in our models was \u003cem\u003eFaculty Standard\u003c/em\u003e, which represented the dichotomized outcomes of students scoring at or above faculty passing standards versus those who scored below. Model predictors also included Medical College Admissions Test scores (\u003cem\u003eMCAT\u003c/em\u003e) and pre-matriculating grade-point-averages for biology, chemistry, physics, and math courses (\u003cem\u003eBCPM-GPA\u003c/em\u003e).\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eFirst, we calculated predicted probabilities and overall classification accuracy for \u003cem\u003eAcademic Success\u003c/em\u003e. Then, we plotted these predicted probabilities against observed outcomes for \u003cem\u003eAcademic Success\u003c/em\u003e in order to construct ROC-curves. Inferring validity evidence from ROC-curves, we applied conventional guidelines for evaluating area under the curve (AUC): AUC of 0.5 to 0.59 is poor, 0.6 to 0.69 is fair, 0.7 to 0.8 is good, 0.8 to 0.9 is excellent, and more than 0.9 is considered outstanding.\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eIn plotting predicted probabilities against actual outcomes, ROC-curves illustrate simultaneously the true and false positive rates of students who should and should not have passed, showing the probability of success or failure across all probability thresholds between 0 and 1. For each pre-clinical block, Fig.\u0026nbsp;1 illustrates the ROC curves with the Yes-No Angoff method for predicting academic success, with all ROC-curves per cohort bending toward the upper-left corner. ROC curves closer to the upper-left corner indicate better evidence for validity. Thus, with \u003cem\u003eMCAT\u003c/em\u003e, \u003cem\u003eBCPM-GPA\u003c/em\u003e, and \u003cem\u003eFaculty Standard\u003c/em\u003e as predictors, these curves visually demonstrate the goodness of fit of our models for predicting \u003cem\u003eAcademic Success\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eComplementing our study\u0026rsquo;s ROC-curves, Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e documents the AUC for all pre-clinical blocks and cohorts. The 2018 cohort\u0026rsquo;s Block 3 AUC was somewhat low but acceptable at 0.7941. All other models were excellent or outstanding with AUC results ranging between 0.8176 and 0.9238.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eArea Under the Curve for Block Totals\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBlock 1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBlock 2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBlock 3\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eBlock 4\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2017 Cohort\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.8697\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.9153\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.8962\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9195\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2018 Cohort\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.8176\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.8961\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.7941\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9069\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2019 Cohort\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.8992\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.9238\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.9096\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.9111\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003eNote. AUC\u0026thinsp;=\u0026thinsp;0.7 to 0.799 is acceptable; AUC\u0026thinsp;=\u0026thinsp;0.8 to 0.899 is excellent; AUC\u0026thinsp;=\u0026thinsp;0.9\u0026thinsp;+\u0026thinsp;is outstanding.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAs referenced above, we defined \u003cem\u003eAcademic Success\u003c/em\u003e in terms of students\u0026rsquo; unimpeded progress through the pre-clinical curriculum and their first-attempt passage of the USMLE Step 1. Across all cohorts, the actual rates of unimpeded progression through pre-clinical blocks were 90% for Blocks 1, 2, and 3, and 95% for Block 4. Our school\u0026rsquo;s first-time pass rate for USMLE Step 1 across all cohorts was 91%, compared with 94 to 96% for U.S. and Canadian medical school examinees in 2017-2019.\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e The rate of \u003cem\u003eAcademic Success\u003c/em\u003e, as defined for binary logistic regression, was 79%. As a measure of overall accuracy, the rates of correct classification in our models ranged between 89% and 96%.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eAssessment requires validity evidence that affirms academic standards, gives meaning to test scores, and informs high stakes decisions for advancing students through a curriculum.\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e,\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e,\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e In our study, \u0026ldquo;consequential validity\u0026rdquo; was defined in terms of the binary outcome, \u003cem\u003eAcademic Success\u003c/em\u003e. This outcome and our method for investigating validity are based on the work of Dunleavy et al who defined academic success in terms of \u0026ldquo;unimpeded progress\u0026rdquo; and investigated predictive validity with binary logistic regression and ROC-curve analyses.\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e In this way of thinking about validity: (1) the probabilities for achieving \u003cem\u003eAcademic Success\u003c/em\u003e were predicated, in large part, on faculty-set passing standards; and (2) AUC indicated goodness of fit and thus the quality of consequential validity evidence. Per our study\u0026rsquo;s method, AUC findings indicated adequate goodness of fit in our predictive models. Moreover, rates of classification accuracy affirmed the consequential outcomes of students progressing through the curriculum unimpeded and passing Step 1 on their first try.\u003c/p\u003e \u003cp\u003eSignificantly, validity ultimately concerns both the consequences of test use and the accuracy of inferences made about psychological constructs.\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e,\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e Thus, in addition to validating standard setting outcomes, our study\u0026rsquo;s results provided validity evidence for the foundational psychological construct that defines the Yes-No Angoff procedure, namely \u0026ldquo;borderline competence.\u0026rdquo;\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e Accordingly, our results pointed to consequential validity evidence that most students had developed (or not) the necessary competence in biomedical science. However, our models were not always helpful in predicting the likelihood that students who failed truly lacked the necessary competence for medical knowledge. Our school anticipated this limitation when it adopted the Yes-No Angoff procedure and therefore developed a protocol for adjusting faculty standards with an approach suggested by Camara et al.\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e As such, we addressed the possible classification error of unwarranted impeded progress by adjusting cut-off scores downward according to our calculations for standard error of measurement (SEM).\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e,\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e Concurrently, we adjusted faculty standards upward per SEM to address the other kind of classification error: that of passing incompetent students. For this kind of classification error, upward adjustments also helped us identify students who needed academic support. Ultimately, we see our application of SEM as sound protocol in standard setting that mitigates undesirable consequences while supporting the appropriate use of academic standards.\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eIt should be noted that the borderline student is not an honors student or even an average student; but neither is the borderline student a failing student. The borderline competent student may struggle to pass at a given point in the curriculum, but having once passed, will have demonstrated the sufficient competence required for advancing through the pre-clinical curriculum. Faculty development aimed at honing expert raters\u0026rsquo; shared understandings of borderline competence is foundational for supporting accurate inferences.\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e,\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e Moreover, we hold that validation of passing standards should be continuous in order to reflect the dynamic nature of testing for locally-developed assessments. Accordingly, we advocate for annual reviews of predictive models and classification statistics that affirm the competence of high achievers but also help schools of medicine identify students who need additional academic support. As regards our school\u0026rsquo;s dynamic application of the Yes-No Angoff method across multiple cohorts we observed a profoundly important value-added to teaching and learning: the systematic review of exams, item by item, was a necessary task fulfilled by multiple faculty colleagues. In reviewing items, faculty provided feedback on the accuracy of test items, their formatting and relevance to medical practice, and their alignment with curriculum goals and objectives. These reviews supported the quality improvement of new MCQ-items developed for new knowledge content in our pre-clinical curriculum.\u003c/p\u003e \u003cp\u003eCertainly, our cohorts included a number of students who were not coded for \u003cem\u003eAcademic Success\u003c/em\u003e according to our study\u0026rsquo;s stringent rules for defining the binary outcome. And yet, most of our students who failed to achieve so-called \u003cem\u003eAcademic Success\u003c/em\u003e did in fact achieve ultimate success after remediation in the pre-clinical curriculum or after their passage of Step 1 on a second try. We would note that all of our matriculated students have gone through a rigorous evaluation process for selection and that our school\u0026rsquo;s curriculum is stepwise such that faculty do not expect students to have achieved full expertise all at once. It should also be noted that the typical medical school incorporates six competencies into a curriculum, not just medical knowledge. Thus, an MD graduate must also demonstrate competency in interpersonal and communication skills, professional behavior, patient care, systems-based practice, and practice-based learning and improvement.\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e With this broader view of the knowledge, skills, and attitudes students need for a successful professional career, the importance of consequential validity evidence comes into sharper focus as schools of medicine assess for more than just a single competency. Ultimately, validity studies for all competencies are required for affirming students\u0026rsquo; probable success in graduate medical education and their career readiness for the entry-level responsibilities of first-year residents.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eOut study showed that the Yes-No Angoff method was adequate for predicting academic success and therefore indicative of consequential validity. Moreover, the method provided a means for setting valid standards with substantial faculty input. This input supported test development, contributed to MCQ-item accuracy for inferring mastery of knowledge, and improved the clinical relevance of our exams. Significantly, our school\u0026rsquo;s systematic application of the Yes-No Angoff method aided in identifying borderline performing students in need of academic support. Future studies might consider the range of values in binary logistic cut-points for a more comprehensive discussion of positive and negative predictive values and the tradeoff associated with varying cut-point thresholds.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eROC\u003c/p\u003e\n\u003cp\u003eAUC\u003c/p\u003e\n\u003cp\u003eUSMLE\u003c/p\u003e\n\u003cp\u003eMCQ\u003c/p\u003e\n\u003cp\u003eMCAT\u003c/p\u003e\n\u003cp\u003eBCPM-GPA\u003c/p\u003e\n\u003cp\u003eSEM\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e: \u0026nbsp;This study was reviewed by the Mercer University School of Medicine Institutional Review Board\u0026nbsp;and deemed not to involve health care intervention, and need for consent waived, and ethics approval granted, with reference number\u0026nbsp;H2212295.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e: \u0026nbsp; Not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e: The datasets used and analyzed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e: \u0026nbsp;The authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e: \u0026nbsp;None.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; contributions\u003c/strong\u003e: \u0026nbsp;All authors contributed equally to the analysis of the data, the preparation of the manuscript, and read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e: \u0026nbsp;Not applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eDowning SM, Tekian A, Yudkowsky R. Procedures for establishing defensible absolute passing scores on performance examinations in health professions education. Teach Learn Med. 2006;18(1):50\u0026ndash;7. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1207/s15328015tlm1801_11\u003c/span\u003e\u003cspan address=\"10.1207/s15328015tlm1801_11\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYudkowsky R, Downing SM, Wirth S. Simpler standards for local performance examinations: The Yes/No Angoff and Whole Test Ebel. Teach Learn Med. 2008;20(3):212\u0026ndash;7. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/10401330802199450\u003c/span\u003e\u003cspan address=\"10.1080/10401330802199450\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMessick S. Validity of psychological assessment: validation of inferences from persons\u0026rsquo; responses and performances as scientific inquiry into score meaning. ETS Res Rep Ser. 1994;1994(2):i\u0026ndash;28. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/j.2333-8504.1994.tb01618.x\u003c/span\u003e\u003cspan address=\"10.1002/j.2333-8504.1994.tb01618.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDunleavy DM, Kroopnick MH, Dowd KW, Searcy CA, Zhao X. The predictive validity of the MCAT exam in relation to academic performance through medical school: A national cohort study of 2001\u0026ndash;2004 matriculants. Acad Med. 2013;88(5):666\u0026ndash;71. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1097/ACM.0b013e3182864299\u003c/span\u003e\u003cspan address=\"10.1097/ACM.0b013e3182864299\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKane M. Validating the performance standards associated with passing scores. Rev Educ Res. 1994;64:425\u0026ndash;61. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.jstor.org/stable/1170678\u003c/span\u003e\u003cspan address=\"https://www.jstor.org/stable/1170678\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKane MT. Validating the Interpretations and Uses of Test Scores. J Educ Meas. 2013;50(1):1\u0026ndash;73. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/jedm.12000\u003c/span\u003e\u003cspan address=\"10.1111/jedm.12000\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHambleton RK. Setting performance standards on educational assessments and criteria for evaluating the process. In: Cizek GJ, editor. Setting Performance Standards: Theory and Applications. New York: Routledge; 2001.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBachman LF. Building and supporting a case for test use. Lang Assess Q. 2005;2(1):1\u0026ndash;34. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1207/s15434311laq0201_1\u003c/span\u003e\u003cspan address=\"10.1207/s15434311laq0201_1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBerwick V, Cheek L, Ball J. Statistics review 14: Logistic regression. Crit Care. 2005;9(1):112\u0026ndash;8. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/cc3045\u003c/span\u003e\u003cspan address=\"10.1186/cc3045\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEng J. Receiver operating characteristic analysis: A primer. Acad Radiol. 2005;12(7):909\u0026ndash;16. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.acra.2005.04.005\u003c/span\u003e\u003cspan address=\"10.1016/j.acra.2005.04.005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKleshinski J, Khuder SA, Shapiro JI, Gold JP. Impact of preadmission variables on USMLE step 1 and step 2 performance. Adv Health Sci Educ. 2009;14:69\u0026ndash;78. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s10459-007-9087-x\u003c/span\u003e\u003cspan address=\"10.1007/s10459-007-9087-x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMandrekar JN. Receiver operating characteristic curve in diagnostic test assessment. J Thoracic Oncology. 2010;5(9):1315\u0026ndash;1316. https://doi.org10.1097/JTO.0b013e3181ec173d.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUSMLE Performance Data. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.usmle.org/performance-data\u003c/span\u003e\u003cspan address=\"https://www.usmle.org/performance-data\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed 9 August 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLoevinger J. Objective tests as instruments of psychological theory. Psychol Rep. 1957;3(3):635\u0026ndash;94. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.2466/pr0.1957.3.3.635\u003c/span\u003e\u003cspan address=\"10.2466/pr0.1957.3.3.635\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCamara WJ, Allen JM, Moore JL. Empirically based college- and career-readiness cut scores and performance standards. In: McClarty KL, Mattern KD, Gaertner MN, editors. Preparing Students for College and Careers. 1st ed. New York, NY: Routledge; 2018. pp. 70\u0026ndash;81.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHays R, Gupta TS, Veitch J. The practical value of the standard error of measurement in borderline pass/fail decisions. Med Educ. 2008;42(8):810\u0026ndash;5. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/j.1365-2923.2008.03103.x\u003c/span\u003e\u003cspan address=\"10.1111/j.1365-2923.2008.03103.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBiddle RE. How to set cutoff scores for knowledge tests used in promotion, training, certification, and licensing. Public Personnel Manage. 1993;229(1):63\u0026ndash;79. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/009102609302200105\u003c/span\u003e\u003cspan address=\"10.1177/009102609302200105\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eClauser JC, Clauser BE, Hambleton RK. Increasing the validity of Angoff standards through analysis of judge-level internal consistency. Appl Measur Educ. 2014;27:19\u0026ndash;30. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/08957347.2013.853071\u003c/span\u003e\u003cspan address=\"10.1080/08957347.2013.853071\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTannenbaum RJ, Kannan P. Consistency of Angoff-based standard setting judgments: Are item judgments and passing scores replicable across different panels of experts? Educational Assess. 2015;20:66\u0026ndash;78. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/10627197.2015.997619\u003c/span\u003e\u003cspan address=\"10.1080/10627197.2015.997619\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCoalition for Physician Accountability. Consensus Statement on a Framework for Professional Competence by the Coalition for Physician Accountability. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://physicianaccountability.org/wp-content/uploads/2020/05/Coalition-Competencies-Consensus-Statement-FINAL.pdf\u003c/span\u003e\u003cspan address=\"https://physicianaccountability.org/wp-content/uploads/2020/05/Coalition-Competencies-Consensus-Statement-FINAL.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Updated. 2014. Accessed 22 May 2024.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"medical education, standard setting, Yes-No Angoff, consequential validity","lastPublishedDoi":"10.21203/rs.3.rs-4889026/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4889026/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eStandard setting is consequential to student outcomes for defining sufficient mastery of knowledge and thus accurately classifying medical students as competent (or not) for advancing through a pre-clinical curriculum. In particular, for multiple-choice medical knowledge exams the Yes-No Angoff method for standard setting yields consequential cut-off scores based on faculty-experts\u0026rsquo; item-level judgments of students possessing borderline but sufficient competence required for demonstrating sufficient mastery. Adapting a construct validity framework and defining consequential academic success in terms of unimpeded progress, we investigated consequential validity evidence of passing standards derived from the Yes-No Angoff method.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eWe analyzed academic success for four pre-clinical semesters across three student cohorts. First, we identified passing standards for pre-clinical courses using the Yes-No Angoff method. Then, we applied binary logistic regression and receiver-operator characteristic (ROC) analyses with area under the curve (AUC) to evaluate passing standards. For binary outcomes, we defined academic success in terms of unimpeded progress through the curriculum and students\u0026rsquo; first-attempt passage of the United States Medical Licensing Examination (USMLE) Step 1. Model predictors included Yes-No Angoff passing standards, Medical College Admissions test scores, and grade point averages for math and science courses.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eROC analyses showed a low but acceptable area under the curve for a single semester in one cohort and excellent or outstanding AUCs for the remaining 11 semesters. With rates ranging between 89% and 96%, overall classification accuracy for predicting academic success was adequate for all pre-clinical semesters across all cohorts.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eAffirming that Yes-No Angoff standards accurately predicted academic success, study results indicated consequential validity evidence for our school\u0026rsquo;s passing standards in a pre-clinical medical curriculum.\u003c/p\u003e","manuscriptTitle":"Consequential Validity Evidence of Yes-No Angoff Standard Setting in a Pre-Clinical Medical School Curriculum","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-09-16 07:12:12","doi":"10.21203/rs.3.rs-4889026/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-08-20T07:10:54+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-08-19T13:02:37+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-08-19T12:59:48+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Education","date":"2024-08-09T20:36:32+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"382d957b-54ce-4a0c-83c5-260823fd7442","owner":[],"postedDate":"September 16th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-03-17T16:02:09+00:00","versionOfRecord":{"articleIdentity":"rs-4889026","link":"https://doi.org/10.1186/s12909-025-06948-8","journal":{"identity":"bmc-medical-education","isVorOnly":false,"title":"BMC Medical Education"},"publishedOn":"2025-03-14 15:57:45","publishedOnDateReadable":"March 14th, 2025"},"versionCreatedAt":"2024-09-16 07:12:12","video":"","vorDoi":"10.1186/s12909-025-06948-8","vorDoiUrl":"https://doi.org/10.1186/s12909-025-06948-8","workflowStages":[]},"version":"v1","identity":"rs-4889026","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4889026","identity":"rs-4889026","version":["v1"]},"buildId":"-HB7Z8yhvgn0wM9Nzuekk","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.