Standard Setting Very Short Answer Questions (VSAQs) Relative to Single Best Answer Questions (SBAQs): Does Having Access to the Answers Make a Difference? | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Standard Setting Very Short Answer Questions (VSAQs) Relative to Single Best Answer Questions (SBAQs): Does Having Access to the Answers Make a Difference? Amir H Sam, Kate R Millar, Rachel Westacott, Colin R Melville, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1431427/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 10 You are reading this latest preprint version Abstract Background Weinvestigated whether question formatand access to the correct answers affect the pass mark set by standard-setters on written examinations. Methods Trained educatorsused the Angoff method to standard set two 50-item tests with identical vignettes, one in a single best answer question (SBAQ) format (with five answer options) and the other in a very short answer question (VSAQ) format (requiring free text responses). Half the participants had access to the correct answers and half did not. The data for each group were analysed to determine if the question format or having access to the answers affected the pass mark set. Results A lower pass mark was set for the VSAQ test than the SBAQ testby the standard setters who had access to the answers (median difference of 13.85 percentage points, Z=-2.82, p=0.002). Comparable pass marks were set for the SBAQ test by standard setters with and without access to the correct answers (60.65% and 60.90% respectively).A lower pass mark was set for the VSAQ testwhen participants hadaccess to the correct answers (difference in medians -13.75 percentage points, Z=2.46, p=0.014). Conclusions When given access to the potential correct answers, standard setters appear to appreciate the increased difficulty of VSAQs compared to SBAQs. Assessment undergraduate standard setting Figures Figure 1 Background Single Best Answer Questions (SBAQs) are widely used in medical assessment including high stakes licensing exams such as the US Medical Licensing Examination, membership examinations of many of the UK Royal Colleges, and final examinations of UK medical schools. However there has been criticism of this question format, for being subject to cueing and not reflecting real life clinical practice 1 , 2 . Very Short Answer Questions (VSAQs) are a novel assessment method that has been proposed as a solution to this problem. Like SBAQs, VSAQs have a clinical vignette followed by a lead-in question. However instead of having a list of answer options to choose from, the candidate provides their own answer of between one and five words in length. The candidate’s answers are marked against a set of preapproved answers 2 – 6 and, due to recent advances in technology, they can be delivered and marked electronically 3 – 6 . Any answers that do not match the preapproved options can then be reviewed to consider if they should be marked correct and added to the future lists of approved answers 2 – 6 . The cueing associated with SBAQs is mitigated with VSAQs as the answer options are removed 3 . Student performance in exams has been shown to be affected by using this question format, indicating that students find this question type more challenging 3 – 6 . VSAQs have been shown to be a better representation of candidates’ unprompted level of knowledge, with a recent study showing that the average student scored 21 percentage points lower on the VSAQ compared to the SBAQ of the same stem 3 . There is considerable variation in the means of assessment and methods of standard setting across medical schools 7 – 9 . Whilst several standard setting methods have been examined empirically for SBAQs 10 , standard setting methodology for VSAQs has not been studied to date. As VSAQs have been successfully introduced into undergraduate assessment 3 – 6 , 11 , the question of how to set the pass mark for this assessment method needs to considered. It has previously been shown that standard setting estimates for SBAQs are significantly affected by a judge knowing, or not knowing the answer to the item 12 . Verheggen et al. found that a judge’s knowledge of the subject and their stringency as a judge impacted on the standard set for an item and therefore the standard was not purely a reflection of the difficulty of the item 12 . As VSAQs are designed to have a range of accepted answers, provision of these for the judges when standard setting may be even more significant. Bourque et al (2020) looked at standard setting SBAQs for a national postgraduate exam (using the Ebel method) and found no difference in scoring regardless of whether the answers were provided to the judges or not 13 . Using a set of common stem items, we set out to study whether the question format (VSAQs versus SBAQs) affects the pass mark. We also investigated whether having access to the answers had an effect on the pass marks set for both SBAQ and VSAQ formats. Methods Two 50-question assessment papers were created using the same question vignettes, one paper using the VSAQ and the other paper in an SBAQ format with five answer options. These items had previously been used in a formative assessment of 1,417 volunteer final year medical students 3 . The papers were standard set using the Angoff method 14 . Twenty three teaching faculty from Imperial College School of Medicine were trained on standard setting in undergraduate examinations using the Angoff method, through a face-to-face workshop. This allowed participants to arrive at a common understanding of what constitutes a borderline candidate. Participants were randomised into four groups to standard set the papers in different formats as per Table 1 . Table 1 Session Design by Group Group A (n = 6) Group B (n = 6) Group C (n = 6) Group D (n = 5) Pre-session training Face to Face workshop: Standard Setting in Undergraduate Examinations. Session 1 50 SBAQs (with answers) 50 VSAQs (with answers) 50 SBAQs (without answers) 50 VSAQs (without answers) Session 2 50 VSAQs (with answers) 50 SBAQs (with answers) 50 VSAQs (without answers) 50 SBAQs (without answers) Eleven partcipants judged the paper without access to the answers, as the student would see it, and twelve received the correct answers, as typically happens in standard setting practice (Fig. 1 ). In order to account for the impact of the order in which standard setters saw VSAQs and SBAQs half the groups saw VSAQs first before SBAQs and the other half saw SBAQs before VSAQs. A washout period of six weeks between session 1 and session 2 was created to prevent standard-setters being subject to cueing from previous standard setting sessions. Data Collection Each participant was asked to standard set the paper using the Angoff method by judging each item in the paper using the question “What proportion of consistently just safe newly qualified doctors would get the question correct? Scores were submitted using an electronic survey tool. For the VSAQs they were able to submit a response between 0 and 100%. For the SBAQs this was adjusted to a response between 20 and 100% to allow for the 20% chance that the candidate can guess the correct answer from the five options provided with no prior knowledge. Data Analysis The anonymised standards data for each participant were downloaded from the electronic survey tool into an Excel file. The data were transferred into Stata V.16 for analysis. The mean of each participant’s standard from the Angoff method for the 50 items was calculated to produce their overall pass mark. For each group the median of all members’ pass marks was calculated for the VSAQ exam and the SBAQ exam. To determine if the question format influenced the standard set, a paired-data Wilcoxon Sign Rank test on participants’ overall pass marks by question format was carried out separately for the participants that had the question answers (groups A & B) and the participants that did not (groups C & D). To determine if having access to the answers influenced the standard set, an independent samples Mann Whitney U test was carried out separately on the overall pass marks set for the VSAQ and for the SBAQ papers (Groups A + B vs. Groups C + D). For all statistical significance testing a critical p-value of 0.0125 was used given the use of multiple comparisons. Ethical approval for the study was granted by the Imperial College London Medical Education Ethics Committee (MEEC) (MEEC1920-178). Results Question format Table 2 presents summary statistics comparing standards set for the SBAQ and VSAQ formats of the assessment, with and without the answers. There was a large and statistically significant difference between VSAQ and SBAQ pass marks set by the groups that had access to the answers. The SBAQ pass mark was set higher than the VSAQ pass mark, with a median difference of 13.85 percentage points (Z=-2.82, p = 0.002). There was not a statistically significant difference between the VSAQ and SBAQ pass marks set by the groups that did not have access to the answers (median difference − 1.90 percentage points, Z = 0.45, p = 0.700). Access to Answers For VSAQs, having the answers resulted in a statistically significant reduction in the pass marks set, as shown in Table 2 (difference in medians − 13.75 percentage points, Z = 2.46, p = 0.014). For SBAQs, having the answers did not make a significant difference to the pass mark set (difference in medians − 0.25 percentage points, Z = 0.06, p = 0.952). Table 2 Summary Statistics Comparing Standards Set for SBAQs and VSAQs Median SBAQ Median VSAQ Median Difference (SBAQ-VSAQ) Wilcoxon Sign Rank p-value; Z-score Answers 60.65 49.95 13.85 p = 0.002; Z-score = -2.824 No-Answers 60.90 63.70 -1.90 p = 0.700; Z-score = 0.445 Difference in Medians (Answers - No-Answers) -0.25 -13.75 Mann Whitney U; p-value; Z-score p = 0.952; Z-score = 0.062 p = 0.014; Z-score = 2.462 Discussion In this study, the question format affected the pass mark set by standard setting judges when they were given access to the answers (as is usual in standard setting practice). When standard setters were shown the answers for VSAQs, they produced a significantly lower pass mark for the VSAQ paper. It has been shown that students score an average 21 percentage points lower on VSAQs 3 , and this study suggests that this is taken into account to some extent by standard setters with access to the answers who set an median pass mark of almost 14 percentage points lower for the VSAQs. In addition, we investigated if there are different pass marks set when standard setting judges do or do not have access to the answers in VSAQs vs SBAQs. Standard setters who could see the answers set a lower median pass mark for the VSAQs. We hypothesise that this is related to having access to the range of accepted VSAQ answers, which gives a concrete indication of the degree of difficulty of the question. This is in contrast to studies with multiple choice questions, where it has been suggested that access to the answers is likely to cause judges to underestimate the difficulty of the question 12 , or access to the answers made no difference to the standard set using the Ebel method 13 . A limitation of our study was that we were not able to hold a group discussion, as is considered best practice when standard setting for high stakes assessment. Group standard setting meetings result in sharing of information and discussion of questions that often results in constructive revision of scores 8 . It has also been shown to improve method reliability and reduces the number of judges required 15 . This is a valuable part of the process, especially when members of the standard setting panel may be less familiar with VSAQs, and so are likely to benefit from sharing of experience. Providing standard setters with question facility for VSAQs, or typical performance differences for VSAQs versus SBAQs, may also help when setting standards in this unfamiliar question type. This was not done in our study as the questions were used formatively so performance data is not likely to be a true reflection of summative assessment. A further limitation of our study is the small sample size, which was due to finding suitable members of faculty within one institution. We limited eligiable participants to those who had experience in final medical school examinations and were involved in delivery the undergraduate currciculum, to ensure the highest quality judges in our standard setting panel. A future study of standard setters across a wider cohort of medical schools would allow a larger sample size and could also look at judges’ characteristics (average age, years standard setting, years spent in undergraduate teaching for example) that may affect the standards they set. As far as we are aware, our study is the first to look at standard setting for VSAQs. It is also the first study to demonstrate the importance of the standard setting panel having access to the answers when scoring VSAQs. As VSAQs are increasingly introduced into undergraduate medical assessments, it opens the discussion for what must be considered when identifying the ideal standard setting method for this novel question format. Our study demonstrates the feasibility of using the Angoff method to standard set this novel question type in undergraduate medical education. In addition, it provides a platform for further research, including comparing other recognised methods of standard setting – the Cohen method which would consider setting a standard in relation to the performance of the cohort, and the Ebel method which asks judges to consider the importance of the knowledge tested as well as the difficulty of the question. Conclusions The potential benefit of integrating VSAQs into undergraduate medical assessments has already been demonstrated, so it follows that a validated standard setting method is needed if they are to be used in high stakes examinations. To our knowledge this is the first study comparing standard setting with and without the answers in VSAQs. Further research on a larger scale is warranted to determine if the effect we have seen persists in a larger and more varied population of standard setters. Based on the present study, we recommend that answers should be provided to the standard setters to help them arrive at a valid standard. Abbreviations VSAQs Very Short Answer Questions SBAQs Single Best Answer Questions Declarations Ethics approval and consent to participate Ethical approval for the study was sought from and granted by the Imperial College London Medical Education Ethics Committee (MEEC) on 22/11/2019 (ref number: MEEC1920-178). All methods were performed in accordance with the Declarations of Helsinki and informed consent was obtained from all participants for the study. Consent for publication Not applicable Availability of data and materials The datasets used andanalysed during the current study are available from the corresponding author on reasonable request. Competing interests The authors declare that they have no competing interests Funding Not applicable Authors' contributions AHS, KRM and CAB designed and delivered the standard setting sessions. KRM and CAB carried out the statistical analysis, with support from AHS. All authors were involved in the overall design of the study, and all provided major contributations to writing the manuscript. All authors reviewed and approved the final manuscript. Acknowledgements Not applicable References Veloski JJ, Rabinowitz HK, Robeson MR, Young PR. Patients don’t present with five choices: An alternative to multiple- choice tests in assessing physicians’ competence. Acad Med. 1999;74(5):539–46. Sam AH, Hameed S, Harris J, Meeran K. Validity of very short answer versus single best answer questions for undergraduate assessment. BMC Med Educ. 2016;16(1):266. Sam AH, Westacott R, Gurnell M, Wilson R, Meeran K, Brown C. Comparing single-best-answer and very-short-answer questions for the assessment of applied medical knowledge in 20 UK medical schools: Cross-sectional study. BMJ Open. 2019;9(9). Sam AH, Peleva E, Fung CY, Cohen N, Benbow EW, Meeran K. Very Short Answer Questions: A Novel Approach To Summative Assessments In Pathology. Adv Med Educ Pract. 2019;Volume 10:943–8. Sam AH, Fung CY, Wilson RK, Peleva E, Kluth DC, Lupton M, et al. Using prescribing very short answer questions to identify sources of medication errors: A prospective study in two UK medical schools. BMJ Open. 2019;9(7). Sam AH, Field SM, Collares CF, van der Vleuten CPM, Wass VJ, Melville C, et al. Very-short-answer questions: reliability, discrimination and acceptability. Med Educ. 2018;52(4):447–55. MacDougall M. Variation in assessment and standard setting practices across UK undergraduate medicine and the need for a benchmark. Int J Med Educ. 2015 Oct 31;6:125–35. Yeates P, Cope N, Luksaite E, Hassell A, Dikomitis L. Exploring differences in individual and group judgements in standard setting. Med Educ. 2019 Sep 2;53(9):941–52. Taylor CA, Gurnell M, Melville CR, Kluth DC, Johnson N, Wass V. Variation in passing standards for graduation-level knowledge items at UK medical schools. Med Educ. 2017 Jun 1;51(6):612–20. Bandaranayake RC. Setting and maintaining standards in multiple choice examinations: AMEE Guide No. 37. Med Teach. 2008 Jan 3;30(9–10):836–45. Sam AH, Hameed S, Harris J, Meeran K. Validity of very short answer versus single best answer questions for undergraduate assessment. BMC Med Educ. 2016;16(1):266. Verheggen MM, Muijtjens AMM, Van Os J, Schuwirth LWT. Is an Angoff Standard an Indication of Minimal Competence of Examinees or of Judges? Adv Heal Sci Educ 2006 132. 2006 Oct 17;13(2):203–11. Bourque J, Skinner H, Dupré J, Bacchus M, Ainslie M, Ma IWY, et al. Performance of the Ebel standard-setting method in spring 2019 royal college of physicians and surgeons of canada internal medicine certification examination consisted of multiple-choice questions. J Educ Eval Health Prof. 2020 Apr 20;17. McKinley DW, Norcini JJ. How to set standards on performance-based examinations: AMEE Guide No. 85. Med Teach. 2014 Feb 20;36(2):97–110. Fowell SL, Fewtrell R, McLaughlin PJ. Estimating the Minimum Number of Judges Required for Test-centred Standard Setting on Written Assessments. Do Discussion and Iteration have an Influence? Adv Heal Sci Educ. 2008 Mar 7;13(1):11–24. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Major revision 16 Jun, 2022 Reviews received at journal 08 Apr, 2022 Reviewers agreed at journal 24 Mar, 2022 Reviews received at journal 23 Mar, 2022 Reviewers agreed at journal 21 Mar, 2022 Reviewers invited by journal 21 Mar, 2022 Editor assigned by journal 14 Mar, 2022 Editor invited by journal 11 Mar, 2022 Submission checks completed at journal 11 Mar, 2022 First submitted to journal 08 Mar, 2022 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1431427","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":90095417,"identity":"10dd720e-766e-421d-844b-527c6ea06139","order_by":0,"name":"Amir H Sam","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA3UlEQVRIiWNgGAWjYNCCCgkwdYChAEwzE6HlDEyLAbFaGNtgLGK08DewP/x0c55FtDl778MDPwwY5PkbeIwN8GmROMBjLJ27TSJ3Z89xg4M9BgyGM4AiCXitOcDDANay4UYakG3AwLiBgcf4AD4d8gfYH//OnQPUcv8Zw8E/Bgz2BLUYHGAwk85tANnCxnAYaEsiSAtehxke5jGzzjkG8ksaw2EZA4nkGYfZivF6X+54++PbOTV1udvZjzF/fFNhY9vf3rxZAp8WeBxADZYgLu6RtYyCUTAKRsEowAQAieNDhP5iXTcAAAAASUVORK5CYII=","orcid":"","institution":"Imperial College London","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Amir","middleName":"H","lastName":"Sam","suffix":""},{"id":90095418,"identity":"16b0bf61-5c25-4404-89f9-40b8134dbfae","order_by":1,"name":"Kate R Millar","email":"","orcid":"","institution":"Imperial College London","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Kate","middleName":"R","lastName":"Millar","suffix":""},{"id":90095419,"identity":"ebe3983d-e6fc-4c14-9330-d68461248e99","order_by":2,"name":"Rachel Westacott","email":"","orcid":"","institution":"University of Birmingham","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Rachel","middleName":"","lastName":"Westacott","suffix":""},{"id":90095420,"identity":"5fa659b6-865b-47e1-804b-248aba07007a","order_by":3,"name":"Colin R Melville","email":"","orcid":"","institution":"General Medical Council","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Colin","middleName":"R","lastName":"Melville","suffix":""},{"id":90095421,"identity":"6c1d05c4-2915-48b6-ac29-07ce497231f2","order_by":4,"name":"Celia A Brown","email":"","orcid":"","institution":"University of Warwick","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Celia","middleName":"A","lastName":"Brown","suffix":""}],"badges":[],"createdAt":"2022-03-08 14:14:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1431427/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1431427/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":19241828,"identity":"b7727166-40e0-46e8-82ee-f45b924b5bbe","added_by":"auto","created_at":"2022-03-15 14:38:54","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":125587,"visible":true,"origin":"","legend":"\u003cp\u003eQuestion formats: A - SBAQ without answer; B - SBAQ with answer; C - VSAQ without answer; D - VSAQ with answers\u003c/p\u003e","description":"","filename":"figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-1431427/v1/c2599925d16a5276065779e6.png"},{"id":19241829,"identity":"ee13d571-fb9d-4b8b-a391-a03bb830e28a","added_by":"auto","created_at":"2022-03-15 14:38:56","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":393662,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1431427/v1/eaab33e2-55a9-49b8-9b22-40900ab03053.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003eStandard Setting Very Short Answer Questions (VSAQs) Relative to Single Best Answer Questions (SBAQs): Does Having Access to the Answers Make a Difference?\u003c/p\u003e","fulltext":[{"header":"Background","content":"\u003cp\u003eSingle Best Answer Questions (SBAQs) are widely used in medical assessment including high stakes licensing exams such as the US Medical Licensing Examination, membership examinations of many of the UK Royal Colleges, and final examinations of UK medical schools. However there has been criticism of this question format, for being subject to cueing and not reflecting real life clinical practice \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eVery Short Answer Questions (VSAQs) are a novel assessment method that has been proposed as a solution to this problem. Like SBAQs, VSAQs have a clinical vignette followed by a lead-in question. However instead of having a list of answer options to choose from, the candidate provides their own answer of between one and five words in length. The candidate\u0026rsquo;s answers are marked against a set of preapproved answers \u003csup\u003e\u003cspan additionalcitationids=\"CR3 CR4 CR5\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e and, due to recent advances in technology, they can be delivered and marked electronically \u003csup\u003e\u003cspan additionalcitationids=\"CR4 CR5\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. Any answers that do not match the preapproved options can then be reviewed to consider if they should be marked correct and added to the future lists of approved answers \u003csup\u003e\u003cspan additionalcitationids=\"CR3 CR4 CR5\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. The cueing associated with SBAQs is mitigated with VSAQs as the answer options are removed \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Student performance in exams has been shown to be affected by using this question format, indicating that students find this question type more challenging \u003csup\u003e\u003cspan additionalcitationids=\"CR4 CR5\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. VSAQs have been shown to be a better representation of candidates\u0026rsquo; unprompted level of knowledge, with a recent study showing that the average student scored 21 percentage points lower on the VSAQ compared to the SBAQ of the same stem \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThere is considerable variation in the means of assessment and methods of standard setting across medical schools \u003csup\u003e\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Whilst several standard setting methods have been examined empirically for SBAQs \u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e, standard setting methodology for VSAQs has not been studied to date. As VSAQs have been successfully introduced into undergraduate assessment \u003csup\u003e\u003cspan additionalcitationids=\"CR4 CR5\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e, the question of how to set the pass mark for this assessment method needs to considered.\u003c/p\u003e \u003cp\u003eIt has previously been shown that standard setting estimates for SBAQs are significantly affected by a judge knowing, or not knowing the answer to the item \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. Verheggen et al. found that a judge\u0026rsquo;s knowledge of the subject and their stringency as a judge impacted on the standard set for an item and therefore the standard was not purely a reflection of the difficulty of the item\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e. As VSAQs are designed to have a range of accepted answers, provision of these for the judges when standard setting may be even more significant. Bourque et al (2020) looked at standard setting SBAQs for a national postgraduate exam (using the Ebel method) and found no difference in scoring regardless of whether the answers were provided to the judges or not\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eUsing a set of common stem items, we set out to study whether the question format (VSAQs versus SBAQs) affects the pass mark. We also investigated whether having access to the answers had an effect on the pass marks set for both SBAQ and VSAQ formats.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eTwo 50-question assessment papers were created using the same question vignettes, one paper using the VSAQ and the other paper in an SBAQ format with five answer options. These items had previously been used in a formative assessment of 1,417 volunteer final year medical students \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. The papers were standard set using the Angoff method\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eTwenty three teaching faculty from Imperial College School of Medicine were trained on standard setting in undergraduate examinations using the Angoff method, through a face-to-face workshop. This allowed participants to arrive at a common understanding of what constitutes a borderline candidate. Participants were randomised into four groups to standard set the papers in different formats as per Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e.\u0026nbsp;\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\n\u003ctable id=\"Tab1\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eSession Design by Group\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eGroup A (n\u0026thinsp;=\u0026thinsp;6)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eGroup B (n\u0026thinsp;=\u0026thinsp;6)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eGroup C (n\u0026thinsp;=\u0026thinsp;6)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eGroup D (n\u0026thinsp;=\u0026thinsp;5)\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003ePre-session training\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"4\" align=\"left\"\u003e\n\u003cp\u003eFace to Face workshop: Standard Setting in Undergraduate Examinations.\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eSession 1\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50 SBAQs\u003c/p\u003e\n\u003cp\u003e(with answers)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50 VSAQs\u003c/p\u003e\n\u003cp\u003e(with answers)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50 SBAQs (without answers)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50 VSAQs (without answers)\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eSession 2\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50 VSAQs\u003c/p\u003e\n\u003cp\u003e(with answers)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50 SBAQs\u003c/p\u003e\n\u003cp\u003e(with answers)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50 VSAQs (without answers)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50 SBAQs (without answers)\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eEleven partcipants judged the paper without access to the answers, as the student would see it, and twelve received the correct answers, as typically happens in standard setting practice (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\n\u003cp\u003eIn order to account for the impact of the order in which standard setters saw VSAQs and SBAQs half the groups saw VSAQs first before SBAQs and the other half saw SBAQs before VSAQs. A washout period of six weeks between session 1 and session 2 was created to prevent standard-setters being subject to cueing from previous standard setting sessions.\u003c/p\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n\u003ch2\u003eData Collection\u003c/h2\u003e\n\u003cp\u003eEach participant was asked to standard set the paper using the Angoff method by judging each item in the paper using the question \u0026ldquo;What proportion of consistently just safe newly qualified doctors would get the question correct? Scores were submitted using an electronic survey tool. For the VSAQs they were able to submit a response between 0 and 100%. For the SBAQs this was adjusted to a response between 20 and 100% to allow for the 20% chance that the candidate can guess the correct answer from the five options provided with no prior knowledge.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\n\u003ch2\u003eData Analysis\u003c/h2\u003e\n\u003cp\u003eThe anonymised standards data for each participant were downloaded from the electronic survey tool into an Excel file. The data were transferred into Stata V.16 for analysis. The mean of each participant\u0026rsquo;s standard from the Angoff method for the 50 items was calculated to produce their overall pass mark. For each group the median of all members\u0026rsquo; pass marks was calculated for the VSAQ exam and the SBAQ exam.\u003c/p\u003e\n\u003cp\u003eTo determine if the question format influenced the standard set, a paired-data Wilcoxon Sign Rank test on participants\u0026rsquo; overall pass marks by question format was carried out separately for the participants that had the question answers (groups A \u0026amp; B) and the participants that did not (groups C \u0026amp; D). To determine if having access to the answers influenced the standard set, an independent samples Mann Whitney U test was carried out separately on the overall pass marks set for the VSAQ and for the SBAQ papers (Groups A\u0026thinsp;+\u0026thinsp;B vs. Groups C\u0026thinsp;+\u0026thinsp;D). For all statistical significance testing a critical p-value of 0.0125 was used given the use of multiple comparisons.\u003c/p\u003e\n\u003cp\u003eEthical approval\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003efor the study was granted by the Imperial College London Medical Education Ethics Committee (MEEC) (MEEC1920-178).\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\n\u003ch2\u003eQuestion format\u003c/h2\u003e\n\u003cp\u003eTable\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e presents summary statistics comparing standards set for the SBAQ and VSAQ formats of the assessment, with and without the answers. There was a large and statistically significant difference between VSAQ and SBAQ pass marks set by the groups that had access to the answers. The SBAQ pass mark was set higher than the VSAQ pass mark, with a median difference of 13.85 percentage points (Z=-2.82, p\u0026thinsp;=\u0026thinsp;0.002). There was not a statistically significant difference between the VSAQ and SBAQ pass marks set by the groups that did not have access to the answers (median difference \u0026minus;\u0026thinsp;1.90 percentage points, Z\u0026thinsp;=\u0026thinsp;0.45, p\u0026thinsp;=\u0026thinsp;0.700).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\n\u003ch2\u003eAccess to Answers\u003c/h2\u003e\n\u003cp\u003eFor VSAQs, having the answers resulted in a statistically significant reduction in the pass marks set, as shown in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e (difference in medians \u0026minus;\u0026thinsp;13.75 percentage points, Z\u0026thinsp;=\u0026thinsp;2.46, p\u0026thinsp;=\u0026thinsp;0.014). For SBAQs, having the answers did not make a significant difference to the pass mark set (difference in medians \u0026minus;\u0026thinsp;0.25 percentage points, Z\u0026thinsp;=\u0026thinsp;0.06, p\u0026thinsp;=\u0026thinsp;0.952).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab2\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eSummary Statistics Comparing Standards Set for SBAQs and VSAQs\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eMedian SBAQ\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eMedian\u003c/p\u003e\n\u003cp\u003eVSAQ\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eMedian Difference (SBAQ-VSAQ)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eWilcoxon Sign Rank p-value; Z-score\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAnswers\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e60.65\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e49.95\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e13.85\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.002;\u003c/p\u003e\n\u003cp\u003eZ-score = -2.824\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNo-Answers\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e60.90\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e63.70\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e-1.90\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.700;\u003c/p\u003e\n\u003cp\u003eZ-score\u0026thinsp;=\u0026thinsp;0.445\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eDifference in Medians (Answers - No-Answers)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-0.25\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-13.75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMann Whitney U;\u003c/p\u003e\n\u003cp\u003ep-value;\u003c/p\u003e\n\u003cp\u003eZ-score\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.952;\u003c/p\u003e\n\u003cp\u003eZ-score\u0026thinsp;=\u0026thinsp;0.062\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ep\u0026thinsp;=\u0026thinsp;0.014;\u003c/p\u003e\n\u003cp\u003eZ-score\u0026thinsp;=\u0026thinsp;2.462\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, the question format affected the pass mark set by standard setting judges when they were given access to the answers (as is usual in standard setting practice). When standard setters were shown the answers for VSAQs, they produced a significantly lower pass mark for the VSAQ paper. It has been shown that students score an average 21 percentage points lower on VSAQs \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e, and this study suggests that this is taken into account to some extent by standard setters with access to the answers who set an median pass mark of almost 14 percentage points lower for the VSAQs.\u003c/p\u003e \u003cp\u003eIn addition, we investigated if there are different pass marks set when standard setting judges do or do not have access to the answers in VSAQs vs SBAQs. Standard setters who could see the answers set a lower median pass mark for the VSAQs. We hypothesise that this is related to having access to the range of accepted VSAQ answers, which gives a concrete indication of the degree of difficulty of the question. This is in contrast to studies with multiple choice questions, where it has been suggested that access to the answers is likely to cause judges to underestimate the difficulty of the question \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e, or access to the answers made no difference to the standard set using the Ebel method \u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eA limitation of our study was that we were not able to hold a group discussion, as is considered best practice when standard setting for high stakes assessment. Group standard setting meetings result in sharing of information and discussion of questions that often results in constructive revision of scores \u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. It has also been shown to improve method reliability and reduces the number of judges required \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. This is a valuable part of the process, especially when members of the standard setting panel may be less familiar with VSAQs, and so are likely to benefit from sharing of experience. Providing standard setters with question facility for VSAQs, or typical performance differences for VSAQs versus SBAQs, may also help when setting standards in this unfamiliar question type. This was not done in our study as the questions were used formatively so performance data is not likely to be a true reflection of summative assessment.\u003c/p\u003e \u003cp\u003eA further limitation of our study is the small sample size, which was due to finding suitable members of faculty within one institution. We limited eligiable participants to those who had experience in final medical school examinations and were involved in delivery the undergraduate currciculum, to ensure the highest quality judges in our standard setting panel. A future study of standard setters across a wider cohort of medical schools would allow a larger sample size and could also look at judges\u0026rsquo; characteristics (average age, years standard setting, years spent in undergraduate teaching for example) that may affect the standards they set.\u003c/p\u003e \u003cp\u003eAs far as we are aware, our study is the first to look at standard setting for VSAQs. It is also the first study to demonstrate the importance of the standard setting panel having access to the answers when scoring VSAQs. As VSAQs are increasingly introduced into undergraduate medical assessments, it opens the discussion for what must be considered when identifying the ideal standard setting method for this novel question format. Our study demonstrates the feasibility of using the Angoff method to standard set this novel question type in undergraduate medical education. In addition, it provides a platform for further research, including comparing other recognised methods of standard setting \u0026ndash; the Cohen method which would consider setting a standard in relation to the performance of the cohort, and the Ebel method which asks judges to consider the importance of the knowledge tested as well as the difficulty of the question.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eThe potential benefit of integrating VSAQs into undergraduate medical assessments has already been demonstrated, so it follows that a validated standard setting method is needed if they are to be used in high stakes examinations. To our knowledge this is the first study comparing standard setting with and without the answers in VSAQs. Further research on a larger scale is warranted to determine if the effect we have seen persists in a larger and more varied population of standard setters. Based on the present study, we recommend that answers should be provided to the standard setters to help them arrive at a valid standard.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eVSAQs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eVery Short Answer Questions\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSBAQs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eSingle Best Answer Questions\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003e\u003cem\u003eEthics approval and consent to participate\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eEthical approval for the study was sought from and granted by the Imperial College London Medical Education Ethics Committee (MEEC) on 22/11/2019 (ref number: MEEC1920-178).\u003c/p\u003e\n\u003cp\u003eAll methods were performed in accordance with the Declarations of Helsinki and informed consent was obtained from all participants for the study.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eConsent for publication\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eAvailability of data and materials\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets used andanalysed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eCompeting interests\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eFunding\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eAuthors' contributions\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAHS, KRM and CAB designed and delivered the standard setting sessions. KRM and CAB carried out the statistical analysis, with support from AHS. All authors were involved in the overall design of the study, and all provided major contributations to writing the manuscript. All authors reviewed and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eAcknowledgements\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eVeloski JJ, Rabinowitz HK, Robeson MR, Young PR. Patients don\u0026rsquo;t present with five choices: An alternative to multiple- choice tests in assessing physicians\u0026rsquo; competence. Acad Med. 1999;74(5):539\u0026ndash;46.\u003c/li\u003e\n\u003cli\u003eSam AH, Hameed S, Harris J, Meeran K. Validity of very short answer versus single best answer questions for undergraduate assessment. BMC Med Educ. 2016;16(1):266.\u003c/li\u003e\n\u003cli\u003eSam AH, Westacott R, Gurnell M, Wilson R, Meeran K, Brown C. Comparing single-best-answer and very-short-answer questions for the assessment of applied medical knowledge in 20 UK medical schools: Cross-sectional study. BMJ Open. 2019;9(9).\u003c/li\u003e\n\u003cli\u003eSam AH, Peleva E, Fung CY, Cohen N, Benbow EW, Meeran K. Very Short Answer Questions: A Novel Approach To Summative Assessments In Pathology. Adv Med Educ Pract. 2019;Volume 10:943\u0026ndash;8.\u003c/li\u003e\n\u003cli\u003eSam AH, Fung CY, Wilson RK, Peleva E, Kluth DC, Lupton M, et al. Using prescribing very short answer questions to identify sources of medication errors: A prospective study in two UK medical schools. BMJ Open. 2019;9(7).\u003c/li\u003e\n\u003cli\u003eSam AH, Field SM, Collares CF, van der Vleuten CPM, Wass VJ, Melville C, et al. Very-short-answer questions: reliability, discrimination and acceptability. Med Educ. 2018;52(4):447\u0026ndash;55.\u003c/li\u003e\n\u003cli\u003eMacDougall M. Variation in assessment and standard setting practices across UK undergraduate medicine and the need for a benchmark. Int J Med Educ. 2015 Oct 31;6:125\u0026ndash;35.\u003c/li\u003e\n\u003cli\u003eYeates P, Cope N, Luksaite E, Hassell A, Dikomitis L. Exploring differences in individual and group judgements in standard setting. Med Educ. 2019 Sep 2;53(9):941\u0026ndash;52.\u003c/li\u003e\n\u003cli\u003eTaylor CA, Gurnell M, Melville CR, Kluth DC, Johnson N, Wass V. Variation in passing standards for graduation-level knowledge items at UK medical schools. Med Educ. 2017 Jun 1;51(6):612\u0026ndash;20.\u003c/li\u003e\n\u003cli\u003eBandaranayake RC. Setting and maintaining standards in multiple choice examinations: AMEE Guide No. 37. Med Teach. 2008 Jan 3;30(9\u0026ndash;10):836\u0026ndash;45.\u003c/li\u003e\n\u003cli\u003eSam AH, Hameed S, Harris J, Meeran K. Validity of very short answer versus single best answer questions for undergraduate assessment. BMC Med Educ. 2016;16(1):266.\u003c/li\u003e\n\u003cli\u003eVerheggen MM, Muijtjens AMM, Van Os J, Schuwirth LWT. Is an Angoff Standard an Indication of Minimal Competence of Examinees or of Judges? Adv Heal Sci Educ 2006 132. 2006 Oct 17;13(2):203\u0026ndash;11.\u003c/li\u003e\n\u003cli\u003eBourque J, Skinner H, Dupr\u0026eacute; J, Bacchus M, Ainslie M, Ma IWY, et al. Performance of the Ebel standard-setting method in spring 2019 royal college of physicians and surgeons of canada internal medicine certification examination consisted of multiple-choice questions. J Educ Eval Health Prof. 2020 Apr 20;17.\u003c/li\u003e\n\u003cli\u003eMcKinley DW, Norcini JJ. How to set standards on performance-based examinations: AMEE Guide No. 85. Med Teach. 2014 Feb 20;36(2):97\u0026ndash;110.\u003c/li\u003e\n\u003cli\u003eFowell SL, Fewtrell R, McLaughlin PJ. Estimating the Minimum Number of Judges Required for Test-centred Standard Setting on Written Assessments. Do Discussion and Iteration have an Influence? Adv Heal Sci Educ. 2008 Mar 7;13(1):11\u0026ndash;24.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Assessment, undergraduate, standard setting ","lastPublishedDoi":"10.21203/rs.3.rs-1431427/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1431427/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eWeinvestigated whether question formatand access to the correct answers affect the pass mark set by standard-setters on written examinations. \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eTrained educatorsused the Angoff method to standard set two 50-item tests with identical vignettes, one in a single best answer question (SBAQ) format (with five answer options) and the other in a very short answer question (VSAQ) format (requiring free text responses). Half the participants had access to the correct answers and half did not. The data for each group were analysed to determine if the question format or having access to the answers affected the pass mark set.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eA lower pass mark was set for the VSAQ test than the SBAQ testby the standard setters who had access to the answers (median difference of 13.85 percentage points, Z=-2.82, p=0.002). Comparable pass marks were set for the SBAQ test by standard setters with and without access to the correct answers (60.65% and 60.90% respectively).A lower pass mark was set for the VSAQ testwhen participants hadaccess to the correct answers (difference in medians -13.75 percentage points, Z=2.46, p=0.014).\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusions\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eWhen given access to the potential correct answers, standard setters appear to appreciate the increased difficulty of VSAQs compared to SBAQs.\u0026nbsp;\u003c/p\u003e","manuscriptTitle":"Standard Setting Very Short Answer Questions (VSAQs) Relative to Single Best Answer Questions (SBAQs): Does Having Access to the Answers Make a Difference?","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-03-15 14:38:52","doi":"10.21203/rs.3.rs-1431427/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2022-06-16T06:04:42+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2022-04-08T16:01:08+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"bb49b0ae-9977-4fe8-b533-b5d20e259cb0","date":"2022-03-24T12:50:32+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2022-03-23T20:12:55+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"e0f27dbf-0d20-4751-bee8-56fe68c75af0","date":"2022-03-21T14:13:41+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2022-03-21T12:11:27+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2022-03-14T11:27:08+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2022-03-11T13:58:40+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2022-03-11T13:50:54+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Education","date":"2022-03-08T14:13:08+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"df1c4646-c25c-4b5a-901d-cb6bf3d7643d","owner":[],"postedDate":"March 15th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2022-08-09T10:29:24+00:00","versionOfRecord":[],"versionCreatedAt":"2022-03-15 14:38:52","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-1431427","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1431427","identity":"rs-1431427","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.