Augmenting thesis supervision with generative AI feedback: Evaluating the quality and perceived usefulness GenAI feedback

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Delivering timely, high-quality feedback on long-form academic student writing is hard to scale. Therefore, we developed a generative AI feedback tool for undergraduate medical theses. Feedback was based on 36 criteria, and feedback quality was evaluated on six aspects. Feedback consistency was assessed across five repeated iterations, robustness across thesis types and grades, and compared GenAI with human feedback. Finally, we examined students’ perceived usefulness and added value. Thirteen theses (9 systematic reviews, 2 narrative reviews, 2 empirical studies) were evaluated, yielding 14.040 feedback aspect ratings. Median GenAI feedback quality was 1.0 (IQR=0.167, scale 0-1), with most criteria exceeding the set threshold of 0.8. Consistency surpassed 0.8 for the quality aspects of feedback, feedforward, thesis-specific, and criterion-based; thesis-accurate and criterion-accurate did not meet the threshold. Quality and consistency did not differ by thesis type or grade. Relative to 215 human feedback comments on the same theses, GenAI feedback showed higher median quality (1.00 vs. 0.33). During in-class sessions, students (n = 30) rated the feedback as useful, sufficient, and of added value (median 6–7, scale 1–7) and favored future use as a supplement rather than a replacement for supervisor feedback. Limitations of the developed GenAI-tool include difficulty determining thesis type, inability to assess figures and cross-references across chapters, reducing feedback quality. Overall, findings support a hybrid model: GenAI offers rapid, comprehensive formative feedback, while educators provide contextual expertise, literature awareness, and mentorship. This division of labor expands access to actionable feedback at scale without displacing essential human judgment.
Full text 12,636 characters · extracted from preprint-html · click to expand
Augmenting thesis supervision with generative AI feedback: Evaluating the quality and perceived usefulness GenAI feedback | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Augmenting thesis supervision with generative AI feedback: Evaluating the quality and perceived usefulness GenAI feedback Remco Jongkind, Myrthe Heikens, Lotte Barmentloo, Erik Elings, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8352038/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Delivering timely, high-quality feedback on long-form academic student writing is hard to scale. Therefore, we developed a generative AI feedback tool for undergraduate medical theses. Feedback was based on 36 criteria, and feedback quality was evaluated on six aspects. Feedback consistency was assessed across five repeated iterations, robustness across thesis types and grades, and compared GenAI with human feedback. Finally, we examined students’ perceived usefulness and added value. Thirteen theses (9 systematic reviews, 2 narrative reviews, 2 empirical studies) were evaluated, yielding 14.040 feedback aspect ratings. Median GenAI feedback quality was 1.0 (IQR=0.167, scale 0-1), with most criteria exceeding the set threshold of 0.8. Consistency surpassed 0.8 for the quality aspects of feedback, feedforward, thesis-specific, and criterion-based; thesis-accurate and criterion-accurate did not meet the threshold. Quality and consistency did not differ by thesis type or grade. Relative to 215 human feedback comments on the same theses, GenAI feedback showed higher median quality (1.00 vs. 0.33). During in-class sessions, students (n = 30) rated the feedback as useful, sufficient, and of added value (median 6–7, scale 1–7) and favored future use as a supplement rather than a replacement for supervisor feedback. Limitations of the developed GenAI-tool include difficulty determining thesis type, inability to assess figures and cross-references across chapters, reducing feedback quality. Overall, findings support a hybrid model: GenAI offers rapid, comprehensive formative feedback, while educators provide contextual expertise, literature awareness, and mentorship. This division of labor expands access to actionable feedback at scale without displacing essential human judgment. Generative AI thesis supervision AI feedback quality human-AI collaboration formative assessment higher education Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8352038","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":562004107,"identity":"d7169bc6-26e3-410a-960c-96336e9bf5dc","order_by":0,"name":"Remco Jongkind","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABC0lEQVRIiWNgGAWjYLACCRBmBuIPSAIgwHiAkBbJGUAGD5IWBpxaYNqkeYjRYt7A/IDBouKO3cx23oO3bWq22duz9z688TPncDR/A+8BbFpkDrAZMEiceZY8m5kv2Trn2O3EHp7jxpa92w7nzjjAl4BNiwTIIZJth5PlmHnMpHPYbifwSKSxSfACtWxg4DEgrMXi3217HvlnbJJ/idBiJw3Swth2m7FHgo1NGq8tzGwGByTOHE6QbOYBeqEP6JczaczWstvSc2ccxqGFvfnhY4mKw/YS588Y3vjx7bY9e/sxxptvt1nn9rf3GD7AooUBGIOHgdGQ2IBVChdgBKYTe9zSo2AUjIJRMOIBAEDDXPSrLiqtAAAAAElFTkSuQmCC","orcid":"","institution":"Amsterdam UMC, location AMC","correspondingAuthor":true,"prefix":"","firstName":"Remco","middleName":"","lastName":"Jongkind","suffix":""},{"id":562004111,"identity":"fea84a7e-fb8f-420f-8356-23f106db2727","order_by":1,"name":"Myrthe Heikens","email":"","orcid":"","institution":"Amsterdam UMC, location AMC","correspondingAuthor":false,"prefix":"","firstName":"Myrthe","middleName":"","lastName":"Heikens","suffix":""},{"id":562004115,"identity":"cca86e3c-f69c-41a3-aa8b-c93efaf84259","order_by":2,"name":"Lotte Barmentloo","email":"","orcid":"","institution":"Amsterdam UMC, location AMC","correspondingAuthor":false,"prefix":"","firstName":"Lotte","middleName":"","lastName":"Barmentloo","suffix":""},{"id":562004119,"identity":"080f6c97-208f-4706-8242-e1899ee9a2cb","order_by":3,"name":"Erik Elings","email":"","orcid":"","institution":"Amsterdam UMC, location AMC","correspondingAuthor":false,"prefix":"","firstName":"Erik","middleName":"","lastName":"Elings","suffix":""},{"id":562004123,"identity":"1893a4c4-0d19-4241-8106-8f80fc2dfb03","order_by":4,"name":"Floor van der Steijle","email":"","orcid":"","institution":"Amsterdam UMC, location AMC","correspondingAuthor":false,"prefix":"","firstName":"Floor","middleName":"van der","lastName":"Steijle","suffix":""},{"id":562004127,"identity":"d3802ab4-c3ad-4bd6-9d34-d472cbfe2ab3","order_by":5,"name":"Lisa-Maria van Klaveren","email":"","orcid":"","institution":"Amsterdam UMC, location AMC","correspondingAuthor":false,"prefix":"","firstName":"Lisa-Maria","middleName":"van","lastName":"Klaveren","suffix":""}],"badges":[],"createdAt":"2025-12-13 10:23:31","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-8352038/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8352038/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105035068,"identity":"a800bf3d-0258-4794-8ee0-ec581b31fd5c","added_by":"auto","created_at":"2026-03-20 07:25:25","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1586514,"visible":true,"origin":"","legend":"","description":"","filename":"AI4Feedbackpaperv13122025RemcoJongkindetal.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8352038/v1_covered_4fdc461a-e294-4161-9913-1d28f234b2e5.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Augmenting thesis supervision with generative AI feedback: Evaluating the quality and perceived usefulness GenAI feedback","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Generative AI, thesis supervision, AI feedback quality, human-AI collaboration, formative assessment, higher education","lastPublishedDoi":"10.21203/rs.3.rs-8352038/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8352038/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Delivering timely, high-quality feedback on long-form academic student writing is hard to scale. Therefore, we developed a generative AI feedback tool for undergraduate medical theses. Feedback was based on 36 criteria, and feedback quality was evaluated on six aspects. Feedback consistency was assessed across five repeated iterations, robustness across thesis types and grades, and compared GenAI with human feedback. Finally, we examined students’ perceived usefulness and added value. \nThirteen theses (9 systematic reviews, 2 narrative reviews, 2 empirical studies) were evaluated, yielding 14.040 feedback aspect ratings. Median GenAI feedback quality was 1.0 (IQR=0.167, scale 0-1), with most criteria exceeding the set threshold of 0.8. Consistency surpassed 0.8 for the quality aspects of feedback, feedforward, thesis-specific, and criterion-based; thesis-accurate and criterion-accurate did not meet the threshold. Quality and consistency did not differ by thesis type or grade. Relative to 215 human feedback comments on the same theses, GenAI feedback showed higher median quality (1.00 vs. 0.33). During in-class sessions, students (n = 30) rated the feedback as useful, sufficient, and of added value (median 6–7, scale 1–7) and favored future use as a supplement rather than a replacement for supervisor feedback. \nLimitations of the developed GenAI-tool include difficulty determining thesis type, inability to assess figures and cross-references across chapters, reducing feedback quality. Overall, findings support a hybrid model: GenAI offers rapid, comprehensive formative feedback, while educators provide contextual expertise, literature awareness, and mentorship. This division of labor expands access to actionable feedback at scale without displacing essential human judgment. ","manuscriptTitle":"Augmenting thesis supervision with generative AI feedback: Evaluating the quality and perceived usefulness GenAI feedback","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-10 12:17:29","doi":"10.21203/rs.3.rs-8352038/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"cb671eca-61ad-478a-91aa-2d772488db5b","owner":[],"postedDate":"February 10th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-03-19T08:27:24+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-10 12:17:29","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8352038","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8352038","identity":"rs-8352038","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00