A Novel Retrieval-Augmented Generation Framework Using Large Language Models for Lyrics and Song Composition | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Novel Retrieval-Augmented Generation Framework Using Large Language Models for Lyrics and Song Composition Veerababu Reddy, Veeranjaneyulu N This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6390724/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The use of artificial intelligence (AI) in music composition is now very common, but present models have the flaw of losing lyrical coherence, stylistic coherence, and vocal realism. This paper presents a novel AI-driven framework for automated song composition that integrates Natural Language Processing (NLP), Large Language Models (LLMs), and multimodal synthesis techniques. The system employs a hybrid retrieval pipeline that combines BM25 for sparse lexical matching and FAISS for dense semantic search. This Retrieval-Augmented Generation (RAG) approach, powered by GPT-4o-mini, enables the generation of lyrics that reflect the thematic consistency and stylistic nuances of an artist's previous compositions. For vocal synthesis, speaker embeddings are extracted using models such as Wav2Vec2, ECAPA-TDNN, and DeepSpeaker, enabling high-fidelity voice cloning. Synthesized vocals are further refined using the Bark and Suno models to produce expressive singing voices. Instrumental backgrounds are generated using text-to-audio (TTA) models with synchronized beats-per-minute (BPM) and melodic patterns to maintain musical harmony. The final output undergoes noise reduction, vocal-instrumental alignment, and automated mixing to produce studio-quality audio. Experimental evaluation shows improved lyrical relevance, fluency and vocal expressiveness, with an increase in the BERTScore, a lower perplexity, and stylistic alignment accuracy of 92.1%. This system demonstrates a significant advancement in AI-assisted songwriting, offering a scalable and musically coherent solution for end-to-end song generation. Artificial Intelligence (AI) NaturalLanguage Processing (NLP) Large Language Models (LLMs) AI songwriting Retrieval-Augmented Generation (RAG) AI Lyric Generation Speaker Embedding Text-to-Audio (TTA) Music Composition BM25 FAISS Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6390724","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":439286774,"identity":"14d9222d-6f45-46fe-8de1-3ee64eef0850","order_by":0,"name":"Veerababu Reddy","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABHklEQVRIie3RMWuDQBTA8SeCLgZXhWK+womQDKX4VU6Ezo4dMpwIdpHMfoxMLdkMD5JF6OrgYBFculgCKYUQehYCQkylW6H3H94dx/tNByAS/d0yIN1R8amDxHrvY4QSAib7HeEH+WHtu3nqb+qHRTmdG1QxaHB0ncKP3oJFCfpjJmFwSW6Ke9/Ot429TjtCiPdUePFtum3AyClgekkMLZ+ZTEFpVVCZcEJn+SZ2NAUBCgDUhsjLxyc7oXsmrpN05IQwvUbURJHCGD1OpIoTaaWGUT2JEchVEjtmuGz8dVJ1xPHSXRjLkyVqdu6xQSLLr+/sUN49qxSy9mi5eqTWe+2AlrVD3A+QXnp7vikGH3y596djye34jkgkEv2jvgAni21Z0IDXAQAAAABJRU5ErkJggg==","orcid":"","institution":"Vignan's Foundation for Science, Technology \u0026 Research","correspondingAuthor":true,"prefix":"","firstName":"Veerababu","middleName":"","lastName":"Reddy","suffix":""},{"id":439286776,"identity":"5483c4f7-0938-432a-97a6-2ef76767e0c7","order_by":1,"name":"Veeranjaneyulu N","email":"","orcid":"","institution":"Vignan's Foundation for Science, Technology \u0026 Research","correspondingAuthor":false,"prefix":"","firstName":"Veeranjaneyulu","middleName":"","lastName":"N","suffix":""}],"badges":[],"createdAt":"2025-04-07 06:23:27","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6390724/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6390724/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":80787098,"identity":"b0ad4c32-5512-4e6b-b5c4-05ccef9d857d","added_by":"auto","created_at":"2025-04-17 06:03:09","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2309111,"visible":true,"origin":"","legend":"","description":"","filename":"ANovelRetrievalAugmentedGenerationFrameworkUsingLargeLanguageModelsforLyricsandSongCompo.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6390724/v1_covered_d6130c6f-2b7b-427c-9d9d-6f469d226d4f.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Novel Retrieval-Augmented Generation Framework Using Large Language Models for Lyrics and Song Composition","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Artificial Intelligence (AI), NaturalLanguage Processing (NLP), Large Language Models (LLMs), AI songwriting, Retrieval-Augmented Generation (RAG), AI Lyric Generation, Speaker Embedding, Text-to-Audio (TTA), Music Composition, BM25, FAISS","lastPublishedDoi":"10.21203/rs.3.rs-6390724/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6390724/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe use of artificial intelligence (AI) in music composition is now very common, but present models have the flaw of losing lyrical coherence, stylistic coherence, and vocal realism. This paper presents a novel AI-driven framework for automated song composition that integrates Natural Language Processing (NLP), Large Language Models (LLMs), and multimodal synthesis techniques. The system employs a hybrid retrieval pipeline that combines BM25 for sparse lexical matching and FAISS for dense semantic search. This Retrieval-Augmented Generation (RAG) approach, powered by GPT-4o-mini, enables the generation of lyrics that reflect the thematic consistency and stylistic nuances of an artist's previous compositions. For vocal synthesis, speaker embeddings are extracted using models such as Wav2Vec2, ECAPA-TDNN, and DeepSpeaker, enabling high-fidelity voice cloning. Synthesized vocals are further refined using the Bark and Suno models to produce expressive singing voices. Instrumental backgrounds are generated using text-to-audio (TTA) models with synchronized beats-per-minute (BPM) and melodic patterns to maintain musical harmony. The final output undergoes noise reduction, vocal-instrumental alignment, and automated mixing to produce studio-quality audio. Experimental evaluation shows improved lyrical relevance, fluency and vocal expressiveness, with an increase in the BERTScore, a lower perplexity, and stylistic alignment accuracy of 92.1%. This system demonstrates a significant advancement in AI-assisted songwriting, offering a scalable and musically coherent solution for end-to-end song generation.\u003c/p\u003e","manuscriptTitle":"A Novel Retrieval-Augmented Generation Framework Using Large Language Models for Lyrics and Song Composition","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-08 03:51:07","doi":"10.21203/rs.3.rs-6390724/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b3988f64-c7aa-4c86-bd18-1685f20b5173","owner":[],"postedDate":"April 8th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-04-17T05:38:54+00:00","versionOfRecord":[],"versionCreatedAt":"2025-04-08 03:51:07","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6390724","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6390724","identity":"rs-6390724","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.