Towards Text Contextual Understanding: Text Feature Fusion GAN for Text-to-Image Generation | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Towards Text Contextual Understanding: Text Feature Fusion GAN for Text-to-Image Generation Xiaoyan Jiang, Jize Chen, Juan Zhang, Zhichao Chen, Yongbin Gao, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4551157/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Understanding texts and generating high-resolution realistic images from text descriptions is a meaningful and challenging topic. Most existing text-to-image generation methods follow the multi-stage generative adversarial network (GAN) framework. Therefore, the image quality of current stages relies heavily on the images generated in the previous stages. Additionally, current text-to-image generators do not so well guarantee a consistent relationship between the text description and the generated image. To address the above issues, we present a novel Text Feature Fusion GAN (TF-GAN) architecture, emphasizing on selecting useful local words and extracting semantic consistent global sentence features. First, useful keywords are extracted in the text description by the proposed sentence fusion attention mechanism (SFAttn) to optimize image features in early stages and provide fine-grained details for the images of the later stage. Second, we propose a new Conditional Fusion Block (CFBlock), which combines multiple non-linear affine transformations to constrain images on a global semantic level. This allows the model to more deeply fuse sentence and image features, leading to more semantically consistent generated images. Extensive experiments and Comparison with other state-of-the-art methods on the Caltech-UCSD Birds 200 (CUB) dataset and the Microsoft Common Objects in Context (COCO) dataset show that our generated images are more photo-realistic and closer to the text descriptions. GAN Gate Mechanism Feature Fusion Text-to-Image Synthesis Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4551157","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":315723042,"identity":"8d840b92-d544-4cb4-9620-cdc9baa0f802","order_by":0,"name":"Xiaoyan Jiang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABEklEQVRIiWNgGAWjYDACZijN2AwiK6A8HuK1nCFGCwpgbCNCi8Fx5oePeWruMDC3Mz97+HVeXWI//wHGB2/bGOTNcWiRbGYzNuY59gzoMDZzY9lthxNnzkhgNpzbxmC4swG7Fn5mBjNpHrbDIL+YSUtuO5C44QYDmzRvG0OCwQHsWtiY2b9J8/wDaQEyJOfUJe4/f4D9Nz4t/Mw8ZkAzQVp4zCQ/NjAnbmBIYGPGp0WymafYcG7fYR6gljJphmOHjWfcSGyWnHNOwnADDi0G549vfPDm22E5w/7j2yR/1NTJ9vcfPvjhTZmNPC5bQIAJGAs8hg3AaIVEByOQySCBWz1IyQ8gIQ9jjIJRMApGwShABwCCYFQvCwfJGgAAAABJRU5ErkJggg==","orcid":"","institution":"Shanghai University of Engineering Science","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Xiaoyan","middleName":"","lastName":"Jiang","suffix":""},{"id":315723046,"identity":"774c2cca-67f6-453e-8dad-33eb77943425","order_by":1,"name":"Jize Chen","email":"","orcid":"","institution":"Shanghai University of Engineering Science","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jize","middleName":"","lastName":"Chen","suffix":""},{"id":315723049,"identity":"414baa2b-5a8e-4bad-82a4-42ddf4c97934","order_by":2,"name":"Juan Zhang","email":"","orcid":"","institution":"Shanghai University of Engineering Science","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Juan","middleName":"","lastName":"Zhang","suffix":""},{"id":315723052,"identity":"d186dc3e-57bf-4f3c-b127-80bb1c8f9de4","order_by":3,"name":"Zhichao Chen","email":"","orcid":"","institution":"Shanghai Jiao Tong University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Zhichao","middleName":"","lastName":"Chen","suffix":""},{"id":315723057,"identity":"952be628-db46-4d6d-a988-1f2887cb35c7","order_by":4,"name":"Yongbin Gao","email":"","orcid":"","institution":"Shanghai University of Engineering Science","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yongbin","middleName":"","lastName":"Gao","suffix":""},{"id":315723059,"identity":"a6b66e81-8abb-4d28-9e4f-d4ea72956b6d","order_by":5,"name":"Xiang Wu","email":"","orcid":"","institution":"Xuzhou Medical University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xiang","middleName":"","lastName":"Wu","suffix":""}],"badges":[],"createdAt":"2024-06-08 15:40:03","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4551157/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4551157/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":60937391,"identity":"2c3b8f87-8146-4751-8940-5b36c7ee2dda","added_by":"auto","created_at":"2024-07-23 19:44:32","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1849198,"visible":true,"origin":"","legend":"","description":"","filename":"TFGANMachineVisionandApplication4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4551157/v1_covered_c2142abf-37d8-4114-b5c6-39689846e7fc.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Towards Text Contextual Understanding: Text Feature Fusion GAN for Text-to-Image Generation","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"GAN, Gate Mechanism, Feature Fusion, Text-to-Image Synthesis","lastPublishedDoi":"10.21203/rs.3.rs-4551157/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4551157/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eUnderstanding texts and generating high-resolution realistic images from text descriptions is a meaningful and challenging topic. Most existing text-to-image generation methods follow the multi-stage generative adversarial network (GAN) framework. Therefore, the image quality of current stages relies heavily on the images generated in the previous stages. Additionally, current text-to-image generators do not so well guarantee a consistent relationship between the text description and the generated image. To address the above issues, we present a novel Text Feature Fusion GAN (TF-GAN) architecture, emphasizing on selecting useful local words and extracting semantic consistent global sentence features. First, useful keywords are extracted in the text description by the proposed sentence fusion attention mechanism (SFAttn) to optimize image features in early stages and provide fine-grained details for the images of the later stage. Second, we propose a new Conditional Fusion Block (CFBlock), which combines multiple non-linear affine transformations to constrain images on a global semantic level. This allows the model to more deeply fuse sentence and image features, leading to more semantically consistent generated images. Extensive experiments and Comparison with other state-of-the-art methods on the Caltech-UCSD Birds 200 (CUB) dataset and the Microsoft Common Objects in Context (COCO) dataset show that our generated images are more photo-realistic and closer to the text descriptions.\u003c/p\u003e","manuscriptTitle":"Towards Text Contextual Understanding: Text Feature Fusion GAN for Text-to-Image Generation","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-06-24 14:25:59","doi":"10.21203/rs.3.rs-4551157/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ca9f8b67-ff55-4a68-9f78-8b3fdd056b2d","owner":[],"postedDate":"June 24th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-07-23T19:36:24+00:00","versionOfRecord":[],"versionCreatedAt":"2024-06-24 14:25:59","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4551157","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4551157","identity":"rs-4551157","version":["v1"]},"buildId":"-HB7Z8yhvgn0wM9Nzuekk","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.