Image based Natural Scene Text Segmentation and Classification using Enhanced Retrieval and Optimization Technique

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract The enhancement of Optical Character Recognition (OCR) is a lively research domain for numerous well-known applications. These applications are automation of logistics, reading sign boards, helping visually impaired subjects, and retrieval of textual content from natural scene images. The foremost objective of OCR in these applications is to acquire significant performance in order to convert scene text into machine-readable text with adequate accuracy. Henceforth, an improved technique is proposed based on image processing for segmenting textual content (foreground) from unstructured and unconstrained backgrounds. Further, a multimodal feature fusion and selection is carried out for classification. These features are serially fused and out passed to Crow Search Optimization (CSO) in order to generate a salient feature vector using entropy as fitness function. After that, the optimal features are used by ensemble classifiers for classification. The presented work is comprehensively tested on three mainstream challenging datasets SVT, MSRA-TD500, and KAIST consisting several indoor/outdoor natural images with harsh scenic variabilities. The experimental results with existing benchmark techniques confirm that the presented technique performs better in retrieving textual content. Similarly, in the classification process, the proposed technique works fine on different experiments due to the novel optimization method. Finally, the presented optimization method is compared with its counterparts with the same setup and it is found that it also performs better.
Full text 11,262 characters · extracted from preprint-html · click to expand
Image based Natural Scene Text Segmentation and Classification using Enhanced Retrieval and Optimization Technique | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Image based Natural Scene Text Segmentation and Classification using Enhanced Retrieval and Optimization Technique Ghulam Jillani Ansari, Shahbaz Hassan Wasti, Syed Imran Abbas Qazmi, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6910940/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The enhancement of Optical Character Recognition (OCR) is a lively research domain for numerous well-known applications. These applications are automation of logistics, reading sign boards, helping visually impaired subjects, and retrieval of textual content from natural scene images. The foremost objective of OCR in these applications is to acquire significant performance in order to convert scene text into machine-readable text with adequate accuracy. Henceforth, an improved technique is proposed based on image processing for segmenting textual content (foreground) from unstructured and unconstrained backgrounds. Further, a multimodal feature fusion and selection is carried out for classification. These features are serially fused and out passed to Crow Search Optimization (CSO) in order to generate a salient feature vector using entropy as fitness function. After that, the optimal features are used by ensemble classifiers for classification. The presented work is comprehensively tested on three mainstream challenging datasets SVT, MSRA-TD500, and KAIST consisting several indoor/outdoor natural images with harsh scenic variabilities. The experimental results with existing benchmark techniques confirm that the presented technique performs better in retrieving textual content. Similarly, in the classification process, the proposed technique works fine on different experiments due to the novel optimization method. Finally, the presented optimization method is compared with its counterparts with the same setup and it is found that it also performs better. Guided image filtering Segmentation LTP CSO OCR Natural Scene Text Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6910940","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":475497723,"identity":"4937b38f-8053-4141-8c2a-616547e5fccd","order_by":0,"name":"Ghulam Jillani Ansari","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABDUlEQVRIiWNgGAWjYDACCQbGAwwMbDx87EBO4j8bOZDggQf4tTCAtbAxAzkJbGnGYC0JhLUANTGDycOJDSAanxb56B6DAz9z+GTYmJmfPXjAk5Y+P+zwQ6AtdnK6Ddi1GN45Y3CwdxvIYWzmBgkSNrkbb6cZALUkG5sdwKFlRo7BAV6wFgYziQSDtNyNsxNAWg4kbsOj5eBfsBb2bxIJCYfTDWenf8CrRV4ix+AwxBYeoC0HDifIS+fgt8VAIq3gsCxES5lEYkOa4QbpnIIDCQa4/SI/I3njw7fbjtnzs7dvk/zZYCMvPzt984cPFXZyuLQYQMSPoYsYYFcOtqUBTNVgiIyCUTAKRsEogAMAt9NdZVXGazwAAAAASUVORK5CYII=","orcid":"","institution":"University of Education","correspondingAuthor":true,"prefix":"","firstName":"Ghulam","middleName":"Jillani","lastName":"Ansari","suffix":""},{"id":475497724,"identity":"670e2776-43af-4ece-b95a-1fd3fff96ecf","order_by":1,"name":"Shahbaz Hassan Wasti","email":"","orcid":"","institution":"University of Education","correspondingAuthor":false,"prefix":"","firstName":"Shahbaz","middleName":"Hassan","lastName":"Wasti","suffix":""},{"id":475497725,"identity":"2497116e-951f-44fd-9a31-15f1b6749a3d","order_by":2,"name":"Syed Imran Abbas Qazmi","email":"","orcid":"","institution":"University of Education","correspondingAuthor":false,"prefix":"","firstName":"Syed","middleName":"Imran Abbas","lastName":"Qazmi","suffix":""},{"id":475497726,"identity":"9b315029-2774-4a4b-b87e-ed827583fcdc","order_by":3,"name":"Muhammad Jawad Hussain","email":"","orcid":"","institution":"Nantong University","correspondingAuthor":false,"prefix":"","firstName":"Muhammad","middleName":"Jawad","lastName":"Hussain","suffix":""}],"badges":[],"createdAt":"2025-06-17 06:38:09","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6910940/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6910940/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":98080258,"identity":"8535616c-12cd-4143-9519-8e5a7d6d2888","added_by":"auto","created_at":"2025-12-12 14:39:40","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1423960,"visible":true,"origin":"","legend":"","description":"","filename":"RevisedManuscriptV4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6910940/v1_covered_a273fb34-bc08-4dcf-8078-1a8041efafd9.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Image based Natural Scene Text Segmentation and Classification using Enhanced Retrieval and Optimization Technique","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Guided image filtering, Segmentation, LTP, CSO, OCR, Natural Scene Text","lastPublishedDoi":"10.21203/rs.3.rs-6910940/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6910940/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe enhancement of Optical Character Recognition (OCR) is a lively research domain for numerous well-known applications. These applications are automation of logistics, reading sign boards, helping visually impaired subjects, and retrieval of textual content from natural scene images. The foremost objective of OCR in these applications is to acquire significant performance in order to convert scene text into machine-readable text with adequate accuracy. Henceforth, an improved technique is proposed based on image processing for segmenting textual content (foreground) from unstructured and unconstrained backgrounds. Further, a multimodal feature fusion and selection is carried out for classification. These features are serially fused and out passed to Crow Search Optimization (CSO) in order to generate a salient feature vector using entropy as fitness function. After that, the optimal features are used by ensemble classifiers for classification. The presented work is comprehensively tested on three mainstream challenging datasets SVT, MSRA-TD500, and KAIST consisting several indoor/outdoor natural images with harsh scenic variabilities. The experimental results with existing benchmark techniques confirm that the presented technique performs better in retrieving textual content. Similarly, in the classification process, the proposed technique works fine on different experiments due to the novel optimization method. Finally, the presented optimization method is compared with its counterparts with the same setup and it is found that it also performs better.\u003c/p\u003e","manuscriptTitle":"Image based Natural Scene Text Segmentation and Classification using Enhanced Retrieval and Optimization Technique","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-25 11:20:48","doi":"10.21203/rs.3.rs-6910940/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"76f39885-4321-4498-bf46-9f26f7e11c11","owner":[],"postedDate":"June 25th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-12-12T14:39:07+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-25 11:20:48","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6910940","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6910940","identity":"rs-6910940","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-02T02:00:03.124865+00:00
License: CC-BY-4.0