Research on Unbalanced Sample Classification Algorithm Based on Redundant Data Elimination

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

This study proposes a redundant data elimination method using a subspace fusion filter and autocorrelation feature analysis to improve SVM classification accuracy for unlabeled unbalanced sample data in multi-user social networks.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-17 · read from full text

This preprint studies an unlabeled, unbalanced-sample classification problem in multi-user social network settings, proposing a “redundant data elimination” approach to reduce classification error. Using a subspace fusion filter detection model, the method aims to remove redundant data and suppress interference on prior feature information, then extracts clustering convergence feature parameters via statistical analysis and distributed fusion of autocorrelation features, and estimates classification feature quantity through constrained multi-level iterative regression. Classification is performed with an SVM using characteristic target parameters combined with an adaptive weight distribution control and a dynamic learning algorithm. The authors report simulation results with higher classification fidelity and lower error rates, with the main caveat that the work is presented as an unreviewed preprint. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract The unlabeled data in multi-user social network is unbalanced, which leads to poor classification in the process of data scheduling and detection. In order to reduce the error rate of unlabeled unbalanced sample data classification in multi-user social network, a method of unlabeled unbalanced sample data classification in multi-user social network based on redundant data elimination is proposed. The subspace fusion filter detection model is used to filter redundant data and suppress anti-interference on the prior features of randomly sampled unlabeled unbalanced sample data of multi-user social network, the statistical average analysis and distributed fusion detection method of autocorrelation features are used to extract the clustering convergence feature parameters of unlabeled unbalanced sample data of multi-user social network, and the constrained evolution method of multi-level iterative regression analysis is used to estimate the classification feature quantity of unlabeled unbalanced sample data. The characteristic parameters of the classification target are input into the SVM classifier, and the adaptive weight distribution control of SVM classification is carried out in combination with the dynamic learning algorithm, which improves the adaptability of data classification and realizes the optimal classification of unlabeled unbalanced sample data in multi-user social networks. The simulation results show that this algorithm can effectively reduce the interference of redundant data when classifying unlabeled unbalanced sample data of multi-user social network, and the fidelity rate of data classification is high, and the error rate is low, which improves the dynamic data management ability of multi-user social network.
Full text 11,627 characters · extracted from preprint-html · click to expand
Research on Unbalanced Sample Classification Algorithm Based on Redundant Data Elimination | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Research on Unbalanced Sample Classification Algorithm Based on Redundant Data Elimination TINGTING YU, YU SUN This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3198187/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 4 You are reading this latest preprint version Abstract The unlabeled data in multi-user social network is unbalanced, which leads to poor classification in the process of data scheduling and detection. In order to reduce the error rate of unlabeled unbalanced sample data classification in multi-user social network, a method of unlabeled unbalanced sample data classification in multi-user social network based on redundant data elimination is proposed. The subspace fusion filter detection model is used to filter redundant data and suppress anti-interference on the prior features of randomly sampled unlabeled unbalanced sample data of multi-user social network, the statistical average analysis and distributed fusion detection method of autocorrelation features are used to extract the clustering convergence feature parameters of unlabeled unbalanced sample data of multi-user social network, and the constrained evolution method of multi-level iterative regression analysis is used to estimate the classification feature quantity of unlabeled unbalanced sample data. The characteristic parameters of the classification target are input into the SVM classifier, and the adaptive weight distribution control of SVM classification is carried out in combination with the dynamic learning algorithm, which improves the adaptability of data classification and realizes the optimal classification of unlabeled unbalanced sample data in multi-user social networks. The simulation results show that this algorithm can effectively reduce the interference of redundant data when classifying unlabeled unbalanced sample data of multi-user social network, and the fidelity rate of data classification is high, and the error rate is low, which improves the dynamic data management ability of multi-user social network. Redundant data Eliminate Unbalanced sample Classification Support vector machine Fuzzy clustering Social network Full Text Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 04 Aug, 2023 Reviewers invited by journal 04 Aug, 2023 Editor assigned by journal 01 Aug, 2023 First submitted to journal 31 Jul, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3198187","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":223946709,"identity":"f34c24a3-029e-4b11-912e-20427803c697","order_by":0,"name":"TINGTING YU","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAq0lEQVRIiWNgGAWjYLCCDwwMPCBagkj1zIyNM4BaeEjS0gxSTrwWgxv5xx/bth2WsWdgPnibh8EujwgtyYzNuW1pQIexJVvzMCQXE6llmw1QC4+ZNA/DgcQGorRYbpMAauH/RoIWRogtbMRpkTzz2HBm7z+gXw6zGVvOMUgmrIXveOKDDz/OHLZnb29+eONNhR1hLQoHYCxmsDsJqQcCeYKGjoJRMApGwSgAAGxhNCLGIkCyAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0003-0135-4452","institution":"Wuxi Institute of Technology","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"TINGTING","middleName":"","lastName":"YU","suffix":""},{"id":223946710,"identity":"b4ae1a4e-0394-4fe2-8fe3-8e6b052ba3f6","order_by":1,"name":"YU SUN","email":"","orcid":"","institution":"Wuxi Institute of Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"YU","middleName":"","lastName":"SUN","suffix":""}],"badges":[],"createdAt":"2023-07-24 06:41:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3198187/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3198187/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":41327081,"identity":"00489bc8-a022-470f-8793-4421e4efdcdd","added_by":"auto","created_at":"2023-08-09 18:49:28","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":367649,"visible":true,"origin":"","legend":"","description":"","filename":"TINGTINGYU.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3198187/v1_covered_4249340c-c556-4e79-b2c6-cc8171e2bfe1.pdf"}],"financialInterests":"","formattedTitle":"Research on Unbalanced Sample Classification Algorithm Based on Redundant Data Elimination","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"soft-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"soco","sideBox":"Learn more about [Soft Computing](https://www.springer.com/journal/500)","snPcode":"500","submissionUrl":"https://submission.nature.com/new-submission/500/3","title":"Soft Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Redundant data, Eliminate, Unbalanced sample, Classification, Support vector machine, Fuzzy clustering, Social network","lastPublishedDoi":"10.21203/rs.3.rs-3198187/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3198187/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe unlabeled data in multi-user social network is unbalanced, which leads to poor classification in the process of data scheduling and detection. In order to reduce the error rate of unlabeled unbalanced sample data classification in multi-user social network, a method of unlabeled unbalanced sample data classification in multi-user social network based on redundant data elimination is proposed. The subspace fusion filter detection model is used to filter redundant data and suppress anti-interference on the prior features of randomly sampled unlabeled unbalanced sample data of multi-user social network, the statistical average analysis and distributed fusion detection method of autocorrelation features are used to extract the clustering convergence feature parameters of unlabeled unbalanced sample data of multi-user social network, and the constrained evolution method of multi-level iterative regression analysis is used to estimate the classification feature quantity of unlabeled unbalanced sample data. The characteristic parameters of the classification target are input into the SVM classifier, and the adaptive weight distribution control of SVM classification is carried out in combination with the dynamic learning algorithm, which improves the adaptability of data classification and realizes the optimal classification of unlabeled unbalanced sample data in multi-user social networks. The simulation results show that this algorithm can effectively reduce the interference of redundant data when classifying unlabeled unbalanced sample data of multi-user social network, and the fidelity rate of data classification is high, and the error rate is low, which improves the dynamic data management ability of multi-user social network.\u003c/p\u003e","manuscriptTitle":"Research on Unbalanced Sample Classification Algorithm Based on Redundant Data Elimination","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-08-09 18:41:18","doi":"10.21203/rs.3.rs-3198187/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"","date":"2023-08-04T05:40:42+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-08-04T04:11:32+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-08-02T02:51:51+00:00","index":"","fulltext":""},{"type":"submitted","content":"Soft Computing","date":"2023-07-31T06:41:18+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"soft-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"soco","sideBox":"Learn more about [Soft Computing](https://www.springer.com/journal/500)","snPcode":"500","submissionUrl":"https://submission.nature.com/new-submission/500/3","title":"Soft Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"199388d6-3250-4e9a-a122-3652d303e378","owner":[],"postedDate":"August 9th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2023-09-17T06:23:28+00:00","versionOfRecord":[],"versionCreatedAt":"2023-08-09 18:41:18","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3198187","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3198187","identity":"rs-3198187","version":["v1"]},"buildId":"FbvkV6FR0MCFSLy54lSbu","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-30T02:00:01.510937+00:00
License: CC-BY-4.0