Optimization of Classifiers Performance for Node Embedding on Graph Based Data

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract The Publications regarding the research of embedding the large-scale information that helps in getting networks utilizing neighborhood-aware node representations and low-dimensional communities cover a wide area of research. In graph mining applications, these classification models, and embedding performed better than the conventional approaches. When using different conventional machine learning and data analysis approaches, the display of graphs and their relationship is highly useful in describing features present. Many different embedding approaches are used in machine learning, and a literature review was conducted to determine the best techniques for comparison. This study examines the accuracy scores of different classifiers using the approach on a single dataset. The dataset which is used in this study is CORA, and it is used to import it. After the network has been formed using the dataset, the nodes are embedded since the result of this node embedding will be used as a training set. The machine learns through training of model, for which the Node2vex method is applied in this work. The classifiers are used to train the model. Gradient Boosting, Logistic Regression, Random Forest, K-Neighbors, Decision Tree, Gaussian, and SVC are the classifiers utilized to solve this model's classification problem. To assess performance, the model makes use of two classifiers: Gradient Boosting, Logistic Regression, Random Forest, K-Neighbors, Decision Tree, Gaussian, and SVC. Through experimentation, the accuracy score is used to compare the classifier’s levels of efficiency. From the study, it was clearly observed that for the dataset, it was only the Support Vector Classifier that performed best in the testing and training of dataset for getting desired result. This was achieved by achieving an accuracy of 0.7706 and an MCC score of 0.7200. The optimum classifier for model training tasks and node classification can be chosen with the aid of this paper.
Full text 11,451 characters · extracted from preprint-html · click to expand
Optimization of Classifiers Performance for Node Embedding on Graph Based Data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Optimization of Classifiers Performance for Node Embedding on Graph Based Data Neha Yadav, Dhanalekshmi Gopinathan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4426787/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The Publications regarding the research of embedding the large-scale information that helps in getting networks utilizing neighborhood-aware node representations and low-dimensional communities cover a wide area of research. In graph mining applications, these classification models, and embedding performed better than the conventional approaches. When using different conventional machine learning and data analysis approaches, the display of graphs and their relationship is highly useful in describing features present. Many different embedding approaches are used in machine learning, and a literature review was conducted to determine the best techniques for comparison. This study examines the accuracy scores of different classifiers using the approach on a single dataset. The dataset which is used in this study is CORA, and it is used to import it. After the network has been formed using the dataset, the nodes are embedded since the result of this node embedding will be used as a training set. The machine learns through training of model, for which the Node2vex method is applied in this work. The classifiers are used to train the model. Gradient Boosting, Logistic Regression, Random Forest, K-Neighbors, Decision Tree, Gaussian, and SVC are the classifiers utilized to solve this model's classification problem. To assess performance, the model makes use of two classifiers: Gradient Boosting, Logistic Regression, Random Forest, K-Neighbors, Decision Tree, Gaussian, and SVC. Through experimentation, the accuracy score is used to compare the classifier’s levels of efficiency. From the study, it was clearly observed that for the dataset, it was only the Support Vector Classifier that performed best in the testing and training of dataset for getting desired result. This was achieved by achieving an accuracy of 0.7706 and an MCC score of 0.7200. The optimum classifier for model training tasks and node classification can be chosen with the aid of this paper. Gradient Boosting classifier Node2vec Random Forest classifier Graph embedding machine learning Graph embeddings Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4426787","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":308735263,"identity":"f4c7d612-113f-4acb-87bc-f7436792ce90","order_by":0,"name":"Neha Yadav","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABAklEQVRIiWNgGAWjYBACNvaGhIMfDGzq+4GcAw8Y2HjYIBISOLXw8Rx4+FiiIo1xZgNQSwJICxsBLXISiY8NeM4cZtxwAMhLAFtMyGESyWkSkm3MzMbXDj8E2sInwyffwPjhB4NFHk4tPM/SJArb2NjMbqcZwBzGLNnDIFGMUwt7DsgWHh6z2wlwLQzSQL8kNuDSwpD/TYK3TULCeHb6B7gtv/Fq4UhIBnrfwMBAOgduCxt+W3gOJAIDOSFB4nZOwYEEA5CWxDbLHgPcWuTbwVH5P4F/dvrmDx8qjtnLNx8+fONHRR1OLWjA4BiQYAQqNiBOPQjUEK90FIyCUTAKRgwAAKMhTnsyxnK2AAAAAElFTkSuQmCC","orcid":"","institution":"Jaypee Institute of Information Technology","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Neha","middleName":"","lastName":"Yadav","suffix":""},{"id":308735264,"identity":"292996a3-e1c1-46ff-a0cb-b646602c9de1","order_by":1,"name":"Dhanalekshmi Gopinathan","email":"","orcid":"","institution":"Jaypee Institute of Information Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Dhanalekshmi","middleName":"","lastName":"Gopinathan","suffix":""}],"badges":[],"createdAt":"2024-05-15 17:57:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4426787/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4426787/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":58870052,"identity":"688a860b-dbb6-4a1c-adee-4ac200c63ace","added_by":"auto","created_at":"2024-06-22 17:31:30","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":619234,"visible":true,"origin":"","legend":"","description":"","filename":"journl21.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4426787/v1_covered_9c2a42e4-64c0-4986-88ef-f5dacae6c5f1.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Optimization of Classifiers Performance for Node Embedding on Graph Based Data","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Gradient Boosting classifier, Node2vec, Random Forest classifier, Graph embedding, machine learning, Graph embeddings","lastPublishedDoi":"10.21203/rs.3.rs-4426787/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4426787/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe Publications regarding the research of embedding the large-scale information that helps in getting networks utilizing neighborhood-aware node representations and low-dimensional communities cover a wide area of research. In graph mining applications, these classification models, and embedding performed better than the conventional approaches. When using different conventional machine learning and data analysis approaches, the display of graphs and their relationship is highly useful in describing features present. Many different embedding approaches are used in machine learning, and a literature review was conducted to determine the best techniques for comparison. This study examines the accuracy scores of different classifiers using the approach on a single dataset. The dataset which is used in this study is CORA, and it is used to import it. After the network has been formed using the dataset, the nodes are embedded since the result of this node embedding will be used as a training set. The machine learns through training of model, for which the Node2vex method is applied in this work. The classifiers are used to train the model. Gradient Boosting, Logistic Regression, Random Forest, K-Neighbors, Decision Tree, Gaussian, and SVC are the classifiers utilized to solve this model's classification problem. To assess performance, the model makes use of two classifiers: Gradient Boosting, Logistic Regression, Random Forest, K-Neighbors, Decision Tree, Gaussian, and SVC. Through experimentation, the accuracy score is used to compare the classifier\u0026rsquo;s levels of efficiency. From the study, it was clearly observed that for the dataset, it was only the Support Vector Classifier that performed best in the testing and training of dataset for getting desired result. This was achieved by achieving an accuracy of 0.7706 and an MCC score of 0.7200. The optimum classifier for model training tasks and node classification can be chosen with the aid of this paper.\u003c/p\u003e","manuscriptTitle":"Optimization of Classifiers Performance for Node Embedding on Graph Based Data","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-06-03 09:51:00","doi":"10.21203/rs.3.rs-4426787/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"99bffa30-b55c-480a-a284-c61472c7af3c","owner":[],"postedDate":"June 3rd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-06-22T17:23:22+00:00","versionOfRecord":[],"versionCreatedAt":"2024-06-03 09:51:00","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4426787","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4426787","identity":"rs-4426787","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0