BioDataHub: An Integrated VS Code Extension for Streamlined Bioinformatics Dataset Analysis and Visualization | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article BioDataHub: An Integrated VS Code Extension for Streamlined Bioinformatics Dataset Analysis and Visualization Mubashir Ali This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7861003/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Managing and analyzing large-scale bioinformatics datasets often requires multiple tools and complex workflows, leading to inefficiencies and potential errors. Here, we present BioDataHub, a Visual Studio Code extension designed to streamline dataset discovery, management, visu- alization, and analysis for bioinformatics researchers. BioDataHub integrates local and online dataset search, CSV preview, metadata generation, and interactive data visualization within a single IDE environment. To evaluate its utility, we applied BioDataHub to publicly available RNA-seq and microarray datasets, comparing workflow efficiency and data exploration outcomes against conventional tools. Results demonstrate that BioDataHub significantly reduces the time required for dataset preprocessing and provides intuitive visualizations that facilitate rapid in- sight generation. By combining accessibility, automation, and analytical capability, BioDataHub enhances bioinformatics data analysis workflows and offers a foundation for integrating further machine learning pipelines and advanced visualizations. Bioinformatics Artificial Intelligence and Machine Learning Bioinformatics Dataset Analysis CSV Data Management Data Visualization VS Code Extension Metadata Generation RNA-seq Microarray Workflow Optimization Computational Biology Figures Figure 1 Figure 2 1 Introduction The exponential growth of biological data, driven by high-throughput sequencing technologies such as RNA-seq [1] and microarrays [2], has created significant challenges for data management, explo- ration, and analysis in bioinformatics. Researchers often rely on multiple software tools to search, preprocess, and visualize datasets, which can lead to fragmented workflows, inefficiencies, and in- creased potential for errors. Despite the availability of several standalone bioinformatics tools, few solutions integrate dataset discovery, metadata generation, and interactive visualization within a single development environment [3, 4]. To address these challenges, we developed BioDataHub , a Visual Studio Code extension de- signed to streamline bioinformatics workflows. BioDataHub provides integrated features including local and online dataset search, CSV preview, metadata generation, and data visualization—all accessible within a single IDE. By consolidating these functionalities, BioDataHub reduces the time and complexity associated with dataset management and preliminary analysis, enabling researchers to focus on biological interpretation rather than technical overhead [5, 6]. In this study, we demonstrate the utility of BioDataHub by applying it to publicly available RNA-seq and microarray datasets. We evaluate its effectiveness in terms of workflow efficiency, data exploration, and visualization quality compared to conventional methods. Our results highlight how BioDataHub can facilitate rapid, reproducible, and user-friendly bioinformatics analyses, offering a foundation for future integration with machine learning pipelines and advanced computational tools. 2 Literature Review Effective management and analysis of bioinformatics datasets often require the use of multiple tools and platforms, each with distinct capabilities. One widely used solution is the Galaxy platform, which enables reproducible and collaborative biomedical analyses through a web-based interface [3]. While Galaxy is powerful, its web-based nature limits integration with local development environments, making it less convenient for researchers who prefer working within an IDE. Bioconductor provides a comprehensive suite of R packages for statistical analysis and visu- alization of genomic data [7]. However, it requires proficiency in R programming and does not natively support interactive dataset discovery or management within an IDE, which can present a barrier for non-programmers or researchers seeking streamlined workflows. Traditional spreadsheet-based or standalone CSV viewers offer a simple method to inspect small datasets, but they are insufficient for large-scale bioinformatics datasets and lack automated features such as metadata generation and integrated visualization. For data visualization, libraries such as Matplotlib [5] and Seaborn [6] allow for programmatic creation of static and statistical plots, but they require separate scripts and do not integrate directly with dataset management tools. Similarly, interactive visualization frameworks like Plotly and Dash provide rich visualization capabilities but involve complex setup and coding effort, limiting accessibility for researchers with limited programming experience. BioDataHub addresses these limitations by integrating dataset discovery, CSV preview, meta- data generation, and visualization within a single Visual Studio Code extension. This consolidation reduces workflow fragmentation, simplifies dataset exploration, and allows researchers to focus on data interpretation rather than tool management. By providing an IDE-based environment, Bio- DataHub bridges the gap between powerful analysis tools and user-friendly accessibility, enhancing efficiency and reproducibility in bioinformatics research. 3 Methodology 3.1 Software and Environment • Extension Name: BioDataHub • Platform: Visual Studio Code (VS Code) • Version: 1.4.2 • Programming Languages: TypeScript, JavaScript, HTML, CSS • Dependencies: VS Code API, Webview API • License: MIT 3.2 Installation and Setup 1. Open VS Code. 2. Press Ctrl+P to open Quick Open. 3. Paste the following command and press Enter: ext install Mubashir-Ali.bio-data-hub 4. Wait for the installation to complete. 5. Reload VS Code to activate the extension. 3.3 Usage Workflow 1. Open a folder containing CSV files in VS Code. 2. Click on the BioDataHub icon in the Activity Bar to open the extension. 3. Browse and select a CSV file to load. 4. The extension will parse the CSV file and display its contents in a tabular format. 5. Use the provided buttons to generate visualizations such as scatter plots and histograms. 6. View dataset metadata, including source, size, and tags. 7. Export visualizations and metadata as images or JSON files. 3.4 Dataset Selection • Source: Publicly available gene expression datasets from Kaggle and GEO. • Format: CSV and TSV files. • Size: Ranging from 50 MB to 2.3 GB. • Content: Gene expression data for genes such as BRCA1, BRCA2, and TP53. 3.5 Evaluation Metrics • Time Efficiency: Measured the time taken to load, explore, and visualize datasets. • Usability: Assessed user experience through feedback from bioinformatics students and re- searchers. • Functionality: Evaluated the range of features provided by the extension, including dataset preview, visualization, and metadata generation. 4 Results and Discussion 4.1 Dataset Exploration and Visualization BioDataHub was tested on publicly available RNA-seq and microarray datasets, including BRCA1, BRCA2, and TP53 gene expression datasets from Kaggle and GEO. The extension enabled users to load datasets directly within VS Code and generate interactive visualizations: CSV Preview : Datasets were displayed in tabular form with sortable rows and columns. Scatter Plots and Histograms : Gene expression patterns were visualized interactively (Fig. 1 ). 4.2 Metadata Generation and Dataset Cataloging BioDataHub automatically generated metadata for each dataset, including source, size, publication date, tags, and download options. 4.3 Comparative Analysis with Existing Tools Feature-wise comparison of BioDataHub with Galaxy, Bioconductor, and NCBI web tools is summa- rized in Table 1 . BioDataHub demonstrates superior workflow integration, IDE-based exploration, metadata generation, and beginner-friendly usability. Table 1 Feature-wise comparison of BioDataHub with existing bioinformatics tools. Feature / Tool BioDataHub Galaxy Bioconductor NCBI Web Tools Local Dataset Search Online Dataset Search CSV Preview Metadata Generation Interactive Visualization IDE Integration (VS Code) Ease of Use for Beginners Workflow Consolidation Dataset Catalog / Card View ✓ ✓ ✓ ✓ ✓ ✓ High High ✓ ✗ ✓ ✗ ✓ ✓ ✗ Medium Medium ✗ ✗ ✗ ✓ ✗ ✓ ✗ Medium Low ✗ ✗ ✓ ✗ ✗ ✗ ✗ Medium Low ✗ 4.4 Efficiency and Usability Users reported a significant reduction in time required to load, explore, and visualize datasets. The integrated interface and automation of metadata generation improved user experience and minimized errors compared to conventional workflows involving multiple tools. 4.5 Limitations Currently supports only CSV and TSV formats. Future versions will include additional bioin- formatics file formats such as FASTQ and BAM. Integration with machine learning pipelines for automated analysis and prediction is planned. Web-based repository support can be expanded to include additional public and private databases. 5 Conclusion and Future Work In this study, we presented BioDataHub , a Visual Studio Code extension designed to streamline bioinformatics dataset discovery, management, and visualization. BioDataHub integrates dataset search, CSV preview, metadata generation, and interactive visualization within a single IDE, reduc- ing workflow fragmentation and enhancing usability for both beginners and experienced researchers. Our evaluation demonstrates that BioDataHub provides: Efficient exploration and visualization of large-scale gene expression datasets. Automated metadata generation and dataset cataloging. Improved workflow integration compared to existing tools such as Galaxy, Bioconductor, and NCBI web tools. Reduced time and cognitive load for bioinformatics analyses. 5.1 Future Work Future development of BioDataHub will focus on: Expanding support for additional bioinformatics file formats such as FASTQ, BAM, and VCF. Incorporating machine learning and AI pipelines for automated data analysis and predictive modeling. Integration with more public and private repositories for seamless dataset access. Enhancing interactive visualization capabilities and user interface customization. Overall, BioDataHub aims to provide a unified, user-friendly platform that bridges the gap between bioinformatics data management and analytical workflows, enabling researchers to focus on biological insights rather than tool complexities. References Zhong Wang, Mark Gerstein, and Michael Snyder. Rna-seq: a revolutionary tool for transcrip- tomics. Nature reviews genetics , 10(1):57–63, 2009. Mark Schena, Dari Shalon, Ronald W Davis, and Patrick O Brown. Quantitative monitoring of gene expression patterns with a complementary dna microarray. Science , 270(5235):467–470, 1995. Enis Afgan, Dannon Baker, Bérénice Batut, Marius Van Den Beek, Dave Bouvier, Martin Čech, John Chilton, Dave Clements, Nate Coraor, Björn A Grüning, et al. The galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2018 update. Nucleic acids research , 46(W1):W537–W544, 2018. Björn Grüning, John Chilton, Johannes Köster, Ryan Dale, Nicola Soranzo, Marius Van Den Beek, Jeremy Goecks, Rolf Backofen, Anton Nekrutenko, and James Taylor. Practical computational reproducibility in the life sciences. Cell systems , 6(6):631–635, 2018. John D Hunter. Matplotlib: A 2d graphics environment. Computing in science & engineering , 9(03):90–95, 2007. Michael L Waskom. Seaborn: statistical data visualization. Journal of open source software , 6(60):3021, 2021. Wolfgang Huber, Vincent J Carey, Robert Gentleman, Simon Anders, Marc Carlson, Benilton S Carvalho, Hector Corrada Bravo, Sean Davis, Laurent Gatto, Thomas Girke, et al. Orchestrating high-throughput genomic analysis with bioconductor. Nature methods , 12(2):115–121, 2015. Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7861003","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":529624287,"identity":"ee99863b-fb8b-446f-8fd1-75c38eebad14","order_by":0,"name":"Mubashir Ali","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABR0lEQVRIie2RMUvDQBTHXzi4LFe6BjL0EwjvCISWVPwqVwLt0mrHThoQ6hKcW/RDZHS8ctAuAdcUO9ilk4qlIM3mJaKgIdJRJL/l3r3jx3vHH6Ci4k/CAIzA0oWSKPYWaZgBkM/HrI+lijEXw3XYNnkoD1L0SWLcPtJuHRLxu3I0Gcy26V3zDKjkkWDKNqbPs90QvAZKMntgsHJ+KG5y6tu12GoFTPoomsoxbd2ZQI9Hkvoeg41bUPpoG2MLwZJz1FN846bvEgbKiCRzbV20i4qTppnSWF/tBVUXwTLOlZNI1t9KFNeqZQooQEG7BBKWKx09hWZKYbH4petlCoU5YCdsEx72HXuCPX+qqNO6xU3h+4uBWqbjc6zD/Svf51HGfDccecfXi8t18jRa8aAQzAf0+xV1NHk6KEuEIl9pHq5UVFRU/FfeAWbEcPN3oAAhAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0009-0006-0222-7585","institution":"Quaid-i-Azam University, Islamabad","correspondingAuthor":true,"prefix":"","firstName":"Mubashir","middleName":"","lastName":"Ali","suffix":""}],"badges":[],"createdAt":"2025-10-14 17:24:49","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-7861003/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7861003/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":93641147,"identity":"202abf11-b71a-42bb-9d1b-25e6d66882d9","added_by":"auto","created_at":"2025-10-16 02:45:55","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":189940,"visible":true,"origin":"","legend":"","description":"","filename":"BioDataHub.docx","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/d979c2fdfefb803d989bd034.docx"},{"id":93641142,"identity":"c1d5e221-06bc-4241-9be3-caceb7984319","added_by":"auto","created_at":"2025-10-16 02:45:55","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":342,"visible":true,"origin":"","legend":"","description":"","filename":"rs7861003.json","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/8405f08ab2b051c7a6edc4cd.json"},{"id":93641141,"identity":"f4560e97-e1ba-4515-afc0-cb8a5e1f8ec7","added_by":"auto","created_at":"2025-10-16 02:45:55","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":33195,"visible":true,"origin":"","legend":"","description":"","filename":"rs78610030enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/0bc08d8abec99b3263600986.xml"},{"id":93641144,"identity":"c17bf548-2cd7-4fa2-b3cb-a0ba127fab94","added_by":"auto","created_at":"2025-10-16 02:45:55","extension":"jpeg","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":113761,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/18557ba48c57a013f6ac38cb.jpeg"},{"id":93641730,"identity":"fead3d1c-6689-4b28-9e38-3b87e3068dfa","added_by":"auto","created_at":"2025-10-16 02:53:55","extension":"jpeg","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":96252,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/b1c77307c9fedcd15ebca340.jpeg"},{"id":93641143,"identity":"f849969e-9824-4181-aedb-d57b9e99e656","added_by":"auto","created_at":"2025-10-16 02:45:55","extension":"png","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":46920,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/82e891cad94314879baef459.png"},{"id":93641731,"identity":"fc91eaa8-15d8-41d5-88e0-d667a418b051","added_by":"auto","created_at":"2025-10-16 02:53:55","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":38990,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/9f19fa9e9f2b9e9795e7ab56.png"},{"id":93641728,"identity":"47bc6f70-d706-447c-9f0c-415e41540bf1","added_by":"auto","created_at":"2025-10-16 02:53:55","extension":"xml","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":32503,"visible":true,"origin":"","legend":"","description":"","filename":"rs78610030structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/27ac794535d97b6564a8bc6d.xml"},{"id":93641906,"identity":"7b4be399-51bb-4206-9def-f9980f19e80c","added_by":"auto","created_at":"2025-10-16 03:01:55","extension":"html","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":38613,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/96230bcc611e9f7acab4fde2.html"},{"id":93641140,"identity":"167bf9fa-8cc5-45de-9197-7979f1a44e67","added_by":"auto","created_at":"2025-10-16 02:45:54","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":150396,"visible":true,"origin":"","legend":"\u003cp\u003eVisualization of gene expression dataset within BioDataHub showing scatter plot of BRCA1, BRCA2, and TP53 expression levels under control and treatment conditions.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/d278afc8679e9bd33b3a0903.png"},{"id":93641139,"identity":"1df8cc7d-b0f3-4305-8459-4ef44837f91b","added_by":"auto","created_at":"2025-10-16 02:45:54","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":115057,"visible":true,"origin":"","legend":"\u003cp\u003eMetadata-rich dataset card automatically generated in BioDataHub for a gene expression dataset retrieved from Kaggle.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/4577e1416900fbafa8604a35.png"},{"id":93642445,"identity":"5ac3cf15-f12b-47d9-bcc7-598c9f91a378","added_by":"auto","created_at":"2025-10-16 03:09:55","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":769668,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7861003/v1/69d6fc99-b66d-4376-b2a5-e27ea83b669e.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003ch2\u003eBioDataHub: An Integrated VS Code Extension for Streamlined Bioinformatics Dataset Analysis and Visualization\u003c/h2\u003e","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eThe exponential growth of biological data, driven by high-throughput sequencing technologies such as RNA-seq [1] and microarrays [2], has created significant challenges for data management, explo- ration, and analysis in bioinformatics. Researchers often rely on multiple software tools to search, preprocess, and visualize datasets, which can lead to fragmented workflows, inefficiencies, and in- creased potential for errors. Despite the availability of several standalone bioinformatics tools, few solutions integrate dataset discovery, metadata generation, and interactive visualization within a single development environment [3, 4].\u003c/p\u003e\n\u003cp\u003eTo address these challenges, we developed \u003cstrong\u003eBioDataHub\u003c/strong\u003e, a Visual Studio Code extension de- signed to streamline bioinformatics workflows. BioDataHub provides integrated features including\u003c/p\u003e\n\u003cp\u003elocal and online dataset search, CSV preview, metadata generation, and data visualization\u0026mdash;all accessible within a single IDE. By consolidating these functionalities, BioDataHub reduces the time and complexity associated with dataset management and preliminary analysis, enabling researchers to focus on biological interpretation rather than technical overhead [5, 6].\u003c/p\u003e\n\u003cp\u003eIn this study, we demonstrate the utility of BioDataHub by applying it to publicly available RNA-seq and microarray datasets. We evaluate its effectiveness in terms of workflow efficiency, data exploration, and visualization quality compared to conventional methods. Our results highlight how BioDataHub can facilitate rapid, reproducible, and user-friendly bioinformatics analyses, offering a foundation for future integration with machine learning pipelines and advanced computational tools.\u003c/p\u003e"},{"header":"2 Literature Review","content":"\u003cp\u003eEffective management and analysis of bioinformatics datasets often require the use of multiple tools and platforms, each with distinct capabilities. One widely used solution is the \u003cstrong\u003eGalaxy \u003c/strong\u003eplatform, which enables reproducible and collaborative biomedical analyses through a web-based interface [3]. While Galaxy is powerful, its web-based nature limits integration with local development environments, making it less convenient for researchers who prefer working within an IDE.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eBioconductor \u003c/strong\u003eprovides a comprehensive suite of R packages for statistical analysis and visu- alization of genomic data [7]. However, it requires proficiency in R programming and does not natively support interactive dataset discovery or management within an IDE, which can present a barrier for non-programmers or researchers seeking streamlined workflows.\u003c/p\u003e\n\u003cp\u003eTraditional spreadsheet-based or standalone CSV viewers offer a simple method to inspect small datasets, but they are insufficient for large-scale bioinformatics datasets and lack automated features such as metadata generation and integrated visualization.\u003c/p\u003e\n\u003cp\u003eFor data visualization, libraries such as \u003cstrong\u003eMatplotlib \u003c/strong\u003e[5] and \u003cstrong\u003eSeaborn \u003c/strong\u003e[6] allow for programmatic creation of static and statistical plots, but they require separate scripts and do not integrate directly with dataset management tools. Similarly, interactive visualization frameworks like \u003cstrong\u003ePlotly \u003c/strong\u003eand \u003cstrong\u003eDash \u003c/strong\u003eprovide rich visualization capabilities but involve complex setup and coding effort, limiting accessibility for researchers with limited programming experience.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eBioDataHub \u003c/strong\u003eaddresses these limitations by integrating dataset discovery, CSV preview, meta- data generation, and visualization within a single Visual Studio Code extension. This consolidation reduces workflow fragmentation, simplifies dataset exploration, and allows researchers to focus on data interpretation rather than tool management. By providing an IDE-based environment, Bio- DataHub bridges the gap between powerful analysis tools and user-friendly accessibility, enhancing efficiency and reproducibility in bioinformatics research.\u003c/p\u003e"},{"header":"3 Methodology","content":"\u003ch2\u003e3.1 Software\u0026nbsp;and\u0026nbsp;Environment\u003c/h2\u003e\n\u003cp\u003e• \u003cstrong\u003eExtension\u003c/strong\u003e\u003cstrong\u003eName:\u003c/strong\u003eBioDataHub\u003c/p\u003e\n\u003cp\u003e• \u003cstrong\u003ePlatform:\u0026nbsp;\u003c/strong\u003eVisual Studio Code (VS Code)\u003c/p\u003e\n\u003cp\u003e• \u003cstrong\u003eVersion:\u003c/strong\u003e1.4.2\u003c/p\u003e\n\u003cp\u003e• \u003cstrong\u003eProgramming\u0026nbsp;Languages:\u0026nbsp;\u003c/strong\u003eTypeScript,\u0026nbsp;JavaScript,\u0026nbsp;HTML,\u0026nbsp;CSS\u003c/p\u003e\n\u003cp\u003e• \u003cstrong\u003eDependencies:\u0026nbsp;\u003c/strong\u003eVS\u0026nbsp;Code\u0026nbsp;API,\u0026nbsp;Webview\u0026nbsp;API\u003c/p\u003e\n\u003cp\u003e• \u003cstrong\u003eLicense:\u003c/strong\u003eMIT\u003c/p\u003e\n\u003ch2\u003e3.2 Installation\u0026nbsp;and\u0026nbsp;Setup\u003c/h2\u003e\n\u003cp\u003e1.\u0026nbsp;\u0026nbsp;Open\u0026nbsp;VS Code.\u003c/p\u003e\n\u003cp\u003e2. Press Ctrl+P to open Quick Open.\u003c/p\u003e\n\u003cp\u003e3.\u0026nbsp;\u0026nbsp;Paste\u0026nbsp;the\u0026nbsp;following\u0026nbsp;command\u0026nbsp;and\u0026nbsp;press Enter:\u003c/p\u003e\n\u003cp\u003eext\u0026nbsp;install\u0026nbsp;Mubashir-Ali.bio-data-hub\u003c/p\u003e\n\u003cp\u003e4.\u0026nbsp;\u0026nbsp;Wait\u0026nbsp;for\u0026nbsp;the\u0026nbsp;installation\u0026nbsp;to complete.\u003c/p\u003e\n\u003cp\u003e5.\u0026nbsp;\u0026nbsp;Reload\u0026nbsp;VS\u0026nbsp;Code\u0026nbsp;to\u0026nbsp;activate\u0026nbsp;the extension.\u003c/p\u003e\n\u003ch2\u003e3.3 Usage\u0026nbsp;Workflow\u003c/h2\u003e\n\u003cp\u003e1.\u0026nbsp;\u0026nbsp;Open\u0026nbsp;a\u0026nbsp;folder\u0026nbsp;containing\u0026nbsp;CSV\u0026nbsp;files\u0026nbsp;in\u0026nbsp;VS Code.\u003c/p\u003e\n\u003cp\u003e2.\u0026nbsp;\u0026nbsp;Click\u0026nbsp;on\u0026nbsp;the\u0026nbsp;BioDataHub\u0026nbsp;icon\u0026nbsp;in\u0026nbsp;the\u0026nbsp;Activity\u0026nbsp;Bar\u0026nbsp;to\u0026nbsp;open\u0026nbsp;the extension.\u003c/p\u003e\n\u003cp\u003e3.\u0026nbsp;\u0026nbsp;Browse\u0026nbsp;and\u0026nbsp;select\u0026nbsp;a\u0026nbsp;CSV\u0026nbsp;file\u0026nbsp;to load.\u003c/p\u003e\n\u003cp\u003e4.\u0026nbsp;\u0026nbsp;The\u0026nbsp;extension\u0026nbsp;will\u0026nbsp;parse\u0026nbsp;the\u0026nbsp;CSV\u0026nbsp;file\u0026nbsp;and\u0026nbsp;display\u0026nbsp;its\u0026nbsp;contents\u0026nbsp;in\u0026nbsp;a\u0026nbsp;tabular format.\u003c/p\u003e\n\u003cp\u003e5.\u0026nbsp;\u0026nbsp;Use\u0026nbsp;the\u0026nbsp;provided\u0026nbsp;buttons\u0026nbsp;to\u0026nbsp;generate\u0026nbsp;visualizations\u0026nbsp;such\u0026nbsp;as\u0026nbsp;scatter\u0026nbsp;plots\u0026nbsp;and histograms.\u003c/p\u003e\n\u003cp\u003e6.\u0026nbsp;\u0026nbsp;View\u0026nbsp;dataset\u0026nbsp;metadata,\u0026nbsp;including\u0026nbsp;source,\u0026nbsp;size,\u0026nbsp;and tags.\u003c/p\u003e\n\u003cp\u003e7.\u0026nbsp;\u0026nbsp;Export\u0026nbsp;visualizations\u0026nbsp;and\u0026nbsp;metadata\u0026nbsp;as\u0026nbsp;images\u0026nbsp;or\u0026nbsp;JSON files.\u003c/p\u003e\n\u003ch2\u003e3.4 Dataset\u0026nbsp;Selection\u003c/h2\u003e\n\u003cp\u003e• \u003cstrong\u003eSource:\u0026nbsp;\u003c/strong\u003ePublicly available gene expression datasets from Kaggle and GEO.\u003c/p\u003e\n\u003cp\u003e• \u003cstrong\u003eFormat:\u0026nbsp;\u003c/strong\u003eCSV\u0026nbsp;and\u0026nbsp;TSV\u0026nbsp;files.\u003c/p\u003e\n\u003cp\u003e• \u003cstrong\u003eSize:\u0026nbsp;\u003c/strong\u003eRanging from 50 MB to 2.3 GB.\u003c/p\u003e\n\u003cp\u003e• \u003cstrong\u003eContent:\u0026nbsp;\u003c/strong\u003eGene expression data for genes such as BRCA1, BRCA2, and TP53.\u003c/p\u003e\n\u003ch2\u003e3.5 Evaluation\u0026nbsp;Metrics\u003c/h2\u003e\n\u003cp\u003e• \u003cstrong\u003eTime\u0026nbsp;Efficiency:\u0026nbsp;\u003c/strong\u003eMeasured the time taken to load, explore, and visualize datasets.\u003c/p\u003e\n\u003cp\u003e• \u003cstrong\u003eUsability:\u0026nbsp;\u003c/strong\u003eAssessed user experience through feedback from bioinformatics students and re- searchers.\u003c/p\u003e\n\u003cp\u003e• \u003cstrong\u003eFunctionality:\u0026nbsp;\u003c/strong\u003eEvaluated the range of features provided by the extension, including dataset preview, visualization, and metadata generation.\u003c/p\u003e"},{"header":"4 Results and Discussion","content":"\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003e4.1 Dataset Exploration and Visualization\u003c/h2\u003e\u003cp\u003eBioDataHub was tested on publicly available RNA-seq and microarray datasets, including BRCA1, BRCA2, and TP53 gene expression datasets from Kaggle and GEO.\u003c/p\u003e\u003cp\u003eThe extension enabled users to load datasets directly within VS Code and generate interactive visualizations:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eCSV Preview\u003c/b\u003e: Datasets were displayed in tabular form with sortable rows and columns.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eScatter Plots and Histograms\u003c/b\u003e: Gene expression patterns were visualized interactively (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003e4.2 Metadata Generation and Dataset Cataloging\u003c/h2\u003e\u003cp\u003eBioDataHub automatically generated metadata for each dataset, including source, size, publication date, tags, and download options.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\u003ch2\u003e4.3 Comparative Analysis with Existing Tools\u003c/h2\u003e\u003cp\u003eFeature-wise comparison of BioDataHub with Galaxy, Bioconductor, and NCBI web tools is summa- rized in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. BioDataHub demonstrates superior workflow integration, IDE-based exploration, metadata generation, and beginner-friendly usability.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eFeature-wise comparison of BioDataHub with existing bioinformatics tools.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFeature / Tool\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBioDataHub\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGalaxy\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eBioconductor\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eNCBI Web Tools\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLocal Dataset Search\u003c/p\u003e\u003cp\u003eOnline Dataset Search CSV Preview Metadata Generation\u003c/p\u003e\u003cp\u003eInteractive Visualization IDE Integration (VS Code) Ease of Use for Beginners Workflow Consolidation\u003c/p\u003e\u003cp\u003eDataset Catalog / Card View\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e✓\u003c/p\u003e\u003cp\u003e✓\u003c/p\u003e\u003cp\u003e✓\u003c/p\u003e\u003cp\u003e✓\u003c/p\u003e\u003cp\u003e✓\u003c/p\u003e\u003cp\u003e✓ High High\u003c/p\u003e\u003cp\u003e✓\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e✗\u003c/p\u003e\u003cp\u003e✓\u003c/p\u003e\u003cp\u003e✗\u003c/p\u003e\u003cp\u003e✓\u003c/p\u003e\u003cp\u003e✓\u003c/p\u003e\u003cp\u003e✗ Medium Medium\u003c/p\u003e\u003cp\u003e✗\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e✗\u003c/p\u003e\u003cp\u003e✗\u003c/p\u003e\u003cp\u003e✓\u003c/p\u003e\u003cp\u003e✗\u003c/p\u003e\u003cp\u003e✓\u003c/p\u003e\u003cp\u003e✗ Medium Low\u003c/p\u003e\u003cp\u003e✗\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e✗\u003c/p\u003e\u003cp\u003e✓\u003c/p\u003e\u003cp\u003e✗\u003c/p\u003e\u003cp\u003e✗\u003c/p\u003e\u003cp\u003e✗\u003c/p\u003e\u003cp\u003e✗ Medium Low\u003c/p\u003e\u003cp\u003e✗\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003e4.4 Efficiency and Usability\u003c/h2\u003e\u003cp\u003eUsers reported a significant reduction in time required to load, explore, and visualize datasets. The integrated interface and automation of metadata generation improved user experience and minimized errors compared to conventional workflows involving multiple tools.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003ch2\u003e4.5 Limitations\u003c/h2\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eCurrently supports only CSV and TSV formats. Future versions will include additional bioin- formatics file formats such as FASTQ and BAM.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eIntegration with machine learning pipelines for automated analysis and prediction is planned.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eWeb-based repository support can be expanded to include additional public and private databases.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"5 Conclusion and Future Work","content":"\u003cp\u003eIn this study, we presented \u003cb\u003eBioDataHub\u003c/b\u003e, a Visual Studio Code extension designed to streamline bioinformatics dataset discovery, management, and visualization. BioDataHub integrates dataset search, CSV preview, metadata generation, and interactive visualization within a single IDE, reduc- ing workflow fragmentation and enhancing usability for both beginners and experienced researchers.\u003c/p\u003e\u003cp\u003eOur evaluation demonstrates that BioDataHub provides:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eEfficient exploration and visualization of large-scale gene expression datasets.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eAutomated metadata generation and dataset cataloging.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eImproved workflow integration compared to existing tools such as Galaxy, Bioconductor, and NCBI web tools.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eReduced time and cognitive load for bioinformatics analyses.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\u003ch2\u003e5.1 Future Work\u003c/h2\u003e\u003cp\u003eFuture development of BioDataHub will focus on:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eExpanding support for additional bioinformatics file formats such as FASTQ, BAM, and VCF.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eIncorporating machine learning and AI pipelines for automated data analysis and predictive modeling.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eIntegration with more public and private repositories for seamless dataset access.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eEnhancing interactive visualization capabilities and user interface customization.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eOverall, BioDataHub aims to provide a unified, user-friendly platform that bridges the gap between bioinformatics data management and analytical workflows, enabling researchers to focus on biological insights rather than tool complexities.\u003c/p\u003e\u003c/div\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eZhong Wang, Mark Gerstein, and Michael Snyder. Rna-seq: a revolutionary tool for transcrip- tomics. \u003cem\u003eNature reviews genetics\u003c/em\u003e, 10(1):57\u0026ndash;63, 2009.\u003c/li\u003e\n\u003cli\u003eMark Schena, Dari Shalon, Ronald W Davis, and Patrick O Brown. Quantitative monitoring of gene expression patterns with a complementary dna microarray. \u003cem\u003eScience\u003c/em\u003e, 270(5235):467\u0026ndash;470, 1995.\u003c/li\u003e\n\u003cli\u003eEnis Afgan, Dannon Baker, B\u0026eacute;r\u0026eacute;nice Batut, Marius Van Den Beek, Dave Bouvier, Martin Čech, John Chilton, Dave Clements, Nate Coraor, Bj\u0026ouml;rn A Gr\u0026uuml;ning, et al. The galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2018 update. \u003cem\u003eNucleic acids research\u003c/em\u003e, 46(W1):W537\u0026ndash;W544, 2018.\u003c/li\u003e\n\u003cli\u003eBj\u0026ouml;rn Gr\u0026uuml;ning, John Chilton, Johannes K\u0026ouml;ster, Ryan Dale, Nicola Soranzo, Marius Van Den Beek, Jeremy Goecks, Rolf Backofen, Anton Nekrutenko, and James Taylor. Practical computational reproducibility in the life sciences. \u003cem\u003eCell systems\u003c/em\u003e, 6(6):631\u0026ndash;635, 2018.\u003c/li\u003e\n\u003cli\u003eJohn D Hunter. Matplotlib: A 2d graphics environment. \u003cem\u003eComputing in science \u0026amp; engineering\u003c/em\u003e, 9(03):90\u0026ndash;95, 2007.\u003c/li\u003e\n\u003cli\u003eMichael L Waskom. Seaborn: statistical data visualization. \u003cem\u003eJournal of open source software\u003c/em\u003e, 6(60):3021, 2021.\u003c/li\u003e\n\u003cli\u003eWolfgang Huber, Vincent J Carey, Robert Gentleman, Simon Anders, Marc Carlson, Benilton S Carvalho, Hector Corrada Bravo, Sean Davis, Laurent Gatto, Thomas Girke, et al. Orchestrating high-throughput genomic analysis with bioconductor. \u003cem\u003eNature methods\u003c/em\u003e, 12(2):115\u0026ndash;121, 2015.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Bioinformatics, Dataset Analysis, CSV Data Management, Data Visualization, VS Code Extension, Metadata Generation, RNA-seq, Microarray, Workflow Optimization, Computational Biology","lastPublishedDoi":"10.21203/rs.3.rs-7861003/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7861003/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eManaging and analyzing large-scale bioinformatics datasets often requires multiple tools and complex workflows, leading to inefficiencies and potential errors. Here, we present BioDataHub, a Visual Studio Code extension designed to streamline dataset discovery, management, visu- alization, and analysis for bioinformatics researchers. BioDataHub integrates local and online dataset search, CSV preview, metadata generation, and interactive data visualization within a single IDE environment. To evaluate its utility, we applied BioDataHub to publicly available RNA-seq and microarray datasets, comparing workflow efficiency and data exploration outcomes against conventional tools. Results demonstrate that BioDataHub significantly reduces the time required for dataset preprocessing and provides intuitive visualizations that facilitate rapid in- sight generation. By combining accessibility, automation, and analytical capability, BioDataHub enhances bioinformatics data analysis workflows and offers a foundation for integrating further machine learning pipelines and advanced visualizations.\u003c/p\u003e","manuscriptTitle":"BioDataHub: An Integrated VS Code Extension for Streamlined Bioinformatics Dataset Analysis and Visualization","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-16 02:45:50","doi":"10.21203/rs.3.rs-7861003/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"5f02c3e3-2e4e-49eb-b386-79ab213d1dc3","owner":[],"postedDate":"October 16th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":56295864,"name":"Bioinformatics"},{"id":56295865,"name":"Artificial Intelligence and Machine Learning"}],"tags":[],"updatedAt":"2026-03-31T19:05:43+00:00","versionOfRecord":[],"versionCreatedAt":"2025-10-16 02:45:50","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7861003","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7861003","identity":"rs-7861003","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.