MetaMiner: Streamlined GUI Tool for Retrieving, Normalizing and Exploring Metadata

preprint OA: closed
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-14

MetaMiner is a GUI tool that retrieves, normalizes, and visualizes prokaryotic genome metadata, streamlining data exploration and enhancing re-usability.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-14 · read from full text

MetaMiner is a Python-based GUI tool that retrieves prokaryotic genome metadata from NCBI, normalizes key fields (including geographical location, isolation source, and sequencing technology), removes redundancies, and provides an interactive dashboard for exploration and filtering. To assess usability, the authors applied it to randomly selected Acinetobacter baumannii genome metadata downloaded in JSON-lines format on February 3, 2025, and compared normalization and transformation performance against NCBI tools and manually curated datasets, reporting superior transformation results and high classification accuracy with only minor discrepancies attributed to recent geopolitical changes or ambiguous submissions. A stated limitation is that discrepancies can still occur when metadata entries are ambiguous or affected by changing geopolitical information. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Summary Over the past decade, publicly available genomic data has expanded exponentially. However, the usability of these genomes heavily depends on their associated metadata. Metadata provides context about the data that from where, when and how the data was obtained. When not consistent, it limits the re-usage of the available public data and prevents data’s integration into broader genomic analyses. Here, we introduce MetaMiner, a GUI tool designed to retrieve, normalize, and interactively visualize metadata associated with prokaryotic genomes. MetaMiner streamlines the normalization of key metadata elements such as geographical locations, isolation sources, and sequencing technologies. To assess its usability, we used metadata of Acinetobacter baumannii genomes and compared it with NCBI’s data processing tools and manually curated datasets. MetaMiner demonstrated superior performance in data transformation and showed high classification accuracy, with only minor discrepancies arising from recent geopolitical changes or ambiguous submissions. This cleaned and standardized metadata is integrated into an interactive dashboard, which allows users to explore data holistically through visualizations and custom filtering options. The dashboard has interactive plots related to submission date trend, geographical origins, sequencing methods, genome quality scores (e.g., ANI > 95, BUSCO > 95), coverage, etc. Overall, MetaMiner minimizes manual curation efforts of genomic metadata and enhances the accessibility, and re-usability of publicly available genomes. Availability and Implementation The executables and source code are both freely available to use at https://github.com/prekijpatel/MetaMiner for non-commercial use. Detailed documentation is available at https://github.com/prekijpatel/MetaMiner/wiki .
Full text 31,594 characters · extracted from preprint-html · click to expand
MetaMiner: Streamlined GUI Tool for Retrieving, Normalizing and Exploring Metadata | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results MetaMiner: Streamlined GUI Tool for Retrieving, Normalizing and Exploring Metadata View ORCID Profile Jaykumar Kiritkumar Patel , View ORCID Profile Ravikrishnan Elangovan doi: https://doi.org/10.1101/2025.08.20.666107 Jaykumar Kiritkumar Patel a Department of Biochemical Engineering and Biotechnology, Indian Institute of Technology Delhi , New Delhi, 110016, India Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Jaykumar Kiritkumar Patel Ravikrishnan Elangovan a Department of Biochemical Engineering and Biotechnology, Indian Institute of Technology Delhi , New Delhi, 110016, India Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Ravikrishnan Elangovan For correspondence: elangovan{at}dbeb.iitd.ac.in Abstract Full Text Info/History Metrics Supplementary material Data/Code Preview PDF Abstract Summary Over the past decade, publicly available genomic data has expanded exponentially. However, the usability of these genomes heavily depends on their associated metadata. Metadata provides context about the data that from where, when and how the data was obtained. When not consistent, it limits the re-usage of the available public data and prevents data’s integration into broader genomic analyses. Here, we introduce MetaMiner, a GUI tool designed to retrieve, normalize, and interactively visualize metadata associated with prokaryotic genomes. MetaMiner streamlines the normalization of key metadata elements such as geographical locations, isolation sources, and sequencing technologies. To assess its usability, we used metadata of Acinetobacter baumannii genomes and compared it with NCBI’s data processing tools and manually curated datasets. MetaMiner demonstrated superior performance in data transformation and showed high classification accuracy, with only minor discrepancies arising from recent geopolitical changes or ambiguous submissions. This cleaned and standardized metadata is integrated into an interactive dashboard, which allows users to explore data holistically through visualizations and custom filtering options. The dashboard has interactive plots related to submission date trend, geographical origins, sequencing methods, genome quality scores (e.g., ANI > 95, BUSCO > 95), coverage, etc. Overall, MetaMiner minimizes manual curation efforts of genomic metadata and enhances the accessibility, and re-usability of publicly available genomes. Availability and Implementation The executables and source code are both freely available to use at https://github.com/prekijpatel/MetaMiner for non-commercial use. Detailed documentation is available at https://github.com/prekijpatel/MetaMiner/wiki . Background High-throughput sequencing technologies has revolutionised various fields of biology, leading to exponential increase in omics data generation ( Reuter et al ., 2015 ). Resultingly, genome submissions to public databases, such as GenBank and the European Nucleotide Archive (ENA), have increased significantly in recent years, highlighting the rapid rise of sequencing projects worldwide ( Sayers et al ., 2025 ; O’Cathail et al ., 2025 ). The value of this vast genomic datasets depends not only on the sequences themselves but also on the quality and completeness of the accompanying metadata. Metadata—information about the conditions and context in which the data was collected—is extremely helpful in designing new hypotheses and experiments by re-using old data ( Caliskan et al ., 2023 ). This is especially important in fields like comparative genomics. Conducting comparative genomic studies and inferring from them often depends on well-organized and standardised metadata ( Bornstein et al ., 2023 ; Moustafa et al ., 2020 ). Due to the adoption of FAIR (Findable, Accessible, Interoperable, and Reusable) principles ( Wilkinson et al ., 2016 ; Musen et al ., 2022 ) and recent developments, metadata retrieval from public repositories has become easier. The retrieval is usually through command-line tools or FTP sites ( O’Leary et al ., 2024 ; Gálvez-Merchán et al ., 2023 ). Despite these advancements in availability of metadata, the quality of metadata still remains a limiting factor in its applications. To address this, the Genomic Standards Consortium (GSC) has established standards such as MIGS (Minimum Information about a Genome Sequence), MIMS (Minimum Information about a Metagenome Sequence), MISAG (Minimum Information about a Single Amplified Genome), and MIMAG (Minimum Information about a Metagenome-Assembled Genome) ( Field et al ., 2008 ; Bowers et al ., 2017 ). While these frameworks offer valuable guidance for metadata standardisation, many fields still allow freeform text entries without validation. This often leads to unintentional errors, such as misspellings, missing entries, inconsistency, and over-explanation, which diminishes the usability and reliability of the data. These errors are even more common when multiple researchers and/or studies are involved. In addition to this, genomic data’s high dimensionality makes it even more complex as each dataset, varying from human-associated pathogens to environmental symbiont, may have its own unique metadata needs and accompanying challenges–which makes the standardisation/normalisation of metadata even more perplexing. Improving metadata standardisation and quality has the potential to advance genomics research to a great extent. High-quality metadata allows researchers to combine data from different labs and studies, facilitating collaborative research and enabling experiments that are beyond the scope of individual labs ( Ryan et al ., 2021 ; Rajesh et al ., 2021 ). Standardisation of metadata also makes it machine-actionable, which enables automated processing and further simplification of metadata without manual labour ( Batista et al ., 2022 ). Overall, standardised metadata immensely aids in efficient data discovery, reuse, and interpretation. There are resources which are widely recognised for providing uniformly analysed genomes and/or curated metadata, such as EnteroBase ( Dyer et al ., 2024 ), the Genomes OnLine Database ( Mukherjee et al ., 2025 ), the National Microbiome Data Collaborative (NMDC) (National Microbiome Data Collaborative), and BakRep ( Fenske et al ., 2024 ). However, as these repositories often focus on distinct taxa or specific research goals, their metadata may have limitations in general usability for broader comparative studies. To address these gaps, we introduce MetaMiner, a GUI-based tool designed to streamline metadata retrieval, cleaning, and standardisation for prokaryotic genomes from NCBI datasets. It systematically normalizes key metadata elements such as geographical location, isolation source, sequencing technology, and other assembly parameters. Additionally, it identifies and removes redundancies within datasets, and outputs a cleaner, programmatically curated metadata set, which, with minimal cross-checking, results in reliable and standardised metadata. Beyond metadata cleaning, MetaMiner features an interactive visualisation dashboard powered by Dash, allowing users to dynamically explore, filter, and analyse metadata in an intuitive interface. Methodology Metadata Retrieval and Processing The Acinetobacter baumannii metadata was obtained on February 3, 2025. The selection of Acinetobacter baumannii was random and serves as a representative example. The command-line downloaded metadata was downloaded in JSON-lines format using NCBI’s datasets . exe (v16 . 40 . 1) . The dataformat . exe (v16 . 40 . 1) from NCBI was used to format the JSON-line from a webpage, JSON-line from command-line tool, and JSON from MetaMiner. This was then compared to MetaMiner formatted versions. The log for metadata retrieval and formatting is available in Supplementary file 1 . The time estimation for loading formatted TSV files was calculated using the timeit module for five iterations. Comparing Metadata Normalization MetaMiner normalised data was compared with the manually/peer-curated normalisations. The raw metadata that was normalised was same as the one retrieved using MetaMiner on 3 rd February 2025. The code used for comparison and plotting the results is available in Supplementary file 2 . Results Design MetaMiner is a Python-based tool developed for the retrieval, standardisation, and visualisation of prokaryotic metadata. It makes use of various libraries and custom-made databases to process metadata efficiently. The tool is designed to handle both scenarios where pre-downloaded metadata is already available with the user and another where the user needs to download the needed metadata from the NCBI. The metadata retrieval from NCBI is enabled by the datasets command-line tool. MetaMiner transforms the raw JSON/JSONL to a structured, tabular format, making it easier to work with in large-scale analyses. Once the metadata is transformed, preprocessing steps are applied to remove redundant entries from the formatted metadata, specifically, for assemblies that include both RefSeq and GenBank annotations, only the metadata from RefSeq assembly is retained for subsequent processing while retaining the raw metadata in a separate file. To standardise geographical location data, MetaMiner utilises a combination of text-based normalisation and geospatial tools such as PyCountry and Geopy’s Nominatim . The tool extracts and maps geographical information to country- and state-levels based on raw metadata entries. It is important to note that MetaMiner does not perform sub-state level normalisation; therefore, even if district-level information is available, it is aggregated at the state level to ensure uniformity. Normalisation of isolation source metadata is done by utilising rapidfuzz library on certain keywords and localised reference database. The algorithm leverages three metadata fields, ` host `, ` host_disease `, and `isolation_source` , to categorize the isolation sources. Each raw metadata entry, defined as a combination of the ` host `, ` host_disease `, and ` isolation_source ` fields, is matched against an extensive internal dictionary. This dictionary comprises standardized terms, synonyms, common misspellings, unusual abbreviations, and other irregular forms frequently encountered in raw datasets. Matching is conducted using rapidfuzz , which applies fuzzy string-matching techniques guided by empirically determined thresholds to account for near matches and minor variations. Based on the input, the algorithm categorizes isolation source into a four-level classification. The first level broadly categorizes the entry by host type such as, Hospital, Animal, Environmental, Laboratory, among others. The remaining three levels of classification are host-specific and vary depending on the assigned host category. The detailed overview of levelled classification categories and algorithm is illustrated in Figure 1 and 2 , respectively. Download figure Open in new tab Figure 1. Classification scheme illustrating the hierarchical levels in the normalized isolation source data. The four levels are depicted using distinct colors. In case of Hospital-associated isolates, unlike others, the level 2 has 3 separate and independent categories which are classified differently in proceeding levels . Download figure Open in new tab Figure 2. Flowchart depicting the algorithm for normalizing isolation source metadata. The different level of information is retrieved at various steps of the algorithm. The examples of final retrieved results are shown in purple parallelograms. The text-background colour corresponds to the process step (box color) from which the information was obtained. Sequencing technology used to generate assembly data is classified by identifying key terms such as Illumina, Nanopore, DNBseq, MiSeq, etc. Based on them, MetaMiner categorises sequencing technologies into various short-read, long-read, or hybrid sequencing platforms. The representative examples of these normalizations are depicted in Table 1 . In addition to this, other assembly-related parameters, such as gene counts, ANI, BUSCO, coverage, etc., are also re-formatted. All of this is streamed into an in-built dashboard which allows user to interact with metadata visually, providing functionalities such as filtering specific subsets of data, exploring trends through dynamic plots, and saving customized views or filtered datasets for downstream analysis and reporting purposes. View this table: View inline View popup Download powerpoint Table 1. Examples of Normalization performed by MetaMiner. The table illustrates the transformation of heterogeneous metadata into standard formats across three metadata fields: Geographical locations, Isolation source and Sequencing Technology. Validation of Metadata retrieval and processing In order to test the efficiency of MetaMiner to retrieve and transform the data, raw metadata was retrieved for Acinetobacter baumannii assemblies using three different methods: (1) from the NCBI Datasets webpage, (2) via the NCBI Datasets command-line tool, and (3) using MetaMiner. Despite the varying retrieval methods and negligible differences in file size, the number of records in each dataset was identical ( Figure 3a ). This indicates that the data retrieved using MetaMiner is equivalent to that from NCBI, which corroborates with MetaMiner’s use of NCBI Datasets in the background. Download figure Open in new tab Figure 3. a) Number of assemblies for which metadata is available in json/jsonl file downloaded using various modalities. b) Size of the transformed metadata by NCBI dataformat and MetaMiner. The text near the bar represent the shape (row, column) of the transformed metadata table. c)Time taken to load the transformed metadata for further processing. d) Number of assemblies’ normalized geographical locations that matched with manually curated data. The manually curated data bar is shown in grey in back of each bar. e) Number of assemblies’ normalized sequencing technology data that matched with manually curated data. f) Number of assemblies’ categorized isolation source data that matched with manually curated data. The normalization was performed with and without the supporting database. Next, to compare how MetaMiner processes raw metadata in JSON/JSONL format, we transformed the retrieved data into tab-separated values using both NCBI’s Dataformat and MetaMiner. The data formatted by MetaMiner is considerably smaller than the one formatted by NCBI Dataformat ( Figure 3b ). This difference arises from the distinct structuring methods; Dataformat creates a new row for each BioSample attribute, while MetaMiner consolidates all attributes into a single row. This approach significantly enhances the loading and processing times of the transformed data for further use ( Figure 3c ).. Additionally, unlike Dataformat, MetaMiner can handle both JSON and JSON-lines formats. While NCBI Datasets and Dataformat offer many other features which MetaMiner does not, the advantages of MetaMiner in terms of transforming metadata are evident. Evaluation of Metadata Normalisation A total of 2,500 assemblies were randomly selected from the retrieved and processed A. baumannii metadata. For these assemblies, data related to Geographical Location, Isolation Source, and Sequencing Technology were normalized using both manual curation and MetaMiner. Geographical location normalization categorized the ‘geo_loc’ data from the raw data into Country and State information. MetaMiner demonstrated performance comparable to manual curation, correctly identifying 2,496 countries out of 2,500 assemblies ( Figure 3d ). The minor discrepancies stemmed from recent geopolitical changes (e.g., separation of certain countries) and ambiguous entries submitted by submitters, such as “Korea”, which could not be resolved by MetaMiner as North or South. Like geographical locations, the normalization of sequencing technology used for genome sequencing also showed high level of accuracy ( Figure 3e ). The only exception was its inability to identify a single new sequencing technology, GenoLab M. In terms of isolation source normalization, due to the high variability of isolation source entries, a curated database is used by default to support the categorization. This database is manually compiled from existing metadata of similar genomes and provides a reference set for normalization. However, to evaluate MetaMiner’s performance objectively, normalization was tested with and without the curated database. The without-database scenario simulates how the categorization would be handled when entirely new entries are encountered. Figure 3f presents the performance data, highlighting that MetaMiner successfully categorized most entries with the database across all levels. However, it encountered minor challenges in classifying ‘Source’ and ‘Sample’ categories in the absence of database support. Notably, even when MetaMiner is not able to categorise ‘Source’ and ‘Sample’ correctly, it still correctly identified the ‘host’ and ‘source category’ from which the sample was collected, demonstrating its robustness in metadata normalization at those levels if not more. Dash-based Dashboard Once the data is cleaned, normalised, and standardised, it is processed to generate an interactive dashboard ( Figure 4 ). This dashboard serves as the central interface for exploring the normalized dataset, offering both visual summaries and filtering capabilities. For example, normalized geographical locations are plotted into choropleth map ( Figure 5 ), allowing users to identify regional sampling densities and trends. Sequencing platforms associated with each assembly are visualized through an interactive scatter plot ( Figure 6 ), which along with coverage bar plots could be helpful in selecting highly accurate assemblies. To assess genome assembly quality, a scatter plot comparing L50 and N50 values is provided ( Figure 7 ), helping to distinguish between high- and low-contiguity assemblies. The diversity of isolation sources is captured in a hierarchical tree map ( Figure 8 ), where genomes are organized based on host association, isolation category, and sample type, offering an intuitive overview of the ecological and clinical contexts of the dataset. Download figure Open in new tab Figure 4. Top section of the Meta-Mined dashboard interface. The screenshot displays the upper portion of the dashboard, including features such as LED showing total assembly number, radio buttons for including/excluding atypical and suppressed assemblies, and geographical distribution of isolates. Users can apply these filters along with other additional (not seen in the screenshot) ones—such as, sequencing technology, gene counts, isolation source, and genomic quality metrics, etc . —to explore and refine the dataset interactively . Download figure Open in new tab Figure 5. Choropleth map showing geographical distribution of assembles in Meta-Mined dashboard. The static image represents the interactive choropleth map, which displays number of assemblies per country. In the bottom, it shows the number of total assemblies and number of assemblies for which geographical data is not available. In live dashboard, choropleth map is combined with a drop-down menu to select assemblies from specific country . Download figure Open in new tab Figure 6. Scatter map showing sequencing technologies utilized over the years for given assemblies in Meta-Mined dashboard. The visualization shows number of assemblies sequenced using specific technologies in given year. The circle size indicates the log count of the assemblies. In the dashboard, this is coupled with dropdown menu that allows users to select assemblies sequenced by one or multiple sequencing technologies . Download figure Open in new tab Figure 7. Scatter plot of scaffold N50 vs. L50 values for given assemblies in the Meta-Mined dashboard. The plot illustrates structural quality of the assemblies in relation with N50 and L50 values. Each point in the plot represents an assembly, with color indicating the sequencing technology used. This visualization is combined with sliders to filter assemblies based-on certain value of L50 and/or N50 . Download figure Open in new tab Figure 8. Treemap of normalized isolation sources in the Meta-Mined dashboard. The image presents a static treemap (on right) showing the category-wise distribution of assemblies within various isolation sources. Each block corresponds to a specific source, with size proportional to its frequency in the dataset. In live dashboard, user can refine the view by selecting or deselecting a certain block and explore the distribution better. The accompanying dropdown filters (on left) further enables users to select the assemblies from one or more isolation sources based-on their need . Along with these, the other options allow users to subset assemblies based on annotation provider, assembly level, gene count, BioSample or BioProject IDs, and more. For example, one can retrieve genomes submitted between 2013 and 2024, originating from a specific U.S. state, sequenced using a particular technology, and meeting thresholds such as ANI % identity above 95 and N50 value above 1 Mb, among others. The comprehensive visualizations provided by dashboard makes it possible to tailor data exploration and retrieval to specific research questions. The filtered data then can be used to fetch other genomic data for comparative analysis. The detailed description of the dashboard components is available at https://github.com/prekijpatel/MetaMiner/wiki/Components-of-Dashboard . Conclusion MetaMiner enhances the efficiency of genomic metadata processing by automating data retrieval, normalization, and analysis. Its advanced algorithms standardize diverse metadata fields, enabling comprehensive visualization and exploration. The interactive dashboard allows users to filter and analyse data based on custom criteria, supporting more focused and insightful research. By streamlining the process, MetaMiner improves the accessibility and usability of large-scale genomic datasets, making it an essential tool for data-driven genomic research. Declarations Availability of data and materials MetaMiner is available freely at https://github.com/prekijpatel/MetaMiner . Competing interests The authors declare that they have no competing interests. Funding This research was not supported by any funding agencies. Author’s contributions JP conceptualised the study, designed the MetaMiner tool, and performed data curation and analysis. JP also wrote the initial draft of the manuscript. RKE supervised the project, provided critical insights and guided the research framework. RKE also contributed significantly to manuscript revisions and final editing. Both read, reviewed and approved the final manuscript. Acknowledgements JP is supported by the DBT-JRF Fellowship program. The authors sincerely thank all colleagues, particularly Prabhu Balasubramanian sir, Ada Zwetlana, Lily Singh, Ratnabali Ghosh, Sinad Mohamed and Saunak Dasgupta for their valuable time, insightful discussions, and constructive inputs that greatly contributed to the development and refinement of this work. Footnotes https://github.com/prekijpatel/MetaMiner References ↵ Batista , D. et al. ( 2022 ) Machine actionable metadata models . Sci Data , 9 . ↵ Bornstein , K. et al. ( 2023 ) The NIH Comparative Genomics Resource: addressing the promises and challenges of comparative genomics on human health . BMC Genomics , 24 . ↵ Bowers , R.M. et al. ( 2017 ) Minimum information about a single amplified genome (MISAG) and a metagenome-assembled genome (MIMAG) of bacteria and archaea . Nat Biotechnol , 35 , 725 – 731 . OpenUrl CrossRef PubMed ↵ Caliskan , A. et al. ( 2023 ) Metadata integrity in bioinformatics: Bridging the gap between data and knowledge . Comput Struct Biotechnol J , 21 , 4895 – 4913 . OpenUrl CrossRef PubMed ↵ Dyer , N.P. et al. ( 2024 ) EnteroBase in 2025: exploring the genomic epidemiology of bacterial pathogens . Nucleic Acids Res . ↵ Fenske , L. et al. ( 2024 ) BakRep - a searchable large-scale web repository for bacterial genomes, characterizations and metadata . Microb Genom , 10 . ↵ Field , D. et al. ( 2008 ) Claude dePamphilis 18 Phillip Lord . Gálvez-Merchán , Á. et al. ( 2023 ) Metadata retrieval from sequence databases with ffq . Bioinformatics , 39 . ↵ Moustafa , A.M. et al. ( 2020 ) Comparative genomics in infectious disease . Curr Opin Microbiol , 53 , 61 – 70 . OpenUrl PubMed ↵ Mukherjee , S. et al. ( 2025 ) Genomes OnLine Database (GOLD) v.10: new features and updates . Nucleic Acids Res , 53 , D989 – D997 . OpenUrl PubMed ↵ Musen , M.A. et al. ( 2022 ) Modeling community standards for metadata as templates makes data FAIR . Sci Data , 9 . National Microbiome Data Collaborative https://github.com/microbiomedata/nmdc-schema . O’Cathail , C. et al. ( 2025 ) The European Nucleotide Archive in 2024 . Nucleic Acids Res , 53 , D49 – D55 . OpenUrl PubMed O’Leary , N.A. et al. ( 2024 ) Exploring and retrieving sequence and metadata for species across the tree of life with NCBI Datasets . Sci Data , 11 . ↵ Rajesh , A. et al. ( 2021 ) Improving the completeness of public metadata accompanying omics studies . Genome Biol , 22 . ↵ Reuter , J.A. et al. ( 2015 ) High-Throughput Sequencing Technologies . Mol Cell , 58 , 586 – 597 . OpenUrl CrossRef PubMed ↵ Ryan , M.J. et al. ( 2021 ) Towards a unified data infrastructure to support European and global microbiome research: a call to action . Environ Microbiol , 23 , 372 – 375 . OpenUrl CrossRef ↵ Sayers , E.W. et al. ( 2025 ) GenBank 2025 update . Nucleic Acids Res , 53 , D56 – D61 . OpenUrl CrossRef PubMed ↵ Wilkinson , M.D. et al. ( 2016 ) Comment: The FAIR Guiding Principles for scientific data management and stewardship . Sci Data , 3 . View the discussion thread. Back to top Previous Next Posted August 23, 2025. Download PDF Supplementary Material Data/Code Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following MetaMiner: Streamlined GUI Tool for Retrieving, Normalizing and Exploring Metadata Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share MetaMiner: Streamlined GUI Tool for Retrieving, Normalizing and Exploring Metadata Jaykumar Kiritkumar Patel , Ravikrishnan Elangovan bioRxiv 2025.08.20.666107; doi: https://doi.org/10.1101/2025.08.20.666107 Share This Article: Copy Citation Tools MetaMiner: Streamlined GUI Tool for Retrieving, Normalizing and Exploring Metadata Jaykumar Kiritkumar Patel , Ravikrishnan Elangovan bioRxiv 2025.08.20.666107; doi: https://doi.org/10.1101/2025.08.20.666107 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Bioinformatics Subject Areas All Articles Animal Behavior and Cognition (7635) Biochemistry (17697) Bioengineering (13895) Bioinformatics (41951) Biophysics (21456) Cancer Biology (18594) Cell Biology (25520) Clinical Trials (138) Developmental Biology (13381) Ecology (19903) Epidemiology (2067) Evolutionary Biology (24323) Genetics (15612) Genomics (22510) Immunology (17738) Microbiology (40401) Molecular Biology (17184) Neuroscience (88622) Paleontology (667) Pathology (2833) Pharmacology and Toxicology (4825) Physiology (7644) Plant Biology (15158) Scientific Communication and Education (2046) Synthetic Biology (4296) Systems Biology (9825) Zoology (2271)

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00