{"paper_id":"28948063-6ccd-4e58-b941-064c9a650955","body_text":"CircoTax: a novel circular representation of differentially abundant microbial ranks | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Method Article CircoTax: a novel circular representation of differentially abundant microbial ranks Leandro Di Gloria, Lorenzo Casbarra, Matteo Ramazzotti This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6002535/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract One of the most relevant and widely used applications of NGS-based microbial abundance analysis is the discovery of taxa that appear as differentially represented between ecological niches. In the last 10 years several methods have been proposed to accomplish this task. All such analyses invariantly report the differentially represented taxa at the appropriate taxonomic rank as well as a p-value (usually adjusted for multiple testing) and a fold change, representing the confidence on differential call and the amount of variation recorded, usually in log scale, respectively. This work describes CircoTax, an R function implementing a specialized ggplot2 graph to represent a rank-aware collection of differentially abundant taxa in an intuitive and compact radial form. Specifically, the CircoTax builds a circular plot with sectors encoding the taxonomy depth (from kingdom to genus) and, departing from the center, a number of radial bars that reach the appropriate sector and whose color and transparency are proportional to the log fold change intensity and direction. In addition to the manageability in providing custom tables to plot, our implementation also integrates fast and user-friendly ancillary functions which automate both the differential analysis and the display of the results while leaving to the user the choice of further customizations. The visualization of results is both graphically appealing, allowing to accommodate tens of differentially abundant ranks, and biologically informative since the amount of variation, the direction and rank are easily intelligible for each differentially represented taxa at a glance. microbial taxonomy differential abundance analysis microbiota rank abundance graphics Figures Figure 1 Figure 2 Introduction Microbial abundance analysis through NGS is a widely used and largely accepted approach to investigate communities of microorganisms and their relationships or alterations with the different environments [ 1 ]. Quantitative sequencing-based approaches (e.g. RNA-seq, barcode-seq etc.) offer the possibility of quantifying the reads attributed to specific individual molecules in samples, generating a so-called abundance table over which to apply statistical considerations and analytical strategies. Being the observed counts constrained within the total number of reads available for the experiment, the dataset is inherently of compositional nature [ 1 ] and, if left untreated, is hard to be analyzed with commonly used statistics [ 1 ]. In addition, the abnormal inflation of zeros is typical of this discipline [ 2 ]. On the other hand, recent literature proposes several strategies to directly adjust data for classical analyses or to apply non-normal distributions or rank-based analyses to improve data modeling and perform more appropriate investigations [ 3 – 5 ]. Beside statistical considerations, the analysis of abundance tables may be approached from several perspectives that can be generally divided into multivariate (e.g. clustering or principal coordinate analysis) based on beta diversity estimates aimed at describing the relationships among samples [ 6 ], or univariate (e.g. detection of differentially abundant taxa), aimed at detecting specific signals or markers that may at least in part explain the observed differences among samples, if any [ 4 ]. Eventually, certain comparisons reveal several significant signals, which necessitate the use of space-consuming plots to present them in a more human-readable format, rather than relying on tables rich in numerical values distributed across multiple rows. This holds true especially when the microbiota is explored, as the analysis often includes repeated comparisons across multiple taxonomic levels (phylum, class, order, family, genus, and species). Taxonomic information at different ranks can be obtained by direct comparison of each representative sequence against a closed or open set of pre-taxonomized references [ 7 ] or Bayesian inference using k-mers collections [ 8 ]. Motivated by the need of offering an improved yet intuitive representation of the results, we propose a highly informative radial plot, the “CircoTax” plot that exposes differentially abundant signals from abundance tables at the various taxonomic ranks in a compact format. This plot can be generated using publicly available and user-friendly R functions that allow customization. Optionally, we included dedicated functions to offer the flexibility to automate differential analysis, enhancing the efficiency of data handling and results visualization through CircoTax plot. Figure 1 offers a graphical perspective on the workflow we implemented to fuel CircoTax plotting. Although not mandatory for usage, the scheme represents a valid starting point to understand the overall CircoTax capabilities. Methods CircoTax input data The CircoTax plotting function basically receives a text table containing in each row a measure of effect size (usually a log fold change) and a (not-necessarily complete) taxonomic classification from genus/species to phylum. The effect size (and usually other relevant statistics such as the average signal and the adjusted p-value) can be obtained from a plethora of differential analysis methods [ 3 ]. A final table can be assembled in the statistical language R or loaded from an external file (see supplementary material for an example). Adding relevant experimental factors associated with samples, a phyloseq [ 9 ] R object can be optionally created to host all data in a dedicated framework. Taxonomic rank stratification Our CircoTax function can receive both an input table that combines quantitative differential signals (e.g. abundance fold changes) at various taxonomic ranks (usually phylum, class, order, family, genus and, optionally, species) or a phyloseq object that contains, in phyloseq jargon, otu_table, tax_table and sample_data (namely the table indicating that list sample identifiers and their corresponding group). For the scopes of CircoTax, the tax_glom function of phyloseq can be used to aggregate counts at a given taxonomic rank by summing the counts of their descendants. This procedure allows us to stratify the differential analyses across every rank. Table 1 summarizes a typical input table of CircoTax. It can be observed that the selection of significant taxa and their classification depth is different because the analyses are drawn at the different ranks independently. Accordingly, rows 1–2 contain significant results at the genus level, rows 3–4 at the family level and so on. An extended version of the table is available Supplementary Table 1 . Table 1 CircoTax input table format. The log2FC column is needed for coloring. The following 6 columns contain the name of the organisms at the different ranks. The last column is not included in the actual input, rather it has been included to identify the corresponding CircoTax sector. See supplementary material for a complete version of this table. log2FC Domain Phylum Class Order Family Genus Sector -0.741 Bacteria Proteobacteria NA NA NA NA P 5.294 Bacteria Chloroflexi NA NA NA NA P -1.221 Bacteria Proteobacteria Gammaproteobacteria NA NA NA C 2.146 Bacteria Bacteroidetes Flavobacteriia NA NA NA C 1.887 Bacteria Proteobacteria Alphaproteobacteria Caulobacterales NA NA O -3.029 Bacteria Proteobacteria Gammaproteobacteria Pasteurellales NA NA O 2.234 Bacteria Proteobacteria Alphaproteobacteria Caulobacterales Caulobacteraceae NA F -3.173 Bacteria Proteobacteria Gammaproteobacteria Pasteurellales Pasteurellaceae NA F 4.595 Bacteria Proteobacteria Alphaproteobacteria Rhizobiales Rhizobiaceae Shinella G 2.370 Bacteria Proteobacteria Alphaproteobacteria Caulobacterales Caulobacteraceae Brevundimonas G User’s custom input CircoTax is designed to automate the plotting of differentially abundant taxa at the different ranks starting from a phyloseq object. Nevertheless, CircoTax can also be plugged with a custom data frame according to the user's own preferences. In that case, the user only needs to specify three columns in a custom data frame (see Table 1 for an example) i.e. those referred to taxon names, taxonomic ranks and the fold change values. There are no particular requirements in this approach: the characters to plot will not be adjusted according to the classical CircoTax syntax and both the names and ranks order simply depend on the order of rows in the data frame. Integration of differential analysis tools The CircoTax plot function comes in a bundle with functions that facilitate differential abundance analysis (DAA). In fact, this information can be achieved in many ways and a large body of literature is available [ 3 – 5 ]. We chose to implement two well-known, relevant and technically distinct strategies currently used in the microbiota field, namely DESeq2 [ 10 ] and ALDEx2 [ 11 ], both operating in the R framework. The algorithm DESeq2 is designed for RNA-seq counts [ 10 ] but can also be performed on DNA-seq inferred abundances with acceptable robustness and efficiency [ 5 ]. In particular, DESeq2 features high sensitivity [ 3 ] which allows it to find differences between two similar environments despite maintaining a relatively high false discovery rate, probably due to the different bio-statistical assumptions on the variability of RNA and DNA observations [ 1 , 12 ]. On the other hand, the algorithm ALDEx2 features a particularly low false discovery rate and low sensitivity when used on microbiota abundances [ 3 , 5 , 12 ]. Providing functions for both the algorithms, the user has the opportunity to choose the best approach for the data under investigation. Our procedure takes advantage of R phyloseq package [ 9 ] to construct (tax_glom) rank-based abundance tables and store sample features (sample_data). We harmonized the input of the two DAA packages so that users can interrogate them with a common, easy to use and clearly intelligible syntax that respect R formula semantics and accommodate covariates (e.g. for paired analyses). The objective of this effort was to standardize as much as possible (though leaving room for personalization) the retrieval of statistically confident taxonomic ranks to plot. Further details are exposed in the associated manual and example script in the Github repository (see below), which offer practical guidance and step-by-step instructions for applying the described methods. Plot design implementation CircoTax is implemented as an R function that exploits the radial plotting capabilities of the ggplot2 R package for the representation of DAA on microbiota data. The plot is circular, composed by a number of sectors that accommodate the information proposed in the table. The taxonomic classification is used to decide how to sort the bars. Several modes are proposed, the main being (i) alphatax , with members of the same taxonomic lineage that stay close each other; (ii) rank , with bars are sorted by rank; (iii) (abs) fc with the order is decided by the value of the log fold change (optionally in absolute value). All sorting options are presented in Fig. 2 . Bars receive a color that follows a configurable three point color scale (default to the colorblind friendly orange-to-white-to-blue scale) so that the opposite changing directions and their intensities can be clearly visualized. The final graph further accommodates a color-scale bar and a title. Testing and validation CircoTax plots have been already proposed and validated in several datasets previously published by our group [ 13 – 16 ]. We further included in the GitHub repository (see below) a fully working script that reanalyzes the data from Niccolai et al. (2023) [ 14 ] and generates the CircoTax plot. The choice we made of maintaining full compatibility with the most used packages in the respective fields ensures that CircoTax will provide valid results over time. Results A novel information rich plot for microbial abundance analysis The overall result of CircoTax is a radial barplot in which bars represent the differentially abundant taxa that are stratified in 6 or 7 levels (from domain to genus or species) corresponding to the taxonomic depth and whose color and transparency depend on the log fold change of the differential analysis. This way of representing such results is innovative and intuitive and allows to easily recapitulate the relevant taxonomic ranks discriminating two conditions in a compact form. To exemplify CircoTax capability, Fig. 2 proposes a collection of plot variants that highlights the different sorting options. Such variants are not only relevant to offer alternative perspectives on data but can also disclose hidden peculiarities of the differential analysis results. In fact, each radial bar (sector) 1. reaches a sector that corresponds to a specific rank, making the reader immediately aware of its nature, 2. has a color and a transparency that clearly depicts the differential trend (more or less abundant) and intensity (how much more or less abundant) and 3. has a radially oriented label depicting the exact name of the represented element of the rank. Discussion The graphical representation of a differential analysis has been addressed by many strategies (and corresponding packages/programs) before. Most of them use colored bar plots possibly sorted by the (absolute) value of change (e.g. LEfSe [ 17 ]). In this case many panels are needed to represent an entire full rank differential analysis, so part of the results figures is frequently confined into supplementary data. Another frequently used representation is based on multiple paired box plots that represents, for each differentially abundant taxonomic rank, relevant metrics (frequently median and IQRs) of the distributions of the two groups under analysis. This is highly informative but, depending on the number of results to be displayed, it may require faceting of the figure, making the global picture and the overall message less intuitive. Sometimes the differential data is used to simply list the interesting taxa and extract the corresponding abundances from an abundance table, e.g. to elaborate a heatmap, possibly reordering the row/columns using distance-based dendrograms. In this case the amount of variation (so the relevance of the change) is lost and the original information is transformed into a complex image with a completely different scope. Anyway, also in this case many panels must be created since the results for different ranks must be elaborated separately. Lastly, KronaTools proposes since 2011 the Krona graph [ 18 ], that is frequently used in web pages to explore interactively the relevant ranks within a dataset through a radial plot. This tool is also used in a static way in recent published publications [ 19 ], but in that context its power cannot be fully appreciated due to the staticity of the published paper. The basic idea and motivation behind the CircoTax plot was to stick with the results of the differential analysis in a way that can cope with both taxonomic relationships among microbes while maintaining a direct view on the amount of change. Interestingly, the radial nature of the plot allows it to accommodate many results at different taxonomic ranks in the same plot that comprehensively represent the full analysis. For this reason, we are convinced that a CircoTax plot may drive information better than current existing strategies. Future development Currently, CircoTax is released as a free-to-use R script that we routinely use in manuscript production together with ancillary scripts that prepare its input tables. Furthermore, a detailed tutorial regarding installation, uses and technical details is provided. This will be constantly updated to follow CircoTax evolution (see data availability section). In the future the scripts for the DAA and the graphical representations of its results will be further refined according to user suggestions. Furthermore, the entire procedure will be released in the form of an R Shiny website where phyloseq objects can be received and processed in a coherent environment aimed at deciphering and interpreting the information contained in the datasets. Abbreviations DAA differential abundance analysis NGS next generation sequencing FC fold change Declarations Acknowledgments The authors wish to thank Giovanni Bacci for sharing ideas and brainstorming the plot. Ethics approval and consent to participate Not applicable Consent for publication Not applicable. Funding This work has been supported by RICATEN2021 funds from the University of Florence, Italy. Data availability This work reanalyzes data from \"Breast cancer: the first comparative evaluation of oncobiome composition between males and females\" that are already publicly available in NCBI Gene Expression Omnibus (GEO) repository at https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE212891. Code availability CircoTax is released as an open source set of analysis and plotting scripts available at the GitHub repository https://github.com/matteoramazzotti/CircoTax. The main functions have been further integrated with scripts that automate differential abundance analysis and automate the plotting. Furthermore, a tutorial has been prepared to clarify technical aspects, specific procedures and configurations. Competing interests All authors declare no competing interests. Authors' contributions MR conceived the strategy, did initial coding and supervised manuscript development. LDG harmonized the functions and contributed to manuscript drafting. LC optimized the code and contributed to manuscript drafting. References Lin H, Peddada SD. Analysis of microbial compositions: a review of normalization and differential abundance analysis. NPJ biofilms microbiomes. 2020;6:60. Silverman JD, Roche K, Mukherjee S, David LA. Naught all zeros in sequence count data are the same. Comput Struct Biotechnol J. 2020;18:2789–98. Yang L, Chen J. A comprehensive evaluation of microbial differential abundance analysis methods: current status and potential solutions. Microbiome. 2022;10:130. Cappellato M, Baruzzo G, Di Camillo B. Investigating differential abundance methods in microbiome data: A benchmark study. PLoS Comput Biol. 2022;18:e1010467. Calgaro M, Romualdi C, Waldron L, Risso D, Vitulo N. Assessment of statistical methods from single cell, bulk RNA-seq, and metagenomics applied to microbiome data. Genome Biol. 2020;21:191. Paliy O, Shankar V. Application of multivariate statistical techniques in microbial ecology. Mol Ecol. 2016;25:1032–57. Camacho C, Coulouris G, Avagyan V, Ma N, Papadopoulos J, Bealer K, Madden TL. BLAST+: architecture and applications. BMC Bioinformatics. 2009;10:421. Wang Q, Garrity GM, Tiedje JM, Cole JR. Naive Bayesian classifier for rapid assignment of rRNA sequences into the new bacterial taxonomy. Appl Environ Microbiol. 2007;73:5261–7. McMurdie PJ, Holmes S. phyloseq: an R package for reproducible interactive analysis and graphics of microbiome census data. PLoS ONE. 2013;8:e61217. Love MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15:550. Fernandes AD, Macklaim JM, Linn TG, Reid G, Gloor GB. ANOVA-like differential expression (ALDEx) analysis for mixed population RNA-Seq. PLoS ONE. 2013;8:e67019. Lin H, Peddada SD. Analysis of compositions of microbiomes with bias correction. Nat Commun. 2020;11:3514. Baldi S, Fabbrizzi A, Di Gloria L, Pallecchi M, Nannini G, D'Ambrosio M, Luceri C, Bartolucci G, Ramazzotti M, Fontana G, Mannini C, Lavorini F, Amedei A. First Exploration of the Altered Microbial Gut-Lung Axis in the Pathogenesis of Human Refractory Chronic Cough. Lung. 2024;202:107–18. Niccolai E, Baldi S, Nannini G, Gensini F, Papi L, Vezzosi V, Bianchi S, Orzalesi L, Ramazzotti M, Amedei A. Breast cancer: the first comparative evaluation of oncobiome composition between males and females. Biology sex differences. 2023;14:37. Russo E, Gloria LD, Nannini G, Meoni G, Niccolai E, Ringressi MN, Baldi S, Fani R, Tenori L, Taddei A, Ramazzotti M, Amedei A. (2023) From adenoma to CRC stages: the oral-gut microbiome axis as a source of potential microbial and metabolic biomarkers of malignancy, Neoplasia (New York, NY). 40, 100901. Niccolai E, Di Gloria L, Trolese MC, Fabbrizio P, Baldi S, Nannini G, Margotta C, Nastasi C, Ramazzotti M, Bartolucci G, Bendotti C, Nardo G, Amdei A. Host genetics and gut microbiota influence lipid metabolism and inflammation: potential implications for ALS pathophysiology in SOD1(G93A) mice. Acta Neuropathol Commun. 2024;12:174. Segata N, Izard J, Waldron L, Gevers D, Miropolsky L, Garrett WS, Huttenhower C. Metagenomic biomarker discovery and explanation. Genome Biol. 2011;12:R60. Ondov BD, Bergman NH, Phillippy AM. Interactive metagenomic visualization in a Web browser. BMC Bioinformatics. 2011;12:385. Amandito R, Malik A, Rohsiswatmo R. Metagenomic profiles of the early life microbiome of Indonesian inpatient neonates and their influence on clinical characteristics. Sci Rep. 2022;12:9413. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {\"props\":{\"pageProps\":{\"initialData\":{\"identity\":\"rs-6002535\",\"acceptedTermsAndConditions\":true,\"allowDirectSubmit\":true,\"archivedVersions\":[],\"articleType\":\"Method Article\",\"associatedPublications\":[],\"authors\":[{\"id\":415719972,\"identity\":\"ae4e5085-4480-4449-931c-9ec0a6ee3fcd\",\"order_by\":0,\"name\":\"Leandro Di Gloria\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Università degli Studi di Firenze\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Leandro\",\"middleName\":\"Di\",\"lastName\":\"Gloria\",\"suffix\":\"\"},{\"id\":415719973,\"identity\":\"05f6a188-29e5-4855-9465-3ca8deabb5ad\",\"order_by\":1,\"name\":\"Lorenzo Casbarra\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Università degli Studi di Firenze\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Lorenzo\",\"middleName\":\"\",\"lastName\":\"Casbarra\",\"suffix\":\"\"},{\"id\":415719974,\"identity\":\"7fceb254-064c-40f0-9caf-ff3ddb42f0b6\",\"order_by\":2,\"name\":\"Matteo Ramazzotti\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABBUlEQVRIiWNgGAWjYEgANvYGBoYEGwkGNpxKmBkbkHgGEmw8B4Ba0oBacOpB18IgkQCk00D2Ydeg237++IMPFQz5/NMOP/74o+JPHZ/kG7MHDxIsGPjkG7BqMTuTzNg44wyD5YzbaQYGEmeADpPOMTdISMDtMLMDyYzNvG0MBgy3EwwSDNvAWswkEn/g0XL+MWPzX6AW+dvpHw4kgrRInjGTwGvLDaAtjEAtBrdzDBsOgrRI8BDS8thwZs8ZCQPD2znFjA1njCXbeNLKQFp42NgScDgs8cGHHxU2BnK30zcDQ0yOX7798DbJHwl1cvLNB7BbAwESmEI8+NSPglEwCkbBKMAPAE/iUc3rDwKRAAAAAElFTkSuQmCC\",\"orcid\":\"\",\"institution\":\"Università degli Studi di Firenze\",\"correspondingAuthor\":true,\"prefix\":\"\",\"firstName\":\"Matteo\",\"middleName\":\"\",\"lastName\":\"Ramazzotti\",\"suffix\":\"\"}],\"badges\":[],\"createdAt\":\"2025-02-11 00:38:17\",\"currentVersionCode\":1,\"declarations\":\"\",\"doi\":\"10.21203/rs.3.rs-6002535/v1\",\"doiUrl\":\"https://doi.org/10.21203/rs.3.rs-6002535/v1\",\"draftVersion\":[],\"editorialEvents\":[],\"editorialNote\":\"\",\"failedWorkflow\":false,\"files\":[{\"id\":76425101,\"identity\":\"73d57271-eb6a-40a2-98e0-416703c7d9ab\",\"added_by\":\"auto\",\"created_at\":\"2025-02-17 05:26:20\",\"extension\":\"jpg\",\"order_by\":1,\"title\":\"Figure 1\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":677336,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eworkflow implemented to fuel CircoTax plotting. The R phyloseq object may conveniently store ASV or OTU abundance data that can be \\\"glommed\\\" to specific ranks and statistically tested for differential abundance with DESeq2 [10] and ALDEx2 [11], two largely accepted yet complementary methods. CircoTax can receive input from text output of such methods as well as from a properly formatted simple text table (see Table 1 for an example).\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"figure1.tif.jpg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-6002535/v1/3981bee2ba641d69b0c3935e.jpg\"},{\"id\":76425095,\"identity\":\"159e0d85-0d03-48bb-b615-460b91a514f9\",\"added_by\":\"auto\",\"created_at\":\"2025-02-17 05:26:20\",\"extension\":\"jpg\",\"order_by\":2,\"title\":\"Figure 2\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":1301233,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eplotting modes of CircoTax. Plots have a clockwise sorting of bars (sectors) whose ordering can be chosen among \\\"absLogFC\\\" (absolute fold change i.e. with decreasing transparency); \\\"alpha-lineage\\\" (alphabetical order of the highest rank); \\\"LogFC\\\" (log2 fold change); \\\"Alphabetical\\\" (alphabetical order); \\\"Rank\\\" (from genus to phylum). The plots refer to a reanalysis of data from Niccolai et al. (2023) [14] and are exposed in tabular format in Supplementary Table 1.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"Figure2.tif.jpg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-6002535/v1/3ddd76c8c52ad291467d7452.jpg\"},{\"id\":78606530,\"identity\":\"e84a40f0-80a3-40b0-9157-897fa4b1156b\",\"added_by\":\"auto\",\"created_at\":\"2025-03-16 11:46:29\",\"extension\":\"pdf\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":2510002,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"manuscript.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-6002535/v1/aedd96ec-c004-4a8c-97cb-756819ffdf96.pdf\"}],\"financialInterests\":\"No competing interests reported.\",\"formattedTitle\":\"CircoTax: a novel circular representation of differentially abundant microbial ranks\",\"fulltext\":[{\"header\":\"Introduction\",\"content\":\"\\u003cp\\u003eMicrobial abundance analysis through NGS is a widely used and largely accepted approach to investigate communities of microorganisms and their relationships or alterations with the different environments [\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e]. Quantitative sequencing-based approaches (e.g. RNA-seq, barcode-seq etc.) offer the possibility of quantifying the reads attributed to specific individual molecules in samples, generating a so-called abundance table over which to apply statistical considerations and analytical strategies. Being the observed counts constrained within the total number of reads available for the experiment, the dataset is inherently of compositional nature [\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e] and, if left untreated, is hard to be analyzed with commonly used statistics [\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e]. In addition, the abnormal inflation of zeros is typical of this discipline [\\u003cspan citationid=\\\"CR2\\\" class=\\\"CitationRef\\\"\\u003e2\\u003c/span\\u003e]. On the other hand, recent literature proposes several strategies to directly adjust data for classical analyses or to apply non-normal distributions or rank-based analyses to improve data modeling and perform more appropriate investigations [\\u003cspan additionalcitationids=\\\"CR4\\\" citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e3\\u003c/span\\u003e\\u0026ndash;\\u003cspan citationid=\\\"CR5\\\" class=\\\"CitationRef\\\"\\u003e5\\u003c/span\\u003e].\\u003c/p\\u003e \\u003cp\\u003eBeside statistical considerations, the analysis of abundance tables may be approached from several perspectives that can be generally divided into multivariate (e.g. clustering or principal coordinate analysis) based on beta diversity estimates aimed at describing the relationships among samples [\\u003cspan citationid=\\\"CR6\\\" class=\\\"CitationRef\\\"\\u003e6\\u003c/span\\u003e], or univariate (e.g. detection of differentially abundant taxa), aimed at detecting specific signals or markers that may at least in part explain the observed differences among samples, if any [\\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e4\\u003c/span\\u003e]. Eventually, certain comparisons reveal several significant signals, which necessitate the use of space-consuming plots to present them in a more human-readable format, rather than relying on tables rich in numerical values distributed across multiple rows. This holds true especially when the microbiota is explored, as the analysis often includes repeated comparisons across multiple taxonomic levels (phylum, class, order, family, genus, and species). Taxonomic information at different ranks can be obtained by direct comparison of each representative sequence against a closed or open set of pre-taxonomized references [\\u003cspan citationid=\\\"CR7\\\" class=\\\"CitationRef\\\"\\u003e7\\u003c/span\\u003e] or Bayesian inference using k-mers collections [\\u003cspan citationid=\\\"CR8\\\" class=\\\"CitationRef\\\"\\u003e8\\u003c/span\\u003e]. Motivated by the need of offering an improved yet intuitive representation of the results, we propose a highly informative radial plot, the \\u0026ldquo;CircoTax\\u0026rdquo; plot that exposes differentially abundant signals from abundance tables at the various taxonomic ranks in a compact format. This plot can be generated using publicly available and user-friendly R functions that allow customization. Optionally, we included dedicated functions to offer the flexibility to automate differential analysis, enhancing the efficiency of data handling and results visualization through CircoTax plot. Figure\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e offers a graphical perspective on the workflow we implemented to fuel CircoTax plotting. Although not mandatory for usage, the scheme represents a valid starting point to understand the overall CircoTax capabilities.\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e\"},{\"header\":\"Methods\",\"content\":\"\\u003cdiv id=\\\"Sec3\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003eCircoTax input data\\u003c/h2\\u003e \\u003cp\\u003eThe CircoTax plotting function basically receives a text table containing in each row a measure of effect size (usually a log fold change) and a (not-necessarily complete) taxonomic classification from genus/species to phylum. The effect size (and usually other relevant statistics such as the average signal and the adjusted p-value) can be obtained from a plethora of differential analysis methods [\\u003cspan citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e3\\u003c/span\\u003e]. A final table can be assembled in the statistical language R or loaded from an external file (see supplementary material for an example). Adding relevant experimental factors associated with samples, a phyloseq [\\u003cspan citationid=\\\"CR9\\\" class=\\\"CitationRef\\\"\\u003e9\\u003c/span\\u003e] R object can be optionally created to host all data in a dedicated framework.\\u003c/p\\u003e \\u003c/div\\u003e\\n\\u003ch3\\u003eTaxonomic rank stratification\\u003c/h3\\u003e\\n\\u003cp\\u003eOur CircoTax function can receive both an input table that combines quantitative differential signals (e.g. abundance fold changes) at various taxonomic ranks (usually phylum, class, order, family, genus and, optionally, species) or a phyloseq object that contains, in phyloseq jargon, otu_table, tax_table and sample_data (namely the table indicating that list sample identifiers and their corresponding group). For the scopes of CircoTax, the tax_glom function of phyloseq can be used to aggregate counts at a given taxonomic rank by summing the counts of their descendants. This procedure allows us to stratify the differential analyses across every rank.\\u003c/p\\u003e \\u003cp\\u003eTable\\u0026nbsp;\\u003cspan refid=\\\"Tab1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e summarizes a typical input table of CircoTax. It can be observed that the selection of significant taxa and their classification depth is different because the analyses are drawn at the different ranks independently. Accordingly, rows 1\\u0026ndash;2 contain significant results at the genus level, rows 3\\u0026ndash;4 at the family level and so on. An extended version of the table is available \\u003cb\\u003eSupplementary Table\\u0026nbsp;1\\u003c/b\\u003e.\\u003c/p\\u003e \\u003cp\\u003e\\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab1\\\" border=\\\"1\\\"\\u003e\\u003ccaption language=\\\"En\\\"\\u003e\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 1\\u003c/div\\u003e\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\u003cp\\u003eCircoTax input table format. The log2FC column is needed for coloring. The following 6 columns contain the name of the organisms at the different ranks. The last column is not included in the actual input, rather it has been included to identify the corresponding CircoTax sector. See supplementary material for a complete version of this table.\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/caption\\u003e\\u003ccolgroup cols=\\\"8\\\"\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c5\\\" colnum=\\\"5\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c6\\\" colnum=\\\"6\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c7\\\" colnum=\\\"7\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c8\\\" colnum=\\\"8\\\"\\u003e\\u003c/div\\u003e\\u003cthead\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003elog2FC\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eDomain\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003ePhylum\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eClass\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eOrder\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eFamily\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eGenus\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eSector\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003c/thead\\u003e\\u003ctbody\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e-0.741\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eBacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eProteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eP\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e5.294\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eBacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eChloroflexi\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eP\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e-1.221\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eBacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eProteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eGammaproteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eC\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e2.146\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eBacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eBacteroidetes\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eFlavobacteriia\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eC\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e1.887\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eBacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eProteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eAlphaproteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eCaulobacterales\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eO\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e-3.029\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eBacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eProteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eGammaproteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003ePasteurellales\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eO\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e2.234\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eBacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eProteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eAlphaproteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eCaulobacterales\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eCaulobacteraceae\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eF\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e-3.173\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eBacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eProteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eGammaproteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003ePasteurellales\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003ePasteurellaceae\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eNA\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eF\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e4.595\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eBacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eProteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eAlphaproteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eRhizobiales\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eRhizobiaceae\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eShinella\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eG\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e2.370\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eBacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eProteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eAlphaproteobacteria\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eCaulobacterales\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eCaulobacteraceae\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eBrevundimonas\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eG\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003c/tbody\\u003e\\u003c/colgroup\\u003e\\u003c/table\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\n\\u003ch3\\u003eUser’s custom input\\u003c/h3\\u003e\\n\\u003cp\\u003eCircoTax is designed to automate the plotting of differentially abundant taxa at the different ranks starting from a phyloseq object. Nevertheless, CircoTax can also be plugged with a custom data frame according to the user's own preferences. In that case, the user only needs to specify three columns in a custom data frame (see Table\\u0026nbsp;\\u003cspan refid=\\\"Tab1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e for an example) i.e. those referred to taxon names, taxonomic ranks and the fold change values. There are no particular requirements in this approach: the characters to plot will not be adjusted according to the classical CircoTax syntax and both the names and ranks order simply depend on the order of rows in the data frame.\\u003c/p\\u003e\\n\\u003ch3\\u003eIntegration of differential analysis tools\\u003c/h3\\u003e\\n\\u003cp\\u003eThe CircoTax plot function comes in a bundle with functions that facilitate differential abundance analysis (DAA). In fact, this information can be achieved in many ways and a large body of literature is available [\\u003cspan additionalcitationids=\\\"CR4\\\" citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e3\\u003c/span\\u003e\\u0026ndash;\\u003cspan citationid=\\\"CR5\\\" class=\\\"CitationRef\\\"\\u003e5\\u003c/span\\u003e]. We chose to implement two well-known, relevant and technically distinct strategies currently used in the microbiota field, namely DESeq2 [\\u003cspan citationid=\\\"CR10\\\" class=\\\"CitationRef\\\"\\u003e10\\u003c/span\\u003e] and ALDEx2 [\\u003cspan citationid=\\\"CR11\\\" class=\\\"CitationRef\\\"\\u003e11\\u003c/span\\u003e], both operating in the R framework. The algorithm DESeq2 is designed for RNA-seq counts [\\u003cspan citationid=\\\"CR10\\\" class=\\\"CitationRef\\\"\\u003e10\\u003c/span\\u003e] but can also be performed on DNA-seq inferred abundances with acceptable robustness and efficiency [\\u003cspan citationid=\\\"CR5\\\" class=\\\"CitationRef\\\"\\u003e5\\u003c/span\\u003e]. In particular, DESeq2 features high sensitivity [\\u003cspan citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e3\\u003c/span\\u003e] which allows it to find differences between two similar environments despite maintaining a relatively high false discovery rate, probably due to the different bio-statistical assumptions on the variability of RNA and DNA observations [\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e, \\u003cspan citationid=\\\"CR12\\\" class=\\\"CitationRef\\\"\\u003e12\\u003c/span\\u003e]. On the other hand, the algorithm ALDEx2 features a particularly low false discovery rate and low sensitivity when used on microbiota abundances [\\u003cspan citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e3\\u003c/span\\u003e, \\u003cspan citationid=\\\"CR5\\\" class=\\\"CitationRef\\\"\\u003e5\\u003c/span\\u003e, \\u003cspan citationid=\\\"CR12\\\" class=\\\"CitationRef\\\"\\u003e12\\u003c/span\\u003e]. Providing functions for both the algorithms, the user has the opportunity to choose the best approach for the data under investigation. Our procedure takes advantage of R phyloseq package [\\u003cspan citationid=\\\"CR9\\\" class=\\\"CitationRef\\\"\\u003e9\\u003c/span\\u003e] to construct (tax_glom) rank-based abundance tables and store sample features (sample_data). We harmonized the input of the two DAA packages so that users can interrogate them with a common, easy to use and clearly intelligible syntax that respect R formula semantics and accommodate covariates (e.g. for paired analyses). The objective of this effort was to standardize as much as possible (though leaving room for personalization) the retrieval of statistically confident taxonomic ranks to plot. Further details are exposed in the associated manual and example script in the Github repository (see below), which offer practical guidance and step-by-step instructions for applying the described methods.\\u003c/p\\u003e\\n\\u003ch3\\u003ePlot design implementation\\u003c/h3\\u003e\\n\\u003cp\\u003eCircoTax is implemented as an R function that exploits the radial plotting capabilities of the ggplot2 R package for the representation of DAA on microbiota data. The plot is circular, composed by a number of sectors that accommodate the information proposed in the table. The taxonomic classification is used to decide how to sort the bars. Several modes are proposed, the main being (i) \\u003cem\\u003ealphatax\\u003c/em\\u003e, with members of the same taxonomic lineage that stay close each other; (ii) \\u003cem\\u003erank\\u003c/em\\u003e, with bars are sorted by rank; (iii) (abs)\\u003cem\\u003efc\\u003c/em\\u003e with the order is decided by the value of the log fold change (optionally in absolute value). All sorting options are presented in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e. Bars receive a color that follows a configurable three point color scale (default to the colorblind friendly orange-to-white-to-blue scale) so that the opposite changing directions and their intensities can be clearly visualized. The final graph further accommodates a color-scale bar and a title.\\u003c/p\\u003e \\u003cdiv id=\\\"Sec8\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003eTesting and validation\\u003c/h2\\u003e \\u003cp\\u003eCircoTax plots have been already proposed and validated in several datasets previously published by our group [\\u003cspan additionalcitationids=\\\"CR14 CR15\\\" citationid=\\\"CR13\\\" class=\\\"CitationRef\\\"\\u003e13\\u003c/span\\u003e\\u0026ndash;\\u003cspan citationid=\\\"CR16\\\" class=\\\"CitationRef\\\"\\u003e16\\u003c/span\\u003e]. We further included in the GitHub repository (see below) a fully working script that reanalyzes the data from Niccolai et al. (2023) [\\u003cspan citationid=\\\"CR14\\\" class=\\\"CitationRef\\\"\\u003e14\\u003c/span\\u003e] and generates the CircoTax plot. The choice we made of maintaining full compatibility with the most used packages in the respective fields ensures that CircoTax will provide valid results over time.\\u003c/p\\u003e \\u003cp\\u003e \\u003c/p\\u003e \\u003c/div\\u003e\"},{\"header\":\"Results\",\"content\":\"\\u003cdiv id=\\\"Sec10\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003eA novel information rich plot for microbial abundance analysis\\u003c/h2\\u003e \\u003cp\\u003eThe overall result of CircoTax is a radial barplot in which bars represent the differentially abundant taxa that are stratified in 6 or 7 levels (from domain to genus or species) corresponding to the taxonomic depth and whose color and transparency depend on the log fold change of the differential analysis. This way of representing such results is innovative and intuitive and allows to easily recapitulate the relevant taxonomic ranks discriminating two conditions in a compact form. To exemplify CircoTax capability, Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e proposes a collection of plot variants that highlights the different sorting options. Such variants are not only relevant to offer alternative perspectives on data but can also disclose hidden peculiarities of the differential analysis results. In fact, each radial bar (sector) 1. reaches a sector that corresponds to a specific rank, making the reader immediately aware of its nature, 2. has a color and a transparency that clearly depicts the differential trend (more or less abundant) and intensity (how much more or less abundant) and 3. has a radially oriented label depicting the exact name of the represented element of the rank.\\u003c/p\\u003e \\u003c/div\\u003e\"},{\"header\":\"Discussion\",\"content\":\"\\u003cp\\u003eThe graphical representation of a differential analysis has been addressed by many strategies (and corresponding packages/programs) before. Most of them use colored bar plots possibly sorted by the (absolute) value of change (e.g. LEfSe [\\u003cspan citationid=\\\"CR17\\\" class=\\\"CitationRef\\\"\\u003e17\\u003c/span\\u003e]). In this case many panels are needed to represent an entire full rank differential analysis, so part of the results figures is frequently confined into supplementary data.\\u003c/p\\u003e \\u003cp\\u003eAnother frequently used representation is based on multiple paired box plots that represents, for each differentially abundant taxonomic rank, relevant metrics (frequently median and IQRs) of the distributions of the two groups under analysis. This is highly informative but, depending on the number of results to be displayed, it may require faceting of the figure, making the global picture and the overall message less intuitive.\\u003c/p\\u003e \\u003cp\\u003eSometimes the differential data is used to simply list the interesting taxa and extract the corresponding abundances from an abundance table, e.g. to elaborate a heatmap, possibly reordering the row/columns using distance-based dendrograms. In this case the amount of variation (so the relevance of the change) is lost and the original information is transformed into a complex image with a completely different scope. Anyway, also in this case many panels must be created since the results for different ranks must be elaborated separately. Lastly, KronaTools proposes since 2011 the Krona graph [\\u003cspan citationid=\\\"CR18\\\" class=\\\"CitationRef\\\"\\u003e18\\u003c/span\\u003e], that is frequently used in web pages to explore interactively the relevant ranks within a dataset through a radial plot. This tool is also used in a static way in recent published publications [\\u003cspan citationid=\\\"CR19\\\" class=\\\"CitationRef\\\"\\u003e19\\u003c/span\\u003e], but in that context its power cannot be fully appreciated due to the staticity of the published paper.\\u003c/p\\u003e \\u003cp\\u003eThe basic idea and motivation behind the CircoTax plot was to stick with the results of the differential analysis in a way that can cope with both taxonomic relationships among microbes while maintaining a direct view on the amount of change. Interestingly, the radial nature of the plot allows it to accommodate many results at different taxonomic ranks in the same plot that comprehensively represent the full analysis. For this reason, we are convinced that a CircoTax plot may drive information better than current existing strategies.\\u003c/p\\u003e \\u003cdiv id=\\\"Sec12\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003eFuture development\\u003c/h2\\u003e \\u003cp\\u003eCurrently, CircoTax is released as a free-to-use R script that we routinely use in manuscript production together with ancillary scripts that prepare its input tables. Furthermore, a detailed tutorial regarding installation, uses and technical details is provided. This will be constantly updated to follow CircoTax evolution (see data availability section). In the future the scripts for the DAA and the graphical representations of its results will be further refined according to user suggestions. Furthermore, the entire procedure will be released in the form of an R Shiny website where phyloseq objects can be received and processed in a coherent environment aimed at deciphering and interpreting the information contained in the datasets.\\u003c/p\\u003e \\u003c/div\\u003e\"},{\"header\":\"Abbreviations\",\"content\":\"\\u003cdiv class=\\\"DefinitionList\\\"\\u003e \\u003cdiv class=\\\"DefinitionListEntry\\\"\\u003e \\u003cdiv class=\\\"Term\\\"\\u003eDAA\\u003c/div\\u003e \\u003cdiv class=\\\"Description\\\"\\u003e \\u003cp\\u003edifferential abundance analysis\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/div\\u003e \\u003cdiv class=\\\"DefinitionListEntry\\\"\\u003e \\u003cdiv class=\\\"Term\\\"\\u003eNGS\\u003c/div\\u003e \\u003cdiv class=\\\"Description\\\"\\u003e \\u003cp\\u003enext generation sequencing\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/div\\u003e \\u003cdiv class=\\\"DefinitionListEntry\\\"\\u003e \\u003cdiv class=\\\"Term\\\"\\u003eFC\\u003c/div\\u003e \\u003cdiv class=\\\"Description\\\"\\u003e \\u003cp\\u003efold change\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/div\\u003e \\u003c/div\\u003e\"},{\"header\":\"Declarations\",\"content\":\"\\u003cp\\u003e\\u003cstrong\\u003eAcknowledgments\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eThe authors wish to thank Giovanni Bacci for sharing ideas and brainstorming the plot.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eEthics approval and consent to participate\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eNot applicable\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eConsent for publication\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eNot applicable.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eFunding\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eThis work has been supported by RICATEN2021 funds from the University of Florence, Italy.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eData availability\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eThis work reanalyzes data from \\u0026quot;Breast cancer: the first comparative evaluation of oncobiome composition between males and females\\u0026quot; that are already publicly available in NCBI Gene Expression Omnibus (GEO) repository at https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE212891.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eCode availability\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eCircoTax is released as an open source set of analysis and plotting scripts available at the GitHub repository https://github.com/matteoramazzotti/CircoTax. The main functions have been further integrated with scripts that automate differential abundance analysis and automate the plotting. Furthermore, a tutorial has been prepared to clarify technical aspects, specific procedures and configurations.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eCompeting interests\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eAll authors declare no competing interests.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cstrong\\u003eAuthors\\u0026apos; contributions\\u003c/strong\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eMR conceived the strategy, did initial coding and supervised manuscript development. LDG harmonized the functions and contributed to manuscript drafting. LC optimized the code \\u0026nbsp; and contributed to manuscript drafting.\\u003c/p\\u003e\"},{\"header\":\"References\",\"content\":\"\\u003col\\u003e\\u003cli\\u003e\\u003cspan\\u003eLin H, Peddada SD. Analysis of microbial compositions: a review of normalization and differential abundance analysis. NPJ biofilms microbiomes. 2020;6:60.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eSilverman JD, Roche K, Mukherjee S, David LA. Naught all zeros in sequence count data are the same. Comput Struct Biotechnol J. 2020;18:2789\\u0026ndash;98.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eYang L, Chen J. A comprehensive evaluation of microbial differential abundance analysis methods: current status and potential solutions. Microbiome. 2022;10:130.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eCappellato M, Baruzzo G, Di Camillo B. Investigating differential abundance methods in microbiome data: A benchmark study. PLoS Comput Biol. 2022;18:e1010467.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eCalgaro M, Romualdi C, Waldron L, Risso D, Vitulo N. Assessment of statistical methods from single cell, bulk RNA-seq, and metagenomics applied to microbiome data. Genome Biol. 2020;21:191.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003ePaliy O, Shankar V. Application of multivariate statistical techniques in microbial ecology. Mol Ecol. 2016;25:1032\\u0026ndash;57.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eCamacho C, Coulouris G, Avagyan V, Ma N, Papadopoulos J, Bealer K, Madden TL. BLAST+: architecture and applications. BMC Bioinformatics. 2009;10:421.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eWang Q, Garrity GM, Tiedje JM, Cole JR. Naive Bayesian classifier for rapid assignment of rRNA sequences into the new bacterial taxonomy. Appl Environ Microbiol. 2007;73:5261\\u0026ndash;7.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eMcMurdie PJ, Holmes S. phyloseq: an R package for reproducible interactive analysis and graphics of microbiome census data. PLoS ONE. 2013;8:e61217.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eLove MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15:550.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eFernandes AD, Macklaim JM, Linn TG, Reid G, Gloor GB. ANOVA-like differential expression (ALDEx) analysis for mixed population RNA-Seq. PLoS ONE. 2013;8:e67019.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eLin H, Peddada SD. Analysis of compositions of microbiomes with bias correction. Nat Commun. 2020;11:3514.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eBaldi S, Fabbrizzi A, Di Gloria L, Pallecchi M, Nannini G, D'Ambrosio M, Luceri C, Bartolucci G, Ramazzotti M, Fontana G, Mannini C, Lavorini F, Amedei A. First Exploration of the Altered Microbial Gut-Lung Axis in the Pathogenesis of Human Refractory Chronic Cough. Lung. 2024;202:107\\u0026ndash;18.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eNiccolai E, Baldi S, Nannini G, Gensini F, Papi L, Vezzosi V, Bianchi S, Orzalesi L, Ramazzotti M, Amedei A. Breast cancer: the first comparative evaluation of oncobiome composition between males and females. Biology sex differences. 2023;14:37.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eRusso E, Gloria LD, Nannini G, Meoni G, Niccolai E, Ringressi MN, Baldi S, Fani R, Tenori L, Taddei A, Ramazzotti M, Amedei A. (2023) From adenoma to CRC stages: the oral-gut microbiome axis as a source of potential microbial and metabolic biomarkers of malignancy, \\u003cem\\u003eNeoplasia (New York, NY).\\u003c/em\\u003e 40, 100901.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eNiccolai E, Di Gloria L, Trolese MC, Fabbrizio P, Baldi S, Nannini G, Margotta C, Nastasi C, Ramazzotti M, Bartolucci G, Bendotti C, Nardo G, Amdei A. Host genetics and gut microbiota influence lipid metabolism and inflammation: potential implications for ALS pathophysiology in SOD1(G93A) mice. Acta Neuropathol Commun. 2024;12:174.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eSegata N, Izard J, Waldron L, Gevers D, Miropolsky L, Garrett WS, Huttenhower C. Metagenomic biomarker discovery and explanation. Genome Biol. 2011;12:R60.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eOndov BD, Bergman NH, Phillippy AM. Interactive metagenomic visualization in a Web browser. BMC Bioinformatics. 2011;12:385.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eAmandito R, Malik A, Rohsiswatmo R. Metagenomic profiles of the early life microbiome of Indonesian inpatient neonates and their influence on clinical characteristics. Sci Rep. 2022;12:9413.\\u003c/span\\u003e\\u003c/li\\u003e\\u003c/ol\\u003e\"}],\"fulltextSource\":\"\",\"fullText\":\"\",\"funders\":[],\"hasAdminPriorityOnWorkflow\":false,\"hasManuscriptDocX\":true,\"hasOptedInToPreprint\":true,\"hasPassedJournalQc\":\"\",\"hasAnyPriority\":false,\"hideJournal\":true,\"highlight\":\"\",\"institution\":\"\",\"isAcceptedByJournal\":false,\"isAuthorSuppliedPdf\":false,\"isDeskRejected\":\"\",\"isHiddenFromSearch\":false,\"isInQc\":false,\"isInWorkflow\":false,\"isPdf\":false,\"isPdfUpToDate\":true,\"isWithdrawnOrRetracted\":false,\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true},\"keywords\":\"microbial taxonomy, differential abundance analysis, microbiota, rank abundance, graphics\",\"lastPublishedDoi\":\"10.21203/rs.3.rs-6002535/v1\",\"lastPublishedDoiUrl\":\"https://doi.org/10.21203/rs.3.rs-6002535/v1\",\"license\":{\"name\":\"CC BY 4.0\",\"url\":\"https://creativecommons.org/licenses/by/4.0/\"},\"manuscriptAbstract\":\"\\u003cp\\u003eOne of the most relevant and widely used applications of NGS-based microbial abundance analysis is the discovery of taxa that appear as differentially represented between ecological niches. In the last 10 years several methods have been proposed to accomplish this task. All such analyses invariantly report the differentially represented taxa at the appropriate taxonomic rank as well as a p-value (usually adjusted for multiple testing) and a fold change, representing the confidence on differential call and the amount of variation recorded, usually in log scale, respectively. This work describes CircoTax, an R function implementing a specialized ggplot2 graph to represent a rank-aware collection of differentially abundant taxa in an intuitive and compact radial form. Specifically, the CircoTax builds a circular plot with sectors encoding the taxonomy depth (from kingdom to genus) and, departing from the center, a number of radial bars that reach the appropriate sector and whose color and transparency are proportional to the log fold change intensity and direction. In addition to the manageability in providing custom tables to plot, our implementation also integrates fast and user-friendly ancillary functions which automate both the differential analysis and the display of the results while leaving to the user the choice of further customizations. The visualization of results is both graphically appealing, allowing to accommodate tens of differentially abundant ranks, and biologically informative since the amount of variation, the direction and rank are easily intelligible for each differentially represented taxa at a glance.\\u003c/p\\u003e\",\"manuscriptTitle\":\"CircoTax: a novel circular representation of differentially abundant microbial ranks\",\"msid\":\"\",\"msnumber\":\"\",\"nonDraftVersions\":[{\"code\":1,\"date\":\"2025-02-17 05:25:49\",\"doi\":\"10.21203/rs.3.rs-6002535/v1\",\"editorialEvents\":[{\"type\":\"communityComments\",\"content\":0}],\"status\":\"published\",\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true}}],\"origin\":\"\",\"ownerIdentity\":\"f7a7f9f6-24f2-4514-ba2f-3833a70965ae\",\"owner\":[],\"postedDate\":\"February 17th, 2025\",\"published\":true,\"recentEditorialEvents\":[],\"rejectedJournal\":[],\"revision\":\"\",\"amendment\":\"\",\"status\":\"posted\",\"subjectAreas\":[],\"tags\":[],\"updatedAt\":\"2025-03-19T15:38:23+00:00\",\"versionOfRecord\":[],\"versionCreatedAt\":\"2025-02-17 05:25:49\",\"video\":\"\",\"vorDoi\":\"\",\"vorDoiUrl\":\"\",\"workflowStages\":[]},\"version\":\"v1\",\"identity\":\"rs-6002535\",\"journalConfig\":\"researchsquare\"},\"__N_SSP\":true},\"page\":\"/article/[identity]/[[...version]]\",\"query\":{\"redirect\":\"/article/rs-6002535\",\"identity\":\"rs-6002535\",\"version\":[\"v1\"]},\"buildId\":\"8U1c8b4HqxoKbykW_rLl7\",\"isFallback\":false,\"isExperimentalCompile\":false,\"dynamicIds\":[84888],\"gssp\":true,\"scriptLoader\":[]}","source_license":"CC-BY-4.0","license_restricted":false}