Patch seriation to visualize data and model parameters

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract We developed a new seriation merit function for enhancing the visual information of data matrices. A local similarity matrix is calculated, where the average similarity of a neighbouring objects is calculated in a limited variable space and a global function is constructed to maximize the local similarities and cluster them into patches by simple row and column ordering. The method identifies data clusters in a powerful way, if the similarity of objects is caused by some variables and these variables differ for the distinct clusters. The method can be used in the presence of missing data and also on more than two-dimensional data arrays. We show the feasibility of the method on different data sets: on QSAR, chemical, material science, food science, cheminformatics and environmental data in two- and three-dimensional cases. The method can be used during the development and the interpretation of artificial neural network models by seriating different features of the models. It helps to identify interpretable models by elucidating clusters of objects, variables and hidden layer neurons.
Full text 194,800 characters · extracted from preprint-html · click to expand
Patch seriation to visualize data and model parameters | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Patch seriation to visualize data and model parameters Rita Lasfar, Gergely Tóth This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2780120/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 09 Sep, 2023 Read the published version in Journal of Cheminformatics → Version 1 posted 4 You are reading this latest preprint version Abstract We developed a new seriation merit function for enhancing the visual information of data matrices. A local similarity matrix is calculated, where the average similarity of a neighbouring objects is calculated in a limited variable space and a global function is constructed to maximize the local similarities and cluster them into patches by simple row and column ordering. The method identifies data clusters in a powerful way, if the similarity of objects is caused by some variables and these variables differ for the distinct clusters. The method can be used in the presence of missing data and also on more than two-dimensional data arrays. We show the feasibility of the method on different data sets: on QSAR, chemical, material science, food science, cheminformatics and environmental data in two- and three-dimensional cases. The method can be used during the development and the interpretation of artificial neural network models by seriating different features of the models. It helps to identify interpretable models by elucidating clusters of objects, variables and hidden layer neurons. seriation data visualization model interpretation clustering neural network model Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 1. Introduction Seriation? Most of scientist involved in it without knowing the term. If one knows its practical definition, namely, how to do row and/or column permutations to enhance visual perception of a table or heatmap, it is clear for scientist that they have faced with the problem. Its first application goes back to the XIXth century [ 1 ], when it was an explanatory technique to order objects in a way to reveal patterns and regular features easily. Later it spread to all fields of science and the ordering often concern two sequences to be reordered [ 2 – 5 ]. There are, e.g., possibilities to order objects along two axes or one object and one variable sequences in a table. The first applications were connected to fields, where visualization or chronological sequence were natural (archeology, cartography, history, operation research, sociology). Later, especially when information technology is present, different methods and application appeared in many other fields (anthropology, graphics, information visualization, sociometry, psychology, psychometry, ecology, biology, bioinformatics, etc…). The “common” in the methods that they are not common for all fields of science. There is a rather small communication among the fields. The most general review was written by Liiv [ 5 ], where a historical overview of seriation is detailed including the milestones at several application fields. Instead of enumerating here the methods and provide a deficient and scanty list of applications, we forward the reader to the review of Liiv [ 5 ]. Seriation is applied in a latent way in chemistry [ 6 – 9 ], but it is seldom termed. It is often used in many scientific software as a default setting, that, e.g., hierarchical clustering is applied on objects and a visually acceptable sequence is generated using a seriated dendogram [ 10 – 13 ]. At the interdisciplinary cheminformatics, the methods are used consciously according to its large emphasize in bioinformatics, where the number of special methods and the corresponding applications is increasing up to now there. From these we mention only the bi- or co-clustering [ 14 – 15 ]. From the special applications in bioinformatics we may refer to similarity search and alignment methods, where our references are some recent reviews. Similarity searching is widely applied and encompasses many techniques, its principle is based on detecting small molecules that have the same biological activity [ 16 ]. By the same logic, alignment is based on hypothesis of homology [ 17 ]. The real number of alignment algorithms is in hundreds and continue to increase. The alignment can permit, among others, the identification and quantification of conserved regions or functional motifs, profiling of genetic disease. Going back to chemistry, we found up to now only a few articles, where the word seriation is used in the title, abstract or in the keywords [ 18 – 21 ]. One interesting example for seriation in chemistry, while not the main subject of the article, showed its importance in summarising the resulting relationships between production groups and chemical clusters, and has enabled external information to be compared with the cluster results. This has permit eventually, to validate and interpret the clusters [ 18 ]. The main aim of seriation is to get better visualization by introducing some order by appropriate sequencing. The ordered sequence may help to find similar objects or variables close to each other in vectors, tables or in their graphics (e.g., in heatmaps). In some cases, it can be the first visual check of data and it serves as a good starting point to estimate which enhanced data analysis method might be tried. Seriation is a non-destructive method, all information remains in the seriated data. The specialized methods usually outperform seriation, e.g., clustering is usually more efficient to identify similar objects than seriation. Seriation also help to visually detect objects and variables with large amount of missing data and outliers. The usual seriations provide heatmaps with reasonable less striped feature, quite often arrangements around the diagonal or visually detectable clusters. Seriation can be performed not only on measured data, but on, e.g., model parameters, neuron intensities in artificial neural networks, connection data, as well. In these cases, seriation might help in the interpretation of the models and the operations. Liiv started to unify the taxonomy of the different methods [ 5 ]. Theoretically, seriation means the permutation of data stored in one-dimensional vectors up to k-dimensional arrays. There are modes and ways in seriation. A mode means an independent sequence that can be permutated. Way is the dimensionality of the object used in visual perception, during the calculation of a merit function, or during the prescribed operation. The number of modes and ways mostly coincide to the dimension of the data. In chemistry we often have two-dimensional data matrices with N rows connected to the objects and M variables denoting the columns. When we sequence both the objects and the variables in a classical data table, we perform two-mode–two-way seriation. When we sequence only the objects, we usually calculate a symmetric NxN distance matrix and the seriation is one-mode–two-way. If we seriate only the variable sequence, the MxM covariance matrix might be a reasonable choice and the seriation is one-mode–two-way. The number of the methods how to seriate is rather large. There is not any canonical way, the popular methods differ from field to field. Large number of applications can be found usually on the field of bioinformatics, where, e.g., the node deleting algorithm of biclustering is one of the most popular methods [ 15 ]. There are two groups of the methods. In a part of them, a mathematical merit (or loss) function is defined which depends on the sequencing of the modes. In the other cases, set of operational instructions are used. In the first case, the extremum of the merit or loss function can be found by any global optimization scheme, e.g., simulated annealing, genetic algorithm or other specialized solutions. These methods are mostly iterative and they use a stop criteria. The operational methods are usually repeated as long as the condition of the operation is holding. There might be some extra conditions to avoid infinite loops and it is worthwhile to start both methods from several sequences whereof many can be randomized ones. Another aspect of the two-mode seriations whether the two sequences are treated independent from each other, or the merit function/operation contain cross terms. For the independent case an example is the seriation of the objects according to the distance matrix and seriation of the variables according to the covariance matrix. Despite the independence of the two modes, by chance we might get clearly interpretable data, where relations between the two axes are easily readable, e.g., on heatmaps. For having dependent two-mode seriation, we need cross terms between the sequences or geometrical preferences of the matrices. A recipe is sometimes that we have several local function values and the sum or the spatially weighted sum of the local functions provide the merit of loss function. The local functions might be related simply to the increase, to the decrease or to the modality of the data within row or column wise and, e.g., the global loss function is the number of the violated case. For an overview of some of the methods and mathematical details we refer to the review of Liiv [ 5 ] and the study of Hahsler et al. [ 11 ]. For operational algorithms the representation of the problem on graphs can be useful, a part of the algorithm bases on the minimal number of crossings known as Turán’s brick factory problem in mathematics and history. In our previous studies [ 7 , 21 ] a local feature was calculated as the distance of two objects in a limited variable-vector space. We used three-variable spaces and the i-j element of the so called local distance matrix contained the average distance of the i -th object to its sequential neighbours in the local space formed by the j-1,j,j + 1 variables. The global merit function was a weighted sum of these local distances, where the weights were the spatial distance of the i-j matrix element from the diagonal of the data matrix. The algorithm provided that low local distances were sequenced around the diagonal of the data matrix. The ordering was according to one visual feature, it ordered similar objects close to each other and the corresponding variables around the corresponding diagonal parts were suggested to be responsible for the similarity. We experienced that only a part of chemical data is meaningful in the obtained block diagonal forms. For example, there might be a group of variables responsible for two or more clusters of objects what is not easy visually detect on a narrow diagonal-like arrangement. Our first idea was to improve our previous method by introduction of further adaptive lines with similar task as the diagonal had, but during the elaboration we realized that it is easier to think on a seriation forming patches. In this paper we show our new method where the local function is a local average similarity, and the global merit function is the sum of the products of the neighbouring local similarities. We found that this merit function forms patches of the neighbouring objects and variables. A patch means a local space, where the given objects are similar to each other. In our philosophy, the object-variable points outside the patches are not relevant for the similarity patterns. We show it on simple chemical data, as well as on model details of artificial neural networks (ANN). The latter is related to the interpretation [ 22 ] of ANN models, e.g., we were able to interpret the roles of the neurons connected to the variables and to the objects. The former means the seriation of the weights in the network and the latter was managed by seriation of activities caused the different objects on the different neurons. Our method can be easily extended to higher modes and ways seriations. We developed different three-mode-three-way methods. Since it is less easy to interpret three-dimensional data structures than two-dimensional ones, we usually used projections onto two-dimensional heatmaps, here. A part of our results is shown on co-plots elaborated by us, where the original data heatmap and local similarity contour plot are merged. It helps to easily find the responsible variables for the similarity of a cluster of variables. 2. Theory In 2011 we introduced a mathematical merit function for 1-mode-2-way and 2-mode-2-way seriation of matrices [ 7 ]. Two concepts were introduced there. The local distance matrix contained the average distance between the i -th object and its two neighbours in a local three-variable space, where the index of the middle variable assigned to j . The other quantity we called diagonal measure, and it represented the distribution of the elements of the local distance. It was the scalar sum of the local distances weighted with their positional distance from the diagonal of the matrix. In 2-mode-2-way seriation the diagonal measure was maximized to order similar objects close to each other and the corresponding variables around the diagonal were suggested to be responsible for the similarity. If distance matrix of the objects was 1-mode-2 way seriated, the diagonal measure was maximized, as well. If covariance matrix of variables was 1-mode-2-way seriated, the diagonal measure was minimized. In our new research we propose development of our idea both on the local quantity and the global measure. 2.1 Local similarity matrix Distances are unbounded positive numbers what may hinder the interpretation of the actual values. Algorithmically, it is more convenient to use bounded set of values. Similarity is a frequently used concept for that. There is a reciprocal relation between similarity, a value of one denotes perfect similarity of two objects (zero distances of the objects in the variable space) and zero similarity means maximal distance between the objects. There are different definitions of similarity, whereof we finally selected that similarity = 1-(distance/maximal distance) equation. We defined a local similarity matrix (S) similarly to the local distance matrix. s ij shows how the i -th object is similar to its neighbours in a local 3-variable space around variable j. For 2-mode-2way seriation it is calculated as: \({l}_{i,k,j}= \sqrt{\sum _{l=j-1}^{j+1}{\left(\frac{{a}_{kl}-{a}_{il}}{{diff}_{max,l}}\right)}^{2}}\) Eq. 1 \({s}_{ij}=\left(\sum _{k=i-1,i+1}^{}1- \frac{{l}_{i,k,j}}{{D}_{col,j}}\right)/{D}_{row,i}\) Eq. 2. ,where a il and a kl are the elements of the A matrix to be seriated, l ikj is their local distance in the variable space formed by the l = j-1, j and j + 1 variables. diff max,l is the difference between the largest and the smallest elements of the l -th column in A . It is used to scale the distance between [0,sqrt(3)], if the local variable space contains three variables. If the j -th variable is at the first or the last column of the matrix, the local space contains only [0,sqrt(2)]scaled distances. D col,j contains the corresponding upper bounds of the intervals for each variable. D row,I is usually two for the i -th object except the first and the last row, where it is one. These row or column dependent quantities ( diff maxl , D col,j , D row,i ) were introduced to be able to get theoretically s ij ϵ [0,1] values for all i-j positions including the non-bulk matrix elements. 2.2 The global patch function In the case of our previous global scalar (diagonal measure), the seriated matrix placed the variables responsible for object similarities around the corresponding part of the diagonal. It means, only the most important variables were emphasized, and, e.g., there was no possibility to select a variable to be important for several object clusters. In our new method we define a merit function, where forming of several patches is supported by maximising it in 2-mode-2way seriation. If we calculate the sum of the product of two neighbouring local similarity values (P), this quantity reflects the spatial distribution of large and small similarities. If random order of objects and variables is used, the local similarities are distributed randomly in the matrix. If we seriate the matrix to have larger sum of neighbouring products, a higher sum can be reached by clustering high and low local similarities separately. Furthermore, preferential rearrangement is also supported by maximising such a merit function which creates higher similarities by neighbouring similar objects. In the high similarity patches the objects are similar in the local variable space and both the objects and the variables can be identified. \(P=\sum _{k=1}^{nxm}\sum _{l}^{}{\left({s}_{k}{s}_{l}\right)}^{q} = \sum _{i=1}^{n-1}\sum _{j=1}^{m}{2\left({s}_{ij}{s}_{i+1,j}\right)}^{q}+\sum _{i=1}^{n}\sum _{j=1}^{m-1}{2\left({s}_{ij}{s}_{i,j+1}\right)}^{q}\) Eq. 3. , where k goes over all elements of the local similarity matrix, l denotes the given neighbours of k with one common index and an index differing with +/- 1. q is an arbitrary contrast. The two effects of maximising P - spatial ordering and creation of high similarities - can be justified separately. The simple rearrangement of any matrix by clustering large and small values provides large P : it is similar to a negative local entropy. We performed several test calculations supported this and there is also a thought experiment in the supporting material. The other effect is straightforward, that placing similar objects and variables close to each other increases the sum of the local similarities. The exponent q is an empirical contrast factor. At high q values the positioning of highest similarities close to each other is extremely preferential and it may cause compact and small clusters, while small q -s do not penalize so strictly the less large values, it may cause slightly larger patches. We used q = 2 and q = 3 in our calculation. Depending on the dataset, the visual results was sometimes better, for one of the q choices, but it did not seem to be a decisive parameter of the merit function. We note, that our patch function was obtained after several trials, where at first we focused on entropy or Gini-index like approximations. We found, that P defined as in Eq. 3 is a simple and feasible merit function. 2.3 Local similarities in higher dimensions The generalization of the local similarity matrix and the patch function to higher dimensions can be easily done, if we follow the idea that we are interested in the average similarity of an object to its sequential neighbours in a local three-variable space. The calculation of the possible cases, e.g., the dimension of the original data, the dimension of the local similarity matrix, the number of possible local variable vectors are detailed in the Results and Discussion section together with some examples. We show three possibilities for 3-dimensional local similarities, where the three axes are formed by one object and two variable vectors (OVV case, original data are 2D), by two object and one variable vectors (OOV-independent, the original data are two dimensional) and by another two object and one variable vectors case (OOV-dependent, the original data are three dimensional). 2.4 Missing data and noninformative zeros There are several datasets in chemistry, where part of the data is missing. The causes might be different, e.g., lack of general experimental methods for all objects, operational break down, or the given variable is not relevant for that object. In the case of cheminformatics it also common, that several extra variables are added to the database where most of the objects provides a zero value. An example is the presence of chemical groups, if close to all the molecules do not contain that functional group. The traditional method to overwrite the missing data with an average or random value might bias the seriation. Therefore, it would be feasible to avoid the replacement of missing data. Also, it is rather misleading, if the unnecessary and irrelevant zeros have crucial effect on the merit function of seriation. We solve the problem of missing data and unnecessary zeros by proposing a different calculation of the local similarities for these cases: \({s}_{ij}=\sum _{k=i-1,i+1}\sum _{l=j-1}^{j=j+1}\left(1-\left|\frac{{a}_{kl}-{a}_{il}}{{diff}_{max,l}}\right|\right)/6\) , Eq. 4 The inner sum is skipped for all data, where any of the data ( a kl or a il ) is non-existent. It can be used for missing data as well as for unnecessary zero values. Using Eq. 4 the local similarity cannot be one, if there are undetermined cases in the sum. Also, if s ij refers to a matrix position at edges or corners, the possibility for the local similarities to be 1 is excluded, there maximum value is 2/3, 1/2, or 1/3. This handling of the borders is different from Eq. 2. The patch function is calculated according to Eq. 3. There is only one difference, there might be a chance that a sij remains undetermined. In that case the undetermined s ij is skipped in Eq. 3. 3. Calculation Details 3.1 Codes The patch seriation was performed using a C code developed in our laboratory. The code reads the datasets, manages data pre-processing as optional normalization, scaling, changing zeros to undetermined values. The maximization of the patch function was obtained with Metropolis Monte Carlo algorithm, where the ordering with the largest P was stored as the best one. The acceptance ratios for the different type of trial changes were set to be around 0.05. Column and row permutations were performed independently. In the case of three-mode-three-way seriation it was performed independently for all modes. The number of the trials was 1–5 million. A calculation took a few minutes on a PC depending on the size of the dataset. During this calculation length, usually the best sequence was detected and stored at any time after the 20% of the calculation time. A few (2–5) seriations were performed for each dataset at q = 2 and q = 3 values, which of the results to be shown were selected visually. The elaboration and visualization of seriation results was done using R [ 12 ]. In the comparison to other methods, here we used the seriation package of Hahsler et al. [ 11 ]. For three-dimensional graphs we used the RGL package [ 23 ]. We developed an overlay plot, where the heatmap coded scaled data are shown together with contour plots of the local similarity. We think, these overlay plots are rather effective to identify object clusters and the variables causing the similarity. The neural network modelling was performed in Python using the scikit learn package [ 13 ]. The partly optimized hyperparameter sets for the models were selected from one of our previous studies where we used the same datasets [ 24 ]. All codes are deposited and freely downloadable at the homepage of the corresponding author after the publication of the article. 3.2 Datasets The tested datasets are mostly freely available ones related to QSAR, chemistry, material science, food science, cheminformatics and environmental chemistry. Several datasets of them are accessible in repositories [ 25 – 27 ]. Some details of the data and the performed type of seriations are collected in Table 1 . The first dataset (SIM[ 28 ]) is a semi-randomly simulated one, its structure is related to our initial idea, what kind of benefit we would like to get using patch seriation. There are 50 objects and 20 variables in this set ordered in 4 clusters and a random group for the objects. Members of the clusters have similar values at some selected variables, but their other data are random. Some of the selected variables are common also with other clusters. At first, we generated [0,1) random numbers for all data and thereafter the groups were recalculated by adding a given random number for that variable of the group biased with white noise. In the case of other datasets, if a dataset was published for modelling a response variable, we omitted it from the seriation and only the predictor variables were used in the seriation process. The RETSIM dataset [ 28 ] is a simulated one, as well. We defined three functional groups and created 4 compounds with random linear combination of the three groups. We set 6 mixtures of the 4 compounds. 6 chromatographic columns were set as well with differently randomized partial retention times for the functional groups. The retention times of the compounds were calculated with linear combination of the functional groups therein. Finally, we added uniform broadening for each compound with integrals related to the concentrations. In this way we had 36 chromatograms of the 6 mixtures on the 6 columns. Table 1 Datasets abbr. row × column description seriations ref. SIM 50×20 4 clusters with common variables for each OV [ 28 ] RETSIM (6x6)x100 and (6x6)x150 simulated retention times of mixtures on different columns (+’fingerprints’) OV; OOV dependent [ 28 ] POL_MONTH (26x12)x9 monthly air pollutant averages at 26 stations in 2017 OV; OOV dependent [ 29 ] POL_YEAR (12 + 14)x9 yearly air pollutant averages at 26 stations in 2017 (12 at Budapest, 14 at countryside) OV; OOV independent [ 29 ] FLASHP1 420x26 for ANN model (N = 4,6,8,10 hidden neurons), 80 objects in the test set flash point estimation of molecules using different QSAR parameters OV: 80x26 (test objects-variables); 26xN (variables neuron weights); Nx80 (neurons, object activities on the neurons) [ 25 , 30 ] DR8 600x28 for ANN model (N = 10–15 neurons), 114 objects in the test set different QSAR parameters originally used to estimate toxicity OV: 114x28 (test objects-variables); 28xN (variables-neuron weights); Nx114 (neurons-object activities on the neurons); OVV independent 114x(N + 28) (objects, neuron activities, original variables) [ 25 , 31 ] FLASHP2 632x(13 + 12), for ANN models (N = 10–15 neurons), 100 object in the test set 13 molecular and 12 general descriptors of molecules originally used for flash point estimation OV of 100x25, 100x13, 100x12; OVV 100x(13 + 12); OVV 114x(N + 28) (objects, neuron activities, original variables) [ 25 , 32 ] POLMET_DAY 56x(7 + 6), subset of original, two weeks from each season (56 days) Daily meterological and airpollutant data set in 2007 OV: 56x13, 56x7, 56x6; 3D OVV: 56x(7 + 6) [ 33 ] ESSOIL 10x(10 + 38) Essential oils in 10 species, 10 chemical and 38 bactericid/fungicide data OV: 10x48, 48x10; 3D OVV: 10x(10 + 38) [ 8 ] CERAMIC 88x17 ceramics with body and glaze data OV [ 27 , 34 ] GLASS 214x9 composition of glasses from different sources OV [ 26 , 27 , 35 ] WINE 178x13 wine analysis OV [ 26 , 36 ] TOXIC 112x8 toxic on 8 data OV [ 37 ] SAND 30x14 sand data radiation OV [ 38 ] MOLDESCRg 500x50 more subsets of the original molecular descriptors for enormous number of molecules to calculate different properties OV [ 25 , 39 ] COIN 257x10 composition of ancient coins from different era of Hungary OV [ 40 – 42 ] REAC 95x32 fuel combustion with reactions and reactants OV [ 28 , 43 ] 4. Results And Discussion 4.1 Simulated dataset for 2-mode-2-way seriation The dataset contained 50 objects and 20 variables. Each of the 4 clusters had 10 objects with similar set of 5–5 variables. These variables were distinct except for group C and D, here two variables were common but with different average values for the two set of objects. 10 objects and two variables had no cluster affiliation. Figure 1 shows the heatmaps of a randomly ordered matrix (used as start), the seriated data matrix, the corresponding local similarity matrix and the corresponding hidden cluster information at the data generation in the seriated order. If we plot only the data matrix as a heatmap, it is not easy to identify the clusters. Therefore, we use the overlay of a contour plot on the local similarity matrix in Fig. 1 b. In more than one third of our trials the variables were seriated perfectly and the number of clusters of the objects was equal or only slightly more than 4. We show a case in Fig. 1 b- 1 d, where the seriation worked perfectly both for variables and objects. We choose the actual contour levels of the local similarity matrix to help the assignment of the clusters in Fig. 1 b. The 10 random objects and the 2 random variables are sometimes between the clusters or sometimes they are neighbouring to each other, but the overlay contour plot does not identify them as a 5th cluster. The number of the object clusters is 4 in this example. We compared it to other seriation methods built in R [ 11 , 12 ]. The number of object clusters were between 14–33 for the other methods. It means, the clusters were split into 3–8 parts in average. We also calculated how many of the variables is found in a cluster for the object clusters. In our seriations the 5 variables were mostly clustered for all object groups correctly. In the case of the other methods, it was between 13–18. It means, there were only 2–7 cases, when two common variables of object clusters were placed to be neighbours in the variable sequence. The supplementary material contains further details on the comparison of the methods. We should emphasize, that our patch seriation worked efficiently both for objects and variables simultaneously and it explores the link between the two modes. This link is missing for most of the other methods. As it can be seen in the supplementary material, most of the methods use only distance matrices of the objects, where the simultaneous sequencing of the two axes is not possible. In our comparison, we calculated also variable sequencing of the other methods by calculating a ‘distance matrix’ of the variables, as well. In this prototype of data different groups of variables are responsible for the different clusters of the objects and the other variables are not important for the object clustering. It seems so, that our method totally outperforms all the other seriation methods. Even more, clustering methods, e.g., hierarchical clustering is not able to find this type of relationship within the objects due to the nonlocal distance calculations. 4.2 2-mode-2-way seriation of other datasets The most of the tested patch seriations concerned the simultaneous ordering of objects and variables in two dimensions. Here we show some examples with original and seriated heatmaps. Figure 2 . shows three datasets. The first is the POL_YEAR one [ 29 ] in Fig. 2 a-b. It contains the yearly averaged concentrations of 9 air pollutants at 12 places in Budapest and 14 places at countryside (mostly in cities) in 2017. There are differences in the measuring stations, not all of them were able to measure all the 9 components and there were several shutdowns at some places. Here, we calculated the local similarity matrix according to Eq. 4. The seriation clearly shows that the dominant variables for similarity are the different nitrogen-oxide and ozone concentrations. The stations with heavy traffic are ordered close to each other. The other stations are separated into two groups, where the nitrogen oxides are dominant pollutants and where ozone pollution is dominant. It is known, that there is a transition cycle of ozone and nitrogen-dioxide. The first and the last stations are somehow thrown out by the seriation. These are the places where the number of missing data was high. In Fig. 2 c-d we show a case, where the component of ceramics (celadons) are the variables [ 34 ]. Here it is known for the objects, what kind of celadon ceramics and what part of them are analysed (body or glaze). The seriation clearly shows, how the randomized data can be turned to form two groups, body and glaze, according to the low concentration of some oxides in the body part and high concentrations of them in the glaze. The seriation also found a subgroup of celadons, which were only imitated Longquan celadon in Jingdezhen civilian kilns in Ming Dynasty. The seriation was not able to differentiate the Longquan celadons of different dynasties. In Fig.re 2e-f we show a set of reactions and reactants used in the combustion modelling of gasoline [ 43 ]. It is not easy to determine an order of the reactions and reactants. Previously we used our diagonal seriation for this reaction set and we were able to arrange them according to a diagonal suggesting a hypothetical pathway. Of course, that simple order was related to a non-realistic graph structure, where it is well known that a proposed way need not coincidence to the real fluxes of the processes, especially the fluxes highly differ for different combustion conditions. In the case of patch seriation, we concentrated on the identification of reaction parts, e.g., reactions using the same components as reactants or products. The presence of a component in a reaction was denoted with 1 irrespectively the components role and stochiometry. The gasoline components, the final CO 2 and H 2 O components are coloured differently. It can be seen in Fig. 2 f, that there is a reasonable re-clustering of the reaction system, where around six clusters of reaction-components are there. Two of them is related to the final products CO 2 and H 2 O, another is related to the CH 4 and CH 3 components, one is formed by different small entities containing H and O, and another contains additionally carbons. The group on the bottom-middle is related to H and H 2 . Such kind of seriation might be interesting if one intends to build reaction mechanism in a modular way. In Fig. 3 we show two cases, where the seriation helps to find some general pattern. In the GLASS and COIN compositional datasets the exchange of species can be easily detected in the seriated data besides the visual simplification of the heatmaps. In the seriated GLASS data (Fig. 3 b) one can detect a negative mirror like difference in the exchange of ions with similar charges, e.g., K + - Na + ; Ca 2+ - Mg 2+ ; Al 3+ - Ba 3+ . Furthermore, it is also easy to detect the relation between Al 3+ and the refractive index. It is not easy to realize these features in the unseriated data according to the rather striped heatmap (Fig. 3 a). The method also seriated the glasses quite well according to their sources or use detailed in the original source [ 26 , 27 , 35 ]. Figure 3 c-d contain a dataset on Hungarian coins from the X-XIII. century. Here we used Eq. 4 and set the zero values to be skipped during the calculation of the local similarity matrix. The heatmap shows a similar exchange of species, like Cu-Ag exchange. It groups the coins where Sn and Sb took part in it, as well. The seriation was done with 0–1 scaled data, therefore, the traces of some metals had a large effect on the ordering. Due to the scaling, it is easy to identify the coins having the same metals from second importance. Our method clusters the objects (coins) according to the era and kings, but here we should add that traditional clustering and classification methods provided better results [ 40 , 41 ]. 4.3 3D seriation: 1-object–2-variables case If we have a two-dimensional data matrix, where the columns contain two separable set of variables, we might perform a seriation, where the order of the objects, the first set variables and the second set of variables can be separately sequenced. A three-dimensional local similarity matrix can be constructed, where an element shows the average similarity of the given object to its neighbours (axis one), but now in two local three-dimensional variable spaces (axis two or axis three). Using the first type of variables and the second type of variables independently, we calculate the two local similarities between the given object and one of its neighbours. The final s ijk local similarity contains the average of four similarities (over the two neighbours times the two local variable spaces). Such an independent set of variables can be, e.g., the chemical content and the biological activities of the essential oils (ESSOIL), the daily averages of air pollutants and the corresponding meteorological data (POLMET-DAY) or the complex QSAR descriptors and the simple enumeration of functional groups in the flash point modelling FLASHP2. We show our results on the latter one, where the QSAR descriptors for a given molecule are called molecular descriptors in the original paper and the enumeration of functional groups is called general descriptors. If we apply separately two-dimensional seriation for the two set of variables, the obtained heatmaps are clearly arranged (Fig. 4 a-b). Here we calculated the local similarities with skipping the zero data to avoid the clustering of molecules due to the lack of functional groups in the set. If we seriated the total data matrix in two-dimensions, many of the clear patches disappeared. The continuous molecular variables were dominant during the seriation, most of the general descriptors was not clearly seriated (Fig. 4 d). If we performed the seriation using a three-dimensional local similarity matrix, the two parts of the data in the original two-dimensional dataset provided clear patches for both set of variables (projected back to two dimensions: Fig. 4 e). The advantage of the three-dimensional seriation over the two independent two-dimensional seriations is the common target function during the sequencing of the three axes. The three-dimensional local similarity array can be directly visualized (see later Fig. 5 b) or two-dimensional projections can be calculated. Figure 4 f is the projection, where the highest local similarities over the objects are collected for a given set one – set two variable pair. It shows that the highest similarities are obtained for which local variable pairs of set one and set two ones. We emphasize here again, that our mathematical target function connects all the modes of the seriation in contrary to the usual biclustering schemes. 4.4 3D seriation: 2-objects-1-variable case For the demonstration of the case, where a two-dimensional data matrix contains two different sets of objects we selected the air pollution data of stations in Budapest and at countryside (POL-YEAR). The two sets of objects might be seriated independently. Here the first axis of the local similarity matrix contains the stations at Budapest, the second one is the stations at countryside and the third axis shows the yearly averaged pollutant concentrations. The s ijk local similarity contains the average of 4 similarities: the similarity of the i-(i-1) and i-(i + 1) object pairs of the first axis (stations in Budapest) and the j-(j-1) and j-(j + 1) object pairs of the second axis (stations at countryside). The local variable space for this element is spanned by the k-1, k, k + 1 variables. The results of the seriation can be shown in the original two-dimensional data (Fig. 5 a), where both set of stations are rather homogeneously sequenced separately (cf. to Fig. 2 b. of 2D seriation, where the cities might be mixed). Using interactive three-dimensional graphics, one might have a look on the local similarity array. We show an example in Fig. 5 b, but it is rather uninformative without the possibility of rotating the graph. The different projections or enumerations on different subspaces of the three axes might be informative, e.g., which pollutant causes locations in Budapest and in countryside to be similar. It is clear from the graphs, that mostly the NO-NO x -NO 2 , and sometimes the SO 2 and PM10 data cause the similarities. Table 2 shows this projection where highly similar locations are ordered into the middle of the local similarity array. The corresponding alphabetical code enumerates all the local variables involved in at a given high similarity, e.g., a similarity over 0.9 at the O 3 -NO 2 -NO x position means that the two neighbouring variables (SO 2 and NO) are also involved therein. Table 2 Local similarities over 0.9 between stations in Budapest and in countryside. The local variables causing the similarities: A: NO 2 ,NO X ,NO B: O 3 ,NO 2 ,NO X ,NO C: NO 2 ,NO X ,NO,PM10 D: NO X ,NO,PM10 E: SO 2 ,O 3 ,NO 2 ,NO X ,NO F: O 3 ,NO 2 ,NO X ,NO,PM10 G: SO 2 ,O 3 ,NO 2 ,NO X ,NO,PM10 H: SO 2 ,O 3 ,NO 2 ,NO X I: O 3 ,NO 2 ,NO X The stations belonging to the columns (Budapest) and the countryside (rows) are listed in the supplementary material. 1 2 3 4 5 6 7 8 9 10 11 12 1 2 A A A A 3 A A B B C D 4 A A B B A 5 A A B E A 6 F B E E F D 7 F E E E G D D 8 A A E E D D 9 D H 10 B B E E I 11 A A A A 12 13 14 4.5 3D seriation: 2-objects-1-variables - dependent case In the case of two dependent object vectors - one variable vector, the data are originally three dimensional, but we usually have an unfolded two-dimensional data matrix. In the case of the RETSIM data unfolded to two dimensions, the first six rows correspond to the retention times of the 6 mixtures on the first chromatographic column, the next six rows for the second column, etc… The variables (columns) were the retention intensities for 1-100 arbitrary time intervals. The three axes of the local similarity matrix could be the chromatographic columns, the mixtures and the retention times. The s ijk local similarity contains the average of 4 similarities: the similarity of the j-th mixture data for the i-(i-1) and i-(i + 1) chromatographic column pairs and the similarity of the i-th chromatographic column data for the j-(j-1) and j-(j + 1) mixture pairs. The local variable space for this element is spanned by the k-1, k, k + 1 retention time intensities. If we seriate the three axes independently (only the order of the columns and separately the order of the mixtures), a rather limited information can be obtained on the new order, e.g., some link between the mixtures and columns. Furthermore, for such a seriation we need to know a priori the mixture and column labels for each spectrum. It would be more interesting, if we seriate in an unrestricted way, where we fill the 6 times 6 object space with maximal freedom without taking care on the order of the columns and the mixtures. In this case, seriation has a strong explanatory statistical feature, if we will be able, e.g., to cluster the common measurements for a column or a given mixture. The simple 2-mode-2-way seriation of the retention spectra was able to classify the spectra for each column perfectly (Fig. 6 b). On contrary, there were no patterns how the mixtures were ordered. If we use the three-dimensional spectra with supposing a 6x6x100 local similarity matrix, we still obtained a good order for the columns (Fig. 6 c), but the interpretation of the result is not easy. In the three-dimensional seriation an object has two-two neighbours in two variable subspaces with distances at three local retention times. There is a chance distribution that which type of objects are neighbouring in which spaces, e.g., there is no driving force that the column like neighbouring is according to the first axis or the second one. Even more, it can be different at the different grids of the local similarity matrix. For example, the three-dimensional seriation provided a good arrangement for the columns, around 70% of the 4 neighbours were measurements on the same columns. Spatially, it occurred along two axes and the unfolded data were partitioned into 3–3 long sequences for each column. It is not easy to see it on the unfolded seriated data. We found that it is not straightforward how to seriate the mixtures using the retention times. Therefore, we added 50 more variables to the data, where the largest 50 data of a row was normalized with the largest in the same row and were sorted in a descending order. These 50 new variables work as mixture specific fingerprints. We performed 2-mode-2-way and 3-mode-3-way seriations using this extended variable set. The results are shown in tabular form over Fig. 6 d-e, since the three-dimensional seriations cannot be unfolded by a visually striking way (c.f. the unambiguity of the assignment of the planes of the ordering and the column-mixture features.) Some features of the seriation are shown in Table 3 . It can be seen that the use of three-dimensional local similarity matrix helped slightly to get better classification on the mixtures, as well. 100 (84 + 16) or 106(84 + 22) correct neighbours were sorted from the altogether 120 possible neighbouring positions. Table 3 Performance of seriation on the RETSIM dataset Seriation type Maximal theoretical number of correct column neighbours + correct mixture neighbours (2D), or of all neighbours (3D) Realization in an example for columns Realization in an example for mixtures 36x100 size, 2D 60 + 60 60 (perfect classification) 0 36x150 size, 2D 60 + 60 60 (perfect classification) 0 6x6x100 size, 3D 120 84 (3–3 groups in 2D unfolded map) 16 6x6x150 size, 3D 120 84 (3–3 groups in 2D unfolded maps) 22 (repetition of mixtures in +-6 shifts) The other example is the air pollutant dataset containing 9 pollutants at 26 stations. The data are the monthly averages in year 2017 (POL-MONTH). The theoretical three axes are the stations, month and the pollutants. Figure 7 a is a randomized data matrix where both rows (stations in a given month) and columns (pollutants) are randomized. If we perform a 2-mode-2-way patch seriation, the heatmap became simpler, e.g., the NO 2 , NO and NO x variables were seriated near to each other (Fig. 7 b). Here we used that zero and missing values were not used in the similarity calculations (Eq. 4). The stations and pollutants with a lot of missing values are out-seriated to the edges of the heatmap. One can see, as in the case of the monthly averages, that the high nitrogen-oxide and ozone data provided a good basis for similarity. Figure 7 c shows the original data, where an arbitrary alphabetical order was used for the stations while the months are in calendar order. If we perform the 3-mode-3-way patch seriation, we obtained an ordered map with regular stripes. The neighbour analysis showed, that 48–59% of the neighbouring objects in the local similarity array belong to the same season, while this is only 39–47% for the 2-mode-2-way seriation. Around 30% of the four neighbours in the similarity array are the same in the station and/or in the month. We note that we do not want to get a perfect classification for these data, because it is not obligatory that objects of different classes (location, month or season) could not be closer to each other than objects from the same classes. The P (Eq. 3) of Fig. 7 c (perfect classification) is around the at the middle of the random and the best P -s. Our method is data driven and it helps to override traditional classification, where the data do not support to clearly perform classification. 4.6 Seriation of neural network model data Artificial neural network is one of the most popular methods to solve classification and regression tasks. The simplest conventional structure contains an input, a hidden and an output layer, where the input and the hidden ones and the hidden and the output ones are connected using sets of weights. In the simplest case, the hidden layer neurons contain activation functions, and the output layer ones only sum their weighted input. In the case of classification and regression, supervised method is used, where the weights are optimised to get a correct output for a training set. In an optimal situation, there is an independent test set to validate the model. There are two basic trends in the evaluation of models. In the novel applications of data science we concentrate on the output performance without restricting the complexity of the ANN models. In the traditional case we would like to have limited complexity of the models with some possibility to interpret the model itself. The order of the neurons in the input, hidden and output layer are usually totally arbitrary. Therefore, it is an open question for seriation, especially, if we would like to interpret and understand a given model. Here we focus on some simple cases. The first one is the visualisation of the weights between the input and the hidden layer. We might choose the hidden layer neurons as objects and the weights are the columns assigned to the input channels. The opposite assignment is also meaningful, where the input channels (input variables) are the objects and the number of the variables is equal to the number of the neurons in the hidden layer. The same data matrix can be used, in one case the original matrix is seriated, in the other case its transpose is the input. In Fig. 8 . we show the seriated results for the dataset FLASHP1. The original data was intended to estimate the flash point of different molecular systems. The predictor matrix contains information on the presence of different functional groups. We built several ANN models using several hyperparameters settings. Here we show 9 ANN models with three different activation functions and 4, 6 and 8 neurons in the hidden layer. The same training and test set was used for each model and we selected models with both R 2 and Q 2 F2 ( R 2 test ) more than 0.9. The graph shows the case, where the neurons were the objects and the local variable spaces were formed by the weights assigned to the input channels. According to the calculation of the patch function (using Eq. 1), we scaled here the variables. It means, in the presence of both negative and positive weights blue colour might denote a large negative weight and red colour denotes large positive weight. Of course, the scaling might bias the interpretations, but in this feasibility study we do not intend to go really into the details of any ANN model. One can see that some of the seriated graphs (models with 4 neurons and models using logistic activation function) are clearly arranged providing the possibility of interpreting the operation of the model. In the case of this dataset, logistic activation seems to be the most interpretable group of models. We checked several high weight values at one-one neurons, and we assigned them as, e.g., F, O, or N containing functional groups. It means, these neurons are the responsible ones for different chemical parts as in ref. [ 44 – 45 ]. This bunch of seriated graphs can be used to select models which are better interpretable. Our other example is the seriation of the objects (molecules) and their corresponding activity on the neurons as variables. One aspect of neural networks, that the original variables are mapped to the neurons of the hidden layer. This can be used as a dimensional reduction. Depending on the dimensionality of the original data and the number of the neurons, several features of the variables space might remain on the low-dimensional maps. A short investigation of it is shown in the supplementary material for hierarchical clustering. The object activities were calculated as the dot product of the variable vector of a molecule and the weight vector of a neuron in the hidden layer. Figure 9 . shows a case, where 80 test objects are mapped on the neurons and molecules - original variables are shown, as well. One can see in Fig. 9 a-b, that the patch seriation orders the molecules according to their activity on the neurons. This graph might be used to visually detect group of objects and details of the model, e.g., activity, inactivity or redundancy of the models. This object activity -neuron seriation graph resembles somehow to unsupervised maps, e.g., a Kohonen map. The seriation in the original variable space is also successful, but here the variable space is 26 dimensional, while the neuron activity space is only 4 dimensional. 5. Conclusions We developed a new seriation method where our previous idea of using a global merit function based on a local quantity was improved. We defined a local similarity matrix containing the average similarity of neighbouring objects in a local 3-dimensional variable space. These local similarities were put into a global merit function, where the permutations of the object and variable vectors were directed to have both increased local similarities and forming patches of the large similarities. The basic idea behind our seriation method is that there are datasets, where different parts of the variables are responsible for the different clusters of the objects. If a set of variables is not concerned in a cluster, they can be easily identified by being outside of the patches. In our method, an overlay contour plot of the local similarity values can be drawn onto the heatmaps of the original data to identify the clusters of the objects and the variables causing the clustering. Both the local similarity matrix and the global patch function can generalize into more than two dimensions. We showed some examples of different three-dimensional cases, where the data were arranged according to two variable and one object axes or to one variable and two object axes. Furthermore, the local similarity and the patch function can be generalized for data with missing values or cases, where zero values need not be accounted as responsible ones for clustering. We show two simulated datasets (see Fig. 1 and Fig. 6 b), where our patch method is especially effective to discover object clusters and the corresponding variables. Here the traditional seriation methods with non-local distances are mostly in trouble, the ad hoc values of the “non-important’ variables hinder the formation of the clusters. In the case of several public datasets, we found always clearly arranged heatmaps compared to the criss-crossed chequered random ones. Depending on the datasets (material science, compositional, air pollution, reaction kinetic data) clusters of objects and/or variables were always detectable in the seriated heatmaps. In the case of sparse matrices, the patch seriation glue together the non-zero variables. In the case of three-dimensional seriation, the interpretation is less straightforward. One needs advanced three-dimensional graphical software or feasible two-dimensional maps to enhance visual perception. If the result is unfolded into two dimensions, the seriated data show periodic changes according to the dimensions merged visually into one axis. In our examples we show the details of three-dimensional air pollution data and retention of different mixtures on different columns. We show some examples, how seriation helps to interpret neural network data. For example, we seriated the variable – hidden neuron weight matrices of different models and there is a striking difference depending on the activation function and the number of the neurons. For example, logistic activation function provided a more interpretable model than the other functions for a flash point dataset, especially at low number of hidden neurons. Also, seriation is a feasible method to detect the neurons responsible for a cluster of objects and to detect inactive parts of the models. We think, that seriation is a powerful non-destructive data evaluation or pre-evaluation method. Our special method forms patches of object and clusters. It is effective, if the non-important variables for a given cluster mask the identification possibility according to their variability. Seriation does not replace the different pattern recognition methods, but it at least helps to detect which methods and task might be successful on a dataset. Declarations Ethics approval and consent to participate: not applicable Consent for publication: We give our consent for the publication of identifiable details, which can include details within the text and figures to be published in the above Journal and Article. Availability of data and materials: The datasets supporting the conclusions of this article are available in the different repositories referenced one by one in Table 1. The C source code is available at the homepage of the corresponding author [46]. Competing interest: The authors declare that they have no competing interests. Funding: The investigation was partly supported for GT by grant NKFI K-128136. Authors’ contributions: All parts of the investigation have been performed with equal load of the authors except the C language seriation software coded by G. Tóth. Acknowledgement: GT thanks the fruitful discussion with prof. György Turán. The authors acknowledge the datasets for prof. Imre Salma and dr. Anita Rácz. References Petrie WM (1899) Flinders Sequences in Prehistoric Remains. The Journal of the Anthropological Institute of Great Britain and Ireland 29: 295-301. Bertin J (1981) Graphics and graphic information processing . Walter de Gruyter, Berlin, Boston. doi:10.1515/9783110854688 Brower JC, Kile KM (1988) Seriation of an original data matrix as applied to palaeoecology. Lethaia, 21:79-93. doi:10.1111/j.1502-3931.1988.tb01756.x Arabie, P, Hubert, LJ (1996). An overview of combinatorial data analysis. In Arabie P, Hubert LJ, De Soete G (eds.), Clustering and classification. World Scientific, River Edge pp. 5-63. Liiv I (2010) Seriation and Matrix Reordering Methods: An Historical Overview. Stat. Anal. Data Min., 3:70-91. doi:10.1002/sam.10071. Van Gyseghem E, Dejaegher B, Put R, Forlay-Frick P, Elkihel A, Daszykowski M, Héberger K, Massart DL, Heyden YV (2006) Evaluation of chemometric techniques to select orthogonal chromatographic systems. J. Pharm. Biomed. Anal. 41(1): 141-151. doi: 10.1016/j.jpba.2005.11.007. Tóth G, Szepesváry P (2010) A diagonal measure and a local distance matrix to display relations between objects and variables P. J. Chemometr . 24 : 14-21. doi: 10.1002/cem.1267. Sekulića TD, Božinb B, Smolińskic A (2016) Chemometric study of biological activities of 10 aromatic Lamiaceae species’ essential oils. J. Chemometr. 30 : 188–196. doi: 10.1002/cem.2786. Pigler C, Fogarassy-Vathy Á, Abonyi J (2016) Scalable co-clustering using a crossing minimization ‒ application to production flow analysis. Act. Polytech. Hung. 13: 209-228. doi:10.12700/APH.13.2.2016.2.12. Hammer Ø, Harper D, Ryan P (2001). PAST: Paleontological Statistics Software Package for Education and Data Analysis. Palaeontologia Electronica. 4:1-9. Hahsler M, Hornik K, Buchta C (2008) Getting Things in Order: An Introduction to the R Package seriation. J. Stat. Soft. 25(3): 1-34. doi: 10.18637/jss.v025.i03 R Core Team (2013). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. http://www.R-project.org/ Accessed March 21, 2023 Pedregosa F (2011) Scikit-learn: Machine Learning in Python, J. Mach. Learn. Res. 12:2825–2830. Hartigan JA (1972) Direct Clustering of a Data Matrix. J. Am. Stat. Assoc. 67: 123–129. Cheng Y, Church GM. (2000) Biclustering of expression data, Proceedings. International Conference on Intelligent Systems for Molecular Biology 8:93-103. Stumpfe D, Bajorath J (2011) Similarity searching. WIREs Comput Mol Sci, 1: 260-282. doi:10.1002/wcms.23 Rosenberg MS (2009) Sequence Alignment: Methods, Models, Concepts, and Strategies. University of California Press: Berkeley, CA. Leese MN, Hughes MJ, Stopford J (1989) The Chemical Composition of Tiles from Bordesley: a Case Study in Data Treatment, in: Rahtz S (ed.), Computer Applications and Quantitative Methods in Archaeology 1989. CAA89 (BAR International Series 548). B.A.R., Oxford, pp. 241-249. Bartel HG (1990) Seriation to describe some aspects of generalized evolution and its application in chemical informatics. Systems Analysis Modelling Simulation, 7:557-565. Forina M, Lanteri S, Casale M, Cerrato Oliveros M (2007). A new algorithm for seriation and its use in similarity dendrograms. Chemometr. Intell. Lab. Syst. 87:262-274. doi:10.1016/j.chemolab.2007.03.004. Tóth G and Amariamir S (2018) Seriation, the method out of a chemist's mind. Journal of Chemometrics 32(3-4):e2995. doi:10.1002/cem.2995. Molnar C (2022) Interpretable Machine Learning. A Guide for Making Black Box Models Explainable, 2 nd ed. Munich, Germany. Available online: https://christophm.github.io/interpretable-ml-book/ (accessed on 7 June 2022). RGL package https://CRAN.R-project.org/package=rgl last accessed 26th March, 2023 Király P, Kiss R, Kovács D, Ballaj A, Tóth G (2022) The Relevance of Goodness-of-fit, Robustness and Prediction Validation Categories of OECD-QSAR Principles with Respect to Sample Size and Model Type, Mol. Inf. 41:2200072 (15 pages). doi: 10.1002/minf.202200072 Ruusmann V, Sild S, Maran U (2015) QSAR DataBank repository: open and linked qualitative and quantitative structure–activity relationship models. J. Cheminf. 7:32. doi:10.1186/s13321-015-0082-6, http://www.qsardb.org. Kaggle Inc. http://kaggle.com Accessed 2018 Nov.–2023 Apr. D. Dua, C. Graff, UCI Machine Learning Repository, Available at http://archive.ics.uci.edu/ml. Irvine, CA: University of California, School of Information and Computer Science, 2019. Tóth G (2023) Benchmark datasets for seriation Mendeley Data, V1, doi - in progress Hungarian Air Quality Network, http://www.levegominoseg.hu (Accessed at June 2017) later it has been transported to http://legszennyezettseg.met.hu/ Tetteh J, Suzuki T, Metcalfe E, Howells S (1999) Quantitative Structure-Property Relationships for the Estimation of Boiling Point and Flash Point Using a Radial Basis Function Neural Network. J. Chem. Inf. Model 39: 491–507. Drgan V, Zuperl S, Vracko M, Como, F, Novic M (2016) Robust modelling of acute toxicity towards fathead minnow (Pimephales promelas) using counter-propagation artificial neural networks and genetic algorithm. SAR QSAR Environ. Res. 27, 501–519. doi:10.1080/1062936X.2016.1196388. Saldana DA, Starck L, Mougin P, Rousseau B, Pidol L, Jeuland N, Creton B (2011) Flash Point and Cetane Number Predictions for Fuel Compounds Using Quantitative Structure Property Relationship (QSPR) Methods. Energy Fuels 2011, 25, 3900–3908. doi: 10.15152/QDB.123 Salma I (2023) Daily air pollution and meteorological data Budapest, 2007. Mendeley Data, V1, doi: 10.17632/2mmwv3j4ms.1 Ziyang He, Maolin Zhang, Haozhe Zhang (2016) Data-driven research on chemical features of Jingdezhen and Longquan celadon by energy dispersive X-ray fluorescence. Ceramics International, 42:5123-5129. doi:10.1016/j.ceramint.2015.12.030. German B (1987) Glass Identification dataset, Central Research Establishment, Home Office Forensic Science Service, Aldermaston, Reading, Berkshire RG7 4PN Wine recognition dataset, Kaggle Inc. https://www.kaggle.com/brynja/wineuci Accessed 2017 March-2023 Apr. Arthur DE, Uzairu A, Mamza P, Stephen AE, Gideon Shallangwa GA. Quantitative structure‐activity and toxicity relationship study of CCRF‐CEM and RPMI 8402 cell line apoptosis with some anticancer compounds. Chem. Data Coll. 2017;7‐8:8‐50. doi:10.1016/j.cdc.2016.12.002. Hariprasath R, Jose MT, Vijayalakshmi I, Rajesh A (2016) Determination of natural radioactivity and radiological hazards of sediment sands in Tiruchirappalli district, Tamil Nadu, India. Chem. Data Coll. 2:1‐9. doi: 10.1016/j.cdc.2016.03.001. Lang A. Data for: Abraham descriptor A. QsarDB repository, QDB.100. 2012. http://dx.doi.org/10.15152/QDB.100 Rácz A., Héberger K., Rajkó R, Elek J (2013) Classification of Hungarian medieval silver coins using X‐ray fluorescent spectroscopy and multivariate data analysis. Heritage Science. 1(1):2 doi:10.1186/2050-7445-1-2. Christie Olav HJ, Rácz A, Elek J, Héberger K (2014) Classification and unscrambling a class‐inside‐class situation by object target rotation: Hungarian silver coins of the Árpád Dynasty, 997‐1301 AD. J. Chemometr. 28:287‐292. doi: 10.1002/cem.2601 Rácz A, Héberger K, Rajko R, Elek J (2023) Composition data of 257 Hungarian medieval silver coins, Mendeley Data, V1, doi: 10.17632/kbjrfkvcs3.1 Juhász G (2015) Reduction of a biodiesel combustion reaction mechanism. BSc thesis Budapest: Institute of Chemistry, Department of Physical Chemistry, Eötvös Loránd University, Budapest Preuer K, Klambauer G, Rippmann F, Hochreiter S, Unterthiner T (2019). Interpretable Deep Learning in Drug Discovery. In: Samek, W., Montavon, G., Vedaldi, A., Hansen, L., Müller, KR. (eds) Explainable AI: Interpreting, Explaining and Visualizing Deep Learning. Lecture Notes in Computer Science(), vol 11700. Springer, Cham. doi:10.1007/978-3-030-28954-6_18 Jiménez-Luna J, Grisoni F, Schneider G (2020) Drug discovery with explainable artificial intelligence. Nat Mach Intell 2:573–584. doi:10.1038/s42256-020-00236-4 Tóth G, Homepage of Gergely Tóth, http://tothgergely.web.elte.hu (Last accessed 2023 Apr) Additional Declarations No competing interests reported. Supplementary Files tocgraph.png patchsupplementarymaterial.docx Cite Share Download PDF Status: Published Journal Publication published 09 Sep, 2023 Read the published version in Journal of Cheminformatics → Version 1 posted Editorial decision: Major revision 10 Apr, 2023 Submission checks completed at journal 07 Apr, 2023 Editor assigned by journal 07 Apr, 2023 First submitted to journal 05 Apr, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2780120","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":189928684,"identity":"bf7e06c7-ff43-4094-8307-5474f96a05ed","order_by":0,"name":"Rita Lasfar","email":"","orcid":"","institution":"Eötvös Loránd University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Rita","middleName":"","lastName":"Lasfar","suffix":""},{"id":189928686,"identity":"cb4ac4a4-d792-4a97-9528-c3cbd0b34753","order_by":1,"name":"Gergely Tóth","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABOElEQVRIie3RMUvDQBTA8RcOnsvTrieV5CtcCAREwdGv0VCISxRdJJMGCukS6RonP4PfIBKoS+rc0RAQhw63maGDl5R2SdTV4f4kx4Xw4x4cgE73b7sGMKLtB2cZgHoO1J79SMSOCCCOXpQpgvgH2e2Ik/0rsc4fqk8pcmDTxUtZw9o8GyZflSzARGuSsZu4Q+zi1T1OFTGSq7FDIBw6WjxH2RIcRByxxx6S+uhQQ6LAHarBvIRfKiLBi5EE2+8hTx/orBsyW7mHdUuCsiH3MQ5kH7E4sgoakgYupw0xmsFGiAR9RJCPRiIuyEhXjppQvUvfTouC2zH6Iqe37inTOZN1eGLas8Au69A099Lxuwznp9aA5WVFt91TMkAOMFF3sbmWbbxdsw5Qp0TAJMAdWD0/dTqdTtf2DXtqZSBY545FAAAAAElFTkSuQmCC","orcid":"","institution":"Eötvös Loránd University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Gergely","middleName":"","lastName":"Tóth","suffix":""}],"badges":[],"createdAt":"2023-04-05 09:14:32","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2780120/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2780120/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s13321-023-00757-1","type":"published","date":"2023-09-09T15:01:21+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":35545021,"identity":"4746caac-e6d0-43a8-96e3-7ce484942f4f","added_by":"auto","created_at":"2023-04-10 17:48:51","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1809111,"visible":true,"origin":"","legend":"\u003cp\u003e2-mode-2-way seriation of simulated data. The data are shown on scaled heatmaps. a) random order (typical start) b) an example of patch seriated data with \u003cem\u003eq\u003c/em\u003e=3. The overlay contour plot shows the local similarity levels around 0.85. c) the corresponding local similarity matrix sorted as figure b d) intended common values of the clusters during the dataset simulation sorted as figure b\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/5b93f91466a3180e98b2fe08.png"},{"id":35544052,"identity":"e987c444-e997-4aea-9291-9775b6de59b6","added_by":"auto","created_at":"2023-04-10 17:40:51","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":851559,"visible":true,"origin":"","legend":"\u003cp\u003e2-mode-2-way seriation. a,c,e: random data order b,d,f: seriated ones. The data are shown on scaled heatmaps. a-b) yearly air pollution data at 26 stations in 2017 (POL_YEAR dataset), the seriated order of the variables (columns): PM2.5, PM10, O\u003csub\u003e3\u003c/sub\u003e, NO\u003csub\u003e2\u003c/sub\u003e, NO\u003csub\u003eX\u003c/sub\u003e, NO, SO\u003csub\u003e2\u003c/sub\u003e, CO, BENZOL c-d) components of celadon ceramics (CERAMIC dataset) e-f) reaction and species in a gasoline combustion model (REAC dataset), blue denotes the three reactants, red ones are H\u003csub\u003e2\u003c/sub\u003eO and CO\u003csub\u003e2\u003c/sub\u003e.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/d979c07edcb7a2a3e404b41b.png"},{"id":35544051,"identity":"8576c23b-74c9-4a96-8b4e-95c252891737","added_by":"auto","created_at":"2023-04-10 17:40:50","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":384868,"visible":true,"origin":"","legend":"\u003cp\u003e2-mode-2-way seriation. The data are shown on scaled heatmaps. a-b) glass compositions (GLASS dataset) a-random, b-seriated c-d) coin compositions (COIN) c-random, d-seriated\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/f30caa063a67f81f650c90b9.png"},{"id":35544054,"identity":"cadc8e02-193b-41f9-af74-a454f40cea73","added_by":"auto","created_at":"2023-04-10 17:40:51","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1597999,"visible":true,"origin":"","legend":"\u003cp\u003e2-mode-2-way seriation and 3-mode-3-way seriation of the FLASHP2 dataset. a) 2-mode-2- way seriation of the objects with molecular descriptors b) 2-mode-2-way seriation of the objects with general descriptors c) Random start of the whole data matrix used in 2-mode-2-way seriation d) 2-mode-2-way seriated whole data matrix e) 3-mode-3-way seriated whole data matrix unfolded to 2D f) projection of the highest three-dimensional local similarity matrix values on the two variable sets subspace\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/534cffe0fd47ca7cf17d952e.png"},{"id":35546650,"identity":"6f4c45f9-cec2-4268-ac4f-924b1b562e84","added_by":"auto","created_at":"2023-04-10 18:04:51","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":748425,"visible":true,"origin":"","legend":"\u003cp\u003e3-mode-3-way seriation of the POL-YEAR dataset. Right: seriated data matrix (stations in Budapest and at countryside form the two independent object sets. Left: three-dimensional view of the local similarity array. B1-B12: stations at Budapest in alphabetical order, C1-C14: stations at countryside in alphabetical order, PM2 = PM2.5, BE = benzene\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/71390cea3b1c5c1bf5287676.png"},{"id":35545820,"identity":"de717885-972a-416c-b057-6b947326ed5b","added_by":"auto","created_at":"2023-04-10 17:56:51","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":1719587,"visible":true,"origin":"","legend":"\u003cp\u003eseriation of RETSIM data a) example of the simulated spectra b) 2-mode-2way seriation using retention intensities c) 3-mode-3-way seriation using retention intensities d) 2-mode-2-way seriation using retention intensities and fingerprints e) 3-mode-3way simulation using retention intensities and fingerprints\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/e59c3104a035332cf4f19bde.png"},{"id":35545025,"identity":"0e181453-e640-48c6-98e7-95568e7d6893","added_by":"auto","created_at":"2023-04-10 17:48:51","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":500609,"visible":true,"origin":"","legend":"\u003cp\u003eSeriation of monthly air pollutant averages at 26 stations in 2017 POL-MONTH. a-b) 2-mode-2-way seriation a - random b - seriated. c-d) 3-mode-3-way seriation c - ordered by hand, 12 months/station d - seriated.\u003c/p\u003e","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/3a36d4f3008df7e36e30e080.png"},{"id":35544058,"identity":"171f2a20-a0f8-468e-aef7-76688f76fcc0","added_by":"auto","created_at":"2023-04-10 17:40:51","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":592133,"visible":true,"origin":"","legend":"\u003cp\u003eSeriated details of neural network models on the FLASHP1 dataset. The objects are the neurons and the variables are the scaled weights of the original input variables. Three activation function are used (tangent hyperbolic, relu and logistic) with 4, 6 or 8 neurons in the hidden layer.\u003c/p\u003e","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/7e070d0d623160ee914fd22e.png"},{"id":35545822,"identity":"b21c8f15-96b4-4b97-b650-227c2b96a42d","added_by":"auto","created_at":"2023-04-10 17:56:51","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":160174,"visible":true,"origin":"","legend":"\u003cp\u003eSeriation of the FLASHP1 data (test set). a-b). Object – object activities on the hidden layer neurons (model: logistic function with 4 neurons) a- original b-seriated c-d) Object – original variable data c-random order d-seriated\u003c/p\u003e","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/6cd2fcfd6c33a7b9e149775d.png"},{"id":42947196,"identity":"6c0f6325-1bfc-49c9-a4dc-41679e4e8848","added_by":"auto","created_at":"2023-09-11 15:07:16","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5072767,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/96bd66d9-4510-49b9-8de7-ecbb6eaac15e.pdf"},{"id":35544049,"identity":"a5b5f80a-1e57-45e2-8d54-24821cc07330","added_by":"auto","created_at":"2023-04-10 17:40:50","extension":"png","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":140468,"visible":true,"origin":"","legend":"","description":"","filename":"tocgraph.png","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/30f74345cb27e65ac0c8f107.png"},{"id":35545023,"identity":"9943dcd1-353d-474f-8f18-d3ee820b7e72","added_by":"auto","created_at":"2023-04-10 17:48:51","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":575396,"visible":true,"origin":"","legend":"","description":"","filename":"patchsupplementarymaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-2780120/v1/1e37e98ba248ab1a3a50dfb9.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Patch seriation to visualize data and model parameters","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eSeriation? Most of scientist involved in it without knowing the term. If one knows its practical definition, namely, how to do row and/or column permutations to enhance visual perception of a table or heatmap, it is clear for scientist that they have faced with the problem. Its first application goes back to the XIXth century [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e], when it was an explanatory technique to order objects in a way to reveal patterns and regular features easily. Later it spread to all fields of science and the ordering often concern two sequences to be reordered [\u003cspan additionalcitationids=\"CR3 CR4\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. There are, e.g., possibilities to order objects along two axes or one object and one variable sequences in a table. The first applications were connected to fields, where visualization or chronological sequence were natural (archeology, cartography, history, operation research, sociology). Later, especially when information technology is present, different methods and application appeared in many other fields (anthropology, graphics, information visualization, sociometry, psychology, psychometry, ecology, biology, bioinformatics, etc\u0026hellip;). The \u0026ldquo;common\u0026rdquo; in the methods that they are not common for all fields of science. There is a rather small communication among the fields. The most general review was written by Liiv [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], where a historical overview of seriation is detailed including the milestones at several application fields. Instead of enumerating here the methods and provide a deficient and scanty list of applications, we forward the reader to the review of Liiv [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSeriation is applied in a latent way in chemistry [\u003cspan additionalcitationids=\"CR7 CR8\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], but it is seldom termed. It is often used in many scientific software as a default setting, that, e.g., hierarchical clustering is applied on objects and a visually acceptable sequence is generated using a seriated dendogram [\u003cspan additionalcitationids=\"CR11 CR12\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. At the interdisciplinary cheminformatics, the methods are used consciously according to its large emphasize in bioinformatics, where the number of special methods and the corresponding applications is increasing up to now there. From these we mention only the bi- or co-clustering [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. From the special applications in bioinformatics we may refer to similarity search and alignment methods, where our references are some recent reviews. Similarity searching is widely applied and encompasses many techniques, its principle is based on detecting small molecules that have the same biological activity [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. By the same logic, alignment is based on hypothesis of homology [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. The real number of alignment algorithms is in hundreds and continue to increase. The alignment can permit, among others, the identification and quantification of conserved regions or functional motifs, profiling of genetic disease.\u003c/p\u003e \u003cp\u003eGoing back to chemistry, we found up to now only a few articles, where the word seriation is used in the title, abstract or in the keywords [\u003cspan additionalcitationids=\"CR19 CR20\" citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. One interesting example for seriation in chemistry, while not the main subject of the article, showed its importance in summarising the resulting relationships between production groups and chemical clusters, and has enabled external information to be compared with the cluster results. This has permit eventually, to validate and interpret the clusters [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe main aim of seriation is to get better visualization by introducing some order by appropriate sequencing. The ordered sequence may help to find similar objects or variables close to each other in vectors, tables or in their graphics (e.g., in heatmaps). In some cases, it can be the first visual check of data and it serves as a good starting point to estimate which enhanced data analysis method might be tried. Seriation is a non-destructive method, all information remains in the seriated data. The specialized methods usually outperform seriation, e.g., clustering is usually more efficient to identify similar objects than seriation. Seriation also help to visually detect objects and variables with large amount of missing data and outliers. The usual seriations provide heatmaps with reasonable less striped feature, quite often arrangements around the diagonal or visually detectable clusters. Seriation can be performed not only on measured data, but on, e.g., model parameters, neuron intensities in artificial neural networks, connection data, as well. In these cases, seriation might help in the interpretation of the models and the operations.\u003c/p\u003e \u003cp\u003eLiiv started to unify the taxonomy of the different methods [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Theoretically, seriation means the permutation of data stored in one-dimensional vectors up to k-dimensional arrays. There are modes and ways in seriation. A mode means an independent sequence that can be permutated. Way is the dimensionality of the object used in visual perception, during the calculation of a merit function, or during the prescribed operation. The number of modes and ways mostly coincide to the dimension of the data. In chemistry we often have two-dimensional data matrices with \u003cem\u003eN\u003c/em\u003e rows connected to the objects and \u003cem\u003eM\u003c/em\u003e variables denoting the columns. When we sequence both the objects and the variables in a classical data table, we perform two-mode\u0026ndash;two-way seriation. When we sequence only the objects, we usually calculate a symmetric \u003cem\u003eNxN\u003c/em\u003e distance matrix and the seriation is one-mode\u0026ndash;two-way. If we seriate only the variable sequence, the \u003cem\u003eMxM\u003c/em\u003e covariance matrix might be a reasonable choice and the seriation is one-mode\u0026ndash;two-way.\u003c/p\u003e \u003cp\u003eThe number of the methods how to seriate is rather large. There is not any canonical way, the popular methods differ from field to field. Large number of applications can be found usually on the field of bioinformatics, where, e.g., the node deleting algorithm of biclustering is one of the most popular methods [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. There are two groups of the methods. In a part of them, a mathematical merit (or loss) function is defined which depends on the sequencing of the modes. In the other cases, set of operational instructions are used. In the first case, the extremum of the merit or loss function can be found by any global optimization scheme, e.g., simulated annealing, genetic algorithm or other specialized solutions. These methods are mostly iterative and they use a stop criteria. The operational methods are usually repeated as long as the condition of the operation is holding. There might be some extra conditions to avoid infinite loops and it is worthwhile to start both methods from several sequences whereof many can be randomized ones.\u003c/p\u003e \u003cp\u003eAnother aspect of the two-mode seriations whether the two sequences are treated independent from each other, or the merit function/operation contain cross terms. For the independent case an example is the seriation of the objects according to the distance matrix and seriation of the variables according to the covariance matrix. Despite the independence of the two modes, by chance we might get clearly interpretable data, where relations between the two axes are easily readable, e.g., on heatmaps. For having dependent two-mode seriation, we need cross terms between the sequences or geometrical preferences of the matrices. A recipe is sometimes that we have several local function values and the sum or the spatially weighted sum of the local functions provide the merit of loss function. The local functions might be related simply to the increase, to the decrease or to the modality of the data within row or column wise and, e.g., the global loss function is the number of the violated case. For an overview of some of the methods and mathematical details we refer to the review of Liiv [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] and the study of Hahsler et al. [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. For operational algorithms the representation of the problem on graphs can be useful, a part of the algorithm bases on the minimal number of crossings known as Tur\u0026aacute;n\u0026rsquo;s brick factory problem in mathematics and history.\u003c/p\u003e \u003cp\u003eIn our previous studies [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] a local feature was calculated as the distance of two objects in a limited variable-vector space. We used three-variable spaces and the \u003cem\u003ei-j\u003c/em\u003e element of the so called local distance matrix contained the average distance of the \u003cem\u003ei\u003c/em\u003e-th object to its sequential neighbours in the local space formed by the \u003cem\u003ej-1,j,j\u0026thinsp;+\u0026thinsp;1\u003c/em\u003e variables. The global merit function was a weighted sum of these local distances, where the weights were the spatial distance of the \u003cem\u003ei-j\u003c/em\u003e matrix element from the diagonal of the data matrix. The algorithm provided that low local distances were sequenced around the diagonal of the data matrix. The ordering was according to one visual feature, it ordered similar objects close to each other and the corresponding variables around the corresponding diagonal parts were suggested to be responsible for the similarity.\u003c/p\u003e \u003cp\u003eWe experienced that only a part of chemical data is meaningful in the obtained block diagonal forms. For example, there might be a group of variables responsible for two or more clusters of objects what is not easy visually detect on a narrow diagonal-like arrangement. Our first idea was to improve our previous method by introduction of further adaptive lines with similar task as the diagonal had, but during the elaboration we realized that it is easier to think on a seriation forming patches.\u003c/p\u003e \u003cp\u003eIn this paper we show our new method where the local function is a local average similarity, and the global merit function is the sum of the products of the neighbouring local similarities. We found that this merit function forms patches of the neighbouring objects and variables. A patch means a local space, where the given objects are similar to each other. In our philosophy, the object-variable points outside the patches are not relevant for the similarity patterns. We show it on simple chemical data, as well as on model details of artificial neural networks (ANN). The latter is related to the interpretation [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e] of ANN models, e.g., we were able to interpret the roles of the neurons connected to the variables and to the objects. The former means the seriation of the weights in the network and the latter was managed by seriation of activities caused the different objects on the different neurons.\u003c/p\u003e \u003cp\u003eOur method can be easily extended to higher modes and ways seriations. We developed different three-mode-three-way methods. Since it is less easy to interpret three-dimensional data structures than two-dimensional ones, we usually used projections onto two-dimensional heatmaps, here. A part of our results is shown on co-plots elaborated by us, where the original data heatmap and local similarity contour plot are merged. It helps to easily find the responsible variables for the similarity of a cluster of variables.\u003c/p\u003e"},{"header":"2. Theory","content":"\u003cp\u003eIn 2011 we introduced a mathematical merit function for 1-mode-2-way and 2-mode-2-way seriation of matrices [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Two concepts were introduced there. The local distance matrix contained the average distance between the \u003cem\u003ei\u003c/em\u003e-th object and its two neighbours in a local three-variable space, where the index of the middle variable assigned to \u003cem\u003ej\u003c/em\u003e. The other quantity we called diagonal measure, and it represented the distribution of the elements of the local distance. It was the scalar sum of the local distances weighted with their positional distance from the diagonal of the matrix. In 2-mode-2-way seriation the diagonal measure was maximized to order similar objects close to each other and the corresponding variables around the diagonal were suggested to be responsible for the similarity. If distance matrix of the objects was 1-mode-2 way seriated, the diagonal measure was maximized, as well. If covariance matrix of variables was 1-mode-2-way seriated, the diagonal measure was minimized. In our new research we propose development of our idea both on the local quantity and the global measure.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Local similarity matrix\u003c/h2\u003e \u003cp\u003eDistances are unbounded positive numbers what may hinder the interpretation of the actual values. Algorithmically, it is more convenient to use bounded set of values. Similarity is a frequently used concept for that. There is a reciprocal relation between similarity, a value of one denotes perfect similarity of two objects (zero distances of the objects in the variable space) and zero similarity means maximal distance between the objects. There are different definitions of similarity, whereof we finally selected that \u003cem\u003esimilarity\u0026thinsp;=\u0026thinsp;1-(distance/maximal distance)\u003c/em\u003e equation. We defined a local similarity matrix (S) similarly to the local distance matrix. \u003cem\u003es\u003c/em\u003e\u003csub\u003e\u003cem\u003eij\u003c/em\u003e\u003c/sub\u003e shows how the \u003cem\u003ei\u003c/em\u003e-th object is similar to its neighbours in a local 3-variable space around variable j. For 2-mode-2way seriation it is calculated as:\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\({l}_{i,k,j}= \\sqrt{\\sum _{l=j-1}^{j+1}{\\left(\\frac{{a}_{kl}-{a}_{il}}{{diff}_{max,l}}\\right)}^{2}}\\)\u003c/span\u003e \u003c/span\u003e Eq.\u0026nbsp;1\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\({s}_{ij}=\\left(\\sum _{k=i-1,i+1}^{}1- \\frac{{l}_{i,k,j}}{{D}_{col,j}}\\right)/{D}_{row,i}\\)\u003c/span\u003e \u003c/span\u003e Eq.\u0026nbsp;2.\u003c/p\u003e \u003cp\u003e,where \u003cem\u003ea\u003c/em\u003e\u003csub\u003e\u003cem\u003eil\u003c/em\u003e\u003c/sub\u003e and \u003cem\u003ea\u003c/em\u003e\u003csub\u003e\u003cem\u003ekl\u003c/em\u003e\u003c/sub\u003e are the elements of the \u003cem\u003eA\u003c/em\u003e matrix to be seriated, \u003cem\u003el\u003c/em\u003e\u003csub\u003e\u003cem\u003eikj\u003c/em\u003e\u003c/sub\u003e is their local distance in the variable space formed by the \u003cem\u003el\u0026thinsp;=\u0026thinsp;j-1, j\u003c/em\u003e and \u003cem\u003ej\u0026thinsp;+\u0026thinsp;1\u003c/em\u003e variables. \u003cem\u003ediff\u003c/em\u003e\u003csub\u003emax,l\u003c/sub\u003e is the difference between the largest and the smallest elements of the \u003cem\u003el\u003c/em\u003e-th column in \u003cem\u003eA\u003c/em\u003e. It is used to scale the distance between [0,sqrt(3)], if the local variable space contains three variables. If the \u003cem\u003ej\u003c/em\u003e-th variable is at the first or the last column of the matrix, the local space contains only [0,sqrt(2)]scaled distances. \u003cem\u003eD\u003c/em\u003e\u003csub\u003e\u003cem\u003ecol,j\u003c/em\u003e\u003c/sub\u003e contains the corresponding upper bounds of the intervals for each variable. \u003cem\u003eD\u003c/em\u003e\u003csub\u003e\u003cem\u003erow,I\u003c/em\u003e\u003c/sub\u003e is usually two for the \u003cem\u003ei\u003c/em\u003e-th object except the first and the last row, where it is one. These row or column dependent quantities (\u003cem\u003ediff\u003c/em\u003e\u003csub\u003e\u003cem\u003emaxl\u003c/em\u003e\u003c/sub\u003e, \u003cem\u003eD\u003c/em\u003e\u003csub\u003e\u003cem\u003ecol,j\u003c/em\u003e\u003c/sub\u003e, \u003cem\u003eD\u003c/em\u003e\u003csub\u003e\u003cem\u003erow,i\u003c/em\u003e\u003c/sub\u003e) were introduced to be able to get theoretically \u003cem\u003es\u003c/em\u003e\u003csub\u003eij\u003c/sub\u003e ϵ [0,1] values for all \u003cem\u003ei-j\u003c/em\u003e positions including the non-bulk matrix elements.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 The global patch function\u003c/h2\u003e \u003cp\u003eIn the case of our previous global scalar (diagonal measure), the seriated matrix placed the variables responsible for object similarities around the corresponding part of the diagonal. It means, only the most important variables were emphasized, and, e.g., there was no possibility to select a variable to be important for several object clusters. In our new method we define a merit function, where forming of several patches is supported by maximising it in 2-mode-2way seriation. If we calculate the sum of the product of two neighbouring local similarity values (P), this quantity reflects the spatial distribution of large and small similarities. If random order of objects and variables is used, the local similarities are distributed randomly in the matrix. If we seriate the matrix to have larger sum of neighbouring products, a higher sum can be reached by clustering high and low local similarities separately. Furthermore, preferential rearrangement is also supported by maximising such a merit function which creates higher similarities by neighbouring similar objects. In the high similarity patches the objects are similar in the local variable space and both the objects and the variables can be identified.\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(P=\\sum _{k=1}^{nxm}\\sum _{l}^{}{\\left({s}_{k}{s}_{l}\\right)}^{q} = \\sum _{i=1}^{n-1}\\sum _{j=1}^{m}{2\\left({s}_{ij}{s}_{i+1,j}\\right)}^{q}+\\sum _{i=1}^{n}\\sum _{j=1}^{m-1}{2\\left({s}_{ij}{s}_{i,j+1}\\right)}^{q}\\)\u003c/span\u003e \u003c/span\u003e Eq.\u0026nbsp;3.\u003c/p\u003e \u003cp\u003e, where \u003cem\u003ek\u003c/em\u003e goes over all elements of the local similarity matrix, \u003cem\u003el\u003c/em\u003e denotes the given neighbours of \u003cem\u003ek\u003c/em\u003e with one common index and an index differing with +/- 1. q is an arbitrary contrast. The two effects of maximising \u003cem\u003eP\u003c/em\u003e - spatial ordering and creation of high similarities - can be justified separately. The simple rearrangement of any matrix by clustering large and small values provides large \u003cem\u003eP\u003c/em\u003e: it is similar to a negative local entropy. We performed several test calculations supported this and there is also a thought experiment in the supporting material. The other effect is straightforward, that placing similar objects and variables close to each other increases the sum of the local similarities. The exponent \u003cem\u003eq\u003c/em\u003e is an empirical contrast factor. At high q values the positioning of highest similarities close to each other is extremely preferential and it may cause compact and small clusters, while small \u003cem\u003eq\u003c/em\u003e-s do not penalize so strictly the less large values, it may cause slightly larger patches. We used \u003cem\u003eq\u003c/em\u003e\u0026thinsp;=\u0026thinsp;2 and \u003cem\u003eq\u003c/em\u003e\u0026thinsp;=\u0026thinsp;3 in our calculation. Depending on the dataset, the visual results was sometimes better, for one of the q choices, but it did not seem to be a decisive parameter of the merit function. We note, that our patch function was obtained after several trials, where at first we focused on entropy or Gini-index like approximations. We found, that \u003cem\u003eP\u003c/em\u003e defined as in Eq.\u0026nbsp;3 is a simple and feasible merit function.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Local similarities in higher dimensions\u003c/h2\u003e \u003cp\u003eThe generalization of the local similarity matrix and the patch function to higher dimensions can be easily done, if we follow the idea that we are interested in the average similarity of an object to its sequential neighbours in a local three-variable space. The calculation of the possible cases, e.g., the dimension of the original data, the dimension of the local similarity matrix, the number of possible local variable vectors are detailed in the \u003cspan refid=\"Sec10\" class=\"InternalRef\"\u003eResults and Discussion\u003c/span\u003e section together with some examples. We show three possibilities for 3-dimensional local similarities, where the three axes are formed by one object and two variable vectors (OVV case, original data are 2D), by two object and one variable vectors (OOV-independent, the original data are two dimensional) and by another two object and one variable vectors case (OOV-dependent, the original data are three dimensional).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Missing data and noninformative zeros\u003c/h2\u003e \u003cp\u003eThere are several datasets in chemistry, where part of the data is missing. The causes might be different, e.g., lack of general experimental methods for all objects, operational break down, or the given variable is not relevant for that object. In the case of cheminformatics it also common, that several extra variables are added to the database where most of the objects provides a zero value. An example is the presence of chemical groups, if close to all the molecules do not contain that functional group. The traditional method to overwrite the missing data with an average or random value might bias the seriation. Therefore, it would be feasible to avoid the replacement of missing data. Also, it is rather misleading, if the unnecessary and irrelevant zeros have crucial effect on the merit function of seriation. We solve the problem of missing data and unnecessary zeros by proposing a different calculation of the local similarities for these cases:\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\({s}_{ij}=\\sum _{k=i-1,i+1}\\sum _{l=j-1}^{j=j+1}\\left(1-\\left|\\frac{{a}_{kl}-{a}_{il}}{{diff}_{max,l}}\\right|\\right)/6\\)\u003c/span\u003e \u003c/span\u003e, Eq.\u0026nbsp;4\u003c/p\u003e \u003cp\u003eThe inner sum is skipped for all data, where any of the data (\u003cem\u003ea\u003c/em\u003e\u003csub\u003e\u003cem\u003ekl\u003c/em\u003e\u003c/sub\u003e or \u003cem\u003ea\u003c/em\u003e\u003csub\u003e\u003cem\u003eil\u003c/em\u003e\u003c/sub\u003e) is non-existent. It can be used for missing data as well as for unnecessary zero values. Using Eq.\u0026nbsp;4 the local similarity cannot be one, if there are undetermined cases in the sum. Also, if \u003cem\u003es\u003c/em\u003e\u003csub\u003e\u003cem\u003eij\u003c/em\u003e\u003c/sub\u003e refers to a matrix position at edges or corners, the possibility for the local similarities to be 1 is excluded, there maximum value is 2/3, 1/2, or 1/3. This handling of the borders is different from Eq.\u0026nbsp;2. The patch function is calculated according to Eq.\u0026nbsp;3. There is only one difference, there might be a chance that a sij remains undetermined. In that case the undetermined \u003cem\u003es\u003c/em\u003e\u003csub\u003e\u003cem\u003eij\u003c/em\u003e\u003c/sub\u003e is skipped in Eq.\u0026nbsp;3.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Calculation Details","content":"\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Codes\u003c/h2\u003e \u003cp\u003eThe patch seriation was performed using a C code developed in our laboratory. The code reads the datasets, manages data pre-processing as optional normalization, scaling, changing zeros to undetermined values. The maximization of the patch function was obtained with Metropolis Monte Carlo algorithm, where the ordering with the largest \u003cem\u003eP\u003c/em\u003e was stored as the best one. The acceptance ratios for the different type of trial changes were set to be around 0.05. Column and row permutations were performed independently. In the case of three-mode-three-way seriation it was performed independently for all modes. The number of the trials was 1\u0026ndash;5\u0026nbsp;million. A calculation took a few minutes on a PC depending on the size of the dataset. During this calculation length, usually the best sequence was detected and stored at any time after the 20% of the calculation time. A few (2\u0026ndash;5) seriations were performed for each dataset at \u003cem\u003eq\u0026thinsp;=\u0026thinsp;2\u003c/em\u003e and \u003cem\u003eq\u0026thinsp;=\u0026thinsp;3\u003c/em\u003e values, which of the results to be shown were selected visually.\u003c/p\u003e \u003cp\u003eThe elaboration and visualization of seriation results was done using R [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. In the comparison to other methods, here we used the seriation package of Hahsler et al. [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. For three-dimensional graphs we used the RGL package [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. We developed an overlay plot, where the heatmap coded scaled data are shown together with contour plots of the local similarity. We think, these overlay plots are rather effective to identify object clusters and the variables causing the similarity.\u003c/p\u003e \u003cp\u003eThe neural network modelling was performed in Python using the scikit learn package [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. The partly optimized hyperparameter sets for the models were selected from one of our previous studies where we used the same datasets [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. All codes are deposited and freely downloadable at the homepage of the corresponding author after the publication of the article.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Datasets\u003c/h2\u003e \u003cp\u003eThe tested datasets are mostly freely available ones related to QSAR, chemistry, material science, food science, cheminformatics and environmental chemistry. Several datasets of them are accessible in repositories [\u003cspan additionalcitationids=\"CR26\" citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. Some details of the data and the performed type of seriations are collected in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. The first dataset (SIM[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]) is a semi-randomly simulated one, its structure is related to our initial idea, what kind of benefit we would like to get using patch seriation. There are 50 objects and 20 variables in this set ordered in 4 clusters and a random group for the objects. Members of the clusters have similar values at some selected variables, but their other data are random. Some of the selected variables are common also with other clusters. At first, we generated [0,1) random numbers for all data and thereafter the groups were recalculated by adding a given random number for that variable of the group biased with white noise. In the case of other datasets, if a dataset was published for modelling a response variable, we omitted it from the seriation and only the predictor variables were used in the seriation process.\u003c/p\u003e \u003cp\u003eThe RETSIM dataset [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e] is a simulated one, as well. We defined three functional groups and created 4 compounds with random linear combination of the three groups. We set 6 mixtures of the 4 compounds. 6 chromatographic columns were set as well with differently randomized partial retention times for the functional groups. The retention times of the compounds were calculated with linear combination of the functional groups therein. Finally, we added uniform broadening for each compound with integrals related to the concentrations. In this way we had 36 chromatograms of the 6 mixtures on the 6 columns.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDatasets\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eabbr.\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003erow \u0026times; column\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003edescription\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eseriations\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eref.\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSIM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e50\u0026times;20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4 clusters with common variables for each\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRETSIM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(6x6)x100 and (6x6)x150\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003esimulated retention times of mixtures on different columns (+\u0026rsquo;fingerprints\u0026rsquo;)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV; OOV dependent\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePOL_MONTH\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(26x12)x9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003emonthly air pollutant averages at 26 stations in 2017\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV; OOV dependent\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePOL_YEAR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e(12\u0026thinsp;+\u0026thinsp;14)x9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eyearly air pollutant averages at 26 stations in 2017 (12 at Budapest, 14 at countryside)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV; OOV independent\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFLASHP1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e420x26 for ANN model (N\u0026thinsp;=\u0026thinsp;4,6,8,10 hidden neurons), 80 objects in the test set\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eflash point estimation of molecules using different QSAR parameters\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV: 80x26 (test objects-variables); 26xN (variables neuron weights); Nx80 (neurons, object activities on the neurons)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDR8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e600x28 for ANN model (N\u0026thinsp;=\u0026thinsp;10\u0026ndash;15 neurons), 114 objects in the test set\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003edifferent QSAR parameters originally used to estimate toxicity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV: 114x28 (test objects-variables); 28xN (variables-neuron weights); Nx114 (neurons-object activities on the neurons);\u003c/p\u003e \u003cp\u003eOVV independent 114x(N\u0026thinsp;+\u0026thinsp;28) (objects, neuron activities, original variables)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFLASHP2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e632x(13\u0026thinsp;+\u0026thinsp;12), for ANN models (N\u0026thinsp;=\u0026thinsp;10\u0026ndash;15 neurons), 100 object in the test set\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e13 molecular and 12 general descriptors of molecules originally used for flash point estimation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV of 100x25, 100x13, 100x12;\u003c/p\u003e \u003cp\u003eOVV 100x(13\u0026thinsp;+\u0026thinsp;12); OVV 114x(N\u0026thinsp;+\u0026thinsp;28) (objects, neuron activities, original variables)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePOLMET_DAY\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e56x(7\u0026thinsp;+\u0026thinsp;6), subset of original, two weeks from each season (56 days)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDaily meterological and airpollutant data set in 2007\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV: 56x13, 56x7, 56x6; 3D OVV: 56x(7\u0026thinsp;+\u0026thinsp;6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eESSOIL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10x(10\u0026thinsp;+\u0026thinsp;38)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eEssential oils in 10 species, 10 chemical and 38 bactericid/fungicide data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV: 10x48, 48x10;\u003c/p\u003e \u003cp\u003e3D OVV: 10x(10\u0026thinsp;+\u0026thinsp;38)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCERAMIC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e88x17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eceramics with body and glaze data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGLASS\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e214x9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ecomposition of glasses from different sources\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWINE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e178x13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ewine analysis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTOXIC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e112x8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003etoxic on 8 data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSAND\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e30x14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003esand data radiation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMOLDESCRg\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e500x50\u003c/p\u003e \u003cp\u003emore subsets of the original\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003emolecular descriptors for enormous number of molecules to calculate different properties\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOIN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e257x10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ecomposition of ancient coins from different era of Hungary\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan additionalcitationids=\"CR41\" citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eREAC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e95x32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003efuel combustion with reactions and reactants\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eOV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"4. Results And Discussion","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Simulated dataset for 2-mode-2-way seriation\u003c/h2\u003e \u003cp\u003eThe dataset contained 50 objects and 20 variables. Each of the 4 clusters had 10 objects with similar set of 5\u0026ndash;5 variables. These variables were distinct except for group C and D, here two variables were common but with different average values for the two set of objects. 10 objects and two variables had no cluster affiliation. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the heatmaps of a randomly ordered matrix (used as start), the seriated data matrix, the corresponding local similarity matrix and the corresponding hidden cluster information at the data generation in the seriated order. If we plot only the data matrix as a heatmap, it is not easy to identify the clusters. Therefore, we use the overlay of a contour plot on the local similarity matrix in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb. In more than one third of our trials the variables were seriated perfectly and the number of clusters of the objects was equal or only slightly more than 4. We show a case in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb-\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ed, where the seriation worked perfectly both for variables and objects. We choose the actual contour levels of the local similarity matrix to help the assignment of the clusters in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb. The 10 random objects and the 2 random variables are sometimes between the clusters or sometimes they are neighbouring to each other, but the overlay contour plot does not identify them as a 5th cluster.\u003c/p\u003e \u003cp\u003eThe number of the object clusters is 4 in this example. We compared it to other seriation methods built in R [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The number of object clusters were between 14\u0026ndash;33 for the other methods. It means, the clusters were split into 3\u0026ndash;8 parts in average. We also calculated how many of the variables is found in a cluster for the object clusters. In our seriations the 5 variables were mostly clustered for all object groups correctly. In the case of the other methods, it was between 13\u0026ndash;18. It means, there were only 2\u0026ndash;7 cases, when two common variables of object clusters were placed to be neighbours in the variable sequence. The supplementary material contains further details on the comparison of the methods. We should emphasize, that our patch seriation worked efficiently both for objects and variables simultaneously and it explores the link between the two modes. This link is missing for most of the other methods. As it can be seen in the supplementary material, most of the methods use only distance matrices of the objects, where the simultaneous sequencing of the two axes is not possible. In our comparison, we calculated also variable sequencing of the other methods by calculating a \u0026lsquo;distance matrix\u0026rsquo; of the variables, as well.\u003c/p\u003e \u003cp\u003eIn this prototype of data different groups of variables are responsible for the different clusters of the objects and the other variables are not important for the object clustering. It seems so, that our method totally outperforms all the other seriation methods. Even more, clustering methods, e.g., hierarchical clustering is not able to find this type of relationship within the objects due to the nonlocal distance calculations.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e4.2 2-mode-2-way seriation of other datasets\u003c/h2\u003e \u003cp\u003eThe most of the tested patch seriations concerned the simultaneous ordering of objects and variables in two dimensions. Here we show some examples with original and seriated heatmaps. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. shows three datasets.\u003c/p\u003e \u003cp\u003eThe first is the POL_YEAR one [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea-b. It contains the yearly averaged concentrations of 9 air pollutants at 12 places in Budapest and 14 places at countryside (mostly in cities) in 2017. There are differences in the measuring stations, not all of them were able to measure all the 9 components and there were several shutdowns at some places. Here, we calculated the local similarity matrix according to Eq.\u0026nbsp;4. The seriation clearly shows that the dominant variables for similarity are the different nitrogen-oxide and ozone concentrations. The stations with heavy traffic are ordered close to each other. The other stations are separated into two groups, where the nitrogen oxides are dominant pollutants and where ozone pollution is dominant. It is known, that there is a transition cycle of ozone and nitrogen-dioxide. The first and the last stations are somehow thrown out by the seriation. These are the places where the number of missing data was high.\u003c/p\u003e \u003cp\u003eIn Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec-d we show a case, where the component of ceramics (celadons) are the variables [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. Here it is known for the objects, what kind of celadon ceramics and what part of them are analysed (body or glaze). The seriation clearly shows, how the randomized data can be turned to form two groups, body and glaze, according to the low concentration of some oxides in the body part and high concentrations of them in the glaze. The seriation also found a subgroup of celadons, which were only imitated Longquan celadon in Jingdezhen civilian kilns in Ming Dynasty. The seriation was not able to differentiate the Longquan celadons of different dynasties.\u003c/p\u003e \u003cp\u003eIn Fig.re 2e-f we show a set of reactions and reactants used in the combustion modelling of gasoline [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. It is not easy to determine an order of the reactions and reactants. Previously we used our diagonal seriation for this reaction set and we were able to arrange them according to a diagonal suggesting a hypothetical pathway. Of course, that simple order was related to a non-realistic graph structure, where it is well known that a proposed way need not coincidence to the real fluxes of the processes, especially the fluxes highly differ for different combustion conditions. In the case of patch seriation, we concentrated on the identification of reaction parts, e.g., reactions using the same components as reactants or products. The presence of a component in a reaction was denoted with 1 irrespectively the components role and stochiometry. The gasoline components, the final CO\u003csub\u003e2\u003c/sub\u003e and H\u003csub\u003e2\u003c/sub\u003eO components are coloured differently. It can be seen in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ef, that there is a reasonable re-clustering of the reaction system, where around six clusters of reaction-components are there. Two of them is related to the final products CO\u003csub\u003e2\u003c/sub\u003e and H\u003csub\u003e2\u003c/sub\u003eO, another is related to the CH\u003csub\u003e4\u003c/sub\u003e and CH\u003csub\u003e3\u003c/sub\u003e components, one is formed by different small entities containing H and O, and another contains additionally carbons. The group on the bottom-middle is related to H and H\u003csub\u003e2\u003c/sub\u003e. Such kind of seriation might be interesting if one intends to build reaction mechanism in a modular way.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e we show two cases, where the seriation helps to find some general pattern. In the GLASS and COIN compositional datasets the exchange of species can be easily detected in the seriated data besides the visual simplification of the heatmaps. In the seriated GLASS data (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eb) one can detect a negative mirror like difference in the exchange of ions with similar charges, e.g., K\u003csup\u003e+\u003c/sup\u003e - Na\u003csup\u003e+\u003c/sup\u003e; Ca\u003csup\u003e2+\u003c/sup\u003e - Mg\u003csup\u003e2+\u003c/sup\u003e; Al\u003csup\u003e3+\u003c/sup\u003e - Ba\u003csup\u003e3+\u003c/sup\u003e. Furthermore, it is also easy to detect the relation between Al\u003csup\u003e3+\u003c/sup\u003e and the refractive index. It is not easy to realize these features in the unseriated data according to the rather striped heatmap (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea). The method also seriated the glasses quite well according to their sources or use detailed in the original source [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ec-d contain a dataset on Hungarian coins from the X-XIII. century. Here we used Eq.\u0026nbsp;4 and set the zero values to be skipped during the calculation of the local similarity matrix. The heatmap shows a similar exchange of species, like Cu-Ag exchange. It groups the coins where Sn and Sb took part in it, as well. The seriation was done with 0\u0026ndash;1 scaled data, therefore, the traces of some metals had a large effect on the ordering. Due to the scaling, it is easy to identify the coins having the same metals from second importance. Our method clusters the objects (coins) according to the era and kings, but here we should add that traditional clustering and classification methods provided better results [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e4.3 3D seriation: 1-object\u0026ndash;2-variables case\u003c/h2\u003e \u003cp\u003eIf we have a two-dimensional data matrix, where the columns contain two separable set of variables, we might perform a seriation, where the order of the objects, the first set variables and the second set of variables can be separately sequenced. A three-dimensional local similarity matrix can be constructed, where an element shows the average similarity of the given object to its neighbours (axis one), but now in two local three-dimensional variable spaces (axis two or axis three). Using the first type of variables and the second type of variables independently, we calculate the two local similarities between the given object and one of its neighbours. The final \u003cem\u003es\u003c/em\u003e\u003csub\u003eijk\u003c/sub\u003e local similarity contains the average of four similarities (over the two neighbours times the two local variable spaces).\u003c/p\u003e \u003cp\u003eSuch an independent set of variables can be, e.g., the chemical content and the biological activities of the essential oils (ESSOIL), the daily averages of air pollutants and the corresponding meteorological data (POLMET-DAY) or the complex QSAR descriptors and the simple enumeration of functional groups in the flash point modelling FLASHP2.\u003c/p\u003e \u003cp\u003eWe show our results on the latter one, where the QSAR descriptors for a given molecule are called molecular descriptors in the original paper and the enumeration of functional groups is called general descriptors. If we apply separately two-dimensional seriation for the two set of variables, the obtained heatmaps are clearly arranged (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ea-b). Here we calculated the local similarities with skipping the zero data to avoid the clustering of molecules due to the lack of functional groups in the set. If we seriated the total data matrix in two-dimensions, many of the clear patches disappeared. The continuous molecular variables were dominant during the seriation, most of the general descriptors was not clearly seriated (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ed). If we performed the seriation using a three-dimensional local similarity matrix, the two parts of the data in the original two-dimensional dataset provided clear patches for both set of variables (projected back to two dimensions: Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ee). The advantage of the three-dimensional seriation over the two independent two-dimensional seriations is the common target function during the sequencing of the three axes. The three-dimensional local similarity array can be directly visualized (see later Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eb) or two-dimensional projections can be calculated. Figure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ef is the projection, where the highest local similarities over the objects are collected for a given set one \u0026ndash; set two variable pair. It shows that the highest similarities are obtained for which local variable pairs of set one and set two ones. We emphasize here again, that our mathematical target function connects all the modes of the seriation in contrary to the usual biclustering schemes.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e4.4 3D seriation: 2-objects-1-variable case\u003c/h2\u003e \u003cp\u003eFor the demonstration of the case, where a two-dimensional data matrix contains two different sets of objects we selected the air pollution data of stations in Budapest and at countryside (POL-YEAR). The two sets of objects might be seriated independently. Here the first axis of the local similarity matrix contains the stations at Budapest, the second one is the stations at countryside and the third axis shows the yearly averaged pollutant concentrations. The \u003cem\u003es\u003c/em\u003e\u003csub\u003eijk\u003c/sub\u003e local similarity contains the average of 4 similarities: the similarity of the \u003cem\u003ei-(i-1)\u003c/em\u003e and \u003cem\u003ei-(i\u0026thinsp;+\u0026thinsp;1)\u003c/em\u003e object pairs of the first axis (stations in Budapest) and the \u003cem\u003ej-(j-1)\u003c/em\u003e and \u003cem\u003ej-(j\u0026thinsp;+\u0026thinsp;1)\u003c/em\u003e object pairs of the second axis (stations at countryside). The local variable space for this element is spanned by the \u003cem\u003ek-1, k, k\u0026thinsp;+\u0026thinsp;1\u003c/em\u003e variables.\u003c/p\u003e \u003cp\u003eThe results of the seriation can be shown in the original two-dimensional data (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ea), where both set of stations are rather homogeneously sequenced separately (cf. to Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb. of 2D seriation, where the cities might be mixed). Using interactive three-dimensional graphics, one might have a look on the local similarity array. We show an example in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eb, but it is rather uninformative without the possibility of rotating the graph. The different projections or enumerations on different subspaces of the three axes might be informative, e.g., which pollutant causes locations in Budapest and in countryside to be similar. It is clear from the graphs, that mostly the NO-NO\u003csub\u003ex\u003c/sub\u003e-NO\u003csub\u003e2\u003c/sub\u003e, and sometimes the SO\u003csub\u003e2\u003c/sub\u003e and PM10 data cause the similarities. Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e shows this projection where highly similar locations are ordered into the middle of the local similarity array. The corresponding alphabetical code enumerates all the local variables involved in at a given high similarity, e.g., a similarity over 0.9 at the O\u003csub\u003e3\u003c/sub\u003e-NO\u003csub\u003e2\u003c/sub\u003e-NO\u003csub\u003ex\u003c/sub\u003e position means that the two neighbouring variables (SO\u003csub\u003e2\u003c/sub\u003e and NO) are also involved therein.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eLocal similarities over 0.9 between stations in Budapest and in countryside. The local variables causing the similarities: A: NO\u003csub\u003e2\u003c/sub\u003e,NO\u003csub\u003eX\u003c/sub\u003e,NO B: O\u003csub\u003e3\u003c/sub\u003e,NO\u003csub\u003e2\u003c/sub\u003e,NO\u003csub\u003eX\u003c/sub\u003e,NO C: NO\u003csub\u003e2\u003c/sub\u003e,NO\u003csub\u003eX\u003c/sub\u003e,NO,PM10 D: NO\u003csub\u003eX\u003c/sub\u003e,NO,PM10 E: SO\u003csub\u003e2\u003c/sub\u003e,O\u003csub\u003e3\u003c/sub\u003e,NO\u003csub\u003e2\u003c/sub\u003e,NO\u003csub\u003eX\u003c/sub\u003e,NO F: O\u003csub\u003e3\u003c/sub\u003e,NO\u003csub\u003e2\u003c/sub\u003e,NO\u003csub\u003eX\u003c/sub\u003e,NO,PM10 G: SO\u003csub\u003e2\u003c/sub\u003e,O\u003csub\u003e3\u003c/sub\u003e,NO\u003csub\u003e2\u003c/sub\u003e,NO\u003csub\u003eX\u003c/sub\u003e,NO,PM10 H: SO\u003csub\u003e2\u003c/sub\u003e,O\u003csub\u003e3\u003c/sub\u003e,NO\u003csub\u003e2\u003c/sub\u003e,NO\u003csub\u003eX\u003c/sub\u003e I: O\u003csub\u003e3\u003c/sub\u003e,NO\u003csub\u003e2\u003c/sub\u003e,NO\u003csub\u003eX\u003c/sub\u003e The stations belonging to the columns (Budapest) and the countryside (rows) are listed in the supplementary material.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"13\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c11\" colnum=\"11\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c12\" colnum=\"12\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c13\" colnum=\"13\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c10\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c11\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c12\"\u003e \u003cp\u003e11\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c13\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eG\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003eD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eH\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e4.5 3D seriation: 2-objects-1-variables - dependent case\u003c/h2\u003e \u003cp\u003eIn the case of two dependent object vectors - one variable vector, the data are originally three dimensional, but we usually have an unfolded two-dimensional data matrix. In the case of the RETSIM data unfolded to two dimensions, the first six rows correspond to the retention times of the 6 mixtures on the first chromatographic column, the next six rows for the second column, etc\u0026hellip; The variables (columns) were the retention intensities for 1-100 arbitrary time intervals. The three axes of the local similarity matrix could be the chromatographic columns, the mixtures and the retention times. The \u003cem\u003es\u003c/em\u003e\u003csub\u003eijk\u003c/sub\u003e local similarity contains the average of 4 similarities: the similarity of the \u003cem\u003ej-th\u003c/em\u003e mixture data for the \u003cem\u003ei-(i-1)\u003c/em\u003e and \u003cem\u003ei-(i\u0026thinsp;+\u0026thinsp;1)\u003c/em\u003e chromatographic column pairs and the similarity of the \u003cem\u003ei-th\u003c/em\u003e chromatographic column data for the \u003cem\u003ej-(j-1)\u003c/em\u003e and \u003cem\u003ej-(j\u0026thinsp;+\u0026thinsp;1)\u003c/em\u003e mixture pairs. The local variable space for this element is spanned by the \u003cem\u003ek-1, k, k\u0026thinsp;+\u0026thinsp;1\u003c/em\u003e retention time intensities.\u003c/p\u003e \u003cp\u003eIf we seriate the three axes independently (only the order of the columns and separately the order of the mixtures), a rather limited information can be obtained on the new order, e.g., some link between the mixtures and columns. Furthermore, for such a seriation we need to know a priori the mixture and column labels for each spectrum. It would be more interesting, if we seriate in an unrestricted way, where we fill the 6 times 6 object space with maximal freedom without taking care on the order of the columns and the mixtures. In this case, seriation has a strong explanatory statistical feature, if we will be able, e.g., to cluster the common measurements for a column or a given mixture.\u003c/p\u003e \u003cp\u003eThe simple 2-mode-2-way seriation of the retention spectra was able to classify the spectra for each column perfectly (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eb). On contrary, there were no patterns how the mixtures were ordered. If we use the three-dimensional spectra with supposing a 6x6x100 local similarity matrix, we still obtained a good order for the columns (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003ec), but the interpretation of the result is not easy. In the three-dimensional seriation an object has two-two neighbours in two variable subspaces with distances at three local retention times. There is a chance distribution that which type of objects are neighbouring in which spaces, e.g., there is no driving force that the column like neighbouring is according to the first axis or the second one. Even more, it can be different at the different grids of the local similarity matrix. For example, the three-dimensional seriation provided a good arrangement for the columns, around 70% of the 4 neighbours were measurements on the same columns. Spatially, it occurred along two axes and the unfolded data were partitioned into 3\u0026ndash;3 long sequences for each column. It is not easy to see it on the unfolded seriated data.\u003c/p\u003e \u003cp\u003eWe found that it is not straightforward how to seriate the mixtures using the retention times. Therefore, we added 50 more variables to the data, where the largest 50 data of a row was normalized with the largest in the same row and were sorted in a descending order. These 50 new variables work as mixture specific fingerprints. We performed 2-mode-2-way and 3-mode-3-way seriations using this extended variable set. The results are shown in tabular form over Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003ed-e, since the three-dimensional seriations cannot be unfolded by a visually striking way (c.f. the unambiguity of the assignment of the planes of the ordering and the column-mixture features.) Some features of the seriation are shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. It can be seen that the use of three-dimensional local similarity matrix helped slightly to get better classification on the mixtures, as well. 100 (84\u0026thinsp;+\u0026thinsp;16) or 106(84\u0026thinsp;+\u0026thinsp;22) correct neighbours were sorted from the altogether 120 possible neighbouring positions.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance of seriation on the RETSIM dataset\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSeriation type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMaximal theoretical number of correct column neighbours\u0026thinsp;+\u0026thinsp;correct mixture neighbours (2D), or of all neighbours (3D)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRealization in an example for columns\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRealization in an example for mixtures\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e36x100 size, 2D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e60\u0026thinsp;+\u0026thinsp;60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e60 (perfect classification)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e36x150 size, 2D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e60\u0026thinsp;+\u0026thinsp;60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e60 (perfect classification)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6x6x100 size, 3D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e120\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e84 (3\u0026ndash;3 groups in 2D unfolded map)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6x6x150 size, 3D\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e120\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e84 (3\u0026ndash;3 groups in 2D unfolded maps)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e22 (repetition of mixtures in +-6 shifts)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe other example is the air pollutant dataset containing 9 pollutants at 26 stations. The data are the monthly averages in year 2017 (POL-MONTH). The theoretical three axes are the stations, month and the pollutants. Figure\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003ea is a randomized data matrix where both rows (stations in a given month) and columns (pollutants) are randomized. If we perform a 2-mode-2-way patch seriation, the heatmap became simpler, e.g., the NO\u003csub\u003e2\u003c/sub\u003e, NO and NO\u003csub\u003ex\u003c/sub\u003e variables were seriated near to each other (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eb). Here we used that zero and missing values were not used in the similarity calculations (Eq.\u0026nbsp;4). The stations and pollutants with a lot of missing values are out-seriated to the edges of the heatmap. One can see, as in the case of the monthly averages, that the high nitrogen-oxide and ozone data provided a good basis for similarity. Figure\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003ec shows the original data, where an arbitrary alphabetical order was used for the stations while the months are in calendar order. If we perform the 3-mode-3-way patch seriation, we obtained an ordered map with regular stripes. The neighbour analysis showed, that 48\u0026ndash;59% of the neighbouring objects in the local similarity array belong to the same season, while this is only 39\u0026ndash;47% for the 2-mode-2-way seriation. Around 30% of the four neighbours in the similarity array are the same in the station and/or in the month. We note that we do not want to get a perfect classification for these data, because it is not obligatory that objects of different classes (location, month or season) could not be closer to each other than objects from the same classes. The \u003cem\u003eP\u003c/em\u003e (Eq.\u0026nbsp;3) of Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003ec (perfect classification) is around the at the middle of the random and the best \u003cem\u003eP\u003c/em\u003e-s. Our method is data driven and it helps to override traditional classification, where the data do not support to clearly perform classification.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e4.6 Seriation of neural network model data\u003c/h2\u003e \u003cp\u003eArtificial neural network is one of the most popular methods to solve classification and regression tasks. The simplest conventional structure contains an input, a hidden and an output layer, where the input and the hidden ones and the hidden and the output ones are connected using sets of weights. In the simplest case, the hidden layer neurons contain activation functions, and the output layer ones only sum their weighted input. In the case of classification and regression, supervised method is used, where the weights are optimised to get a correct output for a training set. In an optimal situation, there is an independent test set to validate the model. There are two basic trends in the evaluation of models. In the novel applications of data science we concentrate on the output performance without restricting the complexity of the ANN models. In the traditional case we would like to have limited complexity of the models with some possibility to interpret the model itself. The order of the neurons in the input, hidden and output layer are usually totally arbitrary. Therefore, it is an open question for seriation, especially, if we would like to interpret and understand a given model. Here we focus on some simple cases.\u003c/p\u003e \u003cp\u003eThe first one is the visualisation of the weights between the input and the hidden layer. We might choose the hidden layer neurons as objects and the weights are the columns assigned to the input channels. The opposite assignment is also meaningful, where the input channels (input variables) are the objects and the number of the variables is equal to the number of the neurons in the hidden layer. The same data matrix can be used, in one case the original matrix is seriated, in the other case its transpose is the input.\u003c/p\u003e \u003cp\u003eIn Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e. we show the seriated results for the dataset FLASHP1. The original data was intended to estimate the flash point of different molecular systems. The predictor matrix contains information on the presence of different functional groups. We built several ANN models using several hyperparameters settings. Here we show 9 ANN models with three different activation functions and 4, 6 and 8 neurons in the hidden layer. The same training and test set was used for each model and we selected models with both \u003cem\u003eR\u003c/em\u003e\u003csup\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sup\u003e and \u003cem\u003eQ\u003c/em\u003e\u003csub\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sub\u003e\u003csup\u003e\u003cem\u003eF2\u003c/em\u003e\u003c/sup\u003e (\u003cem\u003eR\u003c/em\u003e\u003csup\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sup\u003e\u003csub\u003e\u003cem\u003etest\u003c/em\u003e\u003c/sub\u003e) more than 0.9. The graph shows the case, where the neurons were the objects and the local variable spaces were formed by the weights assigned to the input channels. According to the calculation of the patch function (using Eq.\u0026nbsp;1), we scaled here the variables. It means, in the presence of both negative and positive weights blue colour might denote a large negative weight and red colour denotes large positive weight. Of course, the scaling might bias the interpretations, but in this feasibility study we do not intend to go really into the details of any ANN model. One can see that some of the seriated graphs (models with 4 neurons and models using logistic activation function) are clearly arranged providing the possibility of interpreting the operation of the model. In the case of this dataset, logistic activation seems to be the most interpretable group of models. We checked several high weight values at one-one neurons, and we assigned them as, e.g., F, O, or N containing functional groups. It means, these neurons are the responsible ones for different chemical parts as in ref. [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. This bunch of seriated graphs can be used to select models which are better interpretable.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eOur other example is the seriation of the objects (molecules) and their corresponding activity on the neurons as variables. One aspect of neural networks, that the original variables are mapped to the neurons of the hidden layer. This can be used as a dimensional reduction. Depending on the dimensionality of the original data and the number of the neurons, several features of the variables space might remain on the low-dimensional maps. A short investigation of it is shown in the supplementary material for hierarchical clustering. The object activities were calculated as the dot product of the variable vector of a molecule and the weight vector of a neuron in the hidden layer. Figure\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e. shows a case, where 80 test objects are mapped on the neurons and molecules - original variables are shown, as well. One can see in Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003ea-b, that the patch seriation orders the molecules according to their activity on the neurons. This graph might be used to visually detect group of objects and details of the model, e.g., activity, inactivity or redundancy of the models. This object activity -neuron seriation graph resembles somehow to unsupervised maps, e.g., a Kohonen map. The seriation in the original variable space is also successful, but here the variable space is 26 dimensional, while the neuron activity space is only 4 dimensional.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"5. Conclusions","content":"\u003cp\u003eWe developed a new seriation method where our previous idea of using a global merit function based on a local quantity was improved. We defined a local similarity matrix containing the average similarity of neighbouring objects in a local 3-dimensional variable space. These local similarities were put into a global merit function, where the permutations of the object and variable vectors were directed to have both increased local similarities and forming patches of the large similarities.\u003c/p\u003e \u003cp\u003eThe basic idea behind our seriation method is that there are datasets, where different parts of the variables are responsible for the different clusters of the objects. If a set of variables is not concerned in a cluster, they can be easily identified by being outside of the patches. In our method, an overlay contour plot of the local similarity values can be drawn onto the heatmaps of the original data to identify the clusters of the objects and the variables causing the clustering.\u003c/p\u003e \u003cp\u003eBoth the local similarity matrix and the global patch function can generalize into more than two dimensions. We showed some examples of different three-dimensional cases, where the data were arranged according to two variable and one object axes or to one variable and two object axes. Furthermore, the local similarity and the patch function can be generalized for data with missing values or cases, where zero values need not be accounted as responsible ones for clustering.\u003c/p\u003e \u003cp\u003eWe show two simulated datasets (see Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eb), where our patch method is especially effective to discover object clusters and the corresponding variables. Here the traditional seriation methods with non-local distances are mostly in trouble, the ad hoc values of the \u0026ldquo;non-important\u0026rsquo; variables hinder the formation of the clusters. In the case of several public datasets, we found always clearly arranged heatmaps compared to the criss-crossed chequered random ones. Depending on the datasets (material science, compositional, air pollution, reaction kinetic data) clusters of objects and/or variables were always detectable in the seriated heatmaps. In the case of sparse matrices, the patch seriation glue together the non-zero variables.\u003c/p\u003e \u003cp\u003eIn the case of three-dimensional seriation, the interpretation is less straightforward. One needs advanced three-dimensional graphical software or feasible two-dimensional maps to enhance visual perception. If the result is unfolded into two dimensions, the seriated data show periodic changes according to the dimensions merged visually into one axis. In our examples we show the details of three-dimensional air pollution data and retention of different mixtures on different columns.\u003c/p\u003e \u003cp\u003eWe show some examples, how seriation helps to interpret neural network data. For example, we seriated the variable \u0026ndash; hidden neuron weight matrices of different models and there is a striking difference depending on the activation function and the number of the neurons. For example, logistic activation function provided a more interpretable model than the other functions for a flash point dataset, especially at low number of hidden neurons. Also, seriation is a feasible method to detect the neurons responsible for a cluster of objects and to detect inactive parts of the models.\u003c/p\u003e \u003cp\u003eWe think, that seriation is a powerful non-destructive data evaluation or pre-evaluation method. Our special method forms patches of object and clusters. It is effective, if the non-important variables for a given cluster mask the identification possibility according to their variability. Seriation does not replace the different pattern recognition methods, but it at least helps to detect which methods and task might be successful on a dataset.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cul\u003e\n \u003cli\u003eEthics approval and consent to participate: not applicable\u003c/li\u003e\n \u003cli\u003eConsent for publication: We give our consent for the publication of identifiable details, which can include details within the text and figures to be published in the above Journal and Article.\u003c/li\u003e\n \u003cli\u003eAvailability of data and materials: The datasets supporting the conclusions of this article are available in the different repositories referenced one by one in Table 1. \u0026nbsp;The C source code is available at the homepage of the corresponding author [46].\u003c/li\u003e\n \u003cli\u003eCompeting interest: The authors declare that they have no competing interests.\u003c/li\u003e\n \u003cli\u003eFunding: The investigation was partly supported for GT by grant NKFI K-128136.\u003c/li\u003e\n \u003cli\u003eAuthors\u0026rsquo; contributions: All parts of the investigation have been performed with equal load of the authors except the C language seriation software coded by G. T\u0026oacute;th.\u003c/li\u003e\n \u003cli\u003eAcknowledgement: GT thanks the fruitful discussion with prof. Gy\u0026ouml;rgy Tur\u0026aacute;n. The authors acknowledge the datasets for prof. Imre Salma and dr. Anita R\u0026aacute;cz.\u003c/li\u003e\n\u003c/ul\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003ePetrie WM (1899) Flinders Sequences in Prehistoric Remains. The Journal of the Anthropological Institute of Great Britain and Ireland 29: 295-301.\u003c/li\u003e\n\u003cli\u003eBertin J (1981) Graphics and graphic information processing\u003cem\u003e.\u003c/em\u003e Walter de Gruyter, Berlin, Boston. doi:10.1515/9783110854688\u003c/li\u003e\n\u003cli\u003eBrower JC, Kile KM (1988) Seriation of an original data matrix as applied to palaeoecology. Lethaia, 21:79-93. doi:10.1111/j.1502-3931.1988.tb01756.x\u003c/li\u003e\n\u003cli\u003eArabie, P, Hubert, LJ (1996). An overview of combinatorial data analysis. In Arabie P, Hubert LJ, De Soete G (eds.), Clustering and classification. World Scientific, River Edge pp. 5-63.\u003c/li\u003e\n\u003cli\u003eLiiv I (2010) Seriation and Matrix Reordering Methods: An Historical Overview. Stat. Anal. Data Min., 3:70-91. doi:10.1002/sam.10071.\u003c/li\u003e\n\u003cli\u003eVan Gyseghem E, Dejaegher B, Put R, Forlay-Frick P, Elkihel A, Daszykowski M, H\u0026eacute;berger K, Massart DL, Heyden YV (2006) Evaluation of chemometric techniques to select orthogonal chromatographic systems. J. Pharm. Biomed. Anal. 41(1): 141-151. doi: 10.1016/j.jpba.2005.11.007.\u003c/li\u003e\n\u003cli\u003eT\u0026oacute;th G, Szepesv\u0026aacute;ry P (2010) A diagonal measure and a local distance matrix to display relations between objects and variables P. J. Chemometr\u003cem\u003e.\u003c/em\u003e 24\u003cstrong\u003e: \u003c/strong\u003e14-21. doi: 10.1002/cem.1267.\u003c/li\u003e\n\u003cli\u003eSekulića TD, Božinb B, Smolińskic A (2016) Chemometric study of biological activities of 10 aromatic Lamiaceae species\u0026rsquo; essential oils. J. Chemometr. 30\u003cstrong\u003e:\u003c/strong\u003e 188\u0026ndash;196. doi: 10.1002/cem.2786.\u003c/li\u003e\n\u003cli\u003ePigler C, Fogarassy-Vathy \u0026Aacute;, Abonyi J (2016) Scalable co-clustering using a crossing minimization ‒ application to production flow analysis. Act. Polytech. Hung. 13: 209-228. doi:10.12700/APH.13.2.2016.2.12.\u003c/li\u003e\n\u003cli\u003eHammer \u0026Oslash;, Harper D, Ryan P (2001). PAST: Paleontological Statistics Software Package for Education and Data Analysis. Palaeontologia Electronica. 4:1-9.\u003c/li\u003e\n\u003cli\u003eHahsler M, Hornik K, Buchta C (2008) Getting Things in Order: An Introduction to the R Package seriation. \u003cem\u003eJ. Stat. Soft.\u003c/em\u003e 25(3): 1-34. doi: 10.18637/jss.v025.i03\u003c/li\u003e\n\u003cli\u003eR Core Team (2013). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. http://www.R-project.org/ Accessed March 21, 2023\u003c/li\u003e\n\u003cli\u003ePedregosa F (2011) Scikit-learn: Machine Learning in Python, J. Mach. Learn. Res. 12:2825\u0026ndash;2830.\u003c/li\u003e\n\u003cli\u003eHartigan JA (1972) Direct Clustering of a Data Matrix. J. Am. Stat. Assoc. 67: 123\u0026ndash;129.\u003c/li\u003e\n\u003cli\u003eCheng Y, Church GM. (2000) Biclustering of expression data, Proceedings. International Conference on Intelligent Systems for Molecular Biology 8:93-103.\u003c/li\u003e\n\u003cli\u003eStumpfe D, Bajorath J (2011) Similarity searching. WIREs Comput Mol Sci, 1: 260-282. doi:10.1002/wcms.23\u003c/li\u003e\n\u003cli\u003eRosenberg MS (2009) Sequence Alignment: Methods, Models, Concepts, and Strategies. University of California Press: Berkeley, CA.\u003c/li\u003e\n\u003cli\u003eLeese MN, Hughes MJ, Stopford J (1989) The Chemical Composition of Tiles from Bordesley: a Case Study in Data Treatment, in: Rahtz S (ed.), Computer Applications and Quantitative Methods in Archaeology 1989. CAA89 (BAR International Series 548). B.A.R., Oxford, pp. 241-249.\u003c/li\u003e\n\u003cli\u003eBartel HG (1990) Seriation to describe some aspects of generalized evolution and its application in chemical informatics. Systems Analysis Modelling Simulation, 7:557-565.\u003c/li\u003e\n\u003cli\u003eForina M, Lanteri S, Casale M, Cerrato Oliveros M (2007). A new algorithm for seriation and its use in similarity dendrograms. Chemometr. Intell. Lab. Syst. 87:262-274. doi:10.1016/j.chemolab.2007.03.004.\u003c/li\u003e\n\u003cli\u003eT\u0026oacute;th G and Amariamir S (2018) Seriation, the method out of a chemist\u0026apos;s mind. Journal of Chemometrics 32(3-4):e2995. doi:10.1002/cem.2995.\u003c/li\u003e\n\u003cli\u003eMolnar C (2022) Interpretable Machine Learning. A Guide for Making Black Box Models Explainable, 2\u003csup\u003end\u003c/sup\u003e ed. Munich, Germany. Available online: https://christophm.github.io/interpretable-ml-book/ (accessed on 7 June 2022).\u003c/li\u003e\n\u003cli\u003eRGL package https://CRAN.R-project.org/package=rgl last accessed 26th March, 2023\u003c/li\u003e\n\u003cli\u003eKir\u0026aacute;ly P, Kiss R, Kov\u0026aacute;cs D, Ballaj A, T\u0026oacute;th G (2022) The Relevance of Goodness-of-fit, Robustness and Prediction Validation Categories of OECD-QSAR Principles with Respect to Sample Size and Model Type, Mol. Inf. 41:2200072 (15 pages). doi: 10.1002/minf.202200072\u003c/li\u003e\n\u003cli\u003eRuusmann V, Sild S, Maran U (2015) QSAR DataBank repository: open and linked qualitative and quantitative structure\u0026ndash;activity relationship models. J. Cheminf. 7:32. doi:10.1186/s13321-015-0082-6, http://www.qsardb.org.\u003c/li\u003e\n\u003cli\u003eKaggle Inc. http://kaggle.com Accessed 2018 Nov.\u0026ndash;2023 Apr.\u003c/li\u003e\n\u003cli\u003eD. Dua, C. Graff, UCI Machine Learning Repository, Available at http://archive.ics.uci.edu/ml. Irvine, CA: University of California, School of Information and Computer Science, 2019.\u003c/li\u003e\n\u003cli\u003eT\u0026oacute;th G (2023) Benchmark datasets for seriation Mendeley Data, V1, doi - in progress\u003c/li\u003e\n\u003cli\u003eHungarian Air Quality Network, http://www.levegominoseg.hu (Accessed at June 2017) later it has been transported to http://legszennyezettseg.met.hu/\u003c/li\u003e\n\u003cli\u003eTetteh J, Suzuki T, Metcalfe E, Howells S (1999) Quantitative Structure-Property Relationships for the Estimation of Boiling Point and Flash Point Using a Radial Basis Function Neural Network. J. Chem. Inf. Model 39: 491\u0026ndash;507.\u003c/li\u003e\n\u003cli\u003eDrgan V, Zuperl S, Vracko M, Como, F, Novic M (2016) Robust modelling of acute toxicity towards fathead minnow (Pimephales promelas) using counter-propagation artificial neural networks and genetic algorithm. SAR QSAR Environ. Res. 27, 501\u0026ndash;519. doi:10.1080/1062936X.2016.1196388.\u003c/li\u003e\n\u003cli\u003eSaldana DA, Starck L, Mougin P, Rousseau B, Pidol L, Jeuland N, Creton B (2011) Flash Point and Cetane Number Predictions for Fuel Compounds Using Quantitative Structure Property Relationship (QSPR) Methods. Energy Fuels 2011, 25, 3900\u0026ndash;3908. doi: 10.15152/QDB.123\u003c/li\u003e\n\u003cli\u003eSalma I (2023) Daily air pollution and meteorological data Budapest, 2007. Mendeley Data, V1, doi: 10.17632/2mmwv3j4ms.1\u003c/li\u003e\n\u003cli\u003eZiyang He, Maolin Zhang, Haozhe Zhang (2016) Data-driven research on chemical features of Jingdezhen and Longquan celadon by energy dispersive X-ray fluorescence. Ceramics International, 42:5123-5129. doi:10.1016/j.ceramint.2015.12.030.\u003c/li\u003e\n\u003cli\u003eGerman B (1987) Glass Identification dataset, Central Research Establishment, Home Office Forensic Science Service, Aldermaston, Reading, Berkshire RG7 4PN\u003c/li\u003e\n\u003cli\u003eWine recognition dataset, Kaggle Inc. https://www.kaggle.com/brynja/wineuci Accessed 2017 March-2023 Apr.\u003c/li\u003e\n\u003cli\u003eArthur DE, Uzairu A, Mamza P, Stephen AE, Gideon Shallangwa GA. Quantitative structure‐activity and toxicity relationship study of CCRF‐CEM and RPMI 8402 cell line apoptosis with some anticancer compounds. Chem. Data Coll. 2017;7‐8:8‐50. doi:10.1016/j.cdc.2016.12.002.\u003c/li\u003e\n\u003cli\u003eHariprasath R, Jose MT, Vijayalakshmi I, Rajesh A (2016) Determination of natural radioactivity and radiological hazards of sediment sands in Tiruchirappalli district, Tamil Nadu, India. Chem. Data Coll. 2:1‐9. doi: 10.1016/j.cdc.2016.03.001.\u003c/li\u003e\n\u003cli\u003eLang A. Data for: Abraham descriptor A. QsarDB repository, QDB.100. 2012. http://dx.doi.org/10.15152/QDB.100\u003c/li\u003e\n\u003cli\u003eR\u0026aacute;cz A., H\u0026eacute;berger K., Rajk\u0026oacute; R, Elek J (2013) Classification of Hungarian medieval silver coins using X‐ray fluorescent spectroscopy and multivariate data analysis. Heritage Science. 1(1):2 doi:10.1186/2050-7445-1-2.\u003c/li\u003e\n\u003cli\u003eChristie Olav HJ, R\u0026aacute;cz A, Elek J, H\u0026eacute;berger K (2014) Classification and unscrambling a class‐inside‐class situation by object target rotation: Hungarian silver coins of the \u0026Aacute;rp\u0026aacute;d Dynasty, 997‐1301 AD. J. Chemometr. 28:287‐292. doi: 10.1002/cem.2601\u003c/li\u003e\n\u003cli\u003eR\u0026aacute;cz A, H\u0026eacute;berger K, Rajko R, Elek J (2023) Composition data of 257 Hungarian medieval silver coins, Mendeley Data, V1, doi: 10.17632/kbjrfkvcs3.1\u003c/li\u003e\n\u003cli\u003eJuh\u0026aacute;sz G (2015) Reduction of a biodiesel combustion reaction mechanism. BSc thesis Budapest: Institute of Chemistry, Department of Physical Chemistry, E\u0026ouml;tv\u0026ouml;s Lor\u0026aacute;nd University, Budapest\u003c/li\u003e\n\u003cli\u003ePreuer K, Klambauer G, Rippmann F, Hochreiter S, Unterthiner T (2019). Interpretable Deep Learning in Drug Discovery. In: Samek, W., Montavon, G., Vedaldi, A., Hansen, L., M\u0026uuml;ller, KR. (eds) Explainable AI: Interpreting, Explaining and Visualizing Deep Learning. Lecture Notes in Computer Science(), vol 11700. Springer, Cham. doi:10.1007/978-3-030-28954-6_18\u003c/li\u003e\n\u003cli\u003eJim\u0026eacute;nez-Luna J, Grisoni F, Schneider G (2020) Drug discovery with explainable artificial intelligence. Nat Mach Intell 2:573\u0026ndash;584. doi:10.1038/s42256-020-00236-4\u003c/li\u003e\n\u003cli\u003eT\u0026oacute;th G, Homepage of Gergely T\u0026oacute;th, http://tothgergely.web.elte.hu (Last accessed 2023 Apr)\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"journal-of-cheminformatics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"chin","sideBox":"Learn more about [Journal of Cheminformatics](https://jcheminf.biomedcentral.com/)","snPcode":"13321","submissionUrl":"https://submission.nature.com/new-submission/13321/3","title":"Journal of Cheminformatics","twitterHandle":"@jcheminf","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"seriation, data visualization, model interpretation, clustering, neural network model","lastPublishedDoi":"10.21203/rs.3.rs-2780120/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2780120/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eWe developed a new seriation merit function for enhancing the visual information of data matrices. A local similarity matrix is calculated, where the average similarity of a neighbouring objects is calculated in a limited variable space and a global function is constructed to maximize the local similarities and cluster them into patches by simple row and column ordering. The method identifies data clusters in a powerful way, if the similarity of objects is caused by some variables and these variables differ for the distinct clusters. The method can be used in the presence of missing data and also on more than two-dimensional data arrays. We show the feasibility of the method on different data sets: on QSAR, chemical, material science, food science, cheminformatics and environmental data in two- and three-dimensional cases. The method can be used during the development and the interpretation of artificial neural network models by seriating different features of the models. It helps to identify interpretable models by elucidating clusters of objects, variables and hidden layer neurons.\u003c/p\u003e","manuscriptTitle":"Patch seriation to visualize data and model parameters","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-04-10 17:40:46","doi":"10.21203/rs.3.rs-2780120/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2023-04-10T06:05:00+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2023-04-07T07:23:05+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-04-07T07:23:05+00:00","index":"","fulltext":""},{"type":"submitted","content":"Journal of Cheminformatics","date":"2023-04-05T09:12:42+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"journal-of-cheminformatics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"chin","sideBox":"Learn more about [Journal of Cheminformatics](https://jcheminf.biomedcentral.com/)","snPcode":"13321","submissionUrl":"https://submission.nature.com/new-submission/13321/3","title":"Journal of Cheminformatics","twitterHandle":"@jcheminf","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"35af3525-af98-4a5f-8dab-2bb4df0373ac","owner":[],"postedDate":"April 10th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-09-11T15:04:20+00:00","versionOfRecord":{"articleIdentity":"rs-2780120","link":"https://doi.org/10.1186/s13321-023-00757-1","journal":{"identity":"journal-of-cheminformatics","isVorOnly":false,"title":"Journal of Cheminformatics"},"publishedOn":"2023-09-09 15:01:21","publishedOnDateReadable":"September 9th, 2023"},"versionCreatedAt":"2023-04-10 17:40:46","video":"","vorDoi":"10.1186/s13321-023-00757-1","vorDoiUrl":"https://doi.org/10.1186/s13321-023-00757-1","workflowStages":[]},"version":"v1","identity":"rs-2780120","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2780120","identity":"rs-2780120","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0