GROMODEX: Optimisation of GROMACS Performance through a Design of Experiment Approach

preprint OA: closed
📄 Open PDF Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-04 · read from full text

This paper introduces GROMODEX, a tool that optimizes GROMACS molecular dynamics performance by applying a structured Design of Experiments (DoE) strategy rather than manual or one-factor-at-a-time tuning. Using four molecular systems from the NVIDIA GROMACS benchmark suite, the authors evaluate GROMODEX across four computational platforms spanning varied hardware complexity, optimizing configurations for CPU and GPU architectures while aiming to reduce simulation time and improve resource use. They report that performance improvements can be system-dependent and that performance does not necessarily correlate with hardware complexity or cost, with fine, system-specific tuning delivering substantial gains. The study’s main caveat is that benchmarking was performed on only four benchmark systems and across a limited set of platforms rather than a broad range of molecular simulations. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

We introduce GROMODEX, a novel tool designed to optimise GROMACS molecular dynamics (MD) simulations using a structured Design of Experiments (DoE) approach. GROMACS, though efficient, requires extensive tuning of parameters to perform optimally on different hardware and molecular systems. Manual tuning is tedious and prone to errors, especially for high-performance computing (HPC) environments. GROMODEX automates this process, systematically exploring performance parameters to find the best configurations for each system. By using statistical techniques, it evaluates multiple parameters simultaneously, optimising for CPU and GPU architectures, and reducing simulation times while ensuring efficient resource use. We evaluated GROMODEX on four computational platforms spanning a wide range of cost and hardware complexity using four molecular systems from the NVIDIA GROMACS benchmark suite. Quite surprisingly, our results show that the performance does not always correlate with the hardware complexity and cost in a system-dependent manner. Furthermore, a fine system-dependent tuning assures substantial performance gains and time reductions, highlighting GROMODEX’s ability to improve the performance of MD simulations significantly.
Full text 67,663 characters · extracted from preprint-html · click to expand
GROMODEX: Optimisation of GROMACS Performance through a Design of Experiment Approach | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results GROMODEX: Optimisation of GROMACS Performance through a Design of Experiment Approach View ORCID Profile Marco Savioli , View ORCID Profile Paolo Calligari , View ORCID Profile Ugo Locatelli , View ORCID Profile Gianfranco Bocchinfuso doi: https://doi.org/10.1101/2025.10.08.681202 Marco Savioli 1 Department of Mathematics, University of Rome Tor Vergata , Via della Ricerca Scientifica 1, Rome 00133, Italy Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Marco Savioli Paolo Calligari 2 Department of Chemical Sciences and Technology, University of Rome Tor Vergata , Via della Ricerca Scientifica 1, Rome 00133, Italy Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Paolo Calligari Ugo Locatelli 1 Department of Mathematics, University of Rome Tor Vergata , Via della Ricerca Scientifica 1, Rome 00133, Italy Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Ugo Locatelli Gianfranco Bocchinfuso 2 Department of Chemical Sciences and Technology, University of Rome Tor Vergata , Via della Ricerca Scientifica 1, Rome 00133, Italy Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Gianfranco Bocchinfuso For correspondence: gianfranco.bocchinfuso{at}uniroma2.it Abstract Full Text Info/History Metrics Data/Code Preview PDF Abstract We introduce GROMODEX, a novel tool designed to optimise GROMACS molecular dynamics (MD) simulations using a structured Design of Experiments (DoE) approach. GROMACS, though efficient, requires extensive tuning of parameters to perform optimally on different hardware and molecular systems. Manual tuning is tedious and prone to errors, especially for high-performance computing (HPC) environments. GROMODEX automates this process, systematically exploring performance parameters to find the best configurations for each system. By using statistical techniques, it evaluates multiple parameters simultaneously, optimising for CPU and GPU architectures, and reducing simulation times while ensuring efficient resource use. We evaluated GROMODEX on four computational platforms spanning a wide range of cost and hardware complexity using four molecular systems from the NVIDIA GROMACS benchmark suite. Quite surprisingly, our results show that the performance does not always correlate with the hardware complexity and cost in a system-dependent manner. Furthermore, a fine system-dependent tuning assures substantial performance gains and time reductions, highlighting GROMODEX’s ability to improve the performance of MD simulations significantly. 1. INTRODUCTION In the realm of MD simulations, GROMACS 1 , 2 has firmly established itself as one of the most powerful and widely used tools for simulating molecular systems at the atomic level. Initially developed for biomolecular simulations of proteins, lipids 3 , and nucleic acids, it has expanded its capabilities to simulate diverse molecular systems, including polymers 4 , 5 and inorganic materials 6 . However, this versatility comes with the challenge of performance optimisation, especially in high-performance computing (HPC) environments. One of GROMACS’ key strengths lies in its ability to scale efficiently across CPUs and modern GPU-accelerated systems, offering substantial speedups through parallelisation. Nonetheless, achieving optimal performance is complex, as it depends not only on hardware configurations but also on the system complexity and simulation-specific parameters, such as domain decomposition 7 , Particle Mesh Ewald (PME) 8 , 9 settings, as well as the balance between bonded and non-bonded interactions. A poor configuration of these parameters can introduce inefficiencies, and even small parameter adjustments can lead to significant performance changes, making systematic tuning critical. The diversity of HPC architectures adds further complexity. Hybrid CPU-GPU systems require careful workload balancing to avoid bottlenecks, while factors like memory bandwidth, I/O latency, and interconnects can influence simulation throughput. Manual tuning in such contexts is not only time-consuming but also error-prone, often leading to wasted computational resources. This has sparked interest in automated tools 10 that streamline optimisation and ensure efficient resource use. Performance optimisation in MD simulations takes shape of a multidimensional challenge. Historically, optimisation has often been approached using the one–factor–at–a–time (OFAT) 11 method, which alters a single variable while keeping all others constant. While offering some insights, OFAT has serious limitations in complex multifactorial systems like GROMACS where numerous parameters interact. By contrast, the DoE methodology 12 provides a robust and efficient framework for performance optimisation. Unlike OFAT, which assumes independence between parameters, DoE explores multiple parameters simultaneously, revealing complex, nonlinear interactions that often occur in systems like MD simulations. This approach constructs an experimental matrix with predefined levels, reducing the number of trials needed while offering a comprehensive understanding of the parameter space. By analysing the results, DoE not only identifies optimal settings but also uncovers how interdependent variables influence performance, thus enabling more effective configurations. This approach has proven highly successful across various fields, where multifactorial systems are the norm. In the pharmaceutical industry, it is frequently used to optimise process conditions 13 , 14 and analytical method development and validation 15 . It is also extensively used in food industry 14 , 16 , material and design industry 17 and environmental chemistry 18 . In all these domains, DoE consistently outperforms OFAT, particularly in cases where factor interactions are critical to optimisation. The advantages of DoE are particularly compelling in the context of HPC and MD simulations, where the vast number of variables and possible configurations is overwhelming. Despite its success in other domains, DoE remains underutilised in the optimisation of MD simulations 19 . Currently, most GROMACS users rely on manual tuning or ad hoc methods, which are inherently limited by time constraints and the complexity of the parameter space. There is a clear need for a dedicated tool that systematically applies DoE to optimise GROMACS performance across different HPC platforms: a gap that GROMODEX seeks to address. GROMODEX approaches GROMACS performance optimisation as a well-structured experiment, akin to a scientific study conducted not in vitro , but in silico . Just as chemists design-controlled experiments to isolate specific effects, GROMODEX systematically explores the performance landscape of a MD simulation. By applying DoE principles, it assesses a wide range of configurations, identifying not only the best settings for a given simulation on a specific machine but also how those settings may generalise across other systems. This makes GROMODEX a valuable tool for researchers working with diverse molecular systems and HPC platforms. This structured approach is a significant departure from traditional trial-and-error methods. By bringing the rigor of experimental design to MD simulation optimisation, GROMODEX transforms performance tuning into a scientifically grounded, reproducible and data-driven process. Moreover, GROMODEX significantly reduces the time and computational resources needed for performance optimisation. In large-scale projects, even small improvements in efficiency can translate into substantial time and cost savings. By automating the optimisation process, GROMODEX allows researchers to focus on scientific inquiries rather than technical fine-tuning. Lastly, as demand for HPC grows in fields such as computational chemistry, the environmental impact of these platforms is becoming a key concern 20 , 21 . HPC infrastructures have substantial carbon footprints due to their high energy demands 22–24 .Given the global need to reduce carbon emissions, there is increasing focus on making HPC not only faster but also more sustainable. 25 , 26 By enabling GROMACS to use computational resources more efficiently, GROMODEX reduces the overall energy consumed per simulation, helping to lower the environmental impact of MD research. Optimising performance has a direct correlation to reducing energy usage. By maximising throughput and minimising inefficiencies, GROMODEX enhances both speed and energy efficiency. This improvement in computational efficiency has far-reaching implications, enabling more research to be conducted with the same energy input. By moving away from the traditional OFAT approach and adopting a more holistic DoE framework, GROMODEX offers a deeper understanding of complex parameter interactions and provides a more efficient path to optimisation. With the widespread use of GROMODEX, we aim to establish new best practices in the field, making it a valuable resource for HPC users. As MD simulations continue to grow in scale and complexity, GROMODEX will play an increasingly crucial role in maximising the potential of HPC resources while promoting ecological sustainability. 2. METHODS 2.1. GROMACS Setup and computational resources To conduct a comprehensive evaluation of the effectiveness of GROMODEX across a diverse array of hardware infrastructures, MD simulations were carried out on four distinct computational platforms. In all the cases, the performances were registered in the absence of other running processes. These platforms were carefully selected to encompass a broad spectrum of computational capabilities, ensuring a thorough assessment of GROMODEX’s adaptability and performance under varying conditions. The HPC systems involved for this study include: Majorana : A commodity workstation accessible within our facilities, serving as a baseline for moderate-scale simulations. Franklin : A high-performance workstation featuring an updated architecture, representing a more advanced yet still locally accessible computational resource. Ada : A dedicated HPC node within our departmental facilities, providing enhanced processing power and optimised parallel execution. Leonardo : A state-of-the-art HPC system hosted at CINECA, equipped with cutting-edge computational resources and designed to handle large-scale scientific workloads. It ranks fourth on the TOP500 list 27 Reported data for Leonardo refer to performance on a single node. Table 1 summarises the specifications of the computational platforms used in the present study. Their diversity enables a rigorous evaluation of GROMODEX’s performance, particularly on GPU-accelerated architectures. Majorana, Ada, and Franklin, with multiple GPUs, support the analysis of moderate-scale parallel processing, while Leonardo, with four GPUs, allows us to examine large-scale computational efficiency and performance scaling. This selection ensures a comprehensive assessment of hardware-specific optimisations and parallelisation strategies, making our findings broadly applicable across different computational setups. View this table: View inline View popup Download powerpoint Table 1. HPC configuration evaluated in the present study. 2.2. Molecular Systems and Simulation Parameters To ensure the robustness and validity of our performance assessment, we employed GROMACS version 2022.6, a highly optimised and widely used MD package. The selection of this software was motivated by its extensive adoption in computational biophysics, its sophisticated parallelisation capabilities, and its compatibility with a range of hardware architectures. For benchmarking purposes, we selected four molecular systems from the NVIDIA GROMACS benchmark suite 28 , each representing a different level of complexity and system size. These are: Villin : A small and relatively simple protein, often used as a standard benchmark for MD simulations. RNAse Cubic : A mid-sized molecular system that introduces moderate computational demands. Alcohol Dehydrogenase (ADH Dodec) : A larger biomolecular complex, offering a more computationally intensive challenge. Satellite Tobacco Mosaic Virus (STMV) : A highly complex and large-scale molecular system, representing the upper end of computational feasibility within modern HPC environments. By selecting systems with varying sizes and levels of complexity, we ensured that our results would be broadly applicable across the typical range of MD simulations encountered in biophysical research. Table 2 presents an overview of the molecular parameters associated with each of these systems, paralleling those indicated in the NVIDIA benchmark webpage 28 . View this table: View inline View popup Download powerpoint Table 2. Overview of the molecular system used as performance benchmarks. 2.3. DoE Approach The DoE framework provides a rigorous statistical methodology for systematically analysing the effects of multiple variables on a target response. In computational scenarios such as the performance optimisation of MD engines like GROMACS, DoE enables a structured and efficient exploration of the parameter space, elucidating the individual and joint influence of parameters on key performance indicators. In this study, we employed a full factorial design, one of the most exhaustive and informative experimental strategies. This approach evaluates all possible combinations of factor levels, offering a complete and unbiased mapping of the experimental domain. When the number of factors and associated levels is tractable, the full factorial design is particularly advantageous as it allows for precise estimation of both main and interaction effects, without the risk of confounding. The model structure is contingent mainly upon the nature of the response variable, whether categorical or continuous. Incorporating categorical parameters within a full factorial scheme offers multiple benefits in performance tuning contexts: It ensures comprehensive coverage of the configuration space, avoiding bias introduced by heuristic or sequential designs. It enhances the detection of synergistic or antagonistic interactions between parameters, often overlooked by OFAT or fractional factorial designs. It establishes a solid foundation for subsequent predictive modelling, such as linear regression or surrogate models based on machine learning techniques. 2.4. Experimental Factors and Levels The following factors were chosen as variables in the DoE (Design of Experiments) setup, as they are known to significantly affect the performance of GROMACS simulations across different hardware configurations: ntmpi : Number of MPI threads per rank (5 levels for two GPUs, 15 levels for four GPUs) – This factor controls the parallelization strategy for message passing, which is essential for scaling simulations across multiple processors. The levels represent varying degrees of parallelism, allowing for a detailed analysis of the impact on performance as the number of threads per rank increases. pin : CPU core pinning strategy (2 levels) – CPU core pinning refers to the practice of binding specific threads to specific CPU cores to optimise memory locality and reduce inter-core communication overhead. The two levels represent distinct pinning strategies, one where threads are allowed to float between cores, and another where they are pinned to fixed cores, ensuring consistent memory access patterns. gputasks : Distribution of tasks across GPUs (15 levels for two GPUs, 31 levels for four GPUs) – This factor governs how computational tasks are assigned to GPUs in a multi-GPU setup. The levels vary from distributing tasks evenly across GPUs to more complex configurations where certain GPUs are dedicated to specific types of calculations. This distribution is crucial for optimising GPU utilization and minimizing idle times. bonded : Bonded interaction calculation method (2 levels) – This factor influences the method used for calculating bonded interactions in MD simulations. The two levels represent different hardware (CPU or GPU) for executing the bonded interactions. update : Update method for particle positions (2 levels) – The method by which particle positions are updated during the simulation affects the accuracy and efficiency of the calculations. As in the case of the aforementioned parameter, the two levels identify the hardware to set where to execute update and constraints. Although some of these parameters (such as ntmpi and gputasks ) assume numeric values, all factor levels in this experiment are treated as categorical variables, meaning that the analysis does not rely on any assumed ordering or spacing between the levels. Where: k = 5 is the number of factors of the experiment, L i is the number of levels of the i -th factor. For HPC systems equipped with 2 GPUs—such as Majorana, Ada, and Franklin—the levels for each factor were defined as: Yielding: For 4-GPUs HPC systems gputasks factor was extended to reflect the increased parallelism complexity and task distribution possibilities. The levels for each factor were defined as: Thus, the total number of unique simulation configurations is: Therefore, 600 (2480 in the case of Leonardo) distinct simulation configurations were systematically generated and executed across different computational platforms, enabling a comprehensive analysis of how performance scales with respect to both software-level parameters and hardware-specific characteristics. By its nature, a full factorial DoE systematically explores the performance space of the selected parameters. This approach ensures exhaustive coverage of all possible combinations of factor levels, facilitating a comprehensive analysis of their interactions and effects. However, the full factorial DoE generates all possible combinations of the levels of the independent factors, without considering any logical or technical constraints among configurations. In high-performance computing environments such as GROMACS, this can lead to execution failures, as certain combinations may be inconsistent or invalid with respect to the underlying hardware or the internal logic of the simulation software. For instance, combinations that assign more GPU tasks than physically available devices, or incompatible settings of the -nb, -pme, and -npme options, can result in runtime errors. A representative error message includes: “There were N GPU tasks assigned on node A, but M GPU tasks were identified, and these must match.” In practice, the full factorial design yielded both successful and failed configurations, depending on the compatibility of parameter combinations with the hardware and software constraints. The number of valid and invalid configurations varied across systems and benchmark cases. For the Villin system, 88 configurations completed successfully and 512 failed on Majorana, Franklin, and Ada, while 152 were successful and 2328 failed on Leonardo. For the RNase and ADH Dodecamer systems, the same trend was observed: 120 successful and 480 failed configurations on Majorana, Franklin, and Ada, and 200 successful versus 2,280 failed on Leonardo. Finally, for the STMV system, 120 configurations were successful and 480 failed on Majorana, Franklin, and Ada, while 232 completed successfully and 2248 failed on Leonardo. These results emphasize how, despite the exhaustive exploration provided by a full factorial DoE, only a small subset of the total combinations corresponds to valid configurations. The large proportion of failed runs reflects the presence of strong logical interdependencies among parameters in HPC environments such as GROMACS, where constraints related to GPU task allocation and internal domain decomposition often restrict the feasible configuration space. To address this, combinations violating hardware constraints or software-specific logical requirements were identified and excluded from further analysis. This filtering step ensures that only feasible and meaningful configurations contribute to the performance evaluation, while preserving the methodological rigor of the factorial design. 2.5. Methodology: Automation of DoE and MD Simulations All MD simulations were orchestrated through a custom workflow developed in Python 3.10. This environment allowed for seamless integration of experimental design, simulation execution, performance monitoring, and data post-processing. The workflow was structured to ensure full reproducibility and scalability across different GPU configurations. To efficiently generate and manage the design space, we employed the pyDOE3 library (v1.0.0), which facilitated the construction of a full factorial DoE matrix. Data handling and logging operations were managed using pandas (v1.5.3). Notably, the settings for PME computation—namely -npme , -pme , and -pmefft —were fixed to values that consistently offloaded this calculation to the GPU. This choice was informed by preliminary screening experiments (data not shown), which demonstrated that GPU-based PME calculations outperformed CPU-based alternatives across all configurations tested. The finalized DoE matrix included 256 parameter combinations, which were exported as a CSV file ( gmx_fullfact_doe.csv ) to ensure transparency and traceability. To guarantee consistency across runs, each simulation was initialized using unified topology ( .top ), structure ( .gro ), and parameter ( .mdp ) files, generating via gmx grompp the same MD configuration file ( .tpr ) for all the runs. The simulation runs were executed using subprocess.run() within Python, allowing fine-grained control over each call to gmx mdrun . This approach provided robust handling of standard output, standard error, and runtime exceptions, all of which were logged to a central file ( doe.log ) for auditing and debugging purposes. The execution command for each simulation followed a common template: gmx mdrun -ntmpi X -pin Y -gputasks Z -bonded A -update B -nb gpu -pme gpu -pmefft gpu -npme 1 -nsteps 25000 where the parameters X , Y , Z , A , and B varied according to the experimental design. The number of MD steps was fixed at 25000 to standardise throughput measurement, which was extracted from the GROMACS log file by parsing the line containing the “Performance” keyword. This throughput value (in ns/day) was appended in real time to the DoE matrix, allowing immediate performance comparison across configurations. Upon completion of each simulation, temporary files were removed to minimize storage overhead, preserving only performance logs and GPU usage files ( gpu_usage_runX.csv ). The final results were consolidated into a structured dataset ( gmx_fullfact_doe_results.csv ) combining all experimental parameters with their corresponding throughput values. Although the simulations did not rely on stochastic components, the use of identical input files ensured deterministic behaviour across executions. The total wall-clock time for each simulation was recorded and used as a reference for workflow scalability. 2.6. Data Analysis Given that our response consists of continuous performance metrics, analysis of variance (ANOVA) is the preferred inferential tool for evaluating the significance of each factor and interaction term 29 , 30 . ANOVA was employed to quantify the contribution of each experimental factor to the observed variability of the response variable and to assess whether the apparent effects were statistically distinguishable from background noise. The rationale behind this approach lies in the decomposition of the total variability of the system into systematic sources, attributable to controlled factors and their interactions, and random error, attributable to uncontrolled perturbations and measurement uncertainty. Formally, the model can be expressed as: where y ijk represents the observed response under the i -th level of factor A , the j -th level of factor B, and the k -th replicate. Here, μ denotes the overall mean, α i and β j correspond to the main effects of factors A and B , (αβ) ij captures the interaction effect, and ε ijk denotes the residual error assumed to be independently and normally distributed with constant variance. In the general DoE setting, the model can be extended to accommodate additional factors and higher-order interactions. The inferential procedure of ANOVA relies on partitioning the total sum of squares ( SSQ tot ) into additive components associated with each factor, their interactions, and the residual error. This decomposition is written as: with each term subsequently normalized by its respective degrees of freedom to yield the mean square ( MS ). The central test statistic, the F-ratio, is computed for each factor or interaction as: where df i and df E denote the degrees of freedom of the factor of interest and the error term, respectively. Under the null hypothesis that the factor has no effect, the distribution of F i follows an F-distribution with ( df i , df E ) degrees of freedom, corresponding to the factor of interest and the residual error, respectively. A comparison between the observed F-statistic and the theoretical F-distribution enables the calculation of a p-value, which quantifies the probability of obtaining such an extreme ratio under the null hypothesis. Factors yielding p-values below the pre-specified significance threshold (typically α = 0.05) are considered statistically significant contributors to the response variability. More formally, given an observed value F obs i for the i -th factor, the p-value is defined as: where F dfi , dfE denotes a random variable following the F-distribution. This probability quantifies the degree of compatibility between the observed data and the null model of no effect. The decision-making process in ANOVA is thus anchored in the comparison of the calculated p-value with a pre-established significance level α. If p < α, the null hypothesis is rejected, and the factor is deemed to exert a statistically significant influence on the response variable. Conversely, if p ≥ α, the null hypothesis cannot be rejected, and the data do not provide sufficient evidence to attribute systematic variation to the factor. It is important to underscore that the p-value does not quantify the probability that the null hypothesis is true; rather, it expresses the extremeness of the observed test statistic relative to the null distribution. In the specific setting of DoE, where multiple factors and interactions are simultaneously tested, the p-value serves as a critical criterion for discriminating between genuine effects and random fluctuations. When interpreted alongside effect sizes and confidence intervals, it ensures that the identification of significant parameters is both statistically rigorous and scientifically meaningful. Within the DoE context, this ANOVA-based procedure is of critical importance for disentangling the relative importance of main effects and interaction terms. By systematically evaluating each source of variation, ANOVA provides not only a rigorous statistical test of factor relevance but also a foundation for model reduction and optimization strategies. This methodological framework ensures that the identification of significant parameters is not confounded by random fluctuations, thereby enhancing the robustness of subsequent predictive modelling. Overall, the integration of ANOVA into the DoE workflow enables the systematic exploration of the experimental space with statistical rigor. It transforms raw measurements into actionable insights, establishing whether variations in the studied response are the result of deliberate manipulations of input factors or merely stochastic noise. By doing so, it provides the methodological backbone for optimising factor settings and drawing robust scientific conclusions. 3. RESULTS AND DISCUSSION The results of our study on the development and application of GROMODEX are presented below. For clarity of exposition, we present the results from a molecular-system-centric perspective. Readers are encouraged to refer to the Supplementary Information (SI), where the dataset is organized to vertically isolate each HPC system. This structure provides a clearer framework for analysing how each HPC system adapts to molecular systems of different sizes. 3.1 Performance Analysis 3.1.1 Villin Performance analysis of the simulations revealed substantial variability across runs ( Figure 1 ). The dataset contains a significant proportion of executions yielding zero performance, indicating that these simulations did not execute successfully. The underlying causes can be attributed both to inconsistencies in the input files provided to GROMACS (see section 2.4 ) and to the inability to apply a domain decomposition strategy involving a large number of MPI ranks on a system of limited size. In such cases, the overhead associated with managing excessive parallel tasks exceeds the computational benefits, resulting in aborted runs. Download figure Open in new tab Figure 1. Average performance trend (three replicas for each HPC system) of Villin as a function of parameter settings. (A) Majorana, (B) Franklin, (C) Ada and (D) Leonardo. Download figure Open in new tab Figure 2. Average performance trend (three replicas for each HPC system) of RNAse Cubic as a function of parameter settings. (A) Majorana, (B) Franklin, (C) Ada and (D) Leonardo. Download figure Open in new tab Figure 3. Average performance trend (three replicas for each HPC system) of ADH Dodec as a function of parameter settings. (A) Majorana, (B) Franklin, (C) Ada and (D) Leonardo. Among the successful executions, performance data tend to cluster around specific index values, suggesting a discrete set of configurations with stable computational behaviour. Worth of note, the emergence of a sawtooth-like pattern in the overall trend; a preliminary visual inspection of this behaviour suggests that optimal performance is achieved at configurations with a relatively low number of MPI ranks, which allow for a more balanced load distribution and reduced communication overhead between processes. Table 3 summarises the parameters to which the highest performance corresponds for the Villin system across the four HPC configurations, revealing several non-trivial trends that cannot be directly ascribed to the underlying hardware characteristics. Leonardo achieves the highest simulation speed at 1605.220 ns/day, followed by Majorana (1471.406 ns/day), Franklin (1287.939 ns/day), and Ada (1139.921 ns/day). Despite its comparatively modest hardware configuration, Majorana outperforms both Franklin and Ada, securing second place in overall performance. This result can be partially explained by the hardware specifications: Leonardo is equipped with four custom NVIDIA A100 GPUs, offering not only higher memory bandwidth but also superior computational capabilities. In contrast, Majorana’s better-than-expected performance—despite having fewer and less powerful GPUs—may be attributed to its CPU’s high base frequency, which likely compensates through faster execution of CPU-bound tasks or more efficient GPU synchronisation. Leonardo’s CPU, while offering the highest thread and memory bandwidth potential, operates at a relatively moderate base clock ( Table 1 ). Notably, Leonardo was the only system with pin set to off, suggesting that despite the lack of core binding, the GPUs resources compensate for any potential inefficiencies due to thread migrations. Conversely, the other systems-maintained pin on, likely to maximize CPU affinity and memory locality. View this table: View inline View popup Download powerpoint Table 3. Highest performance parameter settings for Villin. The gputasks distribution further highlights architectural differences. On Majorana and Franklin both GPU were assigned to perform a single task each (i.e., 01). The same happened for Leonardo, which adopt a broader distribution of gputasks (i.e., 0123). In particular, the common feature of these two configurations lies in the fact that they both dedicate exclusively one of the available GPUs to the calculation of PMEs; in particular, in the GROMACS’ syntax, this is identified by the ID of the GPU present as the last value in the string (i.e., 1 for Majorana and Franklin, 3 for Leonardo). Conversely, for Ada the optimal performance was achieved with a 0011 gputasks configuration. The allocation of bonded and update tasks is particularly interesting: while Majorana and Ada kept both tasks on the CPU, Franklin and Leonardo - which use the same type of GPU - instead offloaded the bonded calculations to the GPU. However, for Leonardo, a configuration with the bonded ones performed on the CPU still achieves near-peak performance. This strategy, especially in the case of Franklin, likely contributed significantly to reducing the computational load on the CPU, thus improving the overall throughput. 3.1.2 RNAse Cubic For the RNAse Cubic system, which is computationally more demanding, the performance results are noticeably more abundant in terms of successfully completed runs. This improvement is likely attributable to GROMACS’ ability to access to a higher number of MPI ranks in this context, thus expand the domain decomposition opportunities. The increased computational load appears to better justify the communication overhead, making parallelisation more efficient. The highest performance values are observed in the mid-to-low region of the plots, although this peak emerges slightly later than in the case of the Villin system. In fact, as the dimensionality and complexity of the molecular system increase, HPC architectures tend to achieve optimal performance with different parameter settings, suggesting that alternative spatial decomposition strategies become more effective under these conditions. Moreover, the performance pattern for RNAse Cubic exhibits considerably greater uniformity, with fewer fluctuations and anomalies. Nonetheless, a consistent trend across the four HPC systems is the presence of a pronounced performance degradation at run index 445 for Majorana, Franklin, and Ada, and at index 1527 for Leonardo. This behaviour is associated with the configuration ntmpi : 10, pin : off, gputasks : 0000000001, bonded : cpu, and update : gpu, although this configuration does not necessarily correspond to the absolute minimum performance within the dataset. As for the parameters with the best performance, Leonardo again leads with 828.215 ns/day, followed by Franklin (715.195 ns/day), Majorana (648.367 ns/day), and Ada (627.306 ns/day), as can be seen from Table 4 . The relative performance advantage of Majorana is less pronounced here with respect to the Villin case, possibly indicating a scaling limitation when system size increases and communication overhead becomes more relevant. View this table: View inline View popup Download powerpoint Table 4. Highest performance parameter settings for RNAse Cubic. Regarding ntmpi , Majorana required more MPI ranks than Ada and Franklin, likely due to its weaker hardware. All systems except Leonardo used pin set on . The fact that Leonardo again achieved the best performance with pin set off suggests that in this particular configuration binding cores can sometimes become a bottleneck rather than a benefit, especially when GPU communication and synchronization dominate runtime. The gputasks mappings further explain the performance patterns: Leonardo distributed gputasks across three GPUs (000111222). Majorana, Franklin, and Ada limited the load to two GPUs (000111), with respect to Leonardo. Interestingly, Majorana adopted a non-symmetrical distribution of gputasks , preferring to allocate more computing resources to PME on GPU #1. In all cases bonded and update tasks were computed on the CPU, suggesting a consistent design strategy to balance the CPU–GPU workload. 3.1.3 ADH Dodec Performance analysis for the ADH Dodec system, shown in Figure 3 Table 5 , exhibits a complex and heterogeneous behaviour, reflecting the computational demands of this larger biomolecular system. In contrast to the more compact systems previously analysed, ADH Dodec displays a broader dynamic range in performance values, as well as wider sensitivity to the configuration parameters adopted in each run. The benchmarking results for the ADH Dodec system, a significantly larger molecular system, illustrate once again the critical role of hardware architecture and task distribution in determining GROMACS performance ( Table 5 ). View this table: View inline View popup Download powerpoint Table 5. Highest performance parameter settings for ADH Dodec. Leonardo outperforms all other platforms with a simulation speed of 340.213 ns/day, followed by Franklin (276.369 ns/day) and Ada (217.668 ns/day), while Majorana suffered of poor computational capabilities (181.478 ns/day), almost losing half the performance compared to Leonardo. Regarding parallelization, Leonardo utilized 16 MPI ranks, almost three time the number used on Franklin and Ada (6 each) and twice that of Majorana (8). The higher ntmpi setting on Leonardo facilitated better domain decomposition and load balancing across the multiple GPUs and CPU cores, essential for efficiently handling such a large molecular system. The pin parameter was uniformly set to on across all platforms for this benchmark, reflecting the necessity of strict CPU-GPU affinity to minimize memory access latency and optimise interconnect usage under heavier computational loads. A deeper look into the gputasks assignment showed substantial common features. At this level of dimensionality of molecular systems, a non-symmetrical allocation of GPU tasks was most effective. In fact, Majorana, Franklin and Ada have adopted configurations for this parameter that leave a large part of GPU #1 dedicated to the calculation of PMEs, which, at this level of dimensionality, constitute an important part of the computational load. Leonardo, on the strength of its configuration, distributed the workload across all four GPUs with finer granularity, dedicating one GPU (ID: 3) to exclusively PME computation. The assignment of bonded and update tasks also presented an interesting deviation. While Majorana, Franklin, and Ada offloaded bonded interactions to the GPU to reduce CPU bottlenecks, Leonardo processed bonded interactions on the CPU, without penalizing GPU workload. 3.1.4 STMV The performance evaluation of the STMV system highlights a progressive and well-structured pattern. As a considerably large and complex biomolecular system, STMV presents consistent trends that reflect how hardware is challenged under such intensive demand of computational resources. Figure 4A displays a well-defined staircase-like increase in performance, with three clearly distinguishable regions of growth, suggesting strong saturation in computational scaling. These plateaus, less identifiable in other molecular systems under consideration, fall at a specific combination of the bonded and update levels that unlock higher throughput. The steep performance increase around runs 450 indicates a jump in efficiency, due to progressive computational offload of the GPUs at those configurations. Download figure Open in new tab Figure 4. Average performance trend (three replicas for each HPC system) of STMV as a function of parameter settings. (A) Majorana, (B) Franklin, (C) Ada and (D) Leonardo. Both Franklin and Ada exhibit a trend of gradual performance improvement characterized by multiple local peaks ( Figure 4B and 4C ). This indicates a nuanced scaling behaviour in which many configurations yield similar, though subtly distinct, performance scores. In Franklin’s case, the lack of sharp drops may point to better load balancing or greater system stability, while Ada shows more noticeable variability across individual runs. This variability suggests that, despite the overall improvement, performance remains sensitive to specific configuration choices—highlighting the importance of fine-tuning parallel parameters even within comparable scaling regions. Leonardo in Figure 4D , covering a broader range of runs, reveals a consistent upward trend, confirming a strong sensitivity in performance over a wide range of configurations. This case study corroborates more than ever that, despite mere computing power, the sub-optimal choice of simulation parameters can lead to a drastic degradation of simulation performance, especially in contexts where the molecular system is particularly challenging. Leonardo once again demonstrated superior performance with a simulation speed of 37.197 ns/day, followed by Franklin (26.687 ns/day), Ada (19.745 ns/day), and, considerably behind, Majorana (10.740 ns/day). The results are once again summarised in Table 6 . View this table: View inline View popup Download powerpoint Table 6. Highest performance parameter settings for STMV. Notably, the performance gap widens significantly at this scale, particularly between Majorana and the other systems. Majorana’s relatively modest hardware proved inadequate to sustain competitive throughput. The lower ntmpi on Majorana clearly constrained its ability to decompose the large system into enough parallel tasks, severely bottlenecking performance. By contrast, the higher ntmpi on Franklin, Ada, and Leonardo enabled a much finer domain decomposition, essential for simulations with such a massive number of atoms. The configuration of CPU affinity represents a key differentiating factor in system performance across platforms. On systems such as Franklin, Ada, and Leonardo, CPU affinity was explicitly enabled ( pin on ), a choice that facilitated optimised memory access patterns and reduced cache contention, thereby contributing to overall computational efficiency. In contrast, Majorana operated without CPU pinning ( pin off ), a setting that further magnified its performance limitations under intensive computational workloads. In terms of gputasks assignment, all systems employed non-symmetrical distribution strategies akin to those observed in ADH Dodec. Franklin and Ada deployed a mapping scheme spanning ten GPU tasks (0000001111), reflecting a balanced distribution over two physical GPUs. Leonardo advanced this approach with a more granular and extended mapping scheme (000011112223), with one GPU dedicated to PME calculation only. In contrast, Majorana, constrained by a limited computational capacity, relied on a simpler allocation pattern (0011), lacking the granularity and scalability required for efficiently support the simulation effort. Interestingly, for STMV, all systems offloaded both bonded and update calculations to the GPU, a reflection of the critical need to minimize CPU overhead at such a scale; even Leonardo kept bonded interactions on the GPU here, recognizing that CPU processing would be inefficient compared to the vast parallel capabilities of its GPUs for such a large problem size. 3.1.5 Right-sizing Before proceeding with the statistical analysis, it is useful to consolidate the performance results to highlight the underlying trends across architectures and workloads. This section serves to reframe the data from a broader perspective, emphasizing how computational efficiency emerges from the balance between system size and hardware characteristics. Figure 5 summarises the performance obtained across the four HPC architectures for the four molecular systems. Each curve represents one system, showing how performance scales with the architectural configuration. Download figure Open in new tab Figure 5. Performance comparison across four HPC architectures for the four molecular systems considered. Each curve represents a distinct system, illustrating how computational efficiency scales with architectural characteristics. The figure highlights the right-sizing effect, where optimal performance emerges from the alignment between workload scale and hardware resources. Majorana outperforms larger HPC systems for small-scale molecular dynamics, particularly in the Villin case. This behaviour can be attributed to the limited degree of parallelism involved, which allows the i9 CPU to exploit its higher clock frequency ( Table 1 ) and achieve superior per-thread efficiency. In contrast, the other systems (RNAse, ADH, STMV) exhibit improved scalability on architectures with broader parallel resources, where communication overhead becomes less dominant relative to the total computational workload. These findings indicate that no universally optimal configuration exists; rather, performance depends on a dynamic interplay between workload characteristics, system size, and architectural design. This right-sizing perspective provides a conceptual bridge to the following ANOVA, which quantifies the statistical significance of the factors influencing performance. 3.2 Main Effects and Interactions Two-way ANOVA results ( Figure 6Figure 6 ) across the four high-performance computing (HPC) architectures demonstrate a consistent pattern of both dominant main effects and critical interaction terms influencing computational performance across workloads. Despite architectural and workload-specific variability, the gputasks and the ntmpi systematically emerge as the most impactful parameters shaping execution efficiency. On all platforms, these two factors exhibit exceptionally low p-values (p ≪ 0.001), achieving statistical significance and accounting for a substantial portion of the variance in all simulation performance. This consistency highlights the universal importance of GPU workload partitioning, confirming that GPU-level parallelism remains a core determinant of performance across both cutting-edge and legacy architectures. ntmpi often surpasses gputasks in significance, particularly on architectures such as Leonardo and Franklin, where MPI process distribution critically shapes data locality and interprocess communication overhead. In particular, ntmpi achieves particular relevance in those situations where exceptional sensitivity of these architectures to the granularity of parallel decomposition ( Table 5 and Table 6 ). These results not only reaffirm the central role of hardware-aware parallel configuration – i.e., configurations that explicitly consider the system’s architecture such as core counts, memory hierarchy, and CPU-GPU topology – but also stress the need to tailor process-level and thread-level resource allocation strategies to system-specific architectural constraints. Download figure Open in new tab Figure 6. Two-way ANOVA results: the higher the bar, the more relevant the factor – or combination of two factors – is. From top: Villin, RNAse Cubic, ADH Dodec, STMV. From left: Majorana, Franklin, Ada, Leonardo. Beyond individual factor effects, interaction terms provide key insight into the complex, non-linear nature of performance scaling in heterogeneous systems. Among these, the ntmpi:gputasks interaction repeatedly demonstrates significance across all systems, reinforcing that computational performance cannot be optimised by tuning MPI or GPU parameters in isolation. Instead, a synergistic multidimensional (or multifactorial) adjustment is necessary; this interaction captures a critical performance contour that reflects the load balancing and task scheduling challenges intrinsic to modern HPC environments. Moreover, interactions involving variables that are excluded from the final models as individual factors—specifically pin , update , and bonded —consistently retain statistical significance in combination with ntmpi or gputasks . While their main effects yield p-values well above the 0.05 threshold, rendering them individually non-significant, these variables nonetheless participate in key second-order interactions that modulate performance outcomes. For example, the pin:update interaction displays significant effects on multiple circumstances, suggesting that memory affinity settings influence performance predominantly through their effect on update tasks or data locality during time integration steps. Similarly, bonded:update and gputasks:bonded achieve significance, illustrating that short-range interactions are sensitive to resource allocation primarily in the context of other configuration parameters. These observations underscore an important methodological insight in performance modelling: statistical insignificance of a factor in isolation does not equate to practical irrelevance. Instead, the exclusion of such factors from the main effects may suggest evidence of conditional importance 31 —a phenomenon in which the influence of a parameter becomes apparent only when coupled with others. This is especially relevant in high-dimensional configuration spaces, such as HPC performance tuning, where emergent properties often result from the interplay between multiple control variables. The relevance of these interaction terms varies not only across HPC platforms but also across workloads. For instance, pin , despite lacking individual significance, exhibits crucial interactions, like ntmpi:pin and pin:bonded , which affect memory access and CPU-GPU synchronisation, confirming the conditional relevance. Additionally, bonded and update , although often contributing less to overall variance when considered individually, become more relevant when considered in combination with ntmpi or gputasks . This finding is evident in statistically significant interactions such as ntmpi:bonded and gputasks:bonded on Leonardo, which indicate that the performance sensitivity of bonded calculations is largely determined by their GPU distribution. Across all systems, the explanatory power of the models is uniformly high, with residual variances remaining consistently low. This indicates that the selected main and interaction effects capture the majority of performance determinants and that the factorial design provides a robust framework for exploring configuration sensitivity. However, the findings also reveal that no single configuration strategy is universally optimal across architectures. Instead, the relative importance of factors and interactions shifts depending on HPC-molecular system pair characteristics. 4. CONCLUSIONS This study presented the development and application of GROMODEX, a robust tool for systematically optimising GROMACS MD simulations through a DoE approach. By combining rigorous full-factorial experimental design, high-throughput automation and ANOVA, GROMODEX facilitates the exploration of complex parameter spaces to identify configurations that maximize simulation performance. The results underscore the versatility of the tool and its potential for advancing MD research. It is important to note that the configurations explored in this study were primarily assessed in terms of computational performance and successful execution, rather than long-term physical stability or numerical reproducibility of the molecular dynamics trajectories. Certain parameter combinations—especially those affecting domain decomposition, thread affinity, and GPU task allocation—may influence not only performance but also simulation stability and numerical accuracy. Therefore, the results presented here should be interpreted as indicative of computational efficiency within the tested setups, while additional verification would be required to ensure full physical consistency of the simulations across different configurations. The investigation of four molecular systems across four computational architectures (Majorana, Franklin, Ada and Leonardo) revealed critical insights into the interplay between system– specific requirements and hardware capabilities. The diversity of molecular systems ensured the generalizability of the adopted protocol, while the variation in computational platforms provided a nuanced understanding of architecture-specific optimisation strategies. A key outcome of this work was the identification of optimal parameter configurations for each molecular system and architecture. Factors such as ntmpi and gputasks were shown to have significant impacts on simulation throughput, as evidenced by the two-way ANOVA. These findings highlight the necessity of tailored optimisation approaches that account for both molecular complexity and hardware specifications. Beyond the identification of optimal configurations, the study provided a deeper understanding of the parameter interdependencies that influence performance. The interaction effects revealed by the ANOVA collectively demonstrate that performance optimisation on modern HPC systems cannot rely solely on analysis of primary factors. Instead, attention must be paid to the conditional and often non-additive nature of parameter interactions. While ntmpi and gputasks universally dominate, their impact is magnified or modulated by interactions with seemingly secondary variables such as pin , update , and bonded . Therefore, a comprehensive, interaction-aware approach is essential to unlock the full potential of computational resources. Neglecting such interactions risks oversimplifying the optimisation process and may result in configurations that are suboptimal or misleadingly efficient only under restricted conditions. These findings advocate for the integration of multifactorial analysis methods into performance tuning workflows, promoting a deeper understanding of system behaviour and enabling more informed decisions in production-scale simulations. The adoption of tools like GROMODEX systematic and scalable approach addresses a critical need in the field, enabling researchers to achieve superior performance across diverse systems and architectures. The insights gained from this study not only validate the efficacy of GROMODEX but also provide a foundation for future developments in simulation optimisation, paving the way for more efficient and accurate MD research. 5. DATA AVAILABILITY STATEMENT The GROMODEX source code is available at: https://github.com/MarcoSavioli/GROMODEX.git . ACKNOWLEDGMENTS This research was supported by the Italian Research Center on High Performance Computing, Big Data and Quantum Computing (CN1-HPC) and the Italian SuperComputing Resource Allocation (ISCRA code: HP10C31W4Y). Footnotes https://github.com/MarcoSavioli/GROMODEX.git REFERENCES (1). ↵ Bekker , H. ; Berendsen , H. J. C. ; Dijkstra , E. J. ; Achterop , S. ; van Drunen , R. ; van der Spoel , D. ; Sijbers , H. ; Keegstra , H. ; Renardus , M. K. R. Gromacs: A Parallel Computer for Molecular Dynamics Simulations; DeGroot, R. A., Nadrchal, J., Eds.; World Scientific Publishing: Singapore , 1993 ; pp 252 – 256 . (2). ↵ Berendsen , H. J. C. ; Van Der Spoel , D. ; van Drunen , R. GROMACS: A Message-Passing Parallel Molecular Dynamics Implementation PROGRAM SUMMARY Title of Program: GROMACS Version 1.0 . Comput Phys Commun 1995 , 91 ( 1–3 ), 43 – 56 . doi: 10.1016/0010-4655(95)00042-E . OpenUrl CrossRef (3). ↵ Jiang , F. Y. ; Bouret , Y. ; Kindt , J. T . Molecular Dynamics Simulations of the Lipid Bilayer Edge . Biophys J 2004 , 87 ( 1 ), 182 – 192 . doi: 10.1529/biophysj.103.031054 . OpenUrl CrossRef PubMed (4). ↵ Liu , J. ; Lin , H. ; Li , X . GMXPolymer: A Generated Polymerization Algorithm Based on GROMACS . J Mol Model 2024 , 30 ( 9 ). doi: 10.1007/s00894-024-06119-4 . OpenUrl CrossRef (5). ↵ Degiacomi , M. T. ; Erastova , V. ; Wilson , M. R . Easy Creation of Polymeric Systems for Molecular Dynamics with Assemble! Comput Phys Commun 2016 , 202 , 304 – 309 . doi: 10.1016/j.cpc.2015.12.026 . OpenUrl CrossRef (6). ↵ Aragones , J. L. ; Valeriani , C. ; Vega , C . Note: Free Energy Calculations for Atomic Solids through the Einstein Crystal/Molecule Methodology Using GROMACS and LAMMPS . J Chem Phys 2012 , 137 ( 14 ). doi: 10.1063/1.4758700 . OpenUrl CrossRef PubMed (7). ↵ Hess , B. ; Kutzner , C. ; Van Der Spoel , D. ; Lindahl , E . GROMACS 4: Algorithms for Highly Efficient, Load-Balanced, and Scalable Molecular Simulation . J Chem Theory Comput 2008 , 4 ( 3 ), 435 – 447 . doi: 10.1021/ct700301q . OpenUrl CrossRef PubMed Web of Science (8). ↵ Darden , T. ; York , D. ; Pedersen , L . Particle Mesh Ewald: An N·log(N) Method for Ewald Sums in Large Systems . J Chem Phys 1993 , 98 ( 12 ), 10089 – 10092 . doi: 10.1063/1.464397 . OpenUrl CrossRef Web of Science (9). ↵ Essmann , U. ; Perera , L. ; Berkowitz , M. L. ; Darden , T. ; Lee , H. ; Pedersen , L. G . A Smooth Particle Mesh Ewald Method . J Chem Phys 1995 , 103 ( 19 ), 8577 – 8593 . doi: 10.1063/1.470117 . OpenUrl CrossRef Web of Science (10). ↵ Gecht , M. ; Siggel , M. ; Linke , M. ; Hummer , G. ; Köfinger , J . MDBenchmark: A Toolkit to Optimize the Performance of Molecular Dynamics Simulations . J Chem Phys 2020 , 153 ( 14 ). doi: 10.1063/5.0019045 . OpenUrl CrossRef PubMed (11). ↵ Frey , D. D. ; Engelhardt , F. ; Greitzer , E. M . A Role for “One-Factor-at-a-Time” Experimentation in Parameter Design . Res Eng Des 2003 , 14 ( 2 ), 65 – 74 . doi: 10.1007/s00163-002-0026-9 . OpenUrl CrossRef (12). ↵ Fisher , R. A. The Design of Experiments; Hafner Press: New York , 1935 . (13). ↵ Politis , S. N. ; Colombo , P. ; Colombo , G. ; Rekkas , D. M . Design of Experiments (DoE) in Pharmaceutical Development . Drug Dev Ind Pharm 2017 , 43 ( 6 ), 889 – 901 . doi: 10.1080/03639045.2017.1291672 . OpenUrl CrossRef PubMed (14). ↵ Paulo , F. ; Santos , L . Design of Experiments for Microencapsulation Applications: A Review . Mater Sci Eng C Mater Biol Appl 2017 , 77 , 1327 – 1340 . doi: 10.1016/j.msec.2017.03.219 . OpenUrl CrossRef PubMed (15). ↵ Tome , T. ; Žigart , N. ; Časar , Z. ; Obreza , A. Development and Optimization of Liquid Chromatography Analytical Methods by Using AQbD Principles: Overview and Recent Advances . Org Process Res Dev 2019 , 23 ( 9 ), 1784 – 1802 . doi: 10.1021/acs.oprd.9b00238 . OpenUrl CrossRef (16). ↵ Yu , P. ; Low , M. Y. ; Zhou , W . Design of Experiments and Regression Modelling in Food Flavour and Sensory Analysis: A Review . Trends Food Sci Technol 2018 , 71 , 202 – 215 . doi: 10.1016/j.tifs.2017.11.013 . OpenUrl CrossRef (17). ↵ Tryland , T. ; Hopperstad , O. S. ; Langseth , M . Design of Experiments to Identify Material Properties . Mater Des 2000 , 21 ( 5 ), 477 – 492 . doi: 10.1016/S0261-3069(00)00035-2 . OpenUrl CrossRef (18). ↵ Delgado-Moreno , L. ; Peña , A. ; Mingorance , M. D . Design of Experiments in Environmental Chemistry Studies: Example of the Extraction of Triazines from Soil after Olive Cake Amendment . J Hazard Mater 2009 , 162 ( 2–3 ), 1121 – 1128 . doi: 10.1016/j.jhazmat.2008.05.148 . OpenUrl CrossRef PubMed (19). ↵ Garud , S. S. ; Karimi , I. A. ; Kraft , M . Design of Computer Experiments: A Review . Comput Chem Eng 2017 , 106 , 71 – 95 . doi: 10.1016/j.compchemeng.2017.05.010 . OpenUrl CrossRef (20). ↵ Gupta , U. ; Kim , Y. G. ; Lee , S. ; Tse , J. ; Lee , H. H. S. ; Wei , G. Y. ; Brooks , D. ; Wu , C. J . Chasing Carbon: The Elusive Environmental Footprint of Computing . In Proceedings - International Symposium on High-Performance Computer Architecture; IEEE Computer Society , 2021 ; Vol. 2021-February, pp 854 – 867 . doi: 10.1109/HPCA51647.2021.00076 . OpenUrl CrossRef (21). ↵ Roy , R. B. ; Kanakagiri , R. ; Jiang , Y. ; Tiwari , D . The Hidden Carbon Footprint of Serverless Computing. In SoCC 2024 - Proceedings of the 2024 ACM Symposium on Cloud Computing; Association for Computing Machinery, Inc , 2024 ; pp 570 – 579 . doi: 10.1145/3698038.3698546 . OpenUrl CrossRef (22). Freitag , C. ; Berners-Lee , M. ; Widdicks , K. ; Knowles , B. ; Blair , G. S. ; Friday , A. The Real Climate and Transformative Impact of ICT: A Critique of Estimates, Trends, and Regulations. Patterns . Cell Press September 10 , 2021 . doi: 10.1016/j.patter.2021.100340 . OpenUrl CrossRef PubMed (23). Masciari , E. ; Napolitano , E. V . The Environmental Cost of High Performance Computing System Simulation. In Proceedings - 2024 32nd Euromicro International Conference on Parallel, Distributed and Network-Based Processing, PDP 2024 ; Institute of Electrical and Electronics Engineers Inc ., 2024 ; pp 289 – 292 . doi: 10.1109/PDP62718.2024.00048 . OpenUrl CrossRef (24). Li , B. ; Roy , R. B. ; Wang , D. ; Samsi , S. ; Gadepally , V. ; Tiwari , D . Toward Sustainable HPC: Carbon Footprint Estimation and Environmental Implications of HPC Systems. In International Conference for High Performance Computing, Networking, Storage and Analysis , SC; IEEE Computer Society , 2023 . doi: 10.1145/3581784.3607035 . OpenUrl CrossRef (25). ↵ Silva , C. A. ; Vilaça , R. ; Pereira , A. ; Bessa , R. J . A Review on the Decarbonization of High-Performance Computing Centers . Renewable and Sustainable Energy Reviews 2024 , 189 . doi: 10.1016/j.rser.2023.114019 . OpenUrl CrossRef (26). ↵ Lannelongue , L. ; Grealey , J. ; Bateman , A. ; Inouye , M . Ten Simple Rules to Make Your Computing More Environmentally Sustainable . PLoS Comput Biol 2021 , 17 ( 9 ). doi: 10.1371/journal.pcbi.1009324 . OpenUrl CrossRef PubMed (27). ↵ https://www.top500.org/ . (28). ↵ Páll , S. ; Zhmurov , A. ; Bauer , P. ; Abraham , M. ; Lundborg , M. ; Gray , A. ; Hess , B. ; Lindahl , E . Heterogeneous Parallelization and Acceleration of Molecular Dynamics Simulations in GROMACS . J Chem Phys 2020 , 153 ( 13 ). doi: 10.1063/5.0018516 . OpenUrl CrossRef PubMed (29). ↵ Devore , J. Probability and Statistics for Engineering and the Sciences, 9th ed.; Matthew A. Carlton, Ed.; San Luis Obispo, CA , 2016 . (30). ↵ Montgomery , D. C. Design and Analysis of Experiments; John Wiley & Sons, Inc ., 2017 . (31). ↵ Strobl , C. ; Boulesteix , A. L. ; Kneib , T. ; Augustin , T. ; Zeileis , A . Conditional Variable Importance for Random Forests . BMC Bioinformatics 2008 , 9 ( 1 ). doi: 10.1186/1471-2105-9-307 . OpenUrl CrossRef PubMed View the discussion thread. Back to top Previous Next Posted October 08, 2025. Download PDF Data/Code Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following GROMODEX: Optimisation of GROMACS Performance through a Design of Experiment Approach Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share GROMODEX: Optimisation of GROMACS Performance through a Design of Experiment Approach Marco Savioli , Paolo Calligari , Ugo Locatelli , Gianfranco Bocchinfuso bioRxiv 2025.10.08.681202; doi: https://doi.org/10.1101/2025.10.08.681202 Share This Article: Copy Citation Tools GROMODEX: Optimisation of GROMACS Performance through a Design of Experiment Approach Marco Savioli , Paolo Calligari , Ugo Locatelli , Gianfranco Bocchinfuso bioRxiv 2025.10.08.681202; doi: https://doi.org/10.1101/2025.10.08.681202 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Biophysics Subject Areas All Articles Animal Behavior and Cognition (7621) Biochemistry (17645) Bioengineering (13867) Bioinformatics (41872) Biophysics (21416) Cancer Biology (18549) Cell Biology (25443) Clinical Trials (138) Developmental Biology (13360) Ecology (19866) Epidemiology (2067) Evolutionary Biology (24289) Genetics (15587) Genomics (22470) Immunology (17706) Microbiology (40314) Molecular Biology (17142) Neuroscience (88456) Paleontology (666) Pathology (2826) Pharmacology and Toxicology (4815) Physiology (7634) Plant Biology (15111) Scientific Communication and Education (2042) Synthetic Biology (4285) Systems Biology (9812) Zoology (2268)

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00