Closed-Loop Workflow of High-Entropy Materials Discovery: Efficient and Accurate Synthesizability Prediction via Domain-Specific Local LLMs | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Closed-Loop Workflow of High-Entropy Materials Discovery: Efficient and Accurate Synthesizability Prediction via Domain-Specific Local LLMs Yeongjun Yoon, Geun Ho Gu, Kyeounghak Kim This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8331266/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 24 Apr, 2026 Read the published version in npj Computational Materials → Version 1 posted 10 You are reading this latest preprint version Abstract High-entropy materials (HEMs) offer unprecedented opportunities for superior mechanical, thermal, and catalytic properties, but their vast chemical space makes experimental discovery resource-intensive. State-of-the-art commercial large language models (LLMs) notably fail at HEM synthesizability prediction, a critical bottleneck in materials development. We demonstrate that domain-specific fine-tuning transforms open-weight local LLMs into accurate predictors. Using a dataset of 321,083 inorganic compositions with 2,560 HEM examples, we fine-tuned three 4-bit-quantized models (gpt-oss-20b, Qwen3-14b, and DeepSeek-R1-Distill-Qwen-14b), achieving remarkable balanced accuracy of 0.957, 0.961, and 0.956, respectively. Critically, these models operate efficiently on accessible hardware (< 15GB VRAM), eliminating costly API dependencies while ensuring data privacy and consistent reproducibility. This work could open new pathways toward autonomous closed-loop discovery, where distributed local models enable rapid screening and iterative improvement through experimental feedback. Future collaborative efforts in open data sharing, particularly including negative results, would address current fragmentation in synthesis reporting and accelerate community-wide HEM discovery. Physical sciences/Chemistry Physical sciences/Materials science Physical sciences/Mathematics and computing High-entropy material (HEM) large language model (LLM) AI Prediction Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction High-entropy materials (HEMs), single-phase crystalline solid solutions incorporating multiple principal elements, represent a paradigm shift in materials design. Encompassing classes such as metal alloys, 1–5 borides, 6–8 carbides, 9–11 nitrides, 12–14 oxides, 15–17, and sulfides, 18–20, these HEMs exhibit vast compositional flexibility. This enables the tuning of exceptional mechanical, electrochemical, and catalytic properties, driven by mechanisms like cocktail effects, localized lattice distortion, and entropy-driven phase stabilization. 21 , 22 However, the rational design and screening of new HEMs remain a formidable challenge. The combinatorial complexity of their chemical space makes exhaustive experimental screening intractable. Consequently, discovery has historically relied on empirical or semi-empirical guidelines like the Hume-Rothery rules. 23 , 24 Yet, these rules, developed for binary or ternary alloys, frequently fall short for HEMs, failing to capture the complex multi-body interactions and thermodynamic competition in compositionally complex systems. 25 Traditional machine learning (ML) approaches have shown promise, but they often depend on computationally expensive, hand-crafted descriptors (e.g., from DFT) and are severely constrained by the scarcity of reliable, labeled experimental data. This scarcity is exacerbated by systematic reporting bias: successful syntheses are published, while failed attempts (valuable negative data) are rarely reported, leading to biased models. 26 – 30 Recently, Large Language Models (LLMs) emerged as a new, descriptor-free avenue. Kim et al. demonstrated that fine-tuned LLMs could predict the synthesizability of general inorganic materials. 31 Inspired by this, we sought to address a more complex challenge: extending this LLM-based approach specifically to HEMs, a domain defined by a much vaster compositional space and more nuanced formation rules. Furthermore, we aimed to solve the critical practical bottlenecks of accessibility, cost, and data privacy associated with proprietary, API-based commercial models. Results and Discussion Our approach leverages domain-specific fine-tuning of open-weight LLMs to create an efficient, locally deployable system for HEM synthesizability prediction (Fig. 1 ). The details for our methods are included in SI. For this purpose, our initial dataset comprised 393,053 unique inorganic compositions obtained from the previous study by Kim et al., 31 which integrated data from the Materials Project 32 and Open Quantum Materials Database 33 (OQMD) retrieved in February 2020. Among these, 40,817 compounds with Inorganic Crystal Structure Database (ICSD) references were labeled as positive (P, synthesized), while the remaining 352,236 were initially designated as unlabeled (U, hypothesized). To address the label noise inherent in Positive-Unlabeled (PU) learning, we rigorously screened the unlabeled dataset using the Stoichiometric Crystal Graph Neural Fingerprint (stoi-CGNF) framework, 34 rather than naively treating all unreported data as unsynthesizable (N). This process filtered out "potential hidden positives" to define a set of "Reliable Negatives" (RN), creating a robust Positive-Negative (PN) dataset essential for learning accurate high-entropy compositional rules. Detailed procedures for this data curation strategy are provided in Supporting Information . Furthermore, to enhance the representation of high-entropy materials (HEMs) in our dataset, we collected additional data through three complementary approaches: (1) database literature on high-entropy alloys and complex concentrated alloys, 35 (2) systematic searches using Elsevier's Scopus APIs, and (3) manual curation. These sources contributed substantially to enriching our training set with 2,560 additional synthesized HEM examples. To overcome the significant hardware barrier of deploying huge LLM models, our strategy involves 4-bit quantization, a process schematically illustrated in Fig. 2a . We selected three open-weight models (gpt-oss-20b, 36 Qwen3-14b, 37, and DeepSeek-R1-Distill-Qwen-14b 38 , 39 ). However, even these relatively modest-sized LLMs pose significant computational challenges, requiring substantial GPU memory that exceeds the resources available to most researchers in typical laboratory settings. To address this barrier, we applied 4-bit quantization to compress the high-precision floating-point (FP) parameters of each model (Fig. 2a). This quantization dramatically reduced both memory footprint and memory traffic, enabling deployment on accessible hardware platforms such as consumer GPUs or even cloud-based environments like Google Colaboratory. 40 The quantized models require approximately 15GB of VRAM, making them practical for researchers without access to specialized computing infrastructure. See Supporting Information for detailed model architectures. To evaluate model performance, we employed the confusion matrix framework (Fig. 2b). We treated this task as a binary classification operating on the constructed Positive-Negative (PN) dataset. From this, we calculated the True Positive Rate (TPR = TP/(TP + FN)), which quantifies the model's capability to correctly identify synthesizable candidates. Critically, we also calculated the True Negative Rate (TNR = TN/(TN + FP)), regarding the Reliable Negative (RN) set, which evaluates the model's effectiveness in excluding chemically non-viable compositions. Figure 2. (a) Schematic illustration of 4-bit quantization of LLMs. (b) The confusion matrix and the considered performance metrics. TPR, TNR, and bAcc indicate true positive rate, true negative rate, and balanced accuracy, respectively. (c) Performance of commercial LLMs for synthesizability prediction. For an experimental screening tool, TNR is the most important metric. A model with high TPR but low TNR is practically useless; it would recommend synthesizing nearly everything, leading to 100% experimental cost. A high TNR is what provides economic value, saving experimental resources by confidently filtering out non-viable candidates. The balanced accuracy (bAcc = (TPR + TNR)/2) provides a single, un-skewed metric that balances both goals, where 0.5 signifies the failure of random guessing. First, we established a performance baseline by evaluating state-of-the-art commercial LLMs accessible via API as of December 3, 2025. The assessment included the latest models from OpenAI (gpt-5.1, gpt-5-mini, gpt-4.1, gpt-4.1-mini), Google (gemini-3-pro-preview, gemini-2.5-pro, gemini-2.5-flash), and Anthropic (claude-opus-4.5, claude-sonnet-4.5, claude-haiku-4.5). Due to the significant costs and rate limits associated with commercial APIs, we conducted the evaluation on a stratified subset of 10,000 compositions randomly sampled from our Positive-Negative (PN) dataset. The system prompt defined the role and output constraints: "You are a materials science assistant. Given a chemical composition, answer only with 'P' (synthesizable/positive) or 'N' (non-synthesizable/negative)." Correspondingly, each user query was formatted as: "Is the material {composition} likely synthesizable? Answer with P (positive) or N (negative).". The results revealed that, with the notable exception of Google’s gemini-3-pro-preview, most commercial LLMs exhibited severely unbalanced performance between TPR and TNR, resulting in low balanced accuracy (bAcc) (Fig. 2c). While gemini-3-pro-preview, considered a frontier model based on its performance on the GPQA-Diamond benchmark, 41 a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry, achieved a remarkable bAcc of 0.82 with balanced TPR and TNR, other frontier models struggled to establish a reliable decision boundary. For instance, gpt-5.1 achieved a high TPR of 0.79 but a critically low TNR of 0.31, indicating a persistent bias toward classifying materials as synthesizable. Conversely, claude-sonnet-4.5 showed high TNR but poor TPR. Consequently, excluding gemini-3-pro-preview, most commercial models showed bAcc values hovering near or slightly above the random-guessing threshold of 0.5. This pattern of inadequate discrimination is perfectly mirrored in our 4-bit quantized base open-weight models (gpt-oss-20b, Qwen3-14b, and DeepSeek-R1-Distill-Qwen-14B) ( Figure S1 ). Specifically, gpt-oss-20b exhibited an extreme negative bias (TPR = 0.00, TNR = 1.00), effectively classifying all candidates as non-synthesizable, whereas DeepSeek-R1-Distill-Qwen-14B (r1-Qwen2.5-14b) showed the opposite failure mode with a severe positive bias (TPR = 0.95, TNR = 0.06). Consequently, all base models hovered near the random-guessing threshold with balanced accuracies of approximately 0.50 to 0.61. This demonstrates that despite the advancements in some frontier foundation models like gemini-3-pro-preview, the majority of base LLMs lack the inherent domain-specific knowledge to distinguish chemically non-viable (N) from synthesizable (P) materials without fine-tuning. To address this fundamental knowledge gap, we employed Quantized Low-Rank Adaptation 42 (QLoRA) to fine-tune our open-weight local LLMs using the HEM-existed P|N labeled dataset (Figs. 1 a and 3 a). In QLoRA, trainable low-rank adapter modules, representing less than 1% of total model parameters, are introduced into the frozen 4-bit quantized base model. This approach proved remarkably efficient, requiring only about 9 hours per epoch on a single consumer GPU (NVIDIA RTX A6000) while maintaining the model's general chemical knowledge intact. Importantly, the low-rank constraint acts as a natural regularization mechanism that prevents overfitting to dataset artifacts and instead focuses learning on the discriminative features essential for HEM synthesizability prediction. Details of the fine-tuning methodology and evolution of performance metrics are provided in the Supporting Information . This QLoRA-based fine-tuning dramatically transformed model performance across all three architectures. For gpt-oss-20b, the TPR surged from 0.00 to 0.98 (Fig. 3 b), overcoming its initial inability to identify synthesizable materials, while maintaining a near-perfect TNR of 0.94 (Fig. 3 c). Conversely, for r1-Qwen2.5-14b, which initially suffered from a severe lack of TNR, the TNR improved drastically from 0.06 to 0.94, while retaining a high TPR of 0.98. Qwen3-14b also showed substantial gains, evolving from a biased state to a balanced predictor with a TPR and TNR of 0.96. The transformative impact of fine-tuning becomes clear when examining balanced accuracy (bAcc) (Fig. 3 d). The base models achieved bAcc values near or slightly above 0.50 (red dashed line) due to their extreme biases—either positive or negative—averaging out to chance-level performance. This systematic failure to discriminate between synthesizable (P) and non-synthesizable (N) materials made them impractical for screening. After fine-tuning, however, all models achieved remarkable bAccs of 0.96, demonstrating robust discriminative capability. This result represents a fundamental transformation: from indiscriminate positive classification to genuine discriminative capability. The dramatically improved TNR enables these models to effectively filter out chemically non-viable compositions, steering experimental efforts away from materials unlikely to form stable phases. Most importantly for our primary objective, these fine-tuned open-weight models demonstrated exceptional performance specifically on HEM compositions (Fig. 4 ). To establish a rigorous baseline, we selected the single best-performing model from each commercial provider—OpenAI, Google, and Anthropic—based on the highest balanced accuracy (bAcc) achieved in the general screening task (Fig. 2c). Accordingly, gpt-4.1, gemini-3-pro-preview, and claude-sonnet-4.5 were chosen as the representative benchmarks. When evaluated on the HEM subset, our fine-tuned models achieved TPR values of 0.973, 0.931, and 0.973 for gpt-oss-20b, Qwen3-14B, and r1-Qwen2.5-14b, respectively. Notably, both gpt-oss-20b (TPR = 0.973) and r1-Qwen2.5-14b (TPR = 0.973) outperformed even the top-tier commercial model, Google's gemini-3-pro-preview (TPR = 0.930), while Qwen3-14b achieved comparable performance. These near-perfect scores indicate that our models can identify virtually all synthesizable high-entropy materials. This remarkable accuracy on the most challenging multi-component systems validates our approach of combining general materials data with targeted HEM examples during training. Beyond prediction, our work establishes a practical pathway for autonomous, closed-loop workflows (Fig. 5 a). A locally-deployable model, running on a lab's own hardware, can screen millions of de novo candidates, prioritizing a small, high-probability candidates for synthesis. The experimental outcomes are then fed back to iteratively retrain and improve the model. The true power of this loop lies in the "failures": a composition predicted as 'P' but found to fail in synthesis (a False Positive, FP) is not an error, but rather the most valuable data point. It provides the verified negative data that is currently missing from the literature, allowing the model to continuously correct its biases. This local deployment model ensures laboratories maintain complete data sovereignty over their novel compositions and experimental results. Transformative potential emerges through collaborative frameworks (Fig. 5 b). Multiple research groups operating parallel workflows could contribute to centralized open HEM data repository, pooling successes and failures across diverse synthesis conditions and methods. This addresses the critical weakness of these fields: fragmented reporting, where negative results are often discarded, and positive results often lack reproducibility details. Shared learning at this scale would accelerate discovery beyond what any single laboratory could achieve. Our approach democratizes advanced materials prediction by enabling any research group with standard GPU hardware to deploy these models and participate in collective discovery. As HEMs research expands into unexplored compositional territories, this distributed framework offers the only viable path through the combinatorial explosion of multi-component chemical space. These models establish a foundation for a new paradigm that combines local computation, global collaboration, and continuous learning to accelerate materials discovery. In conclusion, we have demonstrated that domain-specific fine-tuning, when combined with a targeted, data-centric enrichment strategy, transforms modest, 4-bit-quantized open-weight LLMs from non-functional predictors into powerful and accurate tools for HEM synthesizability prediction. Our fine-tuned models achieve balanced accuracy (bAcc) exceeding 0.96 and near-perfect TPR (> 0.97) for HEM compositions, all while running on accessible consumer-grade hardware. This work proves that the future of accelerated materials discovery lies not exclusively in massive, proprietary models, but in specialized, accessible, and collaborative tools. We provide a practical foundation for a new discovery paradigm that combines local computation, global data sharing, and continuous learning from both successes and failures. Declarations Code availability The fine-tuned models used in this study are available at: https://huggingface.co/collections/evenfarther/synthesizability-pn-prediction-balance-tpr-tnr Competing interests The authors declare no competing interests. Funding Open Access funding enabled and organized by the "Regional Innovation System & Education (RISE)" through the Seoul RISE Center, funded by the Ministry of Education (MOE) and the Seoul Metropolitan Government. (2025-RISE-01-027-04). Author Contribution K. K. and Y. Y. designed the research framework and drafted the manuscript. Y. Y. performed all simulations and developed LLM models. K. K. supervised the research. G. G. provided feedback on the manuscript. Acknowledgement This research was supported by the Nano & Material Technology Development Program through the National Research Foundation of Korea (NRF) funded by Ministry of Science and ICT (RS-2024-00448287) and Hyundai Motor Chung Mong-Koo Foundation Data Availability The fine-tuned models used in this study are available at: https://huggingface.co/collections/evenfarther/synthesizability-PN-prediction-balance-tpr-tnr. References Zhang, Y.; Yang, X.; Liaw, P. K. Alloy Design and Properties Optimization of High-Entropy Alloys. JOM 2012, 64 , 830–838. Yang, X.; Chen, S. Y.; Cotton, J. D.; Zhang, Y. Phase stability of low-density, multiprincipal component alloys containing aluminum, magnesium, and lithium. JOM 2014, 66 , 2009 – 2020. Takeuchi, A.; Amiya, K.; Wada, T.; Yubuta, K.; Zhang, W. High-entropy alloys with a hexagonal close-packed structure designed by equi-atomic alloy strategy and binary phase diagrams. JOM 2014, 66 , 1984 – 1992. Laws, K. J.; Crosby, C.; Sridhar, A.; Conway, P.; Koloadin, L. S.; Zhao, M.; Aron-Dine, S.; Bassman, L. C. High entropy brasses and bronzes – Microstructure, phase evolution and properties. J. Alloys Compd. 2015, 650 , 949–961. Sohn, S.; Liu, Y.; Liu, J.; Gong, P.; Prades-Rodel, S.; Blatter, A.; Scanley, B. E.; Broadbridge, C. C.; Schroers, J. Noble metal high entropy alloys. Scr. Mater. 2017, 126 , 29–32. Gild, J.; Zhang, Y.; Harrington, T.; Jiang, S.; Hu, T.; Quinn, M. C.; Mellor, W. M.; Zhou, N.; Vecchio, K.; Luo, J. High-entropy metal diborides: a new class of high-entropy materials and a new type of ultrahigh temperature ceramics. Sci. Rep. 2016, 6 , 37946. Rosenberg, A. A.; Lintz, D. T.; Li, J.; Zhang, Y.; Doane, J. T.; Bristol, M. N.; Kolakji, A.; Wang, T.; Yeung, M. T. Tailoring high-entropy borides for hydrogenation: crystal morphology and catalytic pathways. Inorg. Chem. Front. 2025, 12 , 4828–4834. Wang, X.; Zuo, Y.; Horta, S.; He, R.; Yang, L.; Moghaddam, A. O.; Ibáñez, M.; Qi, X.; Cabot, A. CoFeNiMnZnB as a high-entropy metal boride to boost the oxygen evolution reaction. ACS Appl. Mater. Interfaces 2022, 14 , 48212–48219. Sarker, P.; Harrington, T.; Toher, C.; Oses, C.; Samiee, M.; Maria, J.-P.; Brenner, D. W.; Vecchio, K. S.; Curtarolo, S. High-entropy high-hardness metal carbides discovered by entropy descriptors. Nat. Commun. 2018, 9 , 4980. Li, Y.; Zhao, S.; Wu, Z. Uncovering the effects of chemical disorder on the irradiation resistance of high-entropy carbide ceramics. Acta Mater. 2024, 277 , 120187. Hossain, M. D.; Borman, T.; Oses, C.; Esters, M.; Toher, C.; Feng, L.; Kumar, A.; Fahrenholtz, W. G.; Curtarolo, S.; Brenner, D. W.; LeBeau, J. M.; Maria J.-P. Entropy landscaping of high-entropy carbides. Adv. Mater. 2021, 33 , 2102904. Moskovskikh, D.; Vorotilo, S.; Buinevich, V.; Sedegov, A.; Kuskov, K.; Khort, A.; Shuck, C.; Zhukovskyi, M.; Mukasyan, A. Extremely hard and tough high entropy nitride ceramics. Sci. Rep. 2020, 10 , 19874. Hei, J.; Wang, N.; Jing, R.; Chen, X.; Yin, X.; Li, J.; Zuo, P.; Yin, Y.; Cui, L. High-entropy nitrides as superior electrocatalysts: unveiling the role of entropy in enhanced performance. Chem. Eur. J. 2025, 31 , e202500039. Li, J.; Chen, Y.; Zhao, Y.; Shi, X.; Wang, S.; Zhang, S. Super-hard (MoSiTiVZr)Nₓ high-entropy nitride coatings. J. Alloys Compd. 2022, 926 , 166807. Qian, F.; Cao, D.; Chen, S.; Yuan, Y.; Chen, K.; Chimtali, P. J.; Liu, H.; Jiang, W.; Sheng, B.; Yi, L.; Huang, J.; Hu, C.; Lei, H.; Wu, X.; Wen, Z.; Chen, Q.; Song, L. High-entropy RuO₂ catalyst with dual-site oxide path for durable acidic oxygen evolution reaction. Nat. Commun. 2025, 16 , 6894. Bao, W.; Shen, H.; Zhang, Y.; Qian, C.; Zeng, G.; Jing, K.; Cui, D.; Xia, J.; Liu, H.; Guo, C.; Yu, F.; Sun, K.; Li, J. High-entropy oxides for energy storage and conversion. J. Mater. Chem. A 2024, 12 , 23179–23201. Iwase, K.; Honma, I. High-entropy spinel oxide nanoparticles synthesized via supercritical hydrothermal processing as oxygen evolution electrocatalysts. ACS Appl. Energy Mater. 2022, 5 , 9292–9296. Lin, L.; Wang, K.; Sarkar, A.; Njel, C.; Karkera, G.; Wang, Q.; Azmi, R.; Fichtner, M.; Hahn, H.; Schweidler, S.; Breitung, B. High-entropy sulfides as electrode materials for Li-ion batteries. Adv. Energy Mater. 2022, 12 , 2103090. Gong, L.; Zhang, W.; Zhuang, Y.; Zhang, K.; Zhao, Q.; Xiao, D.; Liu, S.; Liu, Z.; Zhang, Y. High-entropy sulfides as electrode materials for Li-ion batteries. ACS Appl. Mater. Interfaces 2024, 16 , 66211–66218. Zhang, F.; Gao, T.; Zhang, Y.; Sun, K.; Qu, X.; Luo, Y.; Song, Y.; Fang, F.; Sun, D.; Wang, F.; Liu, Y. High-entropy metal sulfide nanocrystal libraries for highly reversible sodium storage. Adv. Mater. 2025, 37 , 2418890. Schweidler, S.; Botros, M.; Strauss, F.; Wang, Q.; Ma, Y.; Velasco, L.; Marques, G. C.; Sarkar, A.; Kübel, C.; Hahn, H.; Aghassi-Hagmann, J.; Brezesinski, T.; Breitung, B. High-entropy materials for energy and electronic applications. Nat. Rev. Mater. 2024, 9 , 266–281. Gu, X.; Guo, X.-B.; Li, W.-H.; Jiang, Y.-P.; Liu, Q.-X.; Tang, X.-G. High-entropy materials for application: electricity, magnetism, and optics. ACS Appl. Mater. Interfaces 2024, 16 , 53372–53392. Hume-Rothery, W.; Powell, H. M. Z . On the theory of super-lattice structures in alloys. Kristallogr. Cryst. Mater. 1935, 91 , 23–47. Mizutani, U. Hume-Rothery Rules for Structurally Complex Alloy Phases; CRC Press: Boca Raton, FL, 2016. Otto, F.; Yang, Y.; Bei, H.; George, E. P. Relative effects of enthalpy and entropy on the phase stability of equiatomic high-entropy alloys. Acta Mater. 2013, 61 , 2628–2638. Rao, Z.; Tung, P.-Y.; Xie, R.; Wei, Y.; Zhang, H.; Ferrari, A.; Klaver, T. P. C.; Körmann, F.; Sukumar, P. T.; da Silva, A. K.; Chen, Y.; Li, Z.; Ponge, D.; Neugebauer, J.; Gutfleisch, O.; Bauer, S.; Raabe, D. Machine learning–enabled high-entropy alloy discovery. Science 2022, 378 , 78–85. Wen, C.; Zhang, Y.; Wang, C.; Xue, D.; Bai, Y.; Antonov, S.; Dai, L.; Lookman, T.; Su, Y. Machine learning assisted design of high entropy alloys with desired property. Acta Mater. 2019, 170 , 109–117. Zhou, Z.; Zhou, Y.; He, Q.; Ding, Z.; Li, F.; Yang, Y. Machine learning guided appraisal and exploration of phase design for high entropy alloys. npj Comput. Mater. 2019, 5 , 128. Rickman, J. M.; Chan, H. M.; Harmer, M. P.; Smeltzer, J. A.; Marvel, C. J.; Roy, A.; Balasubramanian, G. Materials informatics for the screening of multi-principal elements and high-entropy alloys. Nat. Commun. 2019, 10 , 2618. Liu, X.; Zhang, J.; Pei, Z. Machine learning for high-entropy alloys: progress, challenges and opportunities. Prog. Mater. Sci. 2023, 131 , 101018. Kim, S.; Jung, Y.; Schrier, J. Large Language Models for Inorganic Synthesis Predictions. J. Am. Chem. Soc. 2024, 146 , 19654–19659. Jain, A.; Ong, S. P.; Hautier, G.; Chen, W.; Richards, W. D.; Dacek, S.; Cholia, S.; Gunter, D.; Skinner, D.; Ceder, G.; Persson, K. A. Commentary: The Materials Project: A materials genome approach to accelerating materials innovation. APL Mater. 2013, 1 , 011002. Kirklin, S.; Saal, J. E.; Meredig, B.; Thompson, A.; Doak, J. W.; Aykol, M.; Rühl, S.; Wolverton, C. The Open Quantum Materials Database (OQMD): assessing the accuracy of DFT formation energies. npj Comput. Mater. 2015, 1 , 15010. Jang, J.; Noh, J.; Zhou, L.; Gu, G. H.; Gregoire, J. M.; Jung, Y. Synthesizability of materials stoichiometry using semi-supervised learning. Matter 2024, 7 , 2294–2312. Gorsse, S.; Nguyen, M. H.; Senkov, O. N.; Miracle, D. B. Database on the mechanical properties of high entropy alloys and complex concentrated alloys. Data in Brief 2018, 21 , 2664–2678. OpenAI; Agarwal, S.; Ahmad, L.; Ai, J.; Altman, S.; Applebaum, A.; Arbus, E.; Arora, R. K.; Bai, Y.; Baker, B.; et al. gpt-oss-120b & gpt-oss-20b Model Card. arXiv 2025, arXiv:2508.10925. Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; et al. Qwen3 Technical Report. arXiv 2025, arXiv:2505.09388. Guo, D.; Yang, D.; Zhang, H.; et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 2025, 645 , 633–638. Yang, A.; Yu, B.; Li, C.; Liu, D.; Huang, F.; Huang, H.; Jiang, J.; Tu, J.; Zhang, J.; Zhou, J.; et al. Qwen2.5-1M Technical Report. arXiv 2025, arXiv:2501.15383. Google. Google Colaboratory. https://colab.research.google.com/ (accessed 2025-12-05) Rein, D.; Hou, B. L.; Stickland, A. C.; Petty, J.; Pang, R. Y.; Dirani, J.; Michael, J.; Bowman, S. R. GPQA: A graduate-level Google-proof Q&A benchmark. arXiv 2023, arXiv:2311.12022. Dettmers, T.; Pagnoni, A.; Holtzman, A.; Zettlemoyer, L. QLoRA: Efficient Finetuning of Quantized LLMs. arXiv 2023, arXiv:2305.14314. Additional Declarations No competing interests reported. Supplementary Files SILLM.docx floatimage1.png Cite Share Download PDF Status: Published Journal Publication published 24 Apr, 2026 Read the published version in npj Computational Materials → Version 1 posted Editorial decision: Revision requested 19 Jan, 2026 Reviews received at journal 07 Jan, 2026 Reviews received at journal 03 Jan, 2026 Reviewers agreed at journal 20 Dec, 2025 Reviewers agreed at journal 17 Dec, 2025 Reviewers agreed at journal 17 Dec, 2025 Reviewers invited by journal 17 Dec, 2025 Editor assigned by journal 16 Dec, 2025 Submission checks completed at journal 12 Dec, 2025 First submitted to journal 10 Dec, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8331266","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":562101406,"identity":"0003fdc0-c427-4365-bc6b-30f7136b126c","order_by":0,"name":"Yeongjun Yoon","email":"","orcid":"","institution":"Hanyang University","correspondingAuthor":false,"prefix":"","firstName":"Yeongjun","middleName":"","lastName":"Yoon","suffix":""},{"id":562101407,"identity":"a03d0a78-841a-4550-a5e3-e0575b99f288","order_by":1,"name":"Geun Ho Gu","email":"","orcid":"","institution":"Korea Institute of Energy Technology (KENTECH)","correspondingAuthor":false,"prefix":"","firstName":"Geun","middleName":"Ho","lastName":"Gu","suffix":""},{"id":562101409,"identity":"5a7c3e86-84d8-44ed-858a-e374c11a0c3c","order_by":2,"name":"Kyeounghak Kim","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+0lEQVRIiWNgGAWjYDACCRBxQIKBXwJJkLGBGC2SM0jUwsBgcINYLfyzm49u+HDGItr4dvOzxzw1d+wa2A8/YJy5B48ld46l3ZxxQyJ3251j5sY8x54lN/CkGTBueIZbi4FEjtltng9ALTcSzKR52A4nMzDkMDA+OIBPS/6323+AWjbPSP8mzfMPqIX/DSEtOWy3GYAO2wC0Tpq37bAdgwTQlg14tEjcSDO72XNGInfGjZwyybl9hxPYJJ4ZHJyBRwv/jORnN34cq8vtn5G+TeLNt8P2/PzJDx/24NGCAph4GBgS2xjA8UQkYPzBwGBPtOpRMApGwSgYMQAAPGNa1KL1CU4AAAAASUVORK5CYII=","orcid":"","institution":"Hanyang University","correspondingAuthor":true,"prefix":"","firstName":"Kyeounghak","middleName":"","lastName":"Kim","suffix":""}],"badges":[],"createdAt":"2025-12-11 00:23:11","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8331266/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8331266/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41524-026-02092-z","type":"published","date":"2026-04-24T15:58:42+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":98641713,"identity":"e039c0bb-3ab5-4676-bc70-bd12f30c32dd","added_by":"auto","created_at":"2025-12-19 18:45:45","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1063420,"visible":true,"origin":"","legend":"","description":"","filename":"ManuscrptLLM.docx","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/74463712302f8ab1affbfda6.docx"},{"id":98641711,"identity":"524761e2-c7d9-4fd3-a6a8-879712814ff8","added_by":"auto","created_at":"2025-12-19 18:45:45","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5180,"visible":true,"origin":"","legend":"","description":"","filename":"845bc9e25e39439dbbb648f9941b8298.json","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/52dc8560ae55093c29177c7b.json"},{"id":98775108,"identity":"91c1e0f3-7fc5-4f8a-afcd-099de63d2fcd","added_by":"auto","created_at":"2025-12-22 12:18:31","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":77152,"visible":true,"origin":"","legend":"","description":"","filename":"SILLM.docx","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/90bfc548559b63b4b1b25ca9.docx"},{"id":98775214,"identity":"89a532f0-0df6-484b-891c-14ba95ef9ab1","added_by":"auto","created_at":"2025-12-22 12:18:55","extension":"xml","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":99336,"visible":true,"origin":"","legend":"","description":"","filename":"845bc9e25e39439dbbb648f9941b82981enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/44768c8e3e60ef7de2102430.xml"},{"id":98775705,"identity":"4a6a02de-76f3-463d-887d-0c3d33fcdb6d","added_by":"auto","created_at":"2025-12-22 12:20:51","extension":"png","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":103829,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/b1c2815529637cc7a3fa5ad1.png"},{"id":98641723,"identity":"9416ebdc-16b6-4b07-9068-60a079add84c","added_by":"auto","created_at":"2025-12-19 18:45:45","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":231289,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/a15ed2a12fae61100197d71d.png"},{"id":98775003,"identity":"dd68114e-60b7-4a98-906d-e3a375e14a12","added_by":"auto","created_at":"2025-12-22 12:17:52","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":146137,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/0fdb7d10555f5375e50c0f91.png"},{"id":98641727,"identity":"5ecb99dd-8d59-4bc8-bd95-6953259af105","added_by":"auto","created_at":"2025-12-19 18:45:45","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":55831,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/f648ad0cd106c22c7ca7cdba.png"},{"id":98641718,"identity":"f7210a17-d3ed-4fb8-afa3-1fc1cfc8e96f","added_by":"auto","created_at":"2025-12-19 18:45:45","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":164766,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/051a4ce9f7089bbaaa29ebf8.png"},{"id":98775638,"identity":"b9894fc2-96ec-4ea7-8405-383ac937f95f","added_by":"auto","created_at":"2025-12-22 12:20:38","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":32242,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/13e71c97ab3279a36646684d.png"},{"id":98641732,"identity":"84ce3b41-f419-476c-8b2d-8096e12a70fe","added_by":"auto","created_at":"2025-12-19 18:45:45","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":27144,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/6acf3e52c97f9cc2b86e7a8e.png"},{"id":98774976,"identity":"144fbe13-ea10-43af-9aed-9bb166e817f9","added_by":"auto","created_at":"2025-12-22 12:17:42","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":53419,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/a15d05c6bc0e2de030115748.png"},{"id":98775942,"identity":"717907f0-4f15-4710-932d-cd1777b8e358","added_by":"auto","created_at":"2025-12-22 12:21:40","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":31907,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/a7899bbc3b18c3f5ec8b3438.png"},{"id":98641729,"identity":"2406c2ab-46af-4d26-bb8e-c23132f8f2ac","added_by":"auto","created_at":"2025-12-19 18:45:45","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":14525,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/7f47841b70d63936bb578424.png"},{"id":98641726,"identity":"40e90d67-4fc6-49ac-8698-76a088dcd225","added_by":"auto","created_at":"2025-12-19 18:45:45","extension":"png","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":41088,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/93df64437be938bb3eb4172a.png"},{"id":98775617,"identity":"4472bd4b-b8c7-48e4-a37a-d5de5c37aa16","added_by":"auto","created_at":"2025-12-22 12:20:34","extension":"xml","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":96720,"visible":true,"origin":"","legend":"","description":"","filename":"845bc9e25e39439dbbb648f9941b82981structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/2d852e62ecd59f170602034a.xml"},{"id":98641733,"identity":"854c4ee1-3b7c-4ee7-9dd7-3d6f938bb4ed","added_by":"auto","created_at":"2025-12-19 18:45:45","extension":"html","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":109813,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/93eff6c3fbb9572cbbb0baee.html"},{"id":98775026,"identity":"6367dae3-22e4-4926-88de-967c0ef5a5b0","added_by":"auto","created_at":"2025-12-22 12:17:58","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":205583,"visible":true,"origin":"","legend":"\u003cp\u003eSchematic illustration of workflow for fine-tuning domain specific LLMs to predict synthesizability of HEMs.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/53ca6ea8db7a0cbbac54b151.png"},{"id":98774959,"identity":"c3f8ae44-572f-4dc4-9e5b-be81ef37ccc9","added_by":"auto","created_at":"2025-12-22 12:17:30","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":434327,"visible":true,"origin":"","legend":"\u003cp\u003e(a) Schematic illustration of 4-bit quantization of LLMs. (b) The confusion matrix and the considered performance metrics. TPR, TNR, and bAcc indicate true positive rate, true negative rate, and balanced accuracy, respectively. (c) Performance of commercial LLMs for synthesizability prediction.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/01ba995e3a8a19517cb20fe7.png"},{"id":98775637,"identity":"ab451ea4-be86-4444-a1c5-9f586e49bf6a","added_by":"auto","created_at":"2025-12-22 12:20:37","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":400453,"visible":true,"origin":"","legend":"\u003cp\u003e(a) Schematic illustration of QLoRA process of local LLMs. (b-d) Performance metrics for synthesizability prediction comparing base models with QLoRA-adapted fine-tuned (FT) versions across the local LLMs. (b) True positive rate (TPR), (c) True negative rate (TNR), and (d) balanced accuracy (bAcc).\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/044d09571f2c50db6cb9ac03.png"},{"id":98775279,"identity":"e75fb70c-38ab-4b67-9f6a-6e1e23902b78","added_by":"auto","created_at":"2025-12-22 12:19:07","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":138513,"visible":true,"origin":"","legend":"\u003cp\u003eTrue positive rate (TPR) of HEMs of commercial models (pink) and open-weight fine-tuned models (green).\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/476d2e1d6f5704a30363f75a.png"},{"id":98774934,"identity":"c7e6475e-1395-4106-b89b-06dcae3bc857","added_by":"auto","created_at":"2025-12-22 12:17:15","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":374674,"visible":true,"origin":"","legend":"\u003cp\u003eSchematic illustration of (a) closed-loop workflow of domain-specific open-weight LLM-accelerated HEMs discovery and (b) collaborative framework with multiple parallel closed-loop workflows sharing data through a centralized HEMs open data repository, enabling accelerated HEMs discovery through distributed learning and knowledge sharing.\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/3893da7d51a1018e2b69eb70.png"},{"id":107928050,"identity":"5467e980-51ca-49f7-969a-a4f4ff75fffd","added_by":"auto","created_at":"2026-04-27 16:06:52","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1782940,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/09f5f952-4dc4-4c17-b237-7403deccd13a.pdf"},{"id":98774787,"identity":"57d9a6ad-8cde-4728-bd9b-affa648bfdf5","added_by":"auto","created_at":"2025-12-22 12:14:18","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":77152,"visible":true,"origin":"","legend":"","description":"","filename":"SILLM.docx","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/141ab3b77a11e47d21b26b7f.docx"},{"id":98641715,"identity":"1373b96f-b2ef-4626-ad94-9ee4df54b477","added_by":"auto","created_at":"2025-12-19 18:45:45","extension":"png","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":369619,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8331266/v1/30e16d6e13fd682e4230e358.png"}],"financialInterests":"No competing interests reported.","formattedTitle":"Closed-Loop Workflow of High-Entropy Materials Discovery: Efficient and Accurate Synthesizability Prediction via Domain-Specific Local LLMs","fulltext":[{"header":"Introduction","content":"\u003cp\u003eHigh-entropy materials (HEMs), single-phase crystalline solid solutions incorporating multiple principal elements, represent a paradigm shift in materials design. Encompassing classes such as metal alloys,\u003csup\u003e1\u0026ndash;5\u003c/sup\u003e borides,\u003csup\u003e6\u0026ndash;8\u003c/sup\u003e carbides,\u003csup\u003e9\u0026ndash;11\u003c/sup\u003e nitrides,\u003csup\u003e12\u0026ndash;14\u003c/sup\u003e oxides,\u003csup\u003e15\u0026ndash;17,\u003c/sup\u003e and sulfides,\u003csup\u003e18\u0026ndash;20,\u003c/sup\u003e these HEMs exhibit vast compositional flexibility. This enables the tuning of exceptional mechanical, electrochemical, and catalytic properties, driven by mechanisms like cocktail effects, localized lattice distortion, and entropy-driven phase stabilization.\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eHowever, the rational design and screening of new HEMs remain a formidable challenge. The combinatorial complexity of their chemical space makes exhaustive experimental screening intractable. Consequently, discovery has historically relied on empirical or semi-empirical guidelines like the Hume-Rothery rules.\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e Yet, these rules, developed for binary or ternary alloys, frequently fall short for HEMs, failing to capture the complex multi-body interactions and thermodynamic competition in compositionally complex systems.\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e Traditional machine learning (ML) approaches have shown promise, but they often depend on computationally expensive, hand-crafted descriptors (e.g., from DFT) and are severely constrained by the scarcity of reliable, labeled experimental data. This scarcity is exacerbated by systematic reporting bias: successful syntheses are published, while failed attempts (valuable negative data) are rarely reported, leading to biased models.\u003csup\u003e\u003cspan additionalcitationids=\"CR27 CR28 CR29\" citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eRecently, Large Language Models (LLMs) emerged as a new, descriptor-free avenue. Kim et al. demonstrated that fine-tuned LLMs could predict the synthesizability of general inorganic materials.\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e Inspired by this, we sought to address a more complex challenge: extending this LLM-based approach specifically to HEMs, a domain defined by a much vaster compositional space and more nuanced formation rules. Furthermore, we aimed to solve the critical practical bottlenecks of accessibility, cost, and data privacy associated with proprietary, API-based commercial models.\u003c/p\u003e"},{"header":"Results and Discussion","content":"\u003cp\u003e \u003c/p\u003e \u003cp\u003eOur approach leverages domain-specific fine-tuning of open-weight LLMs to create an efficient, locally deployable system for HEM synthesizability prediction (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The details for our methods are included in SI. For this purpose, our initial dataset comprised 393,053 unique inorganic compositions obtained from the previous study by Kim et al.,\u003csup\u003e31\u003c/sup\u003e which integrated data from the Materials Project\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e and Open Quantum Materials Database\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e (OQMD) retrieved in February 2020. Among these, 40,817 compounds with Inorganic Crystal Structure Database (ICSD) references were labeled as positive (P, synthesized), while the remaining 352,236 were initially designated as unlabeled (U, hypothesized).\u003c/p\u003e \u003cp\u003eTo address the label noise inherent in Positive-Unlabeled (PU) learning, we rigorously screened the unlabeled dataset using the Stoichiometric Crystal Graph Neural Fingerprint (stoi-CGNF) framework,\u003csup\u003e34\u003c/sup\u003e rather than naively treating all unreported data as unsynthesizable (N). This process filtered out \"potential hidden positives\" to define a set of \"Reliable Negatives\" (RN), creating a robust Positive-Negative (PN) dataset essential for learning accurate high-entropy compositional rules. Detailed procedures for this data curation strategy are provided in \u003cb\u003eSupporting Information\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eFurthermore, to enhance the representation of high-entropy materials (HEMs) in our dataset, we collected additional data through three complementary approaches: (1) database literature on high-entropy alloys and complex concentrated alloys,\u003csup\u003e35\u003c/sup\u003e (2) systematic searches using Elsevier's Scopus APIs, and (3) manual curation. These sources contributed substantially to enriching our training set with 2,560 additional synthesized HEM examples.\u003c/p\u003e \u003cp\u003eTo overcome the significant hardware barrier of deploying huge LLM models, our strategy involves 4-bit quantization, a process schematically illustrated in \u003cb\u003eFig.\u0026nbsp;2a\u003c/b\u003e. We selected three open-weight models (gpt-oss-20b,\u003csup\u003e36\u003c/sup\u003e Qwen3-14b,\u003csup\u003e37,\u003c/sup\u003e and DeepSeek-R1-Distill-Qwen-14b\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e,\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e). However, even these relatively modest-sized LLMs pose significant computational challenges, requiring substantial GPU memory that exceeds the resources available to most researchers in typical laboratory settings. To address this barrier, we applied 4-bit quantization to compress the high-precision floating-point (FP) parameters of each model (Fig.\u0026nbsp;2a). This quantization dramatically reduced both memory footprint and memory traffic, enabling deployment on accessible hardware platforms such as consumer GPUs or even cloud-based environments like Google Colaboratory.\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e The quantized models require approximately 15GB of VRAM, making them practical for researchers without access to specialized computing infrastructure. See \u003cb\u003eSupporting Information\u003c/b\u003e for detailed model architectures.\u003c/p\u003e \u003cp\u003eTo evaluate model performance, we employed the confusion matrix framework (Fig.\u0026nbsp;2b). We treated this task as a binary classification operating on the constructed Positive-Negative (PN) dataset. From this, we calculated the True Positive Rate (TPR\u0026thinsp;=\u0026thinsp;TP/(TP\u0026thinsp;+\u0026thinsp;FN)), which quantifies the model's capability to correctly identify synthesizable candidates. Critically, we also calculated the True Negative Rate (TNR\u0026thinsp;=\u0026thinsp;TN/(TN\u0026thinsp;+\u0026thinsp;FP)), regarding the Reliable Negative (RN) set, which evaluates the model's effectiveness in excluding chemically non-viable compositions.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure 2.\u003c/b\u003e (a) Schematic illustration of 4-bit quantization of LLMs. (b) The confusion matrix and the considered performance metrics. TPR, TNR, and bAcc indicate true positive rate, true negative rate, and balanced accuracy, respectively. (c) Performance of commercial LLMs for synthesizability prediction.\u003c/p\u003e \u003cp\u003eFor an experimental screening tool, TNR is the most important metric. A model with high TPR but low TNR is practically useless; it would recommend synthesizing nearly everything, leading to 100% experimental cost. A high TNR is what provides economic value, saving experimental resources by confidently filtering out non-viable candidates. The balanced accuracy (bAcc = (TPR\u0026thinsp;+\u0026thinsp;TNR)/2) provides a single, un-skewed metric that balances both goals, where 0.5 signifies the failure of random guessing.\u003c/p\u003e \u003cp\u003eFirst, we established a performance baseline by evaluating state-of-the-art commercial LLMs accessible via API as of December 3, 2025. The assessment included the latest models from OpenAI (gpt-5.1, gpt-5-mini, gpt-4.1, gpt-4.1-mini), Google (gemini-3-pro-preview, gemini-2.5-pro, gemini-2.5-flash), and Anthropic (claude-opus-4.5, claude-sonnet-4.5, claude-haiku-4.5). Due to the significant costs and rate limits associated with commercial APIs, we conducted the evaluation on a stratified subset of 10,000 compositions randomly sampled from our Positive-Negative (PN) dataset. The system prompt defined the role and output constraints: \"You are a materials science assistant. Given a chemical composition, answer only with 'P' (synthesizable/positive) or 'N' (non-synthesizable/negative).\" Correspondingly, each user query was formatted as: \"Is the material {composition} likely synthesizable? Answer with P (positive) or N (negative).\".\u003c/p\u003e \u003cp\u003eThe results revealed that, with the notable exception of Google\u0026rsquo;s gemini-3-pro-preview, most commercial LLMs exhibited severely unbalanced performance between TPR and TNR, resulting in low balanced accuracy (bAcc) (Fig.\u0026nbsp;2c). While gemini-3-pro-preview, considered a frontier model based on its performance on the GPQA-Diamond benchmark,\u003csup\u003e41\u003c/sup\u003e a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry, achieved a remarkable bAcc of 0.82 with balanced TPR and TNR, other frontier models struggled to establish a reliable decision boundary. For instance, gpt-5.1 achieved a high TPR of 0.79 but a critically low TNR of 0.31, indicating a persistent bias toward classifying materials as synthesizable. Conversely, claude-sonnet-4.5 showed high TNR but poor TPR. Consequently, excluding gemini-3-pro-preview, most commercial models showed bAcc values hovering near or slightly above the random-guessing threshold of 0.5.\u003c/p\u003e \u003cp\u003eThis pattern of inadequate discrimination is perfectly mirrored in our 4-bit quantized base open-weight models (gpt-oss-20b, Qwen3-14b, and DeepSeek-R1-Distill-Qwen-14B) (\u003cb\u003eFigure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e). Specifically, gpt-oss-20b exhibited an extreme negative bias (TPR\u0026thinsp;=\u0026thinsp;0.00, TNR\u0026thinsp;=\u0026thinsp;1.00), effectively classifying all candidates as non-synthesizable, whereas DeepSeek-R1-Distill-Qwen-14B (r1-Qwen2.5-14b) showed the opposite failure mode with a severe positive bias (TPR\u0026thinsp;=\u0026thinsp;0.95, TNR\u0026thinsp;=\u0026thinsp;0.06). Consequently, all base models hovered near the random-guessing threshold with balanced accuracies of approximately 0.50 to 0.61. This demonstrates that despite the advancements in some frontier foundation models like gemini-3-pro-preview, the majority of base LLMs lack the inherent domain-specific knowledge to distinguish chemically non-viable (N) from synthesizable (P) materials without fine-tuning.\u003c/p\u003e \u003cp\u003eTo address this fundamental knowledge gap, we employed Quantized Low-Rank Adaptation\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e (QLoRA) to fine-tune our open-weight local LLMs using the HEM-existed P|N labeled dataset (Figs.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea and \u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003ea). In QLoRA, trainable low-rank adapter modules, representing less than 1% of total model parameters, are introduced into the frozen 4-bit quantized base model. This approach proved remarkably efficient, requiring only about 9 hours per epoch on a single consumer GPU (NVIDIA RTX A6000) while maintaining the model's general chemical knowledge intact. Importantly, the low-rank constraint acts as a natural regularization mechanism that prevents overfitting to dataset artifacts and instead focuses learning on the discriminative features essential for HEM synthesizability prediction. Details of the fine-tuning methodology and evolution of performance metrics are provided in the \u003cb\u003eSupporting Information\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eThis QLoRA-based fine-tuning dramatically transformed model performance across all three architectures. For gpt-oss-20b, the TPR surged from 0.00 to 0.98 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003eb), overcoming its initial inability to identify synthesizable materials, while maintaining a near-perfect TNR of 0.94 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003ec). Conversely, for r1-Qwen2.5-14b, which initially suffered from a severe lack of TNR, the TNR improved drastically from 0.06 to 0.94, while retaining a high TPR of 0.98. Qwen3-14b also showed substantial gains, evolving from a biased state to a balanced predictor with a TPR and TNR of 0.96. The transformative impact of fine-tuning becomes clear when examining balanced accuracy (bAcc) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003ed). The base models achieved bAcc values near or slightly above 0.50 (red dashed line) due to their extreme biases\u0026mdash;either positive or negative\u0026mdash;averaging out to chance-level performance. This systematic failure to discriminate between synthesizable (P) and non-synthesizable (N) materials made them impractical for screening. After fine-tuning, however, all models achieved remarkable bAccs of 0.96, demonstrating robust discriminative capability.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThis result represents a fundamental transformation: from indiscriminate positive classification to genuine discriminative capability. The dramatically improved TNR enables these models to effectively filter out chemically non-viable compositions, steering experimental efforts away from materials unlikely to form stable phases. Most importantly for our primary objective, these fine-tuned open-weight models demonstrated exceptional performance specifically on HEM compositions (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e4\u003c/span\u003e). To establish a rigorous baseline, we selected the single best-performing model from each commercial provider\u0026mdash;OpenAI, Google, and Anthropic\u0026mdash;based on the highest balanced accuracy (bAcc) achieved in the general screening task (Fig.\u0026nbsp;2c). Accordingly, gpt-4.1, gemini-3-pro-preview, and claude-sonnet-4.5 were chosen as the representative benchmarks. When evaluated on the HEM subset, our fine-tuned models achieved TPR values of 0.973, 0.931, and 0.973 for gpt-oss-20b, Qwen3-14B, and r1-Qwen2.5-14b, respectively. Notably, both gpt-oss-20b (TPR\u0026thinsp;=\u0026thinsp;0.973) and r1-Qwen2.5-14b (TPR\u0026thinsp;=\u0026thinsp;0.973) outperformed even the top-tier commercial model, Google's gemini-3-pro-preview (TPR\u0026thinsp;=\u0026thinsp;0.930), while Qwen3-14b achieved comparable performance. These near-perfect scores indicate that our models can identify virtually all synthesizable high-entropy materials. This remarkable accuracy on the most challenging multi-component systems validates our approach of combining general materials data with targeted HEM examples during training.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eBeyond prediction, our work establishes a practical pathway for autonomous, closed-loop workflows (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e5\u003c/span\u003ea). A locally-deployable model, running on a lab's own hardware, can screen millions of de novo candidates, prioritizing a small, high-probability candidates for synthesis. The experimental outcomes are then fed back to iteratively retrain and improve the model. The true power of this loop lies in the \"failures\": a composition predicted as 'P' but found to fail in synthesis (a False Positive, FP) is not an error, but rather the most valuable data point. It provides the verified negative data that is currently missing from the literature, allowing the model to continuously correct its biases. This local deployment model ensures laboratories maintain complete data sovereignty over their novel compositions and experimental results.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTransformative potential emerges through collaborative frameworks (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e5\u003c/span\u003eb). Multiple research groups operating parallel workflows could contribute to centralized open HEM data repository, pooling successes and failures across diverse synthesis conditions and methods. This addresses the critical weakness of these fields: fragmented reporting, where negative results are often discarded, and positive results often lack reproducibility details. Shared learning at this scale would accelerate discovery beyond what any single laboratory could achieve. Our approach democratizes advanced materials prediction by enabling any research group with standard GPU hardware to deploy these models and participate in collective discovery. As HEMs research expands into unexplored compositional territories, this distributed framework offers the only viable path through the combinatorial explosion of multi-component chemical space. These models establish a foundation for a new paradigm that combines local computation, global collaboration, and continuous learning to accelerate materials discovery.\u003c/p\u003e \u003cp\u003eIn conclusion, we have demonstrated that domain-specific fine-tuning, when combined with a targeted, data-centric enrichment strategy, transforms modest, 4-bit-quantized open-weight LLMs from non-functional predictors into powerful and accurate tools for HEM synthesizability prediction. Our fine-tuned models achieve balanced accuracy (bAcc) exceeding 0.96 and near-perfect TPR (\u0026gt;\u0026thinsp;0.97) for HEM compositions, all while running on accessible consumer-grade hardware. This work proves that the future of accelerated materials discovery lies not exclusively in massive, proprietary models, but in specialized, accessible, and collaborative tools. We provide a practical foundation for a new discovery paradigm that combines local computation, global data sharing, and continuous learning from both successes and failures.\u003c/p\u003e "},{"header":"Declarations","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eCode availability\u003c/h2\u003e \u003cp\u003eThe fine-tuned models used in this study are available at: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://huggingface.co/collections/evenfarther/synthesizability-pn-prediction-balance-tpr-tnr\u003c/span\u003e\u003cspan address=\"https://huggingface.co/collections/evenfarther/synthesizability-pn-prediction-balance-tpr-tnr\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/div\u003e\u003cp\u003e \u003ch2\u003eCompeting interests\u003c/h2\u003e \u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eOpen Access funding enabled and organized by the \"Regional Innovation System \u0026amp; Education (RISE)\" through the Seoul RISE Center, funded by the Ministry of Education (MOE) and the Seoul Metropolitan Government. (2025-RISE-01-027-04).\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eK. K. and Y. Y. designed the research framework and drafted the manuscript. Y. Y. performed all simulations and developed LLM models. K. K. supervised the research. G. G. provided feedback on the manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThis research was supported by the Nano \u0026amp; Material Technology Development Program through the National Research Foundation of Korea (NRF) funded by Ministry of Science and ICT (RS-2024-00448287) and Hyundai Motor Chung Mong-Koo Foundation\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe fine-tuned models used in this study are available at: https://huggingface.co/collections/evenfarther/synthesizability-PN-prediction-balance-tpr-tnr.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eZhang, Y.; Yang, X.; Liaw, P. K. Alloy Design and Properties Optimization of High-Entropy Alloys. \u003cem\u003eJOM\u003c/em\u003e 2012, \u003cem\u003e64\u003c/em\u003e, 830\u0026ndash;838.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, X.; Chen, S. Y.; Cotton, J. D.; Zhang, Y. Phase stability of low-density, multiprincipal component alloys containing aluminum, magnesium, and lithium. \u003cem\u003eJOM\u003c/em\u003e 2014, \u003cem\u003e66\u003c/em\u003e, 2009\u0026thinsp;\u0026ndash;\u0026thinsp;2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTakeuchi, A.; Amiya, K.; Wada, T.; Yubuta, K.; Zhang, W. High-entropy alloys with a hexagonal close-packed structure designed by equi-atomic alloy strategy and binary phase diagrams. \u003cem\u003eJOM\u003c/em\u003e 2014, \u003cem\u003e66\u003c/em\u003e, 1984\u0026thinsp;\u0026ndash;\u0026thinsp;1992.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLaws, K. J.; Crosby, C.; Sridhar, A.; Conway, P.; Koloadin, L. S.; Zhao, M.; Aron-Dine, S.; Bassman, L. C. High entropy brasses and bronzes \u0026ndash; Microstructure, phase evolution and properties. \u003cem\u003eJ. Alloys Compd.\u003c/em\u003e 2015, \u003cem\u003e650\u003c/em\u003e, 949\u0026ndash;961.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSohn, S.; Liu, Y.; Liu, J.; Gong, P.; Prades-Rodel, S.; Blatter, A.; Scanley, B. E.; Broadbridge, C. C.; Schroers, J. Noble metal high entropy alloys. \u003cem\u003eScr. Mater.\u003c/em\u003e 2017, \u003cem\u003e126\u003c/em\u003e, 29\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGild, J.; Zhang, Y.; Harrington, T.; Jiang, S.; Hu, T.; Quinn, M. C.; Mellor, W. M.; Zhou, N.; Vecchio, K.; Luo, J. High-entropy metal diborides: a new class of high-entropy materials and a new type of ultrahigh temperature ceramics. \u003cem\u003eSci. Rep.\u003c/em\u003e 2016, \u003cem\u003e6\u003c/em\u003e, 37946.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRosenberg, A. A.; Lintz, D. T.; Li, J.; Zhang, Y.; Doane, J. T.; Bristol, M. N.; Kolakji, A.; Wang, T.; Yeung, M. T. Tailoring high-entropy borides for hydrogenation: crystal morphology and catalytic pathways. \u003cem\u003eInorg. Chem. Front.\u003c/em\u003e 2025, \u003cem\u003e12\u003c/em\u003e, 4828\u0026ndash;4834.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, X.; Zuo, Y.; Horta, S.; He, R.; Yang, L.; Moghaddam, A. O.; Ib\u0026aacute;\u0026ntilde;ez, M.; Qi, X.; Cabot, A. CoFeNiMnZnB as a high-entropy metal boride to boost the oxygen evolution reaction. \u003cem\u003eACS Appl. Mater. Interfaces\u003c/em\u003e 2022, \u003cem\u003e14\u003c/em\u003e, 48212\u0026ndash;48219.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSarker, P.; Harrington, T.; Toher, C.; Oses, C.; Samiee, M.; Maria, J.-P.; Brenner, D. W.; Vecchio, K. S.; Curtarolo, S. High-entropy high-hardness metal carbides discovered by entropy descriptors. \u003cem\u003eNat. Commun.\u003c/em\u003e 2018, \u003cem\u003e9\u003c/em\u003e, 4980.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, Y.; Zhao, S.; Wu, Z. Uncovering the effects of chemical disorder on the irradiation resistance of high-entropy carbide ceramics. \u003cem\u003eActa Mater.\u003c/em\u003e 2024, \u003cem\u003e277\u003c/em\u003e, 120187.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHossain, M. D.; Borman, T.; Oses, C.; Esters, M.; Toher, C.; Feng, L.; Kumar, A.; Fahrenholtz, W. G.; Curtarolo, S.; Brenner, D. W.; LeBeau, J. M.; Maria J.-P. Entropy landscaping of high-entropy carbides. \u003cem\u003eAdv. Mater.\u003c/em\u003e 2021, \u003cem\u003e33\u003c/em\u003e, 2102904.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoskovskikh, D.; Vorotilo, S.; Buinevich, V.; Sedegov, A.; Kuskov, K.; Khort, A.; Shuck, C.; Zhukovskyi, M.; Mukasyan, A. Extremely hard and tough high entropy nitride ceramics. \u003cem\u003eSci. Rep.\u003c/em\u003e 2020, \u003cem\u003e10\u003c/em\u003e, 19874.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHei, J.; Wang, N.; Jing, R.; Chen, X.; Yin, X.; Li, J.; Zuo, P.; Yin, Y.; Cui, L. High-entropy nitrides as superior electrocatalysts: unveiling the role of entropy in enhanced performance. \u003cem\u003eChem. Eur. J.\u003c/em\u003e 2025, \u003cem\u003e31\u003c/em\u003e, e202500039.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, J.; Chen, Y.; Zhao, Y.; Shi, X.; Wang, S.; Zhang, S. Super-hard (MoSiTiVZr)Nₓ high-entropy nitride coatings. \u003cem\u003eJ. Alloys Compd.\u003c/em\u003e 2022, \u003cem\u003e926\u003c/em\u003e, 166807.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQian, F.; Cao, D.; Chen, S.; Yuan, Y.; Chen, K.; Chimtali, P. J.; Liu, H.; Jiang, W.; Sheng, B.; Yi, L.; Huang, J.; Hu, C.; Lei, H.; Wu, X.; Wen, Z.; Chen, Q.; Song, L. High-entropy RuO₂ catalyst with dual-site oxide path for durable acidic oxygen evolution reaction. \u003cem\u003eNat. Commun.\u003c/em\u003e 2025, \u003cem\u003e16\u003c/em\u003e, 6894.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBao, W.; Shen, H.; Zhang, Y.; Qian, C.; Zeng, G.; Jing, K.; Cui, D.; Xia, J.; Liu, H.; Guo, C.; Yu, F.; Sun, K.; Li, J. High-entropy oxides for energy storage and conversion. \u003cem\u003eJ. Mater. Chem. A\u003c/em\u003e 2024, \u003cem\u003e12\u003c/em\u003e, 23179\u0026ndash;23201.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIwase, K.; Honma, I. High-entropy spinel oxide nanoparticles synthesized via supercritical hydrothermal processing as oxygen evolution electrocatalysts. \u003cem\u003eACS Appl. Energy Mater.\u003c/em\u003e 2022, \u003cem\u003e5\u003c/em\u003e, 9292\u0026ndash;9296.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin, L.; Wang, K.; Sarkar, A.; Njel, C.; Karkera, G.; Wang, Q.; Azmi, R.; Fichtner, M.; Hahn, H.; Schweidler, S.; Breitung, B. High-entropy sulfides as electrode materials for Li-ion batteries. \u003cem\u003eAdv. Energy Mater.\u003c/em\u003e 2022, \u003cem\u003e12\u003c/em\u003e, 2103090.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGong, L.; Zhang, W.; Zhuang, Y.; Zhang, K.; Zhao, Q.; Xiao, D.; Liu, S.; Liu, Z.; Zhang, Y. High-entropy sulfides as electrode materials for Li-ion batteries. \u003cem\u003eACS Appl. Mater. Interfaces\u003c/em\u003e 2024, \u003cem\u003e16\u003c/em\u003e, 66211\u0026ndash;66218.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, F.; Gao, T.; Zhang, Y.; Sun, K.; Qu, X.; Luo, Y.; Song, Y.; Fang, F.; Sun, D.; Wang, F.; Liu, Y. High-entropy metal sulfide nanocrystal libraries for highly reversible sodium storage. \u003cem\u003eAdv. Mater.\u003c/em\u003e 2025, \u003cem\u003e37\u003c/em\u003e, 2418890.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchweidler, S.; Botros, M.; Strauss, F.; Wang, Q.; Ma, Y.; Velasco, L.; Marques, G. C.; Sarkar, A.; K\u0026uuml;bel, C.; Hahn, H.; Aghassi-Hagmann, J.; Brezesinski, T.; Breitung, B. High-entropy materials for energy and electronic applications. \u003cem\u003eNat. Rev. Mater.\u003c/em\u003e 2024, \u003cem\u003e9\u003c/em\u003e, 266\u0026ndash;281.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGu, X.; Guo, X.-B.; Li, W.-H.; Jiang, Y.-P.; Liu, Q.-X.; Tang, X.-G. High-entropy materials for application: electricity, magnetism, and optics. \u003cem\u003eACS Appl. Mater. Interfaces\u003c/em\u003e 2024, \u003cem\u003e16\u003c/em\u003e, 53372\u0026ndash;53392.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHume-Rothery, W.; Powell, H. M. \u003cem\u003eZ\u003c/em\u003e. On the theory of super-lattice structures in alloys. \u003cem\u003eKristallogr. Cryst. Mater.\u003c/em\u003e 1935, \u003cem\u003e91\u003c/em\u003e, 23\u0026ndash;47.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMizutani, U. Hume-Rothery Rules for Structurally Complex Alloy Phases; CRC Press: Boca Raton, FL, 2016.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOtto, F.; Yang, Y.; Bei, H.; George, E. P. Relative effects of enthalpy and entropy on the phase stability of equiatomic high-entropy alloys. \u003cem\u003eActa Mater.\u003c/em\u003e 2013, \u003cem\u003e61\u003c/em\u003e, 2628\u0026ndash;2638.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRao, Z.; Tung, P.-Y.; Xie, R.; Wei, Y.; Zhang, H.; Ferrari, A.; Klaver, T. P. C.; K\u0026ouml;rmann, F.; Sukumar, P. T.; da Silva, A. K.; Chen, Y.; Li, Z.; Ponge, D.; Neugebauer, J.; Gutfleisch, O.; Bauer, S.; Raabe, D. Machine learning\u0026ndash;enabled high-entropy alloy discovery. \u003cem\u003eScience\u003c/em\u003e 2022, \u003cem\u003e378\u003c/em\u003e, 78\u0026ndash;85.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWen, C.; Zhang, Y.; Wang, C.; Xue, D.; Bai, Y.; Antonov, S.; Dai, L.; Lookman, T.; Su, Y. Machine learning assisted design of high entropy alloys with desired property. \u003cem\u003eActa Mater.\u003c/em\u003e 2019, \u003cem\u003e170\u003c/em\u003e, 109\u0026ndash;117.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou, Z.; Zhou, Y.; He, Q.; Ding, Z.; Li, F.; Yang, Y. Machine learning guided appraisal and exploration of phase design for high entropy alloys. \u003cem\u003enpj Comput. Mater.\u003c/em\u003e 2019, \u003cem\u003e5\u003c/em\u003e, 128.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRickman, J. M.; Chan, H. M.; Harmer, M. P.; Smeltzer, J. A.; Marvel, C. J.; Roy, A.; Balasubramanian, G. Materials informatics for the screening of multi-principal elements and high-entropy alloys. \u003cem\u003eNat. Commun.\u003c/em\u003e 2019, \u003cem\u003e10\u003c/em\u003e, 2618.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, X.; Zhang, J.; Pei, Z. Machine learning for high-entropy alloys: progress, challenges and opportunities. \u003cem\u003eProg. Mater. Sci.\u003c/em\u003e 2023, \u003cem\u003e131\u003c/em\u003e, 101018.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim, S.; Jung, Y.; Schrier, J. Large Language Models for Inorganic Synthesis Predictions. \u003cem\u003eJ. Am. Chem. Soc.\u003c/em\u003e 2024, \u003cem\u003e146\u003c/em\u003e, 19654\u0026ndash;19659.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJain, A.; Ong, S. P.; Hautier, G.; Chen, W.; Richards, W. D.; Dacek, S.; Cholia, S.; Gunter, D.; Skinner, D.; Ceder, G.; Persson, K. A. Commentary: The Materials Project: A materials genome approach to accelerating materials innovation. \u003cem\u003eAPL Mater.\u003c/em\u003e 2013, \u003cem\u003e1\u003c/em\u003e, 011002.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKirklin, S.; Saal, J. E.; Meredig, B.; Thompson, A.; Doak, J. W.; Aykol, M.; R\u0026uuml;hl, S.; Wolverton, C. The Open Quantum Materials Database (OQMD): assessing the accuracy of DFT formation energies. \u003cem\u003enpj Comput. Mater.\u003c/em\u003e 2015, \u003cem\u003e1\u003c/em\u003e, 15010.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJang, J.; Noh, J.; Zhou, L.; Gu, G. H.; Gregoire, J. M.; Jung, Y. Synthesizability of materials stoichiometry using semi-supervised learning. \u003cem\u003eMatter\u003c/em\u003e 2024, \u003cem\u003e7\u003c/em\u003e, 2294\u0026ndash;2312.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGorsse, S.; Nguyen, M. H.; Senkov, O. N.; Miracle, D. B. Database on the mechanical properties of high entropy alloys and complex concentrated alloys. \u003cem\u003eData in Brief\u003c/em\u003e 2018, \u003cem\u003e21\u003c/em\u003e, 2664\u0026ndash;2678.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOpenAI; Agarwal, S.; Ahmad, L.; Ai, J.; Altman, S.; Applebaum, A.; Arbus, E.; Arora, R. K.; Bai, Y.; Baker, B.; et al. gpt-oss-120b \u0026amp; gpt-oss-20b Model Card. \u003cem\u003earXiv\u003c/em\u003e 2025, arXiv:2508.10925.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; et al. Qwen3 Technical Report. \u003cem\u003earXiv\u003c/em\u003e 2025, arXiv:2505.09388.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo, D.; Yang, D.; Zhang, H.; et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. \u003cem\u003eNature\u003c/em\u003e 2025, \u003cem\u003e645\u003c/em\u003e, 633\u0026ndash;638.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, A.; Yu, B.; Li, C.; Liu, D.; Huang, F.; Huang, H.; Jiang, J.; Tu, J.; Zhang, J.; Zhou, J.; et al. Qwen2.5-1M Technical Report. \u003cem\u003earXiv\u003c/em\u003e 2025, arXiv:2501.15383.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoogle. Google Colaboratory. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://colab.research.google.com/\u003c/span\u003e\u003cspan address=\"https://colab.research.google.com/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (accessed 2025-12-05)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRein, D.; Hou, B. L.; Stickland, A. C.; Petty, J.; Pang, R. Y.; Dirani, J.; Michael, J.; Bowman, S. R. GPQA: A graduate-level Google-proof Q\u0026amp;A benchmark. \u003cem\u003earXiv\u003c/em\u003e 2023, arXiv:2311.12022.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDettmers, T.; Pagnoni, A.; Holtzman, A.; Zettlemoyer, L. QLoRA: Efficient Finetuning of Quantized LLMs. \u003cem\u003earXiv\u003c/em\u003e 2023, arXiv:2305.14314.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"npj-computational-materials","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"npjcompumats","sideBox":"Learn more about [npj Computational Materials](http://www.nature.com/npjcompumats/)","snPcode":"41524","submissionUrl":"https://mts-npjcompumats.nature.com/","title":"npj Computational Materials","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"High-entropy material (HEM), large language model (LLM), AI, Prediction","lastPublishedDoi":"10.21203/rs.3.rs-8331266/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8331266/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eHigh-entropy materials (HEMs) offer unprecedented opportunities for superior mechanical, thermal, and catalytic properties, but their vast chemical space makes experimental discovery resource-intensive. State-of-the-art commercial large language models (LLMs) notably fail at HEM synthesizability prediction, a critical bottleneck in materials development. We demonstrate that domain-specific fine-tuning transforms open-weight local LLMs into accurate predictors. Using a dataset of 321,083 inorganic compositions with 2,560 HEM examples, we fine-tuned three 4-bit-quantized models (gpt-oss-20b, Qwen3-14b, and DeepSeek-R1-Distill-Qwen-14b), achieving remarkable balanced accuracy of 0.957, 0.961, and 0.956, respectively. Critically, these models operate efficiently on accessible hardware (\u0026lt; 15GB VRAM), eliminating costly API dependencies while ensuring data privacy and consistent reproducibility. This work could open new pathways toward autonomous closed-loop discovery, where distributed local models enable rapid screening and iterative improvement through experimental feedback. Future collaborative efforts in open data sharing, particularly including negative results, would address current fragmentation in synthesis reporting and accelerate community-wide HEM discovery.\u003c/p\u003e","manuscriptTitle":"Closed-Loop Workflow of High-Entropy Materials Discovery: Efficient and Accurate Synthesizability Prediction via Domain-Specific Local LLMs","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-12-19 18:45:38","doi":"10.21203/rs.3.rs-8331266/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-01-19T08:18:50+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-01-07T16:00:41+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-01-03T23:37:37+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"137076019595688963154571697607397772831","date":"2025-12-20T13:17:16+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"184715660909713685095912908647541300598","date":"2025-12-18T03:22:24+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"8068177565674785742893443838608712927","date":"2025-12-17T20:57:12+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-12-17T16:12:28+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-12-16T19:45:48+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-12-12T12:01:37+00:00","index":"","fulltext":""},{"type":"submitted","content":"npj Computational Materials","date":"2025-12-11T00:09:47+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"npj-computational-materials","isNatureJournal":false,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"npjcompumats","sideBox":"Learn more about [npj Computational Materials](http://www.nature.com/npjcompumats/)","snPcode":"41524","submissionUrl":"https://mts-npjcompumats.nature.com/","title":"npj Computational Materials","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"NPJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b4120df6-c4a2-49da-bff0-46fdd26a41af","owner":[],"postedDate":"December 19th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":59855303,"name":"Physical sciences/Chemistry"},{"id":59855304,"name":"Physical sciences/Materials science"},{"id":59855305,"name":"Physical sciences/Mathematics and computing"}],"tags":[],"updatedAt":"2026-04-27T16:04:29+00:00","versionOfRecord":{"articleIdentity":"rs-8331266","link":"https://doi.org/10.1038/s41524-026-02092-z","journal":{"identity":"npj-computational-materials","isVorOnly":false,"title":"npj Computational Materials"},"publishedOn":"2026-04-24 15:58:42","publishedOnDateReadable":"April 24th, 2026"},"versionCreatedAt":"2025-12-19 18:45:38","video":"","vorDoi":"10.1038/s41524-026-02092-z","vorDoiUrl":"https://doi.org/10.1038/s41524-026-02092-z","workflowStages":[]},"version":"v1","identity":"rs-8331266","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8331266","identity":"rs-8331266","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.