Identification and utilization of a novel major QTL and linked markers for enhancing protein content in peanut | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Identification and utilization of a novel major QTL and linked markers for enhancing protein content in peanut Mingjun Wang, Jianbin Guo, Gaorui Jin, Taihua Yang, Weigang Chen, and 8 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6240685/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 22 Aug, 2025 Read the published version in Theoretical and Applied Genetics → Version 1 posted 5 You are reading this latest preprint version Abstract Peanut is a vital source of protein for humans, playing a key role in maintaining a steady protein supply. In this study, the high protein cultivar Zhonghua6 (29.14±1.69%) was crossed with Xuhua13 (24.55±1.84%) to construct a recombinant inbred line (RIL) population. The protein content of the RIL population exhibited significant variation, ranging from 21.05% to 30.28%. To discover genomic regions associated with protein content, four libraries were constructed (two parents and two extreme bulks) for bulked segregant sequencing (BSA-seq). The results revealed significant associations between protein content and the genomic regions on chromosomes A07, A10, B03, and B10. Through linkage analysis, nine QTLs (quantitative trait loci) for protein content were identified, among which the major and stable QTL qPCB03 on chromosomes B03 (125.92-127.94 Mb) explained 11.03%-12.86% of the phenotypic variation. This QTL ( qPCB03 ), simultaneously identified by both BSA-Seq and linkage mapping, had not been documented in prior studies. Within the ~2Mb interval of qPCB03 , a total of 349 genomic variants were discovered, including six single nucleotide polymorphisms (SNPs) that resulted in nonsynonymous mutations in six genes. By assessing allelic effects in peanut germplasm and analyzing transcriptome sequencing data, five candidate genes ( Ah13g469400 , Ah13g473200 , Ah13g475000 , Ah13g477200 and Ah13g477400 ) were identified. According to the marker-assisted selection, favorable genotypes could potentially enhance protein content by 1.23% to 1.57% in the RIL populations. The identification of stable loci and the development of markers facilitate marker-assisted breeding in peanuts, while the discovery of candidate genes lays the groundwork for the fine mapping of key genes regulating protein content. peanut protein content QTL mapping candidate gene marker-assisted breeding Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Key message A novel major QTL qPCB03 for protein content was identified, and a KASP marker linked to qPCB03 was developed and validated to enhance protein content by 1.23% to 1.57% in the RIL populations. Introduction Peanut ( Arachis hypogaea L.) is an important food crop cultivated in more than 100 countries worldwide. In 2023, the global peanut cultivation area reached approximately 30.9 million hectares, yielding a total production of 54.3 million metric tons. China is the leading peanut-producing nation, with a harvested area exceeding 4.8 million hectares and an annual production of 19.2 million tons (FAOSTAT 2024 ). Peanuts contain generally 10.5–22% carbohydrates, 20–33% crude protein and 45%-55% oil as main compounds with larger genetic variation (Upadhyaya et al. 2006 ; Barkley et al. 2016 ; Davis and Dean 2016 ). In many developing countries, peanuts serve as an important source of protein and nutrients, aiding in the fight against malnutrition among impoverished populations (Linneman et al. 2007 ; Lazzerini et al. 2013 ; Matumba et al. 2014 ). With population growth, global protein supplies will face increasing pressure (Aiking 2011 ) and the demand for high-quality peanut as a source of protein is expected to increase significantly in the near future. Thus, improving protein content has become a primary objective of peanut breeding programs. Seed storage protein is a complex trait that is controlled by multiple genes and affected by the environment and genotype × environment interaction (Xu et al. 2012 ; Patil et al. 2017 ; Yang et al. 2023 ). Traditional breeding methods have been inefficient and time-consuming for crop improvement. Marker-assisted selection (MAS) has emerged as a more efficient and precise approach for genetic improvement in various crops, including peanuts (Janila et al. 2016 ; Varshney 2016 ; Pandey et al. 2020 ). Identification and mining of loci affecting peanut protein content could provide a technical guide for MAS in enhancing protein content of peanut seed. QTL mapping is a powerful technique for dissecting complex traits and identifying genomic regions linked to quantitative traits. Current QTL research on peanut has primarily focused on oil content, pod/seed size, and disease resistance (Kassie et al. 2023 ), with few studies on protein content. To date, limited studies have reported QTL identification for seed protein content in peanut. Sarvamangala et al. ( 2011 ) mapped six QTLs for protein content on a genetic map with 45 SSR loci in peanut. Sun and colleagues ( 2022 ) developed a high-density genetic map through whole-genome resequencing. They identified several QTLs associated with the oil, protein, and fatty acid content of peanut seeds, primarily located on chromosomes A05 and A08. However, there is currently a deficiency in the development of molecular markers that can facilitate marker-assisted selection tied to high protein content in peanuts. In this study, we integrated BSA-seq, linkage analysis and RNA-seq approaches to achieve the following objectives: (1) identification of stable major QTLs for protein content, (2) identification of the favorable alleles of candidate genes enhancing protein content in peanut, and (3) development of linked markers suitable for marker-assisted selection breeding. Materials and methods Plant materials and field trials The RIL population, consisting of 257 lines, was derived from a cross between ZH6 and XH13 through the single seed descent method. The generations F 13 –F 14 of the RIL population and parents were planted in three environments, i.e., 2022Yangluo (2022YL), 2023Yangluo (2023YL) and 2023Wuchang (2023WC). In addition, a total of 147 peanut germplasms with diverse protein content were collected from different provinces and phenotyping data in these germplasms was also measured in 2022, Yangluo. The generations F 9 of the RIL population derived from ZH6 and Jihua5 (JH5) were planted in 2024Anyang (2024AY). Each line (approximately 10 plants) was planted in one row, with 2.5 m row length and 0.3 m row spacing. A randomized complete block design was performed with three replicates, and field management followed the standard agricultural practices. Determination of seed protein content and statistical analysis The mature peanut seeds were selected for determination of protein content using the Kjeldahl method. An automatic graphite digestion instrument and a Kjeldahl apparatus (K9860; Hanon Instrument Co., Ltd., Jinan, China) were used for sample digestion and nitrogen determination. Ground samples were weighed (~ 0.3g) and transferred into the digestion instrument containing 1.0 g of catalyst (prepared by mixing 0.6 g of K 2 SO 4 and 0.4 g of CuSO 4 ·5H 2 O) and 10 mL of concentrated H 2 SO 4 for ~ 3.5 hours digestion. The nitrogen content was converted to protein content according to the peanut protein conversion factor (5.46) (Sathe et al. 2009 ). Statistical analysis for the phenotypic data and analysis of variance were performed using IBM SPSS Statistics software version 26 ( https://www.ibm.com/analytics/spss-statistics-software ). The normal distribution of phenotypic data was confirmed by Shapiro-Wilk test. Bulk segregant analysis The lines (19) with high protein content and the lines (19) with low protein content were selected to construct high protein content bulk (bulk-HP) and low protein content bulk (bulk-LP), respectively. The genomic DNA of the lines and the two parents was extracted from the young leaves using Plant Genomic DNA Kit (TianGen, China). Then DNA libraries of two parents and two bulks were constructed and sequenced on an Illumina HiSeq 4000 platform. After removal of low-quality reads, short reads and adapter sequences by Trimmomatic software (Bolger et al. 2014 ), the quality filter reads were aligned to the cultivated peanut reference genome (Bertioli et al. 2019 ) (Tifrunner.gnm2.J5K5; http://www.peanutbase.com ) using Burrows-Wheeler Aligner software (Li and Durbin 2009 ) ( http://bio-bwa.sourceforge.net/ ). Uniquely mapped and properly paired reads were selected with SAMtools (Li et al. 2009 ) ( http://samtools.sourceforge.net/ ). SNPs and short insertions/deletions (InDels) were identified with Genome Analysis Toolkit (GATK) (McKenna et al. 2010 ) ( https://software.broadinstitute.org/gatk/ ). Annotation of variants was performed using ANNOVAR (Wang et al. 2010 ) ( http://www.openbioinformatics.org/annovar/ ). By subtracting SNP-index of bulk-LP from SNP‐index of bulk-HP, the ΔSNP‐index was calculated with a 2 Mb genome window using the R package QTLseqr (Mansfeld and Grumet 2018 ). The details on whole genome re-sequencing data for BSA-seq were shown in Table S1 . The raw sequencing data has been submitted to the National Genomics Data Center ( https://ngdc.cncb.ac.cn/ ) under BioProjectID PRJCA036928. Genotyping and genetic linkage map construction The genomic DNA of the RILs and the parents was extracted from the young leaves using Plant Genomic DNA Kit (TianGen, China). Each line was genotyped using the 50K Genotyping by the Target Sequencing (GBTS) SNP array (MolBreeding Biotechnology Co., Ltd., Shijiazhuang, China). Raw reads were filtered using fastp software ( https://github.com/OpenGene/fastp ). Clean reads were aligned to the cultivated peanut reference genome (Bertioli et al. 2019 ) (Tifrunner.gnm2.J5K5; http://www.peanutbase.com ) using Burrows-Wheeler Aligner (BWA) software (Li and Durbin 2009 ). SNPs and InDels were identified and filtered using UnifiedGenotyper and VariantFiltration commands in GATK (McKenna et al. 2010 ). After excluding the monomorphic variants, polymorphic SNP markers were binned using the BIN function in QTL IciMapping (Meng et al. 2015 ). One marker in each bin was randomly selected to construct the genetic map. Bin markers were organized into corresponding groups with LOD scores ranging from 2 to 10 using JoinMap 4 software (Ooijen et al. 2006) ( https://www.kyazma.nl/index.php/JoinMap/ ). The Kosambi mapping function was employed to evaluate the recombination rate and to calculate the genetic distance between bin markers. Flanking sequences of SNPs were used as queries in a BLAST search against the reference genome to determine their physical locations. Linkage maps were drawn using R package LinkageMapView (Ouellette et al. 2018 ). QTL analysis Genome-wide QTL mapping was conducted using the mean value of protein content across environments. QTLs were detected using Windows QTL Cartographer 2.5 software (Wang et al. 2012 ) ( http://statgen.ncsu.edu/qtlcart/WQTLCart.htm ), applying the composite interval mapping (CIM) method with LOD score values ≥ 2.5. The parameters for control markers, window size, and walking speed were set to 5, 10, and 2 cM, respectively. QTLs were named starting with the letter “q”, followed by the abbreviation of trait name “PC” and the corresponding linkage group. An arabic numeral was added after the linkage group if two or more QTLs were identified in the same linkage group. QTLs detected in more than one environment and phenotypic variant explained > 10% were considered to be major QTLs. RNA-sequencing analysis Two lines with large differences in protein content from RIL populations, namely, QT696 (a high protein content line from bulk-HP) and QT937 (a low protein content line from bulk-LP), were used for transcriptome sequencing and candidate genes analysis. According to the previous studies (Wan et al. 2016 ; Wang et al. 2022 ), developing peanut seeds were collected from approximately 10 plants at five developmental stages (Ⅰ, Ⅱ, Ⅲ, Ⅳ and Ⅴ), corresponding to 20, 30, 40, 50 and 60 days after flowering, respectively. Each stage was represented by three biological replicates. Total RNA was extracted using FastPure Plant Total RNA Isolation Kit RC401-01 (Vazyme, Nanjing, China). The constructed libraries were sequenced on an Illumina Novaseq platform (Novogene Technology Co., Ltd., Beijing, China) and 150 bp paired-end reads were generated. Clean reads were aligned to the cultivated peanut reference genome using Hisat2 v2.0.5 (Kim et al. 2015 ; Bertioli et al. 2019 ). The Gene expression levels were quantified and normalized as Fragments Per Kilobase of transcript per Million mapped reads (FPKM) using StringTie version 1.2.3 (Pertea et al. 2015 ). Differentially expressed genes (DEGs) were identified through DESeq2 version 1.18.1 (Anders and Huber 2010 ) based on criteria that included a minimum FPKM value of 1 in at least one sample, a fold change of at least 2, and an adjusted P -value of 0.05 or lower. TBtools was used for generating gene expression heatmap (Chen et al. 2020a ). The raw sequencing data of RNA-seq have been uploaded to the National Genomics Data Center ( https://ngdc.cncb.ac.cn/ ) under BioProjectID PRJCA036946. Result Phenotypic variation of the RIL population Phenotypic measurements were conducted for 257 recombinant inbred lines (RILs) and two parents across three environments. As shown in Table 1, the male parent ZH6 exhibited consistently higher protein content (27.20% − 30.31%) compared to the female parent XH13 (23.06% − 26.61%) across all three environments. The protein content of the RIL population showed large phenotypic variations ranging from 22.77–30.28% in 2022YL, from 21.05–29.72% in 2023YL, and from 21.56–30.27% in 2023WC. The phenotypic values showed a continuous distribution with transgressive segregation in the RIL population (Fig. 1). The Shapiro-Wilk test was indicated that the phenotypic data of the RIL population were normally distributed in 2022YL and 2023WC, but not in 2023YL (Table 1). Variance analysis was revealed that genetic and environmental factors, as well as genotype × environment interaction significantly influenced protein content in the RIL population (Table 2). Broad-sense heritability for protein content was estimated to be 0.743 (Table 2). Identification of candidate genomic regions associated with protein content using BSAseq The average protein content in the RIL population ranged from 22.02–29.44% across three environments. RILs with the highest protein content (27.52% − 29.31%) were selected to construct the extremely high protein content bulk (bulk-HP), while the RILs with the lowest protein content (22.06% − 25.62%) were used to form extremely low protein content bulk (bulk-LP). The DNA libraries of two parents and two bulks were subjected to whole-genome resequencing using an Illumina HiSeq platform, producing 241.18 million reads (71.45 Gb) for XH13, 212.77 million reads (56.89 Gb) for ZH6, 353.26 million reads (104.36 Gb) for bulk-HP and 343.46 million reads (101.76 Gb) for bulk-LP, respectively. The reads of the parents and two bulks were mapped to the reference genome of cultivated peanut, achieving coverage rates of 93.87% with a depth of 24.03X for XH13, 92.89% coverage with a depth of 19.05X for ZH6, 94.32% coverage with a depth of 34.01X for bulk-HP, and 94.96% coverage with a depth of 33.74X for bulk-LP (Table S1 ). A total of 275,570 genomic variants were identified between bulk-HP and bulk-LP, which included 241,607 SNPs and 33,963 InDels (Table S2 ). To identify the candidate genomic regions controlling protein content, the SNP-index was calculated for each bulk using ZH6 as a reference genome. The ΔSNP-index were calculated by subtracting SNP‐index of bulk-LP from SNP‐index of bulk-HP (Fig. S1 ). Four genomic regions have been identified in association with protein content. These included qPCA07 on chromosome A07 (66.62 Mb − 73.67 Mb), qPCA10 on chromosome A10 (11.35 Mb − 31.17 Mb), qPCB03 on chromosome B03 (124.98 Mb to 128.76 Mb), and qPCB10 on chromosome B10 (136.97 Mb − 138.34 Mb) (Fig. 2; Table 3). The ΔSNP‐index of the four genomic regions was positive, indicating that favorable alleles were from the reference parent ZH6. Identification of QTLs for protein content through linkage analysis The RIL population and two parents were genotyped using the cultivated peanut 50K GBTS SNP array. After excluding the monomorphic loci and loci with partial segregation, a high-density genetic map was constructed. It is consisted of 1881 bin markers distributed on 22 linkage groups, with an average genetic distance of 1.03 cM (Fig. S2 ; Table S3; Table S4). A genome-wide QTL analysis was performed using the genetic map and phenotypic data on protein content across three environments. A total of nine QTLs, with logarithm of odds (LODs) ranging from 2.88 to 7.55, were mapped on six linkage groups (Table 4). The QTLs explained 3.84–12.86% of the phenotypic variance for protein content, with the additive effect values from − 0.33 to 0.82 (Table 4). Among them, qPCB03 was consistently detected in three environments, exhibiting the highest phenotypic variation explained (PVE) of 11.03–12.86%. Additionally, the physical interval of qPCB03 (125.92Mb-127.94Mb) was found to colocate with the candidate genomic region detected by BSA-Seq (Fig. 3). The results indicated that the genomic region on B03 (125.92Mb-127.94Mb) harbored a major and stable locus that regulates protein content. Putative candidate genes for QTL qPCB03 A combined analysis of genomic variation data and gene expression was conducted to identify candidate genes in ~ 2Mb genomic region of qPCB03 . In the physical interval of qPCB03 , 349 genomic variants (301 SNPs and 48 InDels) were detected between ZH6 and XH13 (Table S5). Among the variants, 318 were in intergenic regions, 10 were in introns, 6 non-synonymous were in exons, 15 were in 2kb-upstream/downstream regions. The 6 non-synonymous SNPs were found to affect 6 candidate genes for protein content. For gene expression analysis, developing seeds of the bulk-HP line (QT696) and the bulk-LP line (QT937) at five stages (20, 30, 40, 50 and 60 days after flowering) were collected to perform transcriptome sequencing. The protein content of developing seeds in QT696 was significantly higher than that in QT937 at five stages (Fig. S3). A total of 30 RNA-seq libraries were constructed and 5.68-7.00 Gb clean data per library was obtained (Table S6). On average, 96.97% of clean reads were mapped on the reference genome, and 40,746 to 51,848 genes were identified in each library. Within the physical interval of qPCB03 , 43 genes were found to be expressed in developing seeds (Table S7). Among them, four genes ( Ah13g469400 , Ah13g473200 , Ah13g475000 , and Ah13g477400 ) exhibited nonsynonymous mutants between two parents (Fig. 4A). Furthermore, 22 DEGs were identified in the region of qPCB03 , including one DEG ( Ah13g477200 ) that harbored a variant in the 2kb-upstream (promoter) region (Fig. 4A). Based on the analysis of genomic variant and gene expression, five genes ( Ah13g469400 , Ah13g473200 , Ah13g475000 , Ah13g477200 , and Ah13g477400 ) were identified as candidate genes for the stable QTL qPCB03 (Table 5). To confirm the allelic effects of candidate genes on protein content in peanut germplasm, KASP markers were designed based on five SNP sites of the candidate genes (Table S8). These markers were subsequently utilized to genotype 147 Chinese peanut cultivars which exhibited significant phenotypic variation, with protein content ranging from 24.23–33.58% in the 2022YL. The KASP assays effectively distinguished the allelic variations in the peanut germplasm (Fig. S4), and the average protein content between allelotypes showed statistically significant differences at all five SNP sites (Fig. 4B). These results indicated that the five genes are robust candidates influencing protein content in peanuts. Validation of KASP markers linked to qPCB03 The KASP marker KASP_B03_127452476 , exhibiting the lowest P -value in allelic comparison (Fig. 4B), was chosen for assessing its potential value in breeding programs. Two RIL populations derived from the same high-protein male parent ZH6, ZX_RIL (ZH6 × XH13) and ZJ_RIL (ZH6 × JH5), were genotyped using the KASP_B03 _ 127452476 marker. The high-protein genotype ( B03_127452476_G ) originated from ZH6, while the low-protein genotype ( B03_127452476_A ) was derived from XH13 or JH5. In ZX_RIL population, RILs with the B03_127452476_G/G genotype exhibited significantly higher protein content (27.11 ± 1.39% in 2022YL, 26.83 ± 1.23% in 2023YL, and 26.95 ± 1.57% in 2023WC) compared to RILs with the B03_127452476_A/A genotype (26.11 ± 1.29% in 2022YL, 25.77 ± 1.44% in 2023YL, and 25.71 ± 1.53% in 2023WC). Selection of the favorable allele in ZX_RIL population could lead to an increase in protein content by 1.23–1.57% across the three environments (Fig. 5A). In ZJ_RIL population, the protein content ranged from 19.40–33.78% in 2024AY. RILs with the B03_127452476_G/G genotype demonstrated significantly higher protein content (26.48 ± 3.03% in 2024AY) compared to those with the B03_127452476_A/A genotype (24.95 ± 2.89% in 2024AY) (Fig. 5B). Selecting the favorable allele could enhance protein content by 1.53% in ZJ_RIL population. These findings suggested the potential value of the KASP marker in marker-assisted selection of peanut genotypes associated with high protein content. Discussion A novel and major QTL for seed protein content in peanut The cultivation of high-protein peanut is a primary objective in breeding program. However, limited studies have reported on the identification of QTLs for seed protein content in peanut. In this study, nine QTLs for protein content were identified through linkage analysis. Two of them, qPCA05.1 and qPCA08 , overlapped with or were close to the QTLs for protein content reported by Sun et al. ( 2022 ), indicating they are reliable loci for regulating protein content. In addition, qPCA07 , qPCB03 and qPCB10.2 were repeatedly detected in at least two environments. Notably, qPCB03 were colocated on the genomic regions detected by BSA-Seq. Since qPCB03 exhibited the highest PVE of 11.03–12.86% across three environments and has not been previously reported, it is considered as a novel major QTL for protein content. Five candidate genes were identified in the genomic region of qPCB03 (~ 2 Mb), and their allelic effects were shown to be significantly associated with protein content in peanut germplasm. Based on the SNP marker corresponding to the candidate gene, the selection of favorable genotypes could potentially boost protein content by 1.23–1.57% in RIL populations. The results indicate that qPCB03 is a reliable major QTL for seed protein content, as well as a valuable locus in MAS for breeding high seed protein content in peanut. Genetic correlation between protein content and oil content It is widely recognized that there is an inverse relationship between protein and oil content in seeds. This correlation has been evident in the identification of major QTLs for protein and oil content, which share overlapping regions (Sarvamangala et al. 2011 ; Sun et al. 2022 ; Duan et al. 2023 ). In this study, we found a significant negative correlation between protein and oil contents (correlation coefficient = -0.766, Table S9). Additionally, the linkage analysis revealed that QTLs for both protein and oil content were overlapped on the genetic map (Table S10). Interestingly, the KASP marker ( KASP_B03_127452476 ) for protein content could be employed in the selection of low-oil genotype. The theoretical oil content of low-oil genotype could decrease by 0.85–1.31% in ZX_RIL population and by 1.60% in ZJ_RIL population (Fig. S5). These findings suggest that qPCB03 with linked KASP marker could facilitate marker-assisted selection of peanut varieties with high protein content and low oil content simultaneously. Prediction of candidate genes associated with protein content The high-resolution genetic map enabled the demarcation of qPCB03 to a 2 Mb interval containing 78 genes. In this study, candidate genes were screened out through a combined analysis of genomic variation, gene expression, and allelic effects of candidates in peanut germplasm. Finally, five of 78 genes ( Ah13g469400 , Ah13g473200 , Ah13g475000 , Ah13g477200 and Ah13g477400 ) were deduced to be candidate genes for qPCB03 . Four candidate genes ( Ah13g469400 , Ah13g473200 , Ah13g475000 and Ah13g477400 ) had nonsynonymous SNVs that encode a violaxanthin de-epoxidase-related (VDR), nucleolin-like isoform 2 protein, aspartate/prephenate aminotransferase and CAAX protease self-immunity protein, respectively. One candidate gene ( Ah13g477200 ) harbored a SNP in its upstream region and was differentially expressed in developing seeds. It encodes a trehalose-6-phosphate phosphatase. Notably, violaxanthin de-epoxidase (VDE) is involved in photoprotective response to high-light stress and abscisic acid (ABA) biosynthesis (Demmig et al. 1987 ; Ye et al. 2012 ; Chen et al. 2020b ). Photosynthesis and nitrogen metabolism are closely linked by their interdependence for fixed carbon, chemical energy, and nitrogen assimilates (Paul and Pellny 2003 ; Zheng 2009 ). ABA has been reported to be involved in nitrogen remobilization and seed storage protein synthesis (Zheng et al. 2019 ; Yang et al. 2022 ; Nan et al. 2023 ; Yang et al. 2023 ). Aspartate aminotransferase has been reported to play an important role in carbon and nitrogen metabolism (de la Torre et al. 2006 ; de la Torre et al. 2014 ) and enhanced aspartate aminotransferase enzyme activity would increase total free amino acid contents in rice seeds (Zhou et al. 2009 ). Although evidence suggests that these putative genes may be involved in regulating protein content, further studies are necessary to validate their functions. In summary, a novel major QTL qPCB03 for protein content was identified using both BSA-Seq and linkage mapping. A total of five candidate genes were discovered in genomic region of qPCB03 and their allelic effect were found to be significantly associated with protein content. A KASP marker linked to qPCB03 was developed and validated to enhance protein content by 1.23–1.57% in the RIL populations. The results provide a valuable locus with the linked marker in MAS for breeding high seed protein content in peanut. Declarations Acknowledgments Not applicable Author contributions MW, JG, BL, HJ, NL and YL conceived and designed the research. HJ and YC developed the RIL population. JG, GJ, TY, WC, LH, HL and XZ planted the materials and conducted field management. MW and JG performed the measurement of protein content and statistical analysis of the phenotyping data. MW, JG and NL performed BSA-seq, QTL analysis, and development and validation of the marker associated with protein content. MW and NL performed SNP analysis. MW and NL wrote the manuscript, HJ and YL revised the manuscript and improved the English writing. All the authors read and approved the final manuscript. Funding This work was supported by the National Key Research and Development Program of China (2022YFD1200400), the National Peanut Industry Technology System Construction, China (CARS13), the National Crop Germplasm Resources Center (NCGRC-2024-036), the National Program for Crop Germplasm Protection of China (19210163), the Agricultural Science and Technology Innovation Program of Chinese Academy of Agricultural Sciences (CAAS-ASTIP-2024-OCRI). Data availability All data supporting the results of this study are available within the paper and its supplementary data published online. Conflict of interest The authors declare no conflicts of interest. Ethical standards The authors state that all experiments in the study comply with the ethical standards. References Aiking H (2011) Future protein supply. Trends Food Sci Technol 22:112-120 Anders S, Huber W (2010) Differential expression analysis for sequence count data. Genome Biol 11:R106 Barkley NA, Upadhyaya HD, Liao B, Holbrook CC (2016) Chapter 3 - Global resources of genetic diversity in peanut. In: Stalker HT, F. Wilson R (eds) Peanuts. AOCS Press, pp 67-109 Bertioli DJ, Jenkins J, Clevenger J, Dudchenko O, Gao D, Seijo G, Leal-Bertioli SCM, Ren L, Farmer AD, Pandey MK, Samoluk SS, Abernathy B, Agarwal G, Ballén-Taborda C, Cameron C, Campbell J, Chavarro C, Chitikineni A, Chu Y, Dash S, El Baidouri M, Guo B, Huang W, Kim KD, Korani W, Lanciano S, Lui CG, Mirouze M, Moretzsohn MC, Pham M, Shin JH, Shirasawa K, Sinharoy S, Sreedasyam A, Weeks NT, Zhang X, Zheng Z, Sun Z, Froenicke L, Aiden EL, Michelmore R, Varshney RK, Holbrook CC, Cannon EKS, Scheffler BE, Grimwood J, Ozias-Akins P, Cannon SB, Jackson SA, Schmutz J (2019) The genome sequence of segmental allotetraploid peanut Arachis hypogaea . Nat Genet 51:877-884 Bolger AM, Lohse M, Usadel B (2014) Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 30:2114-2120 Chen C, Chen H, Zhang Y, Thomas HR, Frank MH, He Y, Xia R (2020a) TBtools: an integrative toolkit developed for interactive analyses of big biological data. Mol Plant 13:1194-1202 Chen K, Li GJ, Bressan RA, Song CP, Zhu JK, Zhao Y (2020b) Abscisic acid dynamics, signaling, and functions in plants. J Integr Plant Biol 62:25-54 Davis JP, Dean LL (2016) Chapter 11 - Peanut composition, flavor and nutrition. In: Stalker HT, F. Wilson R (eds) Peanuts. AOCS Press, pp 289-345 de la Torre F, Cañas RA, Pascual MB, Avila C, Cánovas FM (2014) Plastidic aspartate aminotransferases and the biosynthesis of essential amino acids in plants. J Exp Bot 65:5527-5534 de la Torre F, De Santis L, Suárez MF, Crespillo R, Cánovas FM (2006) Identification and functional analysis of a prokaryotic-type aspartate aminotransferase: implications for plant amino acid metabolism. Plant J 46:414-425 Demmig B, Winter K, Krüger A, Czygan FC (1987) Photoinhibition and zeaxanthin formation in intact leaves : a possible role of the xanthophyll cycle in the dissipation of excess light energy. Plant Physiol 84:218-224 Duan Z, Li Q, Wang H, He X, Zhang M (2023) Genetic regulatory networks of soybean seed size, oil and protein contents. Front Plant Sci 14:1160418 FAOSTAT (2024) Statistical database FAOSTAT. http://faostat3.fao.org Janila P, Variath MT, Pandey MK, Desmae H, Motagi BN, Okori P, Manohar SS, Rathnakumar AL, Radhakrishnan T, Liao B, Varshney RK (2016) Genomic tools in groundnut breeding program: status and perspectives. Front Plant Sci 7:289 Kassie FC, Nguepjop JR, Ngalle HB, Assaha DVM, Gessese MK, Abtew WG, Tossim HA, Sambou A, Seye M, Rami JF, Fonceka D, Bell JM (2023) An overview of mapping quantitative trait loci in peanut ( Arachis hypogaea L .). Genes 14:1176 Kim D, Langmead B, Salzberg SL (2015) HISAT: a fast spliced aligner with low memory requirements. Nat Methods 12:357-360 Lazzerini M, Rubert L, Pani P (2013) Specially formulated foods for treating children with moderate acute malnutrition in low- and middle-income countries. Cochrane Database Syst Rev Cd009584 Li H, Durbin R (2009) Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25:1754-1760 Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R (2009) The sequence alignment/map format and SAMtools. Bioinformatics 25:2078-2079 Linneman Z, Matilsky D, Ndekha M, Manary MJ, Maleta K, Manary MJ (2007) A large-scale operational study of home-based therapy with ready-to-use therapeutic food in childhood malnutrition in Malawi. Matern Child Nutr 3:206-215 Mansfeld BN, Grumet R (2018) QTLseqr: An R package for bulk segregant analysis with next-generation sequencing. Plant Genome 11 Matumba L, Monjerezi M, Biswick T, Mwatseteza J, Makumba W, Kamangira D, Mtukuso A (2014) A survey of the incidence and level of aflatoxin contamination in a range of locally and imported processed foods on Malawian retail market. Food Control 39:87-91 McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A, Garimella K, Altshuler D, Gabriel S, Daly M, DePristo MA (2010) The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res 20:1297-1303 Meng L, Li H, Zhang L, Wang J (2015) QTL IciMapping: integrated software for genetic linkage map construction and quantitative trait locus mapping in biparental populations. Crop J 3:269-283 Nan Y, He H, Xie Y, Li C, Atif A, Hui J, Tian H, Gao Y (2023) The responses of genotypes with contrasting NUtE to exogenous ABA during the flowering stage in Brassica napus. Plant Stress 10:100248 Ooijen JWv, Ooijen JWv, Verlaat Jvt, Ooijen JWv, Tol J, Dalén J, Buren JBV, Meer JWMvd, Krieken JHv, Ooijen JWv, Kessel JSV, Van O, Voorrips RE, Heuvel LP (2006) JoinMap 4, Software for the calculation of genetic linkage maps in experimental populations. Kyazma B.V., Wageningen, Netherlands. Ouellette LA, Reid RW, Blanchard SG, Brouwer CR (2018) LinkageMapView-rendering high-resolution linkage and QTL maps. Bioinformatics 34:306-307 Pandey MK, Pandey AK, Kumar R, Nwosu CV, Guo B, Wright GC, Bhat RS, Chen X, Bera SK, Yuan M, Jiang H, Faye I, Radhakrishnan T, Wang X, Liang X, Liao B, Zhang X, Varshney RK, Zhuang W (2020) Translational genomics for achieving higher genetic gains in groundnut. Theor Appl Genet 133:1679-1702 Patil G, Mian R, Vuong T, Pantalone V, Song Q, Chen P, Shannon GJ, Carter TC, Nguyen HT (2017) Molecular mapping and genomics of soybean seed protein: a review and perspective for the future. Theor Appl Genet 130:1975-1991 Paul MJ, Pellny TK (2003) Carbon metabolite feedback regulation of leaf photosynthesis and development. J Exp Bot 54:539-547 Pertea M, Pertea GM, Antonescu CM, Chang TC, Mendell JT, Salzberg SL (2015) StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat biotechnol 33:290-295 Sarvamangala C, Gowda MVC, Varshney RK (2011) Identification of quantitative trait loci for protein content, oil content and oil quality for groundnut ( Arachis hypogaea L .). Field Crops Res 122:49-59 Sathe SK, Venkatachalam M, Sharma GM, Kshirsagar HH, Teuber SS, Roux KH (2009) Solubilization and electrophoretic characterization of select edible nut seed proteins. J Agric Food Chem 57:7846-7856 Sun Z, Qi F, Liu H, Qin L, Xu J, Shi L, Zhang Z, Miao L, Huang B, Dong W, Wang X, Tian M, Feng J, Zhao R, Zheng Z, Zhang X (2022) QTL mapping of quality traits in peanut using whole-genome resequencing. Crop J 10:177-184 Upadhyaya HD, Reddy LJ, Gowda CLL, Singh S (2006) Identification of diverse groundnut germplasm: sources of early maturity in a core collection. Field Crops Res 97:261-271 Varshney RK (2016) Exciting journey of 10 years from genomes to fields and markets: some success stories of genomics-assisted breeding in chickpea, pigeonpea and groundnut. Plant Sci 242:98-107 Wan L, Li B, Pandey MK, Wu Y, Lei Y, Yan L, Dai X, Jiang H, Zhang J, Wei G, Varshney RK, Liao B (2016) Transcriptome analysis of a new peanut seed coat mutant for the physiological regulatory mechanism involved in seed coat cracking and pigmentation. Front Plant Sci 7:1491 Wang K, Li M, Hakonarson H (2010) ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res 38:e164 Wang S, Basten C, Zeng Z (2012) Windows QTL Cartographer v2.5. Department of Statistics, North Carolina State University; Raleigh, NC. Wang Z, Yan L, Chen Y, Wang X, Huai D, Kang Y, Jiang H, Liu K, Lei Y, Liao B (2022) Detection of a major QTL and development of KASP markers for seed weight by combining QTL-seq, QTL-mapping and RNA-seq in peanut. Theor Appl Genet 135:1779-1795 Xu G, Fan X, Miller AJ (2012) Plant nitrogen assimilation and use efficiency. Annu Rev Plant Biol 63:153-182 Yang T, Wang H, Guo L, Wu X, Xiao Q, Wang J, Wang Q, Ma G, Wang W, Wu Y (2022) ABA-induced phosphorylation of basic leucine zipper 29, ABSCISIC ACID INSENSITIVE 19, and Opaque2 by SnRK2.2 enhances gene transactivation for endosperm filling in maize. Plant Cell 34:1933-1956 Yang T, Wu X, Wang W, Wu Y (2023) Regulation of seed storage protein synthesis in monocot and dicot plants: a comparative review. Mol Plant 16:145-167 Ye N, Jia L, Zhang J (2012) ABA signal in rice under stress conditions. Rice 5:1 Zheng X, Li Q, Li C, An D, Xiao Q, Wang W, Wu Y (2019) Intra-Kernel reallocation of proteins in maize depends on VP1-mediated scutellum development and nutrient assimilation. Plant Cell 31:2613-2635 Zheng ZL (2009) Carbon and nitrogen nutrient balance signaling in plants. Plant Signal Behav 4:584-591 Zhou Y, Cai H, Xiao J, Li X, Zhang Q, Lian X (2009) Over-expression of aspartate aminotransferase genes in rice resulted in altered nitrogen metabolism and increased amino acid content in seeds. Theor Appl Genet 118:1381-1390 Tables Tables 1-5 are available in the Supplementary Files section. Supplementary Files SupplementaryFigures.docx SupplementaryTable.xlsx Table1.xlsx Table 1 Description of phenotype analysis of protein content in the RIL population Table2.xlsx Table 2 Analysis of variance for protein content across three environments Table3.xlsx Table 3 The genomic regions association with protein content through BSA-Seq Table4.xlsx Table 4 The information on identified QTLs in three environments Table5.xlsx Table 5 The detailed information on five candidate genes Cite Share Download PDF Status: Published Journal Publication published 22 Aug, 2025 Read the published version in Theoretical and Applied Genetics → Version 1 posted Editorial decision: Major revisions 20 May, 2025 Reviewers agreed at journal 22 Mar, 2025 Reviewers invited by journal 21 Mar, 2025 Editor assigned by journal 17 Mar, 2025 First submitted to journal 16 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6240685","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":432304378,"identity":"ba8e8e29-78da-492c-bcaa-e60cb242d870","order_by":0,"name":"Mingjun Wang","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Mingjun","middleName":"","lastName":"Wang","suffix":""},{"id":432304379,"identity":"13025405-ce58-4832-a569-725b2381a192","order_by":1,"name":"Jianbin Guo","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Jianbin","middleName":"","lastName":"Guo","suffix":""},{"id":432304380,"identity":"5e6a655d-5cc0-4bea-977b-81b1693c0028","order_by":2,"name":"Gaorui Jin","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Gaorui","middleName":"","lastName":"Jin","suffix":""},{"id":432304381,"identity":"5270bc97-2ea1-45f0-bfe4-a7a7b3103456","order_by":3,"name":"Taihua Yang","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Taihua","middleName":"","lastName":"Yang","suffix":""},{"id":432304382,"identity":"46f616df-77db-491b-a07d-8f464b0c33e9","order_by":4,"name":"Weigang Chen","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Weigang","middleName":"","lastName":"Chen","suffix":""},{"id":432304383,"identity":"7526e635-7d09-46cd-97aa-b8cb77a0a8c8","order_by":5,"name":"Yuning Chen","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Yuning","middleName":"","lastName":"Chen","suffix":""},{"id":432304384,"identity":"b050cef5-ef3b-433a-8ba4-817283e2ddc6","order_by":6,"name":"Li Huang","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Li","middleName":"","lastName":"Huang","suffix":""},{"id":432304385,"identity":"e0b90998-2e2e-4c31-a4e6-035f809e7b70","order_by":7,"name":"Huaiyong Luo","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Huaiyong","middleName":"","lastName":"Luo","suffix":""},{"id":432304386,"identity":"85a2334a-f7fa-46ea-a270-41e6e55b5f66","order_by":8,"name":"Xiaojing Zhou","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Xiaojing","middleName":"","lastName":"Zhou","suffix":""},{"id":432304387,"identity":"606681f0-9557-4ce8-9bf8-455959926cab","order_by":9,"name":"Boshou Liao","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Boshou","middleName":"","lastName":"Liao","suffix":""},{"id":432304388,"identity":"e0dbb750-7109-40b1-9336-a01cef80fce0","order_by":10,"name":"Huifang Jiang","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Huifang","middleName":"","lastName":"Jiang","suffix":""},{"id":432304389,"identity":"d138db2b-75d3-4d10-b9cb-686873b946db","order_by":11,"name":"Nian Liu","email":"","orcid":"","institution":"Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":false,"prefix":"","firstName":"Nian","middleName":"","lastName":"Liu","suffix":""},{"id":432304390,"identity":"15030d08-f4b4-4eaa-8cef-cf58d7ef01ec","order_by":12,"name":"Yong Lei","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAuklEQVRIiWNgGAWjYBACPgbGBhAtR7wWNqgWY1K0QEBiA/Fa2JvbJD7uqE3vbz/8+ANDzR0itPAcbJOceeZ47owzaWYSDMeeEaFFIrFNmrftWO4GhgQzoL8OE6FF/iFYS7oB//PPH4jTIsEI0lKTYCCRYyBBnBaexGbLmW0HDGfceFMmkXCMCC387Mcf3vjYVifP35+++cOHGiK0AAGLBAMDVGUCURoYGJg/MDDUEal2FIyCUTAKRiQAAFJEOLMiWuHSAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0002-3738-4197","institution":"OCRI CAAS: Oil Crops Research Institute Chinese Academy of Agricultural Sciences","correspondingAuthor":true,"prefix":"","firstName":"Yong","middleName":"","lastName":"Lei","suffix":""}],"badges":[],"createdAt":"2025-03-17 03:51:45","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6240685/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6240685/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s00122-025-05000-z","type":"published","date":"2025-08-22T16:29:44+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":79569770,"identity":"c9aca2ed-e9bc-4945-9529-17dbf5c73bba","added_by":"auto","created_at":"2025-03-31 10:18:49","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":191216,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePhenotypic distributions of protein content in the RIL population under three environments\u003c/strong\u003e. The x-axis denotes the protein content of peanut seeds (%), and the y-axis denotes the number of lines.\u003c/p\u003e","description":"","filename":"Fig.1.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/d7bf0486bb4dc402015001f5.jpg"},{"id":79570369,"identity":"31fc0a99-c270-4932-8b07-b93a8962dbdb","added_by":"auto","created_at":"2025-03-31 10:26:49","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":538355,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGenomic regions associated with protein content identified by BSA-seq\u003c/strong\u003e. The plot illustrates the ΔSNP-index using ZH6 as the reference parent. The gray lines represent the statistical confidence intervals at a significant level of 5e-5. The blue intervals represent the significant regions associated with protein content.\u003c/p\u003e","description":"","filename":"Fig.2.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/72c51d49244d3066cf9ff938.jpg"},{"id":79569772,"identity":"60e1ff0b-cc0e-4586-9205-ad9ec12eb934","added_by":"auto","created_at":"2025-03-31 10:18:49","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":592013,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverviews of the consistent genomic region of \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eqPCB03 \u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003edetected by BSA-Seq (A) and linkage mapping approaches (B)\u003c/strong\u003e. QTLs detected in different environments are denoted as the different coloured boxes. The markers linked to QTLs are highlighted in red.\u003c/p\u003e","description":"","filename":"Fig.3.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/9069ae403d226fdac810f59c.jpg"},{"id":79570370,"identity":"95a17e35-f49c-420e-b959-b467bdfe5aab","added_by":"auto","created_at":"2025-03-31 10:26:49","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1309627,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eIllustration of candidate genes within the intervals of the stable QTL \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eqPCB03\u003c/strong\u003e\u003c/em\u003e. \u003cstrong\u003eA\u003c/strong\u003e, expression pattern of candidate genes in developing seeds of QT696 and QT937. Details of the differentially expressed stages and variant locations are indicated with the gene ID. Ⅰ, Ⅱ, Ⅲ, Ⅳ and Ⅴ represent developing peanut seeds at 20, 30, 40, 50 and 60 days after flowering, respectively. The heatmap was generated using Log2 transformed FPKM values. \u003cstrong\u003eB\u003c/strong\u003e, Comparison of seed protein content between cultivars carrying different alleles of five candidate genes. \u003cem\u003eP\u003c/em\u003e-values were determined by a two-tailed Student’s t-test. * and *** represent the statistically significant differences at \u003cem\u003eP \u003c/em\u003e\u0026lt; 0.05 and \u003cem\u003eP \u003c/em\u003e\u0026lt; 0.001 level, respectively.\u003c/p\u003e","description":"","filename":"Fig.4.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/9f9106e560da69ccd8aac42b.jpg"},{"id":79569776,"identity":"deb501d7-16d2-4b69-8c7c-cd4235d17941","added_by":"auto","created_at":"2025-03-31 10:18:49","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":534539,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePhenotypic difference of protein content between different alleles of \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eB03_127452476\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e in the RIL population of ZH6 × XH13 for three environments (A) and in the RIL population of ZH6 × JH5 in 2024AY (B)\u003c/strong\u003e. \u003cem\u003eP\u003c/em\u003e-values were determined by a two-tailed Student’s t-test. ** and *** represent the statistically significant differences at \u003cem\u003eP \u003c/em\u003e\u0026lt; 0.01 and \u003cem\u003eP \u003c/em\u003e\u0026lt; 0.001 level, respectively.\u003c/p\u003e","description":"","filename":"Fig.5.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/b434e8763483e06ae402caf6.jpg"},{"id":89847736,"identity":"ba259a52-5143-4f80-b4fc-25866cf82d2b","added_by":"auto","created_at":"2025-08-25 16:44:11","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":4187342,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/62dee0a1-a2a5-4a97-a738-9c6b93cbbb1c.pdf"},{"id":79570371,"identity":"c6a2aa4f-9055-4b46-bfdb-8f3b9caf6ef3","added_by":"auto","created_at":"2025-03-31 10:26:49","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":1744741,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/bf7334e07a18364d6b446f89.docx"},{"id":79569780,"identity":"9d56e3b8-73f5-407a-98a7-5ce4994c2e15","added_by":"auto","created_at":"2025-03-31 10:18:49","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":1751249,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/b9014e346ef24e6e266f88b5.xlsx"},{"id":79569773,"identity":"1314e06a-083b-4042-ae55-32f5693305ad","added_by":"auto","created_at":"2025-03-31 10:18:49","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":9788,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTable 1 \u003c/strong\u003eDescription of phenotype analysis of protein content in the RIL population\u003c/p\u003e","description":"","filename":"Table1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/140eae17bb7b0c6b1ccbb3e7.xlsx"},{"id":79569779,"identity":"4693e7fd-5176-45b1-9da9-5efb42f3af9b","added_by":"auto","created_at":"2025-03-31 10:18:49","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":9574,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTable 2 \u003c/strong\u003eAnalysis of variance for protein content across three environments\u003c/p\u003e","description":"","filename":"Table2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/787d42cc584e8e4599f78b3a.xlsx"},{"id":79570373,"identity":"caabd538-d1c2-480b-9b6a-d20cbb171a08","added_by":"auto","created_at":"2025-03-31 10:26:49","extension":"xlsx","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":9138,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTable 3 \u003c/strong\u003eThe genomic regions association with protein content through BSA-Seq\u003c/p\u003e","description":"","filename":"Table3.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/698b8239c99607d8295a2853.xlsx"},{"id":79570915,"identity":"c3417504-f6c2-4b7c-8561-6c6d13ad6b57","added_by":"auto","created_at":"2025-03-31 10:34:49","extension":"xlsx","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":10373,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTable 4 \u003c/strong\u003eThe information on identified QTLs in three environments\u003c/p\u003e","description":"","filename":"Table4.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/1816a334cd330ad8fbd86647.xlsx"},{"id":79569782,"identity":"ac730911-1d3d-4d2a-b4a6-de2aa75c4fd5","added_by":"auto","created_at":"2025-03-31 10:18:49","extension":"xlsx","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":9611,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTable 5\u003c/strong\u003e The detailed information on five candidate genes\u003c/p\u003e","description":"","filename":"Table5.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6240685/v1/4fe26083354a8a669aa08410.xlsx"}],"financialInterests":"","formattedTitle":"Identification and utilization of a novel major QTL and linked markers for enhancing protein content in peanut","fulltext":[{"header":"Key message","content":"\u003cp\u003eA novel major QTL \u003cem\u003eqPCB03\u003c/em\u003e for protein content was identified, and a\u0026nbsp;KASP marker linked to\u0026nbsp;\u003cem\u003eqPCB03\u003c/em\u003e was developed and validated to enhance protein content by 1.23% to 1.57% in the RIL populations.\u003c/p\u003e"},{"header":"Introduction","content":"\u003cp\u003ePeanut (\u003cem\u003eArachis hypogaea\u003c/em\u003e L.) is an important food crop cultivated in more than 100 countries worldwide. In 2023, the global peanut cultivation area reached approximately 30.9\u0026nbsp;million hectares, yielding a total production of 54.3\u0026nbsp;million metric tons. China is the leading peanut-producing nation, with a harvested area exceeding 4.8\u0026nbsp;million hectares and an annual production of 19.2\u0026nbsp;million tons (FAOSTAT \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2024\u003c/span\u003e).\u003c/p\u003e \u003cp\u003ePeanuts contain generally 10.5\u0026ndash;22% carbohydrates, 20\u0026ndash;33% crude protein and 45%-55% oil as main compounds with larger genetic variation (Upadhyaya et al. \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e2006\u003c/span\u003e; Barkley et al. \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Davis and Dean \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2016\u003c/span\u003e). In many developing countries, peanuts serve as an important source of protein and nutrients, aiding in the fight against malnutrition among impoverished populations (Linneman et al. \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2007\u003c/span\u003e; Lazzerini et al. \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2013\u003c/span\u003e; Matumba et al. \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2014\u003c/span\u003e). With population growth, global protein supplies will face increasing pressure (Aiking \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2011\u003c/span\u003e) and the demand for high-quality peanut as a source of protein is expected to increase significantly in the near future. Thus, improving protein content has become a primary objective of peanut breeding programs. Seed storage protein is a complex trait that is controlled by multiple genes and affected by the environment and genotype \u0026times; environment interaction (Xu et al. \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e2012\u003c/span\u003e; Patil et al. \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Yang et al. \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Traditional breeding methods have been inefficient and time-consuming for crop improvement. Marker-assisted selection (MAS) has emerged as a more efficient and precise approach for genetic improvement in various crops, including peanuts (Janila et al. \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Varshney \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Pandey et al. \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Identification and mining of loci affecting peanut protein content could provide a technical guide for MAS in enhancing protein content of peanut seed. QTL mapping is a powerful technique for dissecting complex traits and identifying genomic regions linked to quantitative traits. Current QTL research on peanut has primarily focused on oil content, pod/seed size, and disease resistance (Kassie et al. \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), with few studies on protein content. To date, limited studies have reported QTL identification for seed protein content in peanut. Sarvamangala et al. (\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2011\u003c/span\u003e) mapped six QTLs for protein content on a genetic map with 45 SSR loci in peanut. Sun and colleagues (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) developed a high-density genetic map through whole-genome resequencing. They identified several QTLs associated with the oil, protein, and fatty acid content of peanut seeds, primarily located on chromosomes A05 and A08. However, there is currently a deficiency in the development of molecular markers that can facilitate marker-assisted selection tied to high protein content in peanuts.\u003c/p\u003e \u003cp\u003eIn this study, we integrated BSA-seq, linkage analysis and RNA-seq approaches to achieve the following objectives: (1) identification of stable major QTLs for protein content, (2) identification of the favorable alleles of candidate genes enhancing protein content in peanut, and (3) development of linked markers suitable for marker-assisted selection breeding.\u003c/p\u003e"},{"header":"Materials and methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003ePlant materials and field trials\u003c/h2\u003e \u003cp\u003eThe RIL population, consisting of 257 lines, was derived from a cross between ZH6 and XH13 through the single seed descent method. The generations F\u003csub\u003e13\u003c/sub\u003e\u0026ndash;F\u003csub\u003e14\u003c/sub\u003e of the RIL population and parents were planted in three environments, i.e., 2022Yangluo (2022YL), 2023Yangluo (2023YL) and 2023Wuchang (2023WC). In addition, a total of 147 peanut germplasms with diverse protein content were collected from different provinces and phenotyping data in these germplasms was also measured in 2022, Yangluo. The generations F\u003csub\u003e9\u003c/sub\u003e of the RIL population derived from ZH6 and Jihua5 (JH5) were planted in 2024Anyang (2024AY). Each line (approximately 10 plants) was planted in one row, with 2.5 m row length and 0.3 m row spacing. A randomized complete block design was performed with three replicates, and field management followed the standard agricultural practices.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eDetermination of seed protein content and statistical analysis\u003c/h3\u003e\n\u003cp\u003eThe mature peanut seeds were selected for determination of protein content using the Kjeldahl method. An automatic graphite digestion instrument and a Kjeldahl apparatus (K9860; Hanon Instrument Co., Ltd., Jinan, China) were used for sample digestion and nitrogen determination. Ground samples were weighed (~\u0026thinsp;0.3g) and transferred into the digestion instrument containing 1.0 g of catalyst (prepared by mixing 0.6 g of K\u003csub\u003e2\u003c/sub\u003eSO\u003csub\u003e4\u003c/sub\u003e and 0.4 g of CuSO\u003csub\u003e4\u003c/sub\u003e\u0026middot;5H\u003csub\u003e2\u003c/sub\u003eO) and 10 mL of concentrated H\u003csub\u003e2\u003c/sub\u003eSO\u003csub\u003e4\u003c/sub\u003e for ~\u0026thinsp;3.5 hours digestion. The nitrogen content was converted to protein content according to the peanut protein conversion factor (5.46) (Sathe et al. \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2009\u003c/span\u003e). Statistical analysis for the phenotypic data and analysis of variance were performed using IBM SPSS Statistics software version 26 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ibm.com/analytics/spss-statistics-software\u003c/span\u003e\u003cspan address=\"https://www.ibm.com/analytics/spss-statistics-software\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The normal distribution of phenotypic data was confirmed by Shapiro-Wilk test.\u003c/p\u003e\n\u003ch3\u003eBulk segregant analysis\u003c/h3\u003e\n\u003cp\u003eThe lines (19) with high protein content and the lines (19) with low protein content were selected to construct high protein content bulk (bulk-HP) and low protein content bulk (bulk-LP), respectively. The genomic DNA of the lines and the two parents was extracted from the young leaves using Plant Genomic DNA Kit (TianGen, China). Then DNA libraries of two parents and two bulks were constructed and sequenced on an Illumina HiSeq 4000 platform. After removal of low-quality reads, short reads and adapter sequences by Trimmomatic software (Bolger et al. \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2014\u003c/span\u003e), the quality filter reads were aligned to the cultivated peanut reference genome (Bertioli et al. \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) (Tifrunner.gnm2.J5K5; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.peanutbase.com\u003c/span\u003e\u003cspan address=\"http://www.peanutbase.com\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) using Burrows-Wheeler Aligner software (Li and Durbin \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2009\u003c/span\u003e) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://bio-bwa.sourceforge.net/\u003c/span\u003e\u003cspan address=\"http://bio-bwa.sourceforge.net/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Uniquely mapped and properly paired reads were selected with SAMtools (Li et al. \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2009\u003c/span\u003e) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://samtools.sourceforge.net/\u003c/span\u003e\u003cspan address=\"http://samtools.sourceforge.net/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). SNPs and short insertions/deletions (InDels) were identified with Genome Analysis Toolkit (GATK) (McKenna et al. \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2010\u003c/span\u003e) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://software.broadinstitute.org/gatk/\u003c/span\u003e\u003cspan address=\"https://software.broadinstitute.org/gatk/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Annotation of variants was performed using ANNOVAR (Wang et al. \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e2010\u003c/span\u003e) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.openbioinformatics.org/annovar/\u003c/span\u003e\u003cspan address=\"http://www.openbioinformatics.org/annovar/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). By subtracting SNP-index of bulk-LP from SNP‐index of bulk-HP, the ΔSNP‐index was calculated with a 2 Mb genome window using the R package QTLseqr (Mansfeld and Grumet \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). The details on whole genome re-sequencing data for BSA-seq were shown in Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e. The raw sequencing data has been submitted to the National Genomics Data Center (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ngdc.cncb.ac.cn/\u003c/span\u003e\u003cspan address=\"https://ngdc.cncb.ac.cn/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) under BioProjectID PRJCA036928.\u003c/p\u003e\n\u003ch3\u003eGenotyping and genetic linkage map construction\u003c/h3\u003e\n\u003cp\u003eThe genomic DNA of the RILs and the parents was extracted from the young leaves using Plant Genomic DNA Kit (TianGen, China). Each line was genotyped using the 50K Genotyping by the Target Sequencing (GBTS) SNP array (MolBreeding Biotechnology Co., Ltd., Shijiazhuang, China). Raw reads were filtered using fastp software (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/OpenGene/fastp\u003c/span\u003e\u003cspan address=\"https://github.com/OpenGene/fastp\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Clean reads were aligned to the cultivated peanut reference genome (Bertioli et al. \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2019\u003c/span\u003e) (Tifrunner.gnm2.J5K5; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.peanutbase.com\u003c/span\u003e\u003cspan address=\"http://www.peanutbase.com\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) using Burrows-Wheeler Aligner (BWA) software (Li and Durbin \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2009\u003c/span\u003e). SNPs and InDels were identified and filtered using UnifiedGenotyper and VariantFiltration commands in GATK (McKenna et al. \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2010\u003c/span\u003e). After excluding the monomorphic variants, polymorphic SNP markers were binned using the BIN function in QTL IciMapping (Meng et al. \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). One marker in each bin was randomly selected to construct the genetic map. Bin markers were organized into corresponding groups with LOD scores ranging from 2 to 10 using JoinMap 4 software (Ooijen et al. 2006) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.kyazma.nl/index.php/JoinMap/\u003c/span\u003e\u003cspan address=\"https://www.kyazma.nl/index.php/JoinMap/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The Kosambi mapping function was employed to evaluate the recombination rate and to calculate the genetic distance between bin markers. Flanking sequences of SNPs were used as queries in a BLAST search against the reference genome to determine their physical locations. Linkage maps were drawn using R package LinkageMapView (Ouellette et al. \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2018\u003c/span\u003e).\u003c/p\u003e\n\u003ch3\u003eQTL analysis\u003c/h3\u003e\n\u003cp\u003eGenome-wide QTL mapping was conducted using the mean value of protein content across environments. QTLs were detected using Windows QTL Cartographer 2.5 software (Wang et al. \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2012\u003c/span\u003e) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://statgen.ncsu.edu/qtlcart/WQTLCart.htm\u003c/span\u003e\u003cspan address=\"http://statgen.ncsu.edu/qtlcart/WQTLCart.htm\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), applying the composite interval mapping (CIM) method with LOD score values\u0026thinsp;\u0026ge;\u0026thinsp;2.5. The parameters for control markers, window size, and walking speed were set to 5, 10, and 2 cM, respectively. QTLs were named starting with the letter \u0026ldquo;q\u0026rdquo;, followed by the abbreviation of trait name \u0026ldquo;PC\u0026rdquo; and the corresponding linkage group. An arabic numeral was added after the linkage group if two or more QTLs were identified in the same linkage group. QTLs detected in more than one environment and phenotypic variant explained\u0026thinsp;\u0026gt;\u0026thinsp;10% were considered to be major QTLs.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eRNA-sequencing analysis\u003c/h2\u003e \u003cp\u003eTwo lines with large differences in protein content from RIL populations, namely, QT696 (a high protein content line from bulk-HP) and QT937 (a low protein content line from bulk-LP), were used for transcriptome sequencing and candidate genes analysis. According to the previous studies (Wan et al. \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Wang et al. \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2022\u003c/span\u003e), developing peanut seeds were collected from approximately 10 plants at five developmental stages (Ⅰ, Ⅱ, Ⅲ, Ⅳ and Ⅴ), corresponding to 20, 30, 40, 50 and 60 days after flowering, respectively. Each stage was represented by three biological replicates. Total RNA was extracted using FastPure Plant Total RNA Isolation Kit RC401-01 (Vazyme, Nanjing, China). The constructed libraries were sequenced on an Illumina Novaseq platform (Novogene Technology Co., Ltd., Beijing, China) and 150 bp paired-end reads were generated. Clean reads were aligned to the cultivated peanut reference genome using Hisat2 v2.0.5 (Kim et al. \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Bertioli et al. \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). The Gene expression levels were quantified and normalized as Fragments Per Kilobase of transcript per Million mapped reads (FPKM) using StringTie version 1.2.3 (Pertea et al. \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). Differentially expressed genes (DEGs) were identified through DESeq2 version 1.18.1 (Anders and Huber \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2010\u003c/span\u003e) based on criteria that included a minimum FPKM value of 1 in at least one sample, a fold change of at least 2, and an adjusted \u003cem\u003eP\u003c/em\u003e-value of 0.05 or lower. TBtools was used for generating gene expression heatmap (Chen et al. \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2020a\u003c/span\u003e). The raw sequencing data of RNA-seq have been uploaded to the National Genomics Data Center (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ngdc.cncb.ac.cn/\u003c/span\u003e\u003cspan address=\"https://ngdc.cncb.ac.cn/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) under BioProjectID PRJCA036946.\u003c/p\u003e \u003c/div\u003e"},{"header":"Result","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003ePhenotypic variation of the RIL population\u003c/h2\u003e \u003cp\u003ePhenotypic measurements were conducted for 257 recombinant inbred lines (RILs) and two parents across three environments. As shown in Table\u0026nbsp;1, the male parent ZH6 exhibited consistently higher protein content (27.20% \u0026minus;\u0026thinsp;30.31%) compared to the female parent XH13 (23.06% \u0026minus;\u0026thinsp;26.61%) across all three environments. The protein content of the RIL population showed large phenotypic variations ranging from 22.77\u0026ndash;30.28% in 2022YL, from 21.05\u0026ndash;29.72% in 2023YL, and from 21.56\u0026ndash;30.27% in 2023WC. The phenotypic values showed a continuous distribution with transgressive segregation in the RIL population (Fig.\u0026nbsp;1). The Shapiro-Wilk test was indicated that the phenotypic data of the RIL population were normally distributed in 2022YL and 2023WC, but not in 2023YL (Table\u0026nbsp;1). Variance analysis was revealed that genetic and environmental factors, as well as genotype \u0026times; environment interaction significantly influenced protein content in the RIL population (Table\u0026nbsp;2). Broad-sense heritability for protein content was estimated to be 0.743 (Table\u0026nbsp;2).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eIdentification of candidate genomic regions associated with protein content using BSAseq\u003c/h2\u003e \u003cp\u003eThe average protein content in the RIL population ranged from 22.02\u0026ndash;29.44% across three environments. RILs with the highest protein content (27.52% \u0026minus;\u0026thinsp;29.31%) were selected to construct the extremely high protein content bulk (bulk-HP), while the RILs with the lowest protein content (22.06% \u0026minus;\u0026thinsp;25.62%) were used to form extremely low protein content bulk (bulk-LP). The DNA libraries of two parents and two bulks were subjected to whole-genome resequencing using an Illumina HiSeq platform, producing 241.18\u0026nbsp;million reads (71.45 Gb) for XH13, 212.77\u0026nbsp;million reads (56.89 Gb) for ZH6, 353.26\u0026nbsp;million reads (104.36 Gb) for bulk-HP and 343.46\u0026nbsp;million reads (101.76 Gb) for bulk-LP, respectively. The reads of the parents and two bulks were mapped to the reference genome of cultivated peanut, achieving coverage rates of 93.87% with a depth of 24.03X for XH13, 92.89% coverage with a depth of 19.05X for ZH6, 94.32% coverage with a depth of 34.01X for bulk-HP, and 94.96% coverage with a depth of 33.74X for bulk-LP (Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). A total of 275,570 genomic variants were identified between bulk-HP and bulk-LP, which included 241,607 SNPs and 33,963 InDels (Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTo identify the candidate genomic regions controlling protein content, the SNP-index was calculated for each bulk using ZH6 as a reference genome. The ΔSNP-index were calculated by subtracting SNP‐index of bulk-LP from SNP‐index of bulk-HP (Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). Four genomic regions have been identified in association with protein content. These included \u003cem\u003eqPCA07\u003c/em\u003e on chromosome A07 (66.62 Mb \u0026minus;\u0026thinsp;73.67 Mb), \u003cem\u003eqPCA10\u003c/em\u003e on chromosome A10 (11.35 Mb \u0026minus;\u0026thinsp;31.17 Mb), \u003cem\u003eqPCB03\u003c/em\u003e on chromosome B03 (124.98 Mb to 128.76 Mb), and \u003cem\u003eqPCB10\u003c/em\u003e on chromosome B10 (136.97 Mb \u0026minus;\u0026thinsp;138.34 Mb) (Fig.\u0026nbsp;2; Table\u0026nbsp;3). The ΔSNP‐index of the four genomic regions was positive, indicating that favorable alleles were from the reference parent ZH6.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eIdentification of QTLs for protein content through linkage analysis\u003c/h2\u003e \u003cp\u003eThe RIL population and two parents were genotyped using the cultivated peanut 50K GBTS SNP array. After excluding the monomorphic loci and loci with partial segregation, a high-density genetic map was constructed. It is consisted of 1881 bin markers distributed on 22 linkage groups, with an average genetic distance of 1.03 cM (Fig. \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e; Table S3; Table S4). A genome-wide QTL analysis was performed using the genetic map and phenotypic data on protein content across three environments. A total of nine QTLs, with logarithm of odds (LODs) ranging from 2.88 to 7.55, were mapped on six linkage groups (Table\u0026nbsp;4). The QTLs explained 3.84\u0026ndash;12.86% of the phenotypic variance for protein content, with the additive effect values from \u0026minus;\u0026thinsp;0.33 to 0.82 (Table\u0026nbsp;4). Among them, \u003cem\u003eqPCB03\u003c/em\u003e was consistently detected in three environments, exhibiting the highest phenotypic variation explained (PVE) of 11.03\u0026ndash;12.86%. Additionally, the physical interval of \u003cem\u003eqPCB03\u003c/em\u003e (125.92Mb-127.94Mb) was found to colocate with the candidate genomic region detected by BSA-Seq (Fig.\u0026nbsp;3). The results indicated that the genomic region on B03 (125.92Mb-127.94Mb) harbored a major and stable locus that regulates protein content.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003ePutative candidate genes for QTL\u003c/b\u003e \u003cb\u003eqPCB03\u003c/b\u003e\u003c/p\u003e \u003cp\u003eA combined analysis of genomic variation data and gene expression was conducted to identify candidate genes in ~\u0026thinsp;2Mb genomic region of \u003cem\u003eqPCB03\u003c/em\u003e. In the physical interval of \u003cem\u003eqPCB03\u003c/em\u003e, 349 genomic variants (301 SNPs and 48 InDels) were detected between ZH6 and XH13 (Table S5). Among the variants, 318 were in intergenic regions, 10 were in introns, 6 non-synonymous were in exons, 15 were in 2kb-upstream/downstream regions. The 6 non-synonymous SNPs were found to affect 6 candidate genes for protein content. For gene expression analysis, developing seeds of the bulk-HP line (QT696) and the bulk-LP line (QT937) at five stages (20, 30, 40, 50 and 60 days after flowering) were collected to perform transcriptome sequencing. The protein content of developing seeds in QT696 was significantly higher than that in QT937 at five stages (Fig. S3).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eA total of 30 RNA-seq libraries were constructed and 5.68-7.00 Gb clean data per library was obtained (Table S6). On average, 96.97% of clean reads were mapped on the reference genome, and 40,746 to 51,848 genes were identified in each library. Within the physical interval of \u003cem\u003eqPCB03\u003c/em\u003e, 43 genes were found to be expressed in developing seeds (Table S7). Among them, four genes (\u003cem\u003eAh13g469400\u003c/em\u003e, \u003cem\u003eAh13g473200\u003c/em\u003e, \u003cem\u003eAh13g475000\u003c/em\u003e, and \u003cem\u003eAh13g477400\u003c/em\u003e) exhibited nonsynonymous mutants between two parents (Fig.\u0026nbsp;4A). Furthermore, 22 DEGs were identified in the region of \u003cem\u003eqPCB03\u003c/em\u003e, including one DEG (\u003cem\u003eAh13g477200\u003c/em\u003e) that harbored a variant in the 2kb-upstream (promoter) region (Fig.\u0026nbsp;4A). Based on the analysis of genomic variant and gene expression, five genes (\u003cem\u003eAh13g469400\u003c/em\u003e, \u003cem\u003eAh13g473200\u003c/em\u003e, \u003cem\u003eAh13g475000\u003c/em\u003e, \u003cem\u003eAh13g477200\u003c/em\u003e, and \u003cem\u003eAh13g477400\u003c/em\u003e) were identified as candidate genes for the stable QTL \u003cem\u003eqPCB03\u003c/em\u003e (Table\u0026nbsp;5).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo confirm the allelic effects of candidate genes on protein content in peanut germplasm, KASP markers were designed based on five SNP sites of the candidate genes (Table S8). These markers were subsequently utilized to genotype 147 Chinese peanut cultivars which exhibited significant phenotypic variation, with protein content ranging from 24.23\u0026ndash;33.58% in the 2022YL. The KASP assays effectively distinguished the allelic variations in the peanut germplasm (Fig. S4), and the average protein content between allelotypes showed statistically significant differences at all five SNP sites (Fig.\u0026nbsp;4B). These results indicated that the five genes are robust candidates influencing protein content in peanuts.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eValidation of KASP markers linked to\u003c/b\u003e \u003cb\u003eqPCB03\u003c/b\u003e\u003c/p\u003e \u003cp\u003eThe KASP marker \u003cem\u003eKASP_B03_127452476\u003c/em\u003e, exhibiting the lowest \u003cem\u003eP\u003c/em\u003e-value in allelic comparison (Fig.\u0026nbsp;4B), was chosen for assessing its potential value in breeding programs. Two RIL populations derived from the same high-protein male parent ZH6, ZX_RIL (ZH6 \u0026times; XH13) and ZJ_RIL (ZH6 \u0026times; JH5), were genotyped using the \u003cem\u003eKASP_B03\u003c/em\u003e_\u003cem\u003e127452476\u003c/em\u003e marker. The high-protein genotype (\u003cem\u003eB03_127452476_G\u003c/em\u003e) originated from ZH6, while the low-protein genotype (\u003cem\u003eB03_127452476_A\u003c/em\u003e) was derived from XH13 or JH5. In ZX_RIL population, RILs with the \u003cem\u003eB03_127452476_G/G\u003c/em\u003e genotype exhibited significantly higher protein content (27.11\u0026thinsp;\u0026plusmn;\u0026thinsp;1.39% in 2022YL, 26.83\u0026thinsp;\u0026plusmn;\u0026thinsp;1.23% in 2023YL, and 26.95\u0026thinsp;\u0026plusmn;\u0026thinsp;1.57% in 2023WC) compared to RILs with the \u003cem\u003eB03_127452476_A/A\u003c/em\u003e genotype (26.11\u0026thinsp;\u0026plusmn;\u0026thinsp;1.29% in 2022YL, 25.77\u0026thinsp;\u0026plusmn;\u0026thinsp;1.44% in 2023YL, and 25.71\u0026thinsp;\u0026plusmn;\u0026thinsp;1.53% in 2023WC). Selection of the favorable allele in ZX_RIL population could lead to an increase in protein content by 1.23\u0026ndash;1.57% across the three environments (Fig.\u0026nbsp;5A). In ZJ_RIL population, the protein content ranged from 19.40\u0026ndash;33.78% in 2024AY. RILs with the \u003cem\u003eB03_127452476_G/G\u003c/em\u003e genotype demonstrated significantly higher protein content (26.48\u0026thinsp;\u0026plusmn;\u0026thinsp;3.03% in 2024AY) compared to those with the \u003cem\u003eB03_127452476_A/A\u003c/em\u003e genotype (24.95\u0026thinsp;\u0026plusmn;\u0026thinsp;2.89% in 2024AY) (Fig.\u0026nbsp;5B). Selecting the favorable allele could enhance protein content by 1.53% in ZJ_RIL population. These findings suggested the potential value of the KASP marker in marker-assisted selection of peanut genotypes associated with high protein content.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eA novel and major QTL for seed protein content in peanut\u003c/h2\u003e \u003cp\u003eThe cultivation of high-protein peanut is a primary objective in breeding program. However, limited studies have reported on the identification of QTLs for seed protein content in peanut. In this study, nine QTLs for protein content were identified through linkage analysis. Two of them, \u003cem\u003eqPCA05.1\u003c/em\u003e and \u003cem\u003eqPCA08\u003c/em\u003e, overlapped with or were close to the QTLs for protein content reported by Sun et al. (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2022\u003c/span\u003e), indicating they are reliable loci for regulating protein content. In addition, \u003cem\u003eqPCA07\u003c/em\u003e, \u003cem\u003eqPCB03\u003c/em\u003e and \u003cem\u003eqPCB10.2\u003c/em\u003e were repeatedly detected in at least two environments. Notably, \u003cem\u003eqPCB03\u003c/em\u003e were colocated on the genomic regions detected by BSA-Seq.\u0026nbsp;Since \u003cem\u003eqPCB03\u003c/em\u003e exhibited the highest PVE of 11.03\u0026ndash;12.86% across three environments and has not been previously reported, it is considered as a novel major QTL for protein content. Five candidate genes were identified in the genomic region of \u003cem\u003eqPCB03\u003c/em\u003e (~\u0026thinsp;2 Mb), and their allelic effects were shown to be significantly associated with protein content in peanut germplasm. Based on the SNP marker corresponding to the candidate gene, the selection of favorable genotypes could potentially boost protein content by 1.23\u0026ndash;1.57% in RIL populations. The results indicate that \u003cem\u003eqPCB03\u003c/em\u003e is a reliable major QTL for seed protein content, as well as a valuable locus in MAS for breeding high seed protein content in peanut.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eGenetic correlation between protein content and oil content\u003c/h2\u003e \u003cp\u003eIt is widely recognized that there is an inverse relationship between protein and oil content in seeds. This correlation has been evident in the identification of major QTLs for protein and oil content, which share overlapping regions (Sarvamangala et al. \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2011\u003c/span\u003e; Sun et al. \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Duan et al. \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). In this study, we found a significant negative correlation between protein and oil contents (correlation coefficient = -0.766, Table S9). Additionally, the linkage analysis revealed that QTLs for both protein and oil content were overlapped on the genetic map (Table S10). Interestingly, the KASP marker (\u003cem\u003eKASP_B03_127452476\u003c/em\u003e) for protein content could be employed in the selection of low-oil genotype. The theoretical oil content of low-oil genotype could decrease by 0.85\u0026ndash;1.31% in ZX_RIL population and by 1.60% in ZJ_RIL population (Fig. S5). These findings suggest that \u003cem\u003eqPCB03\u003c/em\u003e with linked KASP marker could facilitate marker-assisted selection of peanut varieties with high protein content and low oil content simultaneously.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003ePrediction of candidate genes associated with protein content\u003c/h2\u003e \u003cp\u003eThe high-resolution genetic map enabled the demarcation of \u003cem\u003eqPCB03\u003c/em\u003e to a 2 Mb interval containing 78 genes. In this study, candidate genes were screened out through a combined analysis of genomic variation, gene expression, and allelic effects of candidates in peanut germplasm. Finally, five of 78 genes (\u003cem\u003eAh13g469400\u003c/em\u003e, \u003cem\u003eAh13g473200\u003c/em\u003e, \u003cem\u003eAh13g475000\u003c/em\u003e, \u003cem\u003eAh13g477200\u003c/em\u003e and \u003cem\u003eAh13g477400\u003c/em\u003e) were deduced to be candidate genes for \u003cem\u003eqPCB03\u003c/em\u003e. Four candidate genes (\u003cem\u003eAh13g469400\u003c/em\u003e, \u003cem\u003eAh13g473200\u003c/em\u003e, \u003cem\u003eAh13g475000\u003c/em\u003e and \u003cem\u003eAh13g477400\u003c/em\u003e) had nonsynonymous SNVs that encode a violaxanthin de-epoxidase-related (VDR), nucleolin-like isoform 2 protein, aspartate/prephenate aminotransferase and CAAX protease self-immunity protein, respectively. One candidate gene (\u003cem\u003eAh13g477200\u003c/em\u003e) harbored a SNP in its upstream region and was differentially expressed in developing seeds. It encodes a trehalose-6-phosphate phosphatase. Notably, violaxanthin de-epoxidase (VDE) is involved in photoprotective response to high-light stress and abscisic acid (ABA) biosynthesis (Demmig et al. \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e1987\u003c/span\u003e; Ye et al. \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e2012\u003c/span\u003e; Chen et al. \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2020b\u003c/span\u003e). Photosynthesis and nitrogen metabolism are closely linked by their interdependence for fixed carbon, chemical energy, and nitrogen assimilates (Paul and Pellny \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2003\u003c/span\u003e; Zheng \u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e2009\u003c/span\u003e). ABA has been reported to be involved in nitrogen remobilization and seed storage protein synthesis (Zheng et al. \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Yang et al. \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Nan et al. \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Yang et al. \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Aspartate aminotransferase has been reported to play an important role in carbon and nitrogen metabolism (de la Torre et al. \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2006\u003c/span\u003e; de la Torre et al. \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2014\u003c/span\u003e) and enhanced aspartate aminotransferase enzyme activity would increase total free amino acid contents in rice seeds (Zhou et al. \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e2009\u003c/span\u003e). Although evidence suggests that these putative genes may be involved in regulating protein content, further studies are necessary to validate their functions.\u003c/p\u003e \u003cp\u003eIn summary, a novel major QTL \u003cem\u003eqPCB03\u003c/em\u003e for protein content was identified using both BSA-Seq and linkage mapping. A total of five candidate genes were discovered in genomic region of \u003cem\u003eqPCB03\u003c/em\u003e and their allelic effect were found to be significantly associated with protein content. A KASP marker linked to \u003cem\u003eqPCB03\u003c/em\u003e was developed and validated to enhance protein content by 1.23\u0026ndash;1.57% in the RIL populations. The results provide a valuable locus with the linked marker in MAS for breeding high seed protein content in peanut.\u003c/p\u003e "},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgments\u0026nbsp;\u003c/strong\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions\u0026nbsp;\u003c/strong\u003eMW, JG, BL, HJ, NL and YL conceived and designed the research. HJ and YC developed the RIL population. JG, GJ, TY, WC, LH, HL and XZ planted the materials and conducted field management. MW and JG performed the measurement of protein content and statistical analysis of the phenotyping data. MW, JG and NL performed BSA-seq, QTL analysis, and development and validation of the marker associated with protein content. MW and NL performed SNP analysis. MW and NL wrote the manuscript, HJ and YL revised the manuscript and improved the English writing. All the authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e This work was supported by the National Key Research and Development Program of China (2022YFD1200400), the National Peanut Industry Technology System Construction, China (CARS13), the National Crop Germplasm Resources Center (NCGRC-2024-036), the National Program for Crop Germplasm Protection of China (19210163), the Agricultural Science and Technology Innovation Program of Chinese Academy of Agricultural Sciences (CAAS-ASTIP-2024-OCRI).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u0026nbsp;\u003c/strong\u003eAll data supporting the results of this study are available within the paper and its supplementary data published online.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of interest\u0026nbsp;\u003c/strong\u003eThe authors declare no conflicts of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical standards\u003c/strong\u003e The authors state that all experiments in the study comply with the ethical standards.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAiking H (2011) Future protein supply. Trends Food Sci Technol 22:112-120\u003c/li\u003e\n\u003cli\u003eAnders S, Huber W (2010) Differential expression analysis for sequence count data. Genome Biol 11:R106\u003c/li\u003e\n\u003cli\u003eBarkley NA, Upadhyaya HD, Liao B, Holbrook CC (2016) Chapter 3 - Global resources of genetic diversity in peanut. In: Stalker HT, F. Wilson R (eds) Peanuts. AOCS Press, pp 67-109\u003c/li\u003e\n\u003cli\u003eBertioli DJ, Jenkins J, Clevenger J, Dudchenko O, Gao D, Seijo G, Leal-Bertioli SCM, Ren L, Farmer AD, Pandey MK, Samoluk SS, Abernathy B, Agarwal G, Ball\u0026eacute;n-Taborda C, Cameron C, Campbell J, Chavarro C, Chitikineni A, Chu Y, Dash S, El Baidouri M, Guo B, Huang W, Kim KD, Korani W, Lanciano S, Lui CG, Mirouze M, Moretzsohn MC, Pham M, Shin JH, Shirasawa K, Sinharoy S, Sreedasyam A, Weeks NT, Zhang X, Zheng Z, Sun Z, Froenicke L, Aiden EL, Michelmore R, Varshney RK, Holbrook CC, Cannon EKS, Scheffler BE, Grimwood J, Ozias-Akins P, Cannon SB, Jackson SA, Schmutz J (2019) The genome sequence of segmental allotetraploid peanut \u003cem\u003eArachis hypogaea\u003c/em\u003e. Nat Genet 51:877-884\u003c/li\u003e\n\u003cli\u003eBolger AM, Lohse M, Usadel B (2014) Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 30:2114-2120\u003c/li\u003e\n\u003cli\u003eChen C, Chen H, Zhang Y, Thomas HR, Frank MH, He Y, Xia R (2020a) TBtools: an integrative toolkit developed for interactive analyses of big biological data. Mol Plant 13:1194-1202\u003c/li\u003e\n\u003cli\u003eChen K, Li GJ, Bressan RA, Song CP, Zhu JK, Zhao Y (2020b) Abscisic acid dynamics, signaling, and functions in plants. J Integr Plant Biol 62:25-54\u003c/li\u003e\n\u003cli\u003eDavis JP, Dean LL (2016) Chapter 11 - Peanut composition, flavor and nutrition. In: Stalker HT, F. Wilson R (eds) Peanuts. AOCS Press, pp 289-345\u003c/li\u003e\n\u003cli\u003ede la Torre F, Ca\u0026ntilde;as RA, Pascual MB, Avila C, C\u0026aacute;novas FM (2014) Plastidic aspartate aminotransferases and the biosynthesis of essential amino acids in plants. J Exp Bot 65:5527-5534\u003c/li\u003e\n\u003cli\u003ede la Torre F, De Santis L, Su\u0026aacute;rez MF, Crespillo R, C\u0026aacute;novas FM (2006) Identification and functional analysis of a prokaryotic-type aspartate aminotransferase: implications for plant amino acid metabolism. Plant J 46:414-425\u003c/li\u003e\n\u003cli\u003eDemmig B, Winter K, Kr\u0026uuml;ger A, Czygan FC (1987) Photoinhibition and zeaxanthin formation in intact leaves : a possible role of the xanthophyll cycle in the dissipation of excess light energy. Plant Physiol 84:218-224\u003c/li\u003e\n\u003cli\u003eDuan Z, Li Q, Wang H, He X, Zhang M (2023) Genetic regulatory networks of soybean seed size, oil and protein contents. Front Plant Sci 14:1160418\u003c/li\u003e\n\u003cli\u003eFAOSTAT (2024) Statistical database FAOSTAT. http://faostat3.fao.org\u003c/li\u003e\n\u003cli\u003eJanila P, Variath MT, Pandey MK, Desmae H, Motagi BN, Okori P, Manohar SS, Rathnakumar AL, Radhakrishnan T, Liao B, Varshney RK (2016) Genomic tools in groundnut breeding program: status and perspectives. Front Plant Sci 7:289\u003c/li\u003e\n\u003cli\u003eKassie FC, Nguepjop JR, Ngalle HB, Assaha DVM, Gessese MK, Abtew WG, Tossim HA, Sambou A, Seye M, Rami JF, Fonceka D, Bell JM (2023) An overview of mapping quantitative trait loci in peanut (\u003cem\u003eArachis hypogaea L\u003c/em\u003e.). Genes 14:1176\u003c/li\u003e\n\u003cli\u003eKim D, Langmead B, Salzberg SL (2015) HISAT: a fast spliced aligner with low memory requirements. Nat Methods 12:357-360\u003c/li\u003e\n\u003cli\u003eLazzerini M, Rubert L, Pani P (2013) Specially formulated foods for treating children with moderate acute malnutrition in low- and middle-income countries. Cochrane Database Syst Rev Cd009584\u003c/li\u003e\n\u003cli\u003eLi H, Durbin R (2009) Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25:1754-1760\u003c/li\u003e\n\u003cli\u003eLi H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R (2009) The sequence alignment/map format and SAMtools. Bioinformatics 25:2078-2079\u003c/li\u003e\n\u003cli\u003eLinneman Z, Matilsky D, Ndekha M, Manary MJ, Maleta K, Manary MJ (2007) A large-scale operational study of home-based therapy with ready-to-use therapeutic food in childhood malnutrition in Malawi. Matern Child Nutr 3:206-215\u003c/li\u003e\n\u003cli\u003eMansfeld BN, Grumet R (2018) QTLseqr: An R package for bulk segregant analysis with next-generation sequencing. Plant Genome 11\u003c/li\u003e\n\u003cli\u003eMatumba L, Monjerezi M, Biswick T, Mwatseteza J, Makumba W, Kamangira D, Mtukuso A (2014) A survey of the incidence and level of aflatoxin contamination in a range of locally and imported processed foods on Malawian retail market. Food Control 39:87-91\u003c/li\u003e\n\u003cli\u003eMcKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A, Garimella K, Altshuler D, Gabriel S, Daly M, DePristo MA (2010) The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res 20:1297-1303\u003c/li\u003e\n\u003cli\u003eMeng L, Li H, Zhang L, Wang J (2015) QTL IciMapping: integrated software for genetic linkage map construction and quantitative trait locus mapping in biparental populations. Crop J 3:269-283\u003c/li\u003e\n\u003cli\u003eNan Y, He H, Xie Y, Li C, Atif A, Hui J, Tian H, Gao Y (2023) The responses of genotypes with contrasting NUtE to exogenous ABA during the flowering stage in Brassica napus. Plant Stress 10:100248\u003c/li\u003e\n\u003cli\u003eOoijen JWv, Ooijen JWv, Verlaat Jvt, Ooijen JWv, Tol J, Dalén J, Buren JBV, Meer JWMvd, Krieken JHv, Ooijen JWv, Kessel JSV, Van O, Voorrips RE, Heuvel LP (2006) JoinMap 4, Software for the calculation of genetic linkage maps in experimental populations. Kyazma B.V., Wageningen, Netherlands. \u003c/li\u003e\n\u003cli\u003eOuellette LA, Reid RW, Blanchard SG, Brouwer CR (2018) LinkageMapView-rendering high-resolution linkage and QTL maps. Bioinformatics 34:306-307\u003c/li\u003e\n\u003cli\u003ePandey MK, Pandey AK, Kumar R, Nwosu CV, Guo B, Wright GC, Bhat RS, Chen X, Bera SK, Yuan M, Jiang H, Faye I, Radhakrishnan T, Wang X, Liang X, Liao B, Zhang X, Varshney RK, Zhuang W (2020) Translational genomics for achieving higher genetic gains in groundnut. Theor Appl Genet 133:1679-1702\u003c/li\u003e\n\u003cli\u003ePatil G, Mian R, Vuong T, Pantalone V, Song Q, Chen P, Shannon GJ, Carter TC, Nguyen HT (2017) Molecular mapping and genomics of soybean seed protein: a review and perspective for the future. Theor Appl Genet 130:1975-1991\u003c/li\u003e\n\u003cli\u003ePaul MJ, Pellny TK (2003) Carbon metabolite feedback regulation of leaf photosynthesis and development. J Exp Bot 54:539-547\u003c/li\u003e\n\u003cli\u003ePertea M, Pertea GM, Antonescu CM, Chang TC, Mendell JT, Salzberg SL (2015) StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat biotechnol 33:290-295\u003c/li\u003e\n\u003cli\u003eSarvamangala C, Gowda MVC, Varshney RK (2011) Identification of quantitative trait loci for protein content, oil content and oil quality for groundnut (\u003cem\u003eArachis hypogaea L\u003c/em\u003e.). Field Crops Res 122:49-59\u003c/li\u003e\n\u003cli\u003eSathe SK, Venkatachalam M, Sharma GM, Kshirsagar HH, Teuber SS, Roux KH (2009) Solubilization and electrophoretic characterization of select edible nut seed proteins. J Agric Food Chem 57:7846-7856\u003c/li\u003e\n\u003cli\u003eSun Z, Qi F, Liu H, Qin L, Xu J, Shi L, Zhang Z, Miao L, Huang B, Dong W, Wang X, Tian M, Feng J, Zhao R, Zheng Z, Zhang X (2022) QTL mapping of quality traits in peanut using whole-genome resequencing. Crop J 10:177-184\u003c/li\u003e\n\u003cli\u003eUpadhyaya HD, Reddy LJ, Gowda CLL, Singh S (2006) Identification of diverse groundnut germplasm: sources of early maturity in a core collection. Field Crops Res 97:261-271\u003c/li\u003e\n\u003cli\u003eVarshney RK (2016) Exciting journey of 10 years from genomes to fields and markets: some success stories of genomics-assisted breeding in chickpea, pigeonpea and groundnut. Plant Sci 242:98-107\u003c/li\u003e\n\u003cli\u003eWan L, Li B, Pandey MK, Wu Y, Lei Y, Yan L, Dai X, Jiang H, Zhang J, Wei G, Varshney RK, Liao B (2016) Transcriptome analysis of a new peanut seed coat mutant for the physiological regulatory mechanism involved in seed coat cracking and pigmentation. Front Plant Sci 7:1491\u003c/li\u003e\n\u003cli\u003eWang K, Li M, Hakonarson H (2010) ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res 38:e164\u003c/li\u003e\n\u003cli\u003eWang S, Basten C, Zeng Z (2012) Windows QTL Cartographer v2.5. Department of Statistics, North Carolina State University; Raleigh, NC.\u003c/li\u003e\n\u003cli\u003eWang Z, Yan L, Chen Y, Wang X, Huai D, Kang Y, Jiang H, Liu K, Lei Y, Liao B (2022) Detection of a major QTL and development of KASP markers for seed weight by combining QTL-seq, QTL-mapping and RNA-seq in peanut. Theor Appl Genet 135:1779-1795\u003c/li\u003e\n\u003cli\u003eXu G, Fan X, Miller AJ (2012) Plant nitrogen assimilation and use efficiency. Annu Rev Plant Biol 63:153-182\u003c/li\u003e\n\u003cli\u003eYang T, Wang H, Guo L, Wu X, Xiao Q, Wang J, Wang Q, Ma G, Wang W, Wu Y (2022) ABA-induced phosphorylation of basic leucine zipper 29, ABSCISIC ACID INSENSITIVE 19, and Opaque2 by SnRK2.2 enhances gene transactivation for endosperm filling in maize. Plant Cell 34:1933-1956\u003c/li\u003e\n\u003cli\u003eYang T, Wu X, Wang W, Wu Y (2023) Regulation of seed storage protein synthesis in monocot and dicot plants: a comparative review. Mol Plant 16:145-167\u003c/li\u003e\n\u003cli\u003eYe N, Jia L, Zhang J (2012) ABA signal in rice under stress conditions. Rice 5:1\u003c/li\u003e\n\u003cli\u003eZheng X, Li Q, Li C, An D, Xiao Q, Wang W, Wu Y (2019) Intra-Kernel reallocation of proteins in maize depends on VP1-mediated scutellum development and nutrient assimilation. Plant Cell 31:2613-2635\u003c/li\u003e\n\u003cli\u003eZheng ZL (2009) Carbon and nitrogen nutrient balance signaling in plants. Plant Signal Behav 4:584-591\u003c/li\u003e\n\u003cli\u003eZhou Y, Cai H, Xiao J, Li X, Zhang Q, Lian X (2009) Over-expression of aspartate aminotransferase genes in rice resulted in altered nitrogen metabolism and increased amino acid content in seeds. Theor Appl Genet 118:1381-1390\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTables 1-5 are available in the Supplementary Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"theoretical-and-applied-genetics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"taag","sideBox":"Learn more about [Theoretical and Applied Genetics](https://www.springer.com/journal/122)","snPcode":"122","submissionUrl":"https://submission.nature.com/new-submission/122/3","title":"Theoretical and Applied Genetics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"peanut, protein content, QTL mapping, candidate gene, marker-assisted breeding","lastPublishedDoi":"10.21203/rs.3.rs-6240685/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6240685/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003ePeanut is a vital source of protein for humans, playing a key role in maintaining a steady protein supply. In this study, the high protein cultivar Zhonghua6 (29.14±1.69%) was crossed with Xuhua13 (24.55±1.84%) to construct a recombinant inbred line (RIL) population. The protein content of the RIL population exhibited significant variation, ranging from 21.05% to 30.28%. To discover genomic regions associated with protein content, four libraries were constructed (two parents and two extreme bulks) for bulked segregant sequencing (BSA-seq). The results revealed significant associations between protein content and the genomic regions on chromosomes A07, A10, B03, and B10. Through linkage analysis, nine QTLs (quantitative trait loci) for protein content were identified, among which the major and stable QTL \u003cem\u003eqPCB03\u003c/em\u003eon chromosomes B03 (125.92-127.94 Mb) explained 11.03%-12.86% of the phenotypic variation. This QTL (\u003cem\u003eqPCB03\u003c/em\u003e), simultaneously identified by both BSA-Seq and linkage mapping, had not been documented in prior studies. Within the ~2Mb interval of \u003cem\u003eqPCB03\u003c/em\u003e, a total of 349 genomic variants were discovered, including six single nucleotide polymorphisms (SNPs) that resulted in nonsynonymous mutations in six genes. By assessing allelic effects in peanut germplasm and analyzing transcriptome sequencing data, five candidate genes (\u003cem\u003eAh13g469400\u003c/em\u003e, \u003cem\u003eAh13g473200\u003c/em\u003e, \u003cem\u003eAh13g475000\u003c/em\u003e, \u003cem\u003eAh13g477200\u003c/em\u003eand \u003cem\u003eAh13g477400\u003c/em\u003e) were identified. According to the marker-assisted selection, favorable genotypes could potentially enhance protein content by 1.23% to 1.57% in the RIL populations. The identification of stable loci and the development of markers facilitate marker-assisted breeding in peanuts, while the discovery of candidate genes lays the groundwork for the fine mapping of key genes regulating protein content.\u003c/p\u003e","manuscriptTitle":"Identification and utilization of a novel major QTL and linked markers for enhancing protein content in peanut","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-03-31 10:18:44","doi":"10.21203/rs.3.rs-6240685/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revisions","date":"2025-05-20T15:23:02+00:00","index":"","fulltext":""},{"type":"reviewerAgreed","content":"","date":"2025-03-22T23:52:53+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-03-21T23:35:46+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-03-17T05:16:05+00:00","index":"","fulltext":""},{"type":"submitted","content":"Theoretical and Applied Genetics","date":"2025-03-16T23:50:42+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"theoretical-and-applied-genetics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"taag","sideBox":"Learn more about [Theoretical and Applied Genetics](https://www.springer.com/journal/122)","snPcode":"122","submissionUrl":"https://submission.nature.com/new-submission/122/3","title":"Theoretical and Applied Genetics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"b00551cf-627d-4236-a3d2-bd2e7c48dbc2","owner":[],"postedDate":"March 31st, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-08-25T16:40:45+00:00","versionOfRecord":{"articleIdentity":"rs-6240685","link":"https://doi.org/10.1007/s00122-025-05000-z","journal":{"identity":"theoretical-and-applied-genetics","isVorOnly":false,"title":"Theoretical and Applied Genetics"},"publishedOn":"2025-08-22 16:29:44","publishedOnDateReadable":"August 22nd, 2025"},"versionCreatedAt":"2025-03-31 10:18:44","video":"","vorDoi":"10.1007/s00122-025-05000-z","vorDoiUrl":"https://doi.org/10.1007/s00122-025-05000-z","workflowStages":[]},"version":"v1","identity":"rs-6240685","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6240685","identity":"rs-6240685","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.