Whole-genome resequencing uncovers population structure and candidate molecular markers for litter size in hetian sheep.

OA: gold CC-BY-NC-ND-4.0

Abstract

BACKGROUND: This study leveraged whole-genome resequencing to investigate the genetic architecture, population structure, and kinship dynamics of the Hetian sheep population. By integrating genomic data with reproductive trait analysis, we aimed to identify key candidate genes associated with litter size. The findings provide new insights into the genetic mechanisms underlying litter size and establish a molecular foundation for the application of marker-assisted selection and genetic improvement in sheep breeding. RESULTS: WGRS was performed on 198 Hetian sheep. After stringent quality control, 5,483,923 high-quality SNPs were retained for downstream analysis and functionally annotated using ANNOVAR. Population genetic structure was assessed based on stratification patterns and kinship coefficients. The analysis revealed that the Hetian sheep population harbors substantial genetic diversity and exhibits a generally low level of inbreeding. Among the individuals analyzed, 157 were grouped into 16 families based on third-degree kinship (kinship coefficients between 0.12 and 0.25), while 41 individuals showed no detectable third-degree relationships, suggesting high genetic independence within the population. A genome-wide association study (GWAS) using a general linear model (GLM) identified 11 candidate genes potentially associated with litter size, including LOC101120681, LOC106990143, LOC101114058, GALNTL6, CNTNAP5, SAP130, EFNA5, ANTXR1, SPEF2, ZP2, and TRERF1. Among these, 23 SNPs located within five core candidate genes (LOC101120681, LOC106990143, LOC101114058, GALNTL6, and CNTNAP5) were selected for validation using the Sequenom MassARRAY® genotyping platform. Of the 23 SNPs tested, 22 were confirmed as true variants. However, the majority (17/22) showed no statistically significant association with litter size (P >0.05), highlighting the need for further validation in larger populations. CONCLUSIONS: Although the Hotan sheep population exhibits high genetic diversity and low levels of inbreeding, most SNPs are not significantly associated with litter size, indicating that there are currently some limitations to this study. These results should be considered preliminary and require further validation in larger and more diverse populations.
Full text 51,941 characters · extracted from pmc-nxml · 6 sections · click to expand

Results

After stringent quality control, a total of 49,800,754 SNPs were identified across the genomes of 198 Hetian sheep for WGRS analysis. Among these, 382,437 SNPs were located within 1 kb upstream of genes, and 405,142 SNPs were found within 1 kb downstream. The majority of SNPs (61.447%) were located in intergenic regions, followed by 16,287,286 SNPs in intronic regions. The fewest SNPs were detected at splice sites, with a total of 2,817 SNPs identified in these regions. These splice site variants may potentially affect transcriptional splicing, thereby altering protein products and functions (Table  1 ). Ultimately, 5,483,923 high-quality SNPs, evenly distributed across the genome, were retained for downstream analysis (Fig.  1 ). Table 1 Summary of SNP detection and functional annotation in Hetian sheep term Number of SNPs intergenic 30,614,484 intronic 16,287,286 ncRNA_intronic 1,221,678 downstream 405,142 upstream 382,437 exonic 339,744 UTR3 270,729 ncRNA_exonic 179,709 UTR5 82,555 upstream; downstream 13,234 splicing 2,817 ncRNA_splicing 583 UTR5;UTR3 266 exonic; splicing 74 ncRNA_exonic; splicing 16 Total SNP 49,800,754 Fig. 1 Genome-wide distribution of high-quality SNPs across all chromosomes in Hetian sheep after quality control Summary of SNP detection and functional annotation in Hetian sheep Genome-wide distribution of high-quality SNPs across all chromosomes in Hetian sheep after quality control Analysis of nucleotide heterozygosity (He) revealed an average heterozygous allele frequency of 0.423, indicating no significant segregation at this locus within the population (Fig.  2 a). The linkage disequilibrium (LD) analysis showed a rapid decay within the 0–50 kb range, where the LD coefficient (r²) decreased from 0.6 to 0.1. This suggests a relatively high LD decay rate in the Hetian sheep population, indicative of weak selection pressure and high genetic diversity (Fig.  2 b). The distribution of runs of homozygosity (ROH) across chromosomes showed that the majority of ROH segments within individuals were between 1 and 2 Mb in length. Long ROH segments exceeding 5 Mb were rare. Since longer ROH segments are typically associated with higher inbreeding coefficients, the results suggest a generally low level of inbreeding in the population (Fig.  2 c and d). Fig. 2 Genetic diversity analysis of the Hetian sheep population. A  Nucleotide heterozygosity (He) plot; B Linkage disequilibrium (LD) decay plot; C Genome-wide distribution of runs of homozygosity (ROH); D  ROH and heterozygosity at the individual level Genetic diversity analysis of the Hetian sheep population. A  Nucleotide heterozygosity (He) plot; B Linkage disequilibrium (LD) decay plot; C Genome-wide distribution of runs of homozygosity (ROH); D  ROH and heterozygosity at the individual level The population genetic structure of Hetian sheep is shown in Fig.  3 . Principal component analysis (PCA) revealed that the Hetian sheep population could be divided into three distinct clusters. As shown in (Fig.  3 a), the majority of individuals were grouped into one main cluster, while a smaller subset formed two additional clusters. The neighbor-joining (NJ) tree further supported these findings, with individuals from the same subgroups clustering together, indicating a clear population structure consistent with the PCA results (Fig.  3 b). ADMIXTURE analysis showed that when K = 2, the cross-validation (CV) error was 0.582, suggesting that the population could be assigned to two ancestral genomic components, with limited admixture and low genetic differentiation among individuals. When K = 3, the CV error slightly decreased to 0.578. At this value, a portion of the population still belonged to a distinct ancestral component, while others displayed signs of admixture, indicating shared ancestry and genetic contributions between groups. At K = 4 (CV error = 0.579), the population was subdivided into four ancestral clusters, and at K = 5 (CV error = 0.587), it was further subdivided into five clusters (Fig.  3 c). These results collectively indicate that while the Hetian sheep population exhibits some degree of genetic differentiation, the overall divergence is relatively low. Some individuals share similar genetic backgrounds, suggesting the presence of gene flow and a shared ancestral lineage within the population. Fig. 3 Genetic structure analysis of the Hetian sheep population  A  Principal component analysis (PCA) plot; B  Neighbor-Joining (NJ) tree; C Population structure inferred by ADMIXTURE analysis at different K values. Genetic structure analysis of the Hetian sheep population  A  Principal component analysis (PCA) plot; B  Neighbor-Joining (NJ) tree; C Population structure inferred by ADMIXTURE analysis at different K values. Using the KING software, identity-by-descent (IBD) analysis was performed to infer shared ancestry among individuals. Based on third-degree kinship coefficients (0.12–0.25), 157 individuals were grouped into 16 families. The remaining 41 individuals showed no kinship within the third-degree threshold and were considered unrelated within the population (Fig.  4 ). Fig. 4 Kinship and Pedigree Classification Analysis  A  Identity-by-Descent (IBD) analysis plot; B  Cluster analysis of familial relationships Kinship and Pedigree Classification Analysis  A  Identity-by-Descent (IBD) analysis plot; B  Cluster analysis of familial relationships The GWAS results identified a total of 28 candidate genes, including LOC101120681, LOC106990143, LOC101114058, GALNTL6, CNTNAP5, SPATA21, NECAP2, WDR33, AMMECR1L, SAP130, EFNA5, TKDP5, MGLL, TREX1, LOC101110882, SHISA5, VPS13D, ANTXR1, SPEF2, LEXM, DSCAML1, LOC105612888, WDFY2, ZP2, TRERF1, IQCH, LOC121817066, and LOC101121870 (see Extended Data Table 2 ). Based on literature review and enrichment analyses, 11 of these genes—LOC101120681, LOC106990143, LOC101114058, GALNTL6, CNTNAP5, SAP130, EFNA5, ANTXR1, SPEF2, ZP2, and TRERF1—were found to be potentially associated with key traits in Hetian sheep, including growth, litter size, follicular development and maturation, as well as cell proliferation and apoptosis (Fig.  5 ; Table 2 ). However, it should be pointed out that although 11 candidate genes were highlighted, most of the tested SNPs did not show a statistically significant association with litter size, which limits their application in breeding without further validation. Fig. 5 GWAS Analysis of Litter Size Trait in Hetian Sheep Table 2 Molecular genetic markers associated with litter size in Hetian sheep CHR POS p -value GENO Local 12 67,933,009 5.79101476009951e-07 LOC101120681 (67931461–68024608) Intronic 12 67,935,412 5.79101476009951e-07 12 67,935,544 5.79101476009951e-07 12 67,936,209 5.79101476009951e-07 13 82,656,803 3.79337050044708e-07 LOC106990143 (82362609–82372050) Intergenic 13 82,657,900 3.79337050044708e-07 13 82,658,537 7.5737179243848e-08 13 82,658,542 7.5737179243848e-08 13 82,659,604 3.79337050044708e-07 13 82,660,731 1.25737820299565e-07 13 82,660,845 1.8955659714419e-07 13 82,662,784 8.11711556460958e-08 13 82,664,939 2.98921969497857e-07 13 82,665,093 2.98921969497857e-07 18 23,826,114 3.22061940593708e-07 LOC101114058 (23683796–23685176) Intergenic 18 23,826,215 3.22061940593708e-07 18 23,826,452 7.80985437621257e-07 18 23,826,804 7.80985437621257e-07 18 23,826,966 3.18559868079765e-07 18 23,826,973 1.32464706166609e-07 2 117,616,309 3.52535308283586e-08 GALNTL6 (116184753–118014716) Intronic 2 117,645,369 4.54736519437034e-07 2 203,717,603 8.78425645187341e-07 CNTNAP5 (203610272–204846077) Intronic GWAS Analysis of Litter Size Trait in Hetian Sheep Molecular genetic markers associated with litter size in Hetian sheep LOC101120681 (67931461–68024608) LOC106990143 (82362609–82372050) LOC101114058 (23683796–23685176) GALNTL6 (116184753–118014716) CNTNAP5 (203610272–204846077) Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were conducted on the candidate genes, with the results presented in Fig.  6 . The GO enrichment analysis indicated that the candidate genes are involved in three major categories: Biological Process (BP), Cellular Component (CC), and Molecular Function (MF). Within the biological process category, genes were significantly enriched in carbohydrate derivative metabolic processes (GO:1901135), glycoprotein metabolic processes (GO:0009100), and macromolecule biosynthetic processes (GO:0034645). Regarding cellular components, significant enrichment was observed in the Golgi membrane (GO:0000139) and organelle envelope membrane (GO:0098588). In terms of molecular function, genes showed significant enrichment in carbohydrate binding (GO:0030246) and glycosyltransferase activity (GO:0016757). Fig. 6 GO and KEGG Pathway Enrichment Analysis GO and KEGG Pathway Enrichment Analysis KEGG pathway analysis revealed that these genes are significantly enriched in mucin-type O-glycan biosynthesis (oas00512) and other types of O-glycan biosynthesis pathways (oas00514). The Sequenom MassARRAY ® SNP genotyping technology was employed to validate the detection rates of 23 targeted SNP loci. Among these, the locus g. 13_82658537_542#2 exhibited a detection rate of 0%, while the locus g. 18_23826966_973#1 showed a detection rate of 82.65%. The detection rates for the remaining loci all exceeded 90% (see Extended Data Table 3 . Table 3 Genetic analysis of SNP loci in the Hetian sheep population SNPs Genotypic Frequency Gene Frequency He Ho Ne PIC χ 2 ( P ) 2_117616309 AA 0.512(106) CA 0.488(101) CC 0(0) A 0.756 C 0.244 0.184 0.816 1.226 0.116 21.55 (0) 2_117645369 GG 0.132(29) CG 0.466(102) CC 0.402(88) G 0.365 C 0.635 0.464 0.536 1.866 0.357 0.004 (0.998) 12_67933009 GG 0.175(37) CG 0.377(80) CC 0.448(95) G 0.363 C 0.637 0.462 0.538 1.860 0.355 7.195 (0.027) 12_67935412 TT 0.206(45) CT 0.353(77) CC 0.441(96) T 0.383 C 0.617 0.473 0.527 1.898 0.361 13.918 (0.001) 12_67935544 GG 0.438(96) CG 0.297(65) CC 0.265(58) G 0.587 C 0.413 0.485 0.515 1.942 0.367 32.963 (0) 12_67936209 AA 0.443(97) GA 0.388(85) GG 0.169(37) A 0.637 G 0.363 0.462 0.538 1.859 0.355 5.659 (0.059) 13_82656803 AA 0.132(29) GA 0.502(110) GG 0.365(80) A 0.384 G 0.616 0.473 0.527 1.898 0.361 0.846 (0.655) 13_82659604 AA 0.132(29) GA 0.502(110) GG 0.365(80) A 0.384 G 0.616 0.473 0.527 1.898 0.361 0.846 (0.655) 13_82662784 AA 0.132(29) GA 0.502(110) GG 0.365(80) A 0.384 G 0.616 0.473 0.527 1.898 0.361 0.846 (0.655) 13_82664939 AA 0.220(48) GA 0.372(81) GG 0.408(89) A 0.406 G 0.594 0.482 0.518 1.931 0.366 11.495 (0.003) 13_82665093 GG 0.370(81) AG 0.498(109) AA 0.132(29) G 0.619 A 0.381 0.472 0.528 1.894 0.361 0.660 (0.719) 18_23826114 TT 0.114(25) CT 0.461(101) CC 0.425(93) T 0.345 C 0.655 0.452 0.548 1.825 0.350 0.095 (0.954) 18_23826215 CC 0.114(25) TC 0.457(100) TT 0.429(94) C 0.342 T 0.658 0.450 0.550 1.818 0.349 0.042 (0.979) 18_23826452 CC 0(0) TC 0.539(118) TT 0.461(101) C 0.269 T 0.731 0.393 0.607 1.647 0.316 29.779 (0) 18_23826804 AA 0.735(161) GA 0.174(38) GG 0.091(20) A 0.822 G 0.178 0.293 0.707 1.414 0.250 36.324 (0) 18_23826966_973#1 GG 0.105(19) AG 0.011(2) AA 0.884(160) G 0.110 A 0.890 0.196 0.804 1.244 0.177 161.223 (0) 18_23826966_973#2 CC 0.110(24) TC 0.459(100) TT 0.431(94) C 0.339 T 0.661 0.448 0.552 1.812 0.348 0.114 (0.944) Genetic analysis of SNP loci in the Hetian sheep population AA 0.512(106) CA 0.488(101) CC 0(0) A 0.756 C 0.244 21.55 (0) GG 0.132(29) CG 0.466(102) CC 0.402(88) G 0.365 C 0.635 0.004 (0.998) GG 0.175(37) CG 0.377(80) CC 0.448(95) G 0.363 C 0.637 7.195 (0.027) TT 0.206(45) CT 0.353(77) CC 0.441(96) T 0.383 C 0.617 13.918 (0.001) GG 0.438(96) CG 0.297(65) CC 0.265(58) G 0.587 C 0.413 32.963 (0) AA 0.443(97) GA 0.388(85) GG 0.169(37) A 0.637 G 0.363 5.659 (0.059) AA 0.132(29) GA 0.502(110) GG 0.365(80) A 0.384 G 0.616 0.846 (0.655) AA 0.132(29) GA 0.502(110) GG 0.365(80) A 0.384 G 0.616 0.846 (0.655) AA 0.132(29) GA 0.502(110) GG 0.365(80) A 0.384 G 0.616 0.846 (0.655) AA 0.220(48) GA 0.372(81) GG 0.408(89) A 0.406 G 0.594 11.495 (0.003) GG 0.370(81) AG 0.498(109) AA 0.132(29) G 0.619 A 0.381 0.660 (0.719) TT 0.114(25) CT 0.461(101) CC 0.425(93) T 0.345 C 0.655 0.095 (0.954) CC 0.114(25) TC 0.457(100) TT 0.429(94) C 0.342 T 0.658 0.042 (0.979) CC 0(0) TC 0.539(118) TT 0.461(101) C 0.269 T 0.731 29.779 (0) AA 0.735(161) GA 0.174(38) GG 0.091(20) A 0.822 G 0.178 36.324 (0) GG 0.105(19) AG 0.011(2) AA 0.884(160) G 0.110 A 0.890 161.223 (0) CC 0.110(24) TC 0.459(100) TT 0.431(94) C 0.339 T 0.661 0.114 (0.944) Based on preliminary resequencing results, 23 SNP loci were identified and subsequently validated using the Sequenom MassARRAY ® SNP genotyping technology to confirm their authenticity. Among these, 22 SNPs were verified to be genuinely segregating within the Hetian sheep population (including g.2_117616309 A > C, g.2_117645369 C > G, g.2_203717603 C > T, g.12_67933009 C > G, g.12_67935412 C > T, g.12_67935544 G > C, g.12_67936209 A > G, g.13_82656803 G > A, g.13_82657900 A = C, g.13_82658537_542#1 A > C, g.13_82659604 G > A, g.13_82660731 A > C, g.13_82660845 A > G, g.13_82662784 G > A, g.13_82664939 G > A, g.13_82665093 G > A, g.18_23826114 C > T, g.18_23826215 T > C, g.18_23826452 T > C, g.18_23826804 A > G, g.18_23826966_973#1 A > G, and g.18_23826966_973#2 T > C) (see Extended Data Fig.  1 ). Therefore, subsequent analyses focused exclusively on these 22 SNP loci. Genotyping results revealed the following genotype distributions: For loci g.2_117645369 C > G, g.12_67933009 C > G, and g.12_67935544 G > C, three genotypes were detected: CC, GC, and GG. For loci g.12_67936209 A > G, g.13_82656803 G > A, g.13_82659604 G > A, g.13_82662784 G > A, g.13_82664939 G > A, g.13_82665093 G > A, g.18_23826804 A > G, and g.18_23826966_973#1 A > G, three genotypes were identified: AA, AG, and GG. For loci g.12_67935412 C > T, g.18_23826114 C > T, g.18_23826215 T > C, and g.18_23826966_973#2 T > C, three genotypes were observed: CC, CT, and TT. For locus g.2_117616309 A > C, two genotypes were found: AA and CA. For locus g.18_23826452 T > C, two genotypes were identified: TT and TC. For loci g.13_82658537_542#1 A > C, g.13_82660731 A > C, and g.13_82660845 A > G, only the AA genotype was detected. For locus g.2_203717603 C > T, only the CC genotype was observed. For locus g.13_82657900 A = C, only the AG genotype was detected. In the Hetian sheep population, loci g.13_82658537_542#1 A > C, g.13_82660731 A > C, g.13_82660845 A > G, g.2_203717603 C > T, and g.13_82657900 A = C exhibited only a single genotype and were therefore excluded from further analysis. As shown in Table  3 , at loci g.2_117616309 A > C, g.18_23826804 A > G, g.18_23826966_973#1 A > G, and g.12_67936209 A > G, the allele A was dominant, with the AA genotype being the predominant genotype. For locus g.2_117645369 C > G, allele C was dominant, with the CG genotype as the major genotype. At loci g.12_67933009 C > G and g.12_67935412 C > T, allele C was dominant, with the CC genotype as dominant. Loci g.12_67935544 G > C and g.13_82664939 G > A showed G as the dominant allele, with GG as the dominant genotype. For loci g.13_82656803 G > A, g.13_82659604 G > A, g.13_82662784 G > A, and g.13_82665093 G > A, G was the dominant allele and GA was the predominant genotype. At locus g.18_23826114 C > T, allele C was dominant, with CT as the dominant genotype. Lastly, loci g.18_23826215 T > C, g.18_23826966_973#2 T > C, and g.18_23826452 T > C had T as the dominant allele with CT as the major genotype. Among these loci, g.2_117645369 C > G, g.12_67933009 C > G, g.12_67935412 C > T, g.12_67935544 G > C, g.12_67936209 A > G, g.13_82656803 G > A, g.13_82659604 G > A, g.13_82662784 G > A, g.13_82664939 G > A, g.13_82665093 G > A, g.18_23826114 C > T, g.18_23826215 T > C, and g.18_23826452 T > C exhibited moderate polymorphism in the Hetian sheep population (0.25 < PIC  C, g.18_23826804 A > G, g.18_23826966_973#1 A > G, and g.18_23826966_973#2 T > C showed low polymorphism (PIC ≤ 0.25). Based on Hardy-Weinberg equilibrium tests, loci g.2_117616309 A > C, g.12_67933009 C > G, g.12_67935412 C > T, g.12_67935544 G > C, g.13_82664939 G > A, g.18_23826452 T > C, g.18_23826804 A > G, and g.18_23826966_973#1 A > G deviated significantly from Hardy-Weinberg equilibrium ( P   0.05). As shown in Table  4 , there were no significant differences ( P  > 0.05) in the average litter size among different genotypes at the 17 SNP loci analyzed. Table 4 Association analysis between 22 SNP and litter size in Hetian sheep SNPs Genotype Number Litter size/ (individuals) 2_117616309 AA CA 106 101 1.594 ± 0.048 1.574 ± 0.049 2_117645369 GG CG CC 29 102 88 1.724 ± 0.084 1.556 ± 0.049 1.580 ± 0.053 12_67933009 GG CG CC 37 80 95 1.622 ± 0.033 1.588 ± 0.055 1.558 ± 0.051 12_67935412 TT CT CC 45 77 96 1.622 ± 0.073 1.610 ± 0.056 1.563 ± 0.050 12_67935544 GG CG CC 96 65 58 1.552 ± 0.051 1.631 ± 0.060 1.603 ± 0.065 12_67936209 AA GA GG 97 85 37 1.557 ± 0.051 1.612 ± 0.053 1.622 ± 0.081 13_82656803 AA GA GG 29 110 80 1.448 ± 0.094 1.627 ± 0.046 1.588 ± 0.055 13_82659604 AA GA GG 29 110 80 1.448 ± 0.094 1.627 ± 0.046 1.588 ± 0.055 13_82662784 AA GA GG 29 110 80 1.448 ± 0.094 1.627 ± 0.046 1.588 ± 0.055 13_82664939 AA GA GG 48 81 89 1.521 ± 0.073 1.617 ± 0.054 1.607 ± 0.052 13_82665093 AA GA GG 29 109 81 1.448 ± 0.094 1.624 ± 0.047 1.593 ± 0.055 18_23826114 TT CT CC 25 101 93 1.560 ± 0.101 1.545 ± 0.050 1.645 ± 0.050 18_23826215 TT CT CC 94 100 25 1.638 ± 0.050 1.550 ± 0.050 1.560 ± 0.101 18_23826452 TT CT 101 118 1.614 ± 0.049 1.568 ± 0.046 18_23826804 AA GA GG 161 38 20 1.590 ± 0.039 1.605 ± 0.080 1.550 ± 0.114 18_23826966_973#1 AA GA GG 160 2 19 1.581 ± 0.039 2.00 ± 0.000 1.526 ± 0.118 18_23826966_973#2 TT CT CC 94 100 24 1.638 ± 0.050 1.550 ± 0.050 1.542 ± 0.104 Note: Litter size is presented as mean ± standard error. Different lowercase letters within the same column for a given locus indicate significant differences ( P  < 0.05); different uppercase letters indicate highly significant differences ( P   0.05) Association analysis between 22 SNP and litter size in Hetian sheep AA CA 106 101 1.594 ± 0.048 1.574 ± 0.049 GG CG CC 29 102 88 1.724 ± 0.084 1.556 ± 0.049 1.580 ± 0.053 GG CG CC 37 80 95 1.622 ± 0.033 1.588 ± 0.055 1.558 ± 0.051 TT CT CC 45 77 96 1.622 ± 0.073 1.610 ± 0.056 1.563 ± 0.050 GG CG CC 96 65 58 1.552 ± 0.051 1.631 ± 0.060 1.603 ± 0.065 AA GA GG 97 85 37 1.557 ± 0.051 1.612 ± 0.053 1.622 ± 0.081 AA GA GG 29 110 80 1.448 ± 0.094 1.627 ± 0.046 1.588 ± 0.055 AA GA GG 29 110 80 1.448 ± 0.094 1.627 ± 0.046 1.588 ± 0.055 AA GA GG 29 110 80 1.448 ± 0.094 1.627 ± 0.046 1.588 ± 0.055 AA GA GG 48 81 89 1.521 ± 0.073 1.617 ± 0.054 1.607 ± 0.052 AA GA GG 29 109 81 1.448 ± 0.094 1.624 ± 0.047 1.593 ± 0.055 TT CT CC 25 101 93 1.560 ± 0.101 1.545 ± 0.050 1.645 ± 0.050 TT CT CC 94 100 25 1.638 ± 0.050 1.550 ± 0.050 1.560 ± 0.101 TT CT 101 118 1.614 ± 0.049 1.568 ± 0.046 AA GA GG 161 38 20 1.590 ± 0.039 1.605 ± 0.080 1.550 ± 0.114 AA GA GG 160 2 19 1.581 ± 0.039 2.00 ± 0.000 1.526 ± 0.118 TT CT CC 94 100 24 1.638 ± 0.050 1.550 ± 0.050 1.542 ± 0.104 Note: Litter size is presented as mean ± standard error. Different lowercase letters within the same column for a given locus indicate significant differences ( P  < 0.05); different uppercase letters indicate highly significant differences ( P   0.05)

Materials

In this study, we first performed whole-genome resequencing (WGRS) of 198 Hetian sheep to obtain genome-wide variants. These data were used for population genetic diversity analysis (including ROH, LD, PCA, NJ tree, and ADMIXTURE) and for a genome-wide association study (GWAS) to identify candidate loci related to litter size. Subsequently, 23 SNPs from five candidate genes were selected for targeted genotyping validation using the MassARRAY ® platform in an independent population of 219 sheep. Finally, we performed an association analysis between validated SNPs and litter size. We collected blood samples from 198 healthy female Hetian sheep (aged 2–3 years) raised under natural grazing conditions by private farmers in Hotan County, Xinjiang, China. Approximately 3 mL of blood was collected from the jugular vein into EDTA-K2 anticoagulant tubes and stored at − 20 °C until used for selection signature analysis. To validate the association between candidate SNPs and litter size, an additional 219 female Hetian sheep were sampled from another flock in the same region. All sample collections were conducted with informed consent from the animal owners and were approved by the relevant institutional ethics committee. The integrity, concentration, and purity of genomic DNA extracted from Hetian sheep blood samples were assessed using 1% agarose gel electrophoresis and ultraviolet spectrophotometry. Samples that passed quality control were used to construct sequencing libraries, with 1.5 µg of high-quality genomic DNA from each individual. The insert size of each library was evaluated, and only libraries with fragment sizes meeting the expected criteria were subjected to sequencing on the Illumina NovaSeq PE150 platform (Illumina, San Diego, CA, USA). Raw image data obtained from high-throughput sequencing were converted into raw sequencing reads through base calling. Quality control and filtering were performed using FASTP v0.23.2 [ 23 ], which removed adapter sequences, paired-end reads with more than 10% unidentified bases (N), and reads with over 50% low-quality bases. The resulting clean reads were retained for downstream analyses. The clean reads were aligned to the Ovis aries reference genome (Oar_v4.0) ( https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_000298735.2/ ) using BWA v0.7.17 [ 24 ]. SNP detection and genotyping were performed using the Genome Analysis Toolkit (GATK) v4.1.4.1 [ 25 ]. SNP filtering was conducted with VCFTOOLS v0.1.16 using the following parameters: vcftools --vcf A.vcf --maf 0.05 --max-missing 0.45 --minDP 3. High-confidence SNPs passing the quality filters were functionally annotated using ANNOVAR [ 26 ]. Genetic diversity reflects the degree of genetic variation either among populations within a species or among individuals within a population. In this study, only biallelic SNPs were retained for genetic diversity analysis. Quality control was performed using PLINK v1.90 [ 27 ], with SNPs filtered out based on the following criteria: missing genotype rate > 10%, minor allele frequency (MAF) < 5%, and deviation from Hardy–Weinberg equilibrium (HWE) with P < 1 × 10⁻⁶. Expected heterozygosity (He) refers to the average expected heterozygosity per locus. In this study, He was calculated using PLINK v1.90 [ 27 ] with the command:./plink --hardy. Runs of Homozygosity (ROH) refer to consecutive, relatively long homozygous regions in an individual’s genome [ 28 ]. These regions consist of two identical alleles, often resulting from inbreeding or genetic drift within the population’s history. The length and distribution of ROH can reflect the level of inbreeding, genetic diversity, and population history. The length of ROH segments is indicative of the degree of inbreeding within a population, with longer ROH segments typically associated with inbreeding events. ROH analysis was performed using PLINK v1.90 [ 27 ]. In this study, The analysis parameters were defined as follows: (1) up to one missing SNP and a maximum of one potential heterozygous genotype were allowed within a run of homozygosity (ROH); (2) the minimum ROH length was set to 1,000 kb to exclude short ROHs potentially arising from linkage disequilibrium; (3) the minimum SNP density per ROH segment was required to be 100; (4) the maximum gap between two consecutive SNPs within an ROH was set at 250 kb; and (5) to minimize false positives, the minimum number of SNPs constituting an ROH was determined according to the method proposed by Lentz et al. [ 29 ]. Linkage disequilibrium (LD) refers to the non-random association of alleles at different loci during genetic transmission. LD is present when the observed frequency of allele combinations deviates from the expected frequency under random assortment [ 30 ]. Multiple factors influence LD, including mating patterns, genetic drift, natural selection, migration, and population structure. Generally, populations under selection—such as domesticated breeds—exhibit higher levels of LD across the genome compared to natural populations. In this study, the squared correlation coefficient (r²) was used as the measure of LD, ranging from 0 to 1. An r² value of 0 indicates complete linkage equilibrium, where alleles at two loci assort independently, while an r² of 1 indicates complete linkage disequilibrium, where alleles are perfectly linked. LD analysis was conducted using PopLDdecay v3.40 [ 31 ]with the command:./PopLDdecay -InVCF A.vcf, and visualization was performed using the script: perl./Plot_MultiPop.pl. Principal Component Analysis (PCA) is a dimensionality reduction technique that simplifies complex data by extracting the main components representing the majority of variance. PCA enables the visualization of relationships and genetic distances among samples, which can aid in evolutionary and population genetic analyses. In this study, VCFTOOLS v0.1.16 [ 32 ] was first used to convert VCF files to PLINK format with the command: vcftools --vcf final.vcf --plink --out final. Subsequently, PCA was performed using PLINK v1.90 with the command:./plink --bfile A --pca --out B. The first two principal components (PC1 and PC2) were extracted and visualized using RStudio v4.4.3 [ 33 ]. In this study, a Neighbor-Joining (NJ) tree was constructed to infer the evolutionary relationships among individuals based on the principle of minimum evolution. The NJ method generates a tree-like structure that illustrates the genetic proximity and divergence among samples during the course of evolution. To construct the tree, a pairwise identity-by-state (IBS) distance matrix was generated using VCF2Dis ( https://github.com/BGI-shenzhen/VCF2Dis ) with the command: VCF2Dis -i A.vcf -o B.mat. Subsequently, the distance matrix was used as input for FastME 2.0, accessed via the ATGC platform ( http://www.atgc-montpellier.fr/fastme/ ), to construct the NJ tree and output a tree file in Newick format (C.nwk). The resulting phylogenetic tree was then visualized and annotated using iTOL (Interactive Tree Of Life; https://itol.embl.de/ ). ADMIXTURE v1.30 [ 34 ] was employed to estimate population structure based on a maximum likelihood approach derived from a Bayesian clustering model. The number of ancestral populations (K) was set from 2 to 5, and the optimal K value was determined by minimizing the cross-validation (CV) error. The resulting ancestry proportions were visualized using RStudio v4.4.3, with the aid of the pophelper and gtable packages. In this study, a General Linear Model (GLM) was used for the genome-wide association study (GWAS) [ 35 ]. The first three principal components from PCA analysis were included as covariates to control for population structure. The GLM framework evaluates the association between single nucleotide polymorphism (SNP) genotypes and phenotypic traits through regression analysis. In this model, the phenotype serves as the dependent variable, while the SNP genotype is treated as the independent variable, allowing for the estimation of the genotype’s effect on the trait. To minimize false positives due to population stratification, the first three principal components from PCA were included as covariates in the GLM; genomic control (lambda) was also assessed to check for inflation of test statistics. The model is expressed as follows:  \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$Y=X\alpha+Z\beta+e$$\end{document} Where: Y represents the phenotypic trait, Xα denotes the fixed effects of population structure, Zβ corresponds to the marker (SNP) effect, and e is the residual error. Candidate genes were identified by annotating SNPs and their flanking 1 kb regions based on the reference annotation file (GCF_000298735.2_Oar_v4.0_genomic.gtf). Functional annotation and pathway enrichment analyses were subsequently performed using Gene Ontology (GO: http://geneontology.org/ ) and the Kyoto Encyclopedia of Genes and Genomes (KEGG: https://www.kegg.jp/ ). GO terms and KEGG pathways with a P-value < 0.01 were considered significantly enriched. Data visualization was carried out using RStudio and the online platform Bioinformatics.com.cn ( http://www.bioinformatics.com.cn/ ). To further identify key genetic markers involved in the regulation of litter size in Hetian sheep, SNPs within five candidate reproductive genes were screened based on WGRS variant analysis. A total of 23 SNP loci potentially associated with litter size traits were selected. Based on gene sequence information from the NCBI database (Ovis aries), PCR primers and single-base extension primers were designed using AssayDesigner 3.1 software (Sequenom) (see Extended Data Table 1). All primers were synthesized by Beijing Compass Agricultural Technology Co., Ltd. Polymorphism detection of the 23 SNP loci was conducted in a cohort of 219 Hetian sheep using the Sequenom MassARRAY ® SNP genotyping platform, which utilizes matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass spectrometry. This analysis was performed to verify the authenticity and presence of the candidate variants within the Hetian sheep population.

Conclusion

Although the Hetian sheep population shows high diversity and low inbreeding, most SNPs showed no significant association with litter size. Further validation in larger, diverse populations and functional studies are needed to confirm their relevance for breeding.

Discussion

Population genetic diversity provides an accurate framework for analyzing the ancestral groups and individual lineage composition within livestock populations, encompassing methods such as PCA, phylogenetic analysis, and population structure assessment. The concordance among these results is crucial for the effective utilization and conservation of genetic resources. In this study, the population structure of Hetian sheep was investigated. PCA revealed that a few individuals deviated from the main cluster, which may be attributed to the sensitivity of PCA to missing data or the influence of multiple factors such as seasonality and parity. Phylogenetic tree analysis indicated that the 198 Hetian sheep share a common genetic background, with no distinct separation between groups; however, a subset of individuals exhibited genetic admixture. Population structure analysis based on ADMIXTURE demonstrated that at K = 2, the Hetian sheep population was not completely divided, likely due to shared genomic ancestry among some individuals between the two groups. Analysis of population genetic structure enhances our understanding of the domestication processes in various species or breeds. Natural selection events and reproductive behaviors experienced by different species or populations can be inferred through patterns of linkage disequilibrium (LD). Domestication typically leads to a reduction in genetic diversity and an increase in LD among loci. The greater the degree of domestication and selection intensity, the slower the LD decay within the population. The LD decay results (Fig. 2 D) reveal a rapid decline in LD within 0–50 kb, with the LD coefficient (r²) decreasing from 0.6 to 0.1. This indicates a relatively high LD decay rate in the Hetian sheep population, suggesting weak selection pressure and high genetic diversity. These findings demonstrate that while there is some degree of population genetic differentiation in Hetian sheep, the extent of divergence is limited, potentially due to sampling from a geographically restricted area or because the population has not yet undergone strong artificial selection, which aligns with the actual breeding context. Using KING software based on identity-by-descent (IBD) analysis, 157 individuals were grouped into 16 families under third-degree kinship (0.12–0.25), whereas 41 individuals showed no third-degree kinship relationships within the population. In this study, we employed whole-genome sequencing data from 198 Hetian sheep samples to comprehensively characterize population genetic structure and genomic diversity, thereby elucidating the genome-wide genetic mechanisms involved in the breed’s improvement. Advances in genomics have recently yielded novel genotyping tools, enabling breeders to identify specific genomic segments inherited from parents to offspring, facilitating more precise manual selection of desirable traits [ 36 ]. For example, Reza Talebi et al. [ 13 ] employed genotyping-by-sequencing (GBS) to assess the genetic diversity, admixture patterns, and signatures of selection in Moghani sheep. Their results revealed that purebred Moghani sheep exhibited the highest genetic diversity (HO = 0.521 ± 0.10) and the lowest inbreeding coefficient (FIS = − 0.474), acting as an important genetic bridge among populations. In contrast, the BRM and TTM populations showed lower heterozygosity (HO = 0.410 ± 0.09 and 0.431 ± 0.10, respectively), higher inbreeding levels (FROH), and extended runs of homozygosity (ROH), indicating recent inbreeding events and reduced effective population sizes. In another study, the same research group (Reza Talebi et al. [ 14 ]) systematically compared the genomic distribution of ROH islands across five chicken populations, including layers and broilers (Runs of Homozygosity in Modern Chicken Revealed by Sequence Data). The results demonstrated significant differences in ROH size among populations, with an average total ROH length of 432.1 Mb (618.7) per individual. Layer lines exhibited higher levels of autozygosity compared with broilers and red junglefowl, whereas wild birds carried predominantly shorter ROH segments, which may reflect more ancient evolutionary events. Similarly, Zhipeng Han et al. [ 37 ] analyzed the genetic diversity and homozygosity of Hetian sheep using the SNP50K BeadChip. Their findings indicated that MGTS and GTS shared a common ancestral origin. ROH analysis further showed that in MTS, GTS, and MGTS populations, the majority of ROH segments were concentrated within 10 Mb. In addition, Sara M. Nilson et al. [ 38 ] reported 713,157 ROH segments in U.S. Katahdin hair sheep, averaging 73.5 per individual, with ROH lengths ranging from 0.25 to 179.5 Mb (mean = 5.56 Mb; median = 4.01 Mb). Consistent with these studies, our analysis of Hetian sheep revealed that most ROH segments within individuals were concentrated in the 1–2 Mb range, while segments longer than 5 Mb were relatively rare. Generally, longer ROH segments were associated with higher individual inbreeding coefficients. At the population level, however, Hetian sheep exhibited an overall low degree of inbreeding. Taken together, the primary objective of this study is to identify key genomic regions and molecular markers under selection during the improvement of Hetian sheep. Hetian sheep, a local breed from the Hotan region of Xinjiang, China, have adapted to the arid climate, large temperature fluctuations, and limited forage resources along the southern edge of the Tarim Basin. Field investigations revealed that some ewes in this population are capable of producing twins or even triplets, indicating a naturally high reproductive potential. To uncover the genetic basis of this trait, we performed WGRS of 196 Hetian sheep and conducted a genome-wide association study (GWAS) to identify SNPs associated with litter size. This analysis highlighted 11 genes (LOC101120681, LOC106990143, LOC101114058, GALNTL6, CNTNAP5, SAP130, EFNA5, ANTXR1, SPEF2, ZP2, and TRERF1) that are potentially involved in growth and reproductive traits. Notably, LOC101120681, LOC106990143, and LOC101114058 have not been previously reported, suggesting novel genetic contributors to litter size in sheep. Previous studies have shown that CNTNAP5 is associated with chest width, measured as bilateral rib diameter, in Sudanese goats, whereas GALNTL6 has been reported to influence growth traits in sheep, cattle, and other livestock species [ 39 , 40 ]. In addition, SAP130 plays a critical regulatory role in placental and fetal development, while EFNA5 has been demonstrated to affect both milk yield and reproductive performance in cattle [ 41 , 42 ]. In mouse ovarian granulosa cells (GCs), ANTXR1 may promote cell proliferation and inhibit apoptosis, with mutations also linked to ovarian dysfunction [ 43 ]. SPEF2 has been associated with reproductive performance in Chinese indigenous pigs [ 44 ]. Furthermore, in FecB-mutant sheep, granulosa cells regulate nutrient metabolism by modulating the abundance of the zona pellucida protein ZP2, thereby enhancing oocyte maturation and ovulation, ultimately improving reproductive efficiency [ 45 ]. Mutations in TRERF1 may also contribute to the pathogenesis of ovarian endometriosis [ 46 ]. Collectively, functional evidence across species suggests that these genes play conserved and pivotal roles in regulating growth and reproductive traits. Although 23 SNP loci potentially associated with the twinning trait have been identified, their validity as genetic markers for twinning requires further validation in local breeds and individual Tian sheep populations. Moreover, the underlying molecular mechanisms by which these loci influence ovine litter size remain to be elucidated through comprehensive functional studies. A total of 49,800,754 SNP loci were initially identified from 198 Hetian sheep individuals. After stringent quality control and filtering, 5,483,923 high-quality SNPs were retained, which were evenly distributed across the chromosomes. The sequencing coverage indicated that a substantial amount of high-quality sequencing data was obtained, providing extensive base coverage of the reference genome. High coverage reduces the likelihood of missing variants and sequencing errors, thereby offering a more comprehensive and accurate representation of the genome. Furthermore, such high-depth data enhances the reliability and interpretability of downstream analyses, including variant annotation, mutation detection, and investigations into genome structure and function. Lan Rong et al. [ 47 ] performed 20× WGRS on three representative Yunnan Black goats and aligned the data to the goat reference genome. They identified over seven million SNPs, more than 800,000 InDels, and over 40,000 structural variants (SVs), effectively elucidating the molecular characteristics of Yunnan Black goats and providing a robust data foundation for the identification and functional analysis of candidate genes. As whole-genome sequencing becomes increasingly utilized for detecting adaptive variants, the comprehensive and accurate identification of InDels will become ever more critical [ 48 ]. Moreover, InDels and SVs represent significant forms of genetic variation that contribute substantially to genomic diversity [ 49 ]. SNP genotyping technologies are also widely employed in studies of animal breed diversity and population structure [ 50 ]. Utilizing genotyping methods can assist in identifying animals with superior adaptive traits, such as litter size, thereby ensuring the long-term viability of breeding improvement strategies. These approaches offer a genetic selection model to support breeding programs aimed at producing high-quality and well-adapted breeds. Zhao et al. [ 51 ] utilized population-wide single nucleotide polymorphisms (SNPs) to identify candidate genes associated with sheep traits including high fertility, coat color, tail type, and horn size and morphology. In this study, association analysis identified 23 SNP loci related to twinning traits. Further investigation revealed that 17 of these SNP loci—namely g.2_117616309 A > C, g.2_117645369 C > G, g.12_67933009 C > G, g.12_67935412 C > T, g.12_67935544 G > C, g.12_67936209 A > G, g.13_82656803 G > A, g.13_82659604 G > A, g.13_82662784 G > A, g.13_82664939 G > A, g.13_82665093 G > A, g.18_23826114 C > T, g.18_23826215 T > C, g.18_23826452 T > C, g.18_23826804 A > G, g.18_23826966_973#1 A > G, and g.18_23826966_973#2 T > C—were genuinely segregating within the Hetian sheep population. Using SPSS 26.0 software, the average lambing number for each genotype at these loci was analyzed, and no significant differences were found among genotypes for any of the 17 loci ( P > 0.05). The limited variation observed at these SNP loci within the population may be attributable to the genetic background of the local sheep breed and the degree of selective breeding applied to this population [ 52 ].Nine of the SNP loci were found to be in Hardy–Weinberg equilibrium, indicating that the experimental population is relatively unaffected by external selection pressures. This suggests a high degree of sequence conservation at these loci. Under the influence of factors such as selective breeding, migration, and genetic drift, these loci appear to maintain a dynamic equilibrium, and their inheritance follows a random pattern. The Hardy–Weinberg equilibrium is influenced not only by genetic laws but also by external environmental factors. From a genetic perspective, deviations from equilibrium can result from changes in allele frequencies due to natural selection, genetic drift, or gene flow from neighboring populations through hybridization. Such processes may introduce new genetic components into the studied population or alter its existing genetic composition. This study identified that among the 198 Hetian sheep individuals, two SNP loci (g.2_117616309 A > C and g.18_23826452 T > C) exhibited two genotypes each, while five loci (g.13_82658537_542#1 A > C, g.13_82660731 A > C, g.13_82660845 A > G, g.2_203717603 C > T, and g.13_82657900 A = C) exhibited only one genotype. The remaining 15 loci showed all three possible genotypes. Although 17 SNPs were identified, none exhibited statistically significant differences between the twin-bearing and single-lamb groups. The lack of significant associations for most SNPs may be attributed to the limited sample size, the polygenic nature of litter size, or environmental influences, all of which could obscure true genetic effects. In light of these findings, future studies should prioritize increasing the sample size to more accurately characterize the genetic architecture of indigenous Chinese sheep breeds. Such efforts would provide additional evidence to support marker-assisted selection for litter size traits and further enrich the genetic dataset for genomic diversity studies in Hetian sheep. Understanding this diversity may help elucidate interactions between different breeds within shared geographical regions and offer insights into animal genetic evolution and their historical origins in ancient areas around the world. In rural regions where herd sizes are small and performance recording is challenging, many global and local conservation initiatives emphasize the importance of maintaining genetic diversity [ 53 ]. This diversity is essential to support ongoing breeding and conservation efforts. Ultimately, the findings of this study may contribute valuable references for the molecular marker-assisted breeding of twin-bearing traits in local sheep breeds like the Hetian sheep. While this study provides valuable genomic insights, it has several limitations. First, the sample size of 198 individuals, while reasonable for population genetic analysis, may limit the statistical power of GWAS to detect variants with small effect sizes. Second, the lack of functional annotation for several candidate genes (particularly LOC genes) hinders immediate biological interpretation. Third, environmental factors and management practices that may influence litter size were not accounted for in this genetic analysis. Future studies with larger sample sizes, functional validation experiments, and integrated analyses considering both genetic and environmental factors will be necessary to fully elucidate the genetic architecture of litter size in Hetian sheep.

Introduction

As one of the earliest domesticated livestock species, sheep have played a pivotal role in meat, milk, and wool production. litter size, a key reproductive trait, directly influences the economic return of sheep farming through litter size. However, its genetic regulation is governed by complex polygenic interactions and environmental influences, posing significant challenges to traditional phenotype-based selection strategies [ 1 ]. Hetian sheep, a unique indigenous breed from southern Xinjiang, China, exhibit remarkable adaptation to extreme environments—including tolerance to aridity, heat, and low-quality forage—as well as strong genetic resistance to disease. In addition, their semi-coarse, heterogeneous wool supports the globally renowned Hetian carpet industry [ 2 ]. Notably, although Hetian sheep exhibit year-round estrus cycles (16–19 days per cycle, with peaks in April–May and November), their average lambing rate remains low at only 102.52%. This suboptimal reproductive performance limits both population expansion and genetic improvement efforts [ 3 ]. The resulting consequences are twofold: reduced farmer enthusiasm for breeding and a threat to the sustainable utilization of this valuable genetic resource. Therefore, elucidating the genetic architecture underlying litter size in Hetian sheep is crucial for developing targeted breeding strategies and ensuring the conservation of this important local breed. Although several major litter size genes have been identified in sheep, their functional roles vary across breeds. For instance, Talebi et al. [ 4 ] reported that among four well-known fecundity genes (BMP15, GDF9, BMPR1B, and B4GALNT2) in Mehraban ewes, a novel mutation (g.41840985 C > T) was detected exclusively in the 3′UTR of the GDF9 gene on chromosome 5, and a new SNP in exon 7 of BMPR1B was found to be significantly associated with reproductive traits. Similarly, Ahmadi et al. [ 5 ] sequenced GDF9 and BMP15 and identified 12 polymorphisms, which were further evaluated through structural modeling to assess the potential impact of missense mutations on protein function. Based on these variants, individuals were classified into three haplotypes: the wild-type (without mutation), haplotype A (carrying G1, G2, G3, and G4), and haplotype B (carrying G5 and G6). Three-dimensional structural analysis revealed that mutations such as G1/p.Arg87His, G4/p.Glu241Lys, and G6/p.Val332Ile were benign, with limited effects on protein conformation. In addition, Majd et al. [ 6 ] identified seven SNPs within the KISS1R/GPR54 gene across four sheep breeds. Among them, the g.3431 C > A mutation in exon 4 exerted a significant negative effect on litter size (LS) and birth weight (BW) ( P G, g.456T > C, g.475 C > A, and g.571 A > C showed no significant associations with the traits studied ( P > 0.05). Despite the growing body of evidence suggesting potential associations between polymorphisms in genes such as GDF9, BMPR1B, and BMP15 with reproductive performance, the genetic basis underlying litter size in Hetian sheep remains largely unclear and warrants further investigation [ 7 – 11 ]. the genetic basis underlying litter size in Hetian sheep remains largely unclear and warrants further investigation. However, due to increasing anthropogenic disturbances and environmental changes in recent years, the population size of Hetian sheep has been gradually declining. To ensure the conservation and sustainable development of this breed, it is particularly important to apply advanced genetic technologies for breeding and preservation efforts. Reproductive traits represent one of the most economically important characteristics in sheep production and have long been a primary focus in both scientific research and breeding practices in China. Among these traits, litter size is particularly critical, as its genetic improvement can lead to exponential gains in reproductive efficiency and overall economic returns at the flock and industry levels. The genetic architecture underlying litter size in Hetian sheep remains poorly understood. WGRS, a cutting-edge approach based on high-throughput sequencing platforms, enables deep coverage of the entire genome and the generation of high-quality genomic data [ 12 ]. Application of this technology to the genomic sequencing of the Hetian sheep population enabled accurate characterization of its complete genomic landscape. Subsequent analysis of runs of homozygosity (ROH) provided important insights into population history and inbreeding patterns: short ROHs indicate ancient inbreeding events, whereas long ROHs reflect more recent consanguinity [ 13 – 15 ]. Following alignment of sequencing data to the reference genome, a large number of genetic variants can be identified among individuals, particularly single nucleotide polymorphisms (SNPs), which serve as critical markers for assessing genetic diversity and inferring kinship within the population. Leveraging these SNP datasets enables genome-wide genotyping and computation of genetic distances, facilitating the construction of comprehensive phylogenetic and pedigree relationship maps [ 16 ]. Moreover, by integrating molecular evolutionary models, genetic distances can be calibrated on a temporal scale, thereby elucidating the hierarchical relationships and evolutionary trajectories among individuals [ 17 ]. WGRS is not only a powerful tool for deciphering population genetic structure but also enables high-resolution variant detection and functional site identification. It facilitates the prediction of candidate genes and regulatory elements associated with economically important traits, thereby providing critical data support and theoretical foundations for molecular breeding efforts [ 18 , 19 ]. In the future, with the continuous advancement of sequencing technologies and the ongoing reduction in associated costs, WGRS is expected to play an increasingly pivotal role in the genetic research and practical breeding of Hetian sheep. Beyond Hetian sheep, this technology also holds broad applicability in the genetic studies of other animal populations and even in human genomics. In recent years, WGRS has been widely applied in various fields, including genetic improvement of livestock, conservation of germplasm resources, prediction of disease susceptibility, and novel drug development. With its notable advantages of high throughput, accuracy, and efficiency, WGRS provides powerful technical support for genetic innovation and precision breeding in livestock and poultry [ 20 – 22 ]. The objective of this study is to utilize WGRS technology to comprehensively analyze the genomic structure, gene family evolution, and distribution characteristics of regulatory elements in Hetian sheep. By integrating phylogenetic identification and population genetic analysis, this study aims to identify molecular markers and functional genes closely associated with litter size. The findings are intended to provide scientific evidence and theoretical support for the genetic improvement and efficient breeding of Hetian sheep.

Supplementary Material

Supplementary material 1. Supplementary material 1.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-13T06:15:24.848197+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-NC-ND-4.0