Heterogeneity in an Adeno-Associated Virus Transfection-Based Production Process Limits the Production Efficiency

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-08, 2026-08-01

This study characterized rAAV9 production by transient transfection and found that less than 5% of cells expressed necessary genes and produced assembled rAAV9 capsids, highlighting a significant production bottleneck.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

Abstract The popularity of rAAV vectors in gene therapy are placing a burden on current production systems. To improve the accessibility of these life changing treatments, increases in production yields and a reduction in the cost-of-goods are needed. Transient transfection is the most common way to introduce rAAV-encoding plasmids to producer cells but it suffers from significant drawbacks such as low and inconsistent yields as well as high cost due to its need for plasmid DNA. This study aims to address the low yield of transient transfection-based rAAV production through advanced methods in process characterization. Adherent and suspension cultures of a HEK293T cell line were triple-transfected for rAAV9 production using polyethylenimine (PEI). Samples were taken at various times post-transfection for analysis with bulk and single-cell transcriptomics. It was revealed that 46% of the cells lacked transcripts of genes from at least one plasmid, indicating that a significant proportion of the cells did not have the genes necessary for rAAV9 production. Among the remaining 54% of the cells expressing genes from all three plasmids, only 8% showed high plasmid gene expression. Flow cytometric analysis of intracellular rAAV9 confirmed these results by showing that only ~ 3% of cells contained assembled rAAV9 capsids. Titre analysis by qPCR of the supernatant and lysate of the producer cells indicated an average culture performance of 10 13 vg/L. Analysis of the single-cell transcriptomic data showed that a significant proportion of cells that had high plasmid gene expression were in the S-phase. Trajectory inference highlighted that genes involved in the G2-M phase transition, immune response, and protein unfolding were differentially expressed at the branch point between high and low plasmid expression. This study reveals a significant bottleneck in the transient transfection-based production of rAAV. With less than 5% of cells producing rAAV, significant improvements in titres can be achieved if this fraction can be increased. Moreover, regulation of the cell-cycle, inhibition of the immune response, and alleviating protein misfolding all potentially offer the key to enabling these life changing treatments to reach a wider audience.
Full text 149,031 characters · extracted from preprint-html · click to expand
Heterogeneity in an Adeno-Associated Virus Transfection-Based Production Process Limits the Production Efficiency | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Heterogeneity in an Adeno-Associated Virus Transfection-Based Production Process Limits the Production Efficiency Brian Ladd, Sofia Tunmats, Olalekan Daramola, Torbjörn Gräslund, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6220373/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 04 Nov, 2025 Read the published version in Scientific Reports → Version 1 posted 13 You are reading this latest preprint version Abstract The popularity of rAAV vectors in gene therapy are placing a burden on current production systems. To improve the accessibility of these life changing treatments, increases in production yields and a reduction in the cost-of-goods are needed. Transient transfection is the most common way to introduce rAAV-encoding plasmids to producer cells but it suffers from significant drawbacks such as low and inconsistent yields as well as high cost due to its need for plasmid DNA. This study aims to address the low yield of transient transfection-based rAAV production through advanced methods in process characterization. Adherent and suspension cultures of a HEK293T cell line were triple-transfected for rAAV9 production using polyethylenimine (PEI). Samples were taken at various times post-transfection for analysis with bulk and single-cell transcriptomics. It was revealed that 46% of the cells lacked transcripts of genes from at least one plasmid, indicating that a significant proportion of the cells did not have the genes necessary for rAAV9 production. Among the remaining 54% of the cells expressing genes from all three plasmids, only 8% showed high plasmid gene expression. Flow cytometric analysis of intracellular rAAV9 confirmed these results by showing that only ~ 3% of cells contained assembled rAAV9 capsids. Titre analysis by qPCR of the supernatant and lysate of the producer cells indicated an average culture performance of 10 13 vg/L. Analysis of the single-cell transcriptomic data showed that a significant proportion of cells that had high plasmid gene expression were in the S-phase. Trajectory inference highlighted that genes involved in the G2-M phase transition, immune response, and protein unfolding were differentially expressed at the branch point between high and low plasmid expression. This study reveals a significant bottleneck in the transient transfection-based production of rAAV. With less than 5% of cells producing rAAV, significant improvements in titres can be achieved if this fraction can be increased. Moreover, regulation of the cell-cycle, inhibition of the immune response, and alleviating protein misfolding all potentially offer the key to enabling these life changing treatments to reach a wider audience. Biological sciences/Biotechnology/Gene therapy Biological sciences/Biotechnology/Sequencing/Rna sequencing AAV transient transfection transfection heterogeneity single-cell transcriptomics HEK293 cells Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 1 Introduction Since the emergence of gene therapy products, recombinant adeno-associated virus (rAAV) has become one of the leading viral vectors. This can be attributed to their unique properties of non-pathogenicity, low immunogenicity, episomal transgene expression, long-term persistency and broad tissue tropism ( 1 – 5 ). Presently, there are six FDA-approved rAAV-based gene therapies, and over 200 more are under development ( 6 , 7 ). By 2025, FDA predicts to approve 10 to 20 cell and gene therapies per year ( 3 ). Despite the prevalence of rAAV-based gene therapy, its administration has been limited, partly due to high treatment costs ( 2 , 8 ). The high prices can be attributed to high dosage requirements ( 2 ), typically in the scale of 10 12 to 10 15 vector genomes per systemic dose ( 6 ). This requirement, combined with low manufacturing productivities and high production costs ( 8 ), results in costs per dose in excess of 1 million dollars ( 9 ). The increased adoption of rAAVs places a burden on current manufacturing platforms, and achieving higher productivity is necessary to satisfy the growing clinical and commercial needs. The most common platform for rAAV manufacture is transient expression in human embryonic kidney 293 (HEK293) cells after triple-plasmid transient transfection ( 10 ). Compared to other methods, transient expression after plasmid transfection is more versatile and forgoes helper viruses, eliminating the risks associated with their use. However, this method faces difficulties during scale-up and is costly due to high plasmid consumption and lack of consistency ( 7 , 11 , 12 ). Even with a surplus of plasmids, only 30–80% of cells are successfully transfected ( 3 ). Additionally, plasmid imbalances arising from this method may contribute to inefficient packaging and inconsistent empty-to-full capsid ratios ( 2 ). Hence, titers are generally lower compared to other methods ( 11 ), reaching up to 10 14 vector genomes per liter ( 13 , 14 ), requiring several litres of culture for one single dose. Extensive efforts have been made on process development for efficient transfection and transient expression ( 3 , 15 ). However, the gaps in our biological understanding of rAAV production need to be addressed to overcome the limitations that are hindering the process at a cellular level. Assembly of rAAVs rely heavily on the host cell machinery, a rational approach to closing the gaps in our knowledge concerning alterations of the host-cell metabolism during AAV production is through the characterization of the host cell transcriptome. Indeed, previous studies have performed bulk RNA sequencing on rAAV production using either HEK293 cells or Sf9 insect cells and identified characteristic features of the producer cells. Virgolini et al. studied the transcriptional changes induced by rAAV production using the Sf9 baculovirus expression vector system. They found an enrichment of transcripts involved in the cell cycle, cell growth, protein folding and amino acid metabolism ( 16 ). Wang et al. compared the transcriptomes of viral-producing and non-producing cells from two different HEK293 cell lines. They found that rAAV production led to an upregulation of innate immune response signalling pathways and stress response pathways of the host cells, as well as late-phase downregulation of fatty acid metabolism and neutral amino acid transport ( 15 ). Chung et al. performed a kinetic study of the transcriptome after transfection at high cell density and at manufacturing scale and found an increased antiviral and inflammatory response, suggesting that cells respond to rAAV production as if it was an infection ( 17 ). Lu et al. studied the cellular response of HEK293 cells towards transient transfection using both transcriptomics and proteomics and found an upregulation in immune response, unfolded protein response and p53 signaling ( 18 ). Fundamentally, bulk RNA sequencing (RNA seq) studies the average gene expression in a cell population, which conceals the intrinsic cell-to-cell variability that persists even within a homogeneous population ( 19 – 21 ). Single-cell RNA sequencing (scRNA seq) allows this heterogeneity to be revealed by examining the gene expression at a higher resolution in individual cells ( 22 , 23 ). As the impact of cellular heterogeneity becomes more recognized and as advancements improve the accuracy and cost-effectiveness of this technology ( 24 ), more studies have adopted scRNA seq for exploring the gene expression of cells in various contexts. Most of these studies are within immunology and oncology, revealing tumor heterogeneity and distinguishing immune cell types in clonally distinct cells ( 20 ). However, one study has applied scRNA seq to examine rAAV production using the insect-cell baculovirus expression vector system in Sf9 insect cells ( 25 ). This study found a progressive increase in heterogeneity associated with the expression of viral and viral vector genes, as well as alterations in protein folding and translation, metabolic processes and stress response due to infection. This study presents the first application of scRNA seq to adherent and suspension cultures of HEK293 cells during transient transfection-based rAAV production, which uncovered cell-to-cell variability and potentially its underlying causes. Distinct groups were identified among the cells producing rAAVs and the trajectory of their development was examined. Using this approach, a significant limitation in production was identified and strategies for supporting rAAV production are proposed. 2 Materials and Methods 2.1 Cell culture HEK293T/17 ATCC-CRL-11268 cells (see Ladd et al. ( 26 )) were grown in GIBCO™ FreeStyle™ HEK293 Expression medium (Thermo Fisher Scientific, Waltham, USA) as either adherent cultures in Corning® CellBIND® static tissue culture T-25 flasks (Corning Inc., New York, USA) or suspension cultures in 125 mL shake flasks. All cultures were maintained in a Minitron incubator shaker set to 37°C and 5% CO2. 2.2 rAAV production Prior to transfection, the adherent and suspension cultures reached cell densities of 2 × 10 6 cells mL − 1 . Equal concentrations of the transfection agent PEIpro® (Polyplus, France) diluted in the cell culture medium and a DNA mixture containing a 2:1:1 mass ratio of pHelper, pRepCap (rAAV9), and green fluorescent protein (GFP) plasmids were combined. The PEIpro-DNA mixture was incubated for 7 minutes and then added to the cultures. The adherent and suspension cultures were maintained for up to 48 hours post transfection (hpt). Samples were collected at 0, 12, 24 and 48 hpt and cryopreserved in 7.5% dimethylsulfoxide (DMSO). 2.3 Analytics 2.3.1 Cell density and viability Cell density and viability of the samples were determined with the NORMA XS automated cell counter and cell viability analyzer (IPRASENSE, Clapiers, France), which is based on holographic imaging. 2.3.2 GFP expression Samples taken from the cultures were centrifuged at 180 g for 5 min and resuspended in phosphate-buffered saline (PBS). The resuspended cells were analyzed using the Guava ® easyCyte™ benchtop flow cytometer (Cytek ® Biosciences, Fremont, USA). The forward and side scattering were used to gate single cells in the InCyte™ software (Cytek ® Biosciences, Fremont, USA). GFP fluorescence was measured with a 488 nm excitation laser and a combination of a 512/18 nm and a 575/25 nm detector. 2.3.3 rAAV genome titers rAAV genome titer quantification was performed using the AAVpro ® Titration Kit (for Real Time PCR) Ver.2 (cat#6233) (Takara Bio, Kusatsu, Japan). The samples were prepared following the associated protocol and using the provided kit components. The adherent samples were first washed with Ca 2+ - and Mg 2+ -free PBS, then incubated with TrypLE (Thermo Scientific, Waltham, USA) at 37°C for 5 min. Supernatant samples were obtained by centrifuging at 1000 g for 5 min. Lysate samples were prepared through 3x freeze thaw cycles in a lysis buffer at -80°C and centrifuged at 1000 g for 5 min. The lysis buffer containing 150 mM NaCl, 50 mM Tris-HCl and 2 mM MgCl 2 was adjusted to a pH of 8.5 using NaOH. The lysate and supernatant samples were separately treated with DNAse for 1 h at 37°C to digest free DNA and subsequently heated to 95°C for 30 min to inactivate the DNAse. The lysate and supernatant samples were treated with lysis buffer for 20 min to break down the rAAV9 capsids. Finally, the samples were diluted 50-fold to use as a template for real-time qPCR. TB Green Premix Ex Taq II and the AAV2 ITR Forward and Reverse Titer Primers were combined with either the lysis template, supernatant template or the positive control provided in the kit prior to executing real-time qPCR analysis using the Bio-Rad CFX96 qPCR System (Bio-Rad, Hercules, USA). The program included: an initial denaturation at 95°C for 10 min, a 2-step PCR comprising 35 cycles at 95°C for 5 sec and at 60°C for 30 sec for FAM fluorescence detection and a melt curve analysis. The standard curve was generated through a serial dilution of the positive control to facilitate quantification. 2.4 Intracellular rAAV assembled capsid distribution The distribution of rAAV particles containing correctly assembled capsids in the suspension sample at 48 hpt was investigated using the biotinylated affinity ligand CaptureSelect™ Biotin Anti-AAV9 conjugate (Thermo Fisher Scientific, Waltham, USA) which specifically binds to assembled capsids. At 48 hpt, cells were first washed with PBS, then fixated using 4% paraformaldehyde in PBS for 15 min. The fixated cells washed with PBS and centrifuged at 400 g for 5 min three times to remove the fixative. The fixated cells were subsequently permeabilized with 0.1% Triton-X in PBST (PBS with 0.1% Tween 20) for 7 min and centrifuged at 400 g for 5 min. The permeabilized cells were incubated at room temperature with 0.1 ng/µL of the biotin anti-AAV9 conjugate for 1 hr. This was followed by incubation at room temperature with 0.5 ng/µL Streptavidin, R-Phycoerythrin Conjugate (Thermo Fisher Scientific, Waltham, USA) for 1 h. After incubation, the cells washed with PBST, centrifuged at 400 g for 5 min and then resuspended in PBST. The samples were analyzed with the Guava ® easyCyte™ benchtop flow cytometer (Cytek ® Biosciences, Fremont, USA) using a 532 nm excitation laser and a 575/25 nm detector. The gating was performed in the InCyte™ software (Cytek ® Biosciences, Fremont, USA) using the forward and side scattering. Untransfected cells were used as a negative control. 2.5 Single-cell RNA sequencing 2.5.1 Sample preparation for sequencing The samples for scRNA seq were thawed in a water bath at 37°C prior to multiplexing with the Single cell Multiplexing Kit (Cat. No. 633781) (BD Biosciences, Franklin Lakes, USA) according to the manufacturer’s instructions (Doc ID: 210970). The isolation of single-cells, subsequent capture of the RNA, and then cDNA synthesis was performed with the BD Rhapsody Express (BD Biosciences, Franklin Lakes, USA) according to the manufacturer’s protocol (Doc ID: 210967) using one cartridge. A target of 5 000 cells per sample was used, which resulted in a total of 20 000 cells per cartridge. The Whole Transcriptome Analysis (WTA) Amplification Kit (Cat num: 633801) (BD Biosciences, Franklin Lakes, USA) was used to prepare the libraries from the cDNA containing BD Rhapsody beads according to the manufacturer’s protocol (Doc ID: 23-21712-00). The library quality and fragment length were checked with capillary gel electrophoresis on an Agilent 2100 Bioanalyzer (Agilent Technologies, Santa Clara, USA). The High Sensitivity DNA kit (Cat. No. 5067 − 4626) (Agilent Technologies, Santa Clara, USA) was used according to the manufacturer’s protocol. Paired-end mRNA sequencing and adapter demultiplexing using the Illumina NovaSeq 6000 platform (Illumina, Inc., USA) was performed by a third party (Novogene Co., LTD, China). 2.5.2 Reference transcriptome for alignment Whole plasmid sequencing of the plasmids was performed by a third party (Eurofins, Germany) through real-time long-read sequencing using Oxford Nanopore technology. The sequences for the human adenovirus 5 (Ad5) genome (AC_000008.1) and the simian virus 40 large T antigen (SV40 LT) antigen gene (AY531219.1) were acquired from the National Center to Biotechnology Information (NCBI) ( 27 ). The plasmid sequences, Ad5 genome and SV40 LT gene were converted into transcripts using the program created by Pertea et al. called GFFRead ( 28 ). This program used the obtained FASTA and general feature formats (GFF) files to convert the genes into transcripts. The transcripts for the plasmid sequences, Ad5 genome and SV40 LT gene combined with the human reference transcriptome from the GRCh38.p14 assembly (GCF_000001405.40) in NCBI were used as references for the alignment. 2.5.3 Sequence alignment Sequence alignment to the references was performed in the cloud-based bioinformatics platform Seven Bridges (Seven Bridges Genomics Inc., USA). From this platform, the BD Rhapsody™ Sequence Analysis Pipeline version 2.0 utilizing the STAR alignment algorithm ( 29 ). The output FASTQ files obtained from the sequencing and the reads are filtered by quality, demultiplexed and aligned to the references. Unique molecular identifiers (UMI) were quantified per gene per cell. Finally, UMI correction and cell label filtering were performed to obtain the file containing the molecule counts per cell, the binary alignment map (BAM) and the sequencing metrics. 2.6 Data analysis 2.6.1 Processing and clustering The output file from the alignment was processed using the package Seurat version 5.1.0 ( 30 ), in R version 4.3.1. Multiplets, undetermined cells, low quality cells (> 15% mitochondrial features and < 3000 gene count) were removed from the dataset. The dataset was split into two subsets of either adherent or suspension cells. The counts within these subsets were log normalized and features were linearly scaled using the functions NormalizeData() and ScaleData() with default parameters. To mitigate the effects of cell cycle heterogeneity which may otherwise heavily influence the clustering, the cell cycle was regressed out based on assigned G2M and S-phase scores. PCA dimensionality reduction was performed with 50 principal components. An equivalent number of dimensions was used to construct the nearest-neighbor graph with the function findNeighbors(). The clustering was performed with the function findClusters(). The resolution for the clustering was determined using the framework by Patterson-Cross et al. called chooseR ( 31 ), which chooses the most robust parameter through bootstrapped iterative clustering. A significance analysis of the number of clusters was performed using the package scSHC version 0.1.0 ( 32 ), confirming that the number of principal components and the resolution were optimal for the subsets. 2.6.2 Module scores The module scores for the plasmid gene expression were calculated using the AddModuleScore() function in Seurat version 5.1.0 ( 30 ) with default parameters. 2.6.3 Trajectory Inference Trajectory inference was performed on the adherent and suspension subsets separately with Slingshot version 2.10.0 to follow the progression of cells between different cellular states over an underlying temporal variable or pseudotime ( 33 ). Determination of the starting cluster was guided by the plasmid expression. A module score was calculated for the combined expression of all three plasmid genes, and the cluster with the lowest expression was defined as the starting cluster. Using default parameters in Slingshot, lineages were determined for the remaining clusters. Each lineage branching from the starting cluster represents cells transitioning into different states. From these lineages, principal curves were fitted through the data. Each principal curve is a smooth non-linear representation of the lineage. The cell-level weights and pseudotimes were individually extracted for the lineages for further analysis. 2.6.4 Trajectory-based differential expression and overrepresentation analysis Trajectory-based differential expression analysis was individually performed on the lineages using the packages tradeSeq version 1.16.0 ( 34 ) and SingleCellExperiment version 1.24.0 ( 35 ). A negative binomial general additive model (NB-GAM) was used to model gene expression of each gene as smooth nonlinear functions of pseudotime. The optimal number of knots, at which the smooth nonlinear functions are joined together, was determined to be 6 using the function evaluateK(). The cell-level weights and pseudotimes were used as input values, while keeping all other parameters default. To determine the genes that are differentially expressed between the starting and end points of the lineage, the function startVsEndTest() was individually used on the gene NB-GAMs for each lineage. Genes with a log2 fold change < 2 and a p-value < 0.05 were clustered based on similarities in expression pattern over 100 points along the trajectory using the function clusterExpressionPatterns(). The function RSEC() from the package clusterExperiment version 2.22.0 ( 36 ) was then used to consolidate the clusters into a final consensus cluster. The merging cut-off was set to 0.9 while keeping all other parameters default. The final consensus clusters were ordered based on gene expression over pseudotime. Overrepresentation analysis on each of the final consensus clusters was performed using the packages clusterProfiler version 4.10.1 ( 37 ) and enrichR version 3.2 ( 38 ) to determine the Gene Ontology (GO) biological processes associated with each final consensus cluster. An adjusted p-value threshold of < 0.0001 was applied to show only the most significant terms. 2.7 Comparison of single-cell and bulk plasmid gene expression 2.7.1 Sample preparation for bulk RNA sequencing A second experiment was conducted to produce samples for bulk RNA seq using the same methods described in sections 2.1 , 2.2 and 2.3 . Four replicates were prepared and sampled at 0, 12, 24 and 48 hpt. Samples were cryopreserved in 7.5% DMSO prior to sequencing. Library preparation and paired-end mRNA sequencing was performed on the cryopreserved samples by an external party (Novogene Co., LTD, China) using the Illumina NovaSeq 6000 platform. The external party performed quality control and data filtering by removing reads containing adapters, reads with > 10% indeterminable (N) bases, and reads with low quality (Q score ≤ 5) bases comprising > 50% of the total bases. 2.7.3 Alignment and quality control Quality control of the acquired bulk RNA sequencing data was performed prior to data processing using a tool created by Andrews et al. called FastQC ( 39 ) combined with the tool made by Ewels et al. called MultiQC ( 40 ). Alignment and quantification was performed using a program developed by Bray et al. called Kallisto ( 41 ). The reference index used for pseudoalignment and quantification was created in Kallisto with the same reference transcripts and transcriptomes used in 2.5.2. The output file of Kallisto containing the estimated counts for each transcript was used for data analysis. Quality control of the pseudoalignment was performed using MultiQC ( 40 ). 2.7.4 Data processing and analysis The package DESeq2 ( 42 ) was employed to normalize the raw counts of the bulk RNA seq experiment using a negative binomial model. For comparison with the scRNA seq experiment, the raw counts of the plasmid genes from the bulk RNA seq experiment were normalized to their respective values at 48 hpt. These values were then compared to pseudo bulked raw counts from single-cell samples. Pseudo bulking was achieved using the function AggregateExpression from the package Seurat version 5.1.0 ( 30 ). The raw counts from all the cells were summed for each gene and the summed counts were normalized to their respective values at 48 hpt. 3 Results 3.1 Transient rAAV production through transient triple-plasmid transfection Two HEK293 cell clones, growing either adherent or in suspension, were transfected with three plasmids for the production of rAAV9. Analysis of the cell lysates for rAAV titers revealed broadly similar values for the adherent and suspension cultures, around 2000 vg/cell, shown in Fig. 1 . However, in the supernatants, the suspension cultures produced slightly more with 10 8 vg/mL compared to the 10 7 vg/mL of the adherent cells. There were slight variations among replicates, particularly for R4 and R5 in comparison with the other replicates. 3.2 Single-cell and bulk data quality control and processing Next, the transcriptomes of the cells were determined, both for single cells (R1) and bulk (R2-R5). The quality filtering, alignment, UMI correction and cell label filtering for the scRNA seq samples resulted in an average of 3.8× 10 3 reads per cell for all four samples, adherent growth and suspension growth at 12 and 48 hpt, as shown in Table S1 . The resulting number of cells after filtering is shown in Table S1 . The bulk RNA seq data had an average of 2.6 × 10 7 normalized counts per sample after quality filtering, alignment and normalization, given in Table S2. 3.3 Correlation of single-cell plasmid gene expression to bulk To confirm the reproducibility of the plasmid gene expression observed in the single-cell RNA seq experiment, the single-cell raw counts were pseudo bulked then compared to the bulk RNA seq counts. The raw counts from each plasmid gene from 12 hpt of both experiments were normalized to the corresponding gene expression at 48 hpt, shown in Figure S1 . The relationship of expression of all of the plasmid genes between the bulk and scRNA seq experiments is shown in Fig. 2 . Pearson correlation coefficients (r) were calculated for the three plasmids separately. The genes of the pRepCap plasmid had a high correlation between the bulk and scRNA seq experiments. The genes of the pHelper and pGOI plasmids had, however, a less convincing correlation. This was mostly driven by the few genes in the pRepCap plasmids and by a few outliers in the pGOI plasmid, particularly for the adherent samples. An overall r value of > 0.7 was found for all plasmids for both adherent and suspension samples, demonstrating a good overall correlation of plasmid expression between the two experiments. 3.4 Heterogeneity observed in the plasmid gene expression Individual clustering of the two subsets (adherent and suspension) resulted in four distinct clusters in both cases, seen in Figure S2. Cluster 3 for both subsets was found to have high mitochondrial gene expression, potentially indicating dead or dying cells, and was therefore excluded from the rest of the analysis. A module score was calculated using the plasmid genes as a signature. Only a small fraction of cells was deemed to have high plasmid gene expression, while most cells had low or no plasmid gene expression, as shown in Figure S3. The highest plasmid gene expression was observed in cluster 1 for both subsets, shown in Fig. 3 . Clusters 0 and 3 had comparably low plasmid gene expression for both subsets. The plasmid gene expression of cluster 2 differed between the subsets. Some cells were found to have no expression of the plasmid genes, while other cells expressed all plasmid genes, as seen in Figure S4. The cells with no plasmid gene expression were mainly found in cluster 0, while cells expressing all plasmid genes were ubiquitous in all clusters. Thus, it was concluded that cluster 1 contained cells with high plasmid gene expression and cluster 2 contained cells with low plasmid gene expression. Cluster 0 had low or no plasmid gene expression, likely comprising of “progenitor cells” to the other clusters. A large percentage of the cells (46%) were missing gene expression of at least one plasmid, as shown in Fig. 4 . The percentage of cells missing plasmid gene expression decreased over time and was generally lower in samples taken at 48 hpt, especially for the suspension subset. This indicated that not all cells have started expressing genes to a detectable level from all plasmids at 12 hpt. A marginal increase over time was also observed for cells with high plasmid gene expression for both subsets. The adherent and suspension samples had comparable levels of high plasmid gene expression, 5% and 8%, respectively. To further investigate this observed heterogeneity in plasmid gene expression, intracellular rAAV capsids were stained using a nanobody specific to assembled rAAV9 capsids. This revealed that only a small percentage of cells (2.5%) from the 48 hpt suspension subset had intracellular assembled capsids to any large extent, shown in Fig. 5. 3.5 Different cellular trajectories determine the plasmid gene expression Trajectory inference was performed on the subsets to reveal the progression of cells toward different states. Using cluster 0 as the starting cluster, three lineages (α, β, γ) were identified, Fig. 6 . The distance traveled from the beginning of the lineage is denoted as pseudotime. These lineages show that the progenitor cells, cluster 0, could either progress into cells with high plasmid gene expression, cluster 1, following the α lineage or into cells with low plasmid gene expression, cluster 2, following the β lineage. The γ lineage leading to cluster 3 was excluded from the analysis. Differential expression and overrepresentation analysis on the lineages were performed to identify factors that determined cellular fate. For each of the lineages, genes were grouped based on their expression over pseudotime, as shown in Figs. 7 and S5 for suspension and adherent cells. The number of groups identified varied between 9 to 16 for the lineages of the adherent and suspension subsets. Some of the genes could not be merged to any of the groups and thus formed their own grouping of unmatched genes (denoted − 1). Notably, some groups had cycles of gene expression over pseudotime. Significant GO terms were determined for the gene groups and ordered based on their expression over pseudotime, shown in Figs. 8 and S6. Discussion Transient transfection of plasmids to producer cells suffers from many limitations ( 13 ), some of which may be caused by cellular heterogeneity ( 43 ). Heterogeneity is intrinsic to most cellular systems and can exist at different levels, from genomic to proteomic, and the extent of heterogeneity between these levels may vary ( 20 , 21 ). A recent study investigated the cell-to-cell variability in a limited population of non-transfected adherent HEK293FT cells using scRNA seq and found only low levels of biological variance, which were primarily due to that the cells were at different stages in the cell cycle ( 44 ). However, after viral infection and propagation, high heterogeneity of gene expression has been observed ( 45 , 46 ). Single-cell transcriptomics is a powerful tool for revealing the extent and cause of heterogeneity during transient transfection-based rAAV production ( 43 ). Using single-cell transcriptomics, we were able to identify subpopulations of cells exhibiting high or low rAAV plasmid gene expression and determine the characteristic features for these subpopulations. Samples taken at 12 and 48 hpt samples were used for scRNA seq to reflect the early and later stages of transfection, as viral replication begins at 12 hpt while the majority of virions are produced at 24 hpt ( 47 , 48 ). The normalized expressions of plasmid genes in pseudo bulked single-cell samples and bulk samples generally followed similar dynamics. However, genes of the pGOI plasmid tended to be higher at 12 hpt in single-cell samples than in bulk. Correlation analysis between single-cell and bulk plasmid gene expression resulted in a good correlation for the genes of the pHelper plasmid in both adherent and suspension cell samples. Only three genes were available for comparison for the gene of the pRepCap plasmid, indicating that the observed high negative correlation may lack substantial significance. On the other hand, a low r value for the genes of the pGOI plasmid was obtained, in particular for the samples of adherent cells. A more detailed analysis of the data revealed a significant enhancement of the correlation of genes from the pGOI plasmid for both adherent and suspension cell samples through the omittance of the GOI ORI encoding sequence. In general, a high correlation was found between bulk and single-cell plasmid gene expression. Importantly, the mRNA expression may not reflect the corresponding protein expression for every gene in a cell ( 49 , 50 ). However, it can give indications of changes in cellular response to different stimuli as well as changes in expression level of individual proteins over time. Clustering of the scRNA seq data was found to be greatly influenced by the plasmid gene expression, indicating a separation of subpopulations of cells with high or low plasmid gene expression. Both the adherent and suspension subsets acquired a positively skewed distribution of the plasmid gene expression module score, indicating that high plasmid gene expression is further limited to a small subpopulation of cells. AAV production necessitates the successful transfection of all three plasmids, as it would otherwise lead to low titers or incomplete viral packaging ( 12 ). However, the efficiency of transfection typically suffers from high variability and lack of scalability ( 51 ). scRNA seq revealed differences in plasmid gene expression between individual cells originating from the same sample and showed that a large proportion of cells lacked expression of at least one plasmid. This proportion was predominantly influenced by phenotype and marginally influenced by time post-transfection, with suspension cells and later time points exhibiting a larger proportion containing all three plasmids. The proportion of cells exhibiting high expression of plasmid genes were 5 and 8% for the adherent and suspension samples, respectively. High plasmid gene expression was only observed for the cells expressing all three plasmids. However, despite the higher proportion of cells expressing all three plasmids in suspension, no large difference in the percentage of cells with high expression was observed between adherent and suspension cell samples. An equivalently low proportion of cells containing intracellular assembled capsids was observed, potentially linked to the low percentage of cells expressing high plasmid gene expression. Dash et al. had similar findings with only 5–10% of transiently-transfected HEK293SF cells expressing assembled rAAV capsids, even with high transfection efficiencies and vector yields. These findings, combined with the fact that only a small fraction of the suspension cells (8%) and adherent cells (5%) at 48 hpt had high expression of plasmid genes (Fig. 4 ), indicates that not all the cells were successfully producing rAAVs. This shows that the heterogeneity in plasmid gene expression and the low capsid assembly are bottlenecks that needs to be addressed to improve titers. The clear separation of cells with different plasmid gene expressions during clustering indicated that distinct progression towards different states of plasmid gene expression could be discerned from trajectory inference. Three lineages were identified, all originating from cluster 0 which contained the largest number of cells and displayed low plasmid gene expression. One lineage, lineage β, led to the low-to-moderate plasmid gene expression in cluster 2. Another lineage, lineage α, transitioned into high plasmid gene expression. Genes that were differentially expressed along the lineages were identified and grouped into similar expression patterns, and significant GO terms associated with the groups of genes were determined. The adherent lineage β leading to low plasmid gene expression started with terms associated with the G1 phase of the cell cycle ( 52 – 55 ). Other significant GO terms include canonical Wnt signaling pathway and the regulation of the β-catenin import into the nucleus. The canonical Wnt signaling pathway is responsible for G1- to S-phase progression ( 56 ). The accumulation of β-catenin in the cytoplasm occurs during the S-phase. The cells in the adherent lineage β of adherent cells began with an upregulation of genes from the canonical Wnt signaling followed by an upregulation of genes regulating import of the β-catenin into the nucleus, which was then downregulated at the end. This behavior suggests the progression of the cells from G1 to S to G2/M phase. GO terms associated with the regulation of the immune response and response to incorrect protein folding were upregulated at the end of the lineage. In the lineage α of adherent cells leading to high plasmid gene expression, GO terms relating to the M and G1/S phases of the cell cycle were cyclically expressed in the lineage, and underwent two cycles of upregulation and downregulation. Positive regulation of transmembrane receptor protein serine/threonine kinase signaling pathway, which is highly associated with the M phase ( 57 ), was upregulated in the beginning of the lineage, followed by downregulation. Lineage α displayed a high upregulation of ion transmembrane transport at the end, specifically for potassium, which has been linked to G1/S transition ( 58 ). These findings show the progression of cells from M- to G1- and finally to the S-phase. Similar to the lineage β of adherent cells, GO terms related to the immune response were also upregulated at the end of the lineage. The lineage β of suspension cells leading to low plasmid gene expression showed some cyclical upregulation and downregulation, but most groups were either upregulated at the beginning or the end of the lineage. M- and G1-phase related GO terms such as mitosis, generation of precursor metabolites and cell size regulation were cyclically expressed, and both were downregulated at the end of the lineage. The negative regulation of the cell cycle observed at the end of the lineage could potentially indicate an arrest in the M-phase. Reminiscent of the lineage β of adherent cells, incorrect protein folding response was also observed for the suspension lineage β. The lineages leading to low plasmid gene expression were the only ones displaying incorrect protein folding, suggesting that this could be a potential reason for the low plasmid gene levels. Other omics studies also confirm the association of protein folding with rAAV production ( 15 – 18 , 25 ). The outcome herein shows that protein folding could be optimized further to prevent limitations caused by incorrectly folded protein. In the lineage α of suspension cells leading to high plasmid gene expression, GO terms related to the M-phase such as mitosis and chromosome segregation were upregulated at early pseudotime values for lineage α. In line with this, an upregulation of GO terms related to the G1-phase such as cell growth and metabolism was observed. The gamma aminobutyric acid signaling pathway upregulated at the end of the lineage has been shown to lead to an accumulation of cells in the S-phase ( 59 ). These results indicate the transition of cells from M- to G1- to S-phase for the lineage α of suspension cells, a behavior which is also seen in the lineage α of adherent cells. Generally, an earlier immune response was activated for the β lineages leading to low plasmid gene expression. The α lineages leading to high plasmid gene expression also triggered the immune response, although at a much later time point and nearing the end of the lineage. This outcome indicates that the regulation or the delay of the immune response may be a valid strategy for enhancing rAAV productivity. The link between wild type AAV expansion in vivo and the cell cycle has been previously established. AAVs have been found to modulate the host cell cycle to facilitate viral genome replication, which mainly occurs in the S/G2 phase. ( 60 ) Franzoso et al. showed that AAV2 gene expression occurs exclusively in cells in the S/G2 phases due to the cell cycle dependent expression of the Rep gene ( 61 ). Rep 78 has been shown to activate caspase-3, which induces apoptosis during the G 1 and early S phases of the cell cycle ( 62 ), as well as to decrease the Cdc25A activity and accumulate pRb, which induces a complete arrest of the cells in the S phase ( 63 , 64 ). Additionally, many other genes involved in rAAV production induce entry into the S phase or freeze the cells in the S phase by preventing mitosis, apoptosis such as the E1 and E4 genes ( 65 , 66 ). The host cell DNA replication machinery required by DNA viruses for replication is only accessible during the S phase. To make the environment favorable for viral replication, AAVs must induce S-phase entry of infected cells, which they achieve through the virally-encoded proteins that induce S phase entry ( 67 ). The influence of the cell cycle on rAAV productivity was partly shown by Barnes et al., who showed that modifying the SKA2 and ITPRIP gene expressions, which lead to modulatory effects on the cell cycle increased rAAV genome titers ( 68 ). Another study by Tworig et al. who performed bulk RNA seq during rAAV production found the upregulation of proliferation inhibitors and downregulation of proliferation enhancers among high AAV producer cell lines, and proposed modulation of the cell cycle as a strategy to improve AAV titers ( 69 ). The findings in the present study show a strong link between high plasmid gene expression and the cell cycle. Cell cycle synchronization could be used as a strategy for improving plasmid gene expression and thus rAAV production. Conclusion This study reports the first application of single-cell transcriptomics on transient transfection-based rAAV production with adherent and suspension HEK293T cells. In our effort to characterize the process, we uncovered the substantial heterogeneity in HEK293 cells during production and identified subgroups of cells with high and low rAAV plasmid gene expression. Major differences in plasmid gene expression were observed between the cells, with nearly half of the cells missing expression of genes from at least one plasmid suggesting that successful plasmid gene expression is a significant bottleneck for the production. Additionally, we determined that a low proportion of cells produce intracellular assembled capsids. Key factors that distinguished cells with high and low plasmid gene expression were identified, with cell cycle being the primary contributor. Reducing the heterogeneity through strategies such as cell cycle synchronization could potentially improve the productivity of rAAV production based on transient transfection of HEK293 cells with producer plasmids. This study contributes to our biological understanding of the production process as well as emphasizes the need for further omics studies. Abbreviations rAAV recombinant adeno-associated virus RNA seq RNA sequencing scRNA seq single-cell RNA sequencing PEI polyethylenimine GFP Green fluorescent protein HPT hours post transfection DMSO dimethylsulfoxide PBS phosphate-buffered saline PBST phosphate-buffered saline with 0.1% Tween 20 SV40LT simian virus 40 large T antigen UMI Unique Molecular Identifier ADH adherent SUP suspension Declarations Author contributions Conceptualization, B.L., V.C. Methodology, B.L., S.T. Software, B.L., S.T. Data curation, B.L., S.T. Investigation, B.L., S.T. Formal analysis, B.L., S.T. Supervision, T.G., V.C. Funding acquisition, T.G., V.C., Visualization, B.L., S.T. Project administration, T.G., V.C. Resources, T.G., V.C. Writing—original draft preparation, S.T., B.L. Writing—review and editing, S.T., B.L., T.G., V.C. Acknowledgements This research was supported by the Competence Centre for Advanced BioProduction by Continuous Processing, AdBIOPRO, funded by Sweden’s Innovation Agency VINNOVA (diaries nr. 2022-03170), and by the Marie Skłodowska-Curie European Industrial Doctorate programme EID, STACCATO grant No 813453. The authors would also like to thank the centre GeneNova funded by VINNOVA. Conflict of interest The authors declare no conflict of interest. Data availability The datasets analysed during the current study are available in the Sequencing Read Archive (SRA) with the accession number http://www.ncbi.nlm.nih.gov/bioproject/1236707 References Robbins, P. D., Tahara, H. & Ghivizzani, S. C. Viral vectors for gene therapy. Trends Biotechnol. 16 (1), 35–40 (1998). Wang, D., Tai, P. W. & Gao, G. Adeno-associated virus vector as a platform for gene therapy delivery. Nat. Rev. Drug Discovery . 18 (5), 358–378 (2019). Sha, S. et al. Cellular pathways of recombinant adeno-associated virus production for gene therapy. Biotechnol. Adv. 49 , 107764 (2021). Li, C. & Samulski, R. J. Engineering adeno-associated virus vectors for gene therapy. Nat. Rev. Genet. 21 (4), 255–272 (2020). Daya, S. & Berns, K. I. Gene therapy using adeno-associated virus vectors. Clin. Microbiol. Rev. 21 (4), 583–593 (2008). Destro, F. et al. The state of technological advancement to address challenges in the manufacture of rAAV gene therapies. Biotechnol. Adv. :108433. (2024). Wang, J-H., Gessler, D. J., Zhan, W., Gallagher, T. L. & Gao, G. Adeno-associated virus as a delivery vector for gene therapy of human diseases. Signal. Transduct. Target. Therapy . 9 (1), 78 (2024). Smith, J., Grieger, J. & Samulski, R. J. Overcoming bottlenecks in AAV manufacturing for gene therapy. Cell. Gene Therapy Insights . 4 (8), 815–825 (2018). Kliegman, M. et al. A roadmap for affordable genetic medicines. Nature :1–8. (2024). Wang, Y. et al. Decoding cellular mechanism of recombinant adeno-associated virus (rAAV) and engineering host-cell factories toward intensified viral vector manufacturing. Biotechnol. Adv. :108322. (2024). Naso, M. F., Tomkowicz, B., Perry, W. L. III & Strohl, W. R. Adeno-associated virus (AAV) as a vector for gene therapy. BioDrugs 31 (4), 317–334 (2017). Srivastava, A., Mallela, K. M., Deorkar, N. & Brophy, G. Manufacturing challenges and rational formulation development for AAV viral vectors. J. Pharm. Sci. 110 (7), 2609–2624 (2021). Fu, Q., Polanco, A., Lee, Y. S. & Yoon, S. Critical challenges and advances in recombinant adeno-associated virus (rAAV) biomanufacturing. Biotechnol. Bioeng. 120 (9), 2601–2621 (2023). Zhao, H. et al. Creation of a high-yield AAV vector production platform in suspension cells using a design-of-experiment approach. Mol. Therapy Methods Clin. Dev. 18 , 312–320 (2020). Wang, Y., Fu, Q., Lee, Y. S., Sha, S. & Yoon, S. Transcriptomic features reveal molecular signatures associated with recombinant adeno-associated virus production in HEK293 cells. Biotechnol. Prog. 39 (4), e3346 (2023). Virgolini, N. et al. Transcriptome analysis of Sf9 insect cells during production of recombinant Adeno-associated virus. Biotechnol. J. 18 (2), 2200466 (2023). Chung, C-H. et al. Production of rAAV by plasmid transfection induces antiviral and inflammatory responses in suspension HEK293 cells. Mol. Therapy Methods Clin. Dev. 28 , 272–283 (2023). Lu, M., Lee, Z. & Hu, W. S. Multi-omics kinetic analysis of recombinant adeno‐associated virus production by plasmid transfection of HEK293 cells. Biotechnol. Prog. 40 (2), e3428 (2024). Hwang, B., Lee, J. H. & Bang, D. Single-cell RNA sequencing technologies and bioinformatics pipelines. Exp. Mol. Med. 50 (8), 1–14 (2018). Goldman, S. L. et al. The impact of heterogeneity on single-cell sequencing. Front. Genet. 10 , 8 (2019). Altschuler, S. J. & Wu, L. F. Cellular heterogeneity: do differences make a difference? Cell 141 (4), 559–563 (2010). Kolodziejczyk, A. A., Kim, J. K., Svensson, V., Marioni, J. C. & Teichmann, S. A. The technology and biology of single-cell RNA sequencing. Mol. Cell . 58 (4), 610–620 (2015). Choi, Y. H. & Kim, J. K. Dissecting cellular heterogeneity using single-cell RNA sequencing. Mol. Cells . 42 (3), 189–199 (2019). Jovic, D. et al. Single-cell RNA sequencing technologies and applications: A brief overview. Clin. translational Med. 12 (3), e694 (2022). Virgolini, N. et al. Impact of dual-baculovirus infection on the Sf9 insect cell transcriptome during rAAV production using single‐cell RNA‐seq. Biotechnol. Bioeng. 120 (9), 2588–2600 (2023). Ladd, B., Bowes, K., Lundgren, M., Gräslund, T. & Chotteau, V. Proof-of-concept of continuous transfection for adeno-associated virus production in microcarrier-based culture. Processes 10 (3), 515 (2022). Sayers, E. W. et al. Database resources of the national center for biotechnology information. Nucleic Acids Res. 50 (D1), D20–D6 (2022). Pertea, G. & Pertea, M. GFF utilities: GffRead and GffCompare. F1000Research ;9. (2020). Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29 (1), 15–21 (2013). Hao, Y. et al. Dictionary learning for integrative, multimodal and scalable single-cell analysis. Nat. Biotechnol. 42 (2), 293–304 (2024). Patterson-Cross, R. B., Levine, A. J. & Menon, V. Selecting single cell clustering parameter values using subsampling-based robustness metrics. BMC Bioinform. 22 , 1–13 (2021). Grabski, I. N., Street, K. & Irizarry, R. A. Significance analysis for clustering with single-cell RNA-sequencing data. Nat. Methods . 20 (8), 1196–1202 (2023). Street, K. et al. Slingshot: cell lineage and pseudotime inference for single-cell transcriptomics. BMC Genom. 19 , 1–16 (2018). Van den Berge, K. et al. Trajectory-based differential expression analysis for single-cell sequencing data. Nat. Commun. 11 (1), 1201 (2020). Amezquita, R. A. et al. Orchestrating single-cell analysis with Bioconductor. Nat. Methods . 17 (2), 137–145 (2020). Risso, D. et al. clusterExperiment and RSEC: A Bioconductor package and framework for clustering of single-cell and other large gene expression datasets. PLoS Comput. Biol. 14 (9), e1006378 (2018). Wu, T. et al. clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. Innov. ; 2 (3). (2021). Kuleshov, M. V. et al. Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. Nucleic Acids Res. 44 (W1), W90–W7 (2016). Andrews, S. FastQC: a quality control tool for high throughput sequence data. [ (2017). Ewels, P., Magnusson, M., Lundin, S. & Käller, M. MultiQC: summarize analysis results for multiple tools and samples in a single report. Bioinformatics 32 (19), 3047–3048 (2016). Bray, N. L., Pimentel, H., Melsted, P. & Pachter, L. Near-optimal probabilistic RNA-seq quantification. Nat. Biotechnol. 34 (5), 525–527 (2016). Love, M. I., Huber, W. & Anders, S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 15 , 1–21 (2014). Gurazada, S. G. R. et al. HEK-omics: The promise of omics to optimize HEK293 for recombinant adeno-associated virus (rAAV) gene therapy manufacturing. Biotechnol. Adv. :108506. (2024). Borsi, G. et al. Single-cell RNA sequencing reveals homogeneous transcriptome patterns and low variance in a suspension CHO-K1 and an adherent HEK293FT cell line in culture conditions. J. Biotechnol. 364 , 13–22 (2023). Jones, J. E., Le Sage, V. & Lakdawala, S. S. Viral and host heterogeneity and their effects on the viral life cycle. Nat. Rev. Microbiol. 19 (4), 272–282 (2021). Swaminath, S. & Russell, A. B. The use of single-cell RNA-seq to study heterogeneity at varying levels of virus–host interactions. PLoS Pathog. 20 (1), e1011898 (2024). Nguyen, T. N. et al. Mechanistic model for production of recombinant adeno-associated virus via triple transfection of HEK293 cells. Mol. Therapy Methods Clin. Dev. 21 , 642–655 (2021). Chahal, P. S., Schulze, E., Tran, R., Montes, J. & Kamen, A. A. Production of adeno-associated virus (AAV) serotypes by transient transfection of HEK293 cell suspension cultures for gene delivery. J. Virol. Methods . 196 , 163–173 (2014). Maier, T., Güell, M. & Serrano, L. Correlation of mRNA and protein in complex biological samples. FEBS Lett. 583 (24), 3966–3973 (2009). Edfors, F. et al. Gene-specific correlation of RNA and protein levels in human cells and tissues. Mol. Syst. Biol. 12 (10), 883 (2016). van der Loo, J. C. M. & Wright, J. F. Progress and challenges in viral vector manufacturing. Hum. Mol. Genet. 25 (R1), R42–R52 (2015). Calder, A. et al. Lengthened G1 phase indicates differentiation status in human embryonic stem cells. Stem. Cells. Dev. 22 (2), 279–295 (2013). Ter Huurne, M. & Stunnenberg, H. G. G1-phase progression in pluripotent stem cells. Cell. Mol. Life Sci. 78 (10), 4507–4519 (2021). Soufi, A. & Dalton, S. Cycling through developmental decisions: how cell cycle dynamics control pluripotency, differentiation and reprogramming. Development 143 (23), 4301–4311 (2016). Howell, A. S. & Lew, D. J. Morphogenesis and the cell cycle. Genetics 190 (1), 51–77 (2012). Niehrs, C. & Acebron, S. P. Mitotic and mitogenic Wnt signalling. EMBO J. 31 (12), 2705–2713 (2012). Yang, S-C., Huang, C-H., Chen, N-J., Chou, C-K. & Lin, C-H. Functional implication of human serine/threonine kinase, hAIK, in cell cycle progression. J. Biomed. Sci. 7 (6), 484–493 (2000). Rosendo-Pineda, M. J., Moreno, C. M. & Vaca, L. Role of ion channels during cell division. Cell. Calcium . 91 , 102258 (2020). Wang, D. D., Kriegstein, A. R. & Ben-Ari, Y. GABA regulates stem cell proliferation before nervous system formation. Epilepsy Currents . 8 (5), 137–139 (2008). Winocour, E., Callaham, M. & Huberman, E. Perturbation of the cell cycle by adeno-associated virus. Virology 167 (2), 393–399 (1988). Franzoso, F. D. et al. Cell cycle-dependent expression of adeno-associated virus 2 (AAV2) rep in coinfections with herpes simplex virus 1 (HSV-1) gives rise to a mosaic of cells replicating either AAV2 or HSV-1. J. Virol. 91 (15). 10.1128/jvi (2017). 00357 – 17. Schmidt, M., Afione, S. & Kotin, R. M. Adeno-associated virus type 2 Rep78 induces apoptosis through caspase activation independently of p53. J. Virol. 74 (20), 9441–9450 (2000). Berthet, C., Raj, K., Saudan, P. & Beard, P. How adeno-associated virus Rep78 protein arrests cells completely in S phase. Proceedings of the National Academy of Sciences. ;102(38):13634-9. (2005). Saudan, P., Vlach, J. & Beard, P. Inhibition of S-phase progression by adeno‐associated virus Rep78 protein is mediated by hypophosphorylated pRb. EMBO J. (2000). Nash, K., Chen, W., Salganik, M. & Muzyczka, N. Identification of cellular proteins that interact with the adeno-associated virus rep protein. J. Virol. 83 (1), 454–469 (2009). Ben-Israel, H. & Kleinberger, T. Adenovirus and cell cycle control. Front. Biosci. 7 , d1369–d95 (2002). Bagga, S. & Bouchard, M. J. Cell cycle regulation during viral infection. Cell. Cycle Control: Mech. Protocols :165–227. (2014). Barnes, C. R. et al. Genome-wide activation screens to increase adeno-associated virus production. Mol. Therapy-Nucleic Acids . 26 , 94–103 (2021). Tworig, J. et al. Transcriptomics-informed pharmacology identifies epigenetic and cell cycle regulators that enhance AAV production. Mol. Therapy Methods Clin. Dev. ; 32 (4). (2024). Additional Declarations No competing interests reported. Supplementary Files SupplementaryInformationsubmitted.docx Cite Share Download PDF Status: Published Journal Publication published 04 Nov, 2025 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Revision requested 05 May, 2025 Reviews received at journal 28 Apr, 2025 Reviews received at journal 04 Apr, 2025 Reviewers agreed at journal 27 Mar, 2025 Reviews received at journal 25 Mar, 2025 Reviewers agreed at journal 24 Mar, 2025 Reviewers agreed at journal 23 Mar, 2025 Reviewers agreed at journal 19 Mar, 2025 Reviewers invited by journal 19 Mar, 2025 Editor assigned by journal 19 Mar, 2025 Editor invited by journal 19 Mar, 2025 Submission checks completed at journal 18 Mar, 2025 First submitted to journal 13 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6220373","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":433953378,"identity":"5cbaa34f-a66f-480b-a1dc-fd5986d6f708","order_by":0,"name":"Brian Ladd","email":"","orcid":"","institution":"KTH Royal Institute of Technology","correspondingAuthor":false,"prefix":"","firstName":"Brian","middleName":"","lastName":"Ladd","suffix":""},{"id":433953379,"identity":"2c351c6c-9b23-4646-aba3-922c8a8c3743","order_by":1,"name":"Sofia Tunmats","email":"","orcid":"","institution":"KTH Royal Institute of Technology","correspondingAuthor":false,"prefix":"","firstName":"Sofia","middleName":"","lastName":"Tunmats","suffix":""},{"id":433953380,"identity":"9533aab0-1d62-4f4f-be25-807917f29285","order_by":2,"name":"Olalekan Daramola","email":"","orcid":"","institution":"AstraZeneca (Sweden)","correspondingAuthor":false,"prefix":"","firstName":"Olalekan","middleName":"","lastName":"Daramola","suffix":""},{"id":433953381,"identity":"7bb13188-8255-4248-a6e3-c309ebfd605b","order_by":3,"name":"Torbjörn Gräslund","email":"","orcid":"","institution":"KTH Royal Institute of Technology","correspondingAuthor":false,"prefix":"","firstName":"Torbjörn","middleName":"","lastName":"Gräslund","suffix":""},{"id":433953382,"identity":"1fd76b9e-9a6e-45e2-84e8-d4a76d075239","order_by":4,"name":"Véronique Chotteau","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAs0lEQVRIiWNgGAWjYDACCSjNz8PARrQWxoYDQFqyh5lULQZniNXCP7v5+OMPFXaJm8+cP/aA4Y8NEZbcOZbYcOBMcuK2s83sBoxtaYS1GEjkGDYcbGNO3HaemU2CseEwsVr+1Sdu7gdqYfjzn1gtDYcTN/A2A7WwHSCsReJGWuKMM8eOG884c9hMIrEtmbAW/hnJBz5U1FTL9vckPpP48MeOsBZUkECqhlEwCkbBKBgF2AEAMhQ77PXac6QAAAAASUVORK5CYII=","orcid":"","institution":"KTH Royal Institute of Technology","correspondingAuthor":true,"prefix":"","firstName":"Véronique","middleName":"","lastName":"Chotteau","suffix":""}],"badges":[],"createdAt":"2025-03-13 13:08:39","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6220373/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6220373/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-025-26261-0","type":"published","date":"2025-11-04T15:57:18+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":79386867,"identity":"f24189dc-b195-4db7-ae49-609f33c80dbd","added_by":"auto","created_at":"2025-03-27 18:15:56","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":88934,"visible":true,"origin":"","legend":"\u003cp\u003e(a) Cell specific titer in the lysate and (b) volumetric titer in the \u0026nbsp;\u0026nbsp;supernatant measured through qPCR for adherent (ADH) and suspension (SUP) \u0026nbsp;\u0026nbsp;samples at 48 hpt. R1 to R5 corresponds to five replicates of the same \u0026nbsp;\u0026nbsp;experiment. One replicate (R1) was used for the single-cell RNA sequencing \u0026nbsp;\u0026nbsp;experiment and four replicates (R2-R5) were used for the bulk RNA sequencing \u0026nbsp;\u0026nbsp;experiment.\u003c/p\u003e","description":"","filename":"Picture1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6220373/v1/a2138411e060a6bec3b8d4ca.jpg"},{"id":79387370,"identity":"743751ba-d694-445b-bb29-ee3571aadee7","added_by":"auto","created_at":"2025-03-27 18:23:56","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":94435,"visible":true,"origin":"","legend":"\u003cp\u003eCorrelation between bulk and single-cell normalized counts of all plasmid \u0026nbsp;\u0026nbsp;genes obtained during RNA sequencing expressed in the adherent (a) and \u0026nbsp;\u0026nbsp;suspension (b) cells at 12 hpt. The genes are grouped by plasmid of origin: pGOI \u0026nbsp;\u0026nbsp;(blue), pHelper (red), pRepCap (green). The Pearson correlation coefficients \u0026nbsp;\u0026nbsp;(r) are displayed for each group as well as the combined plasmid expression (black).\u003c/p\u003e","description":"","filename":"Picture2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6220373/v1/491d79e3964725340243a589.jpg"},{"id":79386873,"identity":"b66892c3-cb59-4b19-bd73-e51962bc47f4","added_by":"auto","created_at":"2025-03-27 18:15:56","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":175981,"visible":true,"origin":"","legend":"\u003cp\u003eViolin \u0026nbsp;\u0026nbsp;plot of adherent (a) and suspension (b) subsets of the single-cell RNA \u0026nbsp;\u0026nbsp;sequencing data showing the module score for plasmid gene expression of each \u0026nbsp;\u0026nbsp;cluster. UMAP plot highlighting the plasmid gene expression score in the \u0026nbsp;\u0026nbsp;adherent (c) and suspension (d) subsets.\u003c/p\u003e","description":"","filename":"Picture3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6220373/v1/a4782642d7de84eae28b51e6.jpg"},{"id":79386870,"identity":"5d9cbdb1-8c98-412a-8f0b-7ee3eea5c46d","added_by":"auto","created_at":"2025-03-27 18:15:56","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":65764,"visible":true,"origin":"","legend":"\u003cp\u003ePlasmid gene expression of the cells from the single-cell RNA sequencing \u0026nbsp;\u0026nbsp;experiment. Plasmid expression is attributed to a cell if at least 1 \u0026nbsp;\u0026nbsp;transcript of any gene from that plasmid exists in that cell. Cells \u0026nbsp;\u0026nbsp;expressing all plasmid genes (olive green) and cells with high plasmid gene \u0026nbsp;\u0026nbsp;expression (navy blue) are also shown. To the left, the plasmid gene \u0026nbsp;\u0026nbsp;expression of all samples combined is displayed.\u003c/p\u003e","description":"","filename":"Picture4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6220373/v1/8ea344ae6d91154e6400d2d9.jpg"},{"id":79386869,"identity":"44fb849f-48f3-4af9-bc0f-b31504fdd662","added_by":"auto","created_at":"2025-03-27 18:15:56","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":10273,"visible":true,"origin":"","legend":"\u003cp\u003eFlow \u0026nbsp;\u0026nbsp;cytometry histogram with suspension cells at 48 hpt stained for intracellular \u0026nbsp;\u0026nbsp;assembled rAAV capsids and filtered from debris and multiplets. The negative \u0026nbsp;\u0026nbsp;control comprises stained untransfected cells. The positive gate was set \u0026nbsp;\u0026nbsp;based on the negative control.\u003c/p\u003e","description":"","filename":"Picture5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6220373/v1/005761d9898fc98842becbd6.jpg"},{"id":79386872,"identity":"ace59b0f-9922-4302-9e08-bceb74b61544","added_by":"auto","created_at":"2025-03-27 18:15:56","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":112438,"visible":true,"origin":"","legend":"\u003cp\u003eUMAP plot with the three lineages (α, β, γ) for \u0026nbsp;\u0026nbsp;the adherent (a) and suspension (b) subsets of the single-cell RNA sequencing \u0026nbsp;\u0026nbsp;data generated from the trajectory inference. The cells are colored by \u0026nbsp;\u0026nbsp;cluster: 0 (red), 1 (green), 2 (blue), 3 (purple).\u003c/p\u003e","description":"","filename":"Picture6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6220373/v1/f16cd6f98b87a9122fa35589.jpg"},{"id":79386875,"identity":"1981d59a-60c7-45b7-92ce-b7a2f62bcf56","added_by":"auto","created_at":"2025-03-27 18:15:57","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":157633,"visible":true,"origin":"","legend":"\u003cp\u003eHeatmap plot of the relative expression levels (log2 normalized z-score) of differentially expressed genes associated with the trajectory of suspension lineages towards low (a) or high (b) plasmid gene expression. The genes are grouped based on expression pattern over pseudotime. The group labelled -1 shows genes that were not successfully grouped.\u003c/p\u003e","description":"","filename":"Picture7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6220373/v1/4bb087c2e170e1844a38ce12.jpg"},{"id":79386885,"identity":"6b9c2512-be30-4f47-b71e-3b4766bb6a32","added_by":"auto","created_at":"2025-03-27 18:15:57","extension":"jpg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":286368,"visible":true,"origin":"","legend":"\u003cp\u003eResults from the overrepresentation analysis with Gene Ontology biological processes of groups of differentially expressed genes associated with the suspension lineages towards low (a) or high (b) plasmid gene expression. The genes are grouped based on expression pattern over pseudotime.\u003c/p\u003e","description":"","filename":"Picture8.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6220373/v1/f8997527ca24f00393b9f9fe.jpg"},{"id":95564097,"identity":"415a5de3-9c3e-4bf6-a8de-feba43e79f59","added_by":"auto","created_at":"2025-11-10 16:07:49","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2127680,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6220373/v1/a0420c3c-388e-49af-b583-0643845ffcfe.pdf"},{"id":79387378,"identity":"348f5a81-8d9d-4944-9cab-3c78a65316e2","added_by":"auto","created_at":"2025-03-27 18:23:57","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":4756028,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryInformationsubmitted.docx","url":"https://assets-eu.researchsquare.com/files/rs-6220373/v1/a9f8f72345eb8b506e606e7b.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Heterogeneity in an Adeno-Associated Virus Transfection-Based Production Process Limits the Production Efficiency","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eSince the emergence of gene therapy products, recombinant adeno-associated virus (rAAV) has become one of the leading viral vectors. This can be attributed to their unique properties of non-pathogenicity, low immunogenicity, episomal transgene expression, long-term persistency and broad tissue tropism (\u003cspan additionalcitationids=\"CR2 CR3 CR4\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e). Presently, there are six FDA-approved rAAV-based gene therapies, and over 200 more are under development (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e). By 2025, FDA predicts to approve 10 to 20 cell and gene therapies per year (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). Despite the prevalence of rAAV-based gene therapy, its administration has been limited, partly due to high treatment costs (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). The high prices can be attributed to high dosage requirements (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e), typically in the scale of 10\u003csup\u003e12\u003c/sup\u003e to 10\u003csup\u003e15\u003c/sup\u003e vector genomes per systemic dose (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e). This requirement, combined with low manufacturing productivities and high production costs (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e), results in costs per dose in excess of 1\u0026nbsp;million dollars (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e). The increased adoption of rAAVs places a burden on current manufacturing platforms, and achieving higher productivity is necessary to satisfy the growing clinical and commercial needs.\u003c/p\u003e \u003cp\u003eThe most common platform for rAAV manufacture is transient expression in human embryonic kidney 293 (HEK293) cells after triple-plasmid transient transfection (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e). Compared to other methods, transient expression after plasmid transfection is more versatile and forgoes helper viruses, eliminating the risks associated with their use. However, this method faces difficulties during scale-up and is costly due to high plasmid consumption and lack of consistency (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e). Even with a surplus of plasmids, only 30\u0026ndash;80% of cells are successfully transfected (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). Additionally, plasmid imbalances arising from this method may contribute to inefficient packaging and inconsistent empty-to-full capsid ratios (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). Hence, titers are generally lower compared to other methods (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e), reaching up to 10\u003csup\u003e14\u003c/sup\u003e vector genomes per liter (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e), requiring several litres of culture for one single dose. Extensive efforts have been made on process development for efficient transfection and transient expression (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). However, the gaps in our biological understanding of rAAV production need to be addressed to overcome the limitations that are hindering the process at a cellular level.\u003c/p\u003e \u003cp\u003eAssembly of rAAVs rely heavily on the host cell machinery, a rational approach to closing the gaps in our knowledge concerning alterations of the host-cell metabolism during AAV production is through the characterization of the host cell transcriptome. Indeed, previous studies have performed bulk RNA sequencing on rAAV production using either HEK293 cells or Sf9 insect cells and identified characteristic features of the producer cells. Virgolini et al. studied the transcriptional changes induced by rAAV production using the Sf9 baculovirus expression vector system. They found an enrichment of transcripts involved in the cell cycle, cell growth, protein folding and amino acid metabolism (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e). Wang et al. compared the transcriptomes of viral-producing and non-producing cells from two different HEK293 cell lines. They found that rAAV production led to an upregulation of innate immune response signalling pathways and stress response pathways of the host cells, as well as late-phase downregulation of fatty acid metabolism and neutral amino acid transport (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). Chung et al. performed a kinetic study of the transcriptome after transfection at high cell density and at manufacturing scale and found an increased antiviral and inflammatory response, suggesting that cells respond to rAAV production as if it was an infection (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e). Lu et al. studied the cellular response of HEK293 cells towards transient transfection using both transcriptomics and proteomics and found an upregulation in immune response, unfolded protein response and p53 signaling (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eFundamentally, bulk RNA sequencing (RNA seq) studies the average gene expression in a cell population, which conceals the intrinsic cell-to-cell variability that persists even within a homogeneous population (\u003cspan additionalcitationids=\"CR20\" citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). Single-cell RNA sequencing (scRNA seq) allows this heterogeneity to be revealed by examining the gene expression at a higher resolution in individual cells (\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e). As the impact of cellular heterogeneity becomes more recognized and as advancements improve the accuracy and cost-effectiveness of this technology (\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e), more studies have adopted scRNA seq for exploring the gene expression of cells in various contexts. Most of these studies are within immunology and oncology, revealing tumor heterogeneity and distinguishing immune cell types in clonally distinct cells (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e). However, one study has applied scRNA seq to examine rAAV production using the insect-cell baculovirus expression vector system in Sf9 insect cells (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e). This study found a progressive increase in heterogeneity associated with the expression of viral and viral vector genes, as well as alterations in protein folding and translation, metabolic processes and stress response due to infection.\u003c/p\u003e \u003cp\u003eThis study presents the first application of scRNA seq to adherent and suspension cultures of HEK293 cells during transient transfection-based rAAV production, which uncovered cell-to-cell variability and potentially its underlying causes. Distinct groups were identified among the cells producing rAAVs and the trajectory of their development was examined. Using this approach, a significant limitation in production was identified and strategies for supporting rAAV production are proposed.\u003c/p\u003e"},{"header":"2 Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Cell culture\u003c/h2\u003e \u003cp\u003eHEK293T/17 ATCC-CRL-11268 cells (see Ladd et al. (\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e)) were grown in GIBCO\u0026trade; FreeStyle\u0026trade; HEK293 Expression medium (Thermo Fisher Scientific, Waltham, USA) as either adherent cultures in Corning\u0026reg; CellBIND\u0026reg; static tissue culture T-25 flasks (Corning Inc., New York, USA) or suspension cultures in 125 mL shake flasks. All cultures were maintained in a Minitron incubator shaker set to 37\u0026deg;C and 5% CO2.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 rAAV production\u003c/h2\u003e \u003cp\u003ePrior to transfection, the adherent and suspension cultures reached cell densities of 2 \u0026times; 10\u003csup\u003e6\u003c/sup\u003e cells mL\u003csup\u003e\u0026minus;\u0026thinsp;1\u003c/sup\u003e. Equal concentrations of the transfection agent PEIpro\u0026reg; (Polyplus, France) diluted in the cell culture medium and a DNA mixture containing a 2:1:1 mass ratio of pHelper, pRepCap (rAAV9), and green fluorescent protein (GFP) plasmids were combined. The PEIpro-DNA mixture was incubated for 7 minutes and then added to the cultures. The adherent and suspension cultures were maintained for up to 48 hours post transfection (hpt). Samples were collected at 0, 12, 24 and 48 hpt and cryopreserved in 7.5% dimethylsulfoxide (DMSO).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Analytics\u003c/h2\u003e \u003cdiv id=\"Sec6\" class=\"Section3\"\u003e \u003ch2\u003e2.3.1 Cell density and viability\u003c/h2\u003e \u003cp\u003eCell density and viability of the samples were determined with the NORMA XS automated cell counter and cell viability analyzer (IPRASENSE, Clapiers, France), which is based on holographic imaging.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003ch2\u003e2.3.2 GFP expression\u003c/h2\u003e \u003cp\u003eSamples taken from the cultures were centrifuged at 180 g for 5 min and resuspended in phosphate-buffered saline (PBS). The resuspended cells were analyzed using the Guava\u003csup\u003e\u0026reg;\u003c/sup\u003e easyCyte\u0026trade; benchtop flow cytometer (Cytek\u003csup\u003e\u0026reg;\u003c/sup\u003e Biosciences, Fremont, USA). The forward and side scattering were used to gate single cells in the InCyte\u0026trade; software (Cytek\u003csup\u003e\u0026reg;\u003c/sup\u003e Biosciences, Fremont, USA). GFP fluorescence was measured with a 488 nm excitation laser and a combination of a 512/18 nm and a 575/25 nm detector.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e \u003ch2\u003e2.3.3 rAAV genome titers\u003c/h2\u003e \u003cp\u003erAAV genome titer quantification was performed using the AAVpro\u003csup\u003e\u0026reg;\u003c/sup\u003e Titration Kit (for Real Time PCR) Ver.2 (cat#6233) (Takara Bio, Kusatsu, Japan). The samples were prepared following the associated protocol and using the provided kit components. The adherent samples were first washed with Ca\u003csup\u003e2+\u003c/sup\u003e- and Mg\u003csup\u003e2+\u003c/sup\u003e-free PBS, then incubated with TrypLE (Thermo Scientific, Waltham, USA) at 37\u0026deg;C for 5 min. Supernatant samples were obtained by centrifuging at 1000 g for 5 min. Lysate samples were prepared through 3x freeze thaw cycles in a lysis buffer at -80\u0026deg;C and centrifuged at 1000 g for 5 min. The lysis buffer containing 150 mM NaCl, 50 mM Tris-HCl and 2 mM MgCl\u003csub\u003e2\u003c/sub\u003e was adjusted to a pH of 8.5 using NaOH. The lysate and supernatant samples were separately treated with DNAse for 1 h at 37\u0026deg;C to digest free DNA and subsequently heated to 95\u0026deg;C for 30 min to inactivate the DNAse. The lysate and supernatant samples were treated with lysis buffer for 20 min to break down the rAAV9 capsids. Finally, the samples were diluted 50-fold to use as a template for real-time qPCR. TB Green \u003cem\u003ePremix Ex Taq\u003c/em\u003e II and the AAV2 ITR Forward and Reverse Titer Primers were combined with either the lysis template, supernatant template or the positive control provided in the kit prior to executing real-time qPCR analysis using the Bio-Rad CFX96 qPCR System (Bio-Rad, Hercules, USA). The program included: an initial denaturation at 95\u0026deg;C for 10 min, a 2-step PCR comprising 35 cycles at 95\u0026deg;C for 5 sec and at 60\u0026deg;C for 30 sec for FAM fluorescence detection and a melt curve analysis. The standard curve was generated through a serial dilution of the positive control to facilitate quantification.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Intracellular rAAV assembled capsid distribution\u003c/h2\u003e \u003cp\u003eThe distribution of rAAV particles containing correctly assembled capsids in the suspension sample at 48 hpt was investigated using the biotinylated affinity ligand CaptureSelect\u0026trade; Biotin Anti-AAV9 conjugate (Thermo Fisher Scientific, Waltham, USA) which specifically binds to assembled capsids. At 48 hpt, cells were first washed with PBS, then fixated using 4% paraformaldehyde in PBS for 15 min. The fixated cells washed with PBS and centrifuged at 400 g for 5 min three times to remove the fixative. The fixated cells were subsequently permeabilized with 0.1% Triton-X in PBST (PBS with 0.1% Tween 20) for 7 min and centrifuged at 400 g for 5 min. The permeabilized cells were incubated at room temperature with 0.1 ng/\u0026micro;L of the biotin anti-AAV9 conjugate for 1 hr. This was followed by incubation at room temperature with 0.5 ng/\u0026micro;L Streptavidin, R-Phycoerythrin Conjugate (Thermo Fisher Scientific, Waltham, USA) for 1 h. After incubation, the cells washed with PBST, centrifuged at 400 g for 5 min and then resuspended in PBST. The samples were analyzed with the Guava\u003csup\u003e\u0026reg;\u003c/sup\u003e easyCyte\u0026trade; benchtop flow cytometer (Cytek\u003csup\u003e\u0026reg;\u003c/sup\u003e Biosciences, Fremont, USA) using a 532 nm excitation laser and a 575/25 nm detector. The gating was performed in the InCyte\u0026trade; software (Cytek\u003csup\u003e\u0026reg;\u003c/sup\u003e Biosciences, Fremont, USA) using the forward and side scattering. Untransfected cells were used as a negative control.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Single-cell RNA sequencing\u003c/h2\u003e \u003cdiv id=\"Sec11\" class=\"Section3\"\u003e \u003ch2\u003e2.5.1 Sample preparation for sequencing\u003c/h2\u003e \u003cp\u003eThe samples for scRNA seq were thawed in a water bath at 37\u0026deg;C prior to multiplexing with the Single cell Multiplexing Kit (Cat. No. 633781) (BD Biosciences, Franklin Lakes, USA) according to the manufacturer\u0026rsquo;s instructions (Doc ID: 210970). The isolation of single-cells, subsequent capture of the RNA, and then cDNA synthesis was performed with the BD Rhapsody Express (BD Biosciences, Franklin Lakes, USA) according to the manufacturer\u0026rsquo;s protocol (Doc ID: 210967) using one cartridge. A target of 5 000 cells per sample was used, which resulted in a total of 20 000 cells per cartridge. The Whole Transcriptome Analysis (WTA) Amplification Kit (Cat num: 633801) (BD Biosciences, Franklin Lakes, USA) was used to prepare the libraries from the cDNA containing BD Rhapsody beads according to the manufacturer\u0026rsquo;s protocol (Doc ID: 23-21712-00). The library quality and fragment length were checked with capillary gel electrophoresis on an Agilent 2100 Bioanalyzer (Agilent Technologies, Santa Clara, USA). The High Sensitivity DNA kit (Cat. No. 5067\u0026thinsp;\u0026minus;\u0026thinsp;4626) (Agilent Technologies, Santa Clara, USA) was used according to the manufacturer\u0026rsquo;s protocol. Paired-end mRNA sequencing and adapter demultiplexing using the Illumina NovaSeq 6000 platform (Illumina, Inc., USA) was performed by a third party (Novogene Co., LTD, China).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section3\"\u003e \u003ch2\u003e2.5.2 Reference transcriptome for alignment\u003c/h2\u003e \u003cp\u003eWhole plasmid sequencing of the plasmids was performed by a third party (Eurofins, Germany) through real-time long-read sequencing using Oxford Nanopore technology. The sequences for the human adenovirus 5 (Ad5) genome (AC_000008.1) and the simian virus 40 large T antigen (SV40 LT) antigen gene (AY531219.1) were acquired from the National Center to Biotechnology Information (NCBI) (\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e). The plasmid sequences, Ad5 genome and SV40 LT gene were converted into transcripts using the program created by Pertea et al. called GFFRead (\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e). This program used the obtained FASTA and general feature formats (GFF) files to convert the genes into transcripts. The transcripts for the plasmid sequences, Ad5 genome and SV40 LT gene combined with the human reference transcriptome from the GRCh38.p14 assembly (GCF_000001405.40) in NCBI were used as references for the alignment.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section3\"\u003e \u003ch2\u003e2.5.3 Sequence alignment\u003c/h2\u003e \u003cp\u003eSequence alignment to the references was performed in the cloud-based bioinformatics platform Seven Bridges (Seven Bridges Genomics Inc., USA). From this platform, the BD Rhapsody\u0026trade; Sequence Analysis Pipeline version 2.0 utilizing the STAR alignment algorithm (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e). The output FASTQ files obtained from the sequencing and the reads are filtered by quality, demultiplexed and aligned to the references. Unique molecular identifiers (UMI) were quantified per gene per cell. Finally, UMI correction and cell label filtering were performed to obtain the file containing the molecule counts per cell, the binary alignment map (BAM) and the sequencing metrics.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e2.6 Data analysis\u003c/h2\u003e \u003cdiv id=\"Sec15\" class=\"Section3\"\u003e \u003ch2\u003e2.6.1 Processing and clustering\u003c/h2\u003e \u003cp\u003eThe output file from the alignment was processed using the package Seurat version 5.1.0 (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e), in R version 4.3.1. Multiplets, undetermined cells, low quality cells (\u0026gt;\u0026thinsp;15% mitochondrial features and \u0026lt;\u0026thinsp;3000 gene count) were removed from the dataset. The dataset was split into two subsets of either adherent or suspension cells. The counts within these subsets were log normalized and features were linearly scaled using the functions NormalizeData() and ScaleData() with default parameters. To mitigate the effects of cell cycle heterogeneity which may otherwise heavily influence the clustering, the cell cycle was regressed out based on assigned G2M and S-phase scores. PCA dimensionality reduction was performed with 50 principal components. An equivalent number of dimensions was used to construct the nearest-neighbor graph with the function findNeighbors(). The clustering was performed with the function findClusters(). The resolution for the clustering was determined using the framework by Patterson-Cross et al. called chooseR (\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e), which chooses the most robust parameter through bootstrapped iterative clustering. A significance analysis of the number of clusters was performed using the package scSHC version 0.1.0 (\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e), confirming that the number of principal components and the resolution were optimal for the subsets.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section3\"\u003e \u003ch2\u003e2.6.2 Module scores\u003c/h2\u003e \u003cp\u003eThe module scores for the plasmid gene expression were calculated using the AddModuleScore() function in Seurat version 5.1.0 (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e) with default parameters.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section3\"\u003e \u003ch2\u003e2.6.3 Trajectory Inference\u003c/h2\u003e \u003cp\u003eTrajectory inference was performed on the adherent and suspension subsets separately with Slingshot version 2.10.0 to follow the progression of cells between different cellular states over an underlying temporal variable or pseudotime (\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e). Determination of the starting cluster was guided by the plasmid expression. A module score was calculated for the combined expression of all three plasmid genes, and the cluster with the lowest expression was defined as the starting cluster. Using default parameters in Slingshot, lineages were determined for the remaining clusters. Each lineage branching from the starting cluster represents cells transitioning into different states. From these lineages, principal curves were fitted through the data. Each principal curve is a smooth non-linear representation of the lineage. The cell-level weights and pseudotimes were individually extracted for the lineages for further analysis.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section3\"\u003e \u003ch2\u003e2.6.4 Trajectory-based differential expression and overrepresentation analysis\u003c/h2\u003e \u003cp\u003eTrajectory-based differential expression analysis was individually performed on the lineages using the packages tradeSeq version 1.16.0 (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e) and SingleCellExperiment version 1.24.0 (\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e). A negative binomial general additive model (NB-GAM) was used to model gene expression of each gene as smooth nonlinear functions of pseudotime. The optimal number of knots, at which the smooth nonlinear functions are joined together, was determined to be 6 using the function evaluateK(). The cell-level weights and pseudotimes were used as input values, while keeping all other parameters default. To determine the genes that are differentially expressed between the starting and end points of the lineage, the function startVsEndTest() was individually used on the gene NB-GAMs for each lineage. Genes with a log2 fold change\u0026thinsp;\u0026lt;\u0026thinsp;2 and a p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05 were clustered based on similarities in expression pattern over 100 points along the trajectory using the function clusterExpressionPatterns(). The function RSEC() from the package clusterExperiment version 2.22.0 (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e) was then used to consolidate the clusters into a final consensus cluster. The merging cut-off was set to 0.9 while keeping all other parameters default. The final consensus clusters were ordered based on gene expression over pseudotime. Overrepresentation analysis on each of the final consensus clusters was performed using the packages clusterProfiler version 4.10.1 (\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e) and enrichR version 3.2 (\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e) to determine the Gene Ontology (GO) biological processes associated with each final consensus cluster. An adjusted p-value threshold of \u0026lt;\u0026thinsp;0.0001 was applied to show only the most significant terms.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e2.7 Comparison of single-cell and bulk plasmid gene expression\u003c/h2\u003e \u003cdiv id=\"Sec20\" class=\"Section3\"\u003e \u003ch2\u003e2.7.1 Sample preparation for bulk RNA sequencing\u003c/h2\u003e \u003cp\u003eA second experiment was conducted to produce samples for bulk RNA seq using the same methods described in sections \u003cspan refid=\"Sec3\" class=\"InternalRef\"\u003e2.1\u003c/span\u003e, \u003cspan refid=\"Sec4\" class=\"InternalRef\"\u003e2.2\u003c/span\u003e and \u003cspan refid=\"Sec5\" class=\"InternalRef\"\u003e2.3\u003c/span\u003e. Four replicates were prepared and sampled at 0, 12, 24 and 48 hpt. Samples were cryopreserved in 7.5% DMSO prior to sequencing. Library preparation and paired-end mRNA sequencing was performed on the cryopreserved samples by an external party (Novogene Co., LTD, China) using the Illumina NovaSeq 6000 platform. The external party performed quality control and data filtering by removing reads containing adapters, reads with \u0026gt;\u0026thinsp;10% indeterminable (N) bases, and reads with low quality (Q score\u0026thinsp;\u0026le;\u0026thinsp;5) bases comprising\u0026thinsp;\u0026gt;\u0026thinsp;50% of the total bases.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section3\"\u003e \u003ch2\u003e2.7.3 Alignment and quality control\u003c/h2\u003e \u003cp\u003eQuality control of the acquired bulk RNA sequencing data was performed prior to data processing using a tool created by Andrews et al. called FastQC (\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e) combined with the tool made by Ewels et al. called MultiQC (\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e). Alignment and quantification was performed using a program developed by Bray et al. called Kallisto (\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e). The reference index used for pseudoalignment and quantification was created in Kallisto with the same reference transcripts and transcriptomes used in 2.5.2. The output file of Kallisto containing the estimated counts for each transcript was used for data analysis. Quality control of the pseudoalignment was performed using MultiQC (\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section3\"\u003e \u003ch2\u003e2.7.4 Data processing and analysis\u003c/h2\u003e \u003cp\u003eThe package DESeq2 (\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e) was employed to normalize the raw counts of the bulk RNA seq experiment using a negative binomial model. For comparison with the scRNA seq experiment, the raw counts of the plasmid genes from the bulk RNA seq experiment were normalized to their respective values at 48 hpt. These values were then compared to pseudo bulked raw counts from single-cell samples. Pseudo bulking was achieved using the function AggregateExpression from the package Seurat version 5.1.0 (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e). The raw counts from all the cells were summed for each gene and the summed counts were normalized to their respective values at 48 hpt.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"3 Results","content":"\u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Transient rAAV production through transient triple-plasmid transfection\u003c/h2\u003e \u003cp\u003eTwo HEK293 cell clones, growing either adherent or in suspension, were transfected with three plasmids for the production of rAAV9. Analysis of the cell lysates for rAAV titers revealed broadly similar values for the adherent and suspension cultures, around 2000 vg/cell, shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. However, in the supernatants, the suspension cultures produced slightly more with 10\u003csup\u003e8\u003c/sup\u003e vg/mL compared to the 10\u003csup\u003e7\u003c/sup\u003e vg/mL of the adherent cells. There were slight variations among replicates, particularly for R4 and R5 in comparison with the other replicates.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec25\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Single-cell and bulk data quality control and processing\u003c/h2\u003e \u003cp\u003eNext, the transcriptomes of the cells were determined, both for single cells (R1) and bulk (R2-R5). The quality filtering, alignment, UMI correction and cell label filtering for the scRNA seq samples resulted in an average of 3.8\u0026times; 10\u003csup\u003e3\u003c/sup\u003e reads per cell for all four samples, adherent growth and suspension growth at 12 and 48 hpt, as shown in Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e. The resulting number of cells after filtering is shown in Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e. The bulk RNA seq data had an average of 2.6 \u0026times; 10\u003csup\u003e7\u003c/sup\u003e normalized counts per sample after quality filtering, alignment and normalization, given in Table S2.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec26\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Correlation of single-cell plasmid gene expression to bulk\u003c/h2\u003e \u003cp\u003eTo confirm the reproducibility of the plasmid gene expression observed in the single-cell RNA seq experiment, the single-cell raw counts were pseudo bulked then compared to the bulk RNA seq counts. The raw counts from each plasmid gene from 12 hpt of both experiments were normalized to the corresponding gene expression at 48 hpt, shown in Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e. The relationship of expression of all of the plasmid genes between the bulk and scRNA seq experiments is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. Pearson correlation coefficients (r) were calculated for the three plasmids separately. The genes of the pRepCap plasmid had a high correlation between the bulk and scRNA seq experiments. The genes of the pHelper and pGOI plasmids had, however, a less convincing correlation. This was mostly driven by the few genes in the pRepCap plasmids and by a few outliers in the pGOI plasmid, particularly for the adherent samples. An overall r value of \u0026gt;\u0026thinsp;0.7 was found for all plasmids for both adherent and suspension samples, demonstrating a good overall correlation of plasmid expression between the two experiments.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec27\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Heterogeneity observed in the plasmid gene expression\u003c/h2\u003e \u003cp\u003eIndividual clustering of the two subsets (adherent and suspension) resulted in four distinct clusters in both cases, seen in Figure S2. Cluster 3 for both subsets was found to have high mitochondrial gene expression, potentially indicating dead or dying cells, and was therefore excluded from the rest of the analysis. A module score was calculated using the plasmid genes as a signature. Only a small fraction of cells was deemed to have high plasmid gene expression, while most cells had low or no plasmid gene expression, as shown in Figure S3. The highest plasmid gene expression was observed in cluster 1 for both subsets, shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. Clusters 0 and 3 had comparably low plasmid gene expression for both subsets. The plasmid gene expression of cluster 2 differed between the subsets. Some cells were found to have no expression of the plasmid genes, while other cells expressed all plasmid genes, as seen in Figure S4. The cells with no plasmid gene expression were mainly found in cluster 0, while cells expressing all plasmid genes were ubiquitous in all clusters. Thus, it was concluded that cluster 1 contained cells with high plasmid gene expression and cluster 2 contained cells with low plasmid gene expression. Cluster 0 had low or no plasmid gene expression, likely comprising of \u0026ldquo;progenitor cells\u0026rdquo; to the other clusters.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eA large percentage of the cells (46%) were missing gene expression of at least one plasmid, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. The percentage of cells missing plasmid gene expression decreased over time and was generally lower in samples taken at 48 hpt, especially for the suspension subset. This indicated that not all cells have started expressing genes to a detectable level from all plasmids at 12 hpt. A marginal increase over time was also observed for cells with high plasmid gene expression for both subsets. The adherent and suspension samples had comparable levels of high plasmid gene expression, 5% and 8%, respectively.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo further investigate this observed heterogeneity in plasmid gene expression, intracellular rAAV capsids were stained using a nanobody specific to assembled rAAV9 capsids. This revealed that only a small percentage of cells (2.5%) from the 48 hpt suspension subset had intracellular assembled capsids to any large extent, shown in Fig.\u0026nbsp;5.\u003c/p\u003e \u003cdiv id=\"Sec28\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Different cellular trajectories determine the plasmid gene expression\u003c/h2\u003e \u003cp\u003eTrajectory inference was performed on the subsets to reveal the progression of cells toward different states. Using cluster 0 as the starting cluster, three lineages (α, β, γ) were identified, Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e6\u003c/span\u003e. The distance traveled from the beginning of the lineage is denoted as pseudotime. These lineages show that the progenitor cells, cluster 0, could either progress into cells with high plasmid gene expression, cluster 1, following the α lineage or into cells with low plasmid gene expression, cluster 2, following the β lineage. The γ lineage leading to cluster 3 was excluded from the analysis.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eDifferential expression and overrepresentation analysis on the lineages were performed to identify factors that determined cellular fate. For each of the lineages, genes were grouped based on their expression over pseudotime, as shown in Figs.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e7\u003c/span\u003e and S5 for suspension and adherent cells. The number of groups identified varied between 9 to 16 for the lineages of the adherent and suspension subsets. Some of the genes could not be merged to any of the groups and thus formed their own grouping of unmatched genes (denoted \u0026minus;\u0026thinsp;1). Notably, some groups had cycles of gene expression over pseudotime. Significant GO terms were determined for the gene groups and ordered based on their expression over pseudotime, shown in Figs.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e8\u003c/span\u003e and S6.\u003c/p\u003e "},{"header":"Discussion","content":"\u003cp\u003eTransient transfection of plasmids to producer cells suffers from many limitations (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e), some of which may be caused by cellular heterogeneity (\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e). Heterogeneity is intrinsic to most cellular systems and can exist at different levels, from genomic to proteomic, and the extent of heterogeneity between these levels may vary (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). A recent study investigated the cell-to-cell variability in a limited population of non-transfected adherent HEK293FT cells using scRNA seq and found only low levels of biological variance, which were primarily due to that the cells were at different stages in the cell cycle (\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e). However, after viral infection and propagation, high heterogeneity of gene expression has been observed (\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e, \u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e). Single-cell transcriptomics is a powerful tool for revealing the extent and cause of heterogeneity during transient transfection-based rAAV production (\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e). Using single-cell transcriptomics, we were able to identify subpopulations of cells exhibiting high or low rAAV plasmid gene expression and determine the characteristic features for these subpopulations.\u003c/p\u003e \u003cp\u003eSamples taken at 12 and 48 hpt samples were used for scRNA seq to reflect the early and later stages of transfection, as viral replication begins at 12 hpt while the majority of virions are produced at 24 hpt (\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e, \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e). The normalized expressions of plasmid genes in pseudo bulked single-cell samples and bulk samples generally followed similar dynamics. However, genes of the pGOI plasmid tended to be higher at 12 hpt in single-cell samples than in bulk. Correlation analysis between single-cell and bulk plasmid gene expression resulted in a good correlation for the genes of the pHelper plasmid in both adherent and suspension cell samples. Only three genes were available for comparison for the gene of the pRepCap plasmid, indicating that the observed high negative correlation may lack substantial significance. On the other hand, a low r value for the genes of the pGOI plasmid was obtained, in particular for the samples of adherent cells. A more detailed analysis of the data revealed a significant enhancement of the correlation of genes from the pGOI plasmid for both adherent and suspension cell samples through the omittance of the GOI ORI encoding sequence. In general, a high correlation was found between bulk and single-cell plasmid gene expression. Importantly, the mRNA expression may not reflect the corresponding protein expression for every gene in a cell (\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e, \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e). However, it can give indications of changes in cellular response to different stimuli as well as changes in expression level of individual proteins over time.\u003c/p\u003e \u003cp\u003eClustering of the scRNA seq data was found to be greatly influenced by the plasmid gene expression, indicating a separation of subpopulations of cells with high or low plasmid gene expression. Both the adherent and suspension subsets acquired a positively skewed distribution of the plasmid gene expression module score, indicating that high plasmid gene expression is further limited to a small subpopulation of cells.\u003c/p\u003e \u003cp\u003eAAV production necessitates the successful transfection of all three plasmids, as it would otherwise lead to low titers or incomplete viral packaging (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e). However, the efficiency of transfection typically suffers from high variability and lack of scalability (\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e). scRNA seq revealed differences in plasmid gene expression between individual cells originating from the same sample and showed that a large proportion of cells lacked expression of at least one plasmid. This proportion was predominantly influenced by phenotype and marginally influenced by time post-transfection, with suspension cells and later time points exhibiting a larger proportion containing all three plasmids.\u003c/p\u003e \u003cp\u003eThe proportion of cells exhibiting high expression of plasmid genes were 5 and 8% for the adherent and suspension samples, respectively. High plasmid gene expression was only observed for the cells expressing all three plasmids. However, despite the higher proportion of cells expressing all three plasmids in suspension, no large difference in the percentage of cells with high expression was observed between adherent and suspension cell samples.\u003c/p\u003e \u003cp\u003eAn equivalently low proportion of cells containing intracellular assembled capsids was observed, potentially linked to the low percentage of cells expressing high plasmid gene expression. Dash et al. had similar findings with only 5\u0026ndash;10% of transiently-transfected HEK293SF cells expressing assembled rAAV capsids, even with high transfection efficiencies and vector yields. These findings, combined with the fact that only a small fraction of the suspension cells (8%) and adherent cells (5%) at 48 hpt had high expression of plasmid genes (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e), indicates that not all the cells were successfully producing rAAVs. This shows that the heterogeneity in plasmid gene expression and the low capsid assembly are bottlenecks that needs to be addressed to improve titers.\u003c/p\u003e \u003cp\u003eThe clear separation of cells with different plasmid gene expressions during clustering indicated that distinct progression towards different states of plasmid gene expression could be discerned from trajectory inference. Three lineages were identified, all originating from cluster 0 which contained the largest number of cells and displayed low plasmid gene expression. One lineage, lineage β, led to the low-to-moderate plasmid gene expression in cluster 2. Another lineage, lineage α, transitioned into high plasmid gene expression. Genes that were differentially expressed along the lineages were identified and grouped into similar expression patterns, and significant GO terms associated with the groups of genes were determined.\u003c/p\u003e \u003cp\u003eThe adherent lineage β leading to low plasmid gene expression started with terms associated with the G1 phase of the cell cycle (\u003cspan additionalcitationids=\"CR53 CR54\" citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e). Other significant GO terms include canonical Wnt signaling pathway and the regulation of the β-catenin import into the nucleus. The canonical Wnt signaling pathway is responsible for G1- to S-phase progression (\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e). The accumulation of β-catenin in the cytoplasm occurs during the S-phase. The cells in the adherent lineage β of adherent cells began with an upregulation of genes from the canonical Wnt signaling followed by an upregulation of genes regulating import of the β-catenin into the nucleus, which was then downregulated at the end. This behavior suggests the progression of the cells from G1 to S to G2/M phase. GO terms associated with the regulation of the immune response and response to incorrect protein folding were upregulated at the end of the lineage.\u003c/p\u003e \u003cp\u003eIn the lineage α of adherent cells leading to high plasmid gene expression, GO terms relating to the M and G1/S phases of the cell cycle were cyclically expressed in the lineage, and underwent two cycles of upregulation and downregulation. Positive regulation of transmembrane receptor protein serine/threonine kinase signaling pathway, which is highly associated with the M phase (\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e), was upregulated in the beginning of the lineage, followed by downregulation. Lineage α displayed a high upregulation of ion transmembrane transport at the end, specifically for potassium, which has been linked to G1/S transition (\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e). These findings show the progression of cells from M- to G1- and finally to the S-phase. Similar to the lineage β of adherent cells, GO terms related to the immune response were also upregulated at the end of the lineage.\u003c/p\u003e \u003cp\u003eThe lineage β of suspension cells leading to low plasmid gene expression showed some cyclical upregulation and downregulation, but most groups were either upregulated at the beginning or the end of the lineage. M- and G1-phase related GO terms such as mitosis, generation of precursor metabolites and cell size regulation were cyclically expressed, and both were downregulated at the end of the lineage. The negative regulation of the cell cycle observed at the end of the lineage could potentially indicate an arrest in the M-phase.\u003c/p\u003e \u003cp\u003eReminiscent of the lineage β of adherent cells, incorrect protein folding response was also observed for the suspension lineage β. The lineages leading to low plasmid gene expression were the only ones displaying incorrect protein folding, suggesting that this could be a potential reason for the low plasmid gene levels. Other omics studies also confirm the association of protein folding with rAAV production (\u003cspan additionalcitationids=\"CR16 CR17\" citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e). The outcome herein shows that protein folding could be optimized further to prevent limitations caused by incorrectly folded protein.\u003c/p\u003e \u003cp\u003eIn the lineage α of suspension cells leading to high plasmid gene expression, GO terms related to the M-phase such as mitosis and chromosome segregation were upregulated at early pseudotime values for lineage α. In line with this, an upregulation of GO terms related to the G1-phase such as cell growth and metabolism was observed. The gamma aminobutyric acid signaling pathway upregulated at the end of the lineage has been shown to lead to an accumulation of cells in the S-phase (\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e). These results indicate the transition of cells from M- to G1- to S-phase for the lineage α of suspension cells, a behavior which is also seen in the lineage α of adherent cells.\u003c/p\u003e \u003cp\u003eGenerally, an earlier immune response was activated for the β lineages leading to low plasmid gene expression. The α lineages leading to high plasmid gene expression also triggered the immune response, although at a much later time point and nearing the end of the lineage. This outcome indicates that the regulation or the delay of the immune response may be a valid strategy for enhancing rAAV productivity.\u003c/p\u003e \u003cp\u003eThe link between wild type AAV expansion in vivo and the cell cycle has been previously established. AAVs have been found to modulate the host cell cycle to facilitate viral genome replication, which mainly occurs in the S/G2 phase. (\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e) Franzoso et al. showed that AAV2 gene expression occurs exclusively in cells in the S/G2 phases due to the cell cycle dependent expression of the Rep gene (\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e). Rep 78 has been shown to activate caspase-3, which induces apoptosis during the G\u003csub\u003e1\u003c/sub\u003e and early S phases of the cell cycle (\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e), as well as to decrease the Cdc25A activity and accumulate pRb, which induces a complete arrest of the cells in the S phase (\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e, \u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e). Additionally, many other genes involved in rAAV production induce entry into the S phase or freeze the cells in the S phase by preventing mitosis, apoptosis such as the E1 and E4 genes (\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e, \u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e). The host cell DNA replication machinery required by DNA viruses for replication is only accessible during the S phase. To make the environment favorable for viral replication, AAVs must induce S-phase entry of infected cells, which they achieve through the virally-encoded proteins that induce S phase entry (\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe influence of the cell cycle on rAAV productivity was partly shown by Barnes et al., who showed that modifying the SKA2 and ITPRIP gene expressions, which lead to modulatory effects on the cell cycle increased rAAV genome titers (\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e). Another study by Tworig et al. who performed bulk RNA seq during rAAV production found the upregulation of proliferation inhibitors and downregulation of proliferation enhancers among high AAV producer cell lines, and proposed modulation of the cell cycle as a strategy to improve AAV titers (\u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e). The findings in the present study show a strong link between high plasmid gene expression and the cell cycle. Cell cycle synchronization could be used as a strategy for improving plasmid gene expression and thus rAAV production.\u003c/p\u003e "},{"header":"Conclusion","content":" \u003cp\u003eThis study reports the first application of single-cell transcriptomics on transient transfection-based rAAV production with adherent and suspension HEK293T cells. In our effort to characterize the process, we uncovered the substantial heterogeneity in HEK293 cells during production and identified subgroups of cells with high and low rAAV plasmid gene expression.\u003c/p\u003e \u003cp\u003eMajor differences in plasmid gene expression were observed between the cells, with nearly half of the cells missing expression of genes from at least one plasmid suggesting that successful plasmid gene expression is a significant bottleneck for the production. Additionally, we determined that a low proportion of cells produce intracellular assembled capsids. Key factors that distinguished cells with high and low plasmid gene expression were identified, with cell cycle being the primary contributor. Reducing the heterogeneity through strategies such as cell cycle synchronization could potentially improve the productivity of rAAV production based on transient transfection of HEK293 cells with producer plasmids. This study contributes to our biological understanding of the production process as well as emphasizes the need for further omics studies.\u003c/p\u003e \u003c/div\u003e"},{"header":"Abbreviations","content":"\u003cp\u003e\u003cstrong\u003erAAV\u003c/strong\u003e recombinant adeno-associated virus\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRNA seq\u003c/strong\u003e RNA sequencing\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003escRNA seq\u003c/strong\u003e single-cell RNA sequencing\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePEI\u003c/strong\u003e polyethylenimine\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eGFP\u003c/strong\u003e Green fluorescent protein\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eHPT\u003c/strong\u003e hours post transfection\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDMSO\u003c/strong\u003e dimethylsulfoxide\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePBS\u003c/strong\u003e phosphate-buffered saline\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePBST\u003c/strong\u003e phosphate-buffered saline with 0.1% Tween 20\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSV40LT\u003c/strong\u003e simian virus 40 large T antigen\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eUMI\u003c/strong\u003e Unique Molecular Identifier\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eADH\u003c/strong\u003e adherent\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSUP\u003c/strong\u003e suspension\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch1\u003eAuthor contributions\u003c/h1\u003e\n\u003cp\u003eConceptualization, B.L., V.C.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMethodology, B.L., S.T.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eSoftware, B.L., S.T.\u003c/p\u003e\n\u003cp\u003eData curation, B.L., S.T.\u003c/p\u003e\n\u003cp\u003eInvestigation, B.L., S.T.\u003c/p\u003e\n\u003cp\u003eFormal analysis, B.L., S.T.\u003c/p\u003e\n\u003cp\u003eSupervision, T.G., V.C.\u003c/p\u003e\n\u003cp\u003eFunding acquisition, T.G., V.C.,\u003c/p\u003e\n\u003cp\u003eVisualization, B.L., S.T.\u003c/p\u003e\n\u003cp\u003eProject administration, T.G., V.C.\u003c/p\u003e\n\u003cp\u003eResources, T.G., V.C.\u003c/p\u003e\n\u003cp\u003eWriting\u0026mdash;original draft preparation, S.T., B.L.\u003c/p\u003e\n\u003cp\u003eWriting\u0026mdash;review and editing, S.T., B.L., T.G., V.C.\u003c/p\u003e\n\u003ch1 id=\"_Toc187561756\"\u003eAcknowledgements\u003c/h1\u003e\n\u003cp\u003eThis research was supported by the Competence Centre for Advanced BioProduction by Continuous Processing, AdBIOPRO, funded by Sweden\u0026rsquo;s Innovation Agency VINNOVA (diaries nr. 2022-03170), and by the Marie Skłodowska-Curie European Industrial Doctorate programme EID, STACCATO grant No 813453. The authors would also like to thank the centre GeneNova funded by VINNOVA.\u003c/p\u003e\n\u003ch1 id=\"_Toc187561757\"\u003eConflict of interest\u003c/h1\u003e\n\u003cp\u003eThe authors declare no conflict of interest.\u003c/p\u003e\n\u003ch1 id=\"_Toc187561758\"\u003eData availability\u003c/h1\u003e\n\u003cp\u003eThe datasets analysed during the current study are available in the Sequencing Read Archive (SRA) with the accession number http://www.ncbi.nlm.nih.gov/bioproject/1236707\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eRobbins, P. D., Tahara, H. \u0026amp; Ghivizzani, S. C. Viral vectors for gene therapy. \u003cem\u003eTrends Biotechnol.\u003c/em\u003e \u003cb\u003e16\u003c/b\u003e (1), 35\u0026ndash;40 (1998).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, D., Tai, P. W. \u0026amp; Gao, G. Adeno-associated virus vector as a platform for gene therapy delivery. \u003cem\u003eNat. Rev. Drug Discovery\u003c/em\u003e. \u003cb\u003e18\u003c/b\u003e (5), 358\u0026ndash;378 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSha, S. et al. Cellular pathways of recombinant adeno-associated virus production for gene therapy. \u003cem\u003eBiotechnol. Adv.\u003c/em\u003e \u003cb\u003e49\u003c/b\u003e, 107764 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, C. \u0026amp; Samulski, R. J. Engineering adeno-associated virus vectors for gene therapy. \u003cem\u003eNat. Rev. Genet.\u003c/em\u003e \u003cb\u003e21\u003c/b\u003e (4), 255\u0026ndash;272 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDaya, S. \u0026amp; Berns, K. I. Gene therapy using adeno-associated virus vectors. \u003cem\u003eClin. Microbiol. Rev.\u003c/em\u003e \u003cb\u003e21\u003c/b\u003e (4), 583\u0026ndash;593 (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDestro, F. et al. The state of technological advancement to address challenges in the manufacture of rAAV gene therapies. \u003cem\u003eBiotechnol. Adv.\u003c/em\u003e :108433. (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, J-H., Gessler, D. J., Zhan, W., Gallagher, T. L. \u0026amp; Gao, G. Adeno-associated virus as a delivery vector for gene therapy of human diseases. \u003cem\u003eSignal. Transduct. Target. Therapy\u003c/em\u003e. \u003cb\u003e9\u003c/b\u003e (1), 78 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmith, J., Grieger, J. \u0026amp; Samulski, R. J. Overcoming bottlenecks in AAV manufacturing for gene therapy. \u003cem\u003eCell. Gene Therapy Insights\u003c/em\u003e. \u003cb\u003e4\u003c/b\u003e (8), 815\u0026ndash;825 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKliegman, M. et al. A roadmap for affordable genetic medicines. \u003cem\u003eNature\u003c/em\u003e :1\u0026ndash;8. (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, Y. et al. Decoding cellular mechanism of recombinant adeno-associated virus (rAAV) and engineering host-cell factories toward intensified viral vector manufacturing. \u003cem\u003eBiotechnol. Adv.\u003c/em\u003e :108322. (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNaso, M. F., Tomkowicz, B., Perry, W. L. III \u0026amp; Strohl, W. R. Adeno-associated virus (AAV) as a vector for gene therapy. \u003cem\u003eBioDrugs\u003c/em\u003e \u003cb\u003e31\u003c/b\u003e (4), 317\u0026ndash;334 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSrivastava, A., Mallela, K. M., Deorkar, N. \u0026amp; Brophy, G. Manufacturing challenges and rational formulation development for AAV viral vectors. \u003cem\u003eJ. Pharm. Sci.\u003c/em\u003e \u003cb\u003e110\u003c/b\u003e (7), 2609\u0026ndash;2624 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFu, Q., Polanco, A., Lee, Y. S. \u0026amp; Yoon, S. Critical challenges and advances in recombinant adeno-associated virus (rAAV) biomanufacturing. \u003cem\u003eBiotechnol. Bioeng.\u003c/em\u003e \u003cb\u003e120\u003c/b\u003e (9), 2601\u0026ndash;2621 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, H. et al. Creation of a high-yield AAV vector production platform in suspension cells using a design-of-experiment approach. \u003cem\u003eMol. Therapy Methods Clin. Dev.\u003c/em\u003e \u003cb\u003e18\u003c/b\u003e, 312\u0026ndash;320 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, Y., Fu, Q., Lee, Y. S., Sha, S. \u0026amp; Yoon, S. Transcriptomic features reveal molecular signatures associated with recombinant adeno-associated virus production in HEK293 cells. \u003cem\u003eBiotechnol. Prog.\u003c/em\u003e \u003cb\u003e39\u003c/b\u003e (4), e3346 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVirgolini, N. et al. Transcriptome analysis of Sf9 insect cells during production of recombinant Adeno-associated virus. \u003cem\u003eBiotechnol. J.\u003c/em\u003e \u003cb\u003e18\u003c/b\u003e (2), 2200466 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChung, C-H. et al. Production of rAAV by plasmid transfection induces antiviral and inflammatory responses in suspension HEK293 cells. \u003cem\u003eMol. Therapy Methods Clin. Dev.\u003c/em\u003e \u003cb\u003e28\u003c/b\u003e, 272\u0026ndash;283 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLu, M., Lee, Z. \u0026amp; Hu, W. S. Multi-omics kinetic analysis of recombinant adeno‐associated virus production by plasmid transfection of HEK293 cells. \u003cem\u003eBiotechnol. Prog.\u003c/em\u003e \u003cb\u003e40\u003c/b\u003e (2), e3428 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHwang, B., Lee, J. H. \u0026amp; Bang, D. Single-cell RNA sequencing technologies and bioinformatics pipelines. \u003cem\u003eExp. Mol. Med.\u003c/em\u003e \u003cb\u003e50\u003c/b\u003e (8), 1\u0026ndash;14 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoldman, S. L. et al. The impact of heterogeneity on single-cell sequencing. \u003cem\u003eFront. Genet.\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e, 8 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAltschuler, S. J. \u0026amp; Wu, L. F. Cellular heterogeneity: do differences make a difference? \u003cem\u003eCell\u003c/em\u003e \u003cb\u003e141\u003c/b\u003e (4), 559\u0026ndash;563 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKolodziejczyk, A. A., Kim, J. K., Svensson, V., Marioni, J. C. \u0026amp; Teichmann, S. A. The technology and biology of single-cell RNA sequencing. \u003cem\u003eMol. Cell\u003c/em\u003e. \u003cb\u003e58\u003c/b\u003e (4), 610\u0026ndash;620 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChoi, Y. H. \u0026amp; Kim, J. K. Dissecting cellular heterogeneity using single-cell RNA sequencing. \u003cem\u003eMol. Cells\u003c/em\u003e. \u003cb\u003e42\u003c/b\u003e (3), 189\u0026ndash;199 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJovic, D. et al. Single-cell RNA sequencing technologies and applications: A brief overview. \u003cem\u003eClin. translational Med.\u003c/em\u003e \u003cb\u003e12\u003c/b\u003e (3), e694 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVirgolini, N. et al. Impact of dual-baculovirus infection on the Sf9 insect cell transcriptome during rAAV production using single‐cell RNA‐seq. \u003cem\u003eBiotechnol. Bioeng.\u003c/em\u003e \u003cb\u003e120\u003c/b\u003e (9), 2588\u0026ndash;2600 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLadd, B., Bowes, K., Lundgren, M., Gr\u0026auml;slund, T. \u0026amp; Chotteau, V. Proof-of-concept of continuous transfection for adeno-associated virus production in microcarrier-based culture. \u003cem\u003eProcesses\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e (3), 515 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSayers, E. W. et al. Database resources of the national center for biotechnology information. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e50\u003c/b\u003e (D1), D20\u0026ndash;D6 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePertea, G. \u0026amp; Pertea, M. GFF utilities: GffRead and GffCompare. \u003cem\u003eF1000Research\u003c/em\u003e ;9. (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDobin, A. et al. STAR: ultrafast universal RNA-seq aligner. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e29\u003c/b\u003e (1), 15\u0026ndash;21 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHao, Y. et al. Dictionary learning for integrative, multimodal and scalable single-cell analysis. \u003cem\u003eNat. Biotechnol.\u003c/em\u003e \u003cb\u003e42\u003c/b\u003e (2), 293\u0026ndash;304 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePatterson-Cross, R. B., Levine, A. J. \u0026amp; Menon, V. Selecting single cell clustering parameter values using subsampling-based robustness metrics. \u003cem\u003eBMC Bioinform.\u003c/em\u003e \u003cb\u003e22\u003c/b\u003e, 1\u0026ndash;13 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrabski, I. N., Street, K. \u0026amp; Irizarry, R. A. Significance analysis for clustering with single-cell RNA-sequencing data. \u003cem\u003eNat. Methods\u003c/em\u003e. \u003cb\u003e20\u003c/b\u003e (8), 1196\u0026ndash;1202 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStreet, K. et al. Slingshot: cell lineage and pseudotime inference for single-cell transcriptomics. \u003cem\u003eBMC Genom.\u003c/em\u003e \u003cb\u003e19\u003c/b\u003e, 1\u0026ndash;16 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVan den Berge, K. et al. Trajectory-based differential expression analysis for single-cell sequencing data. \u003cem\u003eNat. Commun.\u003c/em\u003e \u003cb\u003e11\u003c/b\u003e (1), 1201 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAmezquita, R. A. et al. Orchestrating single-cell analysis with Bioconductor. \u003cem\u003eNat. Methods\u003c/em\u003e. \u003cb\u003e17\u003c/b\u003e (2), 137\u0026ndash;145 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRisso, D. et al. clusterExperiment and RSEC: A Bioconductor package and framework for clustering of single-cell and other large gene expression datasets. \u003cem\u003ePLoS Comput. Biol.\u003c/em\u003e \u003cb\u003e14\u003c/b\u003e (9), e1006378 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu, T. et al. clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. \u003cem\u003eInnov.\u003c/em\u003e ;\u003cb\u003e2\u003c/b\u003e(3). (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKuleshov, M. V. et al. Enrichr: a comprehensive gene set enrichment analysis web server 2016 update. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e \u003cb\u003e44\u003c/b\u003e (W1), W90\u0026ndash;W7 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAndrews, S. FastQC: a quality control tool for high throughput sequence data. [ (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEwels, P., Magnusson, M., Lundin, S. \u0026amp; K\u0026auml;ller, M. MultiQC: summarize analysis results for multiple tools and samples in a single report. \u003cem\u003eBioinformatics\u003c/em\u003e \u003cb\u003e32\u003c/b\u003e (19), 3047\u0026ndash;3048 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBray, N. L., Pimentel, H., Melsted, P. \u0026amp; Pachter, L. Near-optimal probabilistic RNA-seq quantification. \u003cem\u003eNat. Biotechnol.\u003c/em\u003e \u003cb\u003e34\u003c/b\u003e (5), 525\u0026ndash;527 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLove, M. I., Huber, W. \u0026amp; Anders, S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. \u003cem\u003eGenome Biol.\u003c/em\u003e \u003cb\u003e15\u003c/b\u003e, 1\u0026ndash;21 (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGurazada, S. G. R. et al. HEK-omics: The promise of omics to optimize HEK293 for recombinant adeno-associated virus (rAAV) gene therapy manufacturing. \u003cem\u003eBiotechnol. Adv.\u003c/em\u003e :108506. (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBorsi, G. et al. Single-cell RNA sequencing reveals homogeneous transcriptome patterns and low variance in a suspension CHO-K1 and an adherent HEK293FT cell line in culture conditions. \u003cem\u003eJ. Biotechnol.\u003c/em\u003e \u003cb\u003e364\u003c/b\u003e, 13\u0026ndash;22 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJones, J. E., Le Sage, V. \u0026amp; Lakdawala, S. S. Viral and host heterogeneity and their effects on the viral life cycle. \u003cem\u003eNat. Rev. Microbiol.\u003c/em\u003e \u003cb\u003e19\u003c/b\u003e (4), 272\u0026ndash;282 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSwaminath, S. \u0026amp; Russell, A. B. The use of single-cell RNA-seq to study heterogeneity at varying levels of virus\u0026ndash;host interactions. \u003cem\u003ePLoS Pathog.\u003c/em\u003e \u003cb\u003e20\u003c/b\u003e (1), e1011898 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNguyen, T. N. et al. Mechanistic model for production of recombinant adeno-associated virus via triple transfection of HEK293 cells. \u003cem\u003eMol. Therapy Methods Clin. Dev.\u003c/em\u003e \u003cb\u003e21\u003c/b\u003e, 642\u0026ndash;655 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChahal, P. S., Schulze, E., Tran, R., Montes, J. \u0026amp; Kamen, A. A. Production of adeno-associated virus (AAV) serotypes by transient transfection of HEK293 cell suspension cultures for gene delivery. \u003cem\u003eJ. Virol. Methods\u003c/em\u003e. \u003cb\u003e196\u003c/b\u003e, 163\u0026ndash;173 (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaier, T., G\u0026uuml;ell, M. \u0026amp; Serrano, L. Correlation of mRNA and protein in complex biological samples. \u003cem\u003eFEBS Lett.\u003c/em\u003e \u003cb\u003e583\u003c/b\u003e (24), 3966\u0026ndash;3973 (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEdfors, F. et al. Gene-specific correlation of RNA and protein levels in human cells and tissues. \u003cem\u003eMol. Syst. Biol.\u003c/em\u003e \u003cb\u003e12\u003c/b\u003e (10), 883 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan der Loo, J. C. M. \u0026amp; Wright, J. F. Progress and challenges in viral vector manufacturing. \u003cem\u003eHum. Mol. Genet.\u003c/em\u003e \u003cb\u003e25\u003c/b\u003e (R1), R42\u0026ndash;R52 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCalder, A. et al. Lengthened G1 phase indicates differentiation status in human embryonic stem cells. \u003cem\u003eStem. Cells. Dev.\u003c/em\u003e \u003cb\u003e22\u003c/b\u003e (2), 279\u0026ndash;295 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTer Huurne, M. \u0026amp; Stunnenberg, H. G. G1-phase progression in pluripotent stem cells. \u003cem\u003eCell. Mol. Life Sci.\u003c/em\u003e \u003cb\u003e78\u003c/b\u003e (10), 4507\u0026ndash;4519 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSoufi, A. \u0026amp; Dalton, S. Cycling through developmental decisions: how cell cycle dynamics control pluripotency, differentiation and reprogramming. \u003cem\u003eDevelopment\u003c/em\u003e \u003cb\u003e143\u003c/b\u003e (23), 4301\u0026ndash;4311 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHowell, A. S. \u0026amp; Lew, D. J. Morphogenesis and the cell cycle. \u003cem\u003eGenetics\u003c/em\u003e \u003cb\u003e190\u003c/b\u003e (1), 51\u0026ndash;77 (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNiehrs, C. \u0026amp; Acebron, S. P. Mitotic and mitogenic Wnt signalling. \u003cem\u003eEMBO J.\u003c/em\u003e \u003cb\u003e31\u003c/b\u003e (12), 2705\u0026ndash;2713 (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, S-C., Huang, C-H., Chen, N-J., Chou, C-K. \u0026amp; Lin, C-H. Functional implication of human serine/threonine kinase, hAIK, in cell cycle progression. \u003cem\u003eJ. Biomed. Sci.\u003c/em\u003e \u003cb\u003e7\u003c/b\u003e (6), 484\u0026ndash;493 (2000).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRosendo-Pineda, M. J., Moreno, C. M. \u0026amp; Vaca, L. Role of ion channels during cell division. \u003cem\u003eCell. Calcium\u003c/em\u003e. \u003cb\u003e91\u003c/b\u003e, 102258 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, D. D., Kriegstein, A. R. \u0026amp; Ben-Ari, Y. GABA regulates stem cell proliferation before nervous system formation. \u003cem\u003eEpilepsy Currents\u003c/em\u003e. \u003cb\u003e8\u003c/b\u003e (5), 137\u0026ndash;139 (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWinocour, E., Callaham, M. \u0026amp; Huberman, E. Perturbation of the cell cycle by adeno-associated virus. \u003cem\u003eVirology\u003c/em\u003e \u003cb\u003e167\u003c/b\u003e (2), 393\u0026ndash;399 (1988).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFranzoso, F. D. et al. Cell cycle-dependent expression of adeno-associated virus 2 (AAV2) rep in coinfections with herpes simplex virus 1 (HSV-1) gives rise to a mosaic of cells replicating either AAV2 or HSV-1. \u003cem\u003eJ. Virol.\u003c/em\u003e \u003cb\u003e91\u003c/b\u003e (15). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1128/jvi\u003c/span\u003e\u003cspan address=\"10.1128/jvi\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2017). 00357\u0026thinsp;\u0026ndash;\u0026thinsp;17.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchmidt, M., Afione, S. \u0026amp; Kotin, R. M. Adeno-associated virus type 2 Rep78 induces apoptosis through caspase activation independently of p53. \u003cem\u003eJ. Virol.\u003c/em\u003e \u003cb\u003e74\u003c/b\u003e (20), 9441\u0026ndash;9450 (2000).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBerthet, C., Raj, K., Saudan, P. \u0026amp; Beard, P. How adeno-associated virus Rep78 protein arrests cells completely in S phase. Proceedings of the National Academy of Sciences. ;102(38):13634-9. (2005).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaudan, P., Vlach, J. \u0026amp; Beard, P. Inhibition of S-phase progression by adeno‐associated virus Rep78 protein is mediated by hypophosphorylated pRb. \u003cem\u003eEMBO J.\u003c/em\u003e (2000).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNash, K., Chen, W., Salganik, M. \u0026amp; Muzyczka, N. Identification of cellular proteins that interact with the adeno-associated virus rep protein. \u003cem\u003eJ. Virol.\u003c/em\u003e \u003cb\u003e83\u003c/b\u003e (1), 454\u0026ndash;469 (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBen-Israel, H. \u0026amp; Kleinberger, T. Adenovirus and cell cycle control. \u003cem\u003eFront. Biosci.\u003c/em\u003e \u003cb\u003e7\u003c/b\u003e, d1369\u0026ndash;d95 (2002).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBagga, S. \u0026amp; Bouchard, M. J. Cell cycle regulation during viral infection. \u003cem\u003eCell. Cycle Control: Mech. Protocols\u003c/em\u003e :165\u0026ndash;227. (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarnes, C. R. et al. Genome-wide activation screens to increase adeno-associated virus production. \u003cem\u003eMol. Therapy-Nucleic Acids\u003c/em\u003e. \u003cb\u003e26\u003c/b\u003e, 94\u0026ndash;103 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTworig, J. et al. Transcriptomics-informed pharmacology identifies epigenetic and cell cycle regulators that enhance AAV production. \u003cem\u003eMol. Therapy Methods Clin. Dev.\u003c/em\u003e ;\u003cb\u003e32\u003c/b\u003e(4). (2024).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"AAV, transient transfection, transfection heterogeneity, single-cell transcriptomics, HEK293 cells","lastPublishedDoi":"10.21203/rs.3.rs-6220373/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6220373/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe popularity of rAAV vectors in gene therapy are placing a burden on current production systems. To improve the accessibility of these life changing treatments, increases in production yields and a reduction in the cost-of-goods are needed. Transient transfection is the most common way to introduce rAAV-encoding plasmids to producer cells but it suffers from significant drawbacks such as low and inconsistent yields as well as high cost due to its need for plasmid DNA. This study aims to address the low yield of transient transfection-based rAAV production through advanced methods in process characterization. Adherent and suspension cultures of a HEK293T cell line were triple-transfected for rAAV9 production using polyethylenimine (PEI). Samples were taken at various times post-transfection for analysis with bulk and single-cell transcriptomics. It was revealed that 46% of the cells lacked transcripts of genes from at least one plasmid, indicating that a significant proportion of the cells did not have the genes necessary for rAAV9 production. Among the remaining 54% of the cells expressing genes from all three plasmids, only 8% showed high plasmid gene expression. Flow cytometric analysis of intracellular rAAV9 confirmed these results by showing that only\u0026thinsp;~\u0026thinsp;3% of cells contained assembled rAAV9 capsids. Titre analysis by qPCR of the supernatant and lysate of the producer cells indicated an average culture performance of 10\u003csup\u003e13\u003c/sup\u003e vg/L. Analysis of the single-cell transcriptomic data showed that a significant proportion of cells that had high plasmid gene expression were in the S-phase. Trajectory inference highlighted that genes involved in the G2-M phase transition, immune response, and protein unfolding were differentially expressed at the branch point between high and low plasmid expression. This study reveals a significant bottleneck in the transient transfection-based production of rAAV. With less than 5% of cells producing rAAV, significant improvements in titres can be achieved if this fraction can be increased. Moreover, regulation of the cell-cycle, inhibition of the immune response, and alleviating protein misfolding all potentially offer the key to enabling these life changing treatments to reach a wider audience.\u003c/p\u003e","manuscriptTitle":"Heterogeneity in an Adeno-Associated Virus Transfection-Based Production Process Limits the Production Efficiency","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-03-27 18:15:52","doi":"10.21203/rs.3.rs-6220373/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-05-06T03:49:06+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-29T03:13:30+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-05T01:27:36+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"207397887215317451827915972127815181285","date":"2025-03-28T02:33:25+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-03-25T20:28:30+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"69975793198134697126745032247377663325","date":"2025-03-24T10:36:08+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"138291820717616224061101988535461411475","date":"2025-03-23T14:00:25+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"336382169786981160957310204686919518414","date":"2025-03-19T20:27:58+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-03-19T19:07:29+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-03-19T18:56:00+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-03-19T18:43:41+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-03-18T11:57:40+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2025-03-13T13:05:55+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"56afc6c6-eef0-48f3-9067-dd7b56eec863","owner":[],"postedDate":"March 27th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":46209269,"name":"Biological sciences/Biotechnology/Gene therapy"},{"id":46209270,"name":"Biological sciences/Biotechnology/Sequencing/Rna sequencing"}],"tags":[],"updatedAt":"2025-11-10T16:03:26+00:00","versionOfRecord":{"articleIdentity":"rs-6220373","link":"https://doi.org/10.1038/s41598-025-26261-0","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2025-11-04 15:57:18","publishedOnDateReadable":"November 4th, 2025"},"versionCreatedAt":"2025-03-27 18:15:52","video":"","vorDoi":"10.1038/s41598-025-26261-0","vorDoiUrl":"https://doi.org/10.1038/s41598-025-26261-0","workflowStages":[]},"version":"v1","identity":"rs-6220373","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6220373","identity":"rs-6220373","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00