Further information and requests for resources and reagents should be directed to and will be fulfilled by the Lead Contact, Ronald Chandler (
[email protected] ).
This study did not generate new unique reagents.
The accession number for the sequencing data generated in this manuscript is GEO: GSE148474 . Previously published datasets analyzed herein are also available at GEO: GSE121198 .
All mice were maintained on an outbred genetic background using CD-1 mice (Charles River). (Gt)R26Pik3ca * H1047R , LtfCre ( Tg(Ltf-iCre)14Mmul) and Ep300 fl alleles were purchased from The Jackson Laboratory and identified by PCR using published methods ( Adams et al., 2011 ; Daikoku et al., 2014 ; Kasper et al., 2006 ). Arid1a fl allele was distinguished by PCR as previously described ( Chandler et al., 2015 ). Genotyping primers are listed in Key Resources Table . Endpoints were vaginal bleeding, severe abdominal distension, and signs of severe illness including dehydration, hunching, jaundice, ruffled fur, signs of infection, or non-responsiveness. Sample sizes for each genotype were chosen based on the proportions of animals with vaginal bleeding between each experimental group and Kaplan-Meyer log rank test for survival differences. All mice analyzed in the study were between 6 and 32 weeks old. In cases where a mobility endpoint occurred, tissues were collected at the time of vaginal bleeding, including LtfCre 0/+ ; (Gt)R26Pik3ca * H1047R ; Arid1a fl/fl (μ 1/2 = 107 days) and LtfCre 0/+ ; (Gt)R26Pik3ca * H1047R ; Arid1a fl/fl ; Ep300 fl/fl (μ 1/2 = 143 days) mice. In cases where the animal did not reach a morbidity endpoint or show reduced survival, tissues were collected at comparable time points (between 90 and 150 days) from age-matched, littermate control mice from the mutant crosses. Uteri were collected at time of sacrifice and placed immediately into neutral-buffered formalin at 4°C. After 24 hr, tissues were washed with PBS and 50% EtOH, placed in 70% EtOH, and weight measurements were recorded. Mice were housed at the Michigan State University Grand Rapids Research Center in accordance with protocols approved by Michigan State University. Michigan State University is registered with the U.S. Department of Agriculture (USDA) and has an approved Animal Welfare Assurance from the NIH Office of Laboratory Animal Welfare (OLAW). MSU is accredited by the Association for Assessment and Accreditation of Laboratory Animal Care (AAALAC).
12Z immortalized human endometrial epithelial cells ( Zeitvogel et al., 2001 ) were maintained in DMEM/F12 media supplemented with 10% fetal bovine serum (FBS), 1% L-glutamine and 1% penicillin/streptomycin (P/S). The 12Z cells were provided by the laboratory of Asgi Fazleabas, and cell line validation was performed by IDEXX BioResearch, finding the result that the 12Z cell line has a unique profile not found in the current public databases. A recent study found 12Z cells to be an authentic and pure endometriosis cell line based on marker analysis and short tandem repeat profiling ( Romano et al., 2020 ). Lenti-X™ 293T (Clontech) cells were maintained in DMEM +110 mg/L Sodium Pyruvate (GIBCO) supplemented with 10% FBS, 1% L-glutamine, 1% P/S. 12Z and Lenti-X 293T cells were regularly tested for mycoplasma using the Mycoplasma PCR Detection Kit (Applied Biological Materials). No commonly mis-identified cell lines were used in this study.
For indirect immunohistochemistry (IHC), 10% neutral buffered formalin (NBF)-fixed paraffin sections were processed for heat-based antigen unmasking in 10 mM sodium citrate [pH 6.0]. Sections were incubated with antibodies at the following dilutions: 1:200 ARID1A (D2A8U) (12354, Cell Signaling); 1:1000 P300 (86377, Cell Signaling); 1:400 Phospho-S6 (4585, Cell Signaling); 1:100 KRT8 (TROMA1, DHSB); 1:200 Cleaved Caspase-3 (9579, Cell Signaling); 1:400 Ki67 (12202, Cell Signaling); 1:200 H3K27ac (39133, Active Motif); 1:200 H3K18ac (ab1191, Abcam); 1:1000 PAI-1 (SERPINE1) (ab66705, Abcam). TROMA-I antibody was deposited to the DSHB by Brulet, P./Kemler, R. (DSHB Hybridoma Product TROMA-I). Biotin-conjugated secondary antibodies were donkey anti-rabbit IgG (711-065-152, Jackson Immuno-research Lab) and donkey anti-rat IgG (#705-065-153, Jackson Immuno-research Lab). VECTASTAIN Elite ABC HRP Kit (Vector) was used for secondary antibody detection. Sections for IHC were lightly counter-stained with Hematoxylin QS or Methyl Green (Vector Labs). Routine Hematoxylin and Eosin (H&E) staining of sections was performed by the Van Andel Research Institute (VARI) Histology and Pathology Core. A VARI animal pathologist reviewed histological tumor assessments.
To determine H-scores from mouse slides, one field of view (20X) on a Nikon Eclipse Ni-U upright microscope per mouse from a slide stained with antibody (SERPINE1, H3K27ac, H3K18ac) was used. Epithelial and stromal cells were assigned a value from 0 to 3 indicating intensity of staining (no staining = 0, low staining = 1, moderate staining = 2, and strong staining = 3) and the proportion of cells was determined for each staining intensity. For Ki67, a value of 1 (positive staining) or 0 (negative staining) was assigned to determine the number of Ki67+ cells.
For human endometriosis tissue samples, SERPINE1 (PAI-1) IHC was carried out as published ( Alotaibi et al., 2019 ). Briefly, IHC using the EnVision+ Dual Link system (Dako) and 3,3-diaminobenzidine (DAB) was performed, using mouse monoclonal PAI-1 antibody C-9 (sc 5297, Santa Cruz). PAI-1 expression was evaluated in endometriotic epithelium and stroma using the Histoscore calculation. Areas of endometriosis epithelium and stroma were first scanned at low power (×10) and then analyzed at high power (×40) to evaluate the staining intensity and estimate the proportion of positive cells. ARID1A IHC was used as a surrogate of loss-of-function alterations ( Khalique et al., 2018 ; Trizzino et al., 2018 ) using a Dako Omnis automated immunostainer (Agilent Technologies) and the anti-ARID1A rabbit monoclonal D2A8U (Cell Signaling Technology).
12Z cells were seeded at a density of 30,000 cells/mL in DMEM/F12 media supplemented with 10% FBS and 1% L-glutamine. The following day, cells were transfected with 50 nM siRNA (Dharmacon, ON-TARGETplus Non-targeting Pool, human ARID1A #8289 SMARTpool, human P300 #3486 SMARTpool, human SERPINE1 #19376 SMARTpool) using the RNAiMax (ThermoFisher) lipofectamine reagent according to the manufacturer’s instructions at a ratio of 1:1 volume:volume in OptiMEM (GIBCO). After 24 hr, the media was replaced. For plasmid co-transfection experiments, cells were transfected the following day with 500ng pBabe vector containing PIK3CA H1047R (pPIK3CA H1047R ) or pBabe empty vector using the FuGene HD transfection reagent (Promega) according to the manufacturers’ instructions at a ratio of 2:1 volume:mass, and media was replaced after 4 hr. The pPIK3CA H1047R was a gift from Jean Zhao (Addgene plasmid 12524) ( Zhao et al., 2005 ). In A-485 co-treatment studies, A-485 was included in the media 24 hr-post transfection in 0.1% DMSO. 48 hr after transfection, media was replaced with DMEM/F-12 media supplemented with 0.5% FBS, 1% P/S and 1% L-glutamine. Cells were collected 72 hr-post siRNA transfection using the Quick-RNA Miniprep Kit (Zymo Research) for RNA, RIPA buffer (Cell Signaling) for whole cell lysate, or histone extraction.
Lentiviral particles expressing shRNAs were produced in 293T cells according to the manufacturers’ instructions. Lenti-X™293T cells were transfected with lentiviral packaging mix composed of pNHP and pVSVG (generous gifts from Dr. Fredric Manfredsson) and MISSION pKLO.1 plasmid containing non-targeting shRNA (control) or pooled ARID1A shRNAs (shARID1A) (Sigma) using poly-ethylenimine (PEI) in DMEM + 4.5g/L D-Glucose, 110mg/L Sodium Pyruvate, 10% FBS, 1% L-glutamine. After 24 hr, media was replaced with DMEM/F12, 10% FBS, 1% L-glutamine, 1% P/S. Viral particles were collected after 48 and 96 hr, and viral titers were calculated using the qPCR Lentiviral Titration Kit (ABM).
For lentiviral transduction of 12Z cells, cells were treated with a multiplicity of infection of 100 units per cell. After 24 hours, media was replaced. For plasmid co-transfection experiments, cells were transfected the following day with 500ng pBabe vector containing PIK3CA H1047R (pPIK3CA H1047R ) or pBabe empty vector using the FuGene HD transfection reagent (Promega) according to the manufacturers’ instructions at a ratio of 2:1 volume:mass, and media was replaced after 4 hr. In A-485 co-treatment studies, A-485 was included in the media 24 hr-post transfection in 0.1% DMSO. To generate stable expression cell lines, transduced cells were treated with 600 ng/mL puromycin (Sigma) for three weeks.
Cells were washed with PBS and scraped in PBS containing 5 mM sodium butyrate. Cells were centrifuged and resuspended in TEB buffer (phosphate buffered saline supplemented with 0.5% Triton X-100, 5 mM sodium butyrate, 2 mM phenylmethylsulfonyl fluoride, 1 × protease inhibitor cocktail) and incubated on a 3D spindle nutator at 4°C for 10 min. Cells were centrifuged at 3,000 RPM for 10 min at 4°C. TEB wash step was repeated once. Following second wash, pellet was resuspended in 0.2 N HCl, and incubated on 3D spindle nutator at 4°C overnight. The following day, samples were neutralized with 1:10 volume 1M Tris-HCl pH 8.3. Sample was centrifuged at 3,000 RPM for 10 min at 4°C, and supernatant containing histone proteins was collected.
Protein whole cell lysates and histone extracts were quantified using the Micro BCA Protein Assay Kit (ThermoFisher) and a FlexSystem3 plate reader. Protein lysates were run on a 4%–15% gradient SDS-PAGE gel (BioRad) and transferred to PVDF membrane using the TransBlot Turbo system (BioRad). Primary antibodies dilutions were 1:1,000 ARID1A (D2A8U) (12354, Cell Signaling); 1:100 P300 (NM11) (sc-32244, Santa Cruz); 1:1,000 β-Actin (8457, Cell Signaling); 1:100 PAI-1 (sc-5297, Santa Cruz); 1:1,000 Akt (4691, Cell Signaling); 1:2,000 Phospho-Akt (Ser473) (4060, Cell Signaling). Horseradish peroxidase (HRP) conjugated secondary antibodies (Cell Signaling) were used at a dilution of 1:2,000. Clarity Western ECL Substrate (BioRad) was used for protein band visualization, and western blot exposures were captured using the ChemiDoc XRS+ imaging system (BioRad).
For histone extracts, samples were run on a 15% SDS-PAGE gel and transferred to nitrocellulose membrane in 20 mM sodium phosphate pH 6.7 at 400 mA for 90 min. Primary antibody dilutions were 1:2,000 Histone H3 (4499, Cell Signaling); 1:1,000 H3K27ac (Active Motif, 39133). Donkey anti-rabbit IgG, IRDye 800CW conjugated secondary antibody (LI-COR Biosciences) was used at a dilution of 1:10,000 and fluorescence imaging was performed using the LI-COR Odyssey CLx imaging system (LI-COR Biosciences). Uncropped western blot images are collected in Figure S7 .
12Z cells were seeded in 6-well dishes at a density of 50,000 cells per well. After 24 hr, cells were transfected with siRNA as described above. For drug treatment experiments, cells were treated drug 24 hr after transfection. At 48 hr post-transfection, cells were trypsinized, and 100 μL of cell mixture containing 30,000 cells and 0.3 mg/mL Matrigel was seeded into transwell plates (8 μm pore polycarbonate membrane, Corning) pre-coated with 100 μL of 0.3 mg/mL Matrigel. After 1 hr, serum-free DMEM/F12 1% P/S, 1% L-glutamine media was added to the top chamber and DMEM/F12, 5% FBS, 1% P/S, 1% L-glutamine was added to the bottom chamber. For drug studies, drug was included in both top and bottom chamber media. After 16 hr, transwell units were transferred to plates containing 2 μg/mL calcein-AM in DMEM/F12. After 1 hr, media was aspirated from the top chamber and unmigrated cells were removed with a cotton swab. Images were collected using a Nikon Eclipse T i microscope in five non-overlapping fields per well. ImageJ software (National Institutes of Health) was used to quantify cells based on size and intensity.
12Z cells were seeded in 6-well dishes at a density of 50,000 cells per well. After 24 hr, cells were transfected with siRNA as described above. For drug treatment experiments, cells were treated drug 24 hr after transfection. At 48 hr post-transfection, cells were trypsinized, and 50 μL of cell mixture containing 10,000 cells and 0.3 mg/mL Matrigel was seeded into 96-well plates pre-coated with 100 μL of 0.3 mg/mL Matrigel. After 1 hr, 50 μL of serum-free DMEM/F12 1% P/S, 1% L-glutamine media was added. For drug studies, 100 nM A-485 or vehicle was included in the media. After 16 or 24 hr, 2 μg/mL calcein-AM and 4 μg/mL ethidium homodimer III were added. Wells were imaged using a Nikon Eclipse T i microscope, and ImageJ software (National Institutes of Health) was used to quantify cells based on size and intensity.
12Z cells were seeded into 35mm dishes containing 4-well culture inserts at a density of 4,000 cells per well. Cells were treated with lentiviral particles expressing non-targeting shRNA (control) or shARID1A at a multiplicity of infection of 100, 24 hr after seeding. Media was replaced with serum-free DMEM/F12 containing 1% L-glutamine and 1% P/S including drug or vehicle after 24 hr. Culture inserts were removed and serum-free media containing vehicle or drug was replenished after 16 hr. At 0 and 24 hr of migration, images were taken using a Nikon Eclipse T i microscope. Distances between migration fronts were measured using NIS Elements Advanced Research software at 16 different points 100 μm apart. Migration distance was calculated by subtracting the average distance across migration fronts at 24 hr from the average distance at 0 hr. Cells were counted within a window surrounding the 1050 mm 2 migration area.
Cells were seeded in 6-well plates at a density of 100 cells/well. After 24h, cells were treated with A-485 at concentrations from 10 nM to 100 μM. After 6 days, cells were stained with crystal violet and counted.
Cells were seeded at a density of 4,000 cells per well in a 96-well plate. After 24 hr, cells were transfected as described above. After 24 hr, cells were treated with drugs for 48–72 hr. Cells were incubated with 2 μg/mL calcein-AM for 1 hr and fluorescence was measured using a SpectraMax i3× (Molecular Devices).
The Caspase-Glo 3/7 Assay (Promega) was used according to the manufacturer’s instructions. Following transfection (48 hr) and drug treatment (24 hr) cells were seeded at 10,000 cells per well in a 96-well Cellstar Cell-Repellent plate (Greiner Bio-one) in serum-free DMEM/F12, 1% L-glutamine, 1% P/S containing A-485 or vehicle. After 24 hr, cells were treated with Caspase-Glo at a ratio of 1:1 and incubated at 37°C for 1 hr. Cells were then transferred to a white 96-well plate (costar) and luminescence was measured using a SpectraMax i3x (Molecular Devices).
Expression of Annexin V was measured by flow cytometry using the Annexin V-FITC Kit (Miltenyi Biotec) according to the manufacturer’s instructions. Flow cytometry was performed using a BD Accuri C6 flow cytometer (BD Biosciences) and analyzed using FlowJo v10 software (BD Biosciences).
The Click-iT Plus EdU Flow cytometry Assay Kit (Invitrogen) was used for cell cycle assays. Cells were treated with 10 μM of EdU for 2 hours in culture media. Cells were harvested by trypsinization and washed in 1% BSA in PBS. Cells were resuspended in 100 μL of ice cold PBS, and 900 μL of ice cold 70% ethanol was added dropwise while vortexing. Cells were incubated on ice for two hours. Cells were washed with 1% BSA in PBS and then treated with the Click-iT Plus reaction cocktail including Alexa Fluor 488 picolyl azide according to the manufacturer’s instructions for 30 min. Cells were washed with 1X Click-iT permeabilization buffer and wash reagent, and then treated with 5 μM of Vybrant Dye Cycle Ruby Stain (ThermoFisher) diluted in 1% BSA in PBS for 30 min at 37°C. Flow cytometry was performed using a BD Accuri C6 flow cytometer (BD Biosciences) and analyzed using FlowJo v10 software (BD Biosciences).
RNA samples were collected 72 hr following siRNA transfection using the Quick-RNA Miniprep Kit (Zymo Research). Libraries were prepared by the VARI Genomics Core from 500 ng of total RNA using the KAPA mRNA HyperPrep kit (v4.17) (Kapa Biosystems). RNA was sheared to 300–400 bp. Prior to PCR amplification, cDNA fragments were ligated to IDT for Illumina unique dual adapters (IDT DNA Inc). Quality and quantity of the finished libraries were assessed using a combination of Agilent DNA High Sensitivity chip (Agilent Technologies), QuantiFluor® dsDNA System (Promega), and Kapa Illumina Library Quantification qPCR assays (Kapa Biosystems). Individually indexed libraries were pooled and 100 bp, single end sequencing was performed on an Illumina NovaSeq6000 sequencer using an SP, 100 cycle sequencing kit (Illumina) and each library was sequenced to an average raw depth of 35M reads. Base calling was done by Illumina RTA3 and output of NCS was demultiplexed and converted to FastQ format with Illumina Bcl2fastq v1.9.0.
12Z cells were crosslinked 72 hr post-transduction with lentiviral particles containing control shRNAs or ARID1A-targeting shRNAs (differential P300, H3K18ac, H3K27ac, H3K27me3, H3K4me3 and H3K4me1 ChIP-seq) or untreated cells were used (wild-type P300 ChIP-seq). For crosslinking, cells were treated with 1% formaldehyde in cell culture media for 15 min at room temp. Formaldehyde was quenched by the addition of 0.125 M Glycine, and cells were washed with PBS. 4×10 6 crosslinked cells were used per IP for H3K4me3 and H3K27me3, and 1×10 7 crosslinked cells were used per IP for all other antibodies. Chromatin from crosslinked cells was fractionated by digestion with micrococcal nuclease using the SimpleChIP Enzymatic Chromatin IP Kit (Cell Signaling) per the manufacturers’ instructions. IPs were performed in duplicate per antibody and condition by adapting established methods ( Boyd and Farnham, 1997 ).
For P300 IPs, nuclei were resuspended in nuclear lysis buffer (50 mM Tris-HCl [pH 8.0], 10 mM EDTA [pH 8.0], 1% SDS) and sonicated for 30 s. Protein G magnetic beads (Cell Signaling) were preconjugated with antibody overnight at 4°C in wash buffer (1X PBS, 0.5% BSA, 0.02% Tween-20). Antibody used was 5 μg P300 (sc-32244, Santa Cruz). Fractionated chromatin was diluted into IP buffer (0.01%SDS, 1.1% TrionX-100, 1.2 mM EDTA [pH 8.0], 16.7 mM Tris-HCl [pH 8.0], 167 mM NaCl) and incubated with preconjugated antibody/Dynabeads overnight at 4°C. Samples were washed at 4°C with high-salt buffer (0.1% SDS, 1% Triton X-100, 2 mM EDTA [pH 8.0], 20 mM Tris-HCl [pH 8.0], 0.5 M NaCl), low-salt buffer (0.1% SDS, 1% Triton X-100, 2 mM EDTA [pH 8.0], 20 mM Tris-HCl [pH 8.0], 150 mM NaCl), dialysis buffer (0.2% Sarcosyl, 2 mM EDTA [pH 8.0], 50mM Tris-HCl [pH 8.0]), IP wash buffer (0.25M LiCl, 1% NP-40, 1% Deoxycholate, 1 mM EDTA [pH 8.0], 10 mM Tris-HCl [pH 8.0]) and TE (10mM Tris-HCL [pH 8.0], 1mM EDTA [pH 8.0]). IP chromatin was eluted for 30 min at 37°C with elution buffer (1% SDS, 0.1 M NaHCO3). Crosslinks were reversed with 0.4 mg/mL Proteinase K (ThermoFisher) and 0.2 M NaCl at 65°C for 2 hr. DNA was purified using the ChIP DNA Clean & Concentrator Kit (Zymo).
For H3K27ac, H3K18ac, H3K27me3, H3K4me3 and H3K4me1 IPs were performed using the SimpleChIP Enzymatic Chromatin IP Kit per the manufacturers’ instructions. For H3K27ac and H3K18ac the addition of 5 mM sodium butyrate included in Buffer A and ChIP Buffer. Antibodies used were 10 μg H3K27ac (Active Motif, 39133), 5 μg H3K18ac (ab1191, Abcam), 10 μL H3K27me3 (Cell Signaling, 9733), 10 μL H3K4me3 (Cell Signaling, 9751), or 4 μg H3K4me1 (Abcam, ab8895) per IP. DNA was purified as described above.
Libraries for Input and IP samples were prepared by the VARI Genomics Core from 10 ng of input and IP material when available, and all material when less than 10 ng available, using the KAPA Hyper Prep Kit (v5.16) (Kapa Biosystems). Prior to PCR amplification, end repaired and A-tailed DNA fragments were ligated to Bioo Scientific NEXTflex Adapters (Bioo Scientific). Quality and quantity of the finished libraries were assessed using a combination of Agilent DNA High Sensitivity chip (Agilent Technologies), QuantiFluor® dsDNA System (Promega), and Kapa Illumina Library Quantification qPCR assays (Kapa Biosystems). Individually indexed libraries were pooled. For P300 ChIP in wild-type cells, 100 bp, single end sequencing was performed on an Illumina NovaSeq6000 sequencer using an SP, 100 cycle sequencing kit (Illumina) and each library was sequenced to minimum read depth of 100M reads per input library, 50M reads per IP library. Base calling was done by Illumina NextSeq Control Software (NCS) v2.0. For differential P300 ChIP, 75 bp, paired end sequencing was performed on an Illumina NextSeq 500 sequencer using 150 cycle HO and MO sequencing kits (v2) (Illumina), with all libraries run across 2 flowcells to return a minimum read depth of 80M reads per input library and 40M read per IP library. Base calling was done by NCS v2.0. For differential H3K27ac, H3K18ac and H3K4me1 ChIP-seq IPs, 50 bp, paired end sequencing was performed on an Illumina NovaSeq6000 sequencer using an S1, 100 cycle sequencing kit and each library was sequenced to minimum read depth of 50M reads per IP library. Input samples were sequenced using 100 bp, single end sequencing to a minimum read depth of 100M reads. Base calling was done by Illumina RTA3. For H3K27me3 and H3K4me3 ChIP, 75 bp, single end sequencing was performed on an Illumina NextSeq 500 sequencer using 75 cycle HO sequencing kits (v2), with all libraries run across two flow cells to return a minimum read depth of 80 M reads per input library and 40 M read per IP library. Base calling was done by Illumina NextSeq Control Software (NCS) v2.0. For all experiments, output data was demultiplexed and converted to FastQ format with Illumina Bcl2fastq v1.9.0.
The CUT&RUN protocol was adapted from established protocols ( Skene et al., 2018 ). BioMag Plus Concanavalin A–coated magnetic beads (Bangs Laboratories) were washed in Binding Buffer (20 mM HEPES-KOH pH 7.9, 10 mM KCl, 1 mM CaCl2, 1 mM MnCl2). 72 hr following siRNA transfection, 500,000 12Z cells were harvested and resuspended in Wash Buffer (20 mM HEPES-NaOH pH 7.5, 150 mM NaCl, 0.5 mM Spermidine, 1X protease inhibitor cocktail) and washed twice by centrifuge at 600 × g for 3 min, and then added to the concanavalin A bead suspension and mixed on a tube rotator for 10 min at room temp. Cell/bead conjugates were resuspended in 500 μL of Antibody Buffer (Wash Buffer with 0.05% Digitonin and 2 mM EDTA) containing 5 μg of H3K27ac antibody (Active Motif, cat# 39133) or Rabbit IgG (Cell Signaling, cat#2729) and incubated in a tube nutator overnight at 4°C. The following day, cells were washed in Digitonin Buffer (Wash Buffer with 0.05% Digitonin) three times, resuspended in 250 μL of Digitonin Buffer and 12.5 μL of CUTANA pAG-MNase (EpiCypher, cat# 15–1016) was added. Cells were mixed on a nutator at room temp for 1 hr, followed by two washes in Digitonin Buffer and one wash with Low-Salt Rinse Buffer (20 mM HEPES-NaOH pH 7.5, 0.5 mM Spermidine, 1X protease inhibitor cocktail). Tubes were chilled on ice, 1 mL of Calcium Incubation Buffer (3.5 mM HEPES-NaOH pH 7.5, 10 mM CaCl2, 0.05% Digitonin) was added, and tubes were nutated at 4°C. After 2.5 min, beads were bound to magnet, supernatant was removed and 250 μL of EGTA-STOP Buffer (170 mM NaCl, 20 mM EGTA, 0.05% Digitonin, 20 μg/mL RNase A, 20 μg/mL Glycogen, 0.8 pg/ml S. cerevisiae fragmented nucleosomal DNA) was added. Beads were nutated at 37°C for 30 min, followed by centrifuge at 16,000 × g for 5 min at 4°C. DNA was purified using the NucleoSpin Gel and PCR Clean-up Kit (Takara, cat# 740609.50).
Libraries for CUT&RUN samples were prepared by the Van Andel Genomics Core from 0.5–1 ng of IP material, using the KAPA Hyper Prep Kit (v5.16) (Kapa Biosystems). Prior to PCR amplification, end-repaired and A-tailed DNA fragments were ligated to Bioo Scientific NEXTflex Adapters (Bioo Scientific) at a concentration of 500 nM. Quality and quantity of the finished libraries were assessed using a combination of Agilent DNA High Sensitivity chip (Agilent Technologies, Inc.), QuantiFluor® dsDNA System (Promega Corp.), and Kapa Illumina Library Quantification qPCR assays (Kapa Biosystems). Individually indexed libraries were pooled and 50 bp, paired end sequencing was performed on an Illumina NovaSeq6000 sequencer using an S1, 100 cycle sequencing kit (Illumina Inc.) Each library was sequenced to an average depth of 75M reads. Base calling was done by Illumina RTA3 and output was demultiplexed and converted to FastQ format with Illumina Bcl2fastq v1.9.0.
For standard mRNA gene-level expression analysis, single-end raw reads were trimmed with cutadapt ( Martin, 2011 ) and Trim Galore! ( http://www.bioinformatics.babraham.ac.uk/projects/trim_galore/ ) followed by quality control analysis via FastQC ( Andrews, 2010 ) and MultiQC ( Ewels et al., 2016 ). Trimmed reads were aligned to GRCh38.p12 genome assembly and indexed to GENCODE ( Frankish et al., 2019 ) v28 GFF3 annotation via STAR ( Dobin et al., 2013 ) aligner with flag ‘–quantMode GeneCounts’ for feature counting. Reverse-stranded, gene-level counts were extracted from the STAR output files and constructed into an experimental read count matrix in R. Low count genes were filtered (1 count per sample on average) prior to DESeq2 ( Love et al., 2014 ) count normalization and differential expression analysis. Modeling design matrices were constructed with a single “condition” variable and included an intercept. Calculated differential expression probabilities were corrected for multiple testing by independent hypothesis weighting ( IHW ) ( Ignatiadis et al., 2016 ) for downstream analysis. Threshold for differential expression significance was set at FDR < 0.0001. Relative expression heatmaps were produced using relative regularized-logarithm (rlog) ( Love et al., 2014 ) counts by subtracting mean rlog counts of the control group. Relative linear expression bar plots were produced from DESeq2 normalized counts table. Previously published GEMM expression data ( Wilson et al., 2019 ) were extracted from (GEO: GSE129784 ).
Intergenic eRNA and associated differential expression analysis were also analyzed similarly. Briefly, previously published, paired end total RNA sequencing data from 12Z cells treated with siARID1A or non-targeting siRNA control ( Wilson et al., 2019 ) were extracted from (GEO: GSE129782 ). Reads were trimmed and aligned as described above. Aligned BAMs were inputted to HOMER ( Heinz et al., 2010 ) in order to count integer RNA-seq reads at each of the 18,050 distal, putatively active enhancer elements described in this study. Counted regions were then excluded which overlapped with any genic regions, including introns, using the genes() function of TxDb.Hsapiens.UCSC.hg38.knownGene (Bioconductor Core Team and Bioconductor Package Maintainer, 2016) R package. Expression status of each eRNA locus was then determined by observation of at least 1 count per sample on average, resulting in 3,668 expressed intergenic eRNAs. The filtered eRNA counts table was then normalized and modeled for differential expression analysis by DESeq2 ( Love et al., 2014 ) as described above.
Wild-type P300 and differential H3K27ac, H3K18ac, H3K27me3, and H3K4me3 ChIP-seq experiments were analyzed as single-end libraries, while differential P300 and H3K4me1 ChIP-seq were analyzed as paired-end. Raw reads for IPs and inputs were trimmed with cutadapt ( Martin, 2011 ) and Trim Galore! followed by quality control analysis via FastQC ( Andrews, 2010 ) and MultiQC ( Ewels et al., 2016 ). Trimmed reads were aligned to GRCh38.p12 reference genome via Bowtie2 ( Langmead and Salzberg, 2012 ) with flag ‘–very-sensitive’. Aligned reads were sorted and indexed with samtools ( Li et al., 2009 ). For paired-end analyses, only properly-paired read fragments were retained by samtools view with flag ‘-f 3’ followed by sorting and indexing. Specifically for libraries with differential comparisons, molecular complexity was then estimated from duplicate rates by ATACseqQC ( Ou et al., 2018 ) and preseqR ( Daley and Smith, 2013 ), and libraries were subsampled to equivalent molecular complexity within an experimental design based on these estimates with samtools . Picard MarkDuplicates ( http://broadinstitute.github.io/picard/ ) was used to remove PCR duplicates, followed by sorting and indexing. MACS2 ( Zhang et al., 2008 ) was used to call peaks on each ChIP replicate against the respective input control. For P300 IPs, MACS2 called broadPeaks with FDR < 0.05 threshold and otherwise default settings. For H3K4me3, H3K4me1, H3K27ac and H3K18ac IPs, MACS2 called narrowPeaks with FDR < 0.05 threshold and flags ‘–nomodel–extsize 146’ to bypass model building. For H3K27me3, MACS2 called broadPeaks with FDR < 0.05 threshold and flags ‘–nomodel–extsize 146’ to bypass model building. The resulting peaks were repeat-masked by ENCODE blacklist filtering and filtered for non-standard contigs ( Amemiya et al., 2019 ). A naive overlapping peak set, as defined by ENCODE, was constructed by calling peaks on pooled replicates followed by bedtools intersect ( Quinlan and Hall, 2010 ) to select for peaks of at least 50% overlap with each biological replicate.
ChIP-seq differential binding or abundance analysis was performed with csaw ( Lun and Smyth, 2016 ). Briefly, a consensus peak set was constructed for each differential experiment from the union of replicate-intersecting, filtered MACS2 peaks called in each condition. The replicate intersection criteria used here are less stringent than the naive overlap; any partial intersect between ChIP replicates was accepted as a query region tested for differential binding/abundance. ChIP reads were counted in these query regions by csaw , then filtered for low abundance peaks with average logCPM > −3. When comparing ChIP libraries, any global differences in IP efficiency observed between the two conditions were considered a result of technical bias to ensure a highly conservative analysis. As such, we employed a non-linear loess-based normalization to the peak count matrix, as is implemented in csaw ( Lun and Smyth, 2016 ), to assume a symmetrical MA distribution. A design matrix was then constructed from one “condition” variable, without an intercept. The count matrix and loess offsets were then supplied to edgeR ( Robinson et al., 2010 ) for estimating dispersions and fitting the quasi-likelihood generalized linear model for hypothesis testing. Nearby query regions were then merged up to 500 bp apart for a maximum merged region width of 5 kb, and the most significant probability was used to represent the merged region. An FDR < 0.05 threshold was used to define significant differentially bound/abundant regions.
Analysis of CUT&RUN data followed a highly similar procedure as paired-end ChIP-seq. Briefly, raw paired-end reads for H3K27ac or IgG CUT&RUN were trimmed and aligned, filtered for only properly-paired reads, then molecular complexity was estimated, and libraries were subsampled to equalize based on complexity estimates. PCR duplicates were removed, and MACS2 was used to call narrowPeaks against the IgG negative control as input, with FDR < 0.05 threshold and flags ‘–nomodel–extsize 146’ to bypass model building. Peaks were then blacklist-filtered, and a naive overlapping peak set was constructed as described above.
Differential H3K27ac CUT&RUN analysis was computed with csaw . In order to promote similarity between the differential H3K27ac CUT&RUN and ChIP-seq experiments, we used the same consensus peak set as defined by our differential H3K27ac ChIP-seq experiment for the CUT&RUN analysis here. Briefly, H3K27ac CUT&RUN reads were counted in these query regions by csaw , then filtered for low abundance peaks with average logCPM > −3. When comparing CUT&RUN libraries, any global differences in CUT&RUN reaction efficiency observed between two conditions were considered a result of technical bias to ensure a highly conservative analysis. As such, we employed a non-linear loess-based normalization to the peak count matrix, as is implemented in csaw, to assume a symmetrical MA distribution. A design matrix was then constructed from one “condition” variable, without an intercept. The count matrix and loess offsets were then supplied to edgeR for estimating dispersions and fitting the quasi-likelihood generalized linear model for hypothesis testing. Nearby query regions were then merged up to 500 bp apart for a maximum merged region width of 5 kilobases, and the most significant probability was used to represent the merged region.
ChromHMM ( Ernst and Kellis, 2017 ) was used to segment the hg38 genome based on combinatorial chromatin features in control and ARID1A-depleted 12Z cell conditions. Briefly, all seven chromatin features (total RNA, ATAC, H3K4me1, H3K4me3, H3K27me3, H3K27ac, and H3K18ac) were binarized from aligned BAM files, and chromatin features were modeled in both conditions simultaneously through the “concatenated” option. The concatenated option was selected because it creates a unified model for direct comparison between control and ARID1A-depleted conditions used to identify chromatin state changes. Chromatin state models were built from 5 to 25 states, and each model was manually curated based on inferred biological function to select one with balance between unique and overlapping combinatorial features. We selected 18 states as our final model for downstream analysis. State emissions were then user-reordered to group based on inferred biological function. BED files containing coordinates for each chromatin state in each condition were constructed into non-overlapping GenomicRanges objects in R for downstream enrichment analyses, differential chromatin state analysis, and plotting.
For RNA-seq experiments, three biological replicates were analyzed (n = 3). For ChIP-seq and CUT&RUN experiments two independent IPs were used (n = 2) and were compared against a condition-respective input chromatin sample or IgG negative control, respectively. For in vivo experiments, n represents number of mice. For cell-based assays, n represents biological replicates or independent experiments as indicated in the figure legend. Multiple hypothesis tests corrections via FDR were employed when appropriate to reduce type I errors. Presented probability ( p ) values are representative of the associated statistical tests as indicated in the figure legends. All boxplots presented for genomic analyses are in the style of Tukey without outliers. The ROSE algorithm ( Lovén et al., 2013 ; Whyte et al., 2013 ) was used to define active super-enhancers from H3K27ac peaks which overlapped with accessibility (ATAC) in control 12Z cells. GeneHancer ( Fishilevich et al., 2017 ) database was used to associate enhancers to genes with a GeneHancer score > 1 threshold. Various HOMER ( Heinz et al., 2010 ) functions were applied to annotate genomic regions of interest, quantify signal and count reads at sites of interest for tag density heatmaps and meta peak plots. Chromatin analyses involving ChIP signal quantification at regions of interest used pooled reads from both IP replicates, per feature. TxDb.Hsapiens.UCSC.hg38.-knownGene ( Bioconductor Core Team and Bioconductor Package Maintainer, 2016 ) was used to define gene promoters for all standard hg38 genes as 3 kilobase regions surrounding the primary TSS. MACS2 ( Zhang et al., 2008 ) was used to produce genome-wide signal log-likelihood ratio (logLR) tracks for IGV ( Robinson et al., 2011 ) visualization. ClusterProfiler ( Yu et al., 2012 ) was used to compute and visualize pathway enrichment from a list of gene symbols with respective gene universes. Hallmark pathways and GO Biological Process gene sets were retrieved from MSigDB ( Liberzon et al., 2015 ). ComplexHeatmap ( Gu et al., 2016 ) was used for hierarchical clustering by Euclidean distance and general heatmap visualization. GenomicRanges ( Lawrence et al., 2013 ) functions were frequently used to intersect and manipulate genomic coordinates e.g., for genome-wide association tests. eulerr ( Larsson, 2020 ) was used to produce proportional Euler diagrams. biomaRt ( Durinck et al., 2005 ; Durinck et al., 2009 ) was used for all gene nomenclature and ortholog conversions. ggplot2 ( Wickham, 2016 ) was used for certain plotting applications. The statistical computing language R ( R Core Team, 2018 ) was used for many applications throughout this manuscript. Mantel-Cox tests and t tests were performed using GraphPad Prism 8 software.