{"paper_id":"18c4c47e-3c0d-438f-b332-92aa671364a7","body_text":"Ovarian cancer has the  highest fatality-to-case ratio of all gynecologic\nmalignancies [ 1 ,  2 ].  This is attributed to the lack of early\nwarning signs and efficacious early detection techniques\n[ 1 ,  3 ]. Another problem hindering the successful management\nof the disease is the paucity in prognosticators that could\nassist the selection of treatment modality. One of the most\npromising routes towards improvement in the detection and\nsurveillance of ovarian cancer is the identification of serum\nmarkers. Utilization of the CA125 as an ovarian cancer serum\nmarker has improved cancer detection rates during the last few\nyears [ 1 ,  2 ,  3 ]. Nevertheless, CA125 does not diagnose\nearly-stage cancers with high accuracy and is prone to false\npositives. Therefore, the need to identify additional serum\nmarkers for ovarian cancer is paramount to the successful\nmanagement of this disease.\nA major obstacle in finding a diagnostic biomarker is the tremendous molecular\nheterogeneity that exists for nearly all human cancer, suggesting that simultaneous\nscreening of a patient specimen for multiple biomarkers will be required to improve the\nearly detection/diagnosis of cancer.  DNA chip technologies address this problem at the\ngenomic level, and provide accessibility to gene expression profiles. However, since\nproteins are, for the most part,  the mediators of  a cell's function, the study of the\nchanges in proteins that result from a pathological lesion, such as cancer, would appear to\nbe a rich source of potential cancer biomarkers.\nMost of the previous studies in search of diagnostic biomarkers\nhave employed two-dimensional electrophoresis (2DE) which can\nresolve hundreds to thousands of proteins present in complex\nprotein mixtures, such as cell lysates and body fluids. Although\nsome successes have been reported in detecting potential ovarian\ncancer-associated biomarkers [ 4 ,  5 ,  6 ,  7 ],  this classical\nproteomic technique is very time consuming, not highly\nreproducible, and  not easily adaptable to a clinical assay\nformat.\nA  recently developed mass spectrometry proteomic approach, the\nSELDI (surface-enhanced laser desorption/ionization) ProteinChip\nSystem (Ciphergen Biosystems, Inc, Fremont,\nCalif), appears to hold  promise for biomarker discovery and as\na potential clinical assay format [ 8 ,  9 ]. (The\nSELDI system and its applications are described in\nthe report by Reddy and Dalmasso [ 10 ]; and a\nrecent review by Wright [ 11 ]). Using this system,\n distinct protein patterns of normal,\npremalignant, and malignant cells were found  for ovarian,\nesophageal, prostate, breast, and hepatic cancers\n[ 12 ,  13 ,  14 ]. Potential biomarkers for breast and bladder\ncancers were also detected in nipple aspirate fluid and urine, see respectively [ 15 ,  16 ], by the SELDI\nsystem.\nRecent reports also support that analysis of the SELDI data by\n“artificial intelligence”  algorithms can lead to the\nidentification of protein  “fingerprints”  specific for\nprostate, ovarian, and breast cancers, significantly increasing\nthe accuracy in differentiating cancer from the noncancer groups\n[ 17 ,  18 ,  19 ,  20 ].  These studies employed different algorithms\nto analyze the SELDI data, including a genetic algorithm\n[ 19 ], a decision tree [ 17 ,  18 ], and a support vector\nmachine algorithm [ 20 ]. Each method appeared to be\neffective in developing accurate classification systems.\nThe high dimensionality of the data generated by SELDI requires a\nmathematical algorithm to analyze the data without overfitting.\nSince the SELDI protein profiling approach is new, it is\ndifficult to determine up-front which algorithm to select for the\ndata analysis and development of a  “diagnostic”  classifier.\nIt is also fair to assume that different bioinformatic tools may\nbe required for different cancer or disease systems. The\nobjective of this study was to evaluate the commercially\navailable classification algorithm (biomarker pattern software\n[BPS]) developed by Ciphergen Biosystems Inc for analysis of the\nSELDI serum protein profiling data from patients with ovarian\ncancer, benign pelvic diseases, and normal women.  The potential,\nadvantages, and drawbacks of this approach as well as suggestions\nfor improvement are discussed.\nDemographics of the cancer and control groups included in the study.\n\nSerum samples were obtained from patients with epithelial ovarian cancer prior to\ntreatment administration\n( n  = 44), benign pelvic diseases ( n  = 61), and from women with no\nevidence of pelvic disease ( n  = 34) enrolled through the Division of Gynecologic\nOncology, University of Texas, Southwestern Medical Center. Informed consent was\nobtained from all patient and control groups. The demographics of the patients and the\nstage distribution of the ovarian cancers are presented in  Table 1 . Benign conditions\nincluded benign pelvic masses (endometriosis, cystadenomas, hydrosalpinx, lipoma,\nBrenner tumor, fibroids, endometrial polyp). The sera were aliquoted and stored at \n−80 ° C.\nSerum samples were applied on the strong anion exchange (SAX) and\nimmobilized-copper (IMAC) chip surfaces. In brief, 21  μ L\nof serum  were mixed with 30  μ L 8M\nurea in 1% CHAPS-PBS pH 7.4 buffer for 30 minutes\nat 4 ° C, followed by the addition of \n100  μ L of 1M\nurea in 0.125% CHAPS-PBS buffer and\n600  μ L of binding buffer compatible with the type of\nsurface in use (PBS for IMAC and 20 mM Hepes\ncontaining 0.1% Triton for SAX).\nFifty  μ L of the diluted\nsamples were then applied onto the chips using a bioprocessor.\nFollowing a 30-minute incubation, nonspecifically bound molecules\nwere removed by 3 brief washes in binding buffer followed by 3\nwashes with HPLC-gradient H 2 O. Sinapinic acid (2X\n1  μ L of 50% SPA in 50% ACN-0.1%TFA) was applied to the\nchip array surface and mass spectrometry was\nperformed using a PBS2 SELDI mass spectrometer (Ciphergen\nBiosystems Inc). Protein data were collected by averaging a total\nof 192 laser shots.  Mass calibration was performed using the\nall-in-one peptide standard (Ciphergen Biosystems Inc) which\ncontains vasopressin (1084.2 daltons), somatostatin (1637.9\ndaltons), bovine insulin   β -chain\n(3495.9 daltons), human insulin recombinant (5807.6 daltons), and\nhirudin (7033.6 daltons). All samples were processed in duplicate.\nProtein spectra of one serum sample processed on the\nIMAC metal binding chip array and on the positively charged SAX\nchip array. Note that several different  proteins are captured by\nthe two different chip chemistries.\nProtein peaks were labeled and their intensities were normalized for total ion current (mass\nrange 2–200 kd) to account for variation in ionization\nefficiencies, using the SELDI software (version 3.1). Peak\nclustering was performed using the Biomarker Wizard software\n(Ciphergen Biosystems) and the following specific settings:\nspectral data from IMAC surface; signal/noise (first pass): 4,\nminimum peak threshold: 10%, mass error: 0.3%, and signal/noise\n(second pass): 2 for the 2–20 kd mass range and signal/noise\n(first pass): 5, minimum peak threshold: 10%, mass error: 0.3%,\nand signal/noise (second pass): 2.5 for the 20–100 kd mass\nrange.  Spectral data from the SAX surface were analyzed with the\nsame set of settings with the difference that the minimum peak\nthreshold was set to 5%. With these labeling parameters, a total\nof 122 protein clusters (45 from the IMAC and 77 from the SAX\nsurface) were generated. Peak mass and intensity were exported to\nan excel file, and the peak intensities from each duplicate\nspectra were averaged. Pattern recognition and sample\nclassification \n were performed using the BPS. The\ndecision tree described in the result section was generated using\nthe Gini method nonlinear combinations. A 10-fold cross-validation\nanalysis was performed as an  initial evaluation of the test error\nof the algorithm. Briefly, this process involves splitting up the\ndataset into 10 random segments and using 9 of them for training\nand the 10th as a test set for the algorithm. Multiple trees were\ninitially generated from the 122 classifiers by varying the\nsplitting factor by increments of 0.1. These trees were evaluated\nby cross-validation analysis. The peaks that formed the main\nsplitters of the tree with the highest prediction rates were then\nselected, the tree was rebuilt based on these peaks alone and\nevaluated by the test set. The values of  P  were calculated\nbased on  t -test (Biomarker Wizard software). The value P < .05\nwas considered to be statistically significant.\nDecision tree classification of the ovarian\ncancer (C) and noncancer (normal and benign or B) groups. The\nblue boxes show the decision nodes with the  peak mass\n(M in kd), the peak\nintensity (I) cutoff levels, and the number of samples. The 5.54,\n6.65, and 11.7 kd masses were detected on the IMAC chip, and\nthe 4.4 and 21.5 kd on the SAX chip. These  five masses form\nthe splitting rules. Cases that follow the rule are placed in the\nleft daughter node. The red boxes are the\nterminal  nodes with the classification being either cancer or\nbenign (normal + benign).\n\nOne hundred \nthirty-nine serum samples were assayed by SELDI mass\nspectrometry.  Both SAX and IMAC surfaces could effectively\nresolve low-mass (< 20 kd) protein peaks, although the SAX\nsurface appeared superior in resolving larger (> 20 kd)\nprotein peaks.  Figure 1  shows representative protein\nspectra from one serum sample processed on SAX and IMAC chips.\nOf a total of 139 serum samples, 124 (85 controls and 39 cancers)\nwere randomly selected to form the learning set and 15 (10\ncontrols and 5 cancers) to form the blinded test set for the\nalgorithm.  Five peaks were selected by the BPS algorithm to\ndiscriminate cancer from the noncancer groups.\n Figure 2  is the decision tree that was generated from\n the learning set to\nclassify the two groups. Three peaks (5.54, 6.65, and\n11.7 kd) detected on the IMAC chip and 2 (4.4 and\n21.5 kd) detected on the SAX surface form the main\nsplitters.  Their mass spectra and gray-scale/gel views are shown\nin  Figures  3 ,  4 ,  5 ,  6 , and  7 . These peaks have significantly\ndifferent intensity levels between the cancer and benign or\nnormal controls with the exception of the 6.65 and 21.5 kd\npeaks, which did not differ significantly between cancers and\nbenigns ( Table 2 ).  A 10-fold cross-validation\nanalysis was performed as an initial evaluation of the accuracy\nof the algorithm in predicting ovarian cancer. A specificity of\n80% and sensitivity of 84.6% were obtained\n( Table 3 ). In the test set, sensitivity and\nspecificity of 80% were  obtained\n( Table 3 ). The misclassified samples in the test set\nincluded one benign (uterine fibroid), one normal, and a stage\nIII C cancer.\nStatistical comparison of the intensity levels of the\npeaks used in the decision tree between the cancer and control\ngroups. C-N: cancer versus normal; C-B: cancer versus benign;\nand C-B/N: cancer versus normal and benign.\nPerformance of the decision tree in predicting ovarian\ncancer. Numbers in parentheses denote the number of correctly\nclassified sample out of total number of samples in the\ngroup.\n\nThe high degree of genetic\nheterogeneity associated with human cancers makes it likely that\npanels of multiple biomarkers will be needed to improve early\ndetection/diagnosis.  This entails the development of\nhigh-throughput proteomic and genetic approaches as well as of\nreliable bioinformatic tools for data analysis.\nThe SELDI proteinChip system offers the advantage of rapid and\nsimultaneous detection of multiple proteins from complex biologic\nmixtures.  We employed this system in combination with the BPS\nclassification algorithm for protein profiling of ovarian cancer\nin serum.  Using this approach, a classifier that was 80%\naccurate in discriminating patients with ovarian cancer from\npatients with benign disease and healthy controls from a blinded\ntest set was generated.  Evaluation of the classifier by\ncross-validation and the analysis of the independent test set\noffers statistical confidence of the potential of this approach\nas an ovarian cancer detection tool.  However, the sample size\nincluded in this study decreases the validity of generalized\nconclusions.  Complete evaluation of this classifier will require\ntesting its prediction rates for larger “blinded”\n and independent serum sets.\nThe BPS software was found to be relatively simple to use. However,\nBPS, like other mathematical algorithms, is prone to data\noverfitting, and also is not reliable when a large number of\nvariables relative to samples sizes are included in the analysis.\nA preselection process of the most significant variables using\nstatistical analysis (eg, ROC curve, ANOVA) may help in\nalleviating this problem.\nSpectra (top) and grey-scale or gel views\n(bottom) of the peaks (arrows) forming the splitting rules. The\nprotein peak was detected on IMAC chip. The peak appears to be\nupregulated in the cancer (C1 C4) compared to the benign (B1-B2)\nand normal (N1-N2) groups.\nSpectra (top) and grey-scale or gel views (bottom) of\nthe peaks (arrows) forming the splitting rules. The protein peak\nwas detected on IMAC chip. The peak appears to be downregulated in\nthe cancers.\nSpectra (top) and grey-scale or gel views (bottom) of\nthe peaks (arrows) forming the splitting rules. The protein peak\nwas detected on IMAC chip. The peak appears to be upregulated in\ncancer (C1–C4) compared to the benign (B1-B2) and normal (N1-N2)\ngroups.\nSpectra (top) and grey-scale or gel views (bottom) of\nthe peaks (arrows) forming the splitting rules. The protein peak\nwas detected on the SAX surface. The peak appears to be\nupregulated in the cancer (C1–C4) compared to the begin (B1-B2)\nand normal (N1-N2) groups.\nSpectra (top) and grey-scale or gel views (bottom) of\nthe peaks (arrows) forming the splitting rules. The protein peak\nwas detected on the SAX surface. The peak appears to be\ndownregulated in the cancers.\nPetricoin et al [ 19 ] recently reported the successful\napplication of a genetic algorithm for the analysis of SELDI\nproteomic data from ovarian cancer patients. In this study, five\ndiscriminatory peptides were detected, moleculalr mass range\n500–2500 daltons, and the accuracy in predicting ovarian cancer\nin a blinded set of samples was 97.4%.  We focused on the\nanalysis of potential biomarkers in higher mass ranges (> 2000\ndaltons). Furthermore, in contrast to the case where BPS algorithm\nis processed, that is,  labeled peak information is analyzed,\nthe  genetic algorithm employed by Petricoin et\nal analyzes time-of-flight “raw”  SELDI data.  In this case,\nprerequisite for the further identification of the potential\ndiscriminatory markers is the coupling of the genetic algorithm\nwith a peak identification system where the raw data are\ntranslated into protein peak information.  BPS employs the peak\nidentification system of the SELDI software facilitating\nbiomarker detection. It should be noted, however, that careful\nand precise selection of the peak labeling settings and\nnormalization of peak intensities are considered critical for\nbiomarker identification and for the  efficient and\nreliable performance of any learning algorithm used in\nconjunction with the SELDI system.\nBesides providing a preliminary evaluation of the suitability of\nBPS for the comparison of SELDI data, our study also demonstrates\nthe potential of combining spectral data from different types of\nsurfaces as a means to increase protein resolution. Although,\ncompared to SELDI, the resolving power of 2D gel electrophoresis\nremains unchallenged, we have found that this combinatorial\napproach can significantly enhance biomarker discovery and\nincrease test accuracy for ovarian and breast cancers from\n70–75% up to 90%  [ 21 ].\nIn conclusion, the BPS software appears to be potentially\nsuitable for analysis of the high-dimensional SELDI spectral data.\nAvenues for improvement of the algorithm performance include\noptimization of the peak labeling process as well as preselection\nof the most significant peaks by statistical approaches.  More\nextended studies will be required to validate the potential and\nreliability of BPS as a bioinformatic tool for proteomic studies.\nIt should also be emphasized that comparative analysis of\ndifferent types of algorithms will be of paramount importance for\nthe better evaluation of their performance and the selection of\nthe bioinformatic features needed for effective biomarker\ndiscovery and discrimination of cancer.","source_license":"CC-BY-4.0","license_restricted":false}