Showing posts with label microarray. Show all posts
Showing posts with label microarray. Show all posts

Monday, October 1, 2012

My DREAM Model for Predicting Breast Cancer Survival

This summer, I have worked on submitting a few models to the DREAM competition for predicting breast cancer survival.

Although I was originally planning on posting about my model after the competition was completely finished, I decided to go ahead and describe my experience because 1) my model honestly didn't radically differ from the example model and 2) I don't think I have enough time to redo the whole model building process on the new data before the 10/15 deadline.

To be clear, the performance isn't all that different for the old and new data, but there are technical details that would have to be worked out to submit the models (and I would want to take time to re-examine the best clinical variables to include in the model).  For example, here are the concordance index values for my three models on the training dataset:

 
 
  New Data
CWexprOnly
 0.64
0.60
CWfullModel
 0.72
 NA
CWreducedModel
 0.71
 0.68

The old models are supposed to be converted to work on the new data.  If this does happen, then I'll be able to see the performance of these models on the future datasets (additional METABRIC test dataset + new, previously unpublished dataset).  That would certainly be cool, but this conversion has not yet happened.

In general, my strategy was to pick the gene expression values that correlated most strongly with survival, and I then averaged the expression of probes either positively or negatively correlated with patient survival.  On top of this, I further filtered the probes to only include those that vary between high and low grade patients.  My qualitative observation with working with breast cancer data has been that genes that vary with multiple clinically relevant variables seem to be more reproducible in independent cohorts.  So, I thought that this might help when examining the true, new validation set.  However, I gave this much smaller weight than the survival correlation (I required the probes to have a survival correlation FDR < 1e-8 and a |correlation coefficient| > 0.25, but I only required the probes to also have a differential grade FDR < 0.01).

So, these three models can be described as:

CWexprOnly: cox regression; positive and negative metagenes only

CWfullModel: cox regression; tumor size + treatment * lymph node positive +  grade + Pam50Subtype + positive metagene + negative metagene

CWreducedModel: cox regression; tumor size + treatment * lymph node positive + positive metagene

The CWreducedModel was used to see how much a difference it made to only include the strongest variables (and to what extent the full model may be subject to over-fitting).  The CWexprOnly model was used to see how well the gene expression could predict survival, even without the assistance of any clinical variables.

I included the treatment * lymph node positive variable because it defined a variable similar to the strongly correlated "group" variable, without making assumptions about which were the most important variables (and, as I would later learn, the "group" variable won't be provided for the new dataset).

Additionally, one observation I made prior to the model building process was how strongly the collection site correlated with survival (see below).  This variable wasn't defined by the individual patient, and  I assumed this should be a technical variation (or at least something that won't be useful in a truly independent validation dataset).  The new data dimenishes the imact of this confounding variable, but the correlation is still there.


 
Old Data
New Data
Collection Site
0.42
 0.23
Group
-0.51
 -0.45
Treatment
0.29
 0.28
Tumor Size
-0.18
 NA
Lymph Node Status
-0.24
 NA


ER, PR, and HER2 status are also important variables.  However, PR and HER2 status was missing in the old data, and I didn't record the original ER correlation.  Therefore, they are among the variables that I don't report in the above table.  Likewise, the representation of the tumor size and lymph node status variables changed between the two datasets.

This was a valuable experience to me, and I'm sure the DREAM papers that come out next year will be worth checking out.  There were some details about the organization that I think can be improved (avoid changing the data throughout the competition, find a way to limit the model of models to avoid cherry picking of over-fitted, non-robust models, and providing rewards for intermediate predictions of data where the users could cheat use the publicly available test dataset).  Nevertheless, I'm sure the process will be streamlined if SAGE assists with the DREAM competition next year, and I think there will be some useful observations about optimal model building from the current competition.

Friday, May 25, 2012

Shared Scripts for Genomic Analysis

I've recently added a section to my personal website that contains some scripts that I have used for bioinformatics analysis:

https://sites.google.com/site/cwarden45/scripts

As of right now, the page contains a handful of scripts for microarray, next-generation sequencing, and qPCR analysis.  I plan on updating this page periodically.

Generally speaking, the scripts aren't organized in a carefully documented package (like what you can find from Bioconductor, etc.).  However, I have found them to be very useful templates for routine analysis, so I thought it might be useful to share them with others.

Friday, July 1, 2011

Review of the TCGA Ovarian Cancer Paper

Initial analysis of the TCGA data for ovarian cancer was recently published in Nature this week.  The Cancer Genome Atlas (TCGA) is a joint project by the NCI and NHGRI to study genomic changes that are associated with many different types of cancer by collecting a large number of patient samples for analysis using mRNA gene expression microarrays, copy number arrays, methylation arrays, miRNA microarrays, and exome sequencing.  This data can be freely downloaded using the TCGA Data Portal.

There is a huge amount of information presented in this paper.  For example, the first TCGA paper provided an overview of the glioblasoma data, mostly focusing on somatic mutations and copy number alternations.  There have been a number number of subsequent papers studying the glioblastoma data, and the subsequent TCGA papers that I am most familar with focused on subtypes defined by gene expression patterns (Verhaak et al. 2010) and methylation patterns (Noushmehr et al. 2010).  The new ovarian cancer TCGA paper provides all of the information provided in the glioblastoma nature paper in addition to the subtyping analysis that was covered in mulitple high-impact, highly cited papers.

I think one of the most important take-home messages was the extremely important role of p53 in ovarian cancer.  For example, 96% of high-grade tumors showed p53 mutations, which has also been shown previously in publications such as Ahmed et al. 2010.  In contrast, the TCGA glioblastoma paper showed a p53 mutation rate of 38% to 58% for untreated and treated tumors, respectively.  Interestingly, the ovarian cancer TCGA paper also revealed a high rate of p53 mutation in ovarian cancers contributes to FOXM1 overexpression by using PARADIGM to identify pathway alterations in the new TCGA data (where pathways were defined using the NCI Pathway Interaction Database).

Another striking result was how consistent the copy-number alterations were within either ovarian tumors or glioblastomas but how different the copy-number alternations were between the two cancer types (as shown in Figure 1a).

Although I was impressed that the study defined separate subtypes for mRNA gene expression, miRNA gene expression, and CpG methylation status, I had mixed feelings about the results.  For example, the subtypes defined by methylaton only had "modest stability" (so, they have limited predictive power), and I thought the overlap between the mesenchymal mRNA subtype and tbe C2 miRNA subtye (and the proliferative mRNA subtype with the C1 miRNA subtype) was overemphasized.  I was also a little disappointed that the integrative analysis didn't substantially enhance the subtype definitions (for example, I think Figure S6.4 in the ovarian cancer paper looks less impressive than Figure 3 in Verhaak et al.).  However, I did find it interesting that both the glioblastoma and ovarian cancers had a "mesenchymal" subtype (although I don't think these subtypes necessarily have the same biological meaning), and I think it will definitely be interesting to further characterize the subtypes defined based upon mRNA gene expression.


I was somewhat surprised at how much the survival curves varied for the 4 data sets shown in Figure 2c.  For example, the TCGA test set (N = 255) and the data from Tothill et al. 2008 (N = 237) had very different Cox p-values (0.02 and 0.00008, respectively).  Nevertheless, it is not trivial to get a statistically significant result in 4 independent data sets, and I think the survival results are certianly strong enough to warrant further investigation in order to understand the cause of this variation.

Overall, I would consider this a must-read for any bioinformatician interested in cancer research.

Monday, May 16, 2011

Modeling Bimodal Gene Expression

Since it is often challenging to estimate parameters for mixture models (such as those used to model bimodal gene expression), I thought it might be useful to discuss some of my successes using non-linear least sequares (NLS) regression to model bimodal gene expression.

Many scientists use maximum likelihood estimation (MLE) to model bimodal gene expression (such as Lim et al. 2002, Fan et al. 2005, Mason et al. 2011, etc.).  My MLE model is based the code provided in this discussion thread, so I used the mle function from the stats4 package (which is a wrapper for the standard optim function).

On simulated data, both the MLE and NLS models estimate 37% of samples show over-expression (i.e. come from the distribution with the higher mean), which is very close to the true value of 36%:


Simulated Data


However, I have found that the mle function often returns error messages when working with real data and models built using the nls function tend to fit the data better than the MLE estimates.  Even on the simulated data above, the NLS model appears to prove an ever so slightly better fit for the data (as might be expected because it directly models the density function).

In order to illustrate my point, I have also analyzed some genes exhibiting bimodal gene expression (as identified by Mason et al. 2011) in the GEO dataseries GSE13070.  For example, here is a gene that I could model relatively well with NLS regression whereas I simply couldn't produce an MLE model:




Of course, both of these tools have their limitations.  For example, the data has to have a pretty clean bimodal distribution (here are 3 examples of distributions that couldn't be modeling using either method: ACTIN3, ERAP2, and MAOA (different probe)).  For the NLS model, I also had to set the variance to be equal for the two samples in order to produce a reasonable estimate of over-expression, but I believe this is usually a safe assumption.

Although I do not present the data in this blog post (because it contains unpublished results), I have also found NLS regression to be useful on other genes in several other datasets,  So,  I know NLS regression works well with more than just the one gene that I show above.

Also, for those that are interested, here is the source code that I used to produce all of the above figures.

Monday, March 14, 2011

Article Review: Epigenetic suppression of the TGF-beta pathway revealed by transcriptome profiling in ovarian cancer

In this paper, Matsumura et al. develop a method to identify methylated genes in ovarian cancer patients using gene expression data from roughly 40 ovarian cancer cell lines and 20 cultured primary tumor samples.  The authors posit that this method provides a unique opportunity to study pathways affected by methylation because it directly examines gene expression.

My overall thoughts on this paper:

Pros:
  • The study produced a relatively large amount of data, which is now available in GEO
  • The study utilized a large amount of publicly available data, providing a very useful list of citations for anyone interested in doing bioinformatics analysis on ovarian cancer (especially those interested in methylation).
  • The authors utilize useful open-source tools for pathway analysis (namely GATHER and the specialized binary regression method)

Cons:
  • I think it is more likely that methylation directly suppresses EMT-related genes (such as those involved with cell adhesion) rather than repressing the TGF-beta pathway (which then regulates EMT genes).
  • Unlike in other cancers, patients with methylated genes do not show a worse prognosis.  In fact, I wouldn't be surprised if patents with methylated genes had a slightly better prognosis because methylation suppresses genes associated with the epithelial-mesenchymal transition (which is associated with a progression to a more aggressive cancer).  This hypothesis is also supported by the stromal response data shown in Figure S9.

I think one of the most useful tools discussed in this paper is GATHER, which is very fast and has a simple user interface.  GATHER provides enrichment analysis for information from various databases, such as Gene Ontology, KEGG Pathways, TRANSFAC, and MEDLINE.  More detailed information about the data mined in GATHER can be found in the associated paper by Chang and Nevins.

In fact, GATHER was immediately useful in helping interpret the results of this study.  For example, I used GATHER to check the enrichment for the list of 378 methylated genes described in this paper.  This revealed that the TGF-beta signaling pathway was not the most significantly enriched pathway in the gene list, and the TGF-beta signaling pathway actually had the smallest number of representative genes in the methylated gene list (out of the significantly enriched pathways).  GATHER was also useful for studying the enrichment of pathways in the more conservative "methyl cluster" gene list (which showed a weaker association with the TGF-beta pathway and a stronger association with other pathways, such as the focal adhesion genes).  These are some of the reasons that I believe the methylation directly suppresses EMT-related genes in these ovarian cancer patients (rather than acting through the TGF-beta pathway).

Another useful open-source tool described in the paper is the binary regression method used to define the TGF-beta gene signature.  The binary regression method is especially useful for biologists without a lot of coding experience because it has MATLAB GUI with a simple, user-friendly interface (and version 2.0 is even better than the original code).  In addition to defining gene and pathway signatures, the Bild lab is also currently using this binary regression algorithm to predict drug sensitivity from patient samples.

That said, there are probably a few things I should warn potential users about before giving this product my complete stamp of approval.  Although I have played around this tool a little bit (with encouraging results), I haven't had a chance to use it as much as the relatively common R packages for SVMs (in the e1071 package) and classification trees (in the tree package).  Therefore, I can't really comment about the practical limitations of this algorithm.

I was also a little bit nervous when I saw that Anil Potti (who I mentioned in my previous blog post) was one of the authors on the original Nature paper by Bild et al. for the binary regression method.  However, Potti wasn't involved with the early framework for this method (described by West et al.), and a retraction request for one of the retracted Potti papers states "although we believe that the underlying approach to developing predictive signatures is valid, a corruption of several validation data sets precludes conclusions regarding these signatures."  Therefore, I don't think Anil Potti had any negative influence on the binary regression method.

Overall, I found this paper to be useful and informative, and I would recommend it for anyone interested in microarray analysis.

Sunday, June 27, 2010

Paper on Microarray Analysis of "Watchful Waiting" Prostate Cancer Cohort

Last week, I noticed an interesting paper that was published in BMC Medical Genomics this March.

The authors of this paper wanted to use microarrays to develop an effective prostate cancer diagnostic defined by gene expression patterns.  More specifically, the authors were studying tissue samples taken from the Swedish "Watchful Waiting" cohort.  This large collection of patients developed prostate cancer between 1977 and 1999.  The length of follow-up time for clinical data recorded in this study is significantly longer than has been used in any other attempt to develop a microarray-based prostate cancer diagnostic.  In some cases, clinical information about individuals in this cohort was recorded over 2 decades before the microarray was even invented.

There were two aspects of this study I found particularly interesting.  First, it is pretty rare to find a cohort studied as carefully as the Watchful Waiting cohort.  Second, the authors concluded that "none of the predictive models using molecular profiles significantly improved over models using clinical variables only."

The findings of this study seem to agree with an earlier post where I mentioned two earlier studies to show that GWAS data did not significantly improve risk models for heart disease and type II diabetes.  Although those studies utilized a fundamentally different tools for analysis (the earlier studies looked at genomic sequence whereas this newer study examined gene expression patterns), it was interesting to see examples of cases where genomic technology has not been able to improve upon existing clinical diagnostics.

Of course, these studies leave the reader asking several important questions.  For example, why do these large studies result in negative results?  How long will it take for genomic research to make substantial impacts on clinical diagnostics and therapeutics?  What are the practical limits for developing applications based upon medical genomic research?

I'm not going to even pretend like I know the answers to all of these questions.  Although I'm certain that genomic research will ultimately result in disappointing results for some major studies, this paper did provide some hope that genomic research can still pave the way for future breakthroughs.

For example, the authors discuss how there is significant heterogeneity within and between prostate cancer samples - expression patterns in one region of a given tumor can be significantly different than other regions of that same tumor, and this makes it especially difficult to compare gene expression patterns between different tumors.  It is also important to determine the optimal time to take tissue samples for analysis; diagnostics taken too far in the advance will not yield clinically useful information, and feasible treatments may not even exist for results of a diagnostic applied during a late stage of cancer development.  The authors also point out that several other diagnostic microarray studies resulted in similar lists of prostate cancer biomarkers.  In other words, microarray analysis can probably yield reasonably accurate results - the problem is that the biomarkers aren't a significant improvement over current diagnostics.

I find it encouraging that the authors have a plausible explanation for their negative results and that independent microarray studies have come to similar conclusions, and I continue to be hopeful that genomics research can help achieve important medical breakthroughs in the future.

-------

FYI, Nakagawa et al. 2008 is also an excellent prostate cancer study utilizing microarray data.

Friday, April 2, 2010

Why Do Genomic Cancer Diagnostics Cost So Much?

After reading the introduction to this PLoS ONE article, I started to wonder why there are there several published microarray expression profiles for cancer progression yet relatively few microarray-based diagnostics used in a clinical setting. Although this PLoS ONE paper focuses on analysis of ovarian cancer (and mentions the lack of a clinical microarray diagnostic for ovarian cancer), the paper also cites the current use of a breast cancer diagnostic called MammaPrint.

After reading the wikipedia entry on MammaPrint, I was surprised to learn that it took 5 years for the diagnostic to reach the market following the initial publication showing that the expression profiles for a set of 70 genes could successfully predict the cancer progression. This information is important because more aggressive treatments early in cancer progression may be able to help cancer patients who would otherwise have a high mortality rate (as predicted by their gene expression profile). I was also surprised to learn the high price of both MammaPrint and its competitor Oncotype DX. Although I do not think I can provide a complete answer to why these prices are so high, I would like to take a moment to first demonstrate that the price of these tests far exceeds the cost to conduct the test and then discuss how I think these costs can be offset by decreasing the amount of time and effort that it takes to bring a medical diagnostic tool to the market.

Based upon their wikipedia entries, the MammaPrint diagnostic costs $4,200 and the Oncotype DX test costs $3,978. To give you an idea about how much it actually costs to carry out this test, it costs $350 for a full service microarray analysis (including labor and data analysis) of an Agilent Whole Human Genome Microarray for on-campus customers at the UT-Southwestern Micoarray facility. This is comparable to the cost of most of the microarray facilites that I have worked with, and Agilent produces high quality microarrays. Now, most laboratory kits have a warning that they are “intended for research purposes only,” and this is probably true for the human Agilent array. However, I think this warning is mostly to avoid litigation and not due to a severe lack of technical accuracy, and I expect the actual cost for a clinical microarray test to be in the hundreds (not thousands) of dollars. I’m sure that this high cost is the product of a combination of factors, such as research costs, legal costs, patent law, and the US healthcare system. However, I’m going to focus on ways to potentially cut research costs because that is the area that I know most about.

Now, I want to make clear that the initial publication of a potential diagnostic test is not sufficient to prove the widespread effectiveness of that test. For example, the microarray test for ovarian cancer in the PLoS ONE article had substantially better predictive power on the training dataset than when applied to a new dataset. Therefore, I want to make clear that I do think follow-up studies were necessary to prove the effectiveness of MammaPrint.  However, I still don’t think it should have taken 5 years to test the effectiveness of this diagnostic and I think effectiveness can be determined without as much government regulation.

Before MammaPrint could be put into widespread use, it had to gain FDA approval. This required multiple verification studies, and this is the crucial event that defines the 5 year gap between initial publication and availability on the free market. First off, I don’t think FDA approval should be necessary for diagnostics. I do think physicians need some way to quickly access the effectiveness of a medical diagnostic and/or therapeutic, but I think there are more better ways to determine the effectiveness of a given treatment. For example, a relatively recently posted TED talk by Jamie Heywood discusses how his start-up Patients Like Me, developed by three MIT engineers, can diagnose medical treatments more quickly and effectively than clinical trails. This website analyses a database of information provided by patients, and therapeutic effectiveness can be assessed immediately based upon currently available data. In the very least, I think this company could be an excellent model for a more formal system using data from physicians that does not carry all the restrictions of a clinical trail. These changes should decrease the cost of medical care because companies claim that these price markups are necessary to recoup the costs of research and development, and a more streamlined process for accessing the effectiveness of treatments will decrease research costs.
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.