Showing posts with label predictive models. Show all posts
Showing posts with label predictive models. Show all posts

Monday, October 1, 2012

My DREAM Model for Predicting Breast Cancer Survival

This summer, I have worked on submitting a few models to the DREAM competition for predicting breast cancer survival.

Although I was originally planning on posting about my model after the competition was completely finished, I decided to go ahead and describe my experience because 1) my model honestly didn't radically differ from the example model and 2) I don't think I have enough time to redo the whole model building process on the new data before the 10/15 deadline.

To be clear, the performance isn't all that different for the old and new data, but there are technical details that would have to be worked out to submit the models (and I would want to take time to re-examine the best clinical variables to include in the model).  For example, here are the concordance index values for my three models on the training dataset:

 
 
  New Data
CWexprOnly
 0.64
0.60
CWfullModel
 0.72
 NA
CWreducedModel
 0.71
 0.68

The old models are supposed to be converted to work on the new data.  If this does happen, then I'll be able to see the performance of these models on the future datasets (additional METABRIC test dataset + new, previously unpublished dataset).  That would certainly be cool, but this conversion has not yet happened.

In general, my strategy was to pick the gene expression values that correlated most strongly with survival, and I then averaged the expression of probes either positively or negatively correlated with patient survival.  On top of this, I further filtered the probes to only include those that vary between high and low grade patients.  My qualitative observation with working with breast cancer data has been that genes that vary with multiple clinically relevant variables seem to be more reproducible in independent cohorts.  So, I thought that this might help when examining the true, new validation set.  However, I gave this much smaller weight than the survival correlation (I required the probes to have a survival correlation FDR < 1e-8 and a |correlation coefficient| > 0.25, but I only required the probes to also have a differential grade FDR < 0.01).

So, these three models can be described as:

CWexprOnly: cox regression; positive and negative metagenes only

CWfullModel: cox regression; tumor size + treatment * lymph node positive +  grade + Pam50Subtype + positive metagene + negative metagene

CWreducedModel: cox regression; tumor size + treatment * lymph node positive + positive metagene

The CWreducedModel was used to see how much a difference it made to only include the strongest variables (and to what extent the full model may be subject to over-fitting).  The CWexprOnly model was used to see how well the gene expression could predict survival, even without the assistance of any clinical variables.

I included the treatment * lymph node positive variable because it defined a variable similar to the strongly correlated "group" variable, without making assumptions about which were the most important variables (and, as I would later learn, the "group" variable won't be provided for the new dataset).

Additionally, one observation I made prior to the model building process was how strongly the collection site correlated with survival (see below).  This variable wasn't defined by the individual patient, and  I assumed this should be a technical variation (or at least something that won't be useful in a truly independent validation dataset).  The new data dimenishes the imact of this confounding variable, but the correlation is still there.


 
Old Data
New Data
Collection Site
0.42
 0.23
Group
-0.51
 -0.45
Treatment
0.29
 0.28
Tumor Size
-0.18
 NA
Lymph Node Status
-0.24
 NA


ER, PR, and HER2 status are also important variables.  However, PR and HER2 status was missing in the old data, and I didn't record the original ER correlation.  Therefore, they are among the variables that I don't report in the above table.  Likewise, the representation of the tumor size and lymph node status variables changed between the two datasets.

This was a valuable experience to me, and I'm sure the DREAM papers that come out next year will be worth checking out.  There were some details about the organization that I think can be improved (avoid changing the data throughout the competition, find a way to limit the model of models to avoid cherry picking of over-fitted, non-robust models, and providing rewards for intermediate predictions of data where the users could cheat use the publicly available test dataset).  Nevertheless, I'm sure the process will be streamlined if SAGE assists with the DREAM competition next year, and I think there will be some useful observations about optimal model building from the current competition.

Sunday, February 27, 2011

My 23andMe Results: Using Predictive Models for Genetic Testing

NOTE: Getting Advice About Genetic Testing

There is a big difference between saying that a mutation has a statistically significant association with a trait and saying that a mutation has strong predictive power for a trait.  To illustrate this point, I’ll focus on the 23andMe prediction for eye color.

One of the top 5 traits shown in my trait overview is “eye color.”  My eye color was predicted to be “likely brown” (which was an accurate prediction).  However, I wanted to look at this report more carefully because I remembered Francis Collins mentioning that 23andMe incorrectly predicted his eye color in “The Language of Life".

I don’t know Francis Collins’ genotype, but I did notice a potential problem when I looked at the section for “Your Genetic Data”

Genotypes for rs12913832 (from 23andMe)

Percent Brown
Percent Green
Percent Blue
AA
85%
14%
1%
AG
56%
37%
7%
GG
1%
27%
72%

My genotype is AG (shown in red above).  I was correctly predicted to have brown eyes, but I actually only have a 56% chance of having brown eyes.  This means that the AG genotype is roughly as accurate as a flip of a coin at predicting individuals with brown eye color.  In my opinion, I don’t think this should have come up as a top prediction, and I think it probably would have been better to flag this SNP as “not predictive” for this particular genotype.

Now, don’t get me wrong – I don’t think people should only have access to information about their highly predictive SNPs.  In fact, I would want to be able to know the frequency of my genotype if it significantly varies for different traits.  However, I don’t think information like my predicted eye color should have been something that caught my eye within 5 minutes of viewing my results.

In general, I think it would be useful to give scores to predictions in the same way that stars are given for disease risk.  Ideally, it would be nice to provide all the relevant information (such as overall accuracy, sensitivity, specificity. positive predictive value, negative predictive value, etc.) about these predictive models in a table, but in practice it might be better to focus on one or two features for simplicity (especially when dealing with traits that are not binary).   In order to distinguish between predictive scores and reproducibility/statistical scores (which I what I would call the 4 star system), perhaps SNPs with a positive predictive value (PPV) greater than 75% get a bronze circle, SNPs with a PPV greater than 85% get a silver circle, SNPs with a PPV greater than 95% get gold circle, and all other SNPS get labeled as "not predictive."

These calculations become trickier when considering diseases that are significantly influenced by multiple SNPs, and this is where building a predictive model (using SVM, CART, etc.) could really be helpful in providing individuals with the most accurate predictions.  In fact, these predictive models need not only consider genetic information and can also be used for non-genetic / environmental risk factors like weight, family history, blood sugar, blood pressure, etc. (which I mentioned in my second post).

Unfortunately, there is no absolute best way to build a predictive model using different SNPs (and/or non-genetic information).  Based upon my experience, I think regression-based models do a good job of providing probabilities that individuals have a particular trait, and very strong associations should have similar results regardless of which machine learning technique is used.  However, I don’t know if a single tool will be appropriate for creating all predictive models, and I’m sure there are some traits for which no good predictive model can be created.

Since there is not a lot of precedence for clinically-relevant predictive models incorporating genetic and non-genetic information, I think 23andMe has a great opportunity to experiment with this for their trait predictions.  Since 23andMe is the largest DTC genetic testing company, they will have the most incoming data that they can use as validation sets.  If they are worried about how easily individuals will interpret these results, perhaps a separate “experimental” section can be providing this type of result (just like I think it might be best to test these models for the "trait" section before incorporating these results into disease risk and drug response results, which people are more likely to use when making medical decisions).  Also, I should acknowledge that I don't know what's going on behind the scenes at 23andMe (for example, I don’t know how 23andMe is currently combining SNPs for disease risk etc.), so this may have be something they have already started to investigate.

Finally, I would like to close this 3-part post by emphasizing that I was generally pleased with my 23andMe results, and I have only provided constructive criticism because I want these results to be as clear and accurate as possible because I think 23andMe has the potential to be an invaluable resource to empower patients to utilize their genetic information to the fullest extent.
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.