Showing posts with label diabetes. Show all posts
Showing posts with label diabetes. Show all posts

Thursday, December 5, 2019

PRS Results from my Genomics Data (mostly from impute.me)

I haven't had a whole lot of personal experience with Polygenic Risk Score (PRS) estimates, so I thought it was interesting when I found a couple options for re-analysis of my own genomics data (for selected examples):

Association SNP chip
(impute.me)

(Folkersen et al. 2020)
Other
Re-Analysis Options
23andMe Results
Type 2 Diabetes
(No)
(Type 2 Diabetes, 146 variants)

Average / Above Average
(23andMe-V3, 12/19)

Average / Above Average
(AncestryDNA, 12/19)
MySeq

1.000 risk ratio [error]
(Nebula lcWGS)

0.955 risk ratio
(Genos Exome, 3 variants)

1.089 risk ratio
(Veritas WGS, 6 variants)
"Typical Risk" of 23% (directly from 23andMe, PRS with 1,244 loci)
[actually, slightly lower than normal]

Reduces to less than 1% when age, height, weight, fast food consumption, and exercise rate are taken into consideration (also from 23andMe)
Ulcerative Colitis
(once, so I think really "no")
(23 variants, and 116 variants)

Both Below Average and Above Average Risk, for different PRS
(23andMe-V3, 12/19)

Both Below Average and Above Average Risk, for different PRS
(AncestryDNA, 12/19)
Anxiety Disorder
(Yes, but getting better)
(6 variants)

Average / Above Average
(23andMe-V3, 12/19)

Average / Above Average
(AncestryDNA, 12/19)
Migraine
(Periodic)
(26 variants, and 21 variants)

2 PRS (Average and Above Average)
(23andMe-V3, 12/19)

2 PRS (Average and Above Average)
(AncestryDNA, 12/19)
Eye Color
(Light Brown)
DNA.land

Likely to have Brown Eyes
(23andMe-V3, 12/19)

Likely to have Brown Eyes
(23andMe-V3_V5, 12/19)

Likely to have Brown Eyes
(AncestryDNA, 12/19)
23andMe reports that I am expected to have "brown or hazel eyes" based upon 1 SNP (rs12913832)
Hair Color
(Light Brown)
See Below

(Roughly 25% Red and 50% Blonde)
For "Light or Dark Hair", 23andMe reports that I have "Likely Dark" Hair (using 42 SNPs)

For "Red Hair" 23andMe reports that I am "Unlikely to have red hair" (using 3 MC1R SNPs: rs1805007, rs1805008, and another custom MCR1 probe)
Height
(180 cm)
See Below DNA.land

171 cm: "Likely Taller than Average"
(23andMe-V3, 12/19)

171 cm: "Likely Taller than Average"
(23andMe-V3_V5, 12/19)

171 cm: "Likely Taller than Average"
(AncestryDNA, 12/19)

Individual SNP risks were reported (from impute.me).  While I had a bit of a hard time finding the precise overall risk estimate (without trying to sum / multiply separate risks), this might be OK in terms of getting a sense of whether I was an outlier or not.  For example, being above or below average for "Type 2 Diabetes" seemed to vary (unless you say most people were under something like a null distribution for "average" risk).  In other words, I thought the following plots (which you could see for various traits) were interesting:

impute.me Type 2 Diabetes PRS (23andMe V3)



impute.me Ulcerative Colitis (1st entry, 23andMe V3)


impute.me Ulcerative Colitis (2nd entry, 23andMe V3)

impute.me Anxiety Disorder PRS (23andMe V3)

impute.me Migraine-Broad PRS (23andMe V3)


impute.me Migraine PRS (23andMe V3)

impute.me Hair Color (23andMe V3 + Ancestry DNA, respectively)



impute.me Height (23andMe V3)



I thought the anxiety disorder result was interesting for 2 reasons.  First, I have had issues with anxiety problems (for example, you can click here for notes, even though they are primarily related to PatientsLikeMe).  Second, notice the environmental component is larger than the genetics component.  This matches my concerns that I expressed in this review of "blueprint".  For example, I would say the predictive power from birth has some notable limitations (such as difficulties in the need to take medication at any given point in your life).

While I am not sure if the exact right term was used (since I thought "Ulcerative Colitis" was a condition, rather than a symptom).  However, I was hospitalized for Ulcerative Colitis (even though that was a one time occurrence caused from E. coli with Shiga toxin).

I also get migraines.

I don't have Type 2 Diabetes, but I provided that because I also had other PRS results to compare.  Similarly, if others have suggestions where I can quickly compare to the impute.me PRS results, please let me know and I would be very happy to add them!

For example, I did add DNA.land (and 23andMe) Eye Color and Height based upon a Twitter response.  While I think height is one of the more heritable traits, DNA.land couldn't guess my actual height within a few inches (and there is a noticeable spread of points for the impute.me plot above).  Even though DNA.land gave lower confidence to other predictions, I would say these have been "fair" rather than "high" confidence (and everything else probably should have been "low" confidence).  I am close to the diagonal for the impute.me plot, but I don't know if the scale is 1:1.  For example, my DNA.land height prediction was off by 3-4 inches.  However, to be fair, note that the highest and lowest percentiles for high don't have overlap (there are not any points in the upper-left or bottom-right regions of the scatter plot, even though those make up a smaller fraction of the population).

For comparison, here is the distribution of score for DNA.land (where my true height was greater than anything on the density distribution - perhaps because this was height scaled for female percentiles?):



For impute.me, the predicted hair color shows blondness on the x-axis and redness on the y-axis.  The cyan circle is my actual color (which I filled in), and the while circle is my predicted color.  I think my hair color used to be lighter than it is now (and I think the shade that I reported for myself was a bit too dark), so that is closer to the genetic prediction (perhaps half-way between).

It may be worth noting that 23andMe could predict that I had brown hair and eyes (although I think that covers most people and you need the more rare traits to better calculate accuracy - for example, Francis Collins said that his 23andMe report indicated he had brown eyes when he really had blue eyes, at least 10 years ago).

Again, for comparison, here is the distribution of DNA.land scores for eye color:



I didn't add the AncestryDNA density plots since they looked qualitatively similar to the 23andMe V3 plots (and, on another computer, I had an issue with the percent variance explained appearing in a pie chart that was harder to read).  I also originally intended to test my updated 23andMe genotypes (V3+V5), but I got an error saying that data was already uploaded (from my V3 chip).  However, perhaps I can test those results later, and see if they are still similar.

With a $5 donation, the turn-around time for processing was 1-3 days.

For Genos Exome and Veritas WGS data, I used the BWA-MEM Re-Aligned GATK Variant calls.  However, I think the main conclusion from looking at my diabetes results was that I was of average risk, and I don't believe my own genetic diabetes PRS risk assessment was great without taking additional factors into consideration (for 23andMe, that was a difference between 23% and 1%, after considering BMI, diet, and exercise).

This essentially matches Supplementary Figure S12 for this paper (whose title I respectfully believe can give the reader the wrong impression, and there is at least one objective error that I believe needs to be corrected), where absolute risk explained was usually very low (usually explaining less than 15% of the variation for a trait).  You can also see that the variability explained by "this score" for the impute.me PRS above is estimated to be less than half of the genetic component.

I think the preprint by Brockman et al. 2021 might also have some additional relevant information for this discussion.

In somewhat different contexts, you can also see some notes / concerns about percentiles / indices in the posts on Nebula and basepaws lcWGS results.

Change Log:

12/5/2019 - public post
12/7/2019 - add DNA.land results based upon Twitter reply from Debbie Kennett; revise wording in post
12/8/2019 - mention possible scaling for female height; also fix date for previous log entry.
6/25/2020 - add links to posts with Nebula and basepaws results.  Minor formatting changes.
7/7/2020 - add reference to impute.me paper
4/22/2021 - add reference to another paper
2/4/2024 - change column labels to be more precise

Sunday, February 27, 2011

My 23andMe Results: The Importance of Non-Genetic Risk Factors

NOTE: Getting Advice About Genetic Testing

When I first saw my 23andMe results, I was very glad to see that each genetic association also had a heritability value.  For example, the disease description for Type 1 Diabetes indicates that 72-88% of the disease is determined by genetics whereas the sample description for Type 2 Diabetes indicates that 26% of the disease is determined by genetics (meaning that 74% is determined by environmental factors).  In other words, there is a lot more that can be done to prevent the onset of Type 2 Diabetes than Type 1 Diabetes.

I was also glad to see a “What You Can Do” section describing what actions high risk individuals could potentially take to prevent or manage their disease.

Although these two steps may represent the best way to currently convey this information for most associations, I think it would really help if non-genetic factors could directly be incorporated into risk calculations.

For example, I noticed that one of my Promethease results indicated that a long history of high blood sugar would increase my risk for CAD from 1.7x to 7x (for SNP rs1333049).  If 23andMe could incorporate information about my medical history directly into my risk calculations, then I think that could make the predictions much more powerful.

In addition to providing more precise risk assessments, I think incorporating non-genetic information could also actively help individuals manage their health.  For example, if I knew that losing 30 pounds would cut my risk of developing a particular type of cardiovascular disease by 50% based upon a personalized quantitative model, then I would probably be more inclined to lose that weight than if I simply knew that eating right and exercising was generally a good idea.

That said, I think there are some fundamental changes that may need to take place before such an idea could be implemented (assuming science has progressed to the point where we could provide such models for most diseases).  First, I think users need a more dynamic way to record their medical information in 23andMe.  By this I mean that users need to be able to update their medical information rather than fill out surveys at one point after they create their account.  For example, I responded that I didn’t have any serious illnesses (like cancer, cardiovascular disease, etc.), but I’m sure my answers to those questions will change over the course of my life.

In order to integrate both genetic and environmental risk factors into risk calculations, it may also be helpful to think about other ways to present genetic testing results (which I discuss in greater detail in my third post on predictive models).

My 23andMe Results: Getting a (Free) Second Opinion

NOTE: Getting Advice About Genetic Testing

In order to get an idea about how well the 23andMe risk calculator agrees with other algorithms (when using the same exact same SNP data), I searched for other tools that I could use to analyze my genetic data.

For this post, I have compared my risk assessments from 23andMe to those provided by Promethease (which uses the information available in SNPedia).  I also played around with the free version of Enlis Genome (Personal Edition), but I found the GUI to be a little buggy and they didn’t automatically prioritize risk assessments (unlike 23andMe and Promethease).  So, this post will focus only on comparing my 23andMe assessment with my Promethease assessment.

To be fair, I should point out that I would not necessarily expect 100% concordance between my 23andMe and Promethease results for various reasons.  For example, the “magnitude” score from promethease is a subjective measure, and the curation methods are different for these two tools.  However, I think such a comparison will still be useful because it will still be encouraging to see any predictions that are shared by both tools, and I think both of these tools provide useful information since there is no “standard” way to combine all possible associated SNPs associated with a particular disease.

I will focus on my increased disease risks, but the same principles could be applied to decreased disease risk, drug response, or any other trait.

All Diseases with Increased Risk
NOTE: Percentages refer to percent of individuals with my genotype that have a particular disease, and the relative risk compared to average percentage is given in parentheses.  Percentages are not given for Promethease results because the percentage provided in the summary report refers to the population frequency of the SNP and not the percent of individuals with that SNP that will have a particular disease. Only Promethease results with clearly defined disease names and relative risk values were considered.   Promethease cutoff chosen based upon change in color from pink to red in summary report (and also the number of associations listed at this threshold).  23andMe risk assessment was recorded on 2/26/2011.  Promethease report was generated on 1/27/2011.  Multiple values are provided for Promethease but not 23andMe because promethease provides risk assessments for individual SNPs whereas 23andMe provides a single risk value for each disease.


Overall, I thought that there was pretty good agreement between the two methods.  This may not be apparent from the table above, but that is because my list of “higher importance” SNPs is considerably smaller but with greater overlap.  For example, I would have ideally preferred to look at SNPs with a 1.5x increase in risk and an absolute risk greater than 50%. The absolute cut-off of 50% is because I would prefer to look at SNPs where I am more likely to get the disease than not get the disease.  The 1.5x (or 50% increase in risk) is a somewhat arbitrary cutoff that is loosely based upon my microarray data analysis experience.  Since no SNPs meet both of these criteria, I chose to look at those with a greater than 1.5x relative risk and greater than 5% absolute risk (which, in my opinion, is still quite low).  Now, take a look at my more subjective SNP list.

 “Higher Priority” SNPs



Now, 2 out of the 3 SNPs have similar predictions.  Although there wasn’t a high magnitude SNP in promethease for venous thromboembolism, this could be because I subjectively considered this disease to be less well known than arthritis or diabetes, so I figured less popular diseases may have lower magnitude scores.  For this reason, I decided to look into what SNPs are used by 23andMe and promethease to determine venous thromboembolism risk.  I also checked the Genome-Wide Association (GWAS) Catalog to try and get a idea which SNPs are the best established (according to the US National Human Genome Institute).

SNPs Associated with Venous Thromboembolism Risk

23andMe
Promethease
GWAS Catalog
rs6025
Yes
Yes
No
i3002432
Yes
No
No
rs505922
No
Yes
Yes
NOTE: Promethease lists 19 SNPs associated venous thrombembolism.  In order to simplify the table (and avoid listing some potentially inaccurate and/or low-confidence associations), I have only listed SNPs listed by 23andMe or the GWAS Catalog.


Now there is agreement between the 23andMe and Promethease results because both tools indicate that I have a mutation in rs6025, which results in an increased risk of developing venous thromboembolism.  However, I think it is worth pointing out that the results are not quite as clean as they could be.  For example, this SNP was not listed in the GWAS Catalog, and I couldn’t determine the dbSNP annotation for i3002432 (so it was relatively hard for me to cross-reference this result with other databases).

Another topic that is worth considering is family history.  Before I saw my results, there were 3 diseases that I wanted to check due to family history: type I diabetes, type II diabetes, and macular degeneration.  Thus, it was interesting to see type I diabetes come up in both reports.  Although I doubt that I will get type I diabetes (since the absolute risk is low and this disease usually appears during childhood), this information may still be useful if these mutations have other affects and/or increase the likelihood of children inheriting type I diabetes.

On the other hand, I didn’t see results indicating a increase in risk for type II diabetes and there were some conflicting results about macular degeneration.  Of course, family history is not a gold standard, and I may very well never develop type II diabetes or macular degeneration.  However, I think it is important to think carefully about ambiguous or uncertain results.  For example, this could be done by comparing SNP association to family history as well as considering both genetic and non-genetic risk factors for disease (the later is the topic of my second post).
 

Tuesday, February 23, 2010

Traditional versus Genetic Risk Factors

Yesterday, I started reading this GenomeWeb article discussing the utility of genome-wide association studies (GWAS) in testing drug safety. This article largely agreed with my earlier post. In fact, the article mentions the utility of genetic diagnostics for prescribing abacavir (Ziagen, a HIV drug), as described in “The Language of Life.” The article also mentions new studies to determine genetic risk factors for flucloxacillin (an antibiotic) and clozapine (an antipsychotic). In general, this article provides good support for my earlier claim that drug sensitivity is the most exciting area of GWAS research.

However, the article also mentions that scientists may have overestimated the strength of GWAS in predicting risk to diseases such as heart disease and type II diabetes. This slightly contradicts my most recent post where I claim that genetic predictors of type II diabetes are probably “pretty good.” Therefore, I decided it might be worthwhile to reevaluate my claims.

I think the recent type II diabetes study is well designed, and I was surprised by the results. The authors calculate multiple models for traditional (e.g. age, sex, family history, waist circumference, body mass index, smoking behavior, cholesterol levels) and genetic risk factors. They found that traditional models identified onset of type II diabetes with a 20-30% success rate (at 5% false positive rate), while the gene-based model had only about a 6.5% success rate (at 5% false positive rate). The GenomeWeb article also mentions an interview with the senior director of research at 23andMe (a genetic diagnostic company), who claims that some rare variants may impart especially high risk to certain individuals and genetic tests should therefore complement risk assessment using traditional risk factors. Readers should also take a look at Web Table B of the diabetes study, which calculates predictive power of individual variants. For example, TCF7L2 had one of the strongest associations with type II diabetes in the GWAS Catalog, and this gene did in fact have a significantly higher rate of incidence for diabetes for individuals with a certain set of variants. However, the risk of developing diabetes only increased from 5.4% to 8.6%. Therefore, I think these time consuming diagnostic studies (lasting 10+ years) are unfortunately essential to determine the practicality of genetic tests for diseases whose genetic basis is either complex or unclear.

On the other hand, I don’t think the results of the heart disease study are very surprising. After checking the GWAS Catalog for genetic associations with myocardial infarction and stroke (the diseases examined in the study), I only found a small amount of information on these diseases. There was only one study on myocardial infarction, which only had one significant variant (rs10757278-G, p-value = 1 x 10-20), and two studies on stroke with weak associations (p-values between 9 x 10-6 and 1 x 10-9) and no overlapping associations between the two studies. The heart disease study also used a gene model with 101 variants, which is much larger than the number of significant associations listed in the GWAS catalog (unless the authors considered diseases not explicitly described in the abstract). Unlike the diabetes study, the authors did not provide a table describing how accurately individual variants could predict the onset of heart disease.

Of course, genetic models could improve when novel variants are discovered and/or better models for complex diseases are developed. I also think genetic models would naturally be more accurate for simple diseases that are associated with mutations in single gene. In the meantime, I think consumers need to currently interpret genetic risk factors for complex diseases with a grain of salt, but I think it is still worth getting excited about using genomic tools to predict novel drug targets and discover genetic sensitivities to drugs that are currently on the market or in clinical trials.

Tuesday, February 16, 2010

Playing with the GWAS Catalog


When I was looking at an article in PLoS Biology, I noticed that the abstract listed a comprehensive government database for genome-wide association studies (the “GWAS Catalog”). This database provides a lot of interesting information. In order to get a feel for the data in the GWAS Catalog, I looked at the data for four specific diseases (autism, prostate cancer, type I diabetes, and type II diabetes).

[If you are a non-scientist looking at this database, stronger genetic associations (which should more accurately predict genetic predisposition to a disease) should have low p-values and should be reproducible between different studies.]

1) Autism - There were 3 studies included in the GWAS Catalog. The first two studies identified the same exact region (but with slightly different variants), and the third study identified a different but nearby region. Although I think that there is probably something interesting going on in this region of chromosome 5, I don’t think it is worth getting very excited about the specific variants identified in these studies. For example, the p-values for the autism studies are the lowest out of the four diseases that I analyzed (meaning autism has the weakest genetic component and/or the genetic component of autism is the most complex to model). Furthermore, the most recent study showed that the expression levels of SEMA5A (one of the genes listed in the GWAS Catalog for autism) are very similar for autistic and normal people (see Fig 2. if you have access to this article). The authors of this study claim that gene expression in autistic patients is significantly lower than in normal patients, but I think the statistical significance may be due to an over-fitting problem because they only look at 20 autism patients and 10 control patients (and I have a hard time believing this was enough data to adjust for “age at brain acquisition, post-mortem interval and sex”). The genes with the strongest genetic association in the first study (CDH10 and CDH9) also have similar expression patterns in both autistic and control patents, and the authors of this first study report that this difference is not statistically significant. Of course, the autism variants may be non-functional yet retain similar gene expression levels, but I would still seriously question the strength of any of the specific variants listed in these studies.

2) Prostate Cancer – I looked at the data for prostate cancer, type I diabetes, and type II diabetes because variants for these three diseases are included in at least two of the three major genomic tests listed in “The Language of Life.”  More specifically, the three major genomic testing companies gave completely different predictions regarding Dr. Collins’ risk of getting prostate cancer. The GWAS Catalog lists 11 studies (10 of which have significant associations), and the genetic associations for prostate cancer were much stronger than for autism (p-values equal 3 x 10-33 vs. 2 x 10-10, respectively). Highly significant genetic associations were found within the 8q24.21 and 17q12 regions in several independent studies, but many associations are only found in individual studies. According to “The Language of Life”, deCODE has 13 variants for prostate cancer, Navigenics has 9 variants, and 23andMe has 5 variants. Based upon what I’ve seen in the GWAS Catalog, I think that there probably are at least 5 strong, reproducible variants that could be used to calculate genetic predisposition to prostate cancer, but I am not certain if there 13 variants with well-established genetic associations. However, calculating genetic association for several variants at the same time can be tricky, and the difference in test results may be a problem with the underlying models for calculating genetic association more so than the individual variants considered for the analysis.

3) Type I Diabetes – Type I diabetes has a very strong genetic component, and the molecular basis for this disease is well understood. In these respects, the data in the GWAS Catalog are a good reflection of what is known about this disease. The strongest associations had the lowest p-value out of all the diseases considered (5 x 10-134 for a variant within the Major Histocompatibility Complex, or MHC), and either MHC or HLA (which is part of the MHC) had the strongest genetic association for 4 out of the 8 studies in the GWAS Catalog. This makes a lot of sense because the MHC displays antigens to immune system (thereby telling the body which cells to attack) and type I diabetes is due to due to an autoimmune response where the immune system attacks and destroys the insulin-producing beta cells in the pancreas. It bothered me that some studies reported pretty different results, but that is why I think that it is necessary to only use reproducible associations for genetic testing.

4) Type II Diabetes – The GWAS Catalog contained 15 studies on type II diabetes (12 of which had significant results), which is the highest number of studies listed for the four diseases that I looked at. The strongest associations for type II diabetes had p-values similar to prostate cancer, but higher than type I diabetes. This makes sense because type I diabetes has a stronger genetic component than type II diabetes, so type II diabetes should have weaker associations than type I diabetes. The 8 genes listed as predictors of type II diabetes in “The Language of Life” (TCF7L2, IGF2BP2, CDKN2A, CDKAL1, KCNJ11, HHEX, SLC20A8, and PPARG) were pretty well represented among the different studies listed in the GWAS Catalog, so I bet the predictors of genetic predisposition to type II diabetes are pretty good.

 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.