Showing posts with label Personal PIT Experiences. Show all posts
Showing posts with label Personal PIT Experiences. Show all posts

Sunday, August 4, 2019

Updated Thoughts on My Genomics Results

I have notes on GitHub, but they are mostly within subfolders and may be a bit messy for some people to read.

The repository has been called "DTC" for "Direct To Consumer" genomics.  However, many of those results had an on-line physican.  At a genomics conference, I heard to term "Patient Initiated Testing" to be more precise and cover those types of results.  So, that is why I have tagged the posts with the term "Personal PIT Experiences."

So, I have some thoughts on the topic of genomics results that are already available to the general public.  I realize that some of these may give the reader an overly negative impression.  However, I want to emphasize that a fair presentation should include both positive and negative perspectives, and I have tried to do that (or at least describe possible solutions to potential problems).

In general, I have genomics / medical data publicly available to download on my Personal Genome Project page (for hu832966) and I have what I would consider a partial electronic medical record on my PatientsLikeMe page (however, I think you have to sign in with a free account to view my profile).

With all of that being said, I want to end on a positive note: among all of my notes, I think Genes for Good is something more people should know about.  They don't really provide as much interpretation about the traits / disease, but they have multiple types of whatever they provide (such as different formats for genotype files, with and without imputation, as well as different ancestry results with somewhat independent methods).  I think this is really great, so you can critically assess your data/results and develop your own opinion (to try and decide what is the most fair representation of the various possible interpretations of your data that can exist).

Change Log:

8/4/2019 - public post date
8/5/2019 - minor changes
8/19/2019 - change title for lcWGS health / trait result (after Twitter DM feedback)
8/21/2019 - add "human" to title for positive lcWGS post
9/16/2019 - change title to emphasize "my" for ancestry result
2/4/2020 - further change title to emphasize "my" for ancestry result

Digging Deeper into my Cystic Fibrosis Carrier Status

One overall goal from the various subfolders on the DTC_Scripts repository was to get an idea about how much the data / results could vary between vendors.

I suppose some people might consider it surprising that the "raw" genotypes/variants could vary, but I previously discussed that in a post about re-processing raw data to get more concordant genotypes (and I also have a post about tools to make HLA assignments among this collection of posts).

Some things, like ancestry, may arguably fall under what I would call "hypothesis generation" results, in that some results may be more robust than others (and limitations to the accuracy of specific ancestry assignments are described in another post).

In contrast, this post focuses on something that I think can be utilized with relatively greater confidence (that I am a cystic fibrosis carrier).

That said, in an sense, making sure you get single-gene, rare-disease genomics analysis consistently correct is more complicated than you might expect.  However, in terms of being confident about any genomics result, I think rare variants associated with Mendelian diseases should be a strong point for genomics benefiting society.

So, here is the outline of what happened:


  • In 2011, I was genotyped (with the V3 chip) by 23andMe 
    • This indicated that I was a cystic fibrosis carrier
  • While some carrier status results have been removed (and added back in), I knew my carrier status before there were any issues with the FDA.
  • In 2016, I got Veritas Whole Genome Sequencing raw data (and a GET-Evidence and ClinVar report from the Personal Genome Project)
  • In 2017, I got Genos Exome raw data with an automated report
    • Update (3/17/2020): When I currently sign into the Genos browser, I see my pathogenic variant annotation in the CFTR gene.  I am not sure when/if this was changed, but the report does now successfully show multiple references that are correct for my cystic fibrosis carrier status.
  • In 2019, I ordered a bunch of extra tests (primarily emphasizing the interpretation over the raw data), but this included Helix Exome+ data from the Mayo GeneGuide (the raw data cost extra, and was a gVCF).
  • So, I had 3 high-throughput sequencing results that covered my cystic fibrosis variant.  However, none of them indicated that I was cystic fibrosis carrier in a way that was immediately obvious, and I think at least one (Mayo GeneGuide) failed to report my cystic fibrosis status (even when covering a smaller number of diseases).
    • You can see my FDA MedWatch / MAUDE report for Mayo GeneGuide in MW5093889.  Helix sent me an e-mail that Mayo GeneGuide was discontinued on 4/30/2020, which you can also see on this website.
    • There are some extra formatting changes that I wasn't expecting, but you can also see my FDA MedWatch / MAUDE report for Veritas Genetics in MW5093888.  That said, I was describing my Personal Genome Project report (since I ordered the sequencing through the PGP) and I don't think Veritas specifically marketed annotating my cystic fibrosis status.  So, it might be OK if it is harder to find this report for Veritas Genetics through the search function.
    • I was particularly surprised by this for GeneGuide, since they limited the number of diseases they officially tested for (which I think was a good idea).  However, their guidelines for defining a pathogenic variant didn't include the variant covered by the 23andMe array.
    • It might also be worth mentioning that an on-line physician signed off of these other 3 results, but that didn't improve the accuracy of my cystic fibrosis carrier status.
  • With the 23andMe result, I could check the details of the variant they used to define me as carrier.  Namely, I could verify my carrier status for rs121908769 in ClinVar.
  • I might be forgetting the exact order of events after that.  However, the following gave me extra confidence that my earliest 23andMe result was in fact the "correct" one.
    • I could visualize my alignment in IGV (for my Veritas WGS and Genos Exome data) to see that I did in fact carry the variant (see below).
    • While a lot less intuitive to visualize, the Helix Exome+ data (which I had to pay extra for, beyond my GeneGuide results) also indicated that I had the variant in question, and IGV does accept a gVCF as an input file (see further below, under the .bam visualization).
    • I used the above data in response to a question on Biostars, I was particularly pleased to discover that I got feedback that helped me gain confidence in my own result.
      • For example, I learned about a website called CFTR2, which provides information unique to cystic fibrosis and the CFTR gene.
      • Specifically, this specialized website indicated that my 394delTT variant should be considered pathogenic for cystic fibrosis (if you have two pathogenic alleles).  Please note that you have to check usage agreement to view the specific result linked above.
      • I also discovered some formatting issues that I believe was responsible for at least one false negative.
    • In other words, all 4 results correctly indicated that I had the variant.  The only issue was with interpretation of that variant (which was "correct" for 1 out of 4 results). 
    • I thought I talked to multiple genetic counselors, but my GeneGuide notes indicate that the genetic counselor from PWNhealth agreed that the above information indicates that I was a cystic fibrosis carrier (even though I believe they were providing guidance for a result that more formally incorrectly indicated that I was not a carrier).

Veritas WGS /  Genos Exome BAM (Provided + BWA-MEM Re-Alignment)



Helix Exome +  / Mayo GeneGuide (gVCF)



In many ways, I still consider this a positive experience.  For example, note the following:


  • Having access to raw data allowed me to determine something that was incorrect / missing in my original report (and I think this should essentially be required)
    • That said, I hope the screenshots above show that FASTQ+BAM+VCF is probably a better format to require providing, rather than gVCF
  • Notice, I got free feedback in a public community forum (Biostars) that helped provide me information that I didn't obtain from any of the companies that I paid for genotyping / sequencing.  This emphasizes the value in having free options for re-analysis / re-processing of your data.
  • While it might require some additional training, sometimes simply viewing your data in IGV (a free genome browser) may be helpful for genetic counselors to assess the accuracy of individual genotypes.
    • While it makes life more difficult, the majority vote (3/4 companies, if you count as I did above) would actually be the wrong answer (falsely indicating that I as not a cystic fibrosis carrier).  So, kind of like I can tell that I need to work on fewer projects more in-depth, I think it probably helps to have specialization for genetic counselors (so, they can have an idea about what questions to ask, beyond what is provided in a short report).
  • I successfully learned (somewhat) more in-depth about a carrier status that could impact offspring (if my partner was also a carrier).  If planning to have a child should be decided on the scale of years (or you are assessing life-time risk for diseases with onset later in life), then taking some time to understand your genome on the scale of years may be OK (although, if you use IVF+PGT, you do need to make sure that the pre-defined variants are missing with high accuracy, on a shorter time-scale)


That said, I do think it is important to have realistic expectations about what can be done in genomics, and the need to spend a non-trivial amount of time sorting out the details for your area of expertise.

Update Log:

8/4/2019 - public post date
8/5/2019 - minor changes
8/6/2019 - minor changes
8/14/2019 - minor changes
8/15/2019 - minor changes
8/16/2019 - add link to IGV
3/17/2020 - list my ability to find CFTR pathogenic variant from Genos
4/24/2020 - add link to FDA MedWatch report (Helix + Mayo GeneGuide)
4/30/2020 - add link for Helix discontinuing Mayo GeneGuide
5/4/2020 - add link to FDA MedWatch report (Veritas Genetics)

Concerns About Using Low-Coverage Sequencing for Trait or Health Results

This is a subset of my notes from my Nebula lcWGS sequencing on GitHub:

NOTE (2/24/2020): Nebula is currently offering 30x sequencing.  So, my concerns about the low coverage Whole Genome Sequencing (lcWGS) at ~0.5x are probably less relevant for that particular company.  However, if you get lcWGS from another company, then this information is probably relevant.

Concerns about Specific Variants

While I very much support providing FASTQ, BAM and VCF data, one of my concerns about the Nebula results was the use of low-coverage sequencing.

So, one of the first things that I did was visualize the alignments for some of my more confidently understood variants from previous data (using IGV).

For the two alignments below, the Genos Exome is the top alignment, the Nebula low-coverage alignment is in the middle, and the Veritas Whole Genome Sequencing (WGS, regular-coverage) is at the bottom.

My cystic fibrosis variant (rs121908769):



My APOE Alzhiemer's risk variant (rs429358, Nebula alignment in middle, variant is red-blue bar in the right-most exon):



For APOE, I zoomed out from the screenshot so that you could get a better perspective of the error rate per-read at other positions around the gene.

You could see my cystic fibrosis variant in the 1 read covered at that position, but you can't see any reads with the APOE variant.  My concern about the use of low-coverage sequencing is due to imputation (at least for traits).  Even though this APOE variant is somewhat common (I believe ~15% of the population), the imputation failed to identify me as having that variant.  You see that from the .vcf files

My APOE Alzhiemer's risk variant:

19      45411941        rs429358        T       C       .       PASS    .       GT:RC:AC:GP:DS  0/0:0:0:0.923102,0.0768962,1.71523e-06:0.0768996

As described in the gVCF header:

GT = Genotype
RC = Count of Reads with Ref Allele
AC = Count of reads with Alt Allele
GP = Genotype Probability: Pr(0/0), Pr(1/0), Pr(1/1)
DS = Estimated Alternate Allele Dosage

The "0/0" (for genotype/GT in the last column) means that low-coverage imputation couldn't detect my APOE variant.  In other words, I believe Nebula incorrectly estimates my genotype to be 0/0 with a probability of 92.3%, and the probably for the true genotype was 7.7%.  I also see a blog post mentioning that these probabilities are provided to users through the web-interface, although I am having difficulty in finding them without the gVCF (and you won't see them in the PDFs that I have uploaded in this section).

Update (8/5): Nebula support got in touch with me and explained that the blog post is in reference to Nebula Research Library (rather than the "Your Traits" section).  I canceled my subscription, I am still able to confirm that I see this under the "Library" section (rather than "Traits," "Ancestry," or "Microbiome").

Likewise, there was no delTT variant in the VCF, so my cystic fibrosis carrier status would also be a false negative (if that was used in the report), even though you could actually see that deletion in the 1 read aligned at that position (because 1 read wasn't sufficient to have confidence in that variant).

Overall Variant Concordance

I can also use my VCF_recovery.pl script to compare recovery of my Veritas WGS variants in my Nebula gVCF.

If you compare SNPs, then the accuracy is noticeably lower than GATK (and even lower than DeepVariant):


3,071,596 / 3,419,611 (89.8%) full SNP recovery
3,184,641 / 3,419,611 (93.1%) partial SNP recovery

The indels are harder to compare (becuase of the freebayes indel format).  So, in the interests of fairness, I am omiting them here (as I did for comparing the provided Genos Exome versus Veritas WGS variants).  However, instead of comparing the provided Veritas WGS .vcf file, I can try comparing the BWA-MEM re-aligned GATK Veritas WGS .vcf (which also had higher concordance between my Exome and WGS datasets):


3,133,635 / 3,419,611 (91.6%) full SNP recovery
3,248,277 / 3,419,611 (95.0%) partial SNP recovery
164,140 / 217,959 (75.3%) full insertion recovery
180,736 / 217,959 (82.9%) partial insertion recovery
190,452 / 266,479 (71.5%) full deletion recovery
213,131 / 266,479 (80.0%) partial deletion recovery

The GATK recovery is a little better.  However, it is very important to emphasize that the gVCF variants do not have 99% accuracy (even for average accuracy, or even for SNPs).  I think whatever benchmark was used for that calculation was probably over-fit on some training data.  To be fair, I think the average SNP chip concordance (with higher coverage WGS data) is also lower than some people might expect, but it is definitely higher than this lcWGS data.

You can also show similar results with precisonFDA (using the BWA-MEM realigned GATK gVCF, which we expect to have better concordance than the provided Veritas gVCF).

For example, the overall file shows noticably low recall when comparing the Nebula imputed gVCF versus the Veritas WGS BWA-MEM re-aligned gVCF:




and, to be more fair for the Exome versus WGS comparison in the blog post, the trend is similar within RefSeq CDS regions:




The screenshots are smaller than in the blog post because there was no precision-recall plot for the Imputed Nebula gVCF comparisons.

So, I  disagree with the use of low-coverage sequencing for traits, and I would respectfully consider removing this section (or only made available to those with higher-coverage sequencing).

When I was trying to upload my raw data to my Personal Genome Project page, I noticed that they had an option called "genetic data - Gencove low pass (e.g. Nebula Genomics)".  This makes me think discouraging low-coverage sequencing is something that needs to be done more broadly (at least for health traits).


Concerns abut Nebula Library Results

My concerns for the previous sections are probably solved when using the higher coverage sequencing data.  So, unless you were an earlier customer and had the lower coverage sequencing data, you probably don't have to be extra careful about possibly overestimated accuracy in your genotype imputations.

However, there is one thing that I think could still be a problem for customers with higher coverage sequencing data (if the reports are the same).  The concept is similar to my concern about the basepaws breed index (described in this blog post) and/or other Polygenic Risk Scores that I have collected for myself, but I think I can explain my concern with the top 3 percentile results that I received from Nebula:



As you can see from this link, the percentile above was calculated using 13 SNPs.  Seven of the thirteen SNPs are on chromosome 6, and 5/7 of those variants didn't have alignments against the main reference chrososome (for hg19) in my higher coverage Veritas Whole Genome Sequencing data.  Nebula predicted 3-4 of those 7 chromosome 6 variants to be homozygous variants, but I am not sure if these are correct or not (and the nucleotide for rs3763312 was different from the variants in dbSNP).  There were a pair of variants on chromosome 10 where I was predicted to be heterozgyous at both sites.  For the remaining 6 non-chr6 variants, I had imputed genotypes for 2 heterozygous variants, 2 homozygous non-reference variants, and 2 homozygous reference variants (and they matched my higher coverage WGS data).

I am only 34, but I definitely don't have hair that looks like the Google images for this disease.  So, I don't know the expected age of onset, but I think I might never get this condition (even though Nebula says that I am at the 100th percentile).



As you can see from this link, the percentile above was calculated using 15 SNPs.  12 of those SNPs had at least 1 variant from the reference genome and 11/12 of those variants matched by Vertias WGS variants.  The discordant variant was rs6910071, which was homozygous for the variant allele on the Nebula lcWGS imputed variants.  So, this could have been consistent with 90% overall accuracy, but I didn't have coverage for either dataset at this position (so, this isn't the same as using a gVCF to make a homozygous reference genotype call).

I have osteoarthritis in my lower back, but I don't believe that I (currently) have rheumatoid arthritis.

While I could believe that I am at increased risk, it is important to note that the summary only describes 4% of variance in disease risk.  I think this should be described for all of the reports, to give a sense of the predictive power (along with other statistics).

I also noticed most of the variants were not present in ClinVar (when I was using dbSNP to check the hg19 genome coordinates and reference allele).



As you can see from this link, the percentile above was calculated using 56 SNPs.

I have blood test results uploaded on my PatientsLikeMe profile, and I thought that I had a normal CRP result.  However, it appears that I might not have remembered that correctly, and I might need to wait until my next checkup to see if I can test my CRP level.  However, this is something where I think it would be relatively easy to show if being at the 99% percentile substantially affects your observed CRP levels (or whether there are limits to what this score represents).




As you can see from this link, the percentile above was calculated using 6 SNPs.  Nebula predicted that I had a homozygous variant for 1 SNP and heterozygous variant for 1 SNP (and reference genotypes for the other 4 variants).  However, all of these variants looked OK in my higher coverage Veritas WGS data.

I believe that I have been previously reported to be at higher risk for restless leg syndrome, but that might have actually been for deep vein thrombosis / venous thromboembolism in the earlier 23andMe reports (before the FDA required approval for a more select set of results).  I do sometimes have difficult sitting perfectly still at night.  However, this does't happen all of the time, and I have never been diagnosed by a doctor for having this condition.  So, I would currently lean towards saying that I don't have restless leg syndrome.

For the 2 sets of SNPs that I checked (for alopecia areata and restless leg syndrome), I also visualized my Genos Exome alignment.  However, most of the variants were not covered by sequencing of coding regions.

In general, the journals where these results are published may make some readers think the results are useful.  However, being able to publish a result in a prestigious journal doesn't mean the associations are predictive enough to be clinically meaningful.  Also, even with a more subtle association, being in a prestigious journal doesn't necessarily mean the result can be reproduced.  For example, there are retractions in high impact journals (you can see some in this blog post), and there are objectively wrong conclusions in papers that haven't been retracted (and the science-wide error rate is mentioned in this blog post).  I don't want to cause unnecessarily alarm, but I think it is important to emphasize that time and large sample sizes (and independent validation) are needed to become comfortable with using genetic results to guide your medical treatment.

Nebula does provide a warning: "Disclaimer: Nebula Library is for research, information, and educational use only. This information is not medical advice, nor is it intended to be used for any diagnostic purpose. Please seek the assistance of a health care provider with any questions regarding your health."  However, I think this is easy to miss and the importance may not be fully understood among all customers.

Additional Note #1: I tried to provide a review for Nebula on Trustpilot, but I have encountered some difficulties.  You can see a screenshot of the current review (that was not accepted) here.  While I still haven't gotten a response for the last attempt to submit a review and get an explanation of what I need to change.  While I am not certain if this is the cause, the link that I received to submit a review initially created a review under another name.  To be fair, Nebula did pay me the $10 Amazon gift card (even when I provided a screenshot of a 2-star review, when most are 4- or 5-star reviews), and I hope that this review can eventually be posted (in which case, I will provide a link to that review, instead of this longer explanation).

Additional Note #2: You can see my report to FDA MedWatch (MW5093887) in MAUDE here.  I received an acknowledgement via mail for another report, but I just looked for this report after waiting a while.

Update Log:

8/4/2019 - public post date
8/5/2019 - add update about Nebula research library
8/6/2019 - minor changes
8/10/2019 - add link for 23andMe SNP chip versus WGS concordance
8/14/2019 - minor changes
8/15/2019 - minor changes; add box around APOE variant
8/16/2019 - minor changes
8/17/2019 - revise title (to better emphasize importance, but also presentation of just my own data)
8/16/2019 - minor changes
11/26/2019 - add arrows to Nebula samples in IGV screenshots
1/26/2020 - add screenshot of Trustpilot review
2/24/2020 - mention that Nebula is currently offering higher coverage sequencing
3/19/2020 - add concerns about Nebula Library (to match what I described in my MedWatch report)
3/20/2020 - add links to SNP details for selected library results
3/22/2020 - add link to other PRS post
4/18/2020 - add notes about checking RA post
4/24/2020 - add notes about FDA MedWatch submission
4/26/2020 - minor changes
7/6/2020 - minor changes "Polygenic Risk Score" label, for the last part of the blog post.

Human Low-Coverage Sequencing is Mostly OK for Broad Ancestry and Relatedness

This is a subset of my notes from my Nebula lcWGS sequencing on GitHub (as well as a couple images from two sections with Genes for Good, for full probe RFMix as well as RFMix SNP-chip down-sampling):

Ancestry Predictions

Even though I think they should only provide continental ancestry results (kind of like the 1000 Genomes "super-populations"), the ancestry was roughly similar to my other results (indicating that I am mostly European, which is accurate).  Plus, I describe limits on the more specific assignments (for SNP chip data) in another blog post.

Nevertheless, if I use my imputed genotypes for RFMix chromosome painting, I get results that look roughly like my SNP chip analysis (which would be an improvement over the Genes for Good imputed SNPs, but comparable to the much smaller number of Genes for Good observed SNPs):


There is no plot for chrX, in part because there are no imputed genotypes for chrX.

For comparison, this is what the full set of observed Genes for Good probes looks like:



and this is what the larger set of imputed Genes for Good probes looks like:



In other words, I think the Nebula results are similar (or perhaps slightly worse) than the genotypes that were directly measured for Genes for Good SNP chip probes (which, by the way, are completely free to obtain),  but imputation process also caused some issues with the ancestry with the Genes for Good SNP chip probes.

However, to be fair, the loss of accuracy with imputation does seem be better than using observed measurements if you decrease the probes and/or reference samples enough.  Shown below is the effect of using only 66 reference samples and 15,924 probes: (20x reduction in 1000 Genomes unrelated reference set, and 18x reduction in probes from my Genes for Good SNP chip):



Although, to be fair again, I think the primary problem is arguably the number of reference samples.  Take a look if I use the same number of probes (15,924 probes), but I have the "full" set 1,329 unrelated reference samples:


However, to get something that looks more like the original result, I would argue you do need to arguably increase the probes as well.  For example, please note a similar plot below with 143,320 probes (2x reduction in the starting amount):



Given that I think this looks kind of similar to the Genes for Good imputed set, I think the original Nebula imputed result (with lcWGS) for broad-level ancestry is a reasonable match to the higher coverage results (all things considered).


Kinship / Identity-By-Descent (Close Family Relationships)

Similar to the IBD estimates that are posted within the Helix/Mayo GeneGuide GitHub section (since I was only provided a gVCF), I can test overall similarity between the imputed Nebula genotypes, 23andMe (CW23), Genes for Good (GFG), Veritas WGS (BWA-MEM Re-Aligned) with 77,072 genomic positions (plotting 1000 Genomes reference samples for comparison):



By this measure, you can also clearly see which samples come from the same individual (me).  However, there is a slight drop in the accuracy for the Nebula imputed values (with kinship values between 0.489181 and 0.489226, instead of between 0.499859 and 0.499962):

FAM1 ID1 FAM2 ID2 nsnp hethet ibs0 kinship
0 CW23 0 GFG 77072 0.605148 0 0.499962
0 Veritas.BWA 0 GFG 76310 0.605032 0 0.499865
0 Veritas.BWA 0 CW23 76310 0.605032 0 0.499859
0 Nebula 0 GFG 77072 0.584596 7.78493e-05 0.489209
0 Nebula 0 CW23 77072 0.584596 7.78493e-05 0.489181
0 Nebula 0 Veritas.BWA 76310 0.58472 6.55222e-05 0.489226
In other words, there is some loss in the genome-wide similarity using low-coverage Whole Genome Sequcing (lcWGS), but you can still clearly tell which samples all same from the same individual (me).

However, if the underlying data is not reliable for traits and health results, then my opinion is that the SNP chip is still the relatively better option (as something that costs less than higher coverage sequencing, while giving ancestry /  relatedness results that are at least as good).


That said, to be fair, I think it really could be best if I could see similar example from those with a different primary (Non-European) ancestry.  For example, there was this New York Times article about someone whose broad-level ancestry assignments were less accurate than mine (although there was also evidence for improvement over time).  For example, most people were predicted to be mostly European (regardless of their actual ancestry) and most customers were European, then you could have a result that looks good even though it wasn't actually a very good predictor (beyond a baseline, like "assume everybody has European ancestry").

Update Log:

8/4/2019 - public post date
8/6/2019 - minor changes
8/15/2019 - minor changes
8/21/2019 - add "human" to the title
9/15/2019 - add warning / note about others with a different main broad ancestry

My Genome-wide, Broad-Level Super-Population Ancestry was Robust, but I Observed Some False Positives in Smaller or Specific Segments

You can get an idea of the specific (country) assignments for ancestry in the various sub-folders on GitHub as well as sometimes in reports that I uploaded to my personal genome project page.

First, the good news: most companies indicate that I am mostly of European Ancestry, which is correct.

Second, the mixed news: while there were some findings that were correct, I had concerns about emphasizing a non-trivial false positive rate for some of the more specific ancestry predictions.  For example, I respectfully believe it is inappropriate for 23andMe to encourage travel destinations based upon their ancestry results.

To some extent, the names themselves sometimes indicate a limit to precision.  For example, if the category is "British & Irish" or "French & German," then you already don't have 1 country for a travel recommendation.  While I do have both British and Irish ancestry (and accordingly, those have the best specific marker evidence), I am a little concerned about the basis of some of the more specific assignments that I currently see.  For example, does overall population affect the density of likeihood that I had relatives from London?  If so, I think that would be kind of like assuming I live in either LA or NYC because I am from the United States (technically, I do live in the greater LA area, but I was born in Cincinnati and raised in Atlanta - plus, I think this is probably sufficient to make my point).  Also, it looks like that density plot is somewhat contradictory with the marker status, even within 23andMe.

While this sort of thing may be hard to firmly prove (for example, convergence between companies does not necessarily indicate the result is accurate, which we saw in a different way for my cystic fibrosis result), I have examples of the sort of things which I did or did not consider to be accurate below.

While some of these could be correct, I think it may sometimes be best to think of them like "hypotheses".

Positive Examples of More Specific (Relatively Recent) Ancestry


  • AncestryDNA predicted that I had more recent relatives in Tennessee, which is correct (on my mother's side).  However, even that may have had some limits to precision, given that the 1925-1950 interval seems less relevant to what I know.
  • 23andMe predicted that I had relatives living in Kingston Parish less than 200 years ago.  This could be correct.  Based upon my other family members, I can tell that this comes from my father's side with a relatively robust prediction of ~2-3% African ancestry (with large segments on multiple chromosomes).
    • I thought I had heard that my Great-Great-Grandfather (my Grandfather's Grandfather) was supposed to have been born from family that moved from the Caribbean to the United States (but I don't currently have confirmation of that).  
    • I also have consistent reports of Y-chromosome lineage E-M123.  While I am not sure if that is completely consistent with what I have described above, the greater African ancestry could be coming from Great-Great-Grandfather's father's mother's side (and/or his mother's side).


Effect of Filtering 23andMe Ancestry for Results with Higher Confidence Threshold


  • While I believe the above explanation for my African ancestry is plausible, there were 2 other specific ancestry predictions that I didn't think were right (and, in fact, those could be filtered by increasing the confidence threshold to 90%)
23andMe V3 Chip Ancestry Results (3/21/2019, 50% Confidence)

23andMe V3 Chip Ancestry Results (3/21/2019, 90% Confidence)

As noted in the GitHub notes, the East Asian & Native American and South Asian results go away with the higher confidence threshold (90%, instead of the default 50%).

What is not as clear from the above plots is that I also have notes of my percent Scandinavian ancestry varying from 11% to 3% (both with the V3 chip, at various times), and this is something that I think should have been called "Broadly European" instead of being assigned to a country that I believe is incorrect).  Accordingly, my Scandinavian ancestry also disappears if I change the confidence interval.  For other 23andMe customers, note the pull-down in the upper-right of the above screenshots.  That is how you can get the more conservative predictions (even though, in my opinion, I think it should be the other way around, where you have to opt-in for more speculative results).

I can also perform chromosome painting re-analysis with RFMix with 1000 Genomes reference samples, which I have shown below:



The overall picture is still that I am of mostly European ancestry.  You now start to get some small SAS (South Asian) predictions, but I think this is consistent with my general suggestion that the smaller segments are more likely to be false positives.

Also, all the segments of African ancestry should be coming from my father's side.  So, even though I think the above plot is good for some sense of overall estimates (for large segments), there is some sort of issue with phasing for my large SHAPEIT/RFMix chr14 segments (all the red should be 1 of my 2 copies of Chromosome 14, similar to my 23andMe results).  However, to be fair, there are chromosome-discordant 50% confidence results in my 23andMe data (with the smaller segments on Chromosome 3), which are on the same chromosome for this particular SHAPEIT/RFMix result (although that may also vary with different random seeds on different days).

Going back to the official 23andMe results, I purchased an upgraded V5 chip, and you can see those results below.

23andMe V5 Chip Ancestry Results (7/11/2019, 50% Confidence)

23andMe V5 Chip Ancestry Results (7/11/2019, 90% Confidence)


You now get a East Asian and Native American segment that remains with the higher confidence threshold.  However, using the same rationale as the SAS RFMix segments, I think the 0.1% segment on chr3 (surrounded by regions that were filtered with the higher confidence threshold) should receive less emphasis based upon the size of the segment.  So, if you ignore that (or just look at the most common ancestry prediction), the results for the V3 and V5 chips are consistent with each other (and other companies) with the broad conclusion that I am of mostly European ancestry.

On the flip side, I should also have some Spanish ancestry, which I don't see with the 90% confidence threshold.  However, I can see that ancestry with 50% confidence on chromosome 3 - in fact, that segment is estimated to be larger (2.1% versus 1.3%) in a later ancestry estimate.  So, unless that ancestry is being represented in another way (defined less precisely), this could be an example of a false negative with the higher confidence threshold.


Free Alternate Ancestry Prediction Options


  • I describe these in more detail on the 1000 Genomes re-analysis page for my 23andMe data (on GitHub).  However, I provide the general links here:
  • Again, it might be possible to have a false positive from multiple programs.  However, I think it is an overall good thing that you have these free options for re-analysis available.


To be clear, I am defining a difference between ancestry and relatedness.  In 23andMe, these are even in different sections ("Ancestry" versus "Family & Friends").  As mentioned in another blog post (please scroll towards the bottom), I believe the close family predictions should be accurate (even though I got a weird result when I uploaded my 23andMe data to FamilyTreeDNA).

However, to be clear, I think "specific" closely related individual predictions should be accurate (and I could in fact verify predicted relatives up to the range of second cousin on 23andMe and AncestryDNA), and this is the different than the more distant "specific" country assignments.  This matches 23andMe's definition of a "close relative."  However, it is hard for me to assess the accuracy of the confidence estimates for increasingly distant "DNA relative" predictions.

Update Log:

8/4/2019 - public post date
8/6/2019 - minor changes
8/14/2019 - minor changes
8/15/2019 - mention issue of RFMix phasing for African ancestry
8/16/2019 - minor changes
9/15/2019 - change title to just refer to myself
9/16/2019 - add link about DNA.land
10/18/2019 - mention aunt with Turner Syndrome (later removed, along with entire section)
10/22/2019 - minor change
12/2/2019 - add links to inpute.me and MySeq
1/27/2020 - add note about Great-Great-Grandfather (later removed)
1/29/2020 - modify notes (based upon what I could verify, even though I will probably have more revisions)
2/1/2020 - further modify notes
2/2/2020 - further modify notes
2/4/2020 - modify content throughout post (including changing the name of the section related to changing the 23andMe confidence thresholds, as well as removing some other details and the section about my mom's chromosome X)
2/5/2020 - additional changes in wording
2/6/2020 - minor changes
3/10/2020 - minor change

Emphasis on "Hypothesis-Generation" for Supplement Recommendations

This is a subset of my notes on 4 "Nutrigenomics" companies on GitHub:

To be fair, I am starting with a negative bias.  However, I will be very happy if I can convince people that the companies / organizations need to i) share your raw data with you, ii) share all the details for coming to their conclusions, and iii) have some way of publicly sharing trends in their own data, with fair representations in limitations in confidence.

In other words, that is a little different than the results being inaccurate, which is harder to show (and I have to be mindful of my prior expectations, since that can cause me to be overly harsh).

That said, my primary concern with most of the Nutrigenomics results is that they weren't the best way to get your genomic data (and/or they provided risk assessments that I thought were less useful than direct measurements, such as routine blood tests covered by insurance).  However, the situation was a little different for Vitagene (who does provide genotype information at 632,150 positions).

While my opinion is that I would probably still prefer 23andMe / Genes for Good / All of US over Vitagene, I think that can largely be considered a personal preference.  However, I think there is one point worth emphasizing more (which I feel more strongly about).

Namely, Vitagene provides supplement recommendations based upon your genotypes (and non-genetic information).

While I did experiment a little with 1 supplement because of these results (and may eventually test 1 or 2 more), I think it was important that I viewed the overall results as something that needed to be critically assessed.

In other words, Vitagene recommended that I take 7 supplements:


  • Bromelain Quercetin Complex (500 mg): Lifestyle (Joint health and Digestive health)
  • Probiotics (40 billion CFU): Genetics (31%, risk of Overweight, Hormonal support, Eczema, Allergies and Blood pressure health, based upon 103 variants, all reported to have "Fair" research quality), Lifesytle (Everyday stress and Digestive health), and Goals (Everyday stress and Overweight)
  • Vitamin D (2000 IU): Genetics (59% Vitamin D Levels, Eczema and Joint health, based upon 36 variants, all reported to have "Fair" research quality), Lifesytle (Everyday stress), and Goals (Everyday stress)
  • Theanine (200 mg): Lifesytle (Everyday stress), and Goals (Everyday stress)
  • Iron Free Multivitamin (10 Multi): Lifestyle (Energy and Nutrient intake levels)
  • Zinc (15 mg): Genetics (50% Overweight, based upon 52 variants, all reported to have "Fair" research quality) and Goals (Overweight)
  • Chromium (200 mcg): Genetics (83% Hormonal support, Overweight and Blood Sugar Health, based upon 203 variants, all reported to have "Fair" research quality) and Goals (Overweight)

To be clear, I believe I currently have a normal BMI (23.7), but I would like to trim down my gut a little bit.  However, I think doing some 8 minute abs and exercise (and avoiding over-eating) is probably going to be more helpful to me than a supplement to lose weight.  I only bring this up because "Overweight" appears more than once above.

More importantly, if I over-estimated the ability to accurately make supplement / drug recommendations, I think this could potentially cause harm.

For example, I would be cautious about taking 7 new drugs all of a sudden.  While they attempt to give some indication of drug interactions (such as listing a potential drug interaction between Bromelain and Indomethican), I didn't see a warning about an interaction between L-Theanine (specifically, Nature's Trove L-Theanine) and SSRIs except on the container after ordering the supplement.

Accordingly, I encountered some minor side effects both times that I took it (as I recorded in my PatientsLikeMe profile).  To re-iterate, there was a warning on the container (but not from Vitagene) said "If you are currently taking prescription antidepressants such as MAOIs or SSRIs, consult your physician before taking this product".

I also tested 50 mg of Zinc from target (which I evaluated on PatientsLikeMe).  However, I don't think it really helped with weight loss (after testing taking it for a little less than 2 weeks).

So, this was not a huge problem for me.  However, I would be a little concerned if the typical patient/customer didn't question the full set of recommendations (and/or didn't consult a doctor prior to adding a large number of supplements into their daily routine).  It may also be worth mention that the supplement I ended up testing (as something novel that I thought could help) did not use any of my genetic information to make that recommendation.

Update (3/8/2020): I submitted an FDA MedWatch report for this collection of blog posts, emphasizing the Vitagene result because of the mild adverse reaction (while providing the FDA with the general link).  Identifying information was removed, but you can see what the public version of report MW5092056 looks like in the MAUDE Adverse Event Report database.

Update (4/27/2020 + 5/24/2020): I don't mention it in the above post (since I made the purchase noticeably later), but I think my GitHub notes on Dante Labs may also be relevant (1 of the 3 reports was a "Nutrigenomics" report).  I have also submitted an FDA MedWatch report, and I will provide the identifier for that as soon as I know it.

I am not sure why the information about Dante Labs is missing (since that seems important), but you can see that report in MAUDE as MW5094322.  You can also see the original draft (without the removed information or reformatted text) here.

Update (5/18/2010): When I thought I heard that Everywell was going to provide COVID-19 tests, I thought that I should report my experiences for what appears to be different metabolite levels for the regular blood draw versus the at-home blood test (MW5094002).  I also mentioned that DNA results were provided, but I didn't think they were the best way to get those results and I thought the blood test results should be getting more emphasis.  As mentioned toward the beginning, this was one of the results that I described in greater detail on the GitHub Nutrigenomics subfolder.


Update Log:

8/4/2019 - public post date
8/16/2019 - minor changes
8/20/2019 - mention zinc testing
8/26/2019 - fix GitHub link
8/30/2019 - add link for PatientsLikeMe zinc evaluation
11/20/2019 - add Chromium (all 7 supplements were on the GitHub page, but I previously only listed 6 in the blog post)
3/8/2020 - add link to FDA MedWatch report
4/27/2020 - add information about Dante Labs
5/18/2020 - add reference to Everywell FDA MedWatch report
5/24/2020 - add reference to Dante Labs FDA MedWatch report

Please Take Time to Critically Assess Anxiety-Inducing Results

I am reproducing part of the content from the FamilyTreeDNA GitHub section here:

I received a result that caused me some anxiety for a couple hours (even though I realized it rationally couldn't be true).

Namely, you can upload your 23andMe data to search the database of FamilyTreeDNA members (or other people who have uploaded data).  I was told that there was someone within the range of "Father/Son", who was also an X-chromosome match (and "X-match").  This is not possible: I am a male, so a father or son would have to have a Y-chromosome match (and not an X-chromosome match).

I e-mailed the contact (you can see the e-mail address for your predicted relatives), but I never heard back.  One possibility is that someone could have created a false account that was based upon my public genome data, but I don't know that for certain.  For example, there are a lot of people listed within the range of "2nd Cousin - 4th Cousin," and I don't recognize any of them.  I also don't have access to the raw data for this individual.  In other words, all I know is that this result can't be precise.

For example, as of 10/3/2019, it looks like other people have encountered a similar result with a self-search (so, it may be the "Father/Son" part is not precise, and instead should have said "Self/Twin").  If that is the case, you can see that explanation on Twitter here.  FamilyTreeDNA also confirmed that kits from the same individual (or monozygotic twins) will also appear as "Parent/Child" with 50% similarity, as described in this table.

More importantly, as a general rule, if you encounter a surprising result, I would like to encourage people to first pause and then try to calm down and critically assess the results.  This can go both ways - for example, I would also recommend waiting at least one day before posting anything negative (as probably would have been wise for a Food Sensitivity test, although I have posted an apology in that section).

To be clear, I would usually expect an IBD calculation for the same individual or parent-child relationship to be reliable.  For example, you can see a clear difference when comparing my own samples versus 1000 Genomes samples and most parent-to-child relationships on this blog post (please scroll towards the bottom of the page).  I have also found known and validated novel relationships on 23andMe and AncestryDNA.

Unfortunately, unexpected family relationships can sometimes be true.  However, there can also be limits to the precision of genomics methods and possibly even some allowance for human error (such as sample mix-ups).  So, please take some time to think about making decisions that could affect the rest of your life and may permanently affect your relationships.  For example, to troubleshoot a possible sample swap, do all people involved have predicted relationships with other people fitting the alternative model?  If one person has expected relatives and the other person doesn't have any matches that can be explained, perhaps that indicates the sample that should be re-processed (either with the same company, and/or a different company).

In other words, if you encounter a result that causes you anxiety, please first try to calm down.  Then, please take some time to evaluate the situation: the result could be real, but please also ask questions (from independent resources) about whether there could be misunderstanding and/or inaccurate information that has caused you concern.  For the potential misunderstanding part, perhaps a good first step would be to contact technical support and/or a genetic counselor.


Update Log:

8/4/2019 - public post date
8/6/2019 - minor changes
8/15/2019 - minor changes
10/3/2019 - add solution suggested from Twitter
10/7/2019 - add link from FamilyTreeDNA
7/22/2020 - add link to discussion about uploaded data
6/21/2024 - fix minor typos

Predicting HLA Types for Array and High-Throughput Sequencing Data

My previous link to my HLA-assignments with varying technologies has the most important table in the middle of the page.  So, I am mostly reproducing that here to make the information easier to view.


SNP2HLA HIBAG bwakit HLAminer
HLA-A A*01, A*02
(23andMe)

A*01, A*02
(Genes for Good)

A*01, A*02
(AncestryDNA)
A*01, A*02
(23andMe)

A*01, A*02
(AncestryDNA)
A*01, A*02
(Genos Exome BWA-MEM)
A*01, A*02
(Genos Exome BWA-MEM)

A*01, A*68
(Genos Exome BWA)
HLA-B B*08, B*40
(23andMe)

B*08, B*40
(Genes for Good)

B*08, B*40
(AncestryDNA)
B*08, B*40
(23andMe)

B*08, B*40
(AncestryDNA)
B*08, B*40
(Genos Exome BWA-MEM)
B*08, B*40
(Genos Exome BWA-MEM)

B*08, B*41
(Genos Exome BWA)
HLA-C C*03, C*07
(23andMe)

C*03, C*07
(Genes for Good)

C*03, C*07
(AncestryDNA)
C*03, C*07
(23andMe)

C*03, C*07
(AncestryDNA)
C*03, C*07
(Genos Exome BWA-MEM)
C*03, C*07
(Genos Exome BWA-MEM)

C*03, C*07
(Genos Exome BWA)
HLA-DRB1 DRB1*01, DRB1*03
(23andMe)

DRB1*01, DRB1*03
(Genes for Good)

DRB1*01, DRB1*03
(AncestryDNA)
DRB1*03, DRB1*11
(23andMe)

DRB1*03, DRB1*15
(AncestryDNA)
DRB1*04, DRB1*04
(Genos Exome BWA-MEM)
DRB1*01, DRB1*15
(Genos Exome BWA-MEM)

DRB1*01, DRB1*15
(Genos Exome BWA)
HLA-DQA1 DQA1*05, DQA1*05
(23andMe)

DQA1*01, DQA1*05
(Genes for Good)

DQA1*01, DQA1*05
(AncestryDNA)
DQA1*05, DQA1*05
(23andMe)

DQA1*01, DQA1*05
(AncestryDNA)
DQA1*03, DQA1*03
(Genos Exome BWA-MEM)
DQA1*02, DQA1*03
(Genos Exome BWA-MEM)

DQA1*02, DQA1*03
(Genos Exome BWA)
HLA-DQB1 DQB1*02, DQB1*05
(23andMe)

DQB1*02, DQB1*02
(Genes for Good)

DQB1*02, DQB1*05
(AncestryDNA)
DQB1*02, DQB1*03
(23andMe)

DQB1*03, DQB1*06
(AncestryDNA)
DQB1*03, DQB1*03
(Genos Exome BWA-MEM)
DQB1*02, DQB1*03
(Genos Exome BWA-MEM)

DQB1*02, DQB1*03
(Genos Exome BWA)

In other words, my HLA-A / HLA-B / HLA-C types could be identified more robustly than the HLA-D genotypes (which I don't know, since I haven't gotten a regular blood test).  However, my understanding is that those types have a greater priority in defining organ transplant matches (although I'm currently encountering some difficulty finding the reference for that).

The GitHub link also goes a little deeper into how 23andMe is using 2 SNPs to represent 2 haplotypes (across genes) for celiac disease (which I found surprising, but that is done for other diagnostics as well).  I am mostly leaving that out of this section, but I did think it was interesting that HLA was used in 23andMe's "Meet Your Genes" when the SNPs are actually intronic / intergenic (with respect to the RefSeq annotations).

My 23andMe report indicated that I was DQ8-positive but DQ2-negative for my celiac disease risk.  In terms of defining the 2 genes used to define my DQ8-positive status I coloring matching assignments above in magenta (HLA-DQA1*03 and HLA-DQB1*0302).

Here is a screenshot for the variants tested by 23andMe (where I have the "C" variant for rs7454108, for the marker described as "HLA-DQ8"):



Again, as described here, a positive HLA-DQ8 status is defined by having HLA-DQA1*03 and HLA-DQB1*0302.

More recently, I collected Illumina Whole Genome Sequencing data where unaligned reads were provided (from Sequencing.com), along with some amount of PacBio HiFi data from Dante Labs.  There are some parts of the results that are not especially clear to me and I am interested to learn about additional options for analysis.  However, I believe those results are consistent with me having at least one DQB1*03 allele.

In terms of what appears to be consistent between the PacBio data and Illumina Whole Genome Sequencing data, the T1K results from the Sequencing.com Illumina reads indicates that the related HLA-DQB1 allele should be HLA-DQB1*03:02:01.

I ordered additional GlutenID testing from Targeted Genomics, with an uploaded subfolder on GitHub.  However, I believe the potential problem with using this for validation is that the DQ8-DQ8 result is based upon the same 1 SNP as my 23andMe result (rs7454108).  So, I will continue to look into additional validation options.

I am interested to learn more about the broader trends if certain HLA types are harder to assign and/or impute than other HLA types.  I have some notes mentioned in this Disqus comment.  Comments containing relevant feedback is also welcome on this blog post.

Update Log:

8/4/2019 - public post date
8/6/2019 - minor changes
8/15/2019 - add coloring for HLA-DQ8
2/11/2024 - add information / links to Whole Genome Sequencing data (Illumina from Sequencing.com and PacBio HiFi from Dante Labs)
2/18/2024 - add screenshot from 23andMe + dbSNP link; add additional HLA-DQ8 sentence; add small paragraph for Illumina WGS T1K result; add link to Disqus comment; fix minor typos + add tags
2/27/2024 - minor change in column header
3/19/2024 - minor change in column header; add link to GlutenID results
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.