Showing posts with label Genes for Good. Show all posts
Showing posts with label Genes for Good. Show all posts

Sunday, August 4, 2019

Human Low-Coverage Sequencing is Mostly OK for Broad Ancestry and Relatedness

This is a subset of my notes from my Nebula lcWGS sequencing on GitHub (as well as a couple images from two sections with Genes for Good, for full probe RFMix as well as RFMix SNP-chip down-sampling):

Ancestry Predictions

Even though I think they should only provide continental ancestry results (kind of like the 1000 Genomes "super-populations"), the ancestry was roughly similar to my other results (indicating that I am mostly European, which is accurate).  Plus, I describe limits on the more specific assignments (for SNP chip data) in another blog post.

Nevertheless, if I use my imputed genotypes for RFMix chromosome painting, I get results that look roughly like my SNP chip analysis (which would be an improvement over the Genes for Good imputed SNPs, but comparable to the much smaller number of Genes for Good observed SNPs):


There is no plot for chrX, in part because there are no imputed genotypes for chrX.

For comparison, this is what the full set of observed Genes for Good probes looks like:



and this is what the larger set of imputed Genes for Good probes looks like:



In other words, I think the Nebula results are similar (or perhaps slightly worse) than the genotypes that were directly measured for Genes for Good SNP chip probes (which, by the way, are completely free to obtain),  but imputation process also caused some issues with the ancestry with the Genes for Good SNP chip probes.

However, to be fair, the loss of accuracy with imputation does seem be better than using observed measurements if you decrease the probes and/or reference samples enough.  Shown below is the effect of using only 66 reference samples and 15,924 probes: (20x reduction in 1000 Genomes unrelated reference set, and 18x reduction in probes from my Genes for Good SNP chip):



Although, to be fair again, I think the primary problem is arguably the number of reference samples.  Take a look if I use the same number of probes (15,924 probes), but I have the "full" set 1,329 unrelated reference samples:


However, to get something that looks more like the original result, I would argue you do need to arguably increase the probes as well.  For example, please note a similar plot below with 143,320 probes (2x reduction in the starting amount):



Given that I think this looks kind of similar to the Genes for Good imputed set, I think the original Nebula imputed result (with lcWGS) for broad-level ancestry is a reasonable match to the higher coverage results (all things considered).


Kinship / Identity-By-Descent (Close Family Relationships)

Similar to the IBD estimates that are posted within the Helix/Mayo GeneGuide GitHub section (since I was only provided a gVCF), I can test overall similarity between the imputed Nebula genotypes, 23andMe (CW23), Genes for Good (GFG), Veritas WGS (BWA-MEM Re-Aligned) with 77,072 genomic positions (plotting 1000 Genomes reference samples for comparison):



By this measure, you can also clearly see which samples come from the same individual (me).  However, there is a slight drop in the accuracy for the Nebula imputed values (with kinship values between 0.489181 and 0.489226, instead of between 0.499859 and 0.499962):

FAM1 ID1 FAM2 ID2 nsnp hethet ibs0 kinship
0 CW23 0 GFG 77072 0.605148 0 0.499962
0 Veritas.BWA 0 GFG 76310 0.605032 0 0.499865
0 Veritas.BWA 0 CW23 76310 0.605032 0 0.499859
0 Nebula 0 GFG 77072 0.584596 7.78493e-05 0.489209
0 Nebula 0 CW23 77072 0.584596 7.78493e-05 0.489181
0 Nebula 0 Veritas.BWA 76310 0.58472 6.55222e-05 0.489226
In other words, there is some loss in the genome-wide similarity using low-coverage Whole Genome Sequcing (lcWGS), but you can still clearly tell which samples all same from the same individual (me).

However, if the underlying data is not reliable for traits and health results, then my opinion is that the SNP chip is still the relatively better option (as something that costs less than higher coverage sequencing, while giving ancestry /  relatedness results that are at least as good).


That said, to be fair, I think it really could be best if I could see similar example from those with a different primary (Non-European) ancestry.  For example, there was this New York Times article about someone whose broad-level ancestry assignments were less accurate than mine (although there was also evidence for improvement over time).  For example, most people were predicted to be mostly European (regardless of their actual ancestry) and most customers were European, then you could have a result that looks good even though it wasn't actually a very good predictor (beyond a baseline, like "assume everybody has European ancestry").

Update Log:

8/4/2019 - public post date
8/6/2019 - minor changes
8/15/2019 - minor changes
8/21/2019 - add "human" to the title
9/15/2019 - add warning / note about others with a different main broad ancestry

Predicting HLA Types for Array and High-Throughput Sequencing Data

My previous link to my HLA-assignments with varying technologies has the most important table in the middle of the page.  So, I am mostly reproducing that here to make the information easier to view.


SNP2HLA HIBAG bwakit HLAminer
HLA-A A*01, A*02
(23andMe)

A*01, A*02
(Genes for Good)

A*01, A*02
(AncestryDNA)
A*01, A*02
(23andMe)

A*01, A*02
(AncestryDNA)
A*01, A*02
(Genos Exome BWA-MEM)
A*01, A*02
(Genos Exome BWA-MEM)

A*01, A*68
(Genos Exome BWA)
HLA-B B*08, B*40
(23andMe)

B*08, B*40
(Genes for Good)

B*08, B*40
(AncestryDNA)
B*08, B*40
(23andMe)

B*08, B*40
(AncestryDNA)
B*08, B*40
(Genos Exome BWA-MEM)
B*08, B*40
(Genos Exome BWA-MEM)

B*08, B*41
(Genos Exome BWA)
HLA-C C*03, C*07
(23andMe)

C*03, C*07
(Genes for Good)

C*03, C*07
(AncestryDNA)
C*03, C*07
(23andMe)

C*03, C*07
(AncestryDNA)
C*03, C*07
(Genos Exome BWA-MEM)
C*03, C*07
(Genos Exome BWA-MEM)

C*03, C*07
(Genos Exome BWA)
HLA-DRB1 DRB1*01, DRB1*03
(23andMe)

DRB1*01, DRB1*03
(Genes for Good)

DRB1*01, DRB1*03
(AncestryDNA)
DRB1*03, DRB1*11
(23andMe)

DRB1*03, DRB1*15
(AncestryDNA)
DRB1*04, DRB1*04
(Genos Exome BWA-MEM)
DRB1*01, DRB1*15
(Genos Exome BWA-MEM)

DRB1*01, DRB1*15
(Genos Exome BWA)
HLA-DQA1 DQA1*05, DQA1*05
(23andMe)

DQA1*01, DQA1*05
(Genes for Good)

DQA1*01, DQA1*05
(AncestryDNA)
DQA1*05, DQA1*05
(23andMe)

DQA1*01, DQA1*05
(AncestryDNA)
DQA1*03, DQA1*03
(Genos Exome BWA-MEM)
DQA1*02, DQA1*03
(Genos Exome BWA-MEM)

DQA1*02, DQA1*03
(Genos Exome BWA)
HLA-DQB1 DQB1*02, DQB1*05
(23andMe)

DQB1*02, DQB1*02
(Genes for Good)

DQB1*02, DQB1*05
(AncestryDNA)
DQB1*02, DQB1*03
(23andMe)

DQB1*03, DQB1*06
(AncestryDNA)
DQB1*03, DQB1*03
(Genos Exome BWA-MEM)
DQB1*02, DQB1*03
(Genos Exome BWA-MEM)

DQB1*02, DQB1*03
(Genos Exome BWA)

In other words, my HLA-A / HLA-B / HLA-C types could be identified more robustly than the HLA-D genotypes (which I don't know, since I haven't gotten a regular blood test).  However, my understanding is that those types have a greater priority in defining organ transplant matches (although I'm currently encountering some difficulty finding the reference for that).

The GitHub link also goes a little deeper into how 23andMe is using 2 SNPs to represent 2 haplotypes (across genes) for celiac disease (which I found surprising, but that is done for other diagnostics as well).  I am mostly leaving that out of this section, but I did think it was interesting that HLA was used in 23andMe's "Meet Your Genes" when the SNPs are actually intronic / intergenic (with respect to the RefSeq annotations).

My 23andMe report indicated that I was DQ8-positive but DQ2-negative for my celiac disease risk.  In terms of defining the 2 genes used to define my DQ8-positive status I coloring matching assignments above in magenta (HLA-DQA1*03 and HLA-DQB1*0302).

Here is a screenshot for the variants tested by 23andMe (where I have the "C" variant for rs7454108, for the marker described as "HLA-DQ8"):



Again, as described here, a positive HLA-DQ8 status is defined by having HLA-DQA1*03 and HLA-DQB1*0302.

More recently, I collected Illumina Whole Genome Sequencing data where unaligned reads were provided (from Sequencing.com), along with some amount of PacBio HiFi data from Dante Labs.  There are some parts of the results that are not especially clear to me and I am interested to learn about additional options for analysis.  However, I believe those results are consistent with me having at least one DQB1*03 allele.

In terms of what appears to be consistent between the PacBio data and Illumina Whole Genome Sequencing data, the T1K results from the Sequencing.com Illumina reads indicates that the related HLA-DQB1 allele should be HLA-DQB1*03:02:01.

I ordered additional GlutenID testing from Targeted Genomics, with an uploaded subfolder on GitHub.  However, I believe the potential problem with using this for validation is that the DQ8-DQ8 result is based upon the same 1 SNP as my 23andMe result (rs7454108).  So, I will continue to look into additional validation options.

I am interested to learn more about the broader trends if certain HLA types are harder to assign and/or impute than other HLA types.  I have some notes mentioned in this Disqus comment.  Comments containing relevant feedback is also welcome on this blog post.

Update Log:

8/4/2019 - public post date
8/6/2019 - minor changes
8/15/2019 - add coloring for HLA-DQ8
2/11/2024 - add information / links to Whole Genome Sequencing data (Illumina from Sequencing.com and PacBio HiFi from Dante Labs)
2/18/2024 - add screenshot from 23andMe + dbSNP link; add additional HLA-DQ8 sentence; add small paragraph for Illumina WGS T1K result; add link to Disqus comment; fix minor typos + add tags
2/27/2024 - minor change in column header
3/19/2024 - minor change in column header; add link to GlutenID results
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.