Showing posts with label 16S rRNA. Show all posts
Showing posts with label 16S rRNA. Show all posts

Sunday, December 12, 2021

Human Metagenomics Comp 2021/2022: Overall Summary

Inspired by an earlier post that I don't believe is currently accessible (from several years ago), I tested collecting stool samples at the same time for multiple companies.  This sometimes includes the same sample submitted the same company at the same time.

The sample collection can be summarized as follows:

 

Psomagen

Thryve

Viome

Stool 1

(3/11/21)

1 Gene & GutBiome

1 GutBiome+

 

2 samples

Stool   2

(5/3/21)

1 GutBiome+

1 sample

1 sample

Stool   3

(6/27/21)

1 GutBiome+

2 samples

1 sample

Stool   4

(10/6/21)

1 GutBiome+

1 sample

1 sample

Stool   5

(12/11/21)

1 Kean Gut

1 Ombre

 

Stool   6

(5/6/22)

1 Ombre

 1 sample

 There are additional details (including the reports and data) on GitHub.


The sequencing performed can be described as follows:

Psomagen GutBiome (within combined "Gene & Gutbiome"): 16S (V3+V4 region, PE300 reads)

Kean Gut: 16S (V3+V4 region, PE300 reads)

thyrve/Ombre: 16S (V4 region, PE150 reads)


Psomagen GutBiome+: "Shotgun" Metagenomics (DNA-Seq, PE150 reads)

Kean GutBiome+: "Shotgun" Metagenomics (DNA-Seq, PE150 reads)

Viome: Metatranscriptomics (RNA-Seq, unknown read length)


As verified by Ombre technical support, you can view an alignment from my 5th paired sample here.  The other V4 sequences were downloaded from the GitHub repository for code associated with Johnson et al. 2019.

I have a couple paired posts (here and here), but I hope this can provide a fairly good summary of my overall experiences:

Individual assignments are not provided within the web interface for Kean Gut, but assignments can be made with re-analysis of the raw data and you can see some of such analysis with mothur here and Kraken2/Bracken here.



Raw Data Return:

1) thryve/Ombre - automatically provides raw FASTQ files as well as table with classifications at various levels

2) Psomagen/Kean - provides FASTQ files if e-mailed (but not automatically)

3) Viome - does not currently provide raw data (even if e-mailed) and classifications have discrete assignments (not percentages of reads)


When raw data was available (either automatically or by request), I have uploaded the data in public links on Google Cloud.  You can see a table of files to download here.

The samples and companies / organizations are different.  However, if it might help to see metagenomics samples that were collected before any of the samples described in this blog post, then you can see links to download raw data here.



Post-Collection Bacterial Growth Suppression:

1) thryve/Ombre - A liquid is included, but I have not yet verified the contents of that liquid.

2) Viome - liquid with preservative ("[bacteria] are not being killed nor growing").  I am not 100% sure about the implications for using RNA to study "active" bacteria, but I think it should help over adding nothing.

3) Psomagen/Kean - no liquid in collection tube to prevent / suppress bacterial growth

I only collected 1 BIOHM sample (10/6/2021), but there was no liquid (and therefore nothing to prevent / suppress post-collection bacterial growth)



Sample Collection Options:

1) Viome - originally provided 2 sizes of stool collector, but I think this is now reduced to 1 stool collector.

2a) Psomagen - 1 stool collector

2b) Kean -  tissue paper (0 stool collectors, if collected by itself)

3) thryve/Ombre - tissue paper (0 stool collectors, if collected by itself)



Cost:

1) Psomagen - $149 (with discounts - I paid $83.49, including $8.99 shipping, for 1 of my samples)

Kean splits options into ability to purchase separate Gut Health (for $99) and Gut+ Health (for $169)

2) thryve/Ombre - $199 (with discounts - I paid $99 before taxes, with free shipping, for at least 1 of my samples)

If you count the re-test discount (and that is still offered through Ombre), then I paid as low as $74.32 for 1 kit (with taxes).

3) Viome -  $299 (with discounts - I paid $129 before taxes, with free shipping, for at least 1 of my samples)



Result Turn-Around Time:

I think there is a limitation or lowest raking for each company, depending upon whether you define "turn-around time" for the kit, the results, the raw data, or for answering questions.

If it takes weeks to receive results that I think should mostly be considered hypotheses, then I don't think the slightly faster arrival of materials is really helping very much.  However, I consider returning raw data for re-analysis and answering questions from consumers to be very important.



Robustness of Results:

1) thryve/Ombre

2) Psomagen/Kean

3) Viome - noticeable variability results for samples collected at the same time

The explanation here is somewhat complicated.

As explained in a paired post, the signature/scores for collections from the same sample were better for thryve than Viome.  There are also some extra examples related to variability in the Viome results in this other paired post, but that is only for Viome.


Viome does not provide raw data and the data collected is different. So, only 2 signature/scores were available for thryve.  The Psomagen Gene & GutBiome kit used a different library design than the Psomagen GutBiome+ kit, so I don't have the replicates from the same sample that I intended.

There is also some information that I manually extracted from the 3 reports here (as well as on the highest level subfolder).  As mentioned earlier, you can also see some re-analysis of raw data (for Psomagen/Kean and thryve/Ombre, here and here).

Overall, I submitted an FDA MedWatch report Viome (for the currently commercially available tests described in this blog post), where the full draft is available to view here.  You can see the de-identified version in the MAUDE database under MW5106218.


To be fair, thryve/Ombre also have some food predictions that I would not place too much emphasis on.  However, Viome clearly has less consistency than the thryve/Ombre (positive and negative) recommendations.  If you download the thryve/Ombre PDF summary tables, then you should be able to access the links to the reports on Google Cloud (because they were too large for GitHub).

If you specifically purchase Kean Gut+, then there are some food recommendations.  I have not made changes based upon the results from any company, and I am not making any changes based upon any specific feedback from Kean either.  However, there were at least no recommendations to avoid eating food that I already find helpful or favorable.  I think the food recommendations also seem like mostly good ideas, regardless of any particular metagenomic result.  On the other hand, I am not certain what to say about the supplement recommendations for either Kean Gut or Kean Gut+.



A Note About Probiotic Detection:

The reason that I ordered a Viome sample for my 6th (but not 5th) paired sample is that I better understood the difference between the Kean Gut and Kean Gut+ kits, and I wondered if it was possibly useful to have a paired sample with untargeted DNA-Seq (for Kean Gut+) along with a Viome sample.

Viome still does not provide raw data, and I believe they only provide discrete specific assignments.  The Viome interface has changed recently, but I was still able to see those in a PDF file when I e-mailed myself to "Share My Results"  (under "Scores," and then at the bottom of the details for one of the scores)

Nevertheless, I could compare Kean Gut+ assignments and the "Active" status for Viome, and most assignments matched in that sense (everything above 1% listed in the table on this page match).

The one exception where I can have both Kean Gut and Kean Gut+ information for individual bacteria is for probiotics.  Lactobacillus was copied over from earlier collections, but I thought that might be interesting in that Ombre and Viome reported detection when both Kean Gut and Kean Gut+ reported a lack of detection.

I thought this was somewhat interesting because I have Activia with lunch on Monday to Friday.  The Wikipedia page mentions that other common probiotics are also present, but Bifidobacterium is mentioned as something specific to Activia.  So, I went back to check the original reports and update the GitHub table for the 6th paired collection, and all 3 companies report Bifidobacterium detection.

There are non-zero mothur read fraction assignments for Bifidobacterium for all samples except Psomagen for the 3rd paired collection (Psomagen3), and I can tell that Kraken2/Braken is also capable of making those assignments at the genus level.  I don't think this indicates what is present is specifically the proprietary strain, but I might guess eating yogurt on a regular basis might have some level of contribution.  I also don't think the bacteria has to be that proprietary strain in order to be helpful.



Closing Thoughts:

In general, my opinion is that is better to focus on a smaller number of things that are truly predictive than attempt to make a large number of claims (and have a lot of them not be valid).  I also think having access to raw data for re-analysis is very important.


I selected Psomagen because it was the company that purchased assets from uBiome.  However, there are notable differences in what Psomagen provides (such as no longer using a liquid to prevent post-collection bacterial growth).  Even though there were other ethical problems for uBiome, I think stopping post-collection growth was a good idea.  I also previously had the ability to sample multiple sites from uBiome, but that was not currently an option.  So, I don't think these company purchases/acquisitions necessarily mean that you can expect the same product.


I think I might have purchased a Viome kit because of an advertisement, but I believe that I am less likely to make future Viome purchases (if similar to what is currently provided).  I also do not plan to make additional BIOHM purchases.  My overall impression for Psomagen/Kean was somewhere in the middle, but I would guess that I am most likely to purchase Ombre in future.  For example, I posted Trustpilot reviews for all 3 companies: 4 stars for Ombre, 3 stars for Kean, and 2 stars for Viome.  However, to be clear, that is largely because of the automatic return of the raw data and the presence of some sort of liquid that I assume helps reduce post-collection growth.  That does not mean that I like or approve everything within the Ombre results.


For example, I am not sure if I support the Ombre supplement recommendations, and I don't remember seeing anything that I considered particularly helpful among the food recommendations.


From a technical standpoint, my critiques were mostly for Viome.  Of course, anybody can submit FDA MedWatch reports for a technical issue with a diagnostic (as I did).  However, if other consumers do plan on taking supplements recommended by the company, then I hope that they use FDA MedWatch if they encounter any adverse events.  I did this for an earlier human genomic test, but I am hesitant to try recommendations from the multiple other companies.  So, I hope that customers for all 3 companies are aware of resources like this and take the time to provide important feedback as relevant.


Unfortunately, I don't think have enough specialized background to gauge the relative importance of various options for specific diseases (such as those offered by a given company, versus other options that might even be free).  So, if you are well aware of the common best practices (as either a physician or possibly as a patient with a chronic condition), then please provide feedback if you notice any claims that might either over-emphasize the benefit of a supplement from a given company or if enough emphasis on other available/established options is not adequately described.  I think there are reporting systems for advertising, but I am not sure of those can be recommended by somebody else.  I think critical evaluation of claims is important, and I think e-mails to the company and discussion with your general or specialized physician are probably a good starting place.


Again, I did not make any changes based upon any results from any of the companies.  I only used this for research purposes to gain a better appreciation for the bacterial metagenomic analysis, which would allow me to do things like study microbiome changes over time.  In the future, I might also make changes based upon independent physician advice and see what happens in my metagenome (mostly as a matter of curiosity), but I consider that different than starting from recommendations based upon the metagenomic data.


I hope others find this helpful, and I would certainly encourage feedback! I think the blog post comments section has been OK in the past.  However, if there is any possible benefit to using the GitHub discussion (which can include images and code), then that is also an option.


Change Log:
12/12/2021 - public post date
12/13/2021 - various wording revisions/corrections
12/14/2021 - change wording for last change log entry; additional revisions
12/19/2021 - add Viome response + additional links
12/29/2021 - thryve/Ombre food recommendations
1/22/2022 - add FDA MedWatch report
5/2/2022 - change title and add additional sample information
6/18/2022 - after receiving results from all companies, add collection date for 6th sample.
6/19/2022 - add links for raw data
6/27/2022 - revise closing thoughts and add some additional details (such as re-analysis and Bifidobacterium detection)
6/28/2022 - add link to GitHub disucssion
7/16/2022 - add Trustpilot review links, and re-arrange content.

For example, I have moved the content below out of the main post for brevity, but I am keeping the comments below for reference:

When I asked Viome for feedback regarding the extra discordance that I believe I saw in my samples (including those collected at the same time), I was directed to this page.  I am not sure if I see the assessments currently being provided to consumers.  However, if that represents the response, then perhaps it can be mentioned that I think it is important to emphasize that the "FDA Breakthrough Device Designation" for a different application that I believe is not available to consumers is not in fact FDA approved (matching the bar chart in the provided link, if you look carefully - however, it was an issue that I separately reported for regulatory misconduct, with draft available here that is a follow-up from discussion with other FDA that informed me about the regulatory misconduct reporting system). That said, I wish to emphasize that it is important that each claim be evaluated independently (within and between tests).

...

I believe this article references low consumer reviewers for Viome.   As of 12/19/2021, I see 1.28 out of 5 from the BBB, and 3.2 out of 5 from Trustpilot.

Human Metagenomics Comp 2021/2022: Divergence for Signatures / Scores for Same Company and Same Sample

 This is a subset of content available on GitHub, focusing on samples collected from the same company at the same time.  This also includes code to reproduce plots and statistics.

Again, I took no action based upon any of these reports.



Intuitively, I thought something about the variability for the same sample for the Viome collections seemed high.  I think the plot above might help illustrate this.

You can also see more about the underlying details below:

Company: Signature / Score

Mean

SD

Viome: Active Microbial Diversity (Percentile)

6.5

7.8

Viome: Digestive Efficacy

64.0

9.9

Viome: Gas Production

38.0

11.3

Viome: Gut Lining Health

67.5

0.7

Viome: Gut Microbiome Health

45.0

5.7

Viome: Inflammatory Activity

38.5

3.5

Viome: Metabolic Fitness

28.5

0.7

Viome: Microbiome-Induced Stress

43.5

10.6

Viome: Protein Fermentation

35.0

9.9

thryve: Gut Diversity Score

93.5

0.7

thryve: Gut Wellness Score

81.0

0.0

 To be fair, some Viome signatures / scores may be relatively robust.  Also, in general, every individual claim needs to be investigated independently (and you not should draw conclusions about one analysis based upon what you saw for a different analysis, whether that is good or bad).  However, I think there is evidence to back up my conclusion that you should not follow all recommendations, and treat the overall set of scores/signatures as hypotheses that may or may not be absolutely true for specific sample.

On the original GitHub page, I think there is some additional evidence (looking at measurements for all 3 companies) that increased variability in similar samples makes it harder to detect a noticeable phenotypic difference in the 4th sample.  However, to be fair, I think that is less rigorous than the comparisons that I am making above.


Please click here to return to the overall summary.

Change Log:
12/12/2021 - public post date
12/13/2021 - various wording revisions/corrections
12/14/2021 - change wording for last change log entry; additional revision
5/2/2022 - modify title to match new main blog post

Monday, November 18, 2013

My American Gut Individual Report

I recently received my American Gut Individual Report for my fecal sample in the mail.  Click here to see this report as a PDF.

The first thing that I noticed was that the phylum distributions were very different than what I calculated from my raw data (click here to see those results).  Namely, ~75% of my reads aligned to Proteobacteria 16S rRNAs, but my individual report said that ~10% of my sample was from Proteobacteria.  So, I contacted the American Gut team to ask why the results were so different.  They said that it appears the shipping process allows differential growth of certain bacteria (especially Gammaprotoebacteria) that would not normally appear in fresh samples.  So, they filter out likely contaminants for the report.

This is something that I would like to learn more about, and the American Gut team also said that they are actively investigating this.  The filtering code is available on the American Gut GitHub website, but I am most curious in seeing population-level metrics about this differential growth.  Namely, I would like to see how the reduction in false positives is affecting the true positive rate.  At least in my case, ~70% of my reads are being filtered out as potential contaminants, which seems like a loss of a lot of information. Additionally, my filtered sample seems to cluster with samples that have relatively low Firmicutes counts (much closer to the original value of ~15%, see PCA plot in lower-right hand corner), when the report says the percentage should be ~65% Firmicutes (bar plot).  In other words, my guess is that the "best" interpretation of my results may lie somewhere between these two reports (where the true Firmuicutes abundance may be lower and the true Proteobacteria abundance may be higher).

That said, I know the American Gut analysis is ongoing, and there will likely be additional future reports.  For example, I know for certain that there will eventually be a report for my oral sample.  It will be interesting to see if the results of future reports change as new scientific findings are discovered.

Monday, October 21, 2013

Analyze Your 16S rRNA Data Using MG-RAST

Step #1: Register for an Account

You can use this link to register:http://metagenomics.anl.gov/?page=Register

Registration is not automated, so registration is not immediate.  It took about a day to create an account for me.

Once you receive an e-mail saying "MG-RAST - account request approved", you can sign-in for the next step.

Step #2: Upload Your Data

Go to the MG-RAST website: http://metagenomics.anl.gov/

I would recommend using Firefox - you will see a pop-up if you do not.

Sign into MG-RAST (username and password are entered in the upper-right hand corner of the screen).

Choose "Upload".  This is represented by a green arrow pointing upwards.  There should be one in the middle of your screen (which says "Upload") as well as in the upper-right hand corner of the screen (although this one is not specifically labeled).

The metadata step is not required.  I skipped this because I figure the American Gut data should eventually be entered into this database, and I didn't want to produce a duplicate dataset (and I probably didn't know all of the details regarding funding, sample processing, etc.).  However, I contacted the MG-RAST developers, and they actually encouraged me to make the sample public.  If you take the time to fill out the metadata for your sample, it will be processed more quickly.

Under "PREPARE DATA", Click "2. upload files" and browse for your FASTQ that you downloaded from ENA.  You will see a pop-up, but just click "close".  It is not necessary to complete the check.  You can click "3. Manage Inbox" to see when the upload is complete (if you want to wait a few minutes, you can keep clicking "update inbox" until the files are ready).  Otherwise, you can just do something else and come back later.

Step #3: Run the MG-RAST Pipeline

After the data has been uploaded, click "1. select metadata file" under "DATA SUBMISSION".  If you didn't create a metadata file, just click the box saying "I do not want to supply metadata" and click "select".

Click "2. select project".  You probably don't have an existing project, so just type in something like "American Gut" and click "select".

Under "3. select sequence file(s)", click the check marks next to the files that you want to analyze and click "select".

Unless you have some experience with metagenomic analysis, just select "4. choose pipeline options" and click "select";

Finally, choose a data submission option (if you don't provide metadata, you have to keep your data private) and click "submit job".

Step #4: Analyze Your Processed Data

The pipeline may take a while (at least a few hours and possibly as long as a week), especially if you are keeping the data private.  So, I would recommend doing something else and then signing back into MG-RAST.  You can check the status of your samples at any time by clicking the earth icon in the upper-right hand corner of the screen (or "Browse Metagenomes" in the middle of the screen).  There will be numbers next to different stages in the upper-left hand corner.  If you click the number next to "In Progress" and you see your samples, then they are not ready (but you can at least you can see where your samples are in the pipeline).  You need to be able to click the number next to "Available for Analysis" and then be able to see your samples in the next menu that is loaded.

Once your samples are available for analysis, click on the bar-plot icon in the upper-right hand corner of the screen.  There are a lot of options available for metagenomic analysis, but I will walk through what I think is the most useful analysis.

Under "Organism Abundance" on the left-hand side, click "Best Hit Classification".  Under "Data Selection", select your samples by clicking the "+" icon next to "Metagenomes".  If you left your samples as private, then they should be relatively easy to select.

Under "Annotation Sources", the default may be "M5NR".  I would strongly recommend you use a RNA database, such ad RDP, Greengenes, or M5RNA.  M5RNA is a little more interesting because it also contains Eukaroytic sequences, but I will mostly focus on RDP (so that I can compare the results to the RDP-Classifier).  Highlight the desired database click "OK".

At this point, you should have all the necessary configurations set up, and your screen should look something like this:



To analyze your selected data, click a radio button under "Data Visualization" and then click "generate".  I think the tree and table tools are the most useful.

Analyze Your 16S rRNA Data Using RDP-Classifier

Step #1: Convert Files from FASTQ to FASTA

There are lots of ways to do this, but I would recommend using Galaxy if you don't have any programming experience:

Go to the Galaxy website: https://usegalaxy.org/

If you are an academic researcher, your institution might have a local mirror (which should be faster).  However, the link above will work for everybody.

Upload your data using "Get Data" --> "Upload File" (the functions are available on the left-hand side of the screen).  You can set the file type to "fastq", but you probably don't need to.  Updates will appear on the right-hand side of the screen, so you know when each step is complete (the box for the corresponding step will turn green).

Go to "NGS: QC and manipulation" --> "FASTQ Groomer" (should be under "ILLUMINA DATA" in grey font).  Leave all the default settings and click "Execute".  This is technically necessary because of a formatting issue.

Go to "Convert Formats" --> "FASTQ to FASTA".  Once this step is complete, click the appropriate green box on the right-hand side.  Once the box becomes larger (allowing you to see the first few lines of the file), click the purple floppy disk icon to download the FASTA file.  I would recommend renaming the FASTA file after it is downloaded, so it is easier to keep track of.

Step #2: Create an RDP Account

You can sign up using this link: https://rdp.cme.msu.edu/user/createAcct.spr

An account will be created automatically.  You will receive an e-mail with a username and password (you will be asked to change your password the first time you sign in).  Technically, you don't need an account to run the classifier.  However, I think it may be helpful if you want to play around with some other tools.

Step #3: Sign-In and Run RDP-Classifier

Using the link provided by the registration e-mail, sign into myRDP.

Now, go to this link: https://rdp.cme.msu.edu/classifier/classifier.jsp

There will be an option to "Choose a file (unaligned format) to upload:".  Use the browser to select the FASTA (not FASTQ) file that you downloaded from Galaxy.  Next, click "Submit".

The classifier is very fast (you should get your results in a few minutes).  The result page is somewhat hard to parse, but everything is clickable to learn more.  The number of reads is shown in parentheses.

It really helps to remember biological classifications when interpreting these results.  Here is a quick cheat sheet:

phylum > class > order (> suborder) > family > genus

Unfortunately, the classifier won't provide species-specific information.

You can also download the results in a text file.  If you do this, you can use a tool like Notepad++ to search for keywords (like phylum, genus, etc.), but I think the results are a little easier to view on the webpage.

How to Download Your American Gut Data

Step #1: Find Your Barcode(s)

Each sample has a nine digit barcode.  If you see a smaller number, add leading zeros.

For example, my barcodes were 2683 (fecal) and 2684 (oral), so I need to use 000002683 and 000002684 as my sample IDs.

Step #2: Search For Your Sample

Go to the European Nucleotide Archive (ENA) website: http://www.ebi.ac.uk/ena/

Copy and paste your 9-digit barcode into the text search

I get a single result for each of my samples when I do this.  If you get multiple results, choose the metagenome sample (see image below).


If you want to double-check you have the right sample, click on the "Sample accession" link (which will start with "ERS").  If you then click the "Attributes" tab, you should be able to see your metadata.  For example, I know I live in Los Angeles, so my state better not be GA.

Step #3: Download Your FASTQ Files

If you are certain you have the right sample, click the the link for "Fastq files (ftp)" to start the download.  Note that the sample will be labeled based upon the "Run accesssion" (starting with "ERR").  For example, here are the different IDs for my samples:

Fecal: 000002683 --> ERS345317 --> ERR336561
Oral: 000002684 --> ERS344890 --> ERR336138

The .fastq files will be compressed, so you should unzip them.  I would recommend using 7zip for this.  I would also recommend renaming your files (like fecal.fastq and oral.fastq) to make it easier for you to keep track of them.

Open-Source Analysis of My Raw American Gut Data

Just like I used open-source tools to re-analyze my 23andMe data, I wanted to see what I could learn from analyzing the raw 16S rRNA reads from the American Gut Project.  The American Gut team is very well organized, and you can currently access your raw data in public databases.

I've provided a tutorial based entirely on web-based analysis, so you don't need to know any programming to follow these steps in your own data.  You can also skip the tutorial links to just see my own results.

Step #1: Get Your Data
Step #2a: Analyze Your Data in MG-RAST (preferred, but time-consuming)
Step #2b: Analyze Your Data using RDP-Classifier (quick, but less functionality)

Here are some charts that I could quickly create in Excel using data from RDP-Classifier:


If you compare these distributions to the average gut distribution from the American Gut preliminary report, then you can tell it is very different than the average participant.  Clearly, I have more Proteobacteria than the average participant (and less Firmicutes and Bacteroidetes).  However, it is also important to note that there was also a large amount of variation between participants.

I don't know how my class distributions compare to other samples, but it seems like I can at least infer that there is more variety in my oral sample than my fecal sample:


I also found it useful to see what specific genera were most highly abundant.  According to RDP-Classifer, these are the genera with more than 1000 reads:


Fecal Oral
Streptococcaceae Streptococcus 0 10614
Neisseriaceae Neisseria 0 8824
Actinomycetaceae Actinomyces 0 4129
Lachnospiraceae Oribacterium 0 2319
Veillonellaceae Veillonella  0 2124
Bacteroidaceae Bacteroides 1800 5
Enterobacteriaceae Escherichia/Shigella 8039 5

I was able to find reports that Actinomyces was associated with transformation of lymphocytes in patients with periodontal disease (Baker et al. 1976) and Streptococcus mutans plays a role in human dental decay (Loesche 1986), although I don't know if I had this specific strain (based upon the MG-RAST report, it looks like I don't).  Likewise, I showed this list to my dentist, and she recognized these two genera as being associated with dental problems.  Although I don't know how common these are overall, it seems to make sense that they could be found in an oral sample.

Obviously, I recognized Escherichia/Shigella.  The American Gut report points out that the phylum containing this genus is not highly abundant in an average participant.  At one point, I had to be hospitalized with ulcerative colitis (with a strain of E. coli producing shiga toxin), so perhaps this is related to the high abundance of this genus (although that was several years ago).

If you are patient enough to wait for your MG-RAST results, then you can make similar (but slightly cooler looking) pie charts and tables automatically.  For example, you can take a look at the corresponding plots from MG-RAST for my fecal and oral samples.  You can also create plots to compare species in multiple samples (red bars are for my oral samples, green bars are for my fecal sample):





Perhaps most importantly, MG-RAST will provide annotations down to the species level (and strain level, when possible).  The species counts aren't perfectly correlated with the genera counts (predicted from the classifier), but the the most interesting genera appeared in both lists.


metagenome
strain
abundance
oral
uncultured bacterium
19715
fecal
Escherichia coli ED1a
11316
oral
Abiotrophia para-adiacens
6185
oral
Actinomyces odontolyticus
4154
oral
Veillonella dispar
1913
oral
Butyrivibrio fibrisolvens
1431
oral
Blautia sp. Ser8
1075
fecal
Bacteroides stercoris
736
oral
Syntrophococcus sucromutans
661
fecal
Pseudomonas fluorescens
660
fecal
Bacteroides caccae
559
fecal
Bacteroides stercoris ATCC 43183
558
fecal
Bacteroides vulgatus
510
oral
Streptococcus
499
oral
Rothia mucilaginosa
450
fecal
uncultured bacterium
370
oral
Haemophilus haemolyticus
295
fecal
Prevotella buccalis
287
oral
Ruminococcus gauvreauii
282
oral
Streptococcus sanguinis
275
oral
Ruminococcus torques L2-14
272
fecal
Escherichia coli
232
oral
Abiotrophia defectiva
213
oral
Gemella morbillorum
204
fecal
Dialister propionicifaciens
184
oral
Veillonella parvula
175
oral
Butyrivibrio hungatei
142
oral
Atopobium minutum
137
oral
Parvimonas micra
126
oral
Leptotrichia shahii
121

This information can allowed to conduct more effective literature searches.  For example, my understanding is that the ED1a strain has not been shown to be associated with ulcerative colitis.  On the other hand, the species information allowed me to find a paper for the discovery of my specific strain of Actinomyces, which was harvested from 450 tooth cavities (Batty 2005).  Likewise, I could confirm that Streptococcus sanguinis was also pathogenic (Xu et al. 2007).

FYI, Galaxy also has some metagenomic tools.  However, running BLAST on Galaxy will take a long time.  If you are comfortable with running BLAST locally, it should be easier (but this requires some comfort using the computer).  You can also analyze your data locally using QIIME upload the results to PICRUSt for functional enrichment, if you don't need the convenience of the web-based tools listed above (MG-RAST can produce QIIME reports, but I think it is better to use a tab-delimited text file to avoid formatting problems).

I am still interested in seeing what my official individual report will look like: although I have general experience with bioinformatics analysis, the folks at American Gut have looked at a lot more metagenomic data than I have.  Likewise, I am interesting in seeing how my profiles change at different time points: once I eventually get my uBiome results, I will put together another post to compare the results.
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.