Showing posts with label eXpress. Show all posts
Showing posts with label eXpress. Show all posts

Thursday, April 17, 2014

Differential Expression Without A Reference Genome

I've noticed a lot of Biostar questions related to conducting differential expression following de novo assembly of RNA-Seq data, so I wanted to create a blog post with a collection of my suggestions.

It is tempting to want to map the assembled transcripts between samples for differential expression, but I wouldn't recommend this because there will often not be a 1:1 mapping between assembled transcripts in different samples.  Instead, these are my suggestions:

Differential Expression Strategies:

  • Use one assembly (either from a control sample or a pooled collection of reads from all samples).  Then, use a reference based alignment (using an aligner like Bowtie or BWA) against this assembly for each sample.  You can perform mRNA quantification using a tool like eXpress, and then you can use your favorite differential expression tool (I would recommend DESeq or limma, among the popular options)
  • Use a kmer-based option (like NIKS, RUFUS, etc.).  Here, the idea is to look for differentially represented kmers and then perform de novo assembly on only the kmers that differ between samples/groups.
  • I haven't tried it, but Corset looking like an interesting option


Tips:

  • I actually found CLC de novo to be the best de novo assembly tool, even though it wasn't specifically designed for RNA-Seq data.  It also automatically provides contig coverage statistics.
    • In my case, I defined the quality of the results based upon the most highly expressed contigs in various tissues (looking for genes that you could expect to be highly expressed in those different tissue types)
  • Among the open-source RNA-Seq de novo assembly options, I would recommend Oases.  In fact, you might find the merged Velvet contigs to be more useful than the transcripts (either way, you will have access to both options)
  • I would recommend against using Trinity, even though that is a popular option.  Based upon my personal experience, I would say that it often stitches together contigs from different genes, producing many very large transcripts (some fusion genes should occur, but not at the rate I saw large transcripts in Trinity)


Relevant Biostar Posts:

Tuesday, February 18, 2014

mRNA Quantification via eXpress

eXpress is a tool that allows mRNA quantification using a set of transcripts as a reference (this is opposed to popular RNA-Seq tools like TopHat, which align reads to a genome and have to model gaps caused by exon junctions).

Using transcripts rather than genomic chromosomes as a reference sequences is actually how I imagined RNA-Seq analysis would be conducted, before I learned about standard practices.  In fact, samtools provides an 'idxstats' function that can be used to calculate normalized RPKM expression values.  So, I was curious if the extra modeling done by eXpress is really any better than this simple sort of RPKM calculation: having a more complicated model can potentially improve accuracy, but more complicated models can also leave extra room for things to go wrong, can lead to over-fitting, etc.  For example, I have used eXpress on some de novo assembly data, and I actually found that normal de novo programs seemed to provide better results than those specifically designed for RNA-Seq data (however, to be clear, I think the results of this blog post emphasize that the problem was with the assembly and not the mRNA quantification, as I would have expected).

The short answer is "Yes" - I think it is better to use eXpress over idxstats for calculating RPKM/FPKM values.

To illustrate this, first take a look at the correlations between the eXpress FPKM values and the RPKM values calculated using idxstats:



The correlation isn't horrible, but you can see a non-trivial amount of genes whose expression levels have consistently lower in eXpress than idxstats.  However, this by itself doesn't really prove one options is better than the other option.  Because I feel comfortable with the gene-level mRNA quantification levels from cufflinks (and the RSEM-like algorithm implemented in Partek; for example, see Figure 5 in this paper or click here to see a direct correlation between these two results), I decided to see how the results compared when using different tools for a transcript-based reference (eXpress, idxstats) versus a genomic/chromosome-based reference (cufflinks, Partek).

Again, you see these outliers if you compare the idxstats results to cufflinks (or to Partek - click here for those results):



However, you don't see these outliers when comparing eXpress to cufflinks (or to Partek - again, click here for those results):



So, eXpress clearly provides more robust results than the simpler idxstats comparison.  You can also see this in box plot below, showing the correlation coefficients for all the mRNA quantification strategies that I tested.



Of course, systematic differences between mRNA quantification methods should (at least partially) be corrected when identifying differentially expressed genes between two groups (because the differences affect both groups).  However, there are some certain circumstances when the mRNA quantification levels may want be used in isolation, such as for ranking the most highly expressed genes in a sample (as was the case for the de novo assembly data that I worked with).  In this situations, I would definitely recommend a tool like eXpress over trying to calculate RPKM values from tools like idxstats.

FYI, here are some details on the methodology for this comparison:
  • MiSeq samples from GSE37703 were used for these comparisons.
  • Correlations were calculated using log2(FPKM/RPKM + 0.1) expression values.
  • eXpress and idxstats were run on Bowtie2 alignments of the same set of RefSeq transcripts (downloaded from the UCSC Genome Browser, with duplicated gene IDs removed).  The Partek EM algorithm used a set of RefSeq sequences used by the vendor and cufflinks used the genes.gtf file downloaded from iGenomes on the TopHat website.  Only commonly represented gene symbols were used for calculating correlations.  Only genes declared "solvable" by eXpress were considered for calculating correlations.  As an example, click here to view a venn diagram of overlapping gene symbols for SRR493372.
P.S. It looks like you may have to be signed into Google Docs to view the image previews properly.  However, you can always download the files to view them locally.
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.