Showing posts with label edgeR. Show all posts
Showing posts with label edgeR. Show all posts

Wednesday, November 6, 2019

Requiring (At Least Some) Methods Testing for Every Project

It may currently be a little hard to find, but I wanted to point out a couple links relevant to showing the value in testing different RNA-Seq methods for every project:


  • SourceForge repository for public data analysis
    • I still have a ways to go before being able to start working on a paper, but you can see how I am progressing here
    • I think the Target_Recovery_Status.xlsx file (for checking recovery of the known genetic perturbation in an experiment) is the most relevant for showing that you could not choose 1 method out of edgeR, DESeq2, and limma-voom to maximally recover the known gene knock-down or over-expression
    • I am also experimenting with having a completely public log for notes and analysis
  • Acknowledgement for GitHub RNA-Seq gene expression template
    • Includes some papers with modified methods
  • While the newer analysis tends to have smaller samples sizes, you can see noticable differences between methods in a much larger cohort in this post


On Biostars (which you can see in a variety of responses, including but not limited to this one), I would generally give the following recommendations:


  • If possible, test calculating p-values with edgeR, DESeq2, and limma-voom
  • I would recommend having an independently calculated expression method (like FPKM, Fragment Per Kilobase per Million), in order to help assess method selection
    • For example, you might see an extremely obvious change in expression for a gene (such as the one that you altered), but it might not have a significant p-value (or have a missing p-value) for one of the methods.
    • While the optimal strategy for discovery may not necessarily be the one that most stringently recovers previous results, you may be able to tell some strategies clearly don't work well on your data.
    • I would also recommend using this gene expression measurement to create heatmaps to compare clustering of replicates
      • I would typically use this instead of exporting normalized counts from the method to calculate the p-value, but testing clustering of replicates (without defining the groups in the normalization) is another possible way to compare strategies.
      • Sometimes this can be a bit qualitative.  However, if you define your gene lists / enrichment as a "hypothesis," then I think this is made up for my having independent validation for your claim.
      • I do realize this treads the line between p-hacking and needing to test methods due to limits in precision (which I mention a little bit in this comment and this Twitter discussion).  However, as scientists, I think this is part of why it is extremely important to be transparent and admit errors as soon as we discover them (in the interests of training ourselves to be as objective as possible).
  • Robustness of identifying a result with different methods may also give you some extra confidence in the results (unless the methods are not really independent, for example)
  • If you test alternative normalization, make sure you have a visualization before and after applying that normalization (to try and assess the likelihood of over-fitting in your adjustment)
  • I also think it is important that these are open-source, freely available programs (so that you can have the ability to determine what works best for your individual project)


In general, these posts may also be relevant to the discussion of limits to precision in the genomics methods:



Again, it is going to be a while, but I do hope to eventually have a preprint to cover the above points (as well as some other observations that I have had from working on a variety of projects for RNA-Seq gene expression analysis).

Change Log:

11/6/2019 - public post
6/3/2020 - add link for earlier (larger) RNA-Seq benchmark
6/7/2022 - minor formatting change

Tuesday, November 19, 2013

RNA-Seq Differential Expression Benchmarks

I recently published a paper whose primary purpose was to serve as a reference for the protocol that I use for RNA-Seq analysis (see main paper and supplemental figures).

The aspect of the paper that I think is most interesting to the genomics community is a comparison of statistical tools for defining differentially expressed genes, which had the greatest influence on the resulting gene lists (at least among the comparisons that I make in the paper).  So, I will review those relevant figures in this blog post.

The plots below show the robustness of the gene lists produced by a given algorithm.  In other words, the higher the "common" line on the graph, the more robust the gene lists (i.e. the higher the proportion of genes commonly called by multiple algorithms).  Most readers will probably not be as interested in the x-axis (rounding factor for RPKM values), and it only changes the gene lists for Partek and sRAP.
Analysis of Patient Cohort (Tumor versus Normal).  1-factor is just tumor versus normal, while 2-factor also includes patient ID (pairing tumor and normal samples).  cuffdiff results not shown because no genes were defined with FDR < 0.05.  sRAP not shown because gene list was very small (see Figure S3 from the paper)
Analysis of Cell Line Comparison (Mutant versus WT)
To be fair, I will certainly admit robustness is not the same as accuracy.   Uniquely identified genes may be true positives that represent a lower false negative rate.  However, this did correspond to some circumstantial evidence I've seen with other datasets where cuffdiff and edgeR have given some weird results.  The results from this paper don't actually contain the clearest examples of this, but you can take a look at the GAGE4 stats to see an example where I would at least argue that edgeR provides inflated statistical significance.

Overall, I think Partek works the best (which is what I use for COH customers), but I was also pleased with DESeq (and sRAP, but I am obviously biased).  In fact, these comparisons support earlier observations that DESeq is conservative in defining lists of differential expressed genes (Robles et al. 2012).

However, my main goal is not to simply tell you what is the single best solution.  In fact, the cell line comparison above also had paired microarray data, and I would say the concordance between the two technologies was roughly similar for most algorithms:
RNA-Seq versus Microarray Gene lists.  "Microarray DEG" = proportion of differentially expressed genes in microarray data also present in RNA-Seq gene list.  "RNA-Seq DEG" = proportion of differentially expressed genes in RNA-Seq data also present in microarray gene list.


The similarity in microarray concordance kind of reminds me of Figure 2a from Rapport et al. 2013, which compares RNA-Seq gene lists to ~1000 qPCR validated genes.  However, I think properly determining accuracy can be difficult.  For example, look at the differences between the qPCR results in Figure 2a and the ERCC spike-ins in Figure S5 for that same paper.

Instead, these are the main take-home points I would like to emphasize:

1) Simple methods comparing RPKM values (in this case, rounded and log2 transformed) for defining differentially expressed genes can work at least as well as more complicated methods that are unique for RNA-Seq analysis (at least for gene-level comparisons).  For example, one claim against count-based methods in general (including edgeR, DESeq, etc.) is that there can be confounding factors, such as changes in splicing patterns.  Although I agree this is a theoretical problem that probably does occur to some extent, it doesn't seem to be a major factor influencing concordance with microarray data, qPCR validation, etc.

2) There is probably not a solution that works best in all situations. In this paper, you can see the results look very different with the patient versus cell line datasets.  For practical reasons, a lot of benchmarks will probably use cell line datasets.  However, it is not safe to assume performance for large patient cohorts will be comparable to cell line data (or patient data with little or no biological replicates).
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.