Showing posts with label template. Show all posts
Showing posts with label template. Show all posts

Wednesday, November 6, 2019

Requiring (At Least Some) Methods Testing for Every Project

It may currently be a little hard to find, but I wanted to point out a couple links relevant to showing the value in testing different RNA-Seq methods for every project:


  • SourceForge repository for public data analysis
    • I still have a ways to go before being able to start working on a paper, but you can see how I am progressing here
    • I think the Target_Recovery_Status.xlsx file (for checking recovery of the known genetic perturbation in an experiment) is the most relevant for showing that you could not choose 1 method out of edgeR, DESeq2, and limma-voom to maximally recover the known gene knock-down or over-expression
    • I am also experimenting with having a completely public log for notes and analysis
  • Acknowledgement for GitHub RNA-Seq gene expression template
    • Includes some papers with modified methods
  • While the newer analysis tends to have smaller samples sizes, you can see noticable differences between methods in a much larger cohort in this post


On Biostars (which you can see in a variety of responses, including but not limited to this one), I would generally give the following recommendations:


  • If possible, test calculating p-values with edgeR, DESeq2, and limma-voom
  • I would recommend having an independently calculated expression method (like FPKM, Fragment Per Kilobase per Million), in order to help assess method selection
    • For example, you might see an extremely obvious change in expression for a gene (such as the one that you altered), but it might not have a significant p-value (or have a missing p-value) for one of the methods.
    • While the optimal strategy for discovery may not necessarily be the one that most stringently recovers previous results, you may be able to tell some strategies clearly don't work well on your data.
    • I would also recommend using this gene expression measurement to create heatmaps to compare clustering of replicates
      • I would typically use this instead of exporting normalized counts from the method to calculate the p-value, but testing clustering of replicates (without defining the groups in the normalization) is another possible way to compare strategies.
      • Sometimes this can be a bit qualitative.  However, if you define your gene lists / enrichment as a "hypothesis," then I think this is made up for my having independent validation for your claim.
      • I do realize this treads the line between p-hacking and needing to test methods due to limits in precision (which I mention a little bit in this comment and this Twitter discussion).  However, as scientists, I think this is part of why it is extremely important to be transparent and admit errors as soon as we discover them (in the interests of training ourselves to be as objective as possible).
  • Robustness of identifying a result with different methods may also give you some extra confidence in the results (unless the methods are not really independent, for example)
  • If you test alternative normalization, make sure you have a visualization before and after applying that normalization (to try and assess the likelihood of over-fitting in your adjustment)
  • I also think it is important that these are open-source, freely available programs (so that you can have the ability to determine what works best for your individual project)


In general, these posts may also be relevant to the discussion of limits to precision in the genomics methods:



Again, it is going to be a while, but I do hope to eventually have a preprint to cover the above points (as well as some other observations that I have had from working on a variety of projects for RNA-Seq gene expression analysis).

Change Log:

11/6/2019 - public post
6/3/2020 - add link for earlier (larger) RNA-Seq benchmark
6/7/2022 - minor formatting change

Friday, July 26, 2019

Personal Thoughts on Collaboration and Long-Term Project Planning

Our opinions can change over time, and some long-term effects may not be noticeable until 5+ years of experience.

While I still don't think I have everything figured out, I am using this page to organize my thoughts on some topics that may be of use to the broader community.  I also hope that the update/change logs may also be helpful for giving credit to feedback from others during discussions.

Nevertheless, for these posts, I am going to try and focus on what I believe I understand most clearly:



Again, it is probably a little early for me to be giving advice (since I don't have a solution worked out for myself yet), but I hope sharing my experiences can be helpful to other people as I sort out the details for figuring out a sustainable workload for myself.  Having the patience to work on agreed processes step-by-step is also important, but I believe some of this information may be important for future changes (even if they don't occur in the immediate future).

To be clear, I very much enjoy working with collaborators as a Bioinformatics Specialist in a Core Facility.  So, while some of what I am saying indicates room for future improvement, I have an overall positive impression of my work environment and the researchers that I have worked with (who are passionate about helping other people).

Plus, even though I think some of this content is important for long-term discussions, I also want to emphasize that you can be genuinely proud for putting in your best effort to help people and there is some need for short-term support (such as a temporary difficulty in getting additional funding) or at least giving people the chance to think carefully about whether a more major transition is necessary.

Update Log:

7/26/2019 - public post date
7/29/2019 - trim down introductory paragraph
7/30/2019 - add link for maintenance / support, and modify preceding sentence
8/4/2019 - minor edit after some proofreading by a family member
8/6/2019 - minor changes
4/30/2020 - add link to code / data sharing details (either required or suggested)

Sunday, June 2, 2019

What's the difference between a "Pipeline" and a "Template"?

The process of understanding each step of analysis is important for presenting the final set of results, and the process of writing the code for that analysis can help you understand the methods better (and identify questions and/or room for improvement in your current code / understanding).

I have some templates for analysis, but I call them "templates" rather than "pipelines" because the code itself usually requires some modification.  While I think it is extremely useful to have packages for specific functions, you may find that having a pre-set pipeline doesn't quite produce publication-quality figures, and having a template for your own code (that is easier for you to change than somebody else) may make it easier to implement changes that come as the result of iterations of project discussions.  I have a note to this effect for most templates (such as the acknowledgement in the README for the RNA-Seq gene expression analysis "template," as well as a the 2nd post-publication comment for COHCAP, which stands for "City of Hope CpG island Analysis Pipeline").

These modifications can be important for semi-automated analysis, and it is possible that other people may find there are some situations where it can be useful to have templates for intermediate results.  However, I also believe there are some other factors for discussion that are worth taking into consideration:

  • Be careful not to increase the turn-around time for an initial result while increasing the total amount of time to get a paper to publication (or skip steps that could decrease the accuracy of the publication).  This can be particularly tricky if it takes a couple years to appreciate all the time required for follow-up requests.
  • Be aware of how other people will view your code, and what is appropriate to include in a publication. For example, even if it becomes appropriate to have "templates" for intermediate results (which I am not saying is necessarily true), taking time to understand your results is important for responsible research practices (regardless of the formality of what you are making public).  So, testing your code on multiple datasets (either within your lab, or public data from other labs) can be important for troubleshooting.
  • Unlike code published with a particular paper, the templates (by definition) are really designed with my own use in mind, and are much more difficult to support for other people (even within the same lab).  So, while looking at portions of the code may be helpful in generating ideas, support for "templates" can't really be provided (at least not in the same way as "pipelines" or packages for a particular step of analysis)
  • While it is increasingly important to provide code to help with reproducible and understanding of analysis, that code likely needs to be different for each paper (and time and energy will be required for supporting the separate code for each paper, for each lab).  While this is not exactly a pipeline (since that code will likely represent a template to be modified for other people's experiments, not run without any changes, or novel modifications), this is also not the same as saying you have a generalized template that you think is good to use on a large number of projects.

In short, I believe it is important to expect some testing for each project, in order to be the most confident in the results that you present.  One possible alternative to a "template" may be having code for public demo datasets (for training), but that is probably also not a complete solution.

Update Log:
6/2/2019 - original public post
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.