Showing posts with label COHCAP. Show all posts
Showing posts with label COHCAP. Show all posts

Tuesday, July 30, 2019

Personal Thoughts on Collaboration and Long-Term Project Planning: Long-Term Maintenance / Support

I decided to go ahead and post this because of an article by Adam Siepel that I read today, describing broader need to take maintenance / support into consideration for projects (and I previously already had most of this content in a draft).  For example, I thought it was interesting that he brought up the R50 Research Specialist Grant.

In terms of my own personal experience, needing to provide support for COHCAP outside of working hours (back in 2018) was one factor that made it clear to me that the "templates" would have issues with support if I continued to expand topics of research at that previous rate (and that I needed to focus on fewer projects more in-depth).

That said, my individual opinion is that it would be inappropriate to convert COHCAP to have a fee-based license, because having a variety of free programs that can be tested for each project has been very helpful to me (and I recommended testing both COHCAP and methylKit for projects, since I couldn't guarantee any strategy would work out for any particular project).  My impression of DNA.land was also somewhat similar: I think it is good as a free option, but I don't think it would be appropriate to charge for the results that I saw (so, I hope I have misunderstood something about their transition plans).  However, that leaves open the question about what should be done for support (to avoid accumulation of over-time hours as more algorithms are developed).

One thought is that maybe suggest a donation to City of Hope for $10 per project, whenever users find the software helpful would be appropriate (or for some possibly larger amount, from non-scientists that want to keep open-source software free but maintained).  However, I would guess that I have already raised a larger amount through (general) matching funds, so I don't think this is especially urgent.

Otherwise, in terms of alternative funding strategies (instead of patents / licences), these are some ideas that I had:

a) Charge for in-person training of open-source programs / databases?  Allow public (and possibly delayed) free support but charge if providing private support via e-mail?

b) Maybe have small grants for software training / support? Perhaps have a target of a 50-70k salary for a bioinformatician in the lab?  If that is not enough, consider a 100-150k grant for 2 support staff for 1 program (and also encourage users to participate in discussions with other analysts, such as on Biostars, to get a variety of opinions).  This was also briefly discussed in this article on how to support open-source software.  More recently, I think this would also be like the Chan Zuckerberg "Essential Open Source Software for Science" grant.

I hope this doesn't become an issue (like for KEGG, or RepBase, as I understand it), but I noticed that there is a NCBI link to OMIM as well as the OMIM.org link that suggests a donation.  So, perhaps a mix of donations and grant funding covers their needs?

Even in terms of the above options, I tested out the "Developer" support for AWS, but I didn't actually get a response within 24 hours (and ended up solving the problem on my own after that, and reverting back to the free "Basic" plan): so, if you do charge for support, you have to be able to be capable of having prompt, daily discussions to work toward being able to solve user difficulties in a variety of contexts.

As another example, I recently canceled my subscription to the New York Times.  At $4/month, the cost was reasonable.  However, due to the extra effort to view the articles on public computers, I essentially stopped reading the articles.  When I thought about it, I already donate $3/month to Wikipedia.  So, the cost isn't really the limitation: the barriers added by the license / subscription (and the relatively good content that I can get for free) are the reason that I stopped supporting the New York Times.  If they did something similar to Wikipedia (at least for some articles), then I would support them (and, likewise, perhaps I should increase my donation to Wikipedia).

I also have these ideas about limits / suggestions for the use of commercial bioinformatics software, which is kind of like an extension for this post (and this was also discussed in the Genome Biology paper that I read when I first made the post public).

Change Log:
7/30/2019 - public post date
7/31/2019 - trim 1st and 2nd paragraph; add NYT example
8/5/2019 - add matching link and de-emphasize grant
8/6/2019 - minor change
8/12/2019 - add link to RepBase subscription
9/11/2019 - change assumption that donation requires non-profit model
9/16/2019 - add link for DNA.land
11/21/2019 - add link for Chan-Zuckerberg open-source software funding

Sunday, June 2, 2019

What's the difference between a "Pipeline" and a "Template"?

The process of understanding each step of analysis is important for presenting the final set of results, and the process of writing the code for that analysis can help you understand the methods better (and identify questions and/or room for improvement in your current code / understanding).

I have some templates for analysis, but I call them "templates" rather than "pipelines" because the code itself usually requires some modification.  While I think it is extremely useful to have packages for specific functions, you may find that having a pre-set pipeline doesn't quite produce publication-quality figures, and having a template for your own code (that is easier for you to change than somebody else) may make it easier to implement changes that come as the result of iterations of project discussions.  I have a note to this effect for most templates (such as the acknowledgement in the README for the RNA-Seq gene expression analysis "template," as well as a the 2nd post-publication comment for COHCAP, which stands for "City of Hope CpG island Analysis Pipeline").

These modifications can be important for semi-automated analysis, and it is possible that other people may find there are some situations where it can be useful to have templates for intermediate results.  However, I also believe there are some other factors for discussion that are worth taking into consideration:

  • Be careful not to increase the turn-around time for an initial result while increasing the total amount of time to get a paper to publication (or skip steps that could decrease the accuracy of the publication).  This can be particularly tricky if it takes a couple years to appreciate all the time required for follow-up requests.
  • Be aware of how other people will view your code, and what is appropriate to include in a publication. For example, even if it becomes appropriate to have "templates" for intermediate results (which I am not saying is necessarily true), taking time to understand your results is important for responsible research practices (regardless of the formality of what you are making public).  So, testing your code on multiple datasets (either within your lab, or public data from other labs) can be important for troubleshooting.
  • Unlike code published with a particular paper, the templates (by definition) are really designed with my own use in mind, and are much more difficult to support for other people (even within the same lab).  So, while looking at portions of the code may be helpful in generating ideas, support for "templates" can't really be provided (at least not in the same way as "pipelines" or packages for a particular step of analysis)
  • While it is increasingly important to provide code to help with reproducible and understanding of analysis, that code likely needs to be different for each paper (and time and energy will be required for supporting the separate code for each paper, for each lab).  While this is not exactly a pipeline (since that code will likely represent a template to be modified for other people's experiments, not run without any changes, or novel modifications), this is also not the same as saying you have a generalized template that you think is good to use on a large number of projects.

In short, I believe it is important to expect some testing for each project, in order to be the most confident in the results that you present.  One possible alternative to a "template" may be having code for public demo datasets (for training), but that is probably also not a complete solution.

Update Log:
6/2/2019 - original public post

Wednesday, March 13, 2013

Bioinformatics 101: DNA Methylation Analysis

Enrichment-Based Analysis Tools:


Bisulfite-Conversion Based Analysis Tools:

Sunday, February 24, 2013

Notes on AGBT 2013

I've organized my notes from AGBT, which you can download here.  I know a lot of people were not able to attend the conference, so I hope this is helpful in addition to all the other comments from the twitter feed (HT @Massgenomics for the Google Docs link).

There were a number of interesting talks on clinical genomics, sequencing technology, and bioinformatics.  There were also a few very interesting metagenomics talks.  I don't know a whole lot about this area, but the talks certainly got me interested in learning more about this topic.

Aside from the interesting science, there were some nice features about the conference (catered meals, lots of freebies, etc.), but I also feel like there was some room for improvement.  For example, I wish they ended the conference a little earlier on the last day in order to make it easier to catch a flight out of Florida.  Also, I personally prefer conferences that take place in major cities (that don't require a 1-hour shuttle to get to the airport), although I'm sure there are varying opinions on this matter.

Finally - all minor complaints aside, I am very grateful to be able to attend this conference and present a poster for my new DNA methylation algorithm (COHCAP - City of Hope CpG Island Analysis Pipeline).  For those that couldn't attend the conference, COHCAP is an algorithm that quantifies the consistency of methylation patterns for CpG sites within CpG islands (for Illumina methylation array data and targeted BS-Seq).  The paper is currently under review, but the algorithm is currently available to download and I'd be glad to answer any questions that you have about it.

 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.