Showing posts with label R. Show all posts
Showing posts with label R. Show all posts

Thursday, November 14, 2019

What do I need to change as an individual?

I can tell that I need to work on fewer-projects in more depth.

I am not sure if additional training is necessary to accomplish this, but that is the focus of my post on "What are the expectations for Individuals with an MS in Bioinformatics versus a PhD in Genomics?"?

Technique-Wise, these are the sort of things that I think I could be comfortable with:


  • I am very comfortable coding in R / Python / Perl
  • If I needed to get back into the lab, I could previously do a PCR and maintain a cell line
    • At least previously, I had some difficulties being able to preform my own microarray experiments
    • However, to be honest, I think the best fit for me would be to keep doing mostly or entirely computational work (as a Bioinformatics Specialist, Bioinformatician, etc.).
  • I think I likely need to reduce the number of new patient samples that I encounter, but I am comfortable working with my own genomics data and I think I have had useful contributions using re-analysis of data deposited by other labs.
    • For example, I thought this was a relatively successful story of my feedback on a pre-print being helpful
    • I also have notes on my human genomics results here
    • I also have on-going analysis of public cell line perturbations to demonstrate method limits for RNA-Seq analysis


Research-Wise, there are topics that I am interested in and/or have some prior experience with:


  • Investigate whether a relatively simple strategy is more robust than a more complicated strategy (for example, try to identify problems with over-fitting)
  • Probably a good idea to either limit the number of samples I work on at a given time and/or make sure that I have enough time for several rounds of analysis / discussion of the same dataset (which is always good for critical assessment of results)
  • Perhaps place more focus on non-human genomics (and I started out doing evolutionary genomics research)?
    • I have some previous virology experience, so perhaps I could study the genetics / genomics of viruses that currently infect other animals (but have not yet evolved the ability to infect people)?  This could even be part of DNA-Seq or RNA-Seq for the host.
    • While I haven't done any such research from a professional standpoint, I have been comparing genomics results for my cat Bastu, and I hope to have a blog post summarizing that soon.
  • If I can get agreement about some things, perhaps there is some value in shared support guidelines?
    • I also truly like the idea of supporting labs with less popular research topics and/or smaller budgets (which may have a relatively greater need for shared support)
    • Essentially, I am emphasizing training (indirect analysis support) and limits on projects for shared staff
    • However, I don't particularly like telling other people what to do (the degree discussion is really more about having the right amount of autonomy and peer respect to perform my own analysis), even though I realize that we sometimes have obligations to society to do things that we may not find enjoyable.
    • So, I think finding a solution for myself what is currently most important, but I think I have some experiences that may be useful to others.
    • In general, if I were to assist with training support, I think I may be able to help with common public datasets, but I think there may need to be a rule that I couldn't help with providing analysis for a new dataset (unless a greater commitment to the project is made, where the limits on shared support would then be important)
    • While helping me stay up-to-date and refreshed on details, perhaps providing local guidance (face-to-face) for a subset of content from selected on-line courses (like Coursera) may be an appropriate way for me to help, but it would be crucial that I not complete exercises for any students (which would violate the honor code, requiring students to complete their own work and demonstrate independent competence).  For example, if I did this before they started the course, perhaps I could then recommend where to learn more and get certification (as well as setting realistic expectations on the likelihood of passing the course).
  • I would need a lot of practice, but outreach / optional education for the general public (such as a book club discussion that I led) can be rewarding 
  • While I am less certain about my role in a professional standpoint, you can see my "speculative opinion" posts about some things that I think could be interesting


Personality-Wise, these are what I believe are my strengths and weaknesses:


  • I have to be fairly independent for my current job, but I do provide a supportive role (where biological / clinical idea usually comes from PI)
  • While the difference between 1 day and 2 weeks turnaround time would be an order of magnitude (for each iteration of analysis/discussion), I have received the good suggestion that I should wait at least an extra day before returning each round of results (to see if I can catch more errors by reviewing the results again the next day).
  • I like the idea of helping provide a "public good," so I think I would prefer to continue working at non-profits
  • Continue becoming better at more mindful when I have a prior assumption (which may or may not be true) and I may not sufficiently understand other perspectives.
    • However, I appreciate those who value the need to take time to be objective and fair
  • While I believe it is important to continue to make future progress, I have some concerns about responsibilities what require excellent communication (and would frequently involve relatively short interactions with individuals where I may not have the chance to correct myself)
    • For example, part of the reason I work on the computer is that I bugs will stay corrected in the code (once I find them).
    • In contrast, if you knew that there was an experimental protocol that required X steps and I was highly like to mess up at least 1/X steps, then that is sufficient for me to not be able to get a protocol to work (for example, I think this is why I had previous difficulty with performing my own microarray experiment).
  • I think I may need to better recognize what I can fix (for myself), versus a concern about the actions of others (which may be solvable if properly communicated, or may be harder to resolve without common agreement)
    • For example, there may be some room for improvement in terms of communicating myself in sensitive situations (such as disagreeing with a policy and/or a superior).  However, this is something that I am actively working on.
  • Most of my family lives on the east coast of the United States (and I currently live in the west coast, in California).  As we get older, this may be something worth taking into consideration.



Change Log:

11/14/2019 - public post
11/15/2019 - public post
11/21/2019 - fix typos + add waiting a little longer to return results
1/7/2020 - add link for on-line course notes
4/30/2020 - update cat link for blog post versus GitHub
10/7/2020 - add note about computational emphasis

Saturday, June 15, 2019

What About Bioinformatics Companies?

A product starting under a grant (such as in academics or a non-profit) but later being becoming a start-up (as a for-profit) is one possibility. However, my post on providing generics through non-profits would bring into question whether such a start-up could continue to be an independent non-profit (and/or a government entity/contract).

While I think I need to learn more before being able to say something strictly can't be provided from a for-profit organization, I think there are some things that may need to be taken into consideration within the current framework of options:


  • Perhaps require free command-line version of software for commercial software with a user interface (and emphasize the education component to learn more coding)?  Otherwise, it isn't really reproducible for most people, and it probably isn't appropriate to only use one program in all situations.
  • In terms of compromises to help with reproducibility, Novoalign has free version that uses fewer threads, and MATLAB allows people to run programs developed in MATLAB without a licence .
  • precisionFDA is designed/supported by a private company (DNAnexus), though what I assume is a government contract.  So, if you have genomics data, it is free to re-analyze / compare your own genomics data (since the FDA is paying for the costs).  I think this is a very good thing for citizens that can help them become more involved (and, hopefully, understand the difficulties of the regulatory process a little better).  I have some notes on on my own experiences with precisionFDA here.  So, I support this strategy (although perhaps there can be discussion about the fact that DNAnexus is currently a for-profit company).

Also, the idea that encouraging the testing multiple free RNA-Seq methods (something else that I would like to be able to show, at some point) may seem at ends with having commercial bioinformatics software.  While there is some truth to this, I would have the following response to such a critique:

1) I think the time frame matters when discussing software recommendations.  If the goal is to increase coding abilities in 5-10 years, then papers that need to be published sooner may need some alternative solution.  In other words, if somebody doesn't know how to code and there is a program with a graphical user interface that helps them do some analysis on their own, I think that can be good.  However, if they get a weird result (and/or a negative result) with that program (which could have a commercial license), I strongly suggest they test other programs before preparing for publication (and those other programs may need to be open-source command line programs)

2) Having extra options (which includes commercial software) gives labs more options of programs that they can test for their project.  So, if you get a weird result with the open-source software, having extra commercial options may help.  My only concern is that the fees may be a barrier to entry for some labs, and I don't want to encourage excessive use of free trials if licenses are not often eventually purchased.

So, I think there is still value in giving suggestions of how to make the most out of available open-source options, even you use use some commercial programs.  This is similar to what I do: the majority of the bioinformatics programs that I use are open-source, but I sometimes also use commercial programs (like IPA).  However, even with IPA, I would also recommend comparing results with free programs (like Enrichr).  Sometimes the free open-source programs end up being a better fit for the individual project than the commercial ones, but that often varies by project.

However, if you can't lock down one particular program to use in all situations (which has definitely been my experience), that is why I am somewhat concerned about the barriers to entry that could be caused by having to purchase licenses (and, if you don't already have a license, I would usually recommend trying out open-source options first).  I think this is also helps with your ability to provide support during transitions.  For example, I code in R rather than MATLAB, so I don't have to worry about keeping a MATLAB license (and a lot of genomics packages are developed in R/Bioconductor, in addition to it being free).


Change Log:
6/15/2019 - public post date

Friday, May 25, 2012

Shared Scripts for Genomic Analysis

I've recently added a section to my personal website that contains some scripts that I have used for bioinformatics analysis:

https://sites.google.com/site/cwarden45/scripts

As of right now, the page contains a handful of scripts for microarray, next-generation sequencing, and qPCR analysis.  I plan on updating this page periodically.

Generally speaking, the scripts aren't organized in a carefully documented package (like what you can find from Bioconductor, etc.).  However, I have found them to be very useful templates for routine analysis, so I thought it might be useful to share them with others.

Monday, May 16, 2011

Modeling Bimodal Gene Expression

Since it is often challenging to estimate parameters for mixture models (such as those used to model bimodal gene expression), I thought it might be useful to discuss some of my successes using non-linear least sequares (NLS) regression to model bimodal gene expression.

Many scientists use maximum likelihood estimation (MLE) to model bimodal gene expression (such as Lim et al. 2002, Fan et al. 2005, Mason et al. 2011, etc.).  My MLE model is based the code provided in this discussion thread, so I used the mle function from the stats4 package (which is a wrapper for the standard optim function).

On simulated data, both the MLE and NLS models estimate 37% of samples show over-expression (i.e. come from the distribution with the higher mean), which is very close to the true value of 36%:


Simulated Data


However, I have found that the mle function often returns error messages when working with real data and models built using the nls function tend to fit the data better than the MLE estimates.  Even on the simulated data above, the NLS model appears to prove an ever so slightly better fit for the data (as might be expected because it directly models the density function).

In order to illustrate my point, I have also analyzed some genes exhibiting bimodal gene expression (as identified by Mason et al. 2011) in the GEO dataseries GSE13070.  For example, here is a gene that I could model relatively well with NLS regression whereas I simply couldn't produce an MLE model:




Of course, both of these tools have their limitations.  For example, the data has to have a pretty clean bimodal distribution (here are 3 examples of distributions that couldn't be modeling using either method: ACTIN3, ERAP2, and MAOA (different probe)).  For the NLS model, I also had to set the variance to be equal for the two samples in order to produce a reasonable estimate of over-expression, but I believe this is usually a safe assumption.

Although I do not present the data in this blog post (because it contains unpublished results), I have also found NLS regression to be useful on other genes in several other datasets,  So,  I know NLS regression works well with more than just the one gene that I show above.

Also, for those that are interested, here is the source code that I used to produce all of the above figures.
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.