Showing posts with label post-publication review. Show all posts
Showing posts with label post-publication review. Show all posts

Wednesday, May 6, 2020

Opinions Related to Gencove Pre-Print Comment


Because a pre-print comment is somewhat formal, I thought that I should separate my opinions from the main feedback.

So, I decided to put those in a blog post.  You can see my pre-print review/comment here, and these are the extra comments:

General Notes / Warnings (Completely Removed from Comment):

My Nebula lcWGS results were OK for some things (like relatedness and broad ancestry), but I found the Gencove accuracy to be unacceptable for specific variants (for myself).

While Nebula has changed to only provide higher coverage sequencing, I previously submitted an FDA MedWatch report for my own data (for the lcWGS Gencove results).

To be fair, there are also general limits to the utility of most of the Polygenic Risk Scores that I was able to test with my own data (with some informal notes in this blog post).  So, while true, mentioning that I still had concerns about the percentiles that I saw from Nebula (even with the higher coverage sequencing data) may be less relevant.

Similarly, while I want to encourage other customers to report anything they find to MedWatch (and/or PatientsLikeMe, etc), I also want to acknowledge my own limitations that this general warning is more about specific issues that I found for myself.  For example, it may help to have an independent analysis with larger sample sizes to gauge my general PRS concerns and/or be more specific in terms of which specific PRS do or do not have clinical utility with sufficient predictive power for the disease association.

Specific Comment #2) I think my own result might match the imputed correlation that is described (in terms of having ~90% accuracy).  However, I would say that is unacceptable for making clinical decisions, especially since more accurate genotypes can be defined.  It is important to be transparent and not over-estimate accuracy, so I think that part is good.  I also realize that something unacceptable for individual variants can be acceptable for other applications.  However, I think something about limits should be mentioned for the general audience, even if they really apply to the same Polygenic Risk Scores in higher coverage sequencing data.

I am not sure if this matters for this particular project, but I have found that it is not unusual to learn about something that may contradict an original funding goal.  I have certainly noticed that it can take me a while to realize I need to question some original assumptions, but sharing those experiences is extremely valuable to the scientific community (if the conclusions then shift to helping others avoid similar mistakes).  I also realize prior assumptions can be hard overlook in comments/reviews as well, and there is definitely more that I can learn.  Given that you have a pre-print and there is a lot of details in supplemental information and external files, I think that is a good sign.

Specific Comment #3) In the future, I hope that this is also the sort of thing that precisionFDA, All of Us, etc. can help with.  In fact, as an individual opinion, this makes we wonder if the SBIR funding mechanism might be able to help with directly providing generics through non-profits (especially for genomics diagnostics).  However, I don’t think that means SBIR for-profit funding would have to be completely ended to preferentially fund non-profits, and I realize that probably can’t affect this particular paper.

If the Gencove code isn’t public, then I am not sure how you could show others could reproduce a freeze of the code before testing application to new samples.  Nevertheless, I applaud that you provided some code for the publication.

Specific Comment #4) There may be a way to revise the current manuscript without adding the independent (public) test data and/or the open-source alternatives.  For example, I don’t think you need additional results for your effective coverage section, but I am more interested in the concordance measures.  If the Gencove / STITCH / GLIMPSE / IMPUTE results are similar in terms of technical replicate concordance (for the same 1000 Genomes samples), then I think that you could skip what is described for specific comment 3) for this paper.

I also noticed that the competing interests statement was in the past tense for the present employees (as I understand it).

Summary: I think the utility for lcWGS to cause additional genomic data types to be considered identifiable information is important (which I have in a different blog post).


Change Log:

5/6/2020 - public post

Thursday, April 30, 2020

Personal Thoughts on Collaboration and Long-Term Project Planning: Reproducibility and Depositing Data / Code

I believe that I broadly need to improve explanations for the need / value to deposit data and have code for the associated paper (even though that takes additional time and effort).

This is already a little different than the other sections, since it more of a question than a suggestion.

Nevertheless, as an individual, this is what I either currently do or I need to learn more about:
  • I am actively trying to better understand the details for proper data deposit for patient data (even though I have previously assisted with GEO and SRA submissions).
    • For example, I am trying to understand how patient consent relates to the need to have a controlled-access submission (even if that increases the time necessary to deposit data, or that certain projects should not be funded if the associated data cannot be deposited appropriately).  So, being involved with a successful dbGaP submission would probably be good experience.
    • I thought the rules were similar for other databases (like ArrayExpress, ENA, EGA, etc.).
    • However, if you know of other ways to appropriately deposit data, then I would certainly be interested in hearing about them!
  • If possible, I always recommend depositing data (and you see several papers where we did in fact do that), but I think different expectations would need to be set for supporting code (hence I have said things like "I cannot provide user support for the templates").
    • This is not to say I don't think code sharing is important.  On the contrary, I think it is important, but you have to plan for the appropriate amount of time to carefully keep track of everything needed to reproduce a result.
    • Also, if are a lot of papers where code has not been provided in the past, then I have to work on figuring out how to explain the need to share code (and spend more time per project, thus reducing the total number of projects that each lab/individual works on).
  • I think it would also be best if I could learn more about IRB/IACUC protocols (for both human and other animal studies).

In terms of how I can think of potentially emphasize the importance of data deposit and code sharing, you can see my notes below.  However, if you have other ideas about how to effectively and politely encourage PIs to deposit data and plan for enough effort to provide reproducible code (and/or help review boards not approve experiments producing data that can't be deposited), then I would certainly appreciate hearing about other experiences!

  • Even if it is not caught during peer review, I think journal data sharing requirements can apply for post-publication review?
  • The NIH has Genomic Data Sharing (GDS) policies regarding when genomics data is expected to be deposited.
    • There is also additional information about submitting genomic data, even if the study was not directly funded by the NIH.
    • There is also information about the NIH data sharing policies here.
  • While it mostly emphasized the need to expectation for data sharing with grants that are greater than $500,000, the NIH Data Sharing Policy and Implementation also mentions the need to code-related information available for reproducibility under "Data Documentation".
  • I also have this blog post on the notes that I have collected about limits to data sharing, but that is more about limiting experiments than data deposit for an experiment that has already been conducted.
  • Eglen et al. 2017 has some guidelines regarding sharing code.
  • This book chapter also discusses data and code sharing in the context of reproducibility.

Change Log:

4/30/2020 - public post
8/5/2020 - public post

Monday, August 26, 2019

Comments on Other Papers (outside Disqus) and Positive Examples of Corrections/Retractions

It is easy to find all of my comments through the Disqus comment system.  However, it is more difficult to find other comments that I have made.  So, I thought it might be good to organize some of those here (and note when there are formal corrections):

Comments with Successful Follow-Up:

Kreutz et al. 2020 (Bioinformatics) - correct typo in title after PubPeer comment

Martin et al. 2019  (Nature Genetics) - formal erratum, and earlier PubPeer comment (one typo, one suggestion that Figure S12 makes absolute accuracy more clear)

Jonsson et al. 2019 (Nature) - formal correction following article comment

Weedon et al. 2019 (bioRxiv pre-print) - extra data (including my own) was added after comment

Zhang et al. 2019 (PLOS Computational Biology) - I had a comment that I believe was able to be corrected before the final version of the paper (although the paper did later have a formal correction)

Pique-Regi et al. 2019 (bioRxiv pre-print) - discussion helped fix links in pre-print

BLAST reference issue reported to NCBI (caused from a number of different papers no recognizing cross contamination, I believe e-mail was sent showing incorrect annotations for a limited number of example sequences).  It looks like problem was not permanently fixed (due to similar problems with new data from other studies).  However, through whatever mechanism, top PhiX BLAST hits that were incorrect were removed - not sure if it was the actual cause for correction, but I posted this on Biostars when it was a problem

Multiple but Primarily Minor Errors (not currently fixed):

Yizhak et al. 2019 - blog post (errors and suggestions) and PubPeer comment (at least one clear typo); I also posted an eLetter describing the 2 most clear errors; Science paper

Arguably Gives Reader Wrong Impression:

Li et al. 2021 - there are a number of things that I am understanding better by participating in the discussion process.  Also, my original intention was to add a Disqus-style comment on the journal article, but that was not provided.  So, that is why I posted a PubPeer comment.

Essentially, I think I would prefer a SNP chip (or Exome, higher coverage WGS, Amplicon-Seq, etc.) over lcWGS (especially the 0.5x to 1x presented in this paper), and I think the value of directly measured genotypes for SNP chips is being underestimated.  I agree with lcWGS imputations often being preferable to SNP chip imputations, but I think inclusion of imputed SNP chip genotypes may not be made sufficiently clear to the broader audience.

I also mention problems with consumer lcWGS products.

Maya et al. 2020 - if I understand the paper correctly, the study does not directly investigate COVID-19 infections (they are looking for associations near ACE2 or TMPRSS2, but in other contexts).  Please scroll to the bottom to see my comment.

Nature summary of PLOS ONE paper - image shows input (not output) of encoding and attempted recovery; unlike Nature research reports, I didn't see a Disqus comment system (but I did mention something on Twitter)

Homburger et al. 2019 - I had enough problems with my lcWGS data (which was exceptionally low coverage for my Color data) that I would not recommend it's use.  I realize that sample size is limited (just my own data).  However, I posted my concerns on a pre-print for this article.  I think this is borderline for the previous category ("Data Contradictory to Main Conclusions", more appropriate for a retraction), but I would need more data to conclude that.  I also posted a similar response on PubPeer.

I think this blog post on being able to self-identify myself can give some sense of the concerns about the very low coverage WGS reads that I received from Color.

Singer et al. 2019 - eDNA paper with multiple comments indicating concerns, and I specifically found there were extra PhiX reads in the NovaSeq samples.  As mentioned in the comment, I have an "answer" on a Biostars discussion more broadly related to PhiX (with a temporary success story about the BLAST database mentioned later in this post).

The author provided a helpful response indicating that the PhiX reads should be removed for DADA2 for downstream analysis.  However, I think we both agree that PhiX spike-ins don't have a barcode.  So, if you view that as extra cross-contamination in some samples, then this leaves the question of what else could be in the NovaSeq samples that might be harder to flag as something that needs to be removed.  Time permitting, I am still looking into this.

Minor Typos:



Yuan and Bar-Joseph 2019 - PubPeer comment about MSigDB citation (as GSEA)

Chen et al. 2013 - comment about typos in PLOS comment system

Doorslaer and Burk 2010 - PubPeer comment

Other:

Börjesson et al. 2022 - PubPeer question/comment

Andrews et al. 2021 - PubPeer question (because Nature Methods doesn't have a Disqus comment system; also, public acknowledgement of misunderstanding on my part)

Antipov et al. 2020 - PubPeer comment (comments left in Word document for Supplemental Figure S2)

Robitaille et al. 2020 - PubPeer comment (question, posted there due to lack of comment system)


Older et al. 2019 - comment in PLOS comment system.

Young 2019 - comment in PLOS comment system

Choonoo et al. 2019 - comment in PLOS comment system

Paz-Zulueta et al. 2018 - PubPeer comment

Mirabello et al. 2017 - PubPeer comment

Shen-Gunther et al. 2017 - PubPeer comment

Robinson and Oshlack 2010 - PubPeer question (in the interests of fairness, I believe this reflects misunderstanding on my part)

Munoz et al. 2003 - PubPeer question (I didn't find any errors in the paper, but I wondered why data was presented in a particular way, and journal didn't have a comment system)

I also have some notes on limits to precision in genomics results available to the public (which I would probably define more as a "product" than a "paper), among this collection of posts.

Given it is harder for me to find these (compared to my Disqus comments), I will probably add more in the future (but the comments in the "other" section are not necessarily bad - I will just probably forget them if I don't save a link here).  For example, I believe increased participation in such comment systems may make them as effective as a "minor comment" (particularly if the journal doesn't go back and change the original publication after a formal correction).

Checking / Correcting Citations of My 1st Author Papers

Silva et al. 2022 - journal comment

Xu et al. 2021 - PubPeer comment

Wojewodzic and Lavender et al. 2021 - preprint comment (hopefully, can be corrected before peer-reviewed publication)

Roudko et al. 2020 - journal comment (please scroll down, past references, to see comment)

Nayak et al. 2019 - Disqus comment (on pre-print)

Borchmann et al. 2019 - Disqus comment (on pre-print)

Youness and Gad 2019 - PubPeer comment

Muller et al. 2019 - Twitter comment

Bogema et al. 2018 - PubPeer comment

Einarsdottir et al. 2018 - PubPeer comment

Pranckeniene et al. 2018 - journal comment (please scroll down, past references, to see comment)

Debniak et al. 2018 - journal comment

Hu et al. 2016 -  PubPeer comment (probably due to confusion on Bioconductor page)

Wong and Chang 2015 - PubPeer comment

Wockner et al. 2015 - PubPeer comment (more of a true comment, than post-publication review for error)

As mentioned in a couple other posts (general and COH-specific), I am trying to correct previous errors in my papers, and I think having 4 first-author (or equivalent) papers in 2013 was probably not ideal (in terms of needing to take more time to carefully review each paper).

However, I also want to show a relatively long list of corrections / retractions (even though they still make up a subset of the total public records), with some greater emphasis on high impact papers and/or papers with a large number of authors. To be fair, I may not know the details of the individual corrections or retractions.  However, I hope this will cause future researchers to be brave and responsible and report / correct problems as soon as they discover them.  As a best case scenario, I believe that it is important for prestigious labs to lead by example (so that everybody represents themselves fairly and finds the best long-term fit).  We do not want to encourage people to hesitate reporting problems out of fear that they will lose funding and/or collaborations with peers (and/or those that do worse work but less frequently admit errors and/or oversell results will be chosen over those who present more realistic expectations).  I am very confident these are achievable goals, but I think some topics may need relatively more frequent discussions.


Other Positive Examples of Corrections:

I am trying to collect examples for high-impact papers/journals and/or consortiums.

Correction for Collins et al. 2020 (Nature paper, The Genome Aggregation Database Consortium paper)

Correction for Karczewski et al. 2020 (Nature paper, The Genome Aggregation Database Consortium paper)

Correction for Minikel et al. 2020 (Nature paper, The Genome Aggregation Database Consortium paper)

Correction for Wang et al. 2020 (Nature Communications paper, The Genome Aggregation Database Consortium paper)

Correction for Whiffin et al. 2020 (Nature Communications paper, The Genome Aggregation Database Consortium paper)

Correction for Cortés-Ciriano et al. 2020 (Nature Genetics paper)

Retraction for Cho et al. 2019 (Science / Nobel Laurete paper: I learned about from Retraction Watch)

Retraction for Wei and Nielson 2019 (Nature Medicine paper)
--> Post-publication view indicated on Twitter in advance
--> Also includes retraction of "News & Views" article

Correction to Kelleher et al. 2019 (Nature Genetics paper)

Correction to Grishin et al. 2019 (Nature Biotechnology paper)

Publisher Correction to Exposito-Alonso et al. 2019 (Nature paper, as well as "500 Genomes Field Experiment Team" consortium paper)

Correction to Bolyen et al. 2019 - (Nature Biotechnology paper; correction for QIIME II paper)

Correction to Kruche et al. 2019 - (Nature Biotechnology paper; GA4GH Small Variant Benchmarking)

Retraction to Kaidi et al. 2019 (Nature paper; positive in the sense author resigned and admitted fabrication, rather than denying fabrication; I learned about from Retraction Watch)

Retraction to Kaidi et al. 2019 (Science paper; positive in the sense author resigned and admitted fabrication, rather than denying fabrication; I learned about from Retraction Watch)

Correction to Ravichandran et al. 2019 (Genetics in Medicine paper; added conflicts of interest)

Erratum to Jiang et al. 2019 (one affiliation issue with Nature Communications paper with a lot of authors)

Correction to Ferdowsi et al. 2018 - (mixed up Figures; Scleroderma Clinical Trials Consortium Damage Index Working Group paper)

Correction to Sanders et al. 2018 - (Nature Neuroscience paper; Whole Genome Sequencing for Psychiatric Disorders (WGSPD) consortium

2 corrections to Gandolfi et al. 2018 paper on cat SNP chip array

Corrigendum to Oh et al. 2018 (99 Lives Consortium paper)

Correction to (a different) Gandolfi et al. 2018

Correction to Vijayakrishnan 2018 (PRACTICAL Consortium included on authors list)

Correction to Matejcic et al. 2018 (another PRACTICAL Consortium paper)

Correction to Went et al. 2018 (another PRACTICAL Consortium paper)

Correction to Schumacher et al. 2018 (Nature Genetics paper, several consortium listed as authors)

Correction to Mancuso et al. 2018 (another PRACTICAL Consortium paper)

Correction to Armenia et al. 2018 (affiliation correction, Nature genetics paper, Stand Up To Cancer, SU2C, Consortium paper)

Withdraw of Werling 2018 Review (I learned about from Retraction Watch)

Correction to Sud et al. 2017 (another PRACTICAL Consortium paper)

Retraction and Replacement of Favini et al. 2017 (JAMA paper ; I learned about from Retraction Watch)

Corrigendum to McHenry et al. 2017 (Nature Neuroscience paper; I learned about from Retraction Watch)

2 Erratums to Teng et al. 2016 (Genome Biology paper; I found from this Tweet, describing error identified and corrected by another scientist from post-publication review; one Erratum is really a subset of the other)

Correction to Nik-Zainal et al. 2016 (Nature paper; typo not fixed in on-line article)

Erratum to Aberdein et al. 2016 (another 99 Lives Consortium paper)

2 corrections to Vinik 2016 (NEJM paper; I learned about from Retraction Watch)

Retraction of Jia et al. 2016 (Nature Chemistry paper + Nobel Laurate Author; I learned about from Retraction Watch)

Retraction of Colla et al. 2016 (JAMA paper; I learned about from Retraction Watch)

Erratum to Kocher et al. 2015 (Genome Biology paper)

Retraction of Zhang et al. 2015 (Nature paper; I learned about from Retraction Watch)

2 Retractions (?) to Akakin et al. 2015 (Journal of Neurosurgery paper; I learned about from Retraction Watch)

Addendum to Koh et al. 2015 (Nature Methods paper; figures were reversed, creating mirror images)

Retraction and Replacement for Hollon et al. 2014 (JAMA Psychiatry paper; I learned about from Retraction Watch)

Retraction to Kitambi et al. 2014 (Cell paper; I learned about from Retraction Watch)

Correction to Huang et al. 2014 (Nature paper; I learned about from Retraction Watch)

Retraction and Replacement of Li et al. 2014 (The Lancet; I learned about from Retraction Watch)

Retraction and Replacement of Siempos et al. 2014 (The Lancet Respiratory Medicine paper; I learned about from Retraction Watch)

Correction to Xu et al. 2014 (eLife paper; I learned about from Retraction Watch)

Retraction of Lortez et al. 2014 (Science paper; I learned about from Retraction Watch)

Retraction of De la Herrán-Arita et al. 2014 (Science Translational Medicine paper; I learned about from Retraction Watch)

Retraction for Dixson et al. 2014 (PNAS paper; I learned about from Retraction Watch)

Retraction of Garcia-Serrano and Frankignoul 2014 (Nature Geoscience paper; I learned about from Retraction Watch)

Retraction of Amara et al. 2013 (JEM paper; I learned about from Retraction Watch)

Retraction of Maisonneuve et al 2013 (Cell paper; I learned about from Retraction Watch)

Retraction of Yi et al. 2013 (Cell paper; I learned about from Retraction Watch)

Retraction of Nandakumar et al. 2013 (PNAS paper; I learned about from Retraction Watch)

Retraction of Venters and Pugh 2013 (Nature paper; I learned about from Retraction Watch, describing 6th Nature retraction in 2014)

Retraction of Maisonneuve et al 2011 (PNAS paper; I learned about from Retraction Watch)

Retraction of Frede et al. 2011 (Blood paper; I learned about from Retraction Watch)

Retraction to Olszewski et al. 2010 (Nature paper; mentioned in this Retraction Watch article)

Correction to Werren et al. 2010 (Science paper; correction not actually indexed in PubMed?)

Retraction of Bill et al. 2010 (Cell paper; I learned about from Retraction Watch)

Retraction to Wang et al. 2009 (Nature paper; I learned about from Retraction Watch)

Retraction of Litovchick and Szostak 2008 (PNAS paper + Nobel Laurate Author; I learned about from Retraction Watch)

Retraction of Okada et al. 2006 (Science paper; I learned about from Retraction Watch)

Withdraw of Ruel et al. 1999 (17 years post publication; JBC paper; I learned about from Retraction Watch)

I think this last category is really important (even though it is still no where near complete - it is just some representative examples that I know about).

Change Log:

8/26/2019 - public post date
8/27/2019 - add sentence about increased comment participating being more like formal comment (in some situations)
8/29/2019 - minor changes
8/29/2019 - added examples tagged with "doing the right thing" in Retraction Watch.  Thank you very much to Dave Fernig!
8/30/2019 - changed tense of previous entry of change log (from future tense to past tense)
9/2/2019 - add another positive example of comments improving pre-print
9/5/2019 - fix typo; minor changes
9/17/2019 - all 2019 correction, which I remembered from this tweet
9/18/2019 - fix typo; add multiple other citations
9/20/2019 - add PLOS Genetics correction as successful example
9/30/2019 - add recent Nature / Nature Medicine corrections
10/8/2019 - add Nature Medicine CCR5 official retraction and Nature Biotechnology / Nature Genetics corrections
10/18/2019 - add another PLOS ONE comment, PhiX example, and started list of incorrect citations for my 1st author papers
10/24/2019 - add PubPeer comment and eDNA paper
10/25/2019 - add PubPeer comment
10/28/2019 - add more PubPeer / citation comments
10/30/2019 - add Twitter comment
11/8/2019 - add a couple more corrections
11/10/2019 - revise sentences about being "brave"
11/11/2019 - add another PubPeer comment
12/2/2019 - add lcWGS concern + PLOS "filleted" typo
12/10/2019 - add positive example for formal Nature correction
12/24/2019 - add another PubPeer comment
12/30/2019 - add another PubPeer comment
1/2/2020 - add PubPeer question + Science retraction
1/14/2019 - add another PubPeer comment
1/29/2020 - add another Disqus comment
4/24/2020 - add comment for citation of one of my papers
5/15/2020 - minor changes
6/4/2020 - add PubPeer entry for article with typo in title
6/9/2020 - add COVID-19 paper that I believe uses confusing wording
6/15/2020 - add F1000 Research comment
6/25/2020 - minor formatting changes
7/30/2020 - MetaviralSPAdes comment
7/30/2020 - add TMM question
11/6/2020 - move eDNA category
11/11/2020 - add Bioinformatics typo correction
12/11/2020 - add PLOS ONE comment
2/2/2021 - add a couple more positive correction notes + BMC Medical Genomics figure label issue
2/3/2021 - add additional corrections for consortium papers published in Nature journals on 5/27/2020
3/13/2021 - add MSigDB citation as GSEA
4/8/2021 - minor changes
4/10/2021 - add lcWGS PubPeer comment
4/14/2021 - add another example of successful post-publication review
7/28/2021 - add Nature Methods multiplet PubPeer question
8/6/2021 - add note to acknowledge my misunderstanding for Nature Methods multiplet PubPeer question
10/30/2021 - try to better explain PhiX issue temporarily fixed by NCBI; also, move a subset of examples to my Google Scholar page (since the RNA-MuTect Science paper eLetter indicating a need for corrections was automatically recognized, but I needed to manually fix some things - this made me realize I may be able to use Google Scholar to help emphasize some post-publication review)
3/15/2022 - add MethReg comment
6/2/2022 - additional COHCAP corrigendum citation references
10/4/2022 - TC-hunter PubPeer question/comment

Friday, July 26, 2019

Personal Thoughts on Collaboration and Long-Term Project Planning

Our opinions can change over time, and some long-term effects may not be noticeable until 5+ years of experience.

While I still don't think I have everything figured out, I am using this page to organize my thoughts on some topics that may be of use to the broader community.  I also hope that the update/change logs may also be helpful for giving credit to feedback from others during discussions.

Nevertheless, for these posts, I am going to try and focus on what I believe I understand most clearly:



Again, it is probably a little early for me to be giving advice (since I don't have a solution worked out for myself yet), but I hope sharing my experiences can be helpful to other people as I sort out the details for figuring out a sustainable workload for myself.  Having the patience to work on agreed processes step-by-step is also important, but I believe some of this information may be important for future changes (even if they don't occur in the immediate future).

To be clear, I very much enjoy working with collaborators as a Bioinformatics Specialist in a Core Facility.  So, while some of what I am saying indicates room for future improvement, I have an overall positive impression of my work environment and the researchers that I have worked with (who are passionate about helping other people).

Plus, even though I think some of this content is important for long-term discussions, I also want to emphasize that you can be genuinely proud for putting in your best effort to help people and there is some need for short-term support (such as a temporary difficulty in getting additional funding) or at least giving people the chance to think carefully about whether a more major transition is necessary.

Update Log:

7/26/2019 - public post date
7/29/2019 - trim down introductory paragraph
7/30/2019 - add link for maintenance / support, and modify preceding sentence
8/4/2019 - minor edit after some proofreading by a family member
8/6/2019 - minor changes
4/30/2020 - add link to code / data sharing details (either required or suggested)

Personal Thoughts on Collaboration and Long-Term Project Planning: Post-Publication Review

I have another post more broadly describing the importance of comments / corrections that I have self-imposed on my papers, as well as thoughts about the science-wide error rate.

However, those are not all from work I did in a shared resource at City of Hope.  So, I thought I should summarize a subset of those points here:


  • COHCAP comment #1: correction of minor typos (now upgraded as a formal corrigendum) 
  • COHCAP comment #2: my personal opinions emphasizing the following points
    • "City of Hope" should not have been used in the algorithm name
    • I've more recently gained better appreciation for the need to have testing of methods for every paper (so, I mention that readers should not consider the best COHCAP results to be completely automated).  Given that COHCAP stands for "City of Hope CpG Island Analysis Pipeline" this is relevant to my discussions of "templates" versus "pipelines"
  • COHCAP comment #3: while the Nucleic Acids Research editors were very helpful in encouraging me to look more closely at a discrepancy in the listing of the machine for processing the 450k array, they declined to post the comment because it was ultimately determined to be an error in the GEO entry rather than the Supplemental Materials for the COHCAP paper.
    • I mention this in a little greater detail on Google Sites; however, I was able to confirm that the HiScanSQ (not the BeadArray) was used to process the samples because i) the BeadArray is not capable of processing a 450k array and ii) City of Hope never owned a BeadArray.
  • 2nd Author Correction: Table #2 was wrong (duplication of table #1, although the table description was correct)
  • 2nd Author Comment #1: Use of the phrase "silhouette plot" was not precise
    • While this could potentially be an example of a concern for a bioinformatician within a biology lab, I worked on this paper when I was in the COH Bioinformatics Core.  So, I think the most important lesson is to develop habits where you stop whenever you encounter something you don't know, and set a pace (and total number of projects) where you expect to have to take some time to learn more about what you see in the literature (and how to ask the right / best questions to collaborators that are likely also busy working on multiple projects).
  • 2nd Author Comment #2: Use of the phrase "silhouette plot" was also used in another paper, which was published before this 2nd author paper (even though this project was started first)
    • I think this is important in terms of better appreciating the interdependence of labs supported by the same staff member (although I have started try and have acknowledgements for templates, and making notes in follow-up analysis whenever code is copied between labs prior to publication).
  • Middle-Author Papers
    • It is important that I am fair to everybody (regardless of whether they are a collaborator).  However, I also realize this is a sensitive issue that requires some additional internal communication.
      • So, I have reduced the amount of details for these examples.  While I think there has been at least 1 correction that was initiated more than a year ago, I am (slowly) continuing to follow-up whenever something is or was not correct.
      • Sorting through the details for corrections is like managing the correct workload for new projects.  If I try to figure out what exactly happened with too many papers at once, I will be more likely to make mistakes.  So, at any given time, I try to focus more on ~3 issues that I know about.
      • In other words, I will be honest if asked about any errors (or potential errors).  However, if I have the advantage of being able to have discussions with people who I know better, then I think it is probably wise to focus on that as much as possible.
      • I am willing to add a link to notes about middle-author papers.  However, if it is possible to wait until everything on that list has been corrected, I think that may be preferable.
    • So far, I don't think that I caused most of the middle-author paper errors, but I made some mistakes for middle author papers.  So, I provide a couple examples omitting some specific details below:
      • GEO Sample Label Update: Since GEO doesn't have a change log, I thought I should mention there was one prostate cancer project that I helped prepare for a GEO upload whose GEO labels were not ideal (even though the patient IDs for sample pairing for the sample were correct, the samples should have been called "sample" rather than "patient," and that has been corrected).  This was not a huge problem, but most other GEO corrections are due to me not knowing about the machine (so, they were errors from somebody that I didn't catch due to a gap in my knowledge).  So, to be fair, I thought I needed to mention this because I was the one who accidentally created the error (rather than passing along somebody else's error).
      • GEO Machine / Base Calling Methods Update: There was at least one submission where machine and methods needed to be updated, both of which involved at least some previous misunderstanding on my part.

If I were to give advice to my previous self, I would say it is important for the project lead to understand the full project (and plan to spend a substantial amount of time revising and critically assessing your results).  If there is something that you don't understand, do everything you can do discuss with the other authors prior to paper submission.  After all, you will likely have to give at least a partial explanation to people asking about your project, such as face-to-face discussions where co-authors may not present.  It is also important to capture the the full amount of work required for a paper (including post-publication support).

You don't necessarily have to be a project lead to need to plan for an appropriate workload, although taking responsibility for a paper is much more difficult if you aren't a project lead (if you caused the mistake, then somebody else may experience more severe consequences for your mistake).

I think it is also important to emphasize personal limits (and the solution it provides).  If your optimum workload is 5 projects and you work on 10 projects, then you are going to encounter difficulties.  However, I think it can then help if you take your time and gain a better intuition about what you don't know (and therefore what you either need to spend more time on or possibly focus less on overall).  I admittedly still have to figure out exactly what produces the best work-life balance, and I think you have to wait to notice some of the accumulation of follow-up requests and/or post-publication support / review.  However, I think I have gotten a better feel for what that "optimal" day is like: I just have to figure out how to consistently have that each day (on the scale of years).  In other words, if you are feeling overwhelmed, then I would recommend focusing on previous positive experiences as hope that you can improve by decreasing your responsibility / workload.  I also needed to learn to recognize and manage stress better (sometimes with medication).

I think a lot of what I described above can also just be simple mistakes.  For example, I can tell that I make more mistakes if I work overtime on a regular basis or if I haven't been well-rested.  While I didn't exactly cause all of the errors that I described above, I think it is necessary for me to take responsibility whenever I was first author (or equivalent).  If I can describe myself as precisely as possible (which I realize is still a work in progress), then I hope that can also help others as well (for every level of collaboration in putting together a paper).

P.S. There were 2 general points (previously under the "middle author" section) that I think may be better to move to another blog.  I have already move that content (and I will provide links here when public).  However, in the meantime, I would say those fall into the categories of i) what is the best way to correct minor errors (which you can see in this ResearchGate discussion) and ii) explain the need and estimate of time required to provide data and code needed for a result to be reproducible.

Update Log:

7/26/2019 - public post date
7/27/2019 - revise concluding paragraph
7/29/2019 - move majority of concluding paragraph back to a draft; try to be more conservative / clear with commentary
7/31/2019 - add link to COHCAP corrigendum
8/1/2019 - mention there will need to be additional corrections
8/5/2019 - minor changes (+ add back in concluding paragraph, followed by additional trimming/revision)
8/6/2019 - minor changes
8/13/2019 - mention GEO update
9/19/2019 - mention data deposit and code sharing
9/20/2019 - expand middle-author section; minor change
9/21/2019 - minor change
9/28/2019 - add experiences learned from IRB / patient consent process
9/29/2019 - fix typos; reword recent changes
10/02/2019 - add Yapeng link
10/15/2019 - add note for gene length calculation, as well as another link to blog post (with some separated content in this post)
11/1/2019 - mention ChIP-Seq issue
1/28/2020 - add intermediate set of ChIP-Seq notes
4/4/2020 - minor changes + reduce middle author content + move general points
4/24/2020 - add link to ResearchGate discussion
9/7/2020 - minor change (removing some specific information)
12/16/2020 - add another middle-author example without any details (shifting from specific to general)

Personal Experiences with Comments and Corrections on Peer-Reviewed Papers

So far, I have at least 5 publications with examples of comments and/or corrections:


  • Coding fRNA Comment (Georgia Tech project, Published While at Princeton).
    • It is important to remember that anything you publish has a certain level of permanence (and you can be contacted about a publication 10+ years later).
    • On the positive site, I think it is worth mentioning that your own desire to work towards helping people and being the best possible scientist is important: for example, peer review can help if you do everything you can before submission, but I recently provided more public data and re-analysis that I think improves the overall message for this paper (from my own initiative; for example, please see this GitHub page, with PDFs for text and comment figures).
    • So, I truly believe peer review can help, but I think personal responsibility and transparency are at least as important.
  • Corrigendum for 2-Word Typo in BAC Review (UMDNJ, Post-Princeton, Pre-COHBIC).
    • More important than the correction, I think it should be emphasized that 6 months of working in a lab (especially without ever doing the experimental protocol emphasized) is not enough time to justify writing a review.
  • As 1st and corresponding author, I have issued two comments (and a corrigendum) for the COHCAP paper describing a method for analysis of DNA methylation data (City of Hope Bioinformatics Core).
    • There was also a third comment that NAR decided not to publish (regarding an error with the machine used in GEO, which is described in more detail on my Google Sites page and briefly mentioned in the related blog post).
  • As a 2nd equal contribution author, there was a correction regarding the 2nd table, as well as 2 comments related to imprecise use of the term "silhouette plot" (City of Hope Bioinformatics Core)  
  • [initiated and completed corrections to middle-author papers and deposited datasets] (City of Hope Integrative Genomics Core, Post-Michigan).
I trimmed down the details above because I think the formatting on my Google Sites page is a little better, and I listed specifics for the details of the City of Hope papers in a separate post.  So, given that only two other papers were pre-COH, I thought it may be better to shift this post more towards the higher-level discussion.

We all have other factors that will contribute to the total amount of time on a project (such as allocating some time for personal life), and I would usually expect work responsibilities to increase over time.  For example, if you are scrambling to complete your work as a graduate student, you may want to be cautious about setting goals for jobs that would have an even greater amount of responsibility.

Some people may be afraid of pointing out similar issues in previous papers.  While possibly somewhat counter-intuitive, I think this can help build trust in the associated papers / researchers: if researchers are not transparent in their actions and overall experience with a set of results, that can contribute to public distrust (and development of bad habits that can become worse over time).  Plus, if research is an on-going process of small steps towards an eventual solution, readers should expect each paper to acknowledge some limitations and/or unsolved problems (and a fair representation of results should help in identifying the most important areas for future research).

One relatively well-known example of the impossibility of being 100% accurate in all predictions is that Nobel Laureate Linus Pauling had a PNAS paper proposing a triple-helix structure for DNA.  There was even a Retraction Watch blog post that brought up the issue of whether this paper should be retracted.  I don't believe anyone currently cites that paper with the believe that DNA has a triple-helix (rather than a double-helix) structure.  However, taking time to correct and address mistakes needs to be taken into consideration for project management, and my point is that I am trying to encourage more self-regulation of corrections (since there are papers in relatively high impact journals whose main conclusion is wrong and they haven't been retracted or corrected).

While it harder to pass judgments on other people's work, I hope that I can be a good example for other people to identify issues with their own previous work.  For example, one counter-argument to the claim that most scientific arguments are wrong is the Jager and Leek 2014 paper where Figure 4 shows a science-wide FDR closer to 15%.  In one sense, this is good (15% is certainly better than 50% or 95%), but I think the correction / retraction rate is probably noticeably less than 15% (so, I think more scientists need to be correcting previous papers).  From my own record (of 1st author or equivalent papers), my correction rate is currently a little higher than that Jager and Leek estimate (3/8, or 37.5%) but my retraction rate is currently lower (0%).  I am not saying I will never have a retraction (or additional corrections).  In fact, I think there probably will be at least a couple additional corrections (among my total publication record).  However, that is my own personal estimate, and I would like to contribute to having discussions to try and reduce this correction rate for future studies.

I believe being open to discussion can cause you to temporarily lean towards agreement, as you try to see things from the other person's perspective (even if you eventually become more confident in your earlier claim).  So, even if a reviewer/editor considers a paper acceptable to publish or a grant is OK to fund (within a relatively short period of review process), the post-publication review is a very important (and I think making funded grants and/or grant submissions public and available for comment may also have value).

I also think that having less formal discussions (on Biostars, blogs, pre-prints, etc.) can help the peer-reviewed version of an article to be more accurate (if people actively comment in a location that is easy to check, reviewers take public comments into consideration, and/or journals use public comments to select reviewers that will provide the most fair assessment).  For multiple platforms, the Disqus comment system provides a centralized way to look for commentary for at least one peer-reviewed and at least one pre-print system.  While not linked directly from the paper, PubPeer also provides an independent commentary on journal articles.  I also have some examples of comments on both those mediums on this blog post.

While not the primary purpose, I think Twitter can also be useful for peer view.  For example, consider the contribution of a Twitter discussion to this Disqus comment.  Likewise, I found out about this article about the limits to peer review from Twitter.

I also have a set of blog posts summarizing experiences that describe the need for correction / qualification of results provided to the public for genomic products (although I think catching errors in papers, or better yet pre-prints, is really the preferable solution).  While maybe having something like the Disqus system for individual results (kind of like ClinVar, GET-Evidence, SNPedia, etc.) may have some advantages, people can currently give feedback in mediums like PatientsLikeMe (where I have described my experiences here) and the FDA MedWatch.

Update Log:

2/2019 - I would like to thank John Storey's tweet to reminding me of Jeff's SWFDR publication (in a draft, prior to the public post)
7/26/2019 - public post date
8/1/2019 - remove middle author link after realizing that there will be additional middle author corrections; add COHCAP corrigendum link; add PubPeer link based upon this tweet.
8/3/2019 - add Twitter links
8/6/2019 - switch link to updated genomics summary
1/16/2020 - add link to Disqus / PubPeer comment list
4/4/2020 - minor changes
7/11/2020 - minor changes
3/20/2022 - add Oncotarget RNA-Seq comment
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.