Showing posts with label cell line. Show all posts
Showing posts with label cell line. Show all posts

Thursday, April 30, 2020

Notes on Limits for Data Sharing

This overlaps with my post showing low-coverage sequencing data was identifiable information (with my own data).  However, I though having separate post to keep track of details still had some value.
  • Institutional Certification is required for patient data.  While some data collected before January 25th, 2015 can be deposited under controlled access without "explicit consent", this is not true for more recently collected samples.
    • For this reason, I would recommend not approving genomics studies with samples collected after this point, if such consent was not obtained (either in the original protocol, or in an amended protocol).
    • This also makes it important to get amendments to your IRB protocols, when you make changes.
    • The website can change over time.  In the event that the current website does not make clear that this applies to cell lines, you can see more explicit mention of cell lines here.
      • I believe the earlier website used the same language as the subheader on this form, saying "data generated from cell lines created or clinical specimens collected".
  • This means that you should not be able to create cell lines using samples collected more recently without "explicit consent" for either public or controlled access data deposit, since it will be extremely hard to enforce the appropriate use of the data after you share the cell lines with other labs.
    • I think that is consent with what is described in this article, which says "[consent] should be requested prior to generation" for cell lines.
      • The NIH GDS Overview says "For studies using cell lines or clinical specimens created or collected after [January 25th, 2015]...Informed consent for future research use and broad data sharing should have been obtained, even if samples are de-identified".
      • The NIH GDS FAQ also says "NIH strongly encourages investigators to transition to the use of specimens that have been consented for future research uses and broad sharing."
      • Additionally, the GEO human subject guidelines say "[it] is your responsibility to ensure that the submitted information does not compromise participant privacy[,] and is in accord with the original consent[,] in addition to all applicable laws, regulations, and institutional policies" (with or without NIH finding).
      • Plus, the NIH GDS FAQ says "investigators who download unrestricted-access data from NIH-designated repositories should not attempt to identify individual human research participants from whom the data were obtained".
    • HeLa cell lines were not obtained with the appropriate consent.  I believe that is why there is a collection of HeLa dbGaP datasets, since they are supposed to be deposited through a controlled access mechanism.  This is not always mentioned on the vendor website, and this is not always immediately enforced.  However, post-publication review applies to datasets and produces (as well as papers, which can be corrected or retracted).
      • In terms of HeLa cells, the genomic data is strictly expected to be deposited as controlled access, as explained in this policy.
    • If there is a way to check consent for cell lines, then I would appreciate learning about that.
    • As far as I know, the only cell lines that are confirmed to have consent to generate genetically identifying data to release publicly are those from the Personal Genome Project participants.  However, again, I would be happy to hear from others.
    • The ATCC website says "Genetic material deposited with ATCC after 12 October 2014 falls under the Convention on Biological Diversity and its Nagoya Protocol...It is the responsibility of end users that these undertakings are complied with and we strongly recommend that customers refer to this prior to purchase."
      • My understanding is that the United States has not joined this agreement.  However, I hope that this matches the sprit of other rules or guidelines from the NIH and HHS.  If I understand everything, I also hope the US joins at a later point in time.
  • In general, I think work done with low-coverage sequencing data can show that a lot of genomic data can be identifiable (which I think matches the need for controlled access and justification for not being allowed to create a cell line without the appropriate consent).
  • There is also this Blay et al. 2019 article describing kinship calculations with RNA-Seq data, also confirming the expectation that the raw FASTQ files contain identifiable information for most common RNA-Seq libraries.
    • The NIH GDS FAQ also includes "transcriptomic" and "gene expression" data as covered under GDS policies
  • I believe the above points may relate to the 2013 Omnibus rule, connecting the GINA and HIPAA laws.  As I understand it, I think you can find an unofficial summary here.
    • I believe that also matches what is described this link from the Health and Human Services (HHS) website (if it related to a health care provider).
    • There are general HIPAA FAQ for Individuals here, including a description of the HIPAA privacy rule here that explains HIPAA is intended to "[set] boundaries on the use and release of health records".
    • The links most directly above are from Health and Human Services (HHS).  However, in the research context, this article mentions the importance of taking genetic information into consideration with HIPAA/PHI/de-identification (which recommends controlled access if there is not appropriate consent for public deposit, since some raw genomic data may not be able to be truly de-identified).
    • At least for someone without a legal background like myself, I think "Under GINA, genetic information is deemed to be ‘health information’ that is protected by the Privacy Rule [citation removed] even if the genetic information is not clinically significant and would not be viewed as health information for other legal purposes." from Clayton et al. 2019 might be worth considering.
    • In other words, I believe that there are both NIH and HHS rules/guidelines that require or recommend care needs to be taken for patient genomic data.
  • I think some of the information from the Design and Interpretation of Clinical Trials Course course from Johns Hopkins University is useful.
    • Even in the research setting, the document from that course for the "Common Rule" includes "Identifiable private information" in the definition of "Human Subject" Research.
    • In the HIPAA privacy rule booklet for that course, it also says "For purposes of the Privacy Rule, genetic information is considered to be health information."  You can also see that posted here.

There are certainly many individuals (at work, as well as at the NIH, NCI, etc.) that have been helping me understand all of this.  So, thank you all very much!

Change Log:

4/30/2020 - public post
7/30/2020 - updates
8/5/2020 - updates
7/9/2021 - add information about RNA-Seq kinship
8/12/2021 - add information about Personal Genome Project cell lines and ATCC / Nagoya Protocol; formatting changes in main text and change log
8/17/2021 - add GDS FAQ and NIH HeLa notes
8/19/2021 - add GEO note
8/27/2021 - add HIPAA notes
11/23/2021 - add HIPAA notes
5/27/2022 - add cell line institutional certification notes
1/15/2023 - add PLOS Computational Biology reference link related to HIPAA/PHI
1/16/2023 - add Common Rule reference from JHU Coursera course + Clayton et al. 2019 reference
1/28/2023 - add note to make link from HHS page more clear + minor formatting changes

Thursday, November 14, 2019

What do I need to change as an individual?

I can tell that I need to work on fewer-projects in more depth.

I am not sure if additional training is necessary to accomplish this, but that is the focus of my post on "What are the expectations for Individuals with an MS in Bioinformatics versus a PhD in Genomics?"?

Technique-Wise, these are the sort of things that I think I could be comfortable with:


  • I am very comfortable coding in R / Python / Perl
  • If I needed to get back into the lab, I could previously do a PCR and maintain a cell line
    • At least previously, I had some difficulties being able to preform my own microarray experiments
    • However, to be honest, I think the best fit for me would be to keep doing mostly or entirely computational work (as a Bioinformatics Specialist, Bioinformatician, etc.).
  • I think I likely need to reduce the number of new patient samples that I encounter, but I am comfortable working with my own genomics data and I think I have had useful contributions using re-analysis of data deposited by other labs.
    • For example, I thought this was a relatively successful story of my feedback on a pre-print being helpful
    • I also have notes on my human genomics results here
    • I also have on-going analysis of public cell line perturbations to demonstrate method limits for RNA-Seq analysis


Research-Wise, there are topics that I am interested in and/or have some prior experience with:


  • Investigate whether a relatively simple strategy is more robust than a more complicated strategy (for example, try to identify problems with over-fitting)
  • Probably a good idea to either limit the number of samples I work on at a given time and/or make sure that I have enough time for several rounds of analysis / discussion of the same dataset (which is always good for critical assessment of results)
  • Perhaps place more focus on non-human genomics (and I started out doing evolutionary genomics research)?
    • I have some previous virology experience, so perhaps I could study the genetics / genomics of viruses that currently infect other animals (but have not yet evolved the ability to infect people)?  This could even be part of DNA-Seq or RNA-Seq for the host.
    • While I haven't done any such research from a professional standpoint, I have been comparing genomics results for my cat Bastu, and I hope to have a blog post summarizing that soon.
  • If I can get agreement about some things, perhaps there is some value in shared support guidelines?
    • I also truly like the idea of supporting labs with less popular research topics and/or smaller budgets (which may have a relatively greater need for shared support)
    • Essentially, I am emphasizing training (indirect analysis support) and limits on projects for shared staff
    • However, I don't particularly like telling other people what to do (the degree discussion is really more about having the right amount of autonomy and peer respect to perform my own analysis), even though I realize that we sometimes have obligations to society to do things that we may not find enjoyable.
    • So, I think finding a solution for myself what is currently most important, but I think I have some experiences that may be useful to others.
    • In general, if I were to assist with training support, I think I may be able to help with common public datasets, but I think there may need to be a rule that I couldn't help with providing analysis for a new dataset (unless a greater commitment to the project is made, where the limits on shared support would then be important)
    • While helping me stay up-to-date and refreshed on details, perhaps providing local guidance (face-to-face) for a subset of content from selected on-line courses (like Coursera) may be an appropriate way for me to help, but it would be crucial that I not complete exercises for any students (which would violate the honor code, requiring students to complete their own work and demonstrate independent competence).  For example, if I did this before they started the course, perhaps I could then recommend where to learn more and get certification (as well as setting realistic expectations on the likelihood of passing the course).
  • I would need a lot of practice, but outreach / optional education for the general public (such as a book club discussion that I led) can be rewarding 
  • While I am less certain about my role in a professional standpoint, you can see my "speculative opinion" posts about some things that I think could be interesting


Personality-Wise, these are what I believe are my strengths and weaknesses:


  • I have to be fairly independent for my current job, but I do provide a supportive role (where biological / clinical idea usually comes from PI)
  • While the difference between 1 day and 2 weeks turnaround time would be an order of magnitude (for each iteration of analysis/discussion), I have received the good suggestion that I should wait at least an extra day before returning each round of results (to see if I can catch more errors by reviewing the results again the next day).
  • I like the idea of helping provide a "public good," so I think I would prefer to continue working at non-profits
  • Continue becoming better at more mindful when I have a prior assumption (which may or may not be true) and I may not sufficiently understand other perspectives.
    • However, I appreciate those who value the need to take time to be objective and fair
  • While I believe it is important to continue to make future progress, I have some concerns about responsibilities what require excellent communication (and would frequently involve relatively short interactions with individuals where I may not have the chance to correct myself)
    • For example, part of the reason I work on the computer is that I bugs will stay corrected in the code (once I find them).
    • In contrast, if you knew that there was an experimental protocol that required X steps and I was highly like to mess up at least 1/X steps, then that is sufficient for me to not be able to get a protocol to work (for example, I think this is why I had previous difficulty with performing my own microarray experiment).
  • I think I may need to better recognize what I can fix (for myself), versus a concern about the actions of others (which may be solvable if properly communicated, or may be harder to resolve without common agreement)
    • For example, there may be some room for improvement in terms of communicating myself in sensitive situations (such as disagreeing with a policy and/or a superior).  However, this is something that I am actively working on.
  • Most of my family lives on the east coast of the United States (and I currently live in the west coast, in California).  As we get older, this may be something worth taking into consideration.



Change Log:

11/14/2019 - public post
11/15/2019 - public post
11/21/2019 - fix typos + add waiting a little longer to return results
1/7/2020 - add link for on-line course notes
4/30/2020 - update cat link for blog post versus GitHub
10/7/2020 - add note about computational emphasis
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.