Showing posts with label public data. Show all posts
Showing posts with label public data. Show all posts

Thursday, April 30, 2020

Personal Thoughts on Collaboration and Long-Term Project Planning: Reproducibility and Depositing Data / Code

I believe that I broadly need to improve explanations for the need / value to deposit data and have code for the associated paper (even though that takes additional time and effort).

This is already a little different than the other sections, since it more of a question than a suggestion.

Nevertheless, as an individual, this is what I either currently do or I need to learn more about:
  • I am actively trying to better understand the details for proper data deposit for patient data (even though I have previously assisted with GEO and SRA submissions).
    • For example, I am trying to understand how patient consent relates to the need to have a controlled-access submission (even if that increases the time necessary to deposit data, or that certain projects should not be funded if the associated data cannot be deposited appropriately).  So, being involved with a successful dbGaP submission would probably be good experience.
    • I thought the rules were similar for other databases (like ArrayExpress, ENA, EGA, etc.).
    • However, if you know of other ways to appropriately deposit data, then I would certainly be interested in hearing about them!
  • If possible, I always recommend depositing data (and you see several papers where we did in fact do that), but I think different expectations would need to be set for supporting code (hence I have said things like "I cannot provide user support for the templates").
    • This is not to say I don't think code sharing is important.  On the contrary, I think it is important, but you have to plan for the appropriate amount of time to carefully keep track of everything needed to reproduce a result.
    • Also, if are a lot of papers where code has not been provided in the past, then I have to work on figuring out how to explain the need to share code (and spend more time per project, thus reducing the total number of projects that each lab/individual works on).
  • I think it would also be best if I could learn more about IRB/IACUC protocols (for both human and other animal studies).

In terms of how I can think of potentially emphasize the importance of data deposit and code sharing, you can see my notes below.  However, if you have other ideas about how to effectively and politely encourage PIs to deposit data and plan for enough effort to provide reproducible code (and/or help review boards not approve experiments producing data that can't be deposited), then I would certainly appreciate hearing about other experiences!

  • Even if it is not caught during peer review, I think journal data sharing requirements can apply for post-publication review?
  • The NIH has Genomic Data Sharing (GDS) policies regarding when genomics data is expected to be deposited.
    • There is also additional information about submitting genomic data, even if the study was not directly funded by the NIH.
    • There is also information about the NIH data sharing policies here.
  • While it mostly emphasized the need to expectation for data sharing with grants that are greater than $500,000, the NIH Data Sharing Policy and Implementation also mentions the need to code-related information available for reproducibility under "Data Documentation".
  • I also have this blog post on the notes that I have collected about limits to data sharing, but that is more about limiting experiments than data deposit for an experiment that has already been conducted.
  • Eglen et al. 2017 has some guidelines regarding sharing code.
  • This book chapter also discusses data and code sharing in the context of reproducibility.

Change Log:

4/30/2020 - public post
8/5/2020 - public post

Notes on Limits for Data Sharing

This overlaps with my post showing low-coverage sequencing data was identifiable information (with my own data).  However, I though having separate post to keep track of details still had some value.
  • Institutional Certification is required for patient data.  While some data collected before January 25th, 2015 can be deposited under controlled access without "explicit consent", this is not true for more recently collected samples.
    • For this reason, I would recommend not approving genomics studies with samples collected after this point, if such consent was not obtained (either in the original protocol, or in an amended protocol).
    • This also makes it important to get amendments to your IRB protocols, when you make changes.
    • The website can change over time.  In the event that the current website does not make clear that this applies to cell lines, you can see more explicit mention of cell lines here.
      • I believe the earlier website used the same language as the subheader on this form, saying "data generated from cell lines created or clinical specimens collected".
  • This means that you should not be able to create cell lines using samples collected more recently without "explicit consent" for either public or controlled access data deposit, since it will be extremely hard to enforce the appropriate use of the data after you share the cell lines with other labs.
    • I think that is consent with what is described in this article, which says "[consent] should be requested prior to generation" for cell lines.
      • The NIH GDS Overview says "For studies using cell lines or clinical specimens created or collected after [January 25th, 2015]...Informed consent for future research use and broad data sharing should have been obtained, even if samples are de-identified".
      • The NIH GDS FAQ also says "NIH strongly encourages investigators to transition to the use of specimens that have been consented for future research uses and broad sharing."
      • Additionally, the GEO human subject guidelines say "[it] is your responsibility to ensure that the submitted information does not compromise participant privacy[,] and is in accord with the original consent[,] in addition to all applicable laws, regulations, and institutional policies" (with or without NIH finding).
      • Plus, the NIH GDS FAQ says "investigators who download unrestricted-access data from NIH-designated repositories should not attempt to identify individual human research participants from whom the data were obtained".
    • HeLa cell lines were not obtained with the appropriate consent.  I believe that is why there is a collection of HeLa dbGaP datasets, since they are supposed to be deposited through a controlled access mechanism.  This is not always mentioned on the vendor website, and this is not always immediately enforced.  However, post-publication review applies to datasets and produces (as well as papers, which can be corrected or retracted).
      • In terms of HeLa cells, the genomic data is strictly expected to be deposited as controlled access, as explained in this policy.
    • If there is a way to check consent for cell lines, then I would appreciate learning about that.
    • As far as I know, the only cell lines that are confirmed to have consent to generate genetically identifying data to release publicly are those from the Personal Genome Project participants.  However, again, I would be happy to hear from others.
    • The ATCC website says "Genetic material deposited with ATCC after 12 October 2014 falls under the Convention on Biological Diversity and its Nagoya Protocol...It is the responsibility of end users that these undertakings are complied with and we strongly recommend that customers refer to this prior to purchase."
      • My understanding is that the United States has not joined this agreement.  However, I hope that this matches the sprit of other rules or guidelines from the NIH and HHS.  If I understand everything, I also hope the US joins at a later point in time.
  • In general, I think work done with low-coverage sequencing data can show that a lot of genomic data can be identifiable (which I think matches the need for controlled access and justification for not being allowed to create a cell line without the appropriate consent).
  • There is also this Blay et al. 2019 article describing kinship calculations with RNA-Seq data, also confirming the expectation that the raw FASTQ files contain identifiable information for most common RNA-Seq libraries.
    • The NIH GDS FAQ also includes "transcriptomic" and "gene expression" data as covered under GDS policies
  • I believe the above points may relate to the 2013 Omnibus rule, connecting the GINA and HIPAA laws.  As I understand it, I think you can find an unofficial summary here.
    • I believe that also matches what is described this link from the Health and Human Services (HHS) website (if it related to a health care provider).
    • There are general HIPAA FAQ for Individuals here, including a description of the HIPAA privacy rule here that explains HIPAA is intended to "[set] boundaries on the use and release of health records".
    • The links most directly above are from Health and Human Services (HHS).  However, in the research context, this article mentions the importance of taking genetic information into consideration with HIPAA/PHI/de-identification (which recommends controlled access if there is not appropriate consent for public deposit, since some raw genomic data may not be able to be truly de-identified).
    • At least for someone without a legal background like myself, I think "Under GINA, genetic information is deemed to be ‘health information’ that is protected by the Privacy Rule [citation removed] even if the genetic information is not clinically significant and would not be viewed as health information for other legal purposes." from Clayton et al. 2019 might be worth considering.
    • In other words, I believe that there are both NIH and HHS rules/guidelines that require or recommend care needs to be taken for patient genomic data.
  • I think some of the information from the Design and Interpretation of Clinical Trials Course course from Johns Hopkins University is useful.
    • Even in the research setting, the document from that course for the "Common Rule" includes "Identifiable private information" in the definition of "Human Subject" Research.
    • In the HIPAA privacy rule booklet for that course, it also says "For purposes of the Privacy Rule, genetic information is considered to be health information."  You can also see that posted here.

There are certainly many individuals (at work, as well as at the NIH, NCI, etc.) that have been helping me understand all of this.  So, thank you all very much!

Change Log:

4/30/2020 - public post
7/30/2020 - updates
8/5/2020 - updates
7/9/2021 - add information about RNA-Seq kinship
8/12/2021 - add information about Personal Genome Project cell lines and ATCC / Nagoya Protocol; formatting changes in main text and change log
8/17/2021 - add GDS FAQ and NIH HeLa notes
8/19/2021 - add GEO note
8/27/2021 - add HIPAA notes
11/23/2021 - add HIPAA notes
5/27/2022 - add cell line institutional certification notes
1/15/2023 - add PLOS Computational Biology reference link related to HIPAA/PHI
1/16/2023 - add Common Rule reference from JHU Coursera course + Clayton et al. 2019 reference
1/28/2023 - add note to make link from HHS page more clear + minor formatting changes

Wednesday, May 22, 2019

Speculative Opinion: Possible Advantages to Directly Providing Generics via Non-Profits


I believe there is a lot more I should learn more about this topic, and I have never been directly involved in a clinical trial.

Nevertheless, these are my current thoughts about the possibility of what might be interesting about providing having generics directly enter the clinic/market through non-profit organizations (admittedly largely influenced by my experiences in genomics, which may be less relevant for some other applications):

Possible Advantages to Patients / Physicians:

  • [data sharing / diagnostic transparency] Maximize public / accessible information available in order to help specialists make "best guesses" about how to proceed with available information
    • I don't believe that sale of access to raw genetic data should be allowed
    • Specialists / physicians should have access to maximal information to help guide decision making process
    • I think it would be nice if some information was completely public, such as population-level data from Color Genomics (even though re-processing data can probably change some variant calls, and this company isn't a non-profit).
    • In general, I think it is important not to place too much emphasis on any one study.  As an example of how that could skew a true estimate of risk, I think there is a useful barplot in this paper.  While over-fitting is not always the explanation, that can be a factor and I think I have a figure in this blog post that I hope can help explain that concern.
      • I am most familiar with this in the context of genomics (which would be for research or diagnostic purposes).  However, I have submitted several FDA MedWatch reports (again, mostly for diagnostics), and there is still a need for surveillance of therapeutics after they have entered the market.
  • [data availability for patient autonomy] Making sure patients have access to all data generated from their samples
    • Having access to your raw data should also help you be capable of getting specialized interpretation as a second opinion.
    • I also think self-reporting (with the ability of the patient to provide raw data) may help with regulation (or at least setting realistic expectations about efficacy / side-effects).
  • I think it may help if there was more judicious use of advertising.
    • Namely, I worry that some advertisements can give a false sense of confidence in the interpretation of results.  For example, I posted this draft a little early because of 23andMe's marketing of travel destinations based upon ancestry, which I don't approve of (although I support other overall goals for 23andMe).
    • That said, I think it can be useful when digital advertisements allow you to comment on them, kind of like a mini self-reporting system.
  • If we are talking about a therapy (rather than a diagnostic), I would expect this should also decrease costs (and is what I most commonly think of when I hear the word "generic").  Otherwise, I am mostly talking about experience with the exchange of information, often dependent upon sequencing/genotyping from another company (like an Illumina sequencer) that frequently makes use of open-source software (or analysis where unnecessarily complexity may sometimes even cause problems).

If this makes production via non-profit preferable, then perhaps a penalty for not meeting the above requirements could be an organization could risk losing it's non-profit status.  Otherwise, I am primarily concerned that the above conditions are met (at least in genomics), and I am just curious if being a non-profit might help in sustainable accomplishing that goal (although I lack knowledge on many of the accounting and legal details, and I don't have experience running a non-profit or for-profit organization).

Possible Advantages to Providers?

  • Assuming expectations are defined clearly and appropriately, participation in "on-going research" may improve understanding (and forgiveness) when there are many unknowns (and possibility even limits to what can be known with high-confidence in the immediate future)?
    • I called this "Decreased liability?" in an earlier version of this post, but I have gotten feedback that makes me question whether this is precisely what I want to describe.
    • If I understand things correctly (and it is possible to show that precisely defining all costs to society is difficult), it seems like forgoing royalties / extra profits in exchange for limited liability (kind of like open-source software, as I understand it) could be appealing in certain situations.
    • Strictly speaking, I see a warning of limited liability within the 23andMe Terms of Service (if you actually read through it).  However, I also know that I am entitled to $40 off purchasing another kit, because of the KCC settlement.  So, I would expect actually enforcing limited liability would be easier for a non-profit (if their profits were limited to begin with, it is harder to get extra money from them).
    • So, even though I believe the concept of limited liability applies in other circumstances, I think public opinion of the organization is important in terms of being patient and understanding when difficulties are encountered.
  • Decreased or lack of taxes paid by non-profit?
    • I think part of the point of having a non-profit is making the primary focus something other than money.  However, I think this link describes some financial advantages and disadvantages to starting a non-profit.
    • There was one person who raised concerns that non-products can't produce products (at least if I understood them correctly).  While I admit that I don't fully understand the tax law, I think connections to research, education, and/or "public goods" qualify for the examples that I am thinking of.  So, I can't tell if any rules need to be changed, but I found some summaries on-line that make me think things may currently be OK (such as here and here).
    • At least from my end, this page says what I thought of when I was saying something should be offered by a non-profit: "Charitable nonprofits typically have these elements:  1) a mission that focuses on activities that benefit society and whose goal is not primarily for profit, 2) public ownership where no person owns shares of the corporation or interests in its property, 3) income that must never be distributed to any owners but recycled back into the nonprofit corporation's public benefit mission and activities....In contrast, a for-profit business seeks to generate income for its founders and employees. Profits, made by sales of products or services, measure the success of for-profit companies and those profits are shared with owners, employees, and shareholders."

I also originally had a bullet point for "If profits are limited, what about refunds?".  However, I decided to place less emphasis on that point after additional feedback.  For example, I recently purchased an upgrade from 23andMe (for their V5 chip, from their V3 chip).  I noticed that I had to acknowledge that the purchase was non-refundable when I purchased the upgrade.  If it is possible (and/or tactful) for the company to provide refunds, then I think there are disadvantages to this style of not providing refunds.  However, this also made me think twice about how such an interaction would look if you were hesitant to give a refund because your profits were limited (and you have things like salary caps).  Most importantly, both non-profits and for-profits have to make sure they are not compromising safety (or unfairly representing their product).
While it is not the only reason why I think something should be provided from a non-profit, I think one characteristic of something that might need to be directly offered by a non-profit is something where there is a need to make sure the experts are in the habit of publicly announcing limitations (and mistakes) on a fairly regular basis.  In other words, if you can get an accurate estimate of a reasonable success rate, you can look more closely at situations where the success rate that either is exceptionally low or exceptionally high (although I would expect gradual improvement over time).

Also, to be fair, I think of "ownership" to be different when you talk about "owning" a pet versus "owning" a product to sell.  However, I think the concept of responsibility for the former is important, and it is also definitely possible that there are misconceptions in my understanding about the ways to provide something through a for-profit organization.

If it doesn't exist already, perhaps there can be some sort of foundation whose goal is to fund diagnostics / therapies that start as generics (without a patent)? If immediately offered as generic, perhaps there could be a non-profit donation suggestion at pharmacy or doctor's office (to a foundation that helps develop medical applications without patents)? Or, if this is not quite the right idea, perhaps another possible option that could be up for discussion could be early development in non-profit could translate into decreased time to become a generic (so, even if the non-profit is not directly providing the product with limited profit margins, the contribution of non-profit can still decrease costs to society).  This relates in part to an earlier post on obligations to publicly funded research, but I believe my current point is a little different.

There is precedent for the polio vaccine not having a patent, but my understanding that came at a great financial cost to the March of Dimes (and that is why more treatments don't enter the market without patents, even though the fundraising strategy was targeted to a large number of individuals that were already on tight budgets).

Genomics Data and Diagnostics

In "The Language of Life" Francis Collins describes the discovery of the CFTR gene.  After describing the invalidation of gene patients for Myriad, he mentions "my own laboratory and that of Lap-Chee Tsui insisted that the discovery of the CF gene, in 1989, be available on a nonexclusive basis to any laboratory that was interested in offering testing" (page 112) as well as saying "I donated all of my own patent royalties from the CF gene discovery to the Cystic Fibrosis Foundation" (page 113).

My understanding is the greatest barrier to having products frequently start out as generics is the cost of conducting the clinical trial.  I need to be careful because I don't have any first-hand experience with clinical trails, but are some possible ideas that I thought might be worth throwing out as ideas:
  1. Allow data sharing to help with providing information to conduct clinical trails.  For example, lets say the infrastructure from a project like All of Us allows people to share raw data from all diagnostics (and electronic medical records), as well as archived blood draws and urine samples.  Now, let's say you have a diagnostic that you want to compare to previously available options.  If the government has access to the previous tests, the original samples, and the ability to test your new diagnostic, maybe use of that information can be combined with an agreement to provide your diagnostic as a generic (with understanding that continued surveillance also serves as an additional type of validation) is a fair trade-off?
    • I'm not sure if this changes how we think of clinical trails, but I think participants should also be allowed to provide notes over the long-term (after you would usually think of the trial as ending).  This would kind of be like post-publication review for papers, and self-reporting in a system like PatientsLikeMe (which I talk about more in another post).
    • Side effects are already monitored for drugs on the market
  2. Define a status for something that can be more easily tested by other scientists if passes safety requirements (Phase I?) but not efficacy requirements?  I guess this would be kind of like a "generic supplement," but it should probably have a little different name.
I also believe that all participants need to have access to their own data (including the ability to look up papers that use their data for publication), but I realize that this doesn't necessarily have to be part of a clinical trail because I have accessed patient genomics data from archived samples and donors/subjects (for which I think the rules are a little different).  Nevertheless, I think it is important and relevant to the points that I am making about patients having access to their raw data.

For some personalized treatments, I would guess you might even have difficulties getting a large enough sample size to get beyond the "experimental" status (equivalent to not being able to complete the clinical trial?). Plus, if some drugs have 6-figure price tags (or even 7-figure price tags), maybe some people would even consider getting a plane ticket to see a specialist for an "experimental" trial / treatment.

Role of the FDA

From what I can read on-line, I believe there is some interest in the FDA helping with generic production, and this NYT article mentions "[the FDA] which has vowed to give priority to companies that want to make generics in markets for which there is little competition", in the context of a hospital producing drugs.  According to this reference, "80 percent of all drugs prescribed are generic, and generic drugs are chosen 94 percent of the time when they are available."

Perhaps it is a bit of a side note, but I was also playing around with the FDA NDC Database (which is an Text / Excel file that you can download and sort).  For example, I could tell my Indomethacin was produced by Camber Pharmaceuticals by one pharmacy (NDC # 31722-543), and my Citalopram from another pharmacy was produced by Aurobindo Pharma Limited (even though Camber Pharmaceuticals also manufactures Citalopram, and Aurobindo also produces Indomethacin Extended-Release, according to the NDC Database).  I thought it was interesting to see how many companies produce the same generic and how many generics are produced by each company.  At least to some extent, this seems kind of like how there may be similar topics studies by labs in different institutes across the world.  So, maybe there can even be some discussions about how there can be both sharing information for the public good as well as independent assessments of a product from different organizations (whether that be a lab in a non-profit or a company specializing in generics).

I also noticed that the FDA has a grant for "complex" generics, but I believe that is for current drugs that are off-patent but there were extra challenges with production that make offering a generic version more difficult.  Nevertheless, it is evidence that there is some belief that academic and non-profit institutes may be able to help bring generics to the market more quickly.

Personal Experience / Open-Source Bioinformatics Software

I believe that I need to work on fewer projects more in-depth.  I wonder if there might be value in having a system for independence that would allow PIs to do the same (with increased responsibility/credit/blame at the level of the individual lab).  If something entered the market as a generic (possibly from a non-profit), perhaps the same individuals can be involved with both development and production of the generic.

Also, for my job as a Bioinformatics Specialist, I mostly use open-source software (but I sometimes use commercial software or software that is only freely available to non-profits/academics).  In particular, I think it is very important to have access to multiple freely available programs, and the topic of limits to precision in genomics methods (at least in the research context) is something I touch on in my post about emphasizing genomics for "hypothesis generation" (at least in the research context).

Concluding Thoughts

Even if is not used in clinical trails (which, as far as I know, was not part of the original plan), I think All of US matches some of what I am describing as a generic from a non-profit (even though it isn't called a "generic," it is a government operation, and free sequencing is not currently guaranteed after sample collection).  Nevertheless, non-profit (or academic) Direct-to-Consumer options that I think more people should know more about include Genes for Good (free genotyping), American Gut (can still be ordered from Indiegogo?), the UC-Davis Veterinary Genetics Lab, and I am excited to learn more about others.  I think this may also be in a similar vein to DIYbio clubs (for example, I believe Biocurious provides a chance to do MiSeq sequencing).  Cores (like where I work) also kind of do this (for labs), but I can tell that I need to work on fewer projects more in-depth (so, I think there would need to be some changes before adopting a "core" model for producing generics).

Finally, I want to make clear that this is something that I would like to gradually learn more about, but that is probably more on the scale of 5-10 years.  That is generally what I am trying to indicate when I add "Speculative Opinion" to a blog post title.  So, I very much welcome feedback, but my ability to have extended discussions on the topic may be limited.

The only things that I feel strongly about in the immediate future is not reversing the Supreme Court decision to not allow genes to be patented, and the limits to predictive power for some genomics methods (such as the concerns I expressed about the 23andMe ancestry results towards the beginning of this post, and how I don't believe it would be appropriate to encourage travel destinations to specific countries).

Change Log:

5/22/2019 - original post date
-I should probably give some amount of credit for the idea of emphasizing decreased health care costs to Ragan Robertson (for his answer to my SABPA/COH Entrepreneur Forum question about generics and providing something in a non-profit versus commercial setting).  However, his answer was admittedly more focused on mentioning how generics could be used for different "off-label" applications after they have entered the clinic (as well as connecting this to decreased health care costs).
5/23/2019 - update some information, after Twitter discussion
5/24/2019 - trim out 1st paragraph
5/25/2019 - move open-source software paragraph towards end.  Also, lots of editing for the overall post.
5/26/2019 - remove sentence with placeholder for shared resources post that is currently only a draft.  Add link to $2.1 million drug treatment tweet (with every interesting comments)
6/1/2019 - remove the word "their" from 23andMe travel sentence
6/27/2019 - update content in response to discussion with family member.  For example, I don't think I was making clear that I was primarily concerned about data sharing / transparency and continuing to not allow genetic testing / information to be patented, at least in the field of genomics (and I am curious if being a non-profit can play a helpful role if those requirements are met).
6/28/2019 - revise explanation for the previous change log entry
6/29/2019 - bring up tax details
7/13/2019 - add explanation for "Speculative Opinion"
7/30/2019 - add comments for Francis Collin's CFTR gene discovery
11/3/2019 - add "speculative opinion" tag
5/6/2020 - minor changes
8/22/2020 - add a couple additional links
10/4/2020 - add FDA "complex" generic grant link + minor changes + add section headers
 
Creative Commons License
Charles Warden's Science Blog by Charles Warden is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 United States License.