Genomic medicine is transforming healthcare, particularly for patients with rare diseases. The rapid evolution of next-generation sequencing has accelerated gene discovery and fueled widespread adoption of genetic testing as a frontline diagnostic tool. For patients with rare diseases and their families, a successful genetic diagnosis is critical. It not only gives a name to a set of medical concerns, but also can inform prognosis, recurrence risk, and patient management. A small but growing number of monogenic conditions now have FDA-approved gene or cellular therapies. One for Sanfilippo syndrome was approved just last week. Such therapies tend to be incredibly expensive, but they can be life-changing.
Despite all of these advances, diagnosis of rare disease patients remains challenging. Even the most comprehensive tests, such as clinical genome sequencing, fail to identify a cause for ~2/3 of patients with suspected genetic disorders. There are many reasons for this, but a major reason is that many gene-disease relationships remain undiscovered. The OMIM database currently has about 7,000 rare genetic diseases with a known molecular basis. There are many more which are known and published but not yet curated by OMIM, but that’s a separate issue. Regardless, the consensus in the field is that there are at least 10,000 distinct rare diseases, if not more. So we have a ways to go.
How Disease Genes Are Discovered
Historically, new gene-disease relationships could be established with the publication of one or a few patients. Now the bar is higher, typically requiring 5-15 patients to establish a new gene-disease relationships. These cohorts are often assembled through international matchmaking using databases like GeneMatcher and Matchmaker Exchange to connect investigators interested in the same gene. The other major source of new gene discovery is from large biobanks, such as Genomics England (GEL), which contain genetic and clinical data for thousands (or hundreds of thousands) of patients. GEL has collected detailed clinical data and whole-genome sequencing for something like 100,000 patients with rare genetic conditions. Their resource is highly secure and requires a daunting vetting/training process, but is available to qualified researchers through controlled access and has enabled many, many studies of new or emerging genetic diseases. The United States, Canada, and various countries are developing similar (albeit currently smaller) biobanks which combine medical and genetic data. Especially in the age of AI and machine learning, biobanks are powerful resources for identifying and establishing new gene-disease relationships.
Case in point: I recently read a study by Zonic et al which used a large biobank to evaluate potential gene-disease relationships that came up during clinical testing. The authors investigated 109 gene-disease pairs that came out of their diagnostic workflows (exome and genome sequencing) during an 11 month period (Jan-Nov 2022). Most of these (68%) were for neurological phenotypes. For each gene-disease pair, the authors collected evidence from public sources: OMIM, HGMD, Orphanet, GenCC, DECIPHER, and other databases. Just over half of gene-disease pairs (59/109) had an associated phenotype in the OMIM database, which means that almost half of them did not.
Establishing Gene-Disease Validity
The ClinGen validity framework, like most ClinGen frameworks, uses a point system. This particular one assigns points for genetic data (patients/families) and experimental data (animal or functional studies) that support a gene-disease relationship. Depending on how many points are assigned, the GDR can be classified as Limited, Moderate, Strong, or Definitive. First, the authors used public information to assign GDR validity. Then, they queried a Biorepository which contained genetic/clinical data for 670,000 patients from 120 different countries. For most of the gene-disease pairs (84/109 or 79%), biorepository data increased the validity score. For a good proportion in fact (21/109), the biobank data pushed the GDR into a higher-confidence category. Some 14 genes went from “Limited” to “Moderate” and 5 went from Moderate to Strong. There were even 2 genes (KCTD3 and NOTCH3) where internal data moved the confidence up two levels, from Limited to Strong.
All told, of the 109 gene-disease pairs investigated, 61 were classified as Moderate, Strong, or Definitive levels of evidence under the ClinGen framework. Do the math, and it’s clear that 1/3 of them got there with help from the internal biorepository data. Over 400 patients received a test report with a clinically relevant variant in one of those 61 genes. That’s great news, right? Applying a rigorous framework and leveraging a large curated biorepository allowed the authors to return useful genetic results to hundreds of patients. It sounds like a tremendous resource that rare disease investigators could use to identify and establish many more new GDRs.
Almost An Incredible Resource for Researchers
There’s just one problem. The wonderful repository of 670,000 patients is not a public resource. The CENTOGENE Biodatabank is proprietary. It is owned by (and the study authors work for) CENTOGENE GmbH, a private company based in Germany. The biobank data which support those 21 strengthened gene-disease relationships are not included in this manuscript, which some might say makes it impossible to replicate this work and thus violates the guidelines of scientific publication. At the very least, it’s not the best look for a paper published in a journal called Genetics in Medicine Open.
You might say hey, that’s fine. Surely CENTOGENE offers avenues to collaborate with rare disease investigators, making little slices of their cohort available to help establish gene-disease relationships. After all, there are companies in the US like GeneDx who have large internal databases but share data through places like GeneMatcher. Well, if the same is true of CENTOGENE, one might expect that many of their novel gene-disease relationships have since been published, either by collaborator groups or the company’s own scientists. And good news, the paper was published in 2023, so we can easily look to see if that’s happened.
Valid Gene-Disease Relationships but Not In OMIM
I took the 109 gene-disease relationships provided in Supplemental Table 1 of this paper, and looked at the 56 GDRs which were not in OMIM at the time of the publication three years ago. Of these, 33 are still not associated with a phenotype according to OMIM, and 12 of those were classified as Moderate or Strong gene-disease relationships in this study:
| Gene | MOI | Phenotype | Classification | References |
| BRSK2 | AD | BRSK2-related neurodevelopmental disorder | Strong | PMIDs: 30879638, 15705853 |
| FAT1 | AR | FAT1-related colobomatous-microphthalmia, ptosis, nephropathy, and syndactyly | Strong | PMIDs: 30862798, 26905694 |
| KCTD3 | AR | KTCD3-related developmental delay and seizure | Strong | PMIDs: 28454995, 29406573, 32552793, 23382386 |
| NCKAP1 | AD | NCKAP1-related neurodevelopmental disorder | Strong | PMIDs: 33157009, 28940097 |
| SCLT1 | AR | SCLT1-related multisystem ciliopathy | Strong | PMIDs: 24285566, 29450879, 27894351, 28005958, 30425282, 32253632, 28486600, 23348840 |
| SHANK1 | AD | SHANK1-related neurodevelopmental disorder | Strong | PMIDs: 34113010, 21695253, 18272690 |
| CBY1 | AR | CBY1-related ciliopathy disorder | Moderate | PMIDs: 33131181, 25103236, 25220153 |
| CTR9 | AD | CTR9-related neurodevelopmental disorder | Moderate | PMIDs: 31785789, 28554332, 28191890, 28714951, 31785789, 27520958, 35717577 |
| FBRSL1 | AD | FBRSL1-related malformation and intellectual disability syndrome | Moderate | PMID: 32424618 |
| FOXA2 | AD | FOXA2-related hyperinsulinism and hypopituitarism. | Moderate | PMIDs: 31294511, 30414530, 30684292, 29329447, 28973288, 11445544 |
| PRICKLE2 | AD | PRICKLE2-related neurodevelopmental disorder) | Moderate | PMIDs: 21276947, 23711981, 34092786, 26942291; |
| RAB11A | AD | RAB11A-related neurodevelopmental disorder | Moderate | PMIDs: 29100083, 28135719, 31785789 |
Several of the above genes have a fairly well-established relationship with disease which is simply not recognized by OMIM. For example, SHANK1 is recognized as a strong candidate gene for autism spectrum disorder by the SFARI database, and has had a supportive mouse model since 2008. KCTD3, one of the genes where CENTOGENE’s internal data significantly supported the GDR, had a pretty convincing mouse model published by Yeonsoo Oh et al in July of this year. Unfortunately, the gene’s OMIM page has not been updated since… *checks notes* 2013. Yikes.
Private Interest versus Public Good
So we’ve established that genetic data in this private biorepository might not be serving to help establish new gene-disease relationships. At least, examples like KCTD3 suggest that three years later, these genetic data which were not truly disclosed in the publication are still being kept privately aside. But hey, it’s perfectly reasonable for CENTOGENE to limit access to their commercial partners in the pharmaceutical industry, right? After all, they built and maintain this valuable asset, and they’re a private company. Surely the hundreds of thousands of patients whose clinical+genetic data are in this database are aware. Everything is good. Oh, but this wouldn’t be a Dan Koboldt blog post if I let them off the hook so easily. Come to think of it, I do wonder what the informed consent document signed by patients who undergo testing with CENTOGENE says about the use of their data.
Well, perhaps we should inquire?
Oh, great news. CENTOGENE’s patient consent document is readily available from their downloads page. In several languages, no less. I’d like to draw your attention to the optional clause labeled “Optional Research Consent to Further use of the Sample and Personal Data”:

Well, I’m not an expert but the highlighted section states pretty clearly the research concerns rare diseases in the interest of the greatest possible benefit to the general public. Strange. When it comes to rare diseases, data sharing is pretty essential. Keeping all of that patient data in a private, pay-to-play biodatarepository doesn’t seem like the greatest possible benefit to me.
Leave a Reply