Articles
The Dark Genome and the Future of Predictive Genomics
By Susanna Zucca, CSO, enGenome
For much of the genomics era, our understanding of the human genome has been shaped by what we could reliably sequence and interpret. But a significant fraction of the genome, sometimes referred to as the “dark genome”, has remained difficult to study. These regions include highly repetitive sequences, structural duplications, and complex chromosomal regions such as centromeres and telomeres that historically resisted standard sequencing and assembly methods.
The emergence of telomere-to-telomere (T2T) genome assemblies is beginning to change that picture. The first complete human genome sequence, published by the Telomere-to-Telomere consortium in 2022, closed the remaining gaps in the reference genome and added nearly 200 million previously unresolved base pairs, including regions that had been invisible to genomic analysis for decades. https://pmc.ncbi.nlm.nih.gov/articles/PMC9336181/
This milestone represents more than a technical achievement. It lays the groundwork for a new phase of predictive genomics, where previously hidden genomic variation can begin to inform disease risk, biological function, and therapeutic decision-making.
Why the “dark genome” matters clinically
Many of the regions historically excluded from the reference genome contain structural variants, segmental duplications, and regulatory sequences that influence gene expression and genome stability. These regions can play important roles in neurological disorders, immune diseases, cancer, and rare genetic conditions.
Because earlier reference genomes omitted or collapsed these repetitive sequences, variant calling and mapping in these regions were often unreliable. The more complete T2T reference improves read mapping accuracy and variant detection, enabling researchers to identify previously hidden variants and potentially uncover new disease mechanisms. https://genome.cshlp.org/content/35/11/2377.full.pdf
For clinical genomics, this has two important implications:
- Improved rare-disease diagnostics: Variants in complex genomic regions that were previously undetectable may now be identifiable.
- More comprehensive structural variant analysis: Long-read sequencing and improved references allow more accurate identification of duplications, inversions, and other structural changes that often drive disease.
As sequencing technologies continue to improve, the expectation is that complete, phased genomes spanning telomere to telomere will eventually become routine in human genetics. But generating more complete genomes is only the first step.
The interpretation challenge
The transition from sequencing to clinical interpretation remains the central challenge of genomic medicine.
New reference assemblies and long-read technologies are revealing vast numbers of variants in previously inaccessible genomic regions. However, translating these discoveries into clinically actionable insights requires new annotation frameworks, better population references, and more sophisticated computational approaches.
Several challenges remain:
1. Reference diversity
The current complete human genome (T2T-CHM13) represents essentially a single haplotype and cannot capture the full structural diversity across human populations.
This is where the emergence of the human pangenome becomes critical. Rather than relying on a single linear reference genome, pangenome approaches aim to represent genomic diversity across multiple individuals, capturing population-specific variation and alternative genomic structures. By incorporating diverse haplotypes into graph-based or multi-reference frameworks, the pangenome provides a more representative foundation for mapping reads and identifying variants, particularly in regions of high structural complexity.
For clinical genomics, this shift is significant. It enables more accurate variant detection across diverse populations and reduces reference bias, particularly in genomic regions that vary substantially between individuals. In the context of the “dark genome,” where structural variation is common, pangenome references are likely to play a key role in making these regions interpretable at scale.
2. Variant interpretation in complex regions
Highly repetitive regions still complicate alignment and structural variant detection, even with long-read sequencing.
3. Functional understanding
Many newly revealed genomic elements lack clear functional annotation, making it difficult to assess their clinical significance. In other words, the dark genome is now becoming visible, but its biological meaning is still largely unexplored.
From sequencing completeness to predictive genomics
For genomic medicine to fully benefit from T2T assemblies and expanded genomic references, the field must move beyond sequencing alone and focus on interpretation at scale.
Predictive genomics requires:
- integrated variant annotation across coding and non-coding regions
- robust interpretation frameworks for structural variation
- population-scale datasets to contextualize rare variants
- AI-driven tools that can prioritize clinically relevant signals within increasingly complex genomes
This is where genomic interpretation platforms become critical. As sequencing technologies illuminate previously hidden regions of the genome, the ability to interpret these signals in a clinical context will determine whether they translate into improved diagnosis and predictive medicine.
The next frontier
The completion of the human genome does not mark the end of the genomics era, it marks the beginning of a more complex one.
For the first time, researchers can explore the full architecture of the human genome, including the regions that were once inaccessible. But unlocking the clinical value of this expanded genomic landscape will depend on combining better reference genomes, richer functional annotation, and scalable interpretation frameworks.
In short, the dark genome is becoming visible, the real challenge now is learning how to read it.