All of Us Research Program

Diversity Drives Discovery: All of Us Research Program Identifies Hidden Genetic Variation

 

Introduction

Imagine a future where healthcare is not just a one-size-fits-all solution, but a personalized experience tailored to each individual’s unique genetic makeup. That future is being paved today by the All of Us Research Program, an ambitious initiative at the cutting edge of scientific innovation and inclusivity.1

With a mission to enroll a diverse group of at least one million individuals across the United States, this longitudinal cohort study is unique in its scope, focusing not on singular diseases or demographic groups but on constructing a comprehensive database to fuel studies across various health conditions. This approach opens the door to many possibilities, including:1

  • Uncovering risk factors for specific ailments
  • Determining which treatments are most effective for individuals from different backgrounds
  • Matching people with clinical studies tailored to their unique profiles
  • Exploring how innovative technologies can guide us toward healthier living

The achievements of the All of Us Research Program are already monumental. To date, the program has identified over one billion genetic variants, including more than 275 million previously unreported genetic variants, of which more than 3.9 million have potential coding consequences. This treasure trove of genetic data is unmatched in its diversity, with 77% of participants belonging to communities historically under-represented in biomedical research and 46% coming from under-represented racial and ethnic minorities.2

 

A Milestone in Genomic Research

In a recently published paper in Nature, the authors explored whole genome sequencing (WGS) data from 245,388 All of Us participants. The team performed a series of data harmonization and quality control procedures, analyzing the data for genetic ancestry and relatedness. The data was validated by replicating well-established genotype-phenotype associations and is available through the All of Us Researcher Workbench.2

What sets this data apart is not just its volume but its quality. Announcing the release of 245,388 clinical-grade genome sequences, All of Us has provided the scientific community with a resource of unparalleled diversity and depth. Linking this genomic data with longitudinal electronic health records has enabled the evaluation of 3,724 genetic variants associated with 117 diseases, demonstrating high replication rates across participants of varied ancestries.2

 

Cutting-Edge Technology for DNA Analysis

Central to the success of the All of Us Research Program was the use of an innovative approach to DNA sequencing. This crucial step in preparing DNA samples involved shearing the DNA by a Covaris focused-ultrasonicator—breaking it down into smaller, uniform pieces—using focused acoustic energy. This process ensured that the DNA was optimally prepared for sequencing, allowing for a higher resolution of genetic insights. Mechanical shearing with Covaris’ Adaptive Focused Acoustics® (AFA®) technology has been shown to improve target enrichment efficiency and provide more uniform batches with improved sequencing quality metrics. By reducing the need for additional sequencing and accommodating more samples on the flow cell, NGS efficiency is enhanced, and the associated costs are lowered while high-quality NGS data is delivered.3

In this study, Covaris’ AFA technology allowed for high confidence with variant calling as shown in Table 1.

 

Table 1. Sensitivity and precision measurements for control samples using the All of Us sequencing protocol)

Variant type NIST ID 1000 Genomes ID Sensitivity Precision
SNV HG-001 NA12878 0.995 >0.999
HG-003 NA24149 0.988 >0.999
HG-004 NA24143 0.988 >0.999
HG-005 NA24631 0.989 >0.999
Indel HG-001 NA12878 0.987 0.996
HG-003 NA24149 0.985 0.997
HG-004 NA24143 0.986 0.998
HG-005 NA24631 0.994 0.999

Source Nature 627, 340–346 (2024). https://doi.org/10.1038/s41586-023-06957-x

In addition, the clear reproducibility of results from across different labs indicates the Covaris workflow is one that customers can trust (Table 2). In fact, it is considered a gold-standard solution for WGS across distinct sample types (including blood, saliva, and urine).

 

Table 2. Batch effects across sequencing centers

Genome Center 1 Genome Center 2 Metric % Difference Cohen’s d (Effect Size) Computed Ancestry
Whole genome metrics with effect size greater than 0.5
Broad UW Indel Count 0.45 0.53 EAS
Broad UW Indel Count 0.49 0.51 EUR
Low mappability metrics with effect size greater than 0.5
Baylor Broad Indel Count -2.49 0.87 AFR
Broad UW Indel Count 2.91 1.05 AFR
Baylor UW SNP Count 1.69 0.52 AMR
Broad UW SNP Count 2.14 0.64 AMR
Baylor Broad Indel Count -2.49 0.65 AMR
Baylor UW Indel Count 2.22 0.65 AMR
Broad UW Indel Count 4.71 1.36 AMR
Broad UW SNP Count 1.12 0.74 EAS
Baylor Broad Indel Count -2.44 1.11 EAS
Baylor UW Indel Count 1.22 0.57 EAS
Broad UW Indel Count 3.65 1.66 EAS
Broad UW SNP Count 0.97 0.56 EUR
Baylor Broad Indel Count -2.5 1.04 EUR
Baylor UW Indel Count 1.21 0.52 EUR
Broad UW Indel Count 3.71 1.51 EUR
Broad UW SNP Count 1.07 0.59 SAS
Baylor Broad Indel Count -2.4 0.96 SAS
Broad UW Indel Count 3.36 1.34 SAS
Identified batch effects for the tandem repeat regions of the genome
Broad UW SNP Count 0.53 0.57 EAS
Broad UW SNP Count 0.61 0.58 EUR
Broad UW SNP Count 0.66 0.52 SAS
Broad UW Indel Count 1.7 0.52 AMR
Broad UW Indel Count 0.51 0.57 EAS
Broad UW Indel Count 0.57 0.56 EUR
Broad UW Indel Count 0.72 0.51 SAS
Batch effects seen in segmental duplication regions
Broad UW Indel Count 3.06 0.72 AMR
Baylor Broad Indel Count -1.66 0.55 EAS
Broad UW Indel Count 1.71 0.59 EAS
Broad UW Indel Count 1.72 0.53 SAS

Source Nature 627, 340–346 (2024). https://doi.org/10.1038/s41586-023-06957-x

 

Conclusion

The All of Us Research Program’s approach to generating diverse clinical-grade genomic data at an unprecedented scale can significantly enhance our understanding of human genetics. The discovery of over 275 million new genetic variants, previously undetected in similar large-scale projects, marks a significant step toward a future where precision medicine is accessible and equitable for everyone.2

In addition, the All of Us Researcher Workbench plays a crucial role in this vision, offering equal access and transparency opportunities to researchers and participants. With plans for regular genomic and phenotypic data updates, the initiative promises to keep the scientific community well-informed and engaged.2

As we look to the future, the collaboration with participants of the All of Us program is poised to transition research efforts from merely uncovering genomic variations to effectively applying genomic medicine on a broad scale. This evolution underscores the potential of this diverse dataset to fulfill the promise of genomic medicine, making it a reality for individuals across all backgrounds.2

 

References

  1. 1. https://allofus.nih.gov/about/program-overview
  2. 2. The All of Us Research Program Genomics Investigators. Genomic data in the All of Us Research Program. Nature 627, 340–346 (2024). https://doi.org/10.1038/s41586-023-06957-x
  3. 3. https://www.covaris.com/wp/wp-content/uploads/m020173.pdf