Breast Cancer Risk: New Multi-Ancestry Genetic Prediction Models Developed

by Grace Chen
Breast Cancer Risk: New Multi-Ancestry Genetic Prediction Models Developed

Researchers have built comprehensive multi-ancestry genetic prediction models for breast tissue gene expression, alternative polyadenylation, and splicing events using normal breast tissue samples from diverse female donors and high-throughput sequencing data, providing critical new resources for investigating breast cancer risk and biology across multiple populations.

Pooling Multi-Ancestry Tissue Resources for Genetic Models

The overall genetic prediction models for gene expression, alternative polyadenylation, and splicing events stem from normal breast tissue data compiled across three key resources: the Susan G. Komen Normal Tissue Bank, the Asia Breast Cancer Consortium, and the Genotype-Tissue Expression project. This research received approval from the Vanderbilt University Medical Center Institutional Review Board. Investigators utilized normal breast tissue samples from 381 female donors to the Komen Tissue Bank alongside samples from 156 Asian-ancestry female participants from the Asia Breast Cancer Consortium to generate fresh genomic and transcriptomic data. These newly gathered datasets were analyzed in tandem with existing normal breast tissue information from the Genotype-Tissue Expression project.

Tissue preservation played a vital role in sample quality. Within the Komen Tissue Bank, normal breast biopsies were harvested from the upper outer quadrant of the breasts and frozen within an average of 6 minutes from the time of biopsy. Participant inclusion criteria focused entirely on female donors due to the study’s orientation around breast cancer risk among women. Genome-wide association study summary statistics used for association analyses excluded male breast cancer cases. Sex determination relied on self-reports confirmed by genotype-based quality-control checks, while ancestry groupings referred strictly to genetic ancestry rather than self-reported race.

High-Throughput Sequencing and Genotyping Platforms

Laboratory processing involved rigorous high-throughput sequencing and genotyping protocols. Genomic DNA samples from all 381 females in the Komen bank and 156 females from the Asia Breast Cancer Consortium were genotyped utilizing the Illumina MEGA platform at Vanderbilt University Medical Center. An additional 73 Asian samples from the Asia Breast Cancer Consortium underwent genotyping via the Affymetrix Array platform.

Model-building filters removed genetic variants possessing a minor allele frequency under 5 percent, a Hardy-Weinberg equilibrium p-value below 10−4, and a missing genotyping rate exceeding 5 percent. Researchers also excluded single nucleotide polymorphisms exhibiting a consistency rate under 98 percent among duplicate samples. Imputation of genotype data utilized the Trans-Omics for Precision Medicine as a reference panel, dropping genetic variants with an imputation quality score below 0.8.

Total RNA extraction and purification followed manufacturer instructions using Qiagen’s AllPrep DNA/RNA/miRNA Universal Kit. Sample quantity and quality underwent verification via Nanodrop ratios and Agilent BioAnalyzer separation, while RNase H removed ribosomal RNA. Each sample underwent pair-ended sequencing with a read length of 100 base pairs using DNBSEQ on BGISeq, securing a minimum of 10M reads per sample.

Mapping Mechanism-Centric Approaches Against Gene-Centric Limits

Broader computational frameworks in cancer research emphasize that biomarker discovery serves as the foundation for personalized treatment planning, disease classification, prognosis, and therapeutic targeting. However, numerous conventional biomarkers represent passenger alterations rather than true driver modifications. This limitation restricts their utility as functional units for therapeutic intervention. Mechanism-centric approaches account for upstream and downstream regulatory mechanisms to identify driver biomarkers.

Classical gene-centric methods typically investigate isolated genes based on previous biological assumptions or select markers via differential behavior without mapping upstream and downstream connections. Such findings often remain dataset-specific and functionally uninterpretable. White-box models, including linear regression and decision trees, provide understandable relationships between input variables and disease outcomes, whereas black-box machine learning models capture complex patterns across various oncology domains.

Participant Demographics and Ethical Consenting Protocols

Participant profiles across the tissue repositories varied by age and ancestry distribution. The mean age of participants at tissue collection stood at 47 years with a standard deviation of 12.7 for the Komen Tissue Bank samples, 46 years with a standard deviation of 7.9 for the Asia Breast Cancer Consortium samples, and 51 years with a standard deviation of 13.1 for the Genotype-Tissue Expression samples. The Komen Tissue Bank samples comprised 182 European-ancestry participants, 49 Asian-ancestry participants, and 150 African-ancestry participants.

Women of African ancestry and the risk of breast cancer.

Acquisition of specimens adhered strictly to ethical guidelines. De-identified participant annotation data and normal breast tissue specimens from the Komen Tissue Bank were secured through an approved tissue request and material transfer agreement, supported by prior written informed consent obtained under institutional review board protocols. Genotype-Tissue Expression data access operated through controlled-access database procedures in accordance with approved data-use terms.

Integrating Cross-Ancestry Analysis Consortia

The integration of multi-ancestry transcriptome-wide association studies relies heavily on summary statistics derived from major breast cancer genetic research consortia. Cross-ancestry multi-study association analyses combined data from the Asia Breast Cancer Consortium, AABCG, and BCAC, with all underlying protocols receiving local institutional ethical committee approval.

Content cover image
Photo: Nature

Future directions in precision therapeutics increasingly emphasize model interpretability, data integration, and advanced single-cell sequencing technologies. Overcoming existing barriers in biomarker reproducibility and tumor heterogeneity requires combining high-resolution multi-omic datasets with robust computational models capable of deciphering complex genetic architectures across diverse global populations.

Breast density, genetics, and breast cancer risk among Black women – Anne Marie McCarthy, SCM, PhD

You may also like