组装未映射的 reads 揭示了南亚基因组中隐藏的变异。

Assembling unmapped reads reveals hidden variation in South Asian genomes.

作者信息

Das Arun, Biddanda Arjun, McCoy Rajiv C, Schatz Michael C

机构信息

Department of Computer Science, Johns Hopkins University, Baltimore, MD, 21218, USA.

Department of Biology, Johns Hopkins University, Baltimore, MD 21218, USA.

出版信息

bioRxiv. 2025 May 14:2025.05.14.653340. doi: 10.1101/2025.05.14.653340.

DOI:10.1101/2025.05.14.653340

PMID:40463162

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC12132312/

Abstract

Conventional genome mapping-based approaches systematically miss genetic variation, particularly in regions that substantially differ from the reference. To explore this hidden variation, we examined unmapped and poorly mapped reads from the genomes of 640 human individuals from South Asian (SAS) populations in the 1000 Genomes Project and the Simons Genome Diversity Project. We assembled tens of megabases of non-redundant sequence in tens of thousands of large contigs, much of which is present in both SAS and non-SAS populations. We demonstrated that much of this sequence is not discovered by traditional variant discovery approaches even when using complete genomes and pangenomes. Across 20,000 placed contigs, we found 8,215 intersections with 106 protein coding genes and >15,000 placements within 1 kbp of a known GWAS hit. We used long read data from a subset of samples to validate the majority of their assembled sequences, aligned RNA-seq data to identify hundreds of unplaced contigs with transcriptional potential, and queried existing nucleotide databases to evaluate the origins of the remaining unplaced sequences. Our results highlight the limitations of even the most complete reference genomes and provide a model for understanding the distribution of hidden variation in any human population.

摘要

基于传统基因组图谱的方法会系统性地遗漏遗传变异，尤其是在与参考序列有显著差异的区域。为了探索这种隐藏的变异，我们检查了来自千人基因组计划和西蒙斯基因组多样性计划中640名南亚（SAS）人群基因组的未映射和映射不佳的读数。我们在数万个大的重叠群中组装了数十兆碱基的非冗余序列，其中许多在SAS和非SAS人群中都存在。我们证明，即使使用完整基因组和泛基因组，传统的变异发现方法也无法发现这些序列中的大部分。在20,000个已定位的重叠群中，我们发现8,215个与106个蛋白质编码基因有交集，并且在已知全基因组关联研究（GWAS）命中位点的1千碱基范围内有超过15,000个定位。我们使用来自一部分样本的长读数据来验证它们组装序列的大部分，比对RNA测序数据以识别数百个具有转录潜力的未定位重叠群，并查询现有的核苷酸数据库以评估其余未定位序列的来源。我们的结果突出了即使是最完整的参考基因组的局限性，并为理解任何人类群体中隐藏变异的分布提供了一个模型。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/02de/12132312/17b8106103be/nihpp-2025.05.14.653340v1-f0001.jpg

相似文献

Assembling unmapped reads reveals hidden variation in South Asian genomes.

bioRxiv. 2025 May 14:2025.05.14.653340. doi: 10.1101/2025.05.14.653340.

The effect of sample site and collection procedure on identification of SARS-CoV-2 infection.

Cochrane Database Syst Rev. 2024 Dec 16;12(12):CD014780. doi: 10.1002/14651858.CD014780.

Survivor, family and professional experiences of psychosocial interventions for sexual abuse and violence: a qualitative evidence synthesis.

Cochrane Database Syst Rev. 2022 Oct 4;10(10):CD013648. doi: 10.1002/14651858.CD013648.pub2.

Signs and symptoms to determine if a patient presenting in primary care or hospital outpatient settings has COVID-19.

Cochrane Database Syst Rev. 2022 May 20;5(5):CD013665. doi: 10.1002/14651858.CD013665.pub3.

Factors that influence parents' and informal caregivers' views and practices regarding routine childhood vaccination: a qualitative evidence synthesis.

Cochrane Database Syst Rev. 2021 Oct 27;10(10):CD013265. doi: 10.1002/14651858.CD013265.pub2.

Falls prevention interventions for community-dwelling older adults: systematic review and meta-analysis of benefits, harms, and patient values and preferences.

Syst Rev. 2024 Nov 26;13(1):289. doi: 10.1186/s13643-024-02681-3.

Cost-effectiveness of using prognostic information to select women with breast cancer for adjuvant systemic therapy.

Health Technol Assess. 2006 Sep;10(34):iii-iv, ix-xi, 1-204. doi: 10.3310/hta10340.

Antidepressants for pain management in adults with chronic pain: a network meta-analysis.

Health Technol Assess. 2024 Oct;28(62):1-155. doi: 10.3310/MKRT2948.

Surgical interventions for treating intracapsular hip fractures in older adults: a network meta-analysis.

Cochrane Database Syst Rev. 2022 Feb 14;2(2):CD013404. doi: 10.1002/14651858.CD013404.pub2.

Surgical interventions for treating extracapsular hip fractures in older adults: a network meta-analysis.

Cochrane Database Syst Rev. 2022 Feb 10;2(2):CD013405. doi: 10.1002/14651858.CD013405.pub2.

本文引用的文献

The distribution of highly deleterious variants across human ancestry groups.

Proc Natl Acad Sci U S A. 2025 May 27;122(21):e2503857122. doi: 10.1073/pnas.2503857122. Epub 2025 May 23.

Mapping genetic diversity with the GenomeIndia project.

Nat Genet. 2025 Apr;57(4):767-773. doi: 10.1038/s41588-025-02153-x.

High-coverage nanopore sequencing of samples from the 1000 Genomes Project to build a comprehensive catalog of human genetic variation.

Genome Res. 2024 Nov 20;34(11):2061-2073. doi: 10.1101/gr.279273.124.

Sources of gene expression variation in a globally diverse human cohort.

Nature. 2024 Aug;632(8023):122-130. doi: 10.1038/s41586-024-07708-2. Epub 2024 Jul 17.

Beyond the Human Genome Project: The Age of Complete Human Genome Sequences and Pangenome References.

Annu Rev Genomics Hum Genet. 2024 Aug;25(1):77-104. doi: 10.1146/annurev-genom-021623-081639. Epub 2024 Aug 6.

The frequency of pathogenic variation in the All of Us cohort reveals ancestry-driven disparities.

Commun Biol. 2024 Feb 19;7(1):174. doi: 10.1038/s42003-023-05708-y.

The complete sequence of a human Y chromosome.

Nature. 2023 Sep;621(7978):344-354. doi: 10.1038/s41586-023-06457-y. Epub 2023 Aug 23.

Addressing Ancestry and Sex Bias in Pharmacogenomics.

Annu Rev Pharmacol Toxicol. 2024 Jan 23;64:53-64. doi: 10.1146/annurev-pharmtox-030823-111731. Epub 2023 Jul 14.

A pangenome reference of 36 Chinese populations.

Nature. 2023 Jul;619(7968):112-121. doi: 10.1038/s41586-023-06173-7. Epub 2023 Jun 14.

A draft human pangenome reference.

Nature. 2023 May;617(7960):312-324. doi: 10.1038/s41586-023-05896-x. Epub 2023 May 10.

文献AI研究员

20分钟写一篇综述，助力文献阅读效率提升50倍。

立即体验

用中文搜PubMed

大模型驱动的PubMed中文搜索引擎

马上搜索

文档翻译

学术文献翻译模型，支持多种主流文档格式。

立即体验

组装未映射的 reads 揭示了南亚基因组中隐藏的变异。

Assembling unmapped reads reveals hidden variation in South Asian genomes.

作者信息

机构信息

出版信息

相似文献

本文引用的文献

文献AI研究员

用中文搜PubMed

文档翻译

Suppr 超能文献

相似文献

本文引用的文献