UMI-tools：对独特分子标识符中的测序错误进行建模以提高定量准确性。

UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy.

作者信息

Smith Tom, Heger Andreas, Sudbery Ian

机构信息

Computational Genomics Analysis and Training Programme, MRC WIMM Centre for Computational Biology, University of Oxford, Oxford OX3 9DS, United Kingdom.

Department of Molecular Biology and Biotechnology, University of Sheffield, Sheffield S10 2TN, United Kingdom.

出版信息

Genome Res. 2017 Mar;27(3):491-499. doi: 10.1101/gr.209601.116. Epub 2017 Jan 18.

DOI:10.1101/gr.209601.116

PMID:28100584

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC5340976/

Abstract

Unique Molecular Identifiers (UMIs) are random oligonucleotide barcodes that are increasingly used in high-throughput sequencing experiments. Through a UMI, identical copies arising from distinct molecules can be distinguished from those arising through PCR amplification of the same molecule. However, bioinformatic methods to leverage the information from UMIs have yet to be formalized. In particular, sequencing errors in the UMI sequence are often ignored or else resolved in an ad hoc manner. We show that errors in the UMI sequence are common and introduce network-based methods to account for these errors when identifying PCR duplicates. Using these methods, we demonstrate improved quantification accuracy both under simulated conditions and real iCLIP and single-cell RNA-seq data sets. Reproducibility between iCLIP replicates and single-cell RNA-seq clustering are both improved using our proposed network-based method, demonstrating the value of properly accounting for errors in UMIs. These methods are implemented in the open source UMI-tools software package.

摘要

独特分子标识符（UMIs）是随机寡核苷酸条形码，在高通量测序实验中越来越常用。通过UMI，可以将源自不同分子的相同拷贝与通过同一分子的PCR扩增产生的拷贝区分开来。然而，利用UMI信息的生物信息学方法尚未正式确立。特别是，UMI序列中的测序错误常常被忽略，或者以临时的方式解决。我们表明，UMI序列中的错误很常见，并引入了基于网络的方法，在识别PCR重复序列时考虑这些错误。使用这些方法，我们在模拟条件以及真实的iCLIP和单细胞RNA测序数据集下均展示了提高的定量准确性。使用我们提出的基于网络的方法，iCLIP重复样本之间的可重复性和单细胞RNA测序聚类均得到改善，证明了正确考虑UMI错误的价值。这些方法在开源的UMI-tools软件包中实现。

https://cdn.ncbi.nlm.nih.gov/pmc/blobs/7c6a/5340976/7003dfad2035/491f01.jpg

相似文献

UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy.

Genome Res. 2017 Mar;27(3):491-499. doi: 10.1101/gr.209601.116. Epub 2017 Jan 18.

Je, a versatile suite to handle multiplexed NGS libraries with unique molecular identifiers.

BMC Bioinformatics. 2016 Oct 8;17(1):419. doi: 10.1186/s12859-016-1284-2.

Accurate estimation of molecular counts from amplicon sequence data with unique molecular identifiers.

Bioinformatics. 2023 Jan 1;39(1). doi: 10.1093/bioinformatics/btad002.

Alignment-free clustering of UMI tagged DNA molecules.

Bioinformatics. 2019 Jun 1;35(11):1829-1836. doi: 10.1093/bioinformatics/bty888.

ATAC-seq with unique molecular identifiers improves quantification and footprinting.

Commun Biol. 2020 Nov 13;3(1):675. doi: 10.1038/s42003-020-01403-4.

AmpUMI: design and analysis of unique molecular identifiers for deep amplicon sequencing.

Bioinformatics. 2018 Jul 1;34(13):i202-i210. doi: 10.1093/bioinformatics/bty264.

Investigation into the genotyping performance of a unique molecular identifier based microhaplotypes MPS panel in complex DNA mixture.

Forensic Sci Int Genet. 2025 Mar;76:103236. doi: 10.1016/j.fsigen.2025.103236. Epub 2025 Feb 5.

Reducing noise and stutter in short tandem repeat loci with unique molecular identifiers.

Forensic Sci Int Genet. 2021 Mar;51:102459. doi: 10.1016/j.fsigen.2020.102459. Epub 2020 Dec 25.

Ultrasensitive sequencing of STR markers utilizing unique molecular identifiers and the SiMSen-Seq method.

Forensic Sci Int Genet. 2024 Jul;71:103047. doi: 10.1016/j.fsigen.2024.103047. Epub 2024 Apr 3.

Incorporation of unique molecular identifiers in TruSeq adapters improves the accuracy of quantitative sequencing.

Biotechniques. 2017 Nov 1;63(5):221-226. doi: 10.2144/000114608.

引用本文的文献

An ancient and essential miRNA family controls cellular interaction pathways in .

Sci Adv. 2025 Sep 5;11(36):eadz1934. doi: 10.1126/sciadv.adz1934. Epub 2025 Sep 3.

DNA methylation insulates genic regions from CTCF loops near nuclear speckles.

Elife. 2025 Sep 3;13:RP102930. doi: 10.7554/eLife.102930.

Therapeutic potential of Desmodium styracifolium polysaccharide in attenuating nano-calcium oxalate induced renal injury and fibrosis.

Commun Biol. 2025 Sep 2;8(1):1330. doi: 10.1038/s42003-025-08757-7.

Food hydrocolloids κ-carrageenan and xanthan gum in processed red meat modify gut health in rats.

Curr Res Food Sci. 2025 Aug 6;11:101162. doi: 10.1016/j.crfs.2025.101162. eCollection 2025.

PYM1 limits non-canonical Exon Junction Complex occupancy in a gene architecture dependent manner to tune mRNA expression.

Nat Commun. 2025 Aug 30;16(1):8138. doi: 10.1038/s41467-025-63455-6.

Skin metatranscriptomics reveals a landscape of variation in microbial activity and gene expression across the human body.

Nat Biotechnol. 2025 Aug 28. doi: 10.1038/s41587-025-02797-4.

Urinary microRNAs as Prognostic Biomarkers for Predicting the Efficacy of Immune Checkpoint Inhibitors in Patients with Urothelial Carcinoma.

Cancers (Basel). 2025 Aug 13;17(16):2640. doi: 10.3390/cancers17162640.

Composition and RNA binding specificity of metazoan RNase MRP.

Nucleic Acids Res. 2025 Aug 27;53(16). doi: 10.1093/nar/gkaf829.

Spatial host-microbiome profiling demonstrates bacterial-associated host transcriptional alterations in pediatric ileal Crohn's disease.

Microbiome. 2025 Aug 23;13(1):189. doi: 10.1186/s40168-025-02178-8.

Unveiling the Future of Infective Endocarditis Diagnosis: The Transformative Role of Metagenomic Next-Generation Sequencing in Culture-Negative Cases.

J Epidemiol Glob Health. 2025 Aug 22;15(1):108. doi: 10.1007/s44197-025-00455-1.

本文引用的文献

Enhancer-promoter interactions are encoded by complex genomic signatures on looping chromatin.

Nat Genet. 2016 May;48(5):488-96. doi: 10.1038/ng.3539. Epub 2016 Apr 4.

Illumina error profiles: resolving fine-scale variation in metagenomic sequencing data.

BMC Bioinformatics. 2016 Mar 11;17:125. doi: 10.1186/s12859-016-0976-y.

SR proteins are NXF1 adaptors that link alternative RNA processing to mRNA export.

Genes Dev. 2016 Mar 1;30(5):553-66. doi: 10.1101/gad.276477.115.

Practical guidelines for B-cell receptor repertoire sequencing analysis.

Genome Med. 2015 Nov 20;7:121. doi: 10.1186/s13073-015-0243-2.

High-throughput and quantitative genome-wide messenger RNA sequencing for molecular phenotyping.

BMC Genomics. 2015 Aug 5;16(1):578. doi: 10.1186/s12864-015-1788-6.

Scalable microfluidics for single-cell RNA printing and sequencing.

Genome Biol. 2015 Jun 6;16(1):120. doi: 10.1186/s13059-015-0684-3.

Highly Parallel Genome-wide Expression Profiling of Individual Cells Using Nanoliter Droplets.

Cell. 2015 May 21;161(5):1202-1214. doi: 10.1016/j.cell.2015.05.002.

Droplet barcoding for single-cell transcriptomics applied to embryonic stem cells.

Cell. 2015 May 21;161(5):1187-1201. doi: 10.1016/j.cell.2015.04.044.

A general method to eliminate laboratory induced recombinants during massive, parallel sequencing of cDNA library.

Virol J. 2015 Apr 9;12:55. doi: 10.1186/s12985-015-0280-x.

ChIP-nexus enables improved detection of in vivo transcription factor binding footprints.

Nat Biotechnol. 2015 Apr;33(4):395-401. doi: 10.1038/nbt.3121. Epub 2015 Mar 9.

文献AI研究员

20分钟写一篇综述，助力文献阅读效率提升50倍。

立即体验

用中文搜PubMed

大模型驱动的PubMed中文搜索引擎

马上搜索

文档翻译

学术文献翻译模型，支持多种主流文档格式。

立即体验

UMI-tools：对独特分子标识符中的测序错误进行建模以提高定量准确性。

UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy.

作者信息

机构信息

出版信息

相似文献

引用本文的文献

本文引用的文献

文献AI研究员

用中文搜PubMed

文档翻译

Suppr 超能文献

相似文献

引用本文的文献

本文引用的文献