A Benchmark of Pre-processing Effect on Single Cell RNA Sequencing Integration Methods

Anaissi, Ali; Zandavi, Seid Miad; Suleiman, Basem; Alyassine, Widad; Braytee, Ali; Vafaee, Fatemeh

doi:10.21203/rs.3.rs-2249309/v1

Cited by 1 publication

(1 citation statement)

References 29 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…An optimality criterion is tested at each iteration to assess whether batch mixing is sufficient, using a local purity metric called Local Inverse Simpson's Index (LISI). Due to its simplicity and availability with both Python and R packages, Harmony is widely used today and still achieves respectable results in benchmarks (Anaissi et al, 2022) despite being limited when facing strong batch effects (Luecken et al, 2022).…”

Section: Horizontal Integration (Hi) Links Batches Anchored By Their ...mentioning

confidence: 99%

Omics data integration in computational biology viewed through the prism of machine learning paradigms

Fouché

Zinovyev²

2023

Front. Bioinform.

View full text Add to dashboard Cite

Important quantities of biological data can today be acquired to characterize cell types and states, from various sources and using a wide diversity of methods, providing scientists with more and more information to answer challenging biological questions. Unfortunately, working with this amount of data comes at the price of ever-increasing data complexity. This is caused by the multiplication of data types and batch effects, which hinders the joint usage of all available data within common analyses. Data integration describes a set of tasks geared towards embedding several datasets of different origins or modalities into a joint representation that can then be used to carry out downstream analyses. In the last decade, dozens of methods have been proposed to tackle the different facets of the data integration problem, relying on various paradigms. This review introduces the most common data types encountered in computational biology and provides systematic definitions of the data integration problems. We then present how machine learning innovations were leveraged to build effective data integration algorithms, that are widely used today by computational biologists. We discuss the current state of data integration and important pitfalls to consider when working with data integration tools. We eventually detail a set of challenges the field will have to overcome in the coming years.

show abstract

Section: Horizontal Integration (Hi) Links Batches Anchored By Their ...mentioning

confidence: 99%

Omics data integration in computational biology viewed through the prism of machine learning paradigms

Fouché

Zinovyev²

2023

Front. Bioinform.

View full text Add to dashboard Cite

show abstract

A Benchmark of Pre-processing Effect on Single Cell RNA Sequencing Integration Methods

Cited by 1 publication

References 29 publications

Omics data integration in computational biology viewed through the prism of machine learning paradigms

Omics data integration in computational biology viewed through the prism of machine learning paradigms

Contact Info

Product

Resources

About