Metadata integrity in bioinformatics: Bridging the gap between data and knowledge

Caliskan, Aylin; Dangwal, Seema; Dandekar, Thomas

doi:10.1016/j.csbj.2023.10.006

Cited by 1 publication

(2 citation statements)

References 100 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…This step, as well as the additional analysis of another suitable data set, are a cost- and time-efficient preliminary step before validating the results in vitro and in vivo . Nevertheless, when analyzing big data in biomedical research, the data quality and data integrity, as well as the comprehensiveness of the metadata, can hugely affect the resulting analyses, which needs to be considered when reusing a data set or publishing a data set [ 69 ].…”

Section: Discussionmentioning

confidence: 99%

“…In our showcase, we will analyze single-cell data provided by Sole ´-Boldo et al (2020) [6] (the RDS file is available via the GEO database: https://www.ncbi.nlm.nih.gov/geo/query/acc. cgi?acc=GSE130973), who analyzed human skin fibroblasts from a sun-protected area of healthy 'young' donors (25 and 27 years old) and healthy 'old' donors (53,69, and 70 years old) using single-cell RNA sequencing and found age-related changes in fibroblast subpopulations [6], using the conditions 'young' and 'old'.…”

Section: Introductionmentioning

confidence: 99%

See 1 more Smart Citation

An orchestra of machine learning methods reveals landmarks in single-cell data exemplified with aging fibroblasts

Rasbach,

Caliskan,

Saderi

et al. 2024

PLoS ONE

Self Cite

View full text Add to dashboard Cite

In this work, a Python framework for characteristic feature extraction is developed and applied to gene expression data of human fibroblasts. Unlabeled feature selection objectively determines groups and minimal gene sets separating groups. ML explainability methods transform the features correlating with phenotypic differences into causal reasoning, supported by further pipeline and visualization tools, allowing user knowledge to boost causal reasoning. The purpose of the framework is to identify characteristic features that are causally related to phenotypic differences of single cells. The pipeline consists of several data science methods enriched with purposeful visualization of the intermediate results in order to check them systematically and infuse the domain knowledge about the investigated process. A specific focus is to extract a small but meaningful set of genes to facilitate causal reasoning for the phenotypic differences. One application could be drug target identification. For this purpose, the framework follows different steps: feature reduction (PFA), low dimensional embedding (UMAP), clustering ((H)DBSCAN), feature correlation (chi-square, mutual information), ML validation and explainability (SHAP, tree explainer). The pipeline is validated by identifying and correctly separating signature genes associated with aging in fibroblasts from single-cell gene expression measurements: PLK3, polo-like protein kinase 3; CCDC88A, Coiled-Coil Domain Containing 88A; STAT3, signal transducer and activator of transcription-3; ZNF7, Zinc Finger Protein 7; SLC24A2, solute carrier family 24 member 2 and lncRNA RP11-372K14.2. The code for the preprocessing step can be found in the GitHub repository https://github.com/AC-PHD/NoLabelPFA, along with the characteristic feature extraction https://github.com/LauritzR/characteristic-feature-extraction.

show abstract

Section: Discussionmentioning

confidence: 99%