Analysis and Interpretation of Hidden Properties in Heterogeneous Biomedical data

Abstract

Modern medicine generates vast amounts of complex and interconnected patient data that often do not align with the assumptions of classical statistical analysis. The traditional approach, based on defining individual variables and examining their relationships, is gradually becoming exhausted, while machine learning and neural network methods are finding ever broader applications. Between these two approaches lies multivariate analysis and its important tool, the study of similarity between objects. In the biomedical context, similarity analysis is most often applied in two types of tasks: classification and clustering. Although these tasks are methodologically distinct, their objectives may overlap from the perspective of a domain expert (a clinician). It is precisely at this intersection that data analysis using similarity networks can be found. Patient Similarity Networks (PSN) combine dimensionality reduction with the analysis of similarity among individual records (patients). Their construction often incorporates information about data classes (e.g., phenotypes or disease severity levels). The aim is to identify clusters of patients with similar characteristics and interpret their key features. This dissertation provides a comprehensive overview of mathematical and computational tools used in PSN analysis, as well as related machine learning methods. The main contributions of the dissertation are as follows: 1. A review of machine learning methods applied to biomedical problems that can be addressed using PSN, and their systematic comparison with the PSN approach. 2. A review of mathematical tools involved in the processes of construction, evaluation, and interpretation of similarity networks. 3. The formulation of a complete methodological framework for the analysis of biomedical data using PSN, from data preprocessing to visualization of results. 4. Proposal of a method of application of the Matthews correlation coefficient (MCC) that enables its use in multi-class multi-cluster problems. 5. The introduction of a novel method for PSN evaluation, cPurity, which complements the information provided by MCC. 6. Implementation of the LRNet algorithm in Python, extended with visualization and evaluation tools. 7. Implementation of a bioinspired extension of the LRNet algorithm for more efficient selection of candidate networks during analysis.

Description

Delayed publication

Available after

Subject(s)

patient similarity network, PSN, similarity, network analysis, similarity network evaluation, visualization, LRNet, rMCC, connection purity, cPurity, augPSN

Citation