Optimizing data integration improves Gene Regulatory Network inference in Arabidopsis thaliana - INRAE - Institut national de recherche pour l’agriculture, l’alimentation et l’environnement Access content directly
Preprints, Working Papers, ... Year : 2023

Optimizing data integration improves Gene Regulatory Network inference in Arabidopsis thaliana

Abstract

Motivations: Gene Regulatory Networks (GRN) are traditionnally inferred from gene expression profiles monitoring a specific condition or treatment. In the last decade, integrative strategies have successfully emerged to guide GRN inference from gene expression with complementary prior data. However, datasets used as prior information and validation gold standards are often related and limited to a subset of genes. This lack of complete and independent evaluation calls for new criteria to robustly estimate the optimal intensity of prior data integration in the inference process. Results: We address this issue for two common regression-based GRN inference models, an integrative Random Forest (weigthedRF) and a generalized linear model with stability selection estimated under a weighted LASSO penalty (weightedLASSO). These approaches are applied to data from the root response to nitrate induction in Arabidopsis thaliana. For each gene, we measure how the integration of transcription factor binding motifs influences model prediction. We propose a new approach, DIOgene, that uses model prediction error and a simulated null hypothesis for optimizing data integration strength in a hypothesisdriven, gene-specific manner. The resulting integration scheme reveals a strong diversity of optimal integration intensities between genes. In addition, it provides a good trade-off between prediction error minimization and validation on experimental interactions, while master regulators of nitrate induction can be accurately retrieved. Availability and implementation The R code and notebooks demonstrating the use of the proposed approaches are available in the repository https://github.com/OceaneCsn/integrative_GRN_N_ induction.
Fichier principal
Vignette du fichier
CassanO.-et al-bioRxiv2-2023.pdf (5.46 Mo) Télécharger le fichier
Origin Files produced by the author(s)
Licence

Dates and versions

hal-04228523 , version 1 (04-10-2023)

Licence

Identifiers

Cite

Océane Cassan, Charles Henri Lecellier, Antoine Martin, Laurent Brehelin, Sophie Lèbre. Optimizing data integration improves Gene Regulatory Network inference in Arabidopsis thaliana. 2023. ⟨hal-04228523⟩
67 View
11 Download

Altmetric

Share

Gmail Mastodon Facebook X LinkedIn More