Semi-artificial datasets as a resource for validation of bioinformatics pipelines for plant virus detection - INRAE - Institut national de recherche pour l’agriculture, l’alimentation et l’environnement Accéder directement au contenu
Article Dans Une Revue Peer Community Journal Année : 2021

Semi-artificial datasets as a resource for validation of bioinformatics pipelines for plant virus detection

Lucie Tamisier
Annelies Haegeman
  • Fonction : Auteur
Yoika Foucart
  • Fonction : Auteur
Nicolas Fouillien
  • Fonction : Auteur
Maher Al Rwahnih
  • Fonction : Auteur
Nihal Buzkan
  • Fonction : Auteur
Thierry Candresse
  • Fonction : Auteur
Michela Chiumenti
  • Fonction : Auteur
Kris de Jonghe
  • Fonction : Auteur
Marie Lefebvre
  • Fonction : Auteur
Paolo Margaria
  • Fonction : Auteur
Jean Sébastien Reynard
  • Fonction : Auteur
Kristian Stevens
  • Fonction : Auteur
Denis Kutnjak
  • Fonction : Auteur
Sébastien Massart
  • Fonction : Auteur

Résumé

The widespread use of High-Throughput Sequencing (HTS) for detection of plant viruses and sequencing of plant virus genomes has led to the generation of large amounts of data and of bioinformatics challenges to process them. Many bioinformatics pipelines for virus detection are available, making the choice of a suitable one difficult. A robust benchmarking is needed for the unbiased comparison of the pipelines, but there is currently a lack of reference datasets that could be used for this purpose. We present 7 semi-artificial datasets composed of real RNA-seq datasets from virus-infected plants spiked with artificial virus reads. Each dataset addresses challenges that could prevent virus detection. We also present 3 real datasets showing a challenging virus composition as well as 8 completely artificial datasets to test haplotype reconstruction software. With these datasets that address several diagnostic challenges, we hope to encourage virologists, diagnosticians and bioinformaticians to evaluate and benchmark their pipeline(s).

Dates et versions

hal-04107125 , version 1 (26-05-2023)

Identifiants

Citer

Lucie Tamisier, Annelies Haegeman, Yoika Foucart, Nicolas Fouillien, Maher Al Rwahnih, et al.. Semi-artificial datasets as a resource for validation of bioinformatics pipelines for plant virus detection. Peer Community Journal, 2021, 1, pp.e53. ⟨10.24072/pcjournal.62⟩. ⟨hal-04107125⟩
10 Consultations
0 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More