Poisson approximation for the number of repeats in a stationary Markov chain - INRAE - Institut national de recherche pour l’agriculture, l’alimentation et l’environnement Access content directly
Journal Articles Journal of Applied Probability Year : 2008

Poisson approximation for the number of repeats in a stationary Markov chain

Abstract

Detection of repeated sequences within complete genomes is a powerful tool to help understanding genome dynamics and species evolutionary history. To distinguish significant repeats from those that can be obtained just by chance, statistical methods have to be developed. In this paper we show that the distribution of the number of long repeats in long sequences generated by stationary Markov chains can be approximated by a Poisson distribution with explicit parameter. Thanks to the Chen-Stein method we provide a bound for the approximation error; this bound converges to 0 as soon as the length n of the sequence tends to ∞ and the length t of the repeats satisfies n2ρt = O(1) for some 0 < ρ < 1. Using this Poisson approximation, p-values can then be easily calculated to determine if a given genome is significantly enriched in repeats of length t.

Dates and versions

hal-02663042 , version 1 (31-05-2020)

Identifiers

Cite

Narjiss Touyar, Sophie S. Schbath, Dominique Cellier, Hélène Dauchel. Poisson approximation for the number of repeats in a stationary Markov chain. Journal of Applied Probability, 2008, 45 (2), pp.440-455. ⟨10.1239/jap/1214950359⟩. ⟨hal-02663042⟩
12 View
0 Download

Altmetric

Share

Gmail Mastodon Facebook X LinkedIn More