PMCCPMCCPMCC

Search tips
Search criteria 

Advanced

 
Logo of bmcbioiBioMed Centralsearchsubmit a manuscriptregisterthis articleBMC Bioinformatics
 
BMC Bioinformatics. 2006; 7: 541.
Published online 2006 December 22. doi:  10.1186/1471-2105-7-541
PMCID: PMC1769404
How repetitive are genomes?
Bernhard Hauboldcorresponding author1 and Thomas Wiehe2
1Department of Biotechnology & Bioinformatics, University of Applied Sciences Weihenstephan, Freising, Germany
2Institute of Genetics, Universität zu Köln, Cologne, Germany
corresponding authorCorresponding author.
Bernhard Haubold: bernhard.haubold/at/fh-weihenstephan.de; Thomas Wiehe: twiehe/at/uni-koeln.de
Received October 26, 2006; Accepted December 22, 2006.
Abstract
Background
Genome sequences vary strongly in their repetitiveness and the causes for this are still debated. Here we propose a novel measure of genome repetitiveness, the index of repetitiveness, Ir, which can be computed in time proportional to the length of the sequences analyzed. We apply it to 336 genomes from all three domains of life.
Results
The expected value of Ir is zero for random sequences of any G/C content and greater than zero for sequences with excess repeats. We find that the Ir of archaea is significantly smaller than that of eubacteria, which in turn is smaller than that of eukaryotes. Mouse chromosomes have a significantly higher Ir than human chromosomes and within each genome the Y chromosome is most repetitive. A sliding window analysis reveals that the human HOXA cluster and two surrounding genes are characterized by local minima in Ir. A program for calculating the Ir is freely available at http://adenine.biz.fh-weihenstephan.de/ir/.
Conclusion
The general measure of DNA repetitiveness proposed in this paper can be efficiently computed on a genomic scale. This reveals a broad spectrum of repetitiveness among diverse genomes which agrees qualitatively with previous studies of repeat content. A sliding window analysis helps to analyze the intragenomic distribution of repeats.
Articles from BMC Bioinformatics are provided here courtesy of
BioMed Central