The Regulatory Repeat Database is a database of repeated sequences found at the binding sites of various Transcription Factors (TFs) and cofactors.
Two types of sequences are in R2-DB: short tracks (ST) of a few dozen bps that are strongly enriched for specific nucleotides and k-mers, and long tracks (LT) that are much longer (several hundred bps) but less biased.
R2-DB can be queried by TF/cofactor, cell type and/or sHMM library.
Sequences in R2-DB have been identified with supervised Hidden Markov Models (sHMMs) trained to discriminate bound vs. unbound sequences in ChIP-seq data.
Specifically, two sHMM libraries allow to identify the STs and LTs associated with peaks of 934 ChIP-seq experiments from 4 cell types (K562, GM12878, HepG2 and HEK293).
R2-DB is developed and maintained by the ATGC platform.
Characterization of repeated sequences around transcription factor binding sites with supervised hidden Markov models
Christophe Menichelli, Oceane Cassan, Sophie Lebre, Charles-Henri Lecellier, Laurent Brehelin.
In preparation (2026).
For scientific purposes: christophe.menichelli@lirmm.fr , laurent.brehelin@lirmm.fr
For technical issues: atgc-contact@lirmm.fr