R2-DB

Regulatory Repeat Database

Repeated sequence tracks associated with transcription factor and cofactor binding sites.

The Regulatory Repeat Database is a database of repeated sequences found at the binding sites of various Transcription Factors (TFs) and cofactors.

Two types of sequences are in R2-DB: short tracks (ST) of a few dozen bps that are strongly enriched for specific nucleotides and k-mers, and long tracks (LT) that are much longer (several hundred bps) but less biased.

Query R2-DB

R2-DB can be queried by TF/cofactor, cell type and/or sHMM library.

TF/cofactor Cell type sHMM library

Supervised HMM libraries

Sequences in R2-DB have been identified with supervised Hidden Markov Models (sHMMs) trained to discriminate bound vs. unbound sequences in ChIP-seq data.

Specifically, two sHMM libraries allow to identify the STs and LTs associated with peaks of 934 ChIP-seq experiments from 4 cell types (K562, GM12878, HepG2 and HEK293).

Two sHMM structures: a three-state C/G track model and a dinucleotide TG track model.
Figure 1. A. Three-state sHMM modelling C/G sequence tracks; the red state models the C/G track, while green and blue states model the surrounding 5' and 3' sequences. B. sHMM modelling sequence tracks enriched for the dinucleotide TG; the gap state models the nucleotide distribution between two TG occurrences.

About this resource

R2-DB is developed and maintained by the ATGC platform.

Citation

Characterization of repeated sequences around transcription factor binding sites with supervised hidden Markov models

Christophe Menichelli, Oceane Cassan, Sophie Lebre, Charles-Henri Lecellier, Laurent Brehelin.

In preparation (2026).

Contact

For scientific purposes: christophe.menichelli@lirmm.fr , laurent.brehelin@lirmm.fr

For technical issues: atgc-contact@lirmm.fr

Protein name
Cell Type
HMM ID
ENCODE ID
N/A
AUC
The figure reports the JASPAR PWM associated with the TF (the version with best accuracy if several versions are available) and different plots related to the sHMM and the identified tracks.
Fig. A ? When the PWM is known, the A plot reports the AUROC achieved by the PWM and the sHMM, followed by the sHMM AUROC when removing the PWM effect (adjusted-AUROC, HMM/P), the adjusted-AUROC when the PWM score is low (Q1) and the adjusted-AUROC when the PWM score is high (Q2).
Fig. B ? Plot B reports the nucleotidic distribution of the state modelling the tracks (for three-states sHMMs), or of the gap-state (for k-mer sHMMs).
Fig. C ? Plot C is the sHMM ROC.
Fig. D ? Plot D reports the length distribution of the identified tracks on positive and negative sequences.
Fig. E-F ? Plot E-F report the distribution of the number of k-mer occurrences and gap proportion for k-mer sHMMs.
Fig. G ? Plot G reports the coverage of the identified tracks on positive and negative sequences, i.e., for each nucleotide position, the proportion of sequences covered by the track at this position. This plot also shows the expected coverages that would be obtained if the identified tracks were uniformly distributed on the sequences.
Fig. H ? Plot H is a representation of the different tracks identified on the sequences. Tracks closer to the 5' end (resp. 3' end) are on the upper (resp. lower) part of the plot. For each set (5' or 3') tracks are drawn from the longest to the smallest, starting from the median line toward the top (5') or the bottom (3') of the plot.
Fig. I ? Plot I shows the AUROC of the selected sHMMs along those of the 10 most common sHMMs of the library. For ST library, sHMMs with a positional bias toward the center of the peaks appear in black.
Tip: use the arrow keys to move between experiments.