This directory contains 150 alignments simulated using the program DAWG (Cartwright, 2005), used for the analysis by Herman et al. (2015).

Ten sequences were simulated on the tree `simulated_tree.nwk', under the GTR substitution model (rates of substitution for AC, AG, AT, CG, CT and GT were set at 1.5, 3.0, 0.9, 1.2, 2.5 and 1.0; equilibrium frequencies for A, C, G and T were set at 0.20, 0.30, 0.30, 0.20), with the G+I rate heterogeneity model ($\gamma = 0.9$, $\iota = 0.05$), and an indel process with lengths distributed according to a negative binomial $NB(3, 0.7)$ distribution. 

The indel rate was set to three different values [$0.01$ (low), $0.02$ (medium) and $0.03$ (high)], to generate datasets of varying alignment uncertainty. For each indel rate, $50$ alignments were generated, yielding $150$ datasets overall. 

StatAlign v1.1 was run using the default settings for nucleotides, with a burnin of $500,000$, and $2$ million sampling steps, taking alignment samples every $2000$ steps, thus producing $1000$ alignment samples for each test case. 

The program WeaveAlign was then used to compute minimum-risk summary alignments for each set of samples (see http://statalign.github.io/WeaveAlign for examples of how to run WeaveAlign).

References
^^^^^^^^^^
Cartwright RA (2005) "DNA assembly with gaps (DAWG): Simulating sequence evolution." Bioinformatics, 21(Suppl 3):31–38

Herman JL, Novák A, Lyngsø R, Szabó A, Miklós I and Hein J (2015) "Efficient representation of uncertainty in multiple sequence alignments using directed acyclic graphs." BMC Bioinformatics (to appear)