Performance graphs with differently sized subsets of GSE6754

Graphs present the dependency of false positive counts given the number of selected best candidate interactions. Curves closer to lower-right corner of the graph indicate better performance. The axes are in logarithmic scale to emphasize the results for smaller numbers of best candidates.

We used the GSE6754 data set, which describes families with two individuals affected by autism spectrum disorders. Individuals were classified as affected (2459 samples) or unaffected (3473 samples) and described with around 10,000 SNPs each. Only the first 2,000 SNPs were used for the analysis.

Curve legend

light gray - theoretically best and worst possible performance curves
black solid - direct scoring
black dashed - scoring with two replication groups
black dotted - scoring with three replication groups

data set size = 100 samples

data set size = 200 samples

data set size = 500 samples

data set size = 1000 samples

data set size = 2000 samples

data set size = 5000 samples