Here are instructions for this version of ISA and memISA. If you have any queries, send me a mail at richardsal1@cf.ac.uk.

1) Take your input file - an expression matrix with genes as rows and columns as samples, presumably unnormalised - and remove any gene or sample names. Keep them available to put back on later!
2) Read your input file into R, into the variable 'Troy'
3) Run ISANormalisationRealMissingB.r on it, changing the output filenames to something suitable (first half of script makes gene-normalised dataset, second half of script makes sample-normalised dataset).
4) Put names back onto gene- and sample- normalised datasets

5) Write the parameter file: first line is ID number of first iteration in output, second and third lines are both number of iterations (10000 or 20000 is typical, will take months on single processor though), subsequent lines are tC and tG parameter pairs (number of iterations on second/third line will be performed for EVERY tC/tG pair)
6) Run perl memISAc.pl GeneNormalised.tsv CondNormalised.tsv ParameterFile.txt (or whatever your files are called) 0.7 5 .
This runs memISA, with f=0.7 (degree of bias against already found clusters) and n=5 (number of rounds of clustering biases are remembered for).
To run standard ISA (doesn't bias against already found clusters), use: memISAc.pl GeneNormalised.tsv CondNormalised.tsv ParameterFile.txt 0 1

This may need to be run in many stages or across many computers - output files with the same parameters should be concatenated after this step. log.txt is the log file that lets you track it. Start.txt are the starting random samples for each iteration (expressed as row numbers in the dataset - currently turned off to reduce size of output).

If you need to pause a run, write a blank file with the name STOP into the directory - the program will then stop at the end of the next iteration, and write out a new parameter file that can be used to restart the process at a later date. One slight downside is that the current biases used by memISA are lost, but the effects of this should be minor unless n is very large.

It is highly recommended that you parallelise this process if you intend to run it on large datasets.
 
7) Use Uniqueify.pl, NormaliseScoresv0.2.exe and DeredundifyBv0.2.exe on the output files, in order.

Example: 
perl Uniqueify.pl OUTPUTtC(0.5)tG(1.0).txt UNIQUEtC(0.5)tG(1.0).txt
perl NormaliseScoresv0.2.pl UNIQUEtC(0.5)tG(1.0).txt NORMtC(0.5)tG(1.0).txt
perl DeredundifyBv0.2.pl NORMtC(0.5)tG(1.0).txt DEREDtC(0.5)tG(1.0).txt

Note : At this stage, if you name your output files in the form DEREDtC(x.x)tG(x.x).txt (where x.x are the tC and tG values), the program CombinerDEREDd.pl will work and run steps 8-10 for you using the standard settings

8) Run RemoveSmallAndFew.pl on files - standard values are minimum number of genes 40 and at least 3 occurrences (works well at 20000 iterations)

Example:
perl RemoveSmallAndFew.pl DEREDtC(0.5)tG(1.0).txt REMtC(0.5)tG(1.0).txt 40 3

9) Run FindClusterOverlapC.pl on files - usual gene overlap is 0.65, sample / condition overlap is 0.55, score correlation is 0.65.

Example:
perl FindClusterOverlapC.pl REMtC(0.5)tG(1.0).txt REMtC(0.5)tG(1.0).txt COMPtC(0.5)tG(1.0).txt OVTABLEtC(0.5)tG(1.0).txt 1 0.65 0.55 0.65

First two arguments are the file to be compared, third argument is output file. 
Fourth argument is output filename of table giving percentage overlaps (gene and sample) and gene/sample score correlation between clusters. 
Fifth argument is 0 if smaller cluster size is to be used as divisor when calculating overlap (giving subsets of clusters 100% overlap with their parent cluster) and 1 if larger cluster sizes is to be used (using 0 will drastically increase the amount of overlap between different-sized clusters). 
Sixth argument is percentage gene overlap (0.65 = 65%) - the proportion of genes the clusters must share to be merged. 
Seventh argument is percentage condition / sample overlap - the proportion of conditions / samples the clusters must share to be merged. 
Eighth argument is correlation of both gene and condition/sample scores needed for clusters to be merged.

10) Run TabGCToNewline.pl and NormaliseScoresv0.2.pl on resulting files. (First script is a minor formatting script, second one is to normalise the gene and sample scores).

Example:
perl TabGCToNewline.pl COMPtC(0.5)tG(1.0).txt COMP2(0.5)tG(1.0).txt
perl NormaliseScoresv0.2.pl COMP2(0.5)tG(1.0).txt NORM4(0.5)tG(1.0).txt

11) You've now got a cluster set for each pair of tC and tG values. To combine them into a single representative cluster set, begin by concatenating them into a single file.

Example:
cat NORM4* >> AllClustersMixed

12) Now use FindClusterOverlapC.pl to merge any similar clusters across parameters. Usually only reasonably small clusters (those produced by tG 2.1 or more on the Dobrin dataset, for example - clusters with less than a thousand genes) are used, as otherwise this merging step will merge everything into a very small number of very large clusters.

Example:
perl FindClusterOverlapC.pl AllClustersMixed AllClustersMixed AllClustersMerged AllClustersOv 1 0.6 0.5 0.6

13) Now use TabGCToNewline.pl to reformat the merged file (replace the tabs with newlines)

Example:
perl TabGCToNewline.pl AllClustersMerged AllClustersNewline

14) Now use OverlapRemover.pl to remove clusters that have a high amount of gene overlap with other clusters but were otherwise not similar enough to be merged (removes the smaller of the two overlapping clusters). Drastically reduces the number of clusters, but appears to leave behind the 'good' clusters (at least, those with better GO enrichments).

Example:
perl OverlapRemover.pl AllClustersNewline AllClustersDeol 70

First argument is input, second argument is output. Third argument is percentage of genes from smaller cluster needed to be present in bigger cluster for the smaller cluster to be removed.

15) Run NamerDeol.pl on the resulting file with the name files to give a final, named set of clusters.

perl NamerDeol.pl AllClustersDeol GeneNames.txt SampleNames.txt AllClustersNamed

First argument is the input file. Second argument is a list of gene names, taken from the original dataset. Third argument is a list of sample / condition names, also taken from original dataset. Fourth argument is output file.