Human Splice DB Workflow

Galaxy Workflow ' Human Splice DB'

Annotation: Create a peptide fasta database with novel splice junctions that are inferred from RNAseq data

StepAnnotation
Step 1: Input dataset
select at runtime
RNA-Seq left mate pair fastq (These should be in fastqsanger format. If not, convert with "Fastq Groomer" tool.)
Step 2: Input dataset
select at runtime
RNA-Seq right mate pair fastq (These should be in fastqsanger format. If not, convert with "Fastq Groomer" tool.)
Step 3: Input dataset
select at runtime
GRCh37_canon.fa Contains only sequences from canonical chromosomes (chr1-22, X, Y, M)
Step 4: Select first
100000
Output dataset 'output' from step 1
Limit the sequence count for demonstration and testing purposes (100000)
Step 5: Select first
100000
Output dataset 'output' from step 2
Limit the sequence count for demonstration and testing purposes (100000)
Step 6: Tophat for Illumina
Output dataset 'out_file1' from step 4
Use a genome from history
Output dataset 'output' from step 3
Paired-end
Output dataset 'out_file1' from step 5
150
Full parameter list
FR Unstranded
20
5
0
70
500000
Yes
3
3
20
50
500000
2
2
25
Yes
Yes
select at runtime
No
No
No
No
No
GTF-guided Tophat alignment. Allow for detection of splice junctions not in the GTF file.
Step 7: Tophat for Illumina
Output dataset 'out_file1' from step 4
Use a genome from history
Output dataset 'output' from step 3
Paired-end
Output dataset 'out_file1' from step 5
150
Full parameter list
FR Unstranded
20
5
0
70
500000
Yes
3
3
20
50
500000
2
2
25
Yes
Yes
select at runtime
No
Yes
No
No
No
GTF-guided Tophat alignment. Do not allow for detection of splice junctions absent from the GTF file.
Step 8: Filter BED on splice junctions
Output dataset 'junctions' from step 6
Output dataset 'junctions' from step 7
66
66
Filter out known splice junctions, thereby only keeping the novel ones.
Step 9: Extract Genomic DNA
Output dataset 'novel_junctions' from step 8
No
History
Output dataset 'output' from step 3
Interval
Retrieve the DNA sequences for the novel splice junctions.
Step 10: Translate BED Sequences
Output dataset 'out_file1' from step 9
pep:splice
depth
Yes
66
66
Yes
10
Translate and output splice-junction peptide sequences from the DNA sequences.