1 Introduction

In this evaluation, there are a total of 9 data tables. Evaluation metrics from the OmicsEV package for these data tables are included in this report, beginning with a summary of the data. The sample distribution by class for each data table is shown in the table below.

class FP-RefRatio FP-RefRatio-MD FP-RefRatio-Weighted-MD FP-VirtualRatio-MD MQ-NoNorm MQ-NoNorm-MD MQ-RefRatio-MD MQ-RefRatio-Weighted-MD MQ-VirtualRatio-MD
Normal 79 79 79 79 79 79 79 79 79
Tumor 103 103 103 103 103 103 103 103 103

Detailed information for each sample included in all data tables is shown below.

sample class batch order
C3L-01287-N Normal 1 1
C3L-00561-N Normal 1 2
C3L-01603-N Normal 1 3
C3N-00834-N Normal 1 4
C3N-01214-N Normal 1 5
C3N-01261-N Normal 1 6
C3L-00917-N Normal 1 7
C3L-00607-N Normal 1 8
C3N-00194-N Normal 1 9
C3L-00010-N Normal 1 10
C3L-01861-N Normal 1 11
C3N-01646-N Normal 1 12
C3L-01281-N Normal 1 13
C3N-00495-N Normal 1 14
C3N-00831-N Normal 1 15
C3L-00183-N Normal 1 16
C3N-00168-N Normal 1 17
C3L-00369-N Normal 1 18
C3L-00791-N Normal 1 19
C3L-00097-N Normal 1 20
C3L-00004-N Normal 1 21
C3N-00953-N Normal 1 22
C3N-00150-N Normal 1 23
C3N-00244-N Normal 1 24
C3N-01178-N Normal 1 25
C3L-00583-N Normal 1 26
C3L-00088-N Normal 1 27
C3L-00814-N Normal 1 28
C3L-01885-N Normal 1 29
C3L-00908-N Normal 1 30
C3L-00026-N Normal 1 31
C3L-00447-N Normal 1 32
C3L-00416-N Normal 1 33
C3N-00246-N Normal 1 34
C3N-00148-N Normal 1 35
C3N-00646-N Normal 1 36
C3L-00902-N Normal 1 37
C3L-00096-N Normal 1 38
C3L-01836-N Normal 1 39
C3N-00177-N Normal 1 40
C3N-01200-N Normal 1 41
C3L-00011-N Normal 1 42
C3N-01808-N Normal 1 43
C3N-01176-N Normal 1 44
C3L-00103-N Normal 1 45
C3N-01649-N Normal 1 46
C3N-00149-N Normal 1 47
C3L-00581-N Normal 1 48
C3L-01313-N Normal 1 49
C3N-00573-N Normal 1 50
C3L-01302-N Normal 1 51
C3N-01361-N Normal 1 52
C3N-01651-N Normal 1 53
C3L-00360-N Normal 1 54
C3L-00448-N Normal 1 55
C3L-00907-N Normal 1 56
C3L-01286-N Normal 1 57
C3L-01607-N Normal 1 58
C3N-01179-N Normal 1 59
C3L-00606-N Normal 1 60
C3N-01648-N Normal 1 61
C3N-00242-N Normal 1 62
C3L-00418-N Normal 1 63
C3N-00577-N Normal 1 64
C3L-01882-N Normal 1 65
C3N-00733-N Normal 1 66
C3L-00910-N Normal 1 67
C3L-00079-N Normal 1 68
C3N-01522-N Normal 1 69
C3N-01220-N Normal 1 70
C3N-00852-N Normal 1 71
C3N-00320-N Normal 1 72
C3N-00437-N Normal 1 73
C3N-00491-N Normal 1 74
C3N-00317-N Normal 1 75
C3N-00310-N Normal 1 76
C3N-00390-N Normal 1 77
C3N-00312-N Normal 1 78
C3N-00494-N Normal 1 79
C3N-01648-T Tumor 1 80
C3N-01214-T Tumor 1 81
C3N-00646-T Tumor 1 82
C3L-00766-T Tumor 1 83
C3L-00360-T Tumor 1 84
C3L-00790-T Tumor 1 85
C3N-00154-T Tumor 1 86
C3L-00765-T Tumor 1 87
C3N-00150-T Tumor 1 88
C3L-01283-T Tumor 1 89
C3N-00315-T Tumor 1 90
C3N-00177-T Tumor 1 91
C3L-00561-T Tumor 1 92
C3L-00800-T Tumor 1 93
C3L-01882-T Tumor 1 94
C3N-00491-T Tumor 1 95
C3N-00577-T Tumor 1 96
C3L-00079-T Tumor 1 97
C3N-01220-T Tumor 1 98
C3L-01607-T Tumor 1 99
C3N-00953-T Tumor 1 100
C3N-00380-T Tumor 1 101
C3N-00317-T Tumor 1 102
C3N-00305-T Tumor 1 103
C3N-00437-T Tumor 1 104
C3N-01651-T Tumor 1 105
C3L-00908-T Tumor 1 106
C3L-00416-T Tumor 1 107
C3N-00733-T Tumor 1 108
C3L-00610-T Tumor 1 109
C3N-01213-T Tumor 1 110
C3N-01200-T Tumor 1 111
C3N-01361-T Tumor 1 112
C3L-00183-T Tumor 1 113
C3L-01861-T Tumor 1 114
C3L-00917-T Tumor 1 115
C3L-00817-T Tumor 1 116
C3L-00581-T Tumor 1 117
C3L-00004-T Tumor 1 118
C3N-01524-T Tumor 1 119
C3L-01302-T Tumor 1 120
C3L-01836-T Tumor 1 121
C3N-00314-T Tumor 1 122
C3L-01287-T Tumor 1 123
C3N-00149-T Tumor 1 124
C3N-00148-T Tumor 1 125
C3N-00390-T Tumor 1 126
C3L-01885-T Tumor 1 127
C3N-00495-T Tumor 1 128
C3N-01646-T Tumor 1 129
C3L-00813-T Tumor 1 130
C3L-00812-T Tumor 1 131
C3N-00494-T Tumor 1 132
C3N-00573-T Tumor 1 133
C3L-01560-T Tumor 1 134
C3N-01649-T Tumor 1 135
C3L-00792-T Tumor 1 136
C3N-01179-T Tumor 1 137
C3N-00168-T Tumor 1 138
C3L-00607-T Tumor 1 139
C3L-00907-T Tumor 1 140
C3N-01178-T Tumor 1 141
C3L-00902-T Tumor 1 142
C3N-00831-T Tumor 1 143
C3L-00096-T Tumor 1 144
C3N-00194-T Tumor 1 145
C3L-00010-T Tumor 1 146
C3N-00852-T Tumor 1 147
C3N-01522-T Tumor 1 148
C3L-01288-T Tumor 1 149
C3L-00799-T Tumor 1 150
C3L-00088-T Tumor 1 151
C3L-00910-T Tumor 1 152
C3N-01261-T Tumor 1 153
C3L-00606-T Tumor 1 154
C3L-01603-T Tumor 1 155
C3L-01553-T Tumor 1 156
C3L-01281-T Tumor 1 157
C3N-00320-T Tumor 1 158
C3L-01557-T Tumor 1 159
C3L-01352-T Tumor 1 160
C3N-01176-T Tumor 1 161
C3N-00834-T Tumor 1 162
C3L-00418-T Tumor 1 163
C3L-00796-T Tumor 1 164
C3L-00103-T Tumor 1 165
C3L-00097-T Tumor 1 166
C3N-00312-T Tumor 1 167
C3L-00583-T Tumor 1 168
C3L-00814-T Tumor 1 169
C3L-01313-T Tumor 1 170
C3L-00026-T Tumor 1 171
C3L-00447-T Tumor 1 172
C3L-00448-T Tumor 1 173
C3N-01808-T Tumor 1 174
C3L-00011-T Tumor 1 175
C3N-00310-T Tumor 1 176
C3N-00244-T Tumor 1 177
C3L-00791-T Tumor 1 178
C3L-01286-T Tumor 1 179
C3N-00246-T Tumor 1 180
C3N-00242-T Tumor 1 181
C3L-00369-T Tumor 1 182

2 Overview

The table below provides an overview about all the quantitative metrics generated in the evaluation. For each metric, the value of the best data table is highlighted in bold and red. The details for each metric can be found in the corresponding sections below.

metric FP-RefRatio FP-RefRatio-MD FP-RefRatio-Weighted-MD FP-VirtualRatio-MD MQ-NoNorm MQ-NoNorm-MD MQ-RefRatio-MD MQ-RefRatio-Weighted-MD MQ-VirtualRatio-MD
#identified features 12210
(0.5989)
12210
(0.5989)
12210
(0.5989)
12210
(0.5989)
10811
(0.5303)
10811
(0.5303)
10811
(0.5303)
10811
(0.5303)
10811
(0.5303)
#quantifiable features 9521
(0.4670)
9521
(0.4670)
9521
(0.4670)
9521
(0.4670)
8461
(0.4150)
8461
(0.4150)
8460
(0.4150)
8460
(0.4150)
8461
(0.4150)
non_missing_value_ratio 0.9521 0.9521 0.9521 0.9521 0.9397 0.9397 0.9397 0.9397 0.9397
data_dist_similarity 0.9526 0.9823 0.9820 0.9812 0.9098 0.9959 0.9752 0.9840 0.9763
complex_auc 0.8167 0.8601 0.8533 0.8715 0.6516 0.6268 0.8461 0.7230 0.8623
func_auc 0.8253 0.8229 0.8310 0.8287 0.6640 0.7046 0.7065 0.7320 0.7352
gene_wise_cor 0.2792 0.4050 0.3939 0.3720 0.1859 0.2111 0.4026 0.2613 0.3625
sample_wise_cor 0.4641 0.4641 0.4664 0.4594 0.4126 0.4126 0.1858 0.3039 0.1695

The radar plot below summarizes results from the overview table above. To generate the radar plot, each metric is scaled from 0 to 1 such that higher values indicate better data quality if necessary. Scaled values are in parentheses in the table.

3 Data depth

3.1 Study-wise

The table below shows the number of identified and quantified proteins or genes for each data table. Identified proteins or genes are those with a measurement in any sample in a data table whereas quantified proteins or genes are those that remain after filtering out those with missing values in more than 50% of the samples in a data table. The values in parentheses are the percentage of proteins or genes identified or quantified based on the total number of proteins or genes (20386) in the study species.

data table #identified features #quantifiable features
FP-RefRatio-MD 12210
(59.89%)
9521
(46.70%)
FP-RefRatio-Weighted-MD 12210
(59.89%)
9521
(46.70%)
FP-RefRatio 12210
(59.89%)
9521
(46.70%)
FP-VirtualRatio-MD 12210
(59.89%)
9521
(46.70%)
MQ-NoNorm-MD 10811
(53.03%)
8461
(41.50%)
MQ-NoNorm 10811
(53.03%)
8461
(41.50%)
MQ-RefRatio-MD 10811
(53.03%)
8460
(41.50%)
MQ-RefRatio-Weighted-MD 10811
(53.03%)
8460
(41.50%)
MQ-VirtualRatio-MD 10811
(53.03%)
8461
(41.50%)

The upset chart below shows overlap between proteins or genes identified in each data table. Numbers of proteins or genes commonly identified in different combinations of data tables are indicated in the top bar chart, and the specific combinations of data tables containing those proteins or genes are indicated with solid points below the bar chart. Total identifications for each data table are indicated on the right as ‘Set size’.

3.2 Sample-wise

The figures below show the number of proteins or genes identified/quantified (non-missing values) in each sample. Samples from different batches are coded with different shapes, and samples from different classes are coded with different colors. A separate figure is shown for each data table.

FP-RefRatio-MDFP-RefRatio-Weighted-MDFP-RefRatioFP-VirtualRatio-MDMQ-NoNorm-MDMQ-NoNormMQ-RefRatio-MDMQ-RefRatio-Weighted-MDMQ-VirtualRatio-MD

3.3 Missing value distribution

The missing value distribution provides an overview of the completeness of the data. The table below shows the percent of missing values for all samples in each data table.

data table non_missing_value_ratio
FP-RefRatio-MD 0.9521
FP-RefRatio-Weighted-MD 0.9521
FP-RefRatio 0.9521
FP-VirtualRatio-MD 0.9521
MQ-NoNorm-MD 0.9397
MQ-NoNorm 0.9397
MQ-RefRatio-MD 0.9397
MQ-RefRatio-Weighted-MD 0.9397
MQ-VirtualRatio-MD 0.9397

The following barplots show missing value distributions for each data table as number (Y axis)/percentage (number above bar) of proteins or genes with missing values in each bin. Genes are binned by proportion of samples with missing values from 0.1 to 1 in increments of 0.1, where 0.1 indicates missing values in no more than 10% of the samples, and 1 indicates missing values in all samples.

FP-RefRatio-MDFP-RefRatio-Weighted-MDFP-RefRatioFP-VirtualRatio-MDMQ-NoNorm-MDMQ-NoNormMQ-RefRatio-MDMQ-RefRatio-Weighted-MDMQ-VirtualRatio-MD

4 Data normalization

4.1 Boxplot

Normalized data is expected to be centered around a similar value and show similar distributions in all samples. The boxplots below show the protein or gene expression measurement distribution across samples in each data table, allowing for qualitative assessment of the normalized data. Samples in input order are indicated on the X axis. The Y axis shows log2 transformed protein or gene values. Samples from different classes are coded with different colors.

FP-RefRatio-MDFP-RefRatio-Weighted-MDFP-RefRatioFP-VirtualRatio-MDMQ-NoNorm-MDMQ-NoNormMQ-RefRatio-MDMQ-RefRatio-Weighted-MDMQ-VirtualRatio-MD

To quantify the normalization effect, we tested for how well the data in the feature set can distinguish between each pair of samples. If the distribution is similar for the two samples in a given pair, the overall feature abundance (levels for all features in one sample vs the other) should not be sufficient to predict which sample is which. Therefore, for each pair of samples, an AUROC test was performed to quantify the ability of feature abundance to distinguish the two samples, and then a data_dist_similarity score was generated: 1-2*abs(AUROC-0.5). This score ranges from 0 to 1, and the higher the score is the better the normalized data quality is (no systematic difference between the two samples). The final metric for each data table is the median of scores from all sample pairs. The column ‘n’ shows the total number of sample pairs in the analysis.

data table data_dist_similarity n
FP-RefRatio-MD 0.9823 16471
FP-RefRatio-Weighted-MD 0.9820 16471
FP-RefRatio 0.9526 16471
FP-VirtualRatio-MD 0.9812 16471
MQ-NoNorm-MD 0.9959 16471
MQ-NoNorm 0.9098 16471
MQ-RefRatio-MD 0.9752 16471
MQ-RefRatio-Weighted-MD 0.9840 16471
MQ-VirtualRatio-MD 0.9763 16471

4.2 Density plot

The density plots below show the expression distributions for all samples (separate line) in each data table. The Y axis shows the density over the range of log2 transformed protein or gene expression values (X axis).

5 Batch effect

5.1 Correlation heatmap

Another way to qualitatively assess batch effect is to visualize the correlations for measurements between samples from the same batch to those in samples from different batches using heatmaps. The following figures show Spearman correlation heatmaps for all pairs of samples (all samples included in both rows and columns) for each data table. The color indicates the correlation between samples. The samples are ordered by batches. Concentration of high correlation values (red color) for pairs of samples from the same batch block compared to other batches indicates the presence of batch effect.

FP-RefRatio-MDFP-RefRatio-Weighted-MDFP-RefRatioFP-VirtualRatio-MDMQ-NoNorm-MDMQ-NoNormMQ-RefRatio-MDMQ-RefRatio-Weighted-MDMQ-VirtualRatio-MD

6 Biological signal

6.1 Correlation among protein complex members

Members of the same protein complex often show greater correlation in gene and protein expression (IntraComplex correlation) than genes or proteins that are in different complexes (InterComplex correlation). Thus, one way to evaluate the quality of the biological signal present in a data table is to compare IntraComplex correlation to InterComplex correlation. Furthermore, because of the need to preserve stoichiometry between protein complex members, the difference between IntraComplex correlation and InterComplex correlation is often greater at the protein level than at the RNA data. If both RNA and protein data tables are available, observing that this difference is more pronounced in the protein data table than the RNA data table serves as an indicator for the quality of the protein data. We use the protein complexes from the CORUM database in this analysis.

The boxplots below show the distributions and ranges for pairwise correlations between genes or proteins from the same complex and for genes and proteins from different complexes for each data table.

The table below shows a summary of the evaluation. ‘diff’ is Cor(intra) - Cor(inter). ‘complex_auc’ is the AUROC value based on correlation of protein pairs from different groups.

data table InterComplex IntraComplex diff complex_auc
FP-RefRatio 0.6583 0.8656 0.2073 0.8167
FP-RefRatio-MD 0.0150 0.4244 0.4094 0.8601
FP-RefRatio-Weighted-MD 0.0188 0.3935 0.3746 0.8533
FP-VirtualRatio-MD -0.0027 0.4916 0.4943 0.8715
MQ-NoNorm 0.3296 0.4467 0.1171 0.6516
MQ-NoNorm-MD 0.0220 0.1128 0.0908 0.6268
MQ-RefRatio-MD 0.0072 0.3675 0.3603 0.8461
MQ-RefRatio-Weighted-MD 0.0052 0.1437 0.1385 0.7230
MQ-VirtualRatio-MD -0.0057 0.4466 0.4523 0.8623
RNA 0.0163 0.1472 0.1310 0.6570

6.2 Gene function prediction

Previous studies have shown that expression correlation is often higher for functionally related genes or proteins than for unrelated genes or proteins and that this correlation is greater when considering protein data than when considering RNA data (Wang, Jing, et al. Molecular & Cellular Proteomics 16.1 (2017): 121-134.). Therefore, we can also evaluate the biological signal present in a data table by evaluating functional category predictions made using a co-expression network generated from each data table.

In this evaluation, each data table was used to build a co-expression network. For a selected network and a selected functional category (such as a selected category from GO or KEGG), proteins/genes annotated to the category and also included in the network were defined as a positive protein/gene set, and other proteins/genes in the network constituted the negative protein/gene set for the category. For a selected functional category, a subset of the proteins/genes were used as seed proteins/genes for random walk through the network to calculate scores for other proteins/genes. A higher score for a protein/gene represents a closer relationship between the protein/gene and the seed proteins/genes. The table below shows AUROCs of the prediction performance using this score for each selected functional category.

FP-RefRatio FP-RefRatio-MD FP-RefRatio-Weighted-MD FP-VirtualRatio-MD MQ-NoNorm MQ-NoNorm-MD MQ-RefRatio-MD MQ-RefRatio-Weighted-MD MQ-VirtualRatio-MD RNA
Acute myeloid leukemia 0.837 0.66 0.707 0.779 0.635 0.589 0.653 0.593 0.608 0.728
Adherens junction 0.82 0.712 0.75 0.743 0.607 0.552 0.741 0.561 0.734 0.575
Adipocytokine signaling pathway 0.668 0.747 0.598 0.691 0.613 0.634 0.607 0.639 0.605 0.59
Allograft rejection 0.91 0.987 0.941 0.989 0.669 0.861 0.967 0.977 0.977 0.965
Alzheimers disease 0.816 0.833 0.839 0.833 0.573 0.676 0.645 0.718 0.677 0.793
Amino sugar and nucleotide sugar metabolism 0.737 0.744 0.798 0.712 0.66 0.62 0.78 0.63 0.805 0.747
Amoebiasis 0.81 0.787 0.773 0.785 0.637 0.645 0.671 0.709 0.76 0.751
Amyotrophic lateral sclerosis (ALS) 0.62 0.608 0.57 0.677 0.584 0.616 0.67 0.654 0.593 0.614
Antigen processing and presentation 0.961 0.971 0.934 0.96 0.628 0.821 0.903 0.832 0.932 0.918
Apoptosis 0.688 0.747 0.639 0.74 0.566 0.536 0.652 0.618 0.707 0.607
Arachidonic acid metabolism 0.67 0.534 0.536 0.608 0.675 0.603 0.627 0.653 0.634 0.616
Arginine and proline metabolism 0.883 0.856 0.828 0.784 0.74 0.715 0.72 0.801 0.73 0.609
Arrhythmogenic right ventricular cardiomyopathy (ARVC) 0.753 0.735 0.783 0.83 0.644 0.62 0.713 0.636 0.834 0.653
Autoimmune thyroid disease 0.9 0.987 0.95 0.989 0.672 0.813 0.954 0.971 0.975 0.954
Axon guidance 0.71 0.726 0.642 0.733 0.594 0.621 0.56 0.586 0.628 0.598
B cell receptor signaling pathway 0.766 0.779 0.754 0.809 0.637 0.705 0.664 0.7 0.708 0.62
Bacterial invasion of epithelial cells 0.795 0.809 0.766 0.755 0.541 0.519 0.702 0.644 0.755 0.593
Bladder cancer 0.798 0.649 0.669 0.653 0.552 0.603 0.711 0.615 0.659 0.578
Calcium signaling pathway 0.793 0.758 0.775 0.724 0.597 0.557 0.719 0.664 0.663 0.556
Carbohydrate digestion and absorption 0.774 0.708 0.677 0.733 0.61 0.64 0.795 0.746 0.612 0.544
Cell adhesion molecules (CAMs) 0.801 0.818 0.797 0.907 0.589 0.613 0.797 0.793 0.788 0.756
Cell cycle 0.863 0.806 0.809 0.858 0.663 0.733 0.839 0.728 0.791 0.752
Chagas disease (American trypanosomiasis) 0.754 0.74 0.802 0.763 0.548 0.551 0.721 0.687 0.732 0.639
Chemokine signaling pathway 0.776 0.648 0.687 0.754 0.603 0.683 0.655 0.609 0.693 0.644
Chronic myeloid leukemia 0.738 0.73 0.607 0.742 0.552 0.656 0.685 0.544 0.655 0.531
Colorectal cancer 0.746 0.687 0.646 0.728 0.568 0.578 0.716 0.615 0.69 0.591
Complement and coagulation cascades 0.951 0.954 0.955 0.922 0.928 0.903 0.938 0.982 0.932 0.686
Cysteine and methionine metabolism 0.743 0.764 0.755 0.762 0.716 0.683 0.747 0.699 0.74 0.602
Cytokine-cytokine receptor interaction 0.78 0.852 0.781 0.833 0.734 0.755 0.622 0.761 0.621 0.682
Cytosolic DNA-sensing pathway 0.657 0.725 0.729 0.71 0.69 0.517 0.775 0.683 0.635 0.664
Dilated cardiomyopathy 0.778 0.705 0.806 0.77 0.622 0.63 0.715 0.589 0.712 0.709
Drug metabolism - cytochrome P450 0.894 0.774 0.731 0.677 0.76 0.676 0.875 0.726 0.802 0.671
ECM-receptor interaction 0.875 0.805 0.847 0.862 0.815 0.789 0.812 0.784 0.84 0.776
Endocytosis 0.841 0.789 0.752 0.773 0.617 0.564 0.619 0.604 0.699 0.616
Endometrial cancer 0.792 0.746 0.714 0.819 0.586 0.567 0.662 0.623 0.613 0.635
Epithelial cell signaling in Helicobacter pylori infection 0.777 0.694 0.812 0.763 0.657 0.641 0.631 0.638 0.593 0.577
ErbB signaling pathway 0.765 0.682 0.639 0.665 0.545 0.541 0.655 0.592 0.616 0.6
Fc epsilon RI signaling pathway 0.853 0.756 0.78 0.794 0.618 0.565 0.761 0.627 0.735 0.558
Fc gamma R-mediated phagocytosis 0.815 0.786 0.73 0.78 0.6 0.678 0.735 0.637 0.764 0.64
Focal adhesion 0.805 0.785 0.793 0.794 0.687 0.728 0.786 0.685 0.771 0.749
Fructose and mannose metabolism 0.82 0.797 0.78 0.793 0.657 0.702 0.765 0.714 0.784 0.713
Galactose metabolism 0.747 0.703 0.741 0.706 0.617 0.691 0.676 0.662 0.774 0.665
Gap junction 0.732 0.713 0.728 0.696 0.63 0.636 0.624 0.578 0.568 0.516
Gastric acid secretion 0.75 0.682 0.624 0.553 0.606 0.576 0.626 0.714 0.66 0.529
Glioma 0.703 0.675 0.631 0.702 0.587 0.539 0.736 0.661 0.604 0.615
Glutathione metabolism 0.603 0.661 0.641 0.592 0.567 0.643 0.784 0.648 0.706 0.512
Glycerolipid metabolism 0.748 0.752 0.75 0.677 0.684 0.685 0.685 0.837 0.592 0.738
Glycerophospholipid metabolism 0.63 0.707 0.596 0.678 0.641 0.603 0.569 0.686 0.658 0.566
Glycolysis / Gluconeogenesis 0.804 0.823 0.831 0.865 0.662 0.744 0.849 0.808 0.83 0.763
GnRH signaling pathway 0.653 0.691 0.623 0.613 0.59 0.583 0.631 0.562 0.702 0.585
Graft-versus-host disease 0.937 0.953 0.977 0.999 0.579 0.873 0.97 0.967 0.975 0.954
Hematopoietic cell lineage 0.662 0.713 0.738 0.7 0.633 0.59 0.757 0.714 0.667 0.676
Hepatitis C 0.787 0.765 0.748 0.842 0.579 0.583 0.706 0.614 0.748 0.577
Huntingtons disease 0.85 0.877 0.86 0.888 0.633 0.697 0.541 0.671 0.611 0.792
Hypertrophic cardiomyopathy (HCM) 0.713 0.696 0.774 0.769 0.584 0.577 0.631 0.645 0.814 0.636
Inositol phosphate metabolism 0.637 0.549 0.606 0.626 0.606 0.589 0.574 0.613 0.64 0.625
Insulin signaling pathway 0.763 0.666 0.667 0.709 0.621 0.656 0.61 0.639 0.746 0.58
Jak-STAT signaling pathway 0.746 0.671 0.756 0.856 0.533 0.629 0.621 0.767 0.601 0.586
Leishmaniasis 0.808 0.845 0.931 0.859 0.626 0.776 0.796 0.798 0.784 0.752
Leukocyte transendothelial migration 0.804 0.699 0.74 0.763 0.643 0.667 0.735 0.713 0.781 0.682
Long-term depression 0.798 0.731 0.663 0.676 0.553 0.702 0.681 0.677 0.695 0.584
Long-term potentiation 0.765 0.634 0.677 0.636 0.632 0.66 0.756 0.582 0.696 0.617
Lysosome 0.902 0.908 0.889 0.87 0.647 0.748 0.799 0.692 0.828 0.782
Malaria 0.665 0.735 0.841 0.76 0.686 0.747 0.79 0.723 0.77 0.645
MAPK signaling pathway 0.72 0.654 0.686 0.636 0.583 0.589 0.539 0.556 0.616 0.609
Melanogenesis 0.659 0.709 0.685 0.641 0.599 0.619 0.787 0.515 0.652 0.633
Melanoma 0.604 0.637 0.718 0.709 0.595 0.544 0.657 0.663 0.631 0.623
Metabolic pathways 0.769 0.786 0.773 0.764 0.654 0.689 0.716 0.736 0.705 0.645
Metabolism of xenobiotics by cytochrome P450 0.849 0.764 0.846 0.7 0.662 0.61 0.658 0.715 0.762 0.645
mRNA surveillance pathway 0.838 0.856 0.87 0.802 0.765 0.694 0.572 0.698 0.646 0.71
mTOR signaling pathway 0.731 0.635 0.771 0.725 0.556 0.547 0.516 0.561 0.652 0.587
N-Glycan biosynthesis 0.981 0.926 0.871 0.891 0.823 0.838 0.915 0.879 0.914 0.711
Natural killer cell mediated cytotoxicity 0.776 0.723 0.781 0.8 0.604 0.717 0.692 0.749 0.724 0.642
Neurotrophin signaling pathway 0.719 0.678 0.728 0.742 0.587 0.675 0.586 0.556 0.724 0.643
NOD-like receptor signaling pathway 0.765 0.747 0.766 0.785 0.636 0.566 0.651 0.732 0.704 0.563
Non-small cell lung cancer 0.773 0.697 0.677 0.718 0.66 0.526 0.755 0.673 0.689 0.55
Nucleotide excision repair 0.866 0.935 0.952 0.818 0.664 0.686 0.861 0.734 0.771 0.749
Oocyte meiosis 0.792 0.765 0.723 0.779 0.603 0.613 0.658 0.562 0.65 0.58
Osteoclast differentiation 0.763 0.802 0.821 0.809 0.583 0.681 0.71 0.708 0.683 0.605
p53 signaling pathway 0.708 0.754 0.712 0.711 0.727 0.67 0.627 0.62 0.676 0.724
Pancreatic cancer 0.731 0.67 0.683 0.79 0.56 0.611 0.58 0.617 0.692 0.564
Pancreatic secretion 0.627 0.743 0.621 0.666 0.569 0.648 0.59 0.571 0.616 0.634
Parkinsons disease 0.886 0.879 0.858 0.887 0.705 0.758 0.64 0.724 0.769 0.821
Pathogenic Escherichia coli infection 0.864 0.743 0.722 0.786 0.567 0.531 0.74 0.552 0.663 0.608
Pathways in cancer 0.653 0.624 0.635 0.686 0.56 0.55 0.618 0.547 0.634 0.602
Pentose phosphate pathway 0.695 0.712 0.801 0.828 0.674 0.687 0.873 0.709 0.798 0.683
Peroxisome 0.832 0.853 0.839 0.797 0.726 0.734 0.793 0.635 0.715 0.732
Phagosome 0.835 0.839 0.851 0.848 0.579 0.704 0.808 0.717 0.788 0.718
Phosphatidylinositol signaling system 0.67 0.655 0.588 0.609 0.581 0.585 0.648 0.592 0.63 0.673
PPAR signaling pathway 0.749 0.658 0.607 0.726 0.629 0.647 0.658 0.587 0.64 0.644
Prion diseases 0.858 0.832 0.796 0.804 0.671 0.689 0.703 0.752 0.789 0.595
Progesterone-mediated oocyte maturation 0.714 0.796 0.809 0.807 0.65 0.693 0.687 0.625 0.596 0.679
Prostate cancer 0.754 0.664 0.692 0.688 0.654 0.604 0.716 0.635 0.651 0.617
Proteasome 0.969 0.99 0.992 0.965 0.726 0.805 0.945 0.874 0.892 0.886
Protein digestion and absorption 0.878 0.851 0.827 0.891 0.634 0.853 0.814 0.872 0.829 0.693
Protein export 0.985 0.98 0.968 0.961 0.749 0.748 0.933 0.807 0.927 0.798
Protein processing in endoplasmic reticulum 0.812 0.823 0.854 0.82 0.63 0.718 0.768 0.755 0.774 0.761
Purine metabolism 0.631 0.743 0.686 0.685 0.579 0.547 0.708 0.561 0.679 0.606
Pyrimidine metabolism 0.63 0.604 0.622 0.614 0.561 0.578 0.689 0.527 0.611 0.596
Regulation of actin cytoskeleton 0.754 0.785 0.743 0.759 0.604 0.617 0.739 0.604 0.713 0.662
Renal cell carcinoma 0.67 0.614 0.555 0.73 0.589 0.563 0.612 0.72 0.616 0.626
Rheumatoid arthritis 0.911 0.977 0.914 0.919 0.703 0.812 0.899 0.883 0.882 0.733
Ribosome 0.949 0.961 0.965 0.98 0.752 0.859 0.952 0.888 0.928 0.925
Ribosome biogenesis in eukaryotes 0.877 0.915 0.919 0.87 0.736 0.786 0.885 0.817 0.801 0.787
RIG-I-like receptor signaling pathway 0.691 0.695 0.758 0.838 0.589 0.626 0.659 0.582 0.706 0.588
RNA degradation 0.855 0.779 0.75 0.839 0.66 0.678 0.706 0.619 0.557 0.678
RNA transport 0.861 0.838 0.791 0.85 0.732 0.764 0.703 0.725 0.635 0.662
Salivary secretion 0.718 0.734 0.742 0.711 0.688 0.537 0.738 0.624 0.754 0.625
Shigellosis 0.734 0.757 0.746 0.664 0.607 0.633 0.661 0.598 0.725 0.692
Small cell lung cancer 0.675 0.628 0.672 0.664 0.65 0.709 0.596 0.678 0.622 0.625
SNARE interactions in vesicular transport 0.838 0.827 0.804 0.807 0.682 0.705 0.82 0.691 0.708 0.58
Spliceosome 0.939 0.952 0.889 0.961 0.749 0.831 0.897 0.819 0.819 0.773
Staphylococcus aureus infection 0.844 0.974 0.983 0.919 0.876 0.898 0.909 0.96 0.919 0.889
Starch and sucrose metabolism 0.796 0.742 0.724 0.654 0.775 0.648 0.734 0.744 0.717 0.702
Systemic lupus erythematosus 0.852 0.952 0.875 0.913 0.803 0.853 0.933 0.939 0.91 0.797
T cell receptor signaling pathway 0.789 0.736 0.726 0.823 0.668 0.687 0.647 0.702 0.672 0.634
TGF-beta signaling pathway 0.873 0.681 0.752 0.832 0.596 0.641 0.772 0.629 0.774 0.615
Tight junction 0.762 0.716 0.696 0.626 0.602 0.69 0.598 0.609 0.631 0.664
Toll-like receptor signaling pathway 0.748 0.683 0.678 0.686 0.528 0.594 0.569 0.662 0.726 0.533
Toxoplasmosis 0.735 0.72 0.714 0.787 0.595 0.63 0.781 0.724 0.716 0.66
Type I diabetes mellitus 0.926 0.928 0.951 0.963 0.641 0.752 0.954 0.945 0.892 0.876
Ubiquitin mediated proteolysis 0.758 0.728 0.731 0.817 0.511 0.562 0.598 0.633 0.594 0.668
Vascular smooth muscle contraction 0.874 0.76 0.782 0.716 0.746 0.703 0.731 0.6 0.722 0.598
Vasopressin-regulated water reabsorption 0.904 0.786 0.836 0.794 0.665 0.577 0.676 0.608 0.781 0.702
VEGF signaling pathway 0.683 0.678 0.739 0.665 0.618 0.571 0.642 0.655 0.649 0.607
Vibrio cholerae infection 0.908 0.828 0.838 0.828 0.671 0.682 0.746 0.732 0.766 0.672
Viral myocarditis 0.74 0.71 0.785 0.82 0.612 0.639 0.707 0.815 0.703 0.852
Wnt signaling pathway 0.786 0.656 0.696 0.666 0.586 0.555 0.675 0.538 0.705 0.565

The rank boxplots below summarize the relative performance of the data tables in the functional prediction analysis. For each functional category, a rank is assigned to each data table based on its AUROC compared to the other data tables, where the best functional prediction rank is 1 and the poorest rank is the number of data tables.

Comparison of each protein (RNA) data table to a designated RNA (protein) data table is also summarized in the scatter plots below. For each point, the AUROC for a given category in the RNA data is plotted on the X-axis whereas the corresponding AUROC in the protein data table is plotted on the Y-axis. The number of categories for which the protein data table outperforms the RNA data table (AUROC(protein) > 1.1 * AUROC (RNA); red dots) and vice versa (AUROC(RNA) > 1.1 * AUROC (protein); blue dots) are also shown.

FP-RefRatioFP-RefRatio-MDFP-RefRatio-Weighted-MDFP-VirtualRatio-MDMQ-NoNormMQ-NoNorm-MDMQ-RefRatio-MDMQ-RefRatio-Weighted-MDMQ-VirtualRatio-MD

6.3 PCA with sample class annotation

Another approach for assessing how well each data table can distinguish between classes is to determine how well each class can be separated by principal component analysis (PCA). In PCA score plots for each data table below, each point is a sample that is colored by class and that has a shape reflecting the batch. For a given sample, the PC2 score is plotted on the Y-axis whereas the PC1 score is plotted on the X-axis. Ellipses highlighting clusters of samples in each class are colored by corresponding class, and the separation between these ellipses indicates how well the variances captured by the first two PCs can distinguish between samples from different classes.

FP-RefRatio-MDFP-RefRatio-Weighted-MDFP-RefRatioFP-VirtualRatio-MDMQ-NoNorm-MDMQ-NoNormMQ-RefRatio-MDMQ-RefRatio-Weighted-MDMQ-VirtualRatio-MD

6.4 Unsupervised clustering

Unsupervised hierarchical clustering can reveal patterns in the data (clusters of genes or samples that behave more similarly to each other than to other genes or samples). Each heatmap below shows the results of hierarchical clustering for a given data table using ComplexHeatmap. Genes/proteins are in rows, while samples are in columns and labeled with corresponding class to visualize any potential associations between classes and clusters.

FP-RefRatio-MDFP-RefRatio-Weighted-MDFP-RefRatioFP-VirtualRatio-MDMQ-NoNorm-MDMQ-NoNormMQ-RefRatio-MDMQ-RefRatio-Weighted-MDMQ-VirtualRatio-MD

7 Multi-omics concordance

The concordance between the protein data and RNA data can be used to assess data quality when both RNA and protein data tables are available. Here, we evaluate gene- and sample-wise correlations between the protein and RNA data tables.

7.1 Gene-wise mRNA-protein correlation

The table below shows the number of genes with measurements (n) in each data table as well as the median of all gene-wise Spearman correlations between mRNA and protein measurements. The columns n5, n6, n7 and n8 show the number of genes with correlation greater than 0.5, 0.6, 0.7 and 0.8, respectively.

data table n n5 n6 n7 n8 gene_wise_cor
FP-RefRatio-MD 9034 3372 2217 1204 467 0.4050
FP-RefRatio-Weighted-MD 9034 3230 2107 1137 432 0.3939
FP-RefRatio 9034 1586 799 319 79 0.2792
FP-VirtualRatio-MD 9034 2921 1735 806 202 0.3720
MQ-NoNorm-MD 8044 917 446 180 50 0.2111
MQ-NoNorm 8044 512 202 70 9 0.1859
MQ-RefRatio-MD 8043 2938 1921 1014 366 0.4026
MQ-RefRatio-Weighted-MD 8043 1217 583 185 22 0.2613
MQ-VirtualRatio-MD 8044 2411 1372 554 123 0.3625

Spearman correlation results are also shown for each gene/protein in the boxplot below.

Another way to visualize the differences between the distributions of all gene-wise RNA-protein correlations is with the cumulative distribution function (CDF) plot shown below. Here each line shows the cumulative distribution for the gene-wise correlations. The further the distribution function is shifted to the right, the more highly correlated the RNA-protein data is.

The histograms below provide another way to visualize the distribution of correlations for each protein (or RNA) data table with the RNA (or protein) data. Here the bars showing binned frequencies of positive correlations are in red, while negative correlations are shown in the blue bins, and summary statistics are also provided.

FP-RefRatio-MDFP-RefRatio-Weighted-MDFP-RefRatioFP-VirtualRatio-MDMQ-NoNorm-MDMQ-NoNormMQ-RefRatio-MDMQ-RefRatio-Weighted-MDMQ-VirtualRatio-MD

7.2 Sample-wise mRNA-protein correlation

Sample-wise RNA-protein correlations are summarized in the table below as the median of Spearman correlations for matched protein and RNA data from all pairs of samples for each data table, while the violin plots below show the distributions of these correlations for each data table.

data table sample_wise_cor
FP-RefRatio 0.4641
FP-RefRatio-MD 0.4641
FP-RefRatio-Weighted-MD 0.4664
FP-VirtualRatio-MD 0.4594
MQ-NoNorm 0.4126
MQ-NoNorm-MD 0.4126
MQ-RefRatio-MD 0.1858
MQ-RefRatio-Weighted-MD 0.3039
MQ-VirtualRatio-MD 0.1695