Additional file 12. Evaluation of domain based approach for the prediction of peptidoglycan hydrolases At the initial stage of this study, the Pfam domain based methodology was also evaluated for prediction of peptidoglycan hydrolases, however, it could not be used in the final model due to the following reasons. In the positive dataset, the functional domains for only 8,601 (37%) out of 23,062 protein sequences could be found from Pfam database. Since a major (~63%) fraction of our sequences did not have functional domains, therefore, domain based classification could not be used. In contrast, machine learning could be implemented since it is based on multiple composition-based features extracted from the complete protein sequence and thus could predict novel peptidoglycan hydrolyses.