Table 1.

Optimization of Neural Network Architecture (Input Parameters vs. Performance on the EGF-Like Domain Type[i])

Training set Test set Total
Parameter No tp fp fn tn C tp fp fn tn C tp fp fn tn C
1NSD, AVS224223493110.771253201370.8536014764600.81
2NSD, AVS, P nfpNSD (NSD), P nfpAVS (AVS)4284473300.97142331370.964268104660.96
3NSD, AVS, P pNSD (NSD),P pAVS (AVS)2291203320.99145301370.98434424700.99
4NSD, AVS, P pNSD (NSD), P pAVS (AVS), P nfpNSD (NSD), P nfpAVS (AVS)4284473300.97142331370.964268104660.96

[i] tp, True positives; fp, false positives; tn, true negatives; fn, false negatives.

[ii] C is the Matthews (Pearson) correlation coefficient (Matthews 1975),

C=tp×tnfp×fn(tp+fn)(tp+fp)(tn+fp)(tn+fn)if fp=fn=0,C=1.000.
Each neural network contained one hidden layer with the same number of elements as the number of input parameters. Note that the Total values were obtained by retraining the ANNs on the entire dataset, so these values are not necessarily equal to the sum of the corresponding Training Set and Test Set values.