Figure 1.

Schematic diagram of the HiDT algorithm. (A) Construction of positive and negative TAD pairs for model training. Positive pairs included differential TAD pairs between GM12878 and K562, whereas negative pairs included GM12878 replicate pairs and nondifferential TAD pairs between GM12878 and K562. The training set comprised 16,280 positive pairs and 17,290 negative pairs. (B) Conversion of paired Hi-C contact matrices into graph representations, with nodes representing genomic bins and edge weights representing interaction frequencies. The resulting graph is not necessarily fully connected, and nodes retain self-interactions. (C) Learning of graph-level embeddings using an edge-enhanced graph neural network with depth-specific normalization and gated MLP aggregation, where G1 and G2 denote the two graphs in a TAD pair, Xi′ denotes the node embedding vector, X^i denotes the normalized node vector, and hG1and hG2 denote the graph representation vectors. (D) Model training with a pairwise loss to separate differential and nondifferential TAD pairs in the embedding space, where d(G1 − G2) denotes the Euclidean distance between graph embeddings, γ denotes the margin, and t ∈ {− 1, 1} indicates whether a pair is negative or positive. (E) TAD pairs with Euclidean distances above the cutoff are classified as differential, where d denotes the Euclidean distance between graph embeddings.

2127f01