From Art to Engineering: An Architecture Evaluation Framework for Deep Graph Transformation Tasks
Paper submitted for peer review. Publication expected in late 2026.
Deep graph transformation lacks evaluation protocols for comparing architectures across structurally different graph transformation tasks. We introduce an evaluation framework for edge-level graph transformation that places nine architectures under a shared evaluation contract so that observed differences reflect architectural suitability rather than incidental experimental variation. Using thirteen graph-computation and real-world tasks drawn from a deep graph transformation taxonomy, we show that no single architecture dominates across tasks: suitability is governed less by model complexity or graph-theoretic complexity class than by the structural computation required. Dense models dominate when canonical or domain-specific features encode the relevant edge signal; graph-native models dominate when multi-hop relational verification is required; and global attention is useful only for some forms of long-range structure. Cross-task clustering recovers architecture-response groups that cut across the original taxonomy, suggesting an empirical task dimension based on four computational primitives. The framework provides reproducible baselines and experimental instruments for studying when relational inductive bias helps, fails, or becomes unnecessary.
Recommended citation: A. Machowczyk, S. De Sabbata and R. Heckel, "From Art to Engineering: An Architecture Evaluation Framework for Deep Graph Transformation Tasks", 2026.
