RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction
Reaction yield prediction is hard due to scarce labeled data and the vast, sparsely populated reaction space. Existing string-, fingerprint-, and graph-based encodings only partially capture chemical transformations, especially for complex substrates. RxnCLF addresses this with a self-supervised contrastive learning framework built on a condensed reaction graph (CRG), which unifies reactant and product information into a single graph, explicitly modeling transformation structure rather than disconnected graphs.
Pretrained on 1.7 million Pistachio reactions, RxnCLF learns a compact, continuous latent space that captures both reaction-center features and side-chain contexts. Fine-tuned on yield prediction benchmarks—Buchwald-Hartwig coupling, Pd-catalyzed BH coupling, and proprietary HTE C-N coupling and amide formation datasets—it consistently outperforms graph- and sequence-based baselines, achieving improved R2 and best overall performance. The results suggest CRG-based foundation models are scalable and potentially generalizable to regioselectivity, enantioselectivity, and reaction condition optimization.