Speaker
Description
Antibody discovery increasingly uses AI structure prediction to select candidate binders before experimental testing. Such selection assumes that a model's confidence score reports on binding specificity, and not merely on the plausibility of the predicted complex.
We benchmarked AlphaFold3, Boltz-2 and Chai-1 on 106 nanobody–antigen and 46 antibody–antigen complexes. Each binder was paired with every antigen (cognate and non-cognate matrices), with 50 predictions per pair (over 15,000 per model). Predictions were scored for confidence (ipTM) and for accuracy against the experimental structure (DockQ).
Cognate pairs were generally well predicted, but separation from non-cognate pairs was incomplete (ROC-AUC 0.77 to 0.87). Miscalibration was model-specific: Boltz-2 was systematically overconfident, Chai-1 underconfident. Repeated sampling within a single seed improved DockQ by 0.2 to 0.3, while confidence remained largely unchanged (variation 0.04 to 0.1), indicating that the models do not register an improved pose. Most trajectories reached a plateau after 10 to 25 samples. Oracle selection of the best of 50 predictions outperforms selection by confidence, but requires knowledge of the correct answer, which is not available in a screening setting.
Confidence metrics describe structural plausibility rather than binding specificity, and are not yet suitable as proxies for affinity.