(C) PLOS One This story was originally published by PLOS One and is unaltered. . . . . . . . . . . AI-first structural identification of pathogenic protein target interfaces [1] ['Mihkel Saluri', 'Department Of Microbiology', 'Tumor', 'Cell Biology', 'Karolinska Institutet', 'Solna', 'Michael Landreh', 'Patrick Bryant', 'Science For Life Laboratory', 'The Department Of Molecular Biosciences'] Date: 2025-07 The risk of pandemics is increasing as global population growth and interconnectedness accelerate. Understanding the structural basis of protein-protein interactions between pathogens and hosts is critical for elucidating pathogenic mechanisms and guiding treatment or vaccine development. Despite 21,064 experimentally supported human-pathogen interactions in the HPIDB, only 52 have resolved structures in the PDB, representing just 0.2%. Advances in protein complex structure prediction, such as AlphaFold, now enable highly accurate modelling of heterodimeric complexes, though their application to host-pathogen interactions, which have distinct evolutionary dynamics, remains underexplored. Here, we investigate the structural protein-protein interaction network between humans and ten pathogens, predicting structures for 9,452 interactions, only 10 of which have known structures. We identify 30 interactions with an expected TM-score ≥0.9, tripling the structural coverage in these networks. A detailed analysis of the Francisella tularensis dihydroprolyl dehydrogenase (IPD) complex with human immunoglobulin kappa constant (IGKC) using homology modelling and native mass spectrometry confirms a predicted 1:2:1 heterotetramer, suggesting potential roles in immune evasion. These findings highlight the transformative potential of structure prediction for rapidly advancing vaccine and drug development against novel pathogenic targets. New infectious diseases are emerging at an increasing pace, and the ability to quickly understand how pathogens interact with the human body is more important than ever. One powerful way to gain this understanding is by examining the three-dimensional structures of the proteins involved, like seeing how puzzle pieces fit together. But while thousands of these interactions have been observed experimentally, only a small fraction have known structures. Recent advances in artificial intelligence, such as AlphaFold, now allow us to predict these structures with high accuracy. In this study, we applied structure prediction to over 9000 protein-protein interactions between humans and ten different pathogens. We generated high-confidence structural models for several interactions that previously lacked any structural information. We further investigated one of these predictions, involving the bacterium Francisella tularensis and a human immune protein, using native mass spectrometry. Our analysis supports the predicted complex and suggests a potential role in immune system evasion. This work demonstrates how structure prediction can rapidly expand our understanding of host-pathogen interactions and help guide the development of new treatments and vaccines against future infectious threats. Funding: This study was supported by the SciLifeLab & Wallenberg Data Driven Life Science Program (grant: KAW 2020.0239, P.B). Computational resources were enabled by the supercomputing resource Berzelius provided by National Supercomputer Centre at Linköping University and the Knut and Alice Wallenberg foundation with project ids berzelius-2021-29, Berzelius-2023-267, Berzelius-2024-78 and Berzelius-2024-292 (P.B.). M.L. is supported by a Karolinska Institutet faculty-funded Career Position, a Cancerfonden Project grant, the Swedish Research Council (VR) Research Environment Grant, a Consolidator Grant from the Swedish Society for Medical Research (SSMF), and the Knut and Alice Wallenberg foundation (2022.0032). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Copyright: © 2025 Saluri et al. This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Here, we predict the structure of 9452 HP-PPIs between human proteins and ten different pathogens from the HPIDB [ 9 ] with AlphaFold [ 1 , 4 ]. We identify novel interfaces and analyse these with native mass spectrometry. The workflow describes how to apply protein structure prediction for the identification of new targets. Previously, it has been shown that mammalian hosts and pathogens are under high evolutionary pressure, leading to positive selection [ 5 ]. In individual cases, it has been shown that HP-PPIs leave evolutionary marks due to this purifying selection [ 6 , 7 ]. However, apparent positive selection between a host and a pathogen protein does not imply a causal link. The positive selection in the host may be due to a previous pathogen that is now extinct or due to several current pathogens acting together, blurring the so-called “Red Queen” hypothesis [ 8 ] of a host-pathogen arms race [ 5 ]. Docking proteins with physics-based methods alone, without the use of evolutionary relationships between proteins, has proven to perform poorly [ 4 ]. This suggests that it is hard to obtain good results using only physical principles. Recent analyses of protein conformational ranking using AlphaFold [ 3 ] suggest that an energy function far better than any previously known has been learned in AlphaFold. This opens up the possibility to predict highly accurate structures of protein complexes even for HP-PPIs that lack orthologs. The recently developed protein structure prediction methods AlphaFold [ 1 ] (AF) and AlphaFold-multimer [ 2 ] (AFM) have been shown to greatly outperform other methods in protein complex prediction [ 3 ]. Both of these methods rely heavily on multiple sequence alignments (MSAs) [ 1 ]. When predicting the structure of protein complexes, these MSAs have to be paired to utilise the evolutionary relationship between orthologous protein sequences. However, none of these methods consider that hosts and pathogens do not have orthologs by definition. As a consequence, it is not known what the performance of AF and AFM is on HP-PPIs compared to predicting interactions within the same species. The interconnectivity of the world is increasing and so is the arms-race between humans and the pathogens we encounter. The same is true for our crops and livestock. Being able to predict host-pathogen protein-protein interactions (HP-PPIs), can yield structural insight and thereby shorten development timelines substantially, generating more profitable and more timely treatments and preventive measures. For the human race to prosper in the coming centuries, we must stay ahead of such evolutionary threats. During the recent pandemic outbreak of SARS-COV-2, the importance of obtaining fast insights into an emerging pathogen and its relationship with the host has become clear. Information about the interaction between the Spike protein and the human ACE2 receptor provided essential structural information for vaccine development and design. If this information could have been obtained earlier, it is possible that the pandemic would have had less of an impact on society due to vaccines and treatments being developed and deployed faster. Results Structure prediction of known host-pathogen interactions The FoldDock protocol, based on AlphaFold (AF) and AlphaFold-multimer (AFM) [10], was used to predict the structure of 111 host-pathogen protein-protein interactions (HP-PPIs). In addition, templates were added to FoldDock due to indications that this can improve the accuracy in some cases [11]. The median TM-score from MMalign [12] is 0.64 for FoldDock, 0.67 for AFM and 0.68 for FoldDock+templates (1a). However, AFM was trained on all proteins with a release date earlier than 2018-04-30. This leaves only 24 out of 111 structures (22%) to test this method, and the median TM-score for AFM is reduced to 0.63, 0.67 for FoldDock+templates and increased to 0.65 for FoldDock on this set (Fig 1b). PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 1. TM-score comparison between FoldDock and AFM. a) The median TM-score is 0.67 for AFM, 0.64 for FoldDock and 0.68 for FoldDock+templates, using all 111 nonredundant host-pathogen interactions in the PDB. b) Selecting only the host-pathogen interactions that are not present in the AFM training set (n = 24), reduced the median TM-score from AFM to 0.63, to 0.67 for FoldDock+templates and increased it for FoldDock to 0.65. c) Comparison of pDockQ and the TM-score using FoldDock (n = 111). The points represent each model, the solid blue line the running average using a step size of 0.1 in pDockQ and the dashed grey line a cutoff of 0.3 in pDockQ. When the pDockQ score is high, so is the TM-score. https://doi.org/10.1371/journal.pcbi.1013168.g001 The pDockQ score from FoldDock has been shown to discriminate true PPIs and their structural accuracy. To test the structural quality correspondence for HP-PPIs, we calculate pDockQ for the 111 HP-PPIs using FoldDock (Fig 1c) and FoldDock+templates (Fig A in S1 Appendix). We find that the TM-score increases with pDockQ for FoldDock, and a cutoff of 0.3 in pDockQ roughly corresponds to an average TM-score of 0.9 and above. Using FoldDock+templates, however, results in reduced resolution between good and bad models and thereby more false positives. Constructing an ROC curve using pDockQ as a separator, with positive examples here having a TM-score over 0.9, 87% of the models above the TM-score 0.9 can be called correct at a FPR of 5% using FoldDock without templates (ROC AUC = 0.97 for FoldDock and 0.95 for FoldDock+templates, Fig A in S1 Appendix). The reason for using the TM-score instead of the DockQ [13] score is because the full genetic sequences were used here, meaning that the predicted and native structures have different lengths, something DockQ handles poorly (see Fig B in S1 Appendix). Due to the better discriminatory ability using FoldDock alone, we abandon the other methods and continue only with the FoldDock protocol in all subsequent analyses. Structure prediction of novel host-pathogen interactions The host-pathogen interaction database [9] 3.0 contains 69’787 curated interactions in total, of which 32’458 are within humans, and 21’064 of these have direct interactions or physical associations. We selected the top ten unique pathogens in terms of interaction prevalence (Fig 2a); Dengue virus type 2 (Dengue), Human T-cell leukemia virus 1 (HTLV-1), Influenza A virus, Coxiella burnetii, Epstein-Barr virus (EBV), Human papillomavirus type 16 (HPV), Hepatitis C virus genotype 1b (HCV), Francisella tularensis subsp. tularensis (F. tularensis), Bacillus anthracis and Yersinia pestis. In total, there are 9576 interactions (45% of all human interactions) with 4051 unique human proteins and 2560 unique pathogenic proteins among these interactions. In total, 8441 HP-PPIs (88%) were successfully predicted. Fig 2b shows the DockQ distributions for each pathogen. Only a small fraction of the total number of HP-PPIs report a pDockQ score above 0.3. This finding is consistent with the evaluation of the known HP-PPIs from PDB (Fig 1c and Table A in S1 Appendix for summary statistics of all predictions and scores). PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 2. Structure prediction for the 10 selected pathogens. a) Number of interactions with human proteins for each pathogen in the HPIDB. The numbers in parentheses are the number of interactions that could be predicted (some proteins are too large to be predicted in complex on GPUs with 40Gb RAM, limit of approximately 3000 residues). b) Distribution of pDockQ scores for the 8441 HP-PPIs that were successfully predicted. c) Distribution of the average plDDT for each chain in the complexes from the HPIDB (n = 8441) and the PDB (n = 111). The PDB set has much higher plDDT on average. There are 93 complexes out of 111 where both chains have over 70 plDDT on average for the PDB set (84%) and 3472 out of 7948 for the HPIDB set (44%). d) The number of HP-PPIs with known complex structure (n = 10) and with high-quality predictions (above 70 plDDT and pDockQ 0.3, n = 30) for each pathogen. Five of these are from HPV, one from EBV, three from F. tularensis, six from B. anthracis and 15 from Y. pestis. https://doi.org/10.1371/journal.pcbi.1013168.g002 The structures in PDB will be highly ordered, even considering the full-length sequences. This is because disordered proteins are hard to crystallise and thereby, infrequent in structural databases. In reality, many proteins are highly flexible, and many viruses have highly disordered or non-structural proteins [14]. To account for this bias, we analyse the protein disorder distribution measured by the predicted lDDT (plDDT) from AlphaFold using the HP-PPIs from PDB and compare it to the set from the HPIDB (Fig 2c). As expected, the PDB set has higher plDDT on average. To account for this discrepancy, we select only the complexes where both chains have an average plDDT ≥ 70 (43% of the HPIDB complex predictions and 84% of the PDB set) and do not contain any clashes between interacting residues (CBs within 1Å from each other). In some cases, one chain is predicted to intersect the other, although there are no clashes (Fig C in S1 Appendix). These cases were removed as well. Fig 2d shows the number of HP-PPIs with known complex structure (n = 10) and the number of high-quality predictions with an expected TM-score of 0.9 (above 70 plDDT and pDockQ 0.3, n = 30). None of the interactions with known structures were accurately predicted. In total, there are 30 new predictions, expanding the number of available structures fourfold. F. tularensis subsp. tularensis and Y. pestis, which have no known HP-PPI structures, here report 3 and 15 high-quality predictions respectively. B. anthracis, Epstein-Barr virus and Human papillomavirus type 16 have 6, 1 and 5 predictions, respectively, compared to 2, 4 and 3 known structures. [END] --- [1] Url: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1013168 Published and (C) by PLOS One Content appears here under this condition or license: Creative Commons - Attribution BY 4.0. via Magical.Fish Gopher News Feeds: gopher://magical.fish/1/feeds/news/plosone/