(C) PLOS One This story was originally published by PLOS One and is unaltered. . . . . . . . . . . Bayesian network imputation methods applied to multi-omics data identify putative causal relationships in a type 2 diabetes dataset containing incomplete data: An IMI DIRECT Study [1] ['Richard Howey', 'Research Software Engineering', 'Newcastle University', 'Newcastle Upon Tyne', 'United Kingdom', 'Population Health Sciences Institute', 'Faculty Of Medical Sciences', 'Jonathan Adam', 'Research Unit Of Molecular Epidemiology', 'Institute Of Epidemiology'] Date: 2025-08 Here we report the results from exploratory analysis using a Bayesian network approach of data originally derived from a large North European study of type 2 diabetes (T2D) conducted by the IMI DIRECT consortium. 3029 individuals (795 with T2D and 2234 without) within 7 different study centres provided data comprising genotypes, proteins, metabolites, gene expression measurements and many different clinical variables. The main aim of the current study was to demonstrate the utility of our previously developed method to fit Bayesian networks by performing exploratory analysis of this dataset to identify possible causal relationships between these variables. The data was analysed using the BayesNetty software package, which can handle mixed discrete/continuous data with missing values. The original dataset consisted of over 16,000 variables, which were filtered down to 260 variables for analysis. Even with this reduction, no individual had complete data for all variables, making it impossible to analyse using standard Bayesian network methodology. However, using the recently proposed novel imputation method implemented in BayesNetty we computed a large average Bayesian network from which we could infer possible associations and causal relationships between variables of interest. Our results confirmed many previous findings in connection with T2D, including possible mediating proteins and genes, some of which have not been widely reported. We also confirmed potential causal relationships with liver fat that were identified in an earlier study that used the IMI DIRECT dataset but was limited to a smaller subset of individuals and variables (namely individuals with complete data at pre-defined variables of interest). In addition to providing valuable confirmation, our analyses thus demonstrate a proof-of-principle of the utility of the method implemented within BayesNetty. The full final average Bayesian network generated from our analysis is freely available and can be easily interrogated further to address specific focussed scientific questions of interest. Bayesian network analysis can be used to identify putative causal relationships between measured variables, including clinical measurements and measurements of genetic and genomic factors. Here we report the results from Bayesian network analysis of data originally derived from a large North European study of type 2 diabetes (T2D). Data were collected for 3029 individuals within 7 different study centres, with the data comprising genotypes, proteins, metabolites, gene expression measurements and many different clinical variables. The original dataset consisted of over 16,000 variables, which were then filtered down to 260 variables for analysis. Even with this reduction, not one individual had complete data for all variables. Using standard methodology it would not be possible to a fit a Bayesian network as it requires complete data. However, using the novel imputation method implemented in the BayesNetty software package, we were able to compute a large average Bayesian network from which we could infer possible associations and causal relationships between variables of interest. Competing interests: I have read the journal’s policy and the authors of this manuscript have the following competing interests: HJC and AV serve on the editorial board of PLOS Genetics. MMcC and AM are currently employees of Genentech and holders of Roche stock. Data Availability: The molecular and clinical raw data as well as the processed data are available under restricted access due to the informed consent given by study participants, the various national ethical approvals for the present study, and the European General Data Protection Regulation (GDPR); individual-level clinical and molecular data cannot be transferred from the centralized IMI DIRECT repository. Requests for access will be informed about how data can be accessed via the IMI DIRECT secure analysis platform following submission of an appropriate application. The IMI DIRECT data access policy is available at https://directdiabetes.org . In this current investigation, we fitted a large average BN (see Methods) to the IMI DIRECT dataset after it had been pre-processed to reduce the number of variables down to a manageable size of 260 variables in 3029 individuals. We then used sub-networks (Markov blankets) to focus on variables of interest, such as body mass index (BMI), liver fat and T2D, replicating and confirming many previous findings. Previous studies focussed on T2D [ 8 , 9 ] have used colocalization analysis [ 10 ] of molecular quantitative trait loci (QTL) to identify a number of potential “effector” variables, including transcripts (gene expression) and protein levels, that may explain the genetic associations seen with T2D. Mahajan et al. (2022) [ 9 ] identified 97 candidate effector genes, including ADCY5 and TCF7L2, based on circulating plasma proteins (pQTL) and gene expression (eQTL) in diverse tissues, while Vi nuela et al. (2020) [ 8 ] focussed on expression in human pancreatic islets and showed colocalization between genetic variants influencing T2D or glycemic traits and 47 islet cis-eQTL, including ADCY5, DGKB and TCF7L2. Colocalization indicates that the two traits in question most likely share a causal genetic factor but does not identify the directionality of the relationship; it is possible that either trait acts as a mediator for the other, or that the genetic variant has independent horizontal pleiotropic effects on both traits. Therefore, Vi nuela et al. (2020) [ 8 ] recommend regarding the genes highlighted by coincident GWAS and eQTL signals as candidate effector transcripts that should be further interrogated through experimental approaches that directly test for causality. In their own analysis of the IMI DIRECT dataset, Brown et al. (2023) [ 4 ] used Bayesian networks applied to triplets of variables (one genetic variant and two molecular phenotypes) to infer the direction of causality, but they did not expand this investigation to include clinical variables such as T2D. The IMI DIRECT (DIabetes REsearCh on patient straTification) consortium [ 1 , 2 ] was set up to gather a wide range of data at different study centres throughout Northern Europe in order to study different aspects of type 2 diabetes (T2D). Findings from this dataset have been on topics as diverse as associations with liver fat [ 3 ] and inferring regulatory networks [ 4 ]. Here, rather than focussing on a specific question related to T2D, we perform exploratory analysis of the dataset to identify possible causal relationships between variables using Bayesian network (BN) methodology as implemented in our own BayesNetty software [ 5 , 6 ]. BNs can be used to infer possible causal relationships between variables based on their conditional dependencies and independencies, which can be particularly useful in complex biological scenarios with many measured variables. Understanding the relationships between measured biological and clinical variables can inform about underlying biological mechanisms, which may ultimately have clinical implications. The main advantage of our approach is that it can handle mixed continuous/discrete data with missing values, which is vital for these types of multi-omics datasets, as such studies often contain a considerable amount of missing data, with no individuals having complete data for every variable. Our approach also leverages the information provided by genetic variables (which can be considered as causal anchors) to help better resolve the direction of non-genetic edges when fitting the BN, in a conceptually similar (but complementary) approach to Mendelian Randomization [ 5 , 7 ]. Results Markov blanket for Liver Fat recapitulates previous findings As a proof-of-principle, we start by investigating a network centred on liver fat, as previously studied by Atabaki et al. [3] using the same IMI DIRECT dataset. The focus of the analyses by Atabaki et al. [3] was to better understand possible relationships of liver fat with T2D and non-alcoholic fatty liver disease (NAFLD), which frequently co-occur. Atabaki et al. [3] also used BN methodology for their analyses focusing on the complete data for their variables of interest including clinical and proteomic measurements (331 individuals with T2D and 964 free from diabetes). In addition, Atabaki et al. conducted two-sample bidirectional Mendelian Randomization [15] on some of the edges identified by BN, with genetic instruments leveraged from publicly available sources [3]. Because standard BN implementations usually require observations without missing data, which can be costly and difficult to produce (particularly in the context of clinical data, including data from Electronic Health Records), methods that can cope with missing data like BayesNetty offer a considerable advantage. Incorporation of genetic variables such as allele scores (as is routinely done in BayesNetty) can also help better resolve the direction of edges when fitting the BN. S1 Fig shows the Markov blanket for liver fat derived from Fig 1, which incorporates data from all 3029 individuals available and makes use of genetic variables in the form of allele scores (as outlined in the Methods). The results of Atabaki et al. [3] had suggested that a higher insulin secretion rate (Basal ISR) and excess visceral fat (VAT) accumulation were the most likely clinical factors in the dataset to cause liver fat accumulation and therefore NAFLD. Our selected network for liver fat (S1 Fig, S3 Table) found strong evidence of a causal relationship from VAT to liver fat, with an edge strength value of 0.99 and a direction value of 0.86. There is also strong evidence of a causal relationship of Basal ISR to liver fat with strength and direction values of 0.86 and 0.95 respectively. Our network identifies both liver iron and VAT to have centre as a parent variable. This node encodes the centre of origin of the samples, thus identifying differences in the measurement of these variables between study centres, but also identifying differences in BMI, as some centres provided samples from individuals with higher values, which introduces a correlation between BMI and centre [1, 2]. We also observed total abdominal adipose tissue (TAAT) and abdominal subcutaneous adipose tissue (ASAT) as causal variables on liver fat (S3 Table), providing an additional causal path from BMI to liver fat. This causal path (Centre BMI TAAT VAT Liver Fat) is not directly visible in S1 Fig, as BMI does not form part of the Markov blanket for liver fat (which is formally defined as the variable of interest and all parent variables, child variables and variables that are also parents of the child variables [13]), however it can be deduced from the connections from Fig 1 listed in S1 Table. In summary, BayesNetty was able to provide similar results to those from traditional BN methods, with the added advantage of allowing the inclusion of data (including genetic data) from all participants, even in datasets that include missing data. Type 2 Diabetes Given that the IMI DIRECT cohort was developed to study T2D, including both individuals diagnosed with T2D and those at risk of diabetes, we next evaluated a Markov blanket for T2D (Fig 2). The aim of this investigation was to see to what extent known relationships with T2D could be recapitulated, along with uncovering any novel or unexpected relationships. There were several incoming arrows for the T2D variable (see a full list of all edges in S4 Table). Not unexpectedly, fasting and mean glucose were found as parent (i.e. potential causal) variables for T2D, as T2D is defined by the glucose level in the blood being too high. The variable coding for Impaired Glucose Regulation (IGR) was also found to be a parent variable for T2D. People with IGR have a high blood-glucose level but not high enough to be diagnosed with T2D; this is sometimes referred to as pre-diabetes and is known to be a precursor to a T2D diagnosis [16]. It is known that males have a higher incidence rate of T2D than females and this is shown by the sex variable also being a parent variable of T2D [17]. However, our network suggested this effect was both direct and indirect by also modulating levels of expression for FADS1. The level of expression for FADS2 was also included in the network, identifying the well known involvement of fatty acid levels and the genomic region around those two genes on T2D [18, 19]. Moreover, the network included the expression of MYRF. A previous study [20] found a variant in MYRF (upstream of FADS2) to be associated with lower LysoPC 20:2 and increased risk of T2D. In addition, a multimorbidity study aiming to identify genes acting on multiple diseases was able to identify a cluster of genes including MYRF and the FADS1-FADS2-FADS3 region to be involved in multiple traits such as T2D, coronary artery, BMI and cholesterol among others [21]. Finally, it is perhaps less well known that height has also been associated with T2D, where taller people have been found to have a lower incidence rate of T2D [22–24], perhaps explaining the identification of height as a parent variable of T2D. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 2. Markov Blanket of type 2 diabetes diagnosis taken from the average BN constructed using imputed data of all variables with strength threshold 0.5. Edges are labelled with the probability that they exist (strength), and, in brackets, the probability that they exist in the shown direction, given that they exist (direction). The thickness of the edges is proportional to the edge strength. The nodes are coloured as follows: red are metabolites; blue (with gene name) are proteins; purple (with gene name) are gene expression measurements; amber are clinical variables; green (prefixed with AS) are allele scores. https://doi.org/10.1371/journal.pgen.1011776.g002 Our network for T2D revealed only a single outgoing edge, suggesting a potential causal influence of T2D on HLA-DRB5 gene expression. While HLA-DRB5 has been linked to type 1 diabetes (T1D) [25], its involvement in T2D remains unclear, with limited evidence supporting an association between the two [26], presumably in the direction of the gene causing the disease rather than vice versa. We observed a potential causal connection between the abundance of the plasma protein KCNQ1 and the expression of HLA-DRB5. The region around the KCNQ1 gene is well known for the association to T2D [27], pointing to the gene as a likely candidate gene mediating the activity of this locus. However, our network does not report a direct connection between KCNQ1 and T2D, and the connection through HLA-DRB5 involves an arrow in the “wrong” direction for suggesting a causal effect of gene on disease (going from T2D towards HLA-DRB5 gene expression rather than vice versa). Overall, these relationships uncovered between HLA-DRB5, KCNQ1 and T2D require further investigation and validation, perhaps using different (more targeted) types of data. It is of interest to investigate to what extent the relationships uncovered by our Bayesian Network approach are supported by evidence from Mendelian Randomization (MR). This is complicated by the fact that many of the relationships we detect are pleoitropic, with the same genetic variant(s) operating through multiple exposures, and with exposures operating through effects on one another. For example, in Fig 2, FADS1 expression is inferred to operate both directly on T2D and indirectly through FADS2, while FADS2 is inferred to operate both directly on T2D and indirectly through MYRF. These types of complicated dependencies between variables causes problems for standard MR approaches[28], which generally require a single route from each exposure to an outcome. This necessitates the use of more sophisticated approaches. We therefore used multivariable MR [29], as well as a recently proposed method, MrDAG [28], to investigate the inferred relationships between FADS1, FADS2, MYRF and T2D, taking advantage of the ability of these methods to operate on summary statistics from large-scale studies, rather than requiring individual-level data. We used genome-wide T2D summary statistics from a large-scale study of 180,834 affected individuals and 1,159,055 controls [9] and downloaded summary statistics for cis-expression quantitative trait loci (cis-eQTL) associations for FADS1, FADS2 and MYRF from the eQTLGen consortium [30]. Filtering to only include independent genetic instruments (not in linkage disequilibrium with one another) identified four SNPs that showed significant association with (that thus could be used as instruments for) gene expression, but only one of these (rs198462) also showed significant association with T2D (Table 2. This in itself suggests that MR is unlikely to provide much evidence for causal effects of gene expression on T2D, as the most “basic” implementation of MR essentially boils down to testing for association between the instrument and the outcome[31]. Consistent with this expectation, multivariable MR using the multivariable inverse-variance weighted method implemented within the R package MendelianRandomization provided no significant evidence of causal effects of gene expression on T2D for FADS1, FADS2 and MYRF (p-values 0.610, 0.381, 0.232 respectively). Interestingly, univariate analysis using the inverse-variance weighted method did show some evidence for a causal effect of MYRF gene expression on T2D (p-value 0.012), although no significant effects were seen at FADS1 and FADS2 (p-values 0.292 and 0.550 respectively). This discrepancy between the multivariable and univariate MR analyses can perhaps be attributed to the fact that the variables of interest may indeed be involved in a complicated network of mutual relationships, of which multivariable MR method only provides evidence of effects over and above the effects that are already included in the model—in the case of causal effects between risk factors, estimates represent the direct causal effect of each risk factor on the outcome by a pathway that is not operating via the other risk factors[29]. PPT PowerPoint slide PNG larger image TIFF original image Download: Table 2. Association test p-values from publicly available summary statistics for association between genetic instruments and gene expression or T2D. https://doi.org/10.1371/journal.pgen.1011776.t002 Analysis using MrDAG also showed little evidence for relationships between FADS1, FADS2, MYRF (considered as exposures) and T2D (considered as the outcome). The only directed edge identified by MrDAG with an estimated probability larger than 0.01 was the edge between FADS1 and FADS2, for which the direction could not be determined (FADS1 to FADS2 had edge probability 0.496 while FADS2 to FADS1 had edge probability 0.486). The various edges between these particular variables implied by our BN analysis of the IMI DIRECT dataset do not, therefore, seem to be recapitulated in MR analysis of summary statistics from large-scale GWAS. It is possible that there are some unique features of the individual-level IMI DIRECT data that are being captured through our BN approach. Alternatively, it may be that the relationships uncovered in the IMI DIRECT dataset are simply false positives. Ultimately, all of these methods (both the network-based methods and the more traditional MR type methods) are perhaps best considered as exploratory analysis tools, generating putative causal relationships between variables that ideally need further investigation/verification by other means (e.g. experimental laboratory work). Body Mass Index Next we explore a Markov blanket for body mass index (BMI) (Fig 3, S5 Table). BMI is defined as weight divided by height squared, so strong associations with both weight and height are expected. While the direction value of 0.5 from weight to BMI does not indicate a clear causal relationship, the direction value of 0.99 from height to BMI suggests a potential causal influence. However, this may partly reflect the mathematical structure of BMI rather than a true causal effect. Despite BMI being directly calculated from weight and height, a causal relationship could still exist if these variables co-vary in specific ways across individuals. For instance, previous studies have shown that taller children tend to have higher BMI, while in adults height is inversely associated with BMI—shorter individuals tend to have higher BMI [32]—a pattern that may be captured in the network analysis. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 3. Markov Blanket of BMI. All edges and nodes show a Markov Blanket of BMI taken from the average BN constructed using imputed data of all variables with strength threshold 0.5. Edges and nodes that are not faded show a Markov Blanket of BMI from the average BN with a strength threshold of 0.85 applied instead of 0.5. The thickness of the edges is proportional to the edge strength. Non-faded edges are highlighted in black and labelled in red with the probability that they exist (strength), and, in brackets, the probability that they exist in the shown direction, given that they exist (direction). Nodes are coloured as follows: red are metabolites; blue are proteins; purple are gene expression measurements; amber are clinical variables; green are allele scores. https://doi.org/10.1371/journal.pgen.1011776.g003 The network suggests that BMI is causal on total abdominal adipose tissue (TAAT) and on abdominal subcutaneous adipose tissue (ASAT), which are essentially measurements of fat around the abdomen. It is generally accepted that obesity causes an increase in these kinds of fat (as discussed, for example, by Verd et al. [33]). A variable for the Stumvoll index is shown; this is designed to measure insulin sensitivity, that is, how sensitive the body is to the effects of insulin, and thus how able to lower blood glucose levels. As BMI is one of the variables used to define the index, a causal relationship from BMI would be expected. However, this connection may also be identifying the known role of BMI on insulin sensitivity. There is also a suggestion of a causal relationship between BMI and insulin from the oral glucose tolerance test (OGTT), which has previously been reported [34]. Overall, the BMI centred network identified the intricate casual relationship between BMI and insulin sensitivity and adipose accumulation, but, with the exception of a genetic variable, the network did not involve any molecular phenotypes. Body Mass Index and Type 2 Diabetes As there is considerable evidence indicating that obesity is a leading cause of some types of T2D, we attempted to find evidence in our average BN of the possible causal path from BMI to T2D. The exact causal mechanisms linking BMI and T2D are not fully understood and, from this dataset, there was no strong evidence linking them directly or indirectly. There is, however, some weak evidence that there is a causal path via fasting glucose and perhaps mean glucose, both of which appear in the Markov blankets for T2D (Fig 2) and BMI (Fig 3). As well as a direct link between fasting glucose and T2D, Fig 2 also shows an indirect link via IGR. We therefore examined the sub-network comprised of T2D, BMI, fasting glucose, mean glucose, IGR and all edges between them (Fig 4, S6 Table). There have been some studies linking BMI to fasting glucose [35], so, despite the weakness of the edges to/from fasting glucose (0.43) and mean glucose (0.45), these may represent genuine relationships. The direction value of BMI to mean glucose is around 0.5, so there is no strong evidence for the directionality of this edge, while the edge from BMI to fasting glucose has a direction value of 0.68, providing some evidence that the direction of causality is in the direction shown. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 4. Sub-network taken from the average BN constructed using imputed data of all variables consisting of variables of interest with respect to T2D and BMI. Edges are labelled with the probability that they exist (strength), and, in brackets, the probability that they exist in the shown direction, given that they exist (direction). The thickness of the edges is proportional to the edge strength. The nodes are coloured as follows: amber are clinical variables and green are allele scores. https://doi.org/10.1371/journal.pgen.1011776.g004 The reasons for the apparent weak evidence linking BMI to T2D in this dataset could be numerous. The T2D variable is already adequately described by variables other than BMI, so if there is a causal relationship from BMI to T2D in the dataset, it may well be captured via other variables. As the mechanisms underpinning the involvement of BMI in T2D are complex, it may not be captured in this dataset other than through multiple variables implicating insulin resistance and glucose management, with weaker or no connections between BMI and -cell function variables. Finally BMI is a constructed variable calculated from weight and height, designed to indicate excess body fat, and it does not take into account other factors implicated in the development of T2D including fat distribution or sex and ethnic differences. For this reason, the use of BMI as a measure of the risk of T2D and other diseases has long been criticised [36]. Ethnic—or other—heterogeneity between participants could potentially be another reason for the lack of strong evidence linking BMI to T2D in this dataset. Thus, BMI may not stand out as a direct biological variable in the causal path with T2D. [END] --- [1] Url: https://journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.1011776 Published and (C) by PLOS One Content appears here under this condition or license: Creative Commons - Attribution BY 4.0. via Magical.Fish Gopher News Feeds: gopher://magical.fish/1/feeds/news/plosone/