Researchers from the Complex Disease Genomics Unit at the Sant Pau Research Institute and the Biomedical Research Networking Center for Rare Diseases have developed an artificial intelligence–based tool that integrates clinical, genetic, and transcriptomic information to identify signals associated with thrombosis, according to a July 6 announcement. The results were published in the Journal of Thrombosis and Haemostasis and reveal hundreds of molecular signals linked to thrombosis while improving characterization of people with different risk profiles.
Venous thrombosis is among the most common cardiovascular diseases and a leading cause of morbidity and mortality. While several factors are known to increase risk, some cases occur without clear triggers—a form known as idiopathic venous thromboembolism—which complicates identification of those at higher predisposition. Previous studies indicate that more than 60% of individual variability in venous thrombosis risk may be influenced by genetics, but current hereditary factors do not fully explain disease development.
To address this gap, researchers analyzed data from 790 individuals in families with a history of venous thromboembolic disease—including 70 who had experienced idiopathic events—using information from the GAIT2 family cohort. They integrated clinical and genetic variables with expression profiles from nearly 13,000 genes to build models identifying patterns associated with disease risk beyond conventional approaches.
"The study's main contribution is not only the identification of new genes associated with thrombosis, but also the demonstration that integrating thousands of biological variables makes it possible to describe risk profiles far more accurately than when traditional factors are analyzed in isolation," said Dr. José Manuel Soria, director of the Complex Disease Genomics Unit at IR Sant Pau and co-senior author.
The research applied machine-learning algorithms capable of analyzing thousands of biological variables simultaneously. Key predictors included established markers such as von Willebrand factor levels, body mass index, age, certain ABO system variants—and identified 494 genes whose activity distinguished those who had experienced thrombosis from those who had not. Many long noncoding RNAs were also highlighted as regulatory molecules previously understudied in this context.
Dr. Pol Ezquerra, first author and researcher at IR Sant Pau, said: "The incorporation of transcriptomic data allowed us to identify disease-associated signals that could not be detected using conventional approaches. This demonstrates the potential of combining artificial intelligence and gene expression to achieve a more precise characterization of patients." The team created a molecular signature based on combined clinical, genetic, and transcriptomic variables—allowing them to generate similarity scores measuring how closely an individual's profile matches those who have already experienced thrombotic events.
Adding gene-expression data refined participant classification: among people without prior history classified as high-risk by traditional models (43%), only 23% remained so after including transcriptomic information; identification accuracy for individuals with prior events increased from 70% to 74%. Signals related to cardiovascular and renal processes previously linked with thrombotic risk were also identified—reinforcing findings' biological relevance.
Authors emphasized further validation is needed before direct clinical application but believe their approach marks progress toward more accurate models for personalized medicine in thrombotic risk stratification. "Our work demonstrates the value of integrating clinical, genetic, and transcriptomic information... In the future these types of strategies could help identify people with a high-risk profile more accurately," Soria concluded.