Boehringer Ingelheim Expands AI and Computational Science for Drug Discovery

by Grace Chen
Boehringer Ingelheim Expands AI and Computational Science for Drug Discovery

The developments aim to bridge the gap between massive real-world datasets and the simplified models required for scientific discovery and drug development.

Scientists have long grappled with the challenge of translating messy, real-world data into actionable equations. Whether dealing with complex biological signals or the progression of rare diseases, researchers face immense hurdles stemming from nonlinear behavior and high-dimensional data systems. Recent computational frameworks are attempting to bridge this gap by combining modern machine learning with century-old mathematical principles and massive data integration. Computational scientists, AI specialists, data scientists and bioinformaticians are at the heart of the new Computational Innovation function.

Transforming Real-World Motion Into Simple Equations at Duke University

Engineers at Duke University are utilizing artificial intelligence to convert complex, real-world motion into concise, manageable rules. Led by Boyuan Chen, director of the General Robotics Lab, along with lead author and PhD candidate Sam Moore, the team published their results in the journal npj Complexity. The newly developed AI framework studies time-series data, meaning measurements taken over time, and then produces compact equations that describe how a system changes across various domains, including weather patterns, electrical circuits, mechanical devices, and biological signals. The goal is not just prediction, but understanding.

“Scientific discovery has always depended on finding simplified representations of complicated processes,” said Chen, the Dickinson Family Assistant Professor of Mechanical Engineering and Materials Science at Duke. “We increasingly have the raw data needed to understand complex systems, but not the tools to turn that information into the kinds of simplified rules scientists rely on. Bridging that gap is essential.”

The Duke approach builds directly upon a mathematical concept introduced by mathematician Bernard Koopman in 1931, which demonstrated that a nonlinear system can sometimes be represented through a linear model, if you describe it in the right coordinates. Linear models are attractive because they let you do global analysis and use tools like spectral decomposition, which can reveal a system’s modes and stability. Traditional Koopman-style modeling often pushes you into a very large, even infinite, space of variables, leading to huge representations for nonlinear systems. Past work has represented the two-dimensional Duffing system with embeddings that reached 100 dimensions, and in some cases far more, with similar inflation appearing for the Van der Pol oscillator.

“The catch is scale. Koopman-style modeling often pushes you into a very large, even infinite, space of variables. That reality has fed a long-running problem in the field. Methods such as Dynamic Mode Decomposition and Extended DMD can be useful, but they often balloon into huge representations for nonlinear systems. Deep learning has also been used to find linear embeddings, yet many approaches still land on latent spaces far larger than the system you started with,” Chen shared with The Brighter Side of News.

To overcome this limitation, the Duke framework tries to keep the linear representation as small as possible while still predicting well over long time windows, avoiding redundancy and decreasing the risk of false modes and overfitting.

Tackling Intractable Diseases Through Computational Innovation

In the pharmaceutical sector, computational science is reshaping how researchers approach complex conditions. Boehringer Ingelheim—a global, privately-held pharmaceutical company, headquartered in Germany—has established a dedicated Computational Innovation function to embed data specialists across every stage of the drug discovery process, moving away from traditional project-by-project outsourcing.

The pharmaceutical industry faces stark efficiency challenges, with industry figures indicating that just one in seven candidate drugs makes it through clinical trials to reach patients, taking an average of ten years and costing more than US$1.3 billion. By integrating advanced computational tools, researchers aim to improve these odds.

“Using AI and other advanced computational tools, we can integrate and analyse datasets from patient samples, biobanks and trials to improve our understanding of disease mechanisms, which will potentially help identify new treatments and concepts,” says immunologist Lamine Mbow, Head of Global Discovery Research at Boehringer Ingelheim.

Uncovering Hidden Biological Links and Overcoming Scientific Bias

Recent breakthroughs highlight the power of this data-driven approach. Researchers at Boehringer Ingelheim recently made a fresh discovery by undertaking the largest meta-analysis in this field to date, focusing on idiopathic pulmonary fibrosis (IPF), a progressive and deadly rare lung disease whose underlying cause is unknown. By integrating results from several genome-wide association studies with whole-genome sequencing data to boost the effects of rare variant analysis, the team increased the statistical power of their analyses and revealed a previously unknown link to a gene involved in human disease.

Data Science Capabilities with Victoria Gamerman of Boehringer Ingelheim
Boehringer Ingelheim Expands AI and Computational Science for Drug Discovery
Photo: Nature

Jan-Nygaard Jensen, Global Head of Computational Innovation at Boehringer Ingelheim, points to advanced data scientist tools being on the brink of delivering major patient benefits by gaining deeper insights into human health to improve the ability to target the root cause of disease. Progress has accelerated in the past few years owing to the rapid growth in capabilities of foundational AI models delivering a wide variety of outputs, specifically large language models trained on biological data (bio-LLMs).

“From very early target identification efforts through to clinical research, our dedicated computational innovation teams are embedded across every step of the process to help maximize the prospects of providing new therapies for patients,” says Jensen.

As one example, Jensen’s team combines spatial transcriptomics data, which identify the genes active in specific tissues down to the level of the individual cell, with images that map gene expression.

“By using bio-LLMs,” he explains, “we can analyse and understand cell char…”

You may also like