Chris Bishop, is Microsoft Technical Fellow and Director of Microsoft Research AI for Science Microsoft Research Podcast
Bishop’s career began in physics, including a PhD in quantum field theory and work on nuclear fusion.
He transitioned to machine learning after being inspired by Geoff Hinton’s backpropagation paper, applying neural networks to fusion data at the JET experiment.
Simulation Paradigm: Using digital computers to solve complex equations (e.g., weather forecasting).
Data-Intensive Paradigm: Utilising large datasets and machine learning (e.g., particle physics at CERN).
Fifth Paradigm: Training machine learning systems using simulation data to create emulators, which are much faster than traditional simulations.
The Fifth Paradigm Explained
The fifth paradigm involves using simulators to generate synthetic training data for deep learning systems.
These trained systems act as “emulators,” providing faster predictions compared to traditional simulations.
Example: Predicting molecule properties by training a model on data from Schrödinger’s equation simulations.
This method can achieve three-to-four-order-of-magnitude acceleration, making it highly disruptive.
Identifying Fields for AI Assistance
The space of opportunity for AI in science is vast.
Factors include data availability (both experimental and synthetic), the potential for real-world impact at scale, and bottlenecks in existing processes.
Microsoft focuses on the molecular level, given the enormous space of potential molecules for drugs, materials, and more.
Impact on Scientific Questions
AI is empowering scientists to think in more expansive and creative ways.
The acceleration provided by AI allows for broader exploration of possibilities.
Microsoft’s partnership with Novartis has demonstrated the impact of AI on drug discovery pipelines.
Microsoft collaborated with the Global Health Drug Discovery Institute (GHDDI) and the Gates Foundation on molecules to combat tuberculosis and coronaviruses.
They used a transformer model trained on SMILES strings (a way to represent molecules as text) to generate new molecule candidates.
The model was informed by the 3D geometry of the protein binding site.
A variational autoencoder was used to optimise existing molecules.
This approach led to the discovery of a molecule with more than two orders of magnitude stronger binding affinity to the target protein in just five months compared to several years using traditional methods.
The molecule was synthesised and tested at GHDDI, confirming its properties.
The team also considers practical aspects such as manufacturability, absorption, metabolism, excretion, and toxicity (ADMET).
AI is used to optimise these factors, focusing on bottlenecks in the drug discovery pipeline.
AI architectures, particularly those inspired by natural language processing (NLP) models, are demonstrating remarkable potential in handling vast datasets and intricate interactions within biological systems. These AI models are poised to accelerate scientific discovery by tackling challenges that traditional methods have struggled to address effectively.
Biological systems are characterized by intricate interplay of DNA, RNA, proteins, and small molecules within cells and organisms. Deciphering the complex interactions between these components is a grand challenge in biomedical science. Traditional analytical methods are not well-suited to the complexity of biological systems, leaving much more unknown than known about how cells, tissues, and bodies function at higher levels.
Modern AI architectures are very well suited to take advantage of biology’s massive datasets, and a new wave of AI models are dramatically accelerating the process of scientific discovery in biology. These models are starting with understanding protein and other cellular structures and the physical interactions between them, and will likely soon zoom out to help understand how cells, tissues, and bodies function at higher levels.
Models and specifics
It is important to keep in mind the distinction between static structural analysis and dynamic conformational analysis, which allows for molecules to change shape as they interact with one another.
Foundation models for biology, along the lines of the large language models (LLMs) that many are familiar with, are just now starting to be trained. Different models are often used for predicting shapes versus predicting sequences, which are two sides of the same coin but are currently modelled separately.
AlphaFold 2, developed by DeepMind, marked a significant breakthrough in protein structure prediction. This AI model can predict a static 3D structure from a protein sequence, providing confidence scores for the predicted structure. However, AlphaFold 2 has limitations in capturing protein dynamics, highlighting the need for models that can predict multiple conformations and their transitions.
Distributional Graph Former, building upon AlphaFold 2’s architecture, predicts ensembles of protein structures and transition pathways. AlphaFlow, a diffusion model trained on molecular dynamics simulation data, generates multiple protein conformations. These advancements demonstrate the potential of AI in capturing the dynamic nature of proteins.
Protein Language Models
Protein language models, built on architectures like BERT (Bidirectional Encoder Representations from Transformers), have emerged as powerful tools for predicting protein structures. ESM-Fold, for example, uses a masked language modelling objective to learn complex patterns and relationships within protein sequences, leading to accurate structural predictions without relying on physics-based simulations.
These models leverage multiple sequence alignments (MSAs) to provide evolutionary information that aids in structure prediction. Attention maps in protein language models correlate with contact maps, representing physical contacts between amino acids, further emphasizing the model’s ability to learn inherent structural information.
Training models to predict binding affinity and protein interactions is a challenging task. Overfitting is a common issue, where models perform well on training data but fail to generalize to new data. Proper data splitting based on sequence and structural similarity is crucial to ensure the model’s ability to generalize to unseen data. This is particularly challenging for protein interaction models, where similar sequences might be present in both training and testing sets, leading to overfitting and poor generalization.
Protein Complexes
RoseTTAFold All Atom extends the capabilities of AlphaFold Multimer to predict the structure of complexes containing proteins with small molecules, nucleic acids, and metals. This represents a significant step forward in modelling interactions beyond protein-protein interactions.
Generative Models for Molecule Design
Generative models based on diffusion and flow-matching approaches enable fine-grained control over the generation of molecules with specific properties. ProteinDT and MoleculeSTM are examples of text-conditioned generative models that allow users to provide natural language prompts to generate molecules with desired properties.
RF Diffusion, a diffusion model built on the RoseTTAFold backbone, offers powerful functionalities for protein engineering. It enables unconditional generation of novel proteins, binder design for high affinity and specificity, partial diffusion for refining existing structures, motif scaffolding for combining functional motifs, symmetric generation of protein complexes, and fold conditioning for generating proteins with specific tertiary structures.
Complementary models like Ligand and PNN (Protein MPNN) are essential for designing amino acid sequences that fold into the desired 3D structures generated by RF Diffusion.
Practical Applications and Workflows
These advanced AI models can be integrated into practical workflows for drug discovery and protein engineering. For example, to design a protein inhibitor, the workflow may involve identifying a problematic protein-protein interaction, extracting the interaction motif using RF Diffusion, scaffolding a new protein structure incorporating the motif, optimizing the structure using Partial Diffusion and Ligand and PNN, and validating the interaction using AlphaFold Multimer.
The speed and efficiency of these models allow for rapid iteration and generation of many potential candidates with high accuracy and effectiveness. Hundreds or thousands of backbones can be designed, and hundreds or thousands of sequences can be designed for each backbone using Ligand and PNN in a matter of minutes. When synthesized and tested, these designed proteins often demonstrate high thermal stability, specificity, and binding affinity.
Bottlenecks in Drug Discovery
Despite the remarkable advancements, there are challenges and bottlenecks in the adoption of AI models in biological research. Target identification remains a significant challenge in drug discovery. The rapid evolution of the field and technical barriers, such as the requirement for programming skills and familiarity with computational environments, can hinder the widespread adoption of these models by researchers without a programming background. This is potentially an opportunity.
Agents in Biological Research
AI agents have the potential to transform biological research by automating tasks such as literature review, hypothesis generation, experimental design, and data analysis. Companies like Future House are developing AI agents that can identify potential drug targets and design experiments, significantly accelerating the process of discovery. These agents, powered by large language models (LLMs) and other AI technologies, can review thousands of research papers, develop targets or hypotheses to test, and even drive autonomous labs.
As these AI agents become more capable, they may play a crucial role in guiding research and helping humans navigate the complex landscape of biological data and interactions. The convergence of AI agents with specific tools for designing molecules, proteins, and nucleic acids could lead to rapid progress in solving challenging problems in biology and medicine.
The development and application of AI-driven biology raise important ethical considerations and concerns related to potential misuse and the need for responsible development. While the computational design of toxic molecules is just one step in a complicated process that requires synthesis and delivery, the increasing capabilities of AI agents and the potential for state-sponsored bad actors highlight the need for oversight and safety measures.
Robust safety protocols, regulations, and ethical frameworks are essential to guide the development and application of these technologies. International collaboration is crucial to address the global nature of biological threats and prevent the proliferation of dangerous technologies. Hiring capable individuals with strong moral grounding and good intentions in companies and organizations working on these technologies is also important.
State of the Art in Medicine:
AI surpasses average human performance: Since mid-2022, top AI systems outperform the average human on average white-collar tasks, signifying a significant milestone. (https://aiindex.stanford.edu/report/)
AI approaches expert-level performance: Current AI systems are nearing expert performance levels in routine medical tasks, like adhering to standard care procedures. (https://aiindex.stanford.edu/report/)
Major Innovations:
GPT evolution: In just five years, GPT models evolved from GPT-2 (2019) to GPT-4, which rivals human experts on challenging cognitive tests like the MMLU (measuring performance across various academic fields). (https://openai.com/research/gpt-4)
Medical licensing exam success: Google’s Med-PaLM AI system passed the US medical licensing exam, with Med-PaLM 2 achieving an impressive 86% score, close to expert-level performance. ([invalid URL removed])
AI outperforms humans in specific medical tasks: Studies show AI outperforming human doctors in medical question answering, differential diagnosis, and interpreting medical images, indicating a potential shift in medical practice.
Virtual tissue staining: AI enables real-time “virtual tissue staining” using intraoperative imaging, significantly faster than traditional biopsy and staining methods, improving surgical precision. ([invalid URL removed])
“Mind reading” through fMRI decoding: AI can reconstruct images viewed by a person during an fMRI scan, demonstrating progress in decoding brain activity and reconstructing visual experiences. (https://www.science.org/doi/10.1126/science.adi1763)
AI impact on different job sectors: Contrary to expectations, current AI primarily impacts high-wage, high-skill jobs (e.g., doctors, lawyers) rather than low-wage, low-skill jobs, reflecting a shift in automation trends. (https://www.nber.org/papers/w31608)
AI nursing: Companies like Hippocratic AI offer AI-powered nursing assistants that perform follow-up tasks, demonstrating a shift towards AI directly competing in the labour market at an hourly rate. (https://www.hippocraticai.com/)
The (near) Future
The integration of AI in biological research holds immense potential for advancing scientific discovery, improving human health, extending lifespan, and enhancing quality of life. As these technologies continue to evolve, they may lead to a paradigm shift in how we think about health, longevity, and our relationship with the environment.
The ability to cure a wide range of diseases and significantly extend human lifespan could potentially lead to a shift in human consciousness, prompting a deeper appreciation for life, health, and interconnectedness.