Computational biology is the application of computational methods, mathematical modelling, statistical inference, and data analysis to understand biological systems and processes at every scale of organisation, from individual molecules and cells through tissues, organisms, populations, and ecosystems. It develops algorithms and models to interpret molecular, cellular, and organismal data, spanning sequence analysis, structural prediction, systems modelling, and simulation. The discipline increasingly relies on machine learning and deep learning to extract patterns from large and complex biological datasets, and is the parent domain of bioinformatics.
Semantic Classification
Content
Compositional Relationships (Components)
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:hasPart ai:Bioinformatics))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:hasPart ai:SystemsBiology))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:hasPart ai:StructuralBiology))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:hasPart ai:PopulationGenetics))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:hasPart ai:EvolutionaryBiology))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:hasPart ai:MolecularDynamics))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:hasPart ai:Phylogenetics))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:hasPart ai:Epigenomics))
Dependency Relationships
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:requires ai:BigData))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:requires ai:GPU))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:requires ai:HighPerformanceComputing))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:requires ai:Algorithm))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:dependsOn ai:Statistics))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:dependsOn ai:Biology))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:dependsOn ai:MachineLearning))
Capability Relationships
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:enables ai:DrugDiscovery))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:enables ai:ProteinStructurePrediction))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:enables ai:PrecisionMedicine))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:enables ai:SyntheticBiology))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:enables ai:CancerGenomics))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:supports ai:Healthcare))
Implementation Relationships
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:uses ai:DeepLearning))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:uses ai:GraphNeuralNetwork))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:uses ai:TransformerArchitecture))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:uses ai:LargeLanguageModels))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:uses ai:BayesianInference))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:uses ai:AlphaFold))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:implements ai:SequenceAlignment))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:implements ai:MolecularDynamics))
Reduction Relationships
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:reducesTo ai:DataScience))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:reducesTo ai:ScientificComputing))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:supports ai:Genomics))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:supports ai:SystemsBiology))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:relatedTo ai:NaturalLanguageProcessing))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:relatedTo ai:KnowledgeGraph))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:bridgesTo ai:ArtificialIntelligence))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:uses ai:HiddenMarkovModel))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:uses ai:VariationalAutoencoder))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:uses ai:ReinforcementLearning))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:enables ai:VaccineDesign))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:enables ai:PandemicPreparedness))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:supports ai:PublicHealth))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:depends ai:CloudComputing))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:hasPart ai:Immunoinformatics))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:implements ai:MarkovChainMonteCarlo))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:implements ai:VariationalInference))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:relatedTo ai:HumanCellAtlas))
SubClassOf(ai:ComputationalBiology
ObjectSomeValuesFrom(ai:supports ai:ClinicalGenomics))
About
- Computational biology is one of the most rapidly expanding scientific disciplines of the early twenty-first century, catalysed by the concurrent explosion in biological data volumes and the maturation of Machine Learning methods capable of extracting biological meaning from high-dimensional observations. The discipline sits at the confluence of Biology, Data Science, Statistics, and Scientific Computing, and increasingly overlaps with the frontier of Artificial Intelligence research. Where traditional experimental biology proceeds hypothesis-first — proposing a mechanism and designing experiments to test it — computational biology operates bidirectionally: it can generate hypothesis-free, data-driven discoveries at scales impossible for bench science alone, while simultaneously providing rigorous mechanistic models that give quantitative bite to classical biological theories such as population genetics, enzyme kinetics, and gene regulatory logic.
- The intellectual lineage of computational biology stretches from the 1953 discovery of the DNA double helix through the 1970s development of the Smith-Waterman Sequence Alignment algorithm, the late 1980s construction of the first molecular sequence databases, and the landmark 2001 draft sequences of the human genome, to the 2021 triumph of AlphaFold2 at the Critical Assessment of Protein Structure Prediction (CASP14) competition, where the Deep Learning model achieved median GDT_TS scores above 92 across all target categories — effectively solving a 50-year grand challenge in structural Biology. AlphaFold3, released by DeepMind in 2024, extended this capability to predict the structures of complexes involving proteins, nucleic acids, small molecules, and modified residues simultaneously, transforming structure prediction into a differentiable framework for the full chemical diversity of the cell. These breakthroughs exemplify the broader pattern in which Deep Learning has displaced classical bioinformatics pipelines across tasks including splice-site detection, transcription factor binding prediction, variant effect scoring, protein–protein interaction modelling, and single-cell trajectory inference.
- The modern computational biology ecosystem encompasses three interacting layers: (1) data infrastructure — biological databases such as UniProt, the Protein Data Bank (PDB), Ensembl, and the NCBI databases that store and version biological knowledge; (2) algorithmic pipelines — workflow managers (Nextflow, Snakemake, WDL) that orchestrate end-to-end analyses from raw sequencing reads to biological interpretation; and (3) model layers — from classical statistical models (hidden Markov models, Bayesian networks) through convolutional and recurrent Neural Network architectures to attention-based Transformer Architecture foundation models (ESMFold, Geneformer, Evo, scGPT) pretrained on hundreds of millions of sequences or millions of single-cell expression profiles. These foundation models, analogous to Large Language Models in Natural Language Processing, enable zero-shot generalisation to novel prediction tasks and are increasingly used in Drug Discovery pipelines for hit identification, lead optimisation, and ADMET property prediction at unprecedented throughput.
Major Subfields Overview
- Bioinformatics
- Sequence analysis: genome assembly, annotation, variant calling, comparative genomics
- Database management: biological databases (UniProt, GenBank, PDB, Ensembl) and their curation pipelines
- Workflow engineering: pipeline orchestration (Nextflow, Snakemake, CWL), reproducibility, containerisation (Docker, Singularity) on High Performance Computing clusters
- NGS analysis: short-read (Illumina), long-read (PacBio, Oxford Nanopore), and hybrid assembly strategies
- Structural Biology
- Experimental structure determination: X-ray crystallography, Cryo-EM, NMR spectroscopy
- Computational structure prediction: AlphaFold2/3, RoseTTAFold, ESMFold, RFdiffusion for de novo design
- Molecular Dynamics simulation: classical and machine-learned force fields, free energy perturbation (FEP) for binding affinity estimation
- Drug–target docking and virtual screening: AutoDock, Glide, DiffDock, generative structure-based design
- Systems Biology
- Gene regulatory network inference: Bayesian Inference networks, mutual information, causal discovery
- Metabolic network analysis: genome-scale metabolic reconstructions (GENRE), flux balance analysis (FBA), elementary flux modes
- Stochastic simulation: Gillespie algorithm for exact stochastic kinetics, tau-leaping approximations, moment closure methods
- Boolean and logical network models: attractor analysis for cell fate determination and differentiation
- Population Genetics and Evolutionary Biology
- Coalescent theory: modelling the genealogical process of random mating populations backwards in time
- Demographic inference: PSMC, MSMC, and simulation-based inference (ABC) for population size history
- Positive and purifying selection detection: FST statistics, dN/dS ratios, haplotype-based selection scans
- Admixture modelling: STRUCTURE, ADMIXTURE, PCA-based ancestry inference for population stratification correction in GWAS
- Transcriptomics
- Bulk RNA-seq: read alignment (STAR, HISAT2), quantification (Salmon, kallisto), differential expression (DESeq2, edgeR, limma)
- Single-cell RNA-seq: cell barcode demultiplexing, count matrix generation (STARsolo, CellRanger), clustering (Louvain, Leiden), trajectory inference (Monocle, scVelo, CellRank)
- Spatial transcriptomics: 10x Visium, Slide-seq, MERFISH, STARmap — mapping gene expression to tissue coordinates
- Proteomics
- Mass spectrometry: data-dependent acquisition (DDA) and data-independent acquisition (DIA) for quantitative proteomics
- Protein identification: database searching (MaxQuant, Mascot, MSFragger), de novo sequencing, spectral library matching
- Post-translational modification (PTM) analysis: phosphoproteomics, ubiquitinomics, glycoproteomics
- Structural proteomics: cross-linking mass spectrometry (XL-MS) for protein interaction mapping at peptide resolution
- Epigenomics
- Chromatin accessibility: ATAC-seq for open chromatin identification; FAIRE-seq, DNase-seq as predecessors
- Histone modification profiling: ChIP-seq for transcription factor binding and histone mark mapping
- DNA methylation: bisulphite sequencing (WGBS, RRBS), nanopore-based direct methylation calling
- Chromatin organisation: Hi-C, Micro-C, and ChIA-PET for 3D genome topology and enhancer-promoter contact mapping
- Metagenomics
- Taxonomic classification: k-mer-based (Kraken2, Bracken), read alignment (MetaPhlAn4), marker gene amplicon sequencing (16S rRNA)
- Metagenome-assembled genome (MAG) reconstruction: binning (MetaBAT2, CONCOCT), quality assessment (CheckM2), functional annotation
- Functional profiling: HUMAnN3 for pathway and gene family abundance; KEGG, COG databases as annotation references
- Viromics: phage discovery, virus–host interaction prediction, and antiviral target identification from metagenomic datasets
- Cancer Genomics
- Somatic mutation calling: single nucleotide variants (SNVs), indels, structural variants (SVs), copy number alterations (CNAs) from tumour-normal pairs
- Mutational signature analysis: COSMIC signatures, SigProfiler for aetiology inference (tobacco, UV, APOBEC, homologous recombination deficiency)
- Tumour heterogeneity: subclonal reconstruction (PHYLOGIC, MOBSTER), clonal dynamics, longitudinal evolution
- Cancer driver gene discovery: dN/dS models, MutSigCV, CDKN2A/EGFR/TP53 as canonical drivers across cancer types
Components / Architecture
- Sequence and Genomics Layer
- Sequence Alignment: pairwise and multiple sequence alignment via dynamic programming (Smith-Waterman, Needleman-Wunsch), heuristic methods (BLAST, DIAMOND), and Deep Learning-based alignment (AlignCLR)
- Genome assembly: overlap-layout-consensus and de Bruijn graph assemblers (SPAdes, Hifiasm, Verkko) for short-read, long-read, and Hi-C data
- Variant calling: genotyping SNPs, indels, CNVs, and structural variants from aligned short-reads (GATK HaplotypeCaller) or long reads (Clair3, DeepVariant)
- Cancer Genomics: somatic mutation detection, tumour phylogeny reconstruction, clonal evolution analysis
- Structural Layer
- Protein Structure Prediction: AlphaFold2/3 for monomers and complexes; RoseTTAFold, ESMFold for rapid inference; Cryo-EM density map fitting via ModelAngelo
- Molecular Dynamics: GROMACS, AMBER, NAMD; Machine Learning force fields (ANI, MACE, NequIP) for accurate atomistic simulation at reduced cost
- Drug–target docking: AutoDock Vina, Glide, and generative Deep Learning approaches (DiffDock, RFDiffusion) for structure-based Drug Discovery
- Systems and Networks Layer
- Systems Biology: ordinary differential equation (ODE) kinetic models, Boolean network models, stochastic simulation algorithms (Gillespie), constraint-based metabolic flux analysis (COBRA)
- Graph Neural Network models for protein–protein interaction networks, gene co-expression graphs, and metabolic network analysis using Knowledge Graph embeddings
- Single-cell multiomics integration: scRNA-seq, ATAC-seq, spatial transcriptomics processed by Seurat, Scanpy, Monocle; foundation models Geneformer and scGPT for cell-type annotation and perturbation prediction
- Evolutionary and Ecological Layer
- Phylogenetics: maximum likelihood (IQ-TREE, RAxML) and Bayesian (BEAST) phylogenetic inference for Evolutionary Biology and molecular clock dating
- Population Genetics: PSMC, MSMC, and simulation-based inference for demographic history; selection scan methods (SweeD, SweepFinder2)
- Metagenomics: taxonomic profiling (Kraken2, MetaPhlAn), metagenome-assembled genome reconstruction, functional annotation via HUMAnN
- Clinical and Translational Layer
- Precision Medicine: polygenic risk scoring (PRS), pharmacogenomics variant interpretation, clinical variant classification (ACMG/AMP criteria)
- Transcriptomics-based patient stratification: bulk RNA-seq differential expression (DESeq2, edgeR) and single-cell subtype deconvolution
- Epigenomics: ATAC-seq, ChIP-seq, bisulphite sequencing, Hi-C chromatin topology; links to Cancer Genomics driver gene identification
Mathematical and Algorithmic Foundations
- Computational biology draws from a diverse toolkit of mathematical frameworks, each matched to a different scale and type of biological question. Understanding the formal foundations clarifies why different methods are applicable in different contexts:
- Dynamic Programming for Sequence Alignment: the Smith-Waterman local alignment algorithm runs in O(mn) time and space for sequences of lengths m and n, using a recurrence relation H(i,j) = max(0, H(i-1,j-1) + s(a_i, b_j), H(i-1,j) - g, H(i,j-1) - g) where s is the substitution score and g is the gap penalty. While exact, it is prohibitive for database search; BLAST’s heuristic — seeding on high-scoring word matches and extending outward — achieves near-linear scalability at the cost of sensitivity. Multiple sequence alignment (MSA) is NP-hard in general (Li et al., 2000) but heuristic progressive methods (Clustal Omega, MAFFT) and consistency-based methods (T-Coffee) make large-scale alignment tractable. The Pfam database of protein families uses profile HMMs — stochastic finite automata with match, insert, and delete states — to represent position-specific residue preferences with exponentially better sensitivity than pairwise methods at detecting distantly homologous sequences.
- Bayesian Inference for Phylogenetics: maximum likelihood phylogenetics (RAxML, IQ-TREE) maximises P(alignment | tree, model) over all tree topologies and branch lengths under a substitution model (GTR+G for DNA, WAG/LG for protein). Bayesian phylogenetics (BEAST, MrBayes) treats tree topology and divergence times as random variables, sampling from the posterior P(tree, model | alignment) via Markov chain Monte Carlo (MCMC). Bayesian methods naturally accommodate rate variation, molecular clock assumptions, and fossil calibrations for absolute dating in Evolutionary Biology, and are the standard approach for SARS-CoV-2 variant phylodynamics and outbreak reconstruction.
- Ordinary Differential Equations in Systems Biology: the Hill equation (v = Vmax × [S]^n / (K_d^n + [S]^n)) models cooperative binding and sigmoidal dose-response; the Michaelis-Menten equation models enzyme kinetics; and the Goodwin oscillator (three coupled ODEs with negative feedback) is the minimal model of circadian rhythm generation. Boolean network models (Kauffman 1969) represent gene regulatory interactions as logical switches and simulate cell type attractors as fixed points of the synchronous update rule. Constraint-based metabolic flux balance analysis (FBA, Orth et al. 2010, Nature Biotechnology) uses linear programming to maximise cellular growth rate subject to stoichiometric constraints derived from genome-scale metabolic network reconstructions (GENRE), producing flux distributions that accurately predict gene essentiality and growth phenotypes in bacteria and yeast.
- Probabilistic Modelling for Single-Cell Data: the negative binomial distribution (NB) is the standard model for scRNA-seq count data, accommodating overdispersion not captured by the Poisson model; DESeq2 and edgeR use empirical Bayes shrinkage of NB dispersion estimates across genes. Variational autoencoders (VAE) — the foundation of scVI (Lopez et al., 2018, Nature Methods) — learn a low-dimensional latent representation of the scRNA-seq data distribution, enabling batch-corrected integration of multiple datasets and uncertainty-aware cell clustering. Optimal transport (Wasserstein distance) has been applied to multi-omic data integration (SCOT, Demetci et al.) and trajectory inference (WOT, Schiebinger et al., 2019, Cell), enabling mechanistic analysis of cellular differentiation without relying on RNA velocity assumptions.
- Attention Mechanisms in Biological Sequence Modelling: the core innovation of Transformer Architecture protein language models is the self-attention mechanism, which computes for each position i a weighted sum of all other positions j, with attention weight proportional to the compatibility of query vector q_i with key vector k_j: Attention(Q,K,V) = softmax(QK^T / sqrt(d_k))V. In protein language models, this allows the model to implicitly learn co-evolutionary constraints between residue pairs — the same information that explicit co-evolution methods (DCA, PseudoLikelihood maximisation) extract through statistical coupling analysis. AlphaFold2’s Evoformer goes further by representing both the MSA (row-wise and column-wise attention) and the pair representation (triangle update and triangle attention operations) in a joint, iteratively refined representation that explicitly encodes pairwise distance and orientation constraints before passing them to the structure module.
- Graph Neural Networks for Biological Networks: protein–protein interaction networks, gene co-expression networks, and metabolic networks are naturally represented as graphs, making Graph Neural Network architectures appropriate for learning molecular and cellular embeddings. Methods include Graph Convolutional Networks (GCN) for node classification, Graph Attention Networks (GAT) for weighted edge propagation, and message-passing neural networks (MPNN) for molecular property prediction. Open Targets and STRING databases provide large-scale biological Knowledge Graphs that ground Graph Neural Network training; Reinforcement Learning on graph-structured spaces enables de novo molecular design subject to structural constraints.
- Formal Genomic Statistics: population genetic inference from genomic data relies on classical statistical tools including: the Hardy-Weinberg equilibrium test for detecting genotyping error or selection; Fst (fixation index) for quantifying population differentiation; linkage disequilibrium (LD) patterns for detecting recombination and selection; coalescent-based likelihood methods (PSMC, SMC++) for inferring effective population size history from single genomes; and polygenic score (PRS) calculation as a linear combination of GWAS effect sizes. Bayesian model comparison (via thermodynamic integration or annealed importance sampling) enables formal comparison of competing demographic models without overfitting to the sample.
- Structural Bioinformatics Metrics: protein structure quality is assessed by Ramachandran plot validation, MolProbity score, real-space correlation to electron density (for Cryo-EM maps), and template modelling score (TM-score, ranging 0-1 where >0.5 implies same fold). RMSD (root mean square deviation of Cα positions after superposition) is the standard metric for structural comparison but is sensitive to domain movements; the GDT_TS (global distance test total score) and lDDT (local difference distance test) are more informative for assessing model accuracy relative to experimental structure in CASP competitions.
Use Cases / Major Families
- Protein Structure and Function
- AlphaFold3 and its successors enable virtual screening of entire proteomes against small-molecule libraries, compressing early-stage Drug Discovery from years to months. By 2025, over 200 pharmaceutical and biotech companies were integrating Protein Structure Prediction into their hit-identification pipelines.
- De novo protein design via RFDiffusion and ProteinMPNN generates novel enzymes, vaccines, and therapeutic proteins with desired structures, leveraging Reinforcement Learning and Deep Learning generative models.
- Genomics and Precision Medicine
- The 100,000 Genomes Project (Genomics England) returned clinically actionable findings for rare-disease patients; by 2026 the NHS Genomic Medicine Service genome sequencing programme had expanded to whole-genome sequencing as standard-of-care for rare diseases and cancer.
- Single-cell atlases (Human Cell Atlas, Tabula Sapiens) catalogue the transcriptional identity of every human cell type; computational biology provides the clustering, trajectory inference, and Transcriptomics integration methods that make these data interpretable.
- Drug Discovery Acceleration
- Machine Learning-driven virtual screening and ADMET prediction (Chemprop, DeepChem) have reduced the cost of early-stage hit identification, with AI-designed small molecules entering phase I clinical trials from 2023 onwards.
- Foundation models trained on protein and chemical structure space (ESM-3, Evo, MolGPT) enable multimodal Drug Discovery by jointly embedding sequence, structure, and function.
- Pandemic Preparedness and Pathogen Genomics
- Real-time Phylogenetics and variant tracking (Nextstrain) during COVID-19 demonstrated the operational value of computational biology pipelines for global public-health surveillance; the infrastructure is now applied to influenza, mpox, and emerging pathogen monitoring.
- Metagenomics-based environmental monitoring programmes detect novel pathogens in wastewater and wildlife without prior sequence knowledge, using de novo assembly followed by taxonomic classification.
- Phylodynamics: combining Phylogenetics and Population Genetics to jointly infer pathogen transmission dynamics, effective reproductive number (R_t), and migration patterns from genomic time series — a critical tool for public health outbreak response.
- Agricultural and Environmental Genomics
- Crop improvement: genomic selection (GS) using whole-genome SNP panels for breeding value estimation in wheat, maize, and rice; CRISPR-guided trait improvement informed by computational Population Genetics of domestication loci.
- Livestock genomics: genomic prediction of disease resistance, milk yield, and feed efficiency; the Roslin Institute (Edinburgh) leads in computational approaches to improving livestock breeds while reducing the environmental footprint of animal agriculture.
- Environmental monitoring: Metagenomics of soil, ocean, and freshwater microbiomes to detect ecological change, antibiotic resistance gene spread, and biodiversity shifts in response to climate change; the Earth Microbiome Project has characterised over 27,000 environmental samples and 300,000 distinct microbial 16S amplicon sequence variants.
- Immunoinformatics and Vaccine Design
- MHC/HLA epitope prediction: predicting peptide-MHC binding affinity (NetMHCpan, MHCflurry) for T-cell epitope identification in vaccine antigen design and personalised neoantigen immunotherapy for Cancer Genomics
- Antibody design: deep learning models for antibody structure prediction (ABodyBuilder2, IgFold), affinity maturation (DiffAb), and developability optimisation; BenevolentAI and Absci deploy these computationally at scale.
- Pan-genome analysis: constructing reference pan-genomes for pathogen populations (bacterial pan-genomes, SARS-CoV-2 pan-genome) to capture diversity not representable by any single reference genome.
Academic Context
- The intellectual history of computational biology is a story of successive technological discontinuities, each triggering a wave of new algorithmic methods. The field’s formal origins are traced to the mid-twentieth century: Francis Crick and James Watson’s 1953 description of the DNA double helix immediately raised the question of how sequence encodes structure and function — a question that was computationally unanswerable with the tools of the time. Margaret Dayhoff’s assembly of the first Atlas of Protein Sequence and Structure in 1965, and her derivation of PAM (Point Accepted Mutation) substitution matrices (1968, 1978) from aligned protein families, established the first mathematical framework for measuring evolutionary divergence between sequences and for scoring pairwise alignments — founding the discipline of protein sequence analysis. Needleman and Wunsch (1970) provided the first rigorous Dynamic Programming algorithm for global Sequence Alignment; Smith and Waterman (1981) extended this to local alignment, enabling detection of conserved domains within otherwise divergent sequences. These algorithms remain in use, implemented in BLAST (Altschul et al., 1990) — which heuristically approximates Smith-Waterman at orders of magnitude greater speed and whose 1997 Nature Genetics paper is among the most cited in the history of Biology.
- The 1990–2003 Human Genome Project was the defining institutional event of computational biology’s maturation as a discipline. The project required the development of genome assembly algorithms (whole-genome shotgun sequencing, Celera Genomics vs. the public consortium), gene prediction methods (ab initio prediction via generalised HMMs, homology-based annotation), comparative genomics tools (whole-genome alignment methods: MUMmer, BLAST-based chains), and large-scale database infrastructure (GenBank, Ensembl). The simultaneous development of the Affymetrix microarray platform (1994) and cDNA microarray technology (Pat Brown, Stanford, 1995) created the first high-throughput gene expression measurement platforms, generating a wave of statistical methodology for differential expression analysis (SAM, Tusher et al. 2001; limma, Smyth 2004; edgeR, Robinson et al. 2010; DESeq2, Love et al. 2014) that became standard practice in molecular Biology laboratories worldwide. Hidden Markov models, developed for speech recognition by Baum et al. in the 1960s and adapted for protein sequences by Krogh et al. (1994), became the dominant framework for protein family modelling (HMMER, Pfam), gene prediction (AUGUSTUS, GENSCAN), and comparative genomics. Bayesian networks for gene regulatory network inference (Friedman et al., 2000) and expectation-maximisation algorithms for latent-variable models (mixture models for expression clustering) extended the statistical toolkit.
- The second discontinuity was next-generation sequencing (NGS), beginning with the Illumina/Solexa platform (2006) and 454 pyrosequencing (2005), which reduced the cost of sequencing a human genome from 1,000 (achieved 2014, Illumina HiSeq X) and eventually to approximately $100 by 2026. This precipitous cost reduction transformed sequencing from a rare research activity into a routine clinical diagnostic tool, generating data volumes that demanded new algorithmic approaches: short-read assembly (Velvet, SOAPdenovo), short-read alignment (BWA, Bowtie), RNA-seq quantification (Cufflinks, kallisto, STAR), variant calling (GATK HaplotypeCaller, FreeBayes, DeepVariant), and eventually metagenomics and single-cell Transcriptomics. The introduction of Pacific Biosciences SMRT sequencing (2011) and Oxford Nanopore Technologies long-read sequencing (2014) further expanded the toolkit, enabling haplotype-resolved chromosome-scale genome assembly and direct detection of base modifications without bisulphite conversion. Long-read assemblers (Hifiasm, Verkko) have enabled the sequencing of complete telomere-to-telomere human genomes (T2T Consortium, 2022 — resolving the approximately 8% of the genome inaccessible to short-read sequencing), opening centromeres and segmental duplications to computational analysis for the first time.
- The third and current discontinuity is the Deep Learning revolution, catalysed within computational biology by the 2021 AlphaFold2 result. AlphaFold2 (Jumper et al., 2021, Nature) achieved a median GDT_TS score of 92.4 on CASP14 targets — effectively matching experimental crystallography accuracy for the first time in the 50-year history of the CASP competition. The model architecture combined multiple sequence alignment-based evolutionary features with an Evoformer attention module that computed pair-wise residue relationships and a structure module that iteratively refined backbone torsion angles. The subsequent AlphaFold Protein Structure Database, jointly released by DeepMind and EMBL-EBI in 2022, provided predicted structures for the entire UniProt proteome — over 200 million proteins — making structural predictions accessible without computational infrastructure for the first time. AlphaFold3 (Abramson et al., 2024, Nature) extended this capability to heteromolecular complexes including proteins, DNA, RNA, and small molecules, achieving 76.4% accuracy on protein-ligand docking benchmarks — a 50% improvement over physics-based methods. The architecture was fundamentally reconceived as a diffusion process, iteratively denoising atomic coordinates within a unified probabilistic framework, reflecting the convergence of computational biology with the generative Artificial Intelligence paradigm. Simultaneously, ESM-2 (Lin et al., 2023, Science) demonstrated that protein language models trained purely on sequence data — without evolutionary information from multiple sequence alignments — could achieve near-AlphaFold2 accuracy via ESMFold, opening a path to ultrafast structural inference using protein language models. ESM-3 (Hayes et al., 2024, EvolutionaryScale; Science 2024) scaled to 98 billion parameters and demonstrated generative design of novel functional proteins, including esmGFP — a fluorescent protein with only 58% sequence identity to any known fluorescent protein, representing the equivalent of hundreds of millions of years of divergent evolution from existing proteins. This aligns with the broader biological foundation model paradigm: just as Large Language Models learn general language understanding from text corpora, ESM-3 and related models learn general biological function from evolutionary sequence corpora.
- The emerging paradigm of biological foundation models in 2024-2026 has spawned a new generation of tools across every omic scale. Geneformer (Theodoris et al., 2023, Nature) pretrained a transformer on 30 million single-cell transcriptomes to learn gene regulatory network context embeddings, enabling zero-shot prediction of perturbation effects in cardiomyocyte disease models. scGPT (Cui et al., 2024, Nature Methods) demonstrated that a GPT-style autoregressive model pretrained on 33 million single cells could be fine-tuned for cell-type annotation, multi-omic integration, and perturbation response prediction. Evo (2024, Arc Institute) applied a hyena-based architecture (avoiding attention’s quadratic scaling) to model the full diversity of prokaryotic genomes at single-nucleotide resolution, enabling prediction of functional consequences of mutations from CRISPR screens without fine-tuning. Nicheformer and related models add spatial context to single-cell representations, capturing cell-cell communication and tissue architecture alongside transcriptional identity. These biological foundation models are directly analogous to Large Language Models in Natural Language Processing but operate over DNA, RNA, protein, and chromatin alphabets; they inherit the same pretraining / fine-tuning paradigm and the same emergent capabilities for zero-shot reasoning about functional relationships not present in the training data. Research is concentrated in groups at the Broad Institute (MIT/Harvard), the European Bioinformatics Institute (EMBL-EBI, Hinxton), the Wellcome Sanger Institute, Stanford’s Genome Technology Center, the Flatiron Institute Center for Computational Biology (New York), EvolutionaryScale (San Francisco), and the Arc Institute (Palo Alto).
Current Landscape (2026)
- As of mid-2026, computational biology occupies a pivotal position as the primary conduit through which Artificial Intelligence enters life-science and clinical practice. Several trends define the current landscape:
- AlphaFold3 and beyond: Released by DeepMind in May 2024, AlphaFold3 uses a diffusion-based architecture to predict the joint structure of protein complexes with DNA, RNA, and small molecules, achieving unprecedented accuracy across molecular interaction types. By 2026, AlphaFold3 is embedded in every major pharmaceutical company’s computational Drug Discovery platform, and several academic groups are developing differentiable extensions of the framework for molecular dynamics and free-energy estimation.
- Single-cell and spatial foundation models: Models such as Geneformer (2023), scGPT (2024), and the Universal Cell Embeddings project (2025) demonstrate that Transformer Architecture models pretrained on millions of single-cell expression profiles transfer effectively to cell-type classification, gene perturbation prediction, and drug response scoring. The Wellcome Sanger Institute renamed its Cellular Genetics programme to Cellular Genomics in 2025, reflecting the integration of single-cell genomics and spatial omics as core computational biology infrastructure.
- Biological large language models: ESM-3 (2024, EvolutionaryScale) and Evo (2024, Arc Institute) are protein and genome foundation models trained on hundreds of millions of sequences, enabling zero-shot functional annotation and multimodal co-generation of sequence, structure, and function. These models are analogous to Large Language Models but operate over biological alphabets.
- NHS Genomic Medicine Service: The NHS in England is the first national health service to implement whole-genome sequencing as a standard diagnostic for rare disease and cancer, generating computational biology workloads at population scale. The 2025 Genomics Futures workshop series (Wellcome Sanger Institute and Wellcome Trust) engaged genomics researchers, clinicians, and policy-makers in defining the computational infrastructure for the next decade of population genomics.
- Industry adoption: AstraZeneca, GSK, Exscientia, BenevolentAI, and Recursion Pharmaceuticals are embedding computational biology at every stage of their Drug Discovery pipelines, from target identification through clinical trial design, with AI-discovered candidate molecules entering phase II trials by 2025. In early 2026, GSK committed 400M, signalling that frontier Artificial Intelligence laboratories are making direct strategic bets on computational biology as a core application domain. By 2026, over 350 distinct biological AI models had been published, spanning protein structure (AlphaFold, ESM-3, Boltz-1), small-molecule Drug Discovery (Chai-1, DiffDock), genomics (Evo, Nucleotide Transformer), and single-cell Transcriptomics (scGPT, Geneformer, Nicheformer), fundamentally restructuring both the academic field and the commercial Drug Discovery landscape.
- Isomorphic Labs IsoDDE and clinical milestones: In February 2026, Isomorphic Labs (the DeepMind spinout focused on AI-driven Drug Discovery) published a 27-page technical report describing the IsoDDE (Isomorphic Drug Design Engine), a diffusion-based drug design model that substantially exceeds AlphaFold3 on difficult protein-ligand structure prediction cases — achieving 50% accuracy on cases with less than 20% similarity to training data, versus AlphaFold3’s 23.3%. As of June 2026, Isomorphic Labs reports 17 active drug development programmes spanning oncology, immunology, and cardiovascular disease, with CEO Demis Hassabis confirming at the January 2026 World Economic Forum that the first AI-designed cancer drug is on track to enter Phase I clinical trials by end of 2026. This milestone — from computational design to first-in-human dosing — represents the most direct validation yet that computational biology and Machine Learning-driven Drug Discovery can translate from algorithmic breakthrough to clinical reality at timescales competitive with traditional pharmaceutical development. The convergence of AlphaFold-derived structural biology, generative chemistry Deep Learning, and industrial-scale High Performance Computing is transforming the Drug Discovery timeline from a decade-long process to a multi-year one.
- Spatial omics: The Wellcome Sanger Institute is at the forefront of deploying spatial transcriptomics to study cell organisation in tissue disease contexts — generating inflammatory skin atlases spanning 22 skin conditions including psoriasis and eczema, using spatial technologies to map how cell type, state, and position combine to drive disease. The integration of spatial and single-cell data at the EMBL-EBI / Sanger Single-Cell Genomics Centre constitutes one of the most computationally demanding data integration challenges in contemporary Bioinformatics, requiring novel multi-modal alignment methods that register cells across technologies, donors, and conditions.
UK Context
- The United Kingdom holds a globally prominent position in computational biology, anchored by a cluster of world-leading institutions in and around Cambridge. The Wellcome Sanger Institute at the Wellcome Genome Campus (Hinxton) is one of the world’s largest genome sequencing centres and a primary source of computational methodology for single-cell Genomics, population Genomics, and pathogen surveillance. EMBL-EBI, co-located at Hinxton, curates the UniProt, Ensembl, PDB, and ArrayExpress databases that underpin global computational biology research. The Francis Crick Institute (London) brings together MRC, Cancer Research UK, and the Wellcome Trust in a large-scale biomedical research environment with significant computational biology groups. The Babraham Institute (Cambridge) leads in Epigenomics and single-cell Transcriptomics methodology, with particular strength in chromatin accessibility and DNA methylation computational methods.
- Edinburgh houses two major computational biology presences: the Roslin Institute (site of the Dolly sheep cloning experiment), whose computational animal and agricultural Genomics group applies Machine Learning to livestock breeding and zoonotic disease; and the MRC Human Genetics Unit, a centre for statistical Genomics and Population Genetics methods. The MRC Institute of Genetics and Cancer at Edinburgh University is a leading centre for cancer computational biology, with particular strengths in single-cell analysis and tumour evolution modelling. UCL’s Genetics Institute and the UCL Institute of Health Informatics apply Bayesian Inference and Machine Learning to population-scale Genomics and Precision Medicine in the NHS context, with projects spanning rare-disease genomics, polygenic risk prediction, and multimodal data linkage across the UK Biobank and NHS Electronic Health Records.
- In Northern England, the Cancer Research UK Manchester Institute hosts a Computational Biology Support team that provides Bioinformatics analysis across the institute’s cancer research portfolio. The University of Manchester offers an MSc in Bioinformatics and Systems Biology ranked 7th in UK Biological Sciences (QS World University Rankings 2025) and has active research groups in metabolic network modelling and single-cell Transcriptomics. The universities of Leeds and Sheffield are developing computational biology and Bioinformatics research communities with strong connections to regional NHS Genomics laboratories, and Leeds’ Astbury Centre for Structural Molecular Biology is a nationally significant facility for Structural Biology and Cryo-EM. Newcastle University’s Biosciences Institute and the Northern Institute for Cancer Research (NICR) apply computational biology methods to regional cancer incidence, mortality, and treatment outcomes in the context of Northern England’s health inequalities — an important social mission that grounds computational biology in Healthcare equity alongside technical excellence.
- The Alan Turing Institute (London) has established a Health and Medical Sciences programme that bridges computational biology, clinical informatics, and Machine Learning, reflecting UK strategy to exploit NHS data for biomedical Artificial Intelligence development. The Crick Data Science Platform supports data management, analysis infrastructure, and Machine Learning method development across the Francis Crick Institute’s 1,500-person research community. Imperial College London’s Department of Bioengineering and Department of Metabolism, Digestion and Reproduction apply computational biology and Systems Biology to metabolic disease, multi-omics data integration, and imaging-genomics convergence, with particular strength in spatial omics and network medicine.
- The UK bioinformatics job market reflects this academic concentration: EMBL-EBI, the Wellcome Sanger Institute, Genomics England, AstraZeneca, and GSK collectively represent the largest employers of computational biology specialists in the country, with demand concentrated in Cambridge, London, and Oxford — the so-called “Golden Triangle” of UK bioscience. Manchester, Sheffield, and Newcastle have also established regional bioinformatics nodes supporting NHS diagnostic Genomics laboratories. The 2026 UK Bioinformatics Jobs survey (Biotechnology Jobs UK) identified Python, R, Nextflow, and Cloud Computing (AWS/GCP) as the dominant skill requirements for computational biology positions, reflecting the industrialisation of the field’s analytical infrastructure. The Health Data Research UK (HDRUK) network coordinates cross-institutional data access for population-scale computational biology research, linking NHS electronic health records, UK Biobank genetic data, and multi-omic research datasets across member institutions including Imperial College, UCL, Edinburgh, Oxford, Cambridge, and Manchester. The UKRI Medical Research Council maintains dedicated funding streams for computational biology methodological innovation through its Skills and Career Development Awards and through the MRC Methodology Research Programme.
Future Directions (2026–2030)
- Whole-cell and multi-scale simulation: The next ambition in computational biology is the simulation of entire cells as dynamical systems, integrating metabolic flux, gene regulation, protein production, membrane mechanics, and mechanical forces in a unified multi-scale model. Projects such as the Virtual Cell (University of Connecticut Health), the Allen Institute’s cell modelling programme, and EMBL’s open science computational biology initiative are early exemplars; Machine Learning force fields (MACE, NequIP, ANI) and differentiable simulation frameworks (JAX-MD, OpenMM, DiffTaichi) are expected to make multi-scale whole-cell models computationally tractable at microsecond-to-second timescales by 2028–2030, closing the gap between molecular simulation and cellular phenotypic prediction.
- Generative biology: Foundation models are moving from prediction to design — generating novel protein sequences, regulatory elements, and genome edits with specified functional properties. By 2028, computationally designed enzymes and therapeutic proteins designed entirely in silico are expected to enter clinical development, closing the loop between Computational Biology, Synthetic Biology, and translational Healthcare.
- AI-guided experiments: Closed-loop laboratory systems that couple computational biology predictions with robotic experimentation (self-driving labs) will shorten the design-build-test-learn cycle in Drug Discovery and materials biology from months to days. The UK Rosalind Franklin Institute and the Manchester Synthetic Biology Research Centre (SYNBIOCHEM) are early exemplars.
- Regulatory genomics and non-coding interpretation: the 98% of the human genome that does not encode proteins — enhancers, promoters, silencers, insulators, non-coding RNAs, and topological domain boundaries — is emerging as a major frontier for computational biology. Models such as Enformer (DeepMind / Avsec et al., 2021) and Borzoi (2024) predict genome-wide gene expression from raw DNA sequence using convolutional and attention architectures trained on ENCODE and GTEx functional genomics data, enabling in silico prediction of how regulatory variants affect gene expression — a key step in interpreting non-coding GWAS hits, which account for the majority of trait-associated variants.
- Federated and privacy-preserving genomics: As population-scale genomics generates sensitive personal health data, federated learning and secure multi-party computation will enable cross-institutional and cross-national computational biology analyses without centralising patient data, supporting Precision Medicine at national scale while respecting data sovereignty. Initiatives such as the European Genomic Data Infrastructure (GDI), the Global Alliance for Genomics and Health (GA4GH) Federated Analysis Framework, and the UK’s Health Data Research UK network are piloting federated analysis systems that allow computation to travel to data rather than vice versa — a paradigm shift for cohort-scale Genomics research that protects participant privacy while enabling discovery that no single cohort can support alone.
- Causal inference in genomics: moving beyond correlation to causal relationship inference is an urgent frontier. Mendelian randomisation (MR) uses genetic variants as instrumental variables to estimate causal effects of modifiable exposures (e.g., LDL cholesterol) on disease outcomes, leveraging the random assortment of alleles at conception as a natural experiment. Bayesian causal discovery (PC algorithm, Fast Causal Inference) learns causal graph structure from observational multi-omic data. Perturbational single-cell genomics (Perturb-seq, CRISPR screens) provides causal gene function data at transcriptome-wide scale, enabling training of causal Deep Learning models that predict gene regulatory consequences of any perturbation.
- Multimodal biology: the integration of genomics, Transcriptomics, Proteomics, Epigenomics, metabolomics, and imaging data in unified computational models — “multi-omics” — will produce unprecedented mechanistic understanding of complex diseases. Foundation models trained on multi-modal biological data (e.g., the Ginkgo Single-Cell Atlas, combining scRNA-seq with spatial transcriptomics and ATAC-seq at single-cell resolution) will enable holistic prediction of cellular phenotype from genotype under any perturbation condition. The UK EMBL-EBI is building the next generation of multi-modal databases to support this convergence, and the Wellcome Sanger Institute Cellular Genomics programme exemplifies the institutional bet on multi-modal computational biology as the future of biomedical research.
- Digital biology and simulation: the convergence of Molecular Dynamics, Systems Biology, and generative AI will enable digital twins of biological systems — computational replicas of a patient’s molecular state that can be used to test therapeutic interventions in silico before committing to clinical trials. Early implementations include personalised tumour growth models calibrated to patient imaging and genomic data (used in adaptive radiotherapy), and mechanistic pharmacokinetics-pharmacodynamics (PKPD) models augmented with Machine Learning to predict inter-patient variability in drug response from genomic and proteomic inputs.
- Quantum computational biology: Quantum Computing platforms offer potential speedups for specific computational biology problems, particularly quantum chemistry calculations for Drug Discovery (ground-state energy estimation for ligand-protein systems using variational quantum eigensolvers), protein folding as a combinatorial optimisation problem (quantum annealing approaches), and quantum machine learning for molecular property prediction. UK investment in quantum computing through the National Quantum Computing Centre (NQCC, established at Harwell, Oxfordshire, 2021) includes a life sciences programme that identifies computational biology as a priority application domain for quantum advantage by the early 2030s.
- Integration with Knowledge Graph infrastructure: Biological Knowledge Graphs linking genes, proteins, diseases, drugs, and clinical outcomes (e.g., Open Targets, STRING, BioKG) will increasingly serve as the structured context layer that grounds Large Language Models operating over biomedical literature, enabling more reliable biomedical AI with traceable evidence chains.
Software Ecosystem and Key Tools
- Genomics and Bioinformatics infrastructure:
- GATK (Genome Analysis Toolkit, Broad Institute): the gold-standard pipeline for germline short-variant calling (HaplotypeCaller) and somatic variant calling (Mutect2); widely used in NHS diagnostic sequencing.
- Samtools / BCFtools / HTSlib: low-level C libraries for manipulating SAM/BAM alignment files and VCF variant files; the lingua franca of short-read genomics pipelines.
- STAR (Spliced Transcripts Alignment to a Reference): the standard short-read RNA aligner supporting splice-aware alignment across exon-intron junctions; used in DESeq2 / edgeR differential expression workflows.
- Nextflow / Snakemake: workflow management systems enabling reproducible, containerised, and cloud-scalable computational biology pipelines with support for SLURM and AWS/GCP.
- nf-core: a community-curated collection of Nextflow bioinformatics analysis pipelines (RNAseq, sarek for cancer genomics, chipseq, scRNAseq) with standardised containers and test datasets.
- Protein Structure Prediction tools:
- AlphaFold2/3 (DeepMind): end-to-end structure prediction; AF3 via webserver only (academic use) or Isomorphic Labs API.
- ESMFold (Meta AI Research): 650M-parameter ESM-2 language model with a structure module; 60× faster than AF2, suitable for proteome-scale batch prediction.
- RoseTTAFold All-Atom (Institute for Protein Design, UW): predicts protein–small molecule and protein–DNA/RNA complexes; open-source.
- OpenFold / UniFold: open-source reimplementations of AlphaFold2 enabling fine-tuning on custom structural datasets.
- Single-cell analysis tools:
- Seurat (Satija Lab, NYGC): R package for scRNA-seq analysis — normalisation, PCA, UMAP, clustering, differential expression, cross-dataset integration.
- Scanpy (Theis Lab, Helmholtz Munich): Python AnnData-based framework for large-scale single-cell analysis; used in the EMBL-EBI Single-Cell Expression Atlas.
- scVI (YosefLab, UCB): variational autoencoder for probabilistic single-cell data integration and batch correction; underlying model for CellxGene Census.
- CellChat / LIANA: tools for inferring cell-cell communication and ligand-receptor interaction networks from single-cell data; important for understanding tissue-level Systems Biology.
- Drug Discovery computational tools:
- RDKit: open-source cheminformatics toolkit for molecular descriptor calculation, fingerprinting, and 2D/3D structure handling.
- Chemprop: message-passing Graph Neural Network for molecular property prediction; state-of-the-art for ADMET property modelling.
- DiffDock (MIT): diffusion-based blind docking model that outperforms classical docking methods (AutoDock Vina) on cross-docking benchmarks; jointly predicts ligand binding pose with confidence.
- ProteinMPNN (Institute for Protein Design, UW): sequence design model for given backbone structures; combined with RFDiffusion enables complete de novo protein design pipeline.
Benchmark Datasets and Canonical Resources
- Computational biology has a rich ecosystem of benchmark datasets that serve as gold standards for method evaluation and as persistent shared resources for the community:
- Protein Structure Benchmarks: the Critical Assessment of Protein Structure Prediction (CASP) competition (biennial since 1994) provides standardised blind assessment of Protein Structure Prediction methods. CASP14 (2020) established AlphaFold2’s landmark accuracy; CASP15 (2022) and CASP16 (2024) have further refined understanding of complex prediction and RNA structure. The Protein Data Bank (PDB, founded 1971, current curation by wwPDB) contains over 220,000 experimentally determined macromolecular structures as of 2026, serving as the ground-truth reference for structural computational biology.
- Sequence and Functional Databases: UniProt/Swiss-Prot (manually curated protein sequences and functions), Pfam (protein domain families using hidden Markov models), InterPro (integrated protein family classification), and RefSeq (NCBI reference sequences for genomes and transcripts). The AlphaFold Database (AlphaFoldDB, EMBL-EBI and DeepMind) provides AI-predicted structures for over 200 million proteins, a scale that dwarfs the PDB.
- Genomics and Variant Resources: gnomAD (Genome Aggregation Database, 125,000+ exome sequences and 15,000+ genome sequences for population variant frequency), ClinVar (clinical significance of genomic variants), and the UK Biobank (500,000 participants with linked genomic, imaging, and health data) provide population-scale variant interpretation benchmarks. GWAS Catalog aggregates thousands of genome-wide association study results across hundreds of traits and diseases.
- Single-Cell Atlases: the Human Cell Atlas (HCA) single-cell RNA-sequencing reference datasets for human tissues, Tabula Sapiens (Tabula Muris Consortium), and CZ CellxGene (Chan Zuckerberg Initiative) serve as benchmark references for cell-type classification, trajectory inference, and perturbation response prediction tasks. CELLxGENE hosts over 1,000 single-cell datasets comprising hundreds of millions of cells as of 2026.
- Drug Discovery Benchmarks: MoleculeNet (Wu et al., 2018) standardises molecular property prediction benchmarks (ESOL, FreeSolv, Lipophilicity, BACE, BBBP, Tox21); DUD-E (Directory of Useful Decoys — Enhanced) and PCBA benchmark virtual screening performance; CADD challenges (computer-aided drug discovery) provide blind prospective docking benchmarks. GuacaMol and MOSES benchmark molecular generation quality and diversity.
- Pathogen and Metagenomics Resources: NCBI RefSeq bacterial and viral genomes; the PATRIC pathogen database; and the Earth Microbiome Project (EMP) metagenomics dataset provide reference collections for Metagenomics classification methods. CAMI (Critical Assessment of Metagenome Interpretation) runs periodic competitions analogous to CASP for metagenomics benchmarking.
Formal Analysis: Complexity and Algorithmic Constraints
- Computational biology problems span a wide range of algorithmic complexity classes, and understanding the formal difficulty of core problems guides method design and interpretation of approximation results:
- NP-hard problems in computational biology:
- Multiple sequence alignment (MSA) is NP-hard for 3 or more sequences under the sum-of-pairs scoring model.
- Protein threading (fold recognition by sequence-to-structure alignment) is NP-hard in general.
- Haplotype assembly from shotgun sequencing reads is NP-hard.
- Minimum spanning tree phylogeny (minimum parsimony) is NP-hard, motivating heuristic algorithms such as TNT, PAUP*, and NNI/SPR tree search.
- Polynomial-time solvable core problems:
- Pairwise sequence alignment (Smith-Waterman, Needleman-Wunsch): O(mn) dynamic programming.
- RNA secondary structure prediction (Nussinov, Zuker): O(n³) dynamic programming for minimum free energy structure.
- Genome-scale metabolic flux balance analysis: linear programming, solvable in polynomial time.
- Neighbour-joining phylogenetic tree construction: O(n³) in number of taxa.
- Approximation and heuristics: many NP-hard problems in computational biology are addressed by:
- Progressive multiple alignment (ClustalW, MAFFT): greedy guide-tree-based heuristic, not provably near-optimal.
- BLAST: heuristic approximation of Smith-Waterman with controlled sensitivity-specificity trade-off.
- Greedy genome assembly: graph-based heuristics (de Bruijn, overlap-layout-consensus) with practical but not worst-case performance guarantees.
- Randomised and Probabilistic Modelling approaches:
- MCMC (Markov chain Monte Carlo): used in Bayesian Phylogenetics (BEAST, MrBayes) and single-cell trajectory inference to sample from posteriors that are intractable analytically.
- Simulated annealing and genetic algorithms: applied to molecular Molecular Dynamics docking, protein design, and genome assembly parameter optimisation.
- Variational inference: scalable Bayesian Inference approximation (mean-field, ELBO maximisation) used in scVI, pyro-based probabilistic models for single-cell data.
- Streaming and online algorithms for Big Data:
- MinHash and locality-sensitive hashing for genome sketching and approximate nearest-neighbour search (Mash, Sourmash).
- Bloom filters for k-mer membership queries in streaming genome assembly (GATB library, Bifrost assembler).
- Approximate string matching with suffix arrays and FM-indexes (BWT-based short-read aligners: BWA-MEM, Bowtie2) enabling O(n log n) genome indexing and O(m) query time.
Key Terminology
- Bioinformatics: the computational branch of computational biology focused primarily on analysis of biological sequence data — DNA, RNA, and protein sequences — including assembly, alignment, annotation, and functional inference. Distinguished from the broader computational biology by its specific focus on sequence-level data; sometimes used interchangeably but formally a subdiscipline.
- Sequence Alignment: the process of arranging two or more sequences to identify regions of similarity, implying functional, structural, or evolutionary relationships. Global alignment (Needleman-Wunsch) aligns full sequences; local alignment (Smith-Waterman) finds the highest-scoring aligned subsequence. BLAST heuristically approximates local alignment at orders-of-magnitude greater speed.
- Hidden Markov Model (HMM): a probabilistic graphical model with hidden states and observed emissions, used extensively in bioinformatics for sequence modelling (HMMER for protein families, AUGUSTUS for gene prediction). HMM-based profiles capture position-specific amino acid preferences across protein domain families far more sensitively than pairwise alignment.
- Protein Structure Prediction: the computational problem of determining the three-dimensional atomic coordinates of a protein from its amino acid sequence. Solved to near-experimental accuracy for single-chain proteins by AlphaFold2 (2021); extended to complexes and multi-molecular assemblies by AlphaFold3 (2024) and RoseTTAFold All-Atom.
- Molecular Dynamics (MD): atomistic simulation of macromolecular motion by numerically integrating Newton’s equations of motion under an empirical or machine-learned force field. Classical MD timesteps are femtoseconds (10⁻¹⁵ s); biologically relevant timescales are microseconds to milliseconds, requiring heroic compute or enhanced sampling methods. Machine Learning force fields (MACE, NequIP) achieve near-quantum accuracy at a fraction of the cost of ab initio methods.
- Genome-Wide Association Study (GWAS): a statistical analysis correlating millions of single nucleotide polymorphisms (SNPs) distributed across the genome with a phenotypic trait or disease in a large population cohort. Identifies loci of statistical association; causal inference requires fine-mapping, colocalization, and functional validation. UK Biobank enables GWAS at scale in diverse ancestries.
- Polygenic Risk Score (PRS): a scalar summary of an individual’s genome-wide genetic liability to a trait or disease, computed as a weighted sum of risk alleles across GWAS loci. Deployed in Precision Medicine for stratified Healthcare interventions and clinical trial enrichment.
- Single-Cell RNA Sequencing (scRNA-seq): a technology profiling the transcriptome (expressed RNA) of individual cells in an unbiased manner, enabling cell-type discovery, trajectory inference, and cell-state characterisation at single-cell resolution. The Wellcome Sanger Institute is applying spatial transcriptomics — which adds spatial coordinates to single-cell expression — to map tissue organisation at molecular resolution.
- Synthetic Biology: the design and construction of new biological parts, devices, and systems, or the redesign of existing natural systems for useful purposes. Computational biology provides the design frameworks (genetic circuit modelling, codon optimisation, metabolic pathway design), and generative AI tools such as ESM-3 and ProteinMPNN provide sequence-level design capabilities that reduce experimental iteration cycles.
- ESM-3 / EvolutionaryScale: a multimodal protein language model (98 billion parameters) jointly reasoning over protein sequence, structure, and function, developed by EvolutionaryScale (launched June 2024). Trained with over 10²⁴ FLOPs on an evolutionary corpus — simulating 500 million years of protein evolution — ESM-3 generated esmGFP, a novel green fluorescent protein with only 58% sequence identity to any known fluorescent protein, demonstrating generative exploration of previously unsampled protein space.
Research & Literature
-
- Jumper, J. et al. (2021). “Highly accurate protein structure prediction with AlphaFold.” Nature, 596, 583–589. https://doi.org/10.1038/s41586-021-03819-2
-
- Varadi, M. et al. (2022). “AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models.” Nucleic Acids Research, 50(D1), D439–D444. https://doi.org/10.1093/nar/gkab1061
-
- Abramson, J. et al. (2024). “Accurate structure prediction of biomolecular interactions with AlphaFold 3.” Nature, 630, 493–500. https://doi.org/10.1038/s41586-024-07487-w
-
- Lin, Z. et al. (2023). “Evolutionary-scale prediction of atomic-level protein structure with a language model.” Science, 379(6637), 1123–1130. https://doi.org/10.1126/science.ade2574
-
- Rives, A. et al. (2021). “Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences.” PNAS, 118(15), e2016239118. https://doi.org/10.1073/pnas.2016239118
-
- Theodoris, C.V. et al. (2023). “Transfer learning enables predictions in network biology.” Nature, 618, 616–624. https://doi.org/10.1038/s41586-023-06139-9
-
- Cui, H. et al. (2024). “scGPT: toward building a foundation model for single-cell multi-omics using generative AI.” Nature Methods, 21, 1470–1480. https://doi.org/10.1038/s41592-024-02201-0
-
- Hie, B.L. et al. (2024). “Efficient evolution of human antibodies from general protein language models.” Nature Biotechnology, 42, 275–283. https://doi.org/10.1038/s41587-023-01763-2
-
- Callaway, E. (2024). “AlphaFold3 — why did DeepMind give it away?” Nature, 629, 728–729. https://doi.org/10.1038/d41586-024-01383-z
-
- Needleman, S.B. and Wunsch, C.D. (1970). “A general method applicable to the search for similarities in the amino acid sequence of two proteins.” Journal of Molecular Biology, 48(3), 443–453. https://doi.org/10.1016/0022-2836(70)90057-4
-
- Smith, T.F. and Waterman, M.S. (1981). “Identification of common molecular subsequences.” Journal of Molecular Biology, 147(1), 195–197. https://doi.org/10.1016/0022-2836(81)90087-5
-
- Krogh, A., Brown, M., Mian, I.S., Sjölander, K. and Haussler, D. (1994). “Hidden Markov models in computational biology: Applications to protein modeling.” Journal of Molecular Biology, 235(5), 1501–1531. https://doi.org/10.1006/jmbi.1994.1104
-
- Lander, E.S. et al. (2001). “Initial sequencing and analysis of the human genome.” Nature, 409, 860–921. https://doi.org/10.1038/35057062
-
- Friedman, N., Linial, M., Nachman, I. and Pe’er, D. (2000). “Using Bayesian networks to analyze expression data.” Journal of Computational Biology, 7(3–4), 601–620. https://doi.org/10.1089/106652700750050961
-
- Love, M.I., Huber, W. and Anders, S. (2014). “Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2.” Genome Biology, 15, 550. https://doi.org/10.1186/s13059-014-0550-8
-
- Satija, R., Farrell, J.A., Gennert, D., Schier, A.F. and Regev, A. (2015). “Spatial reconstruction of single-cell gene expression data.” Nature Biotechnology, 33, 495–502. https://doi.org/10.1038/nbt.3192
-
- Zaheer, M. et al. (2024). “CancerFoundation: A single-cell RNA sequencing foundation model to decipher drug resistance in cancer.” bioRxiv preprint. https://doi.org/10.1101/2024.11.01.621087
-
- Gao, Z. et al. (2025). “Biology-driven insights into the power of single-cell foundation models.” Genome Biology, 26, 88. https://doi.org/10.1186/s13059-025-03781-6
-
- Shi, Y. et al. (2024). “Foundation Model in Biomedicine.” arXiv preprint. https://arxiv.org/abs/2503.02104
-
- Senior, A.W. et al. (2020). “Improved protein structure prediction using potentials from deep learning.” Nature, 577, 706–710. https://doi.org/10.1038/s41586-019-1923-7
-
- Huttenhower, C. et al. (2012). “Structure, function and diversity of the healthy human microbiome.” Nature, 486, 207–214. https://doi.org/10.1038/nature11234
-
- Barabási, A.-L. and Oltvai, Z.N. (2004). “Network biology: understanding the cell’s functional organization.” Nature Reviews Genetics, 5, 101–113. https://doi.org/10.1038/nrg1272
-
- Watson, J.D. and Crick, F.H. (1953). “Molecular structure of nucleic acids: A structure for deoxyribose nucleic acid.” Nature, 171, 737–738. https://doi.org/10.1038/171737a0
-
- Dayhoff, M.O., Schwartz, R.M. and Orcutt, B.C. (1978). “A model of evolutionary change in proteins.” Atlas of Protein Sequence and Structure, 5(Suppl 3), 345–352.
-
- Yang, J. et al. (2025). “Multimodal foundation transformer models for multiscale genomics.” arXiv preprint. https://arxiv.org/abs/2503.02104
-
- Cancer Research UK Manchester Institute. (2026). “Computational Biology Support.” https://www.cruk.manchester.ac.uk/facility/computational-biology-support/
-
- Wellcome Sanger Institute. (2025). “Sanger Institute-EBI Single-Cell Genomics Centre.” https://www.sanger.ac.uk/collaboration/sanger-institute-ebi-single-cell-genomics-centre/
-
- Biotechnology Jobs UK. (2026). “Best Universities for Bioinformatics in the UK 2026.” https://unifresher.co.uk/uni-prep/rankings/best-universities-for-bioinformatics/