The structural design and configuration of neural networks and machine learning systems, encompassing layer arrangements, activation functions, and connection patterns that determine how models process information and learn from data.

Semantic Classification

Content

Core Architectures

Convolutional Neural Networks (CNNs)

  • Image analysis

  • Feature extraction

  • Pooling layers

  • Spatial hierarchies

  • Computer vision

    Recurrent Neural Networks (RNNs)

  • Sequence processing

  • Temporal patterns

  • Hidden states

  • Time series

  • Natural language

    Transformers

  • Attention mechanisms

  • Parallel processing

  • Long-range dependencies

  • Language models

  • Multi-modal learning

    Generative Adversarial Networks (GANs)

  • Generator networks

  • Discriminator networks

  • Adversarial training

  • Synthetic data

  • Image generation

    Advanced Architectures

    LSTM and GRU

  • Memory cells

  • Gating mechanisms

  • Long sequences

  • Speech recognition

  • Time series analysis

    Capsule Networks

  • Nested layers

  • Spatial relationships

  • Viewpoint invariance

  • Hierarchical parsing

  • Dynamic routing

    Graph Neural Networks (GNNs)

  • Node relationships

  • Edge features

  • Message passing

  • Social networks

  • Molecular structures

    Novel Developments (2024)

    Kolmogorov-Arnold Networks (KAN)

  • Enhanced interpretability

  • Mathematical foundation

  • Explainable outputs

  • Research interest

  • Transparent learning

    Neural Architecture Search (NAS)

  • Automated design

  • Algorithm optimisation

  • Model discovery

  • Efficiency gains

  • Performance improvement

    Design Patterns

    Layer Types

  • Dense/Fully connected

  • Convolutional

  • Recurrent

  • Attention

  • Normalisation

    Activation Functions

  • ReLU variants

  • Sigmoid

  • Tanh

  • Softmax

  • GELU

    Regularisation

  • Dropout

  • Batch normalisation

  • Weight decay

  • Data augmentation

  • Early stopping

    Training Techniques

    Learning Methods

  • Supervised learning

  • Self-supervised learning

  • Federated learning

  • Transfer learning

  • Reinforcement learning

    Optimisation

  • Adam optimiser

  • SGD variants

  • Learning rate scheduling

  • Gradient clipping

  • Mixed precision

    Architecture Selection

    Considerations

  • Task requirements

  • Data characteristics

  • Computational resources

  • Latency constraints

  • Accuracy needs

    Trade-offs

  • Depth vs width

  • Accuracy vs speed

  • Memory vs performance

  • Complexity vs interpretability

  • Training vs inference

Current Landscape (2026)

  • Mixture-of-Experts (MoE) has become the default frontier architecture: by 2025-2026 essentially every major open-weight flagship is sparse, including DeepSeek-V3/R1 (671B total / 37B active), Llama 4 Maverick (400B / 17B, 128 routed experts), Qwen3-235B-A22B, Mistral Large 3, Kimi K2 and OpenAI’s gpt-oss-120b, with NVIDIA noting the top ten open models all use MoE.
  • Attention itself is being restructured for long-context economics: DeepSeek’s V3.2 introduced DeepSeek Sparse Attention (DSA), subsequently adopted by Zhipu’s GLM-5 (released February 2026), building on Multi-head Latent Attention (MLA) which compresses the KV cache, while Llama 4 uses interleaved-RoPE (iRoPE) to push context towards 10M tokens.
  • The Transformer-versus-SSM debate resolved into hybrids rather than replacement: production stacks now keep a minority of attention layers and swap the rest for Mamba-2 recurrence, as in NVIDIA’s Nemotron-H (~3x faster inference), IBM’s Granite 4.0 (>70% lower memory, ~2x faster serving), AI21’s Jamba and TII’s Falcon-H1.
  • State-space modelling advanced at the research frontier with Mamba-3 (Princeton’s Goomba Lab, ICLR 2026 Oral, arXiv 2603.15569), adding complex-valued state updates and a MIMO formulation that matches Mamba-2 perplexity at half the state size, alongside RWKV-7 “Goose” from the RNN side.
  • New frontier releases treat long-context efficiency as a first-class architectural objective: DeepSeek-V4-Pro (April 2026, 1.6T total / ~49B active, native 1M context) combines compressed sparse attention with FP4 expert training, reaching roughly 27% of V3.2’s per-token FLOPs and 10% of its KV cache at 1M context.
  • Diffusion Transformers (DiT) consolidated as the backbone for generative image and video systems such as Sora 2 and Veo, keeping the Transformer central outside pure language modelling.
  • Open challenges as of 2026 include the SSM in-context recall gap (exact long-range retrieval remains weak versus attention), MoE routing stability and load balancing, and the serving-infrastructure shift to data-parallel attention plus expert-parallel MoE (e.g. vLLM wide expert-parallelism) needed to run these sparse trillion-parameter models economically.

References

Provenance