A machine learning paradigm that trains models from comparative human judgements (e.g., ‘A is better than B’) rather than absolute labels or demonstrations, enabling alignment with human values. Preference learning underpins reinforcement learning from human feedback and direct preference optimisation, and is the standard technique for aligning large language models.

Semantic Classification

Content

  • A machine learning paradigm that learns from comparative judgments (e.g., “A is better than B”) rather than absolute labels or demonstrations. Preference learning enables training models to align with human values by learning from rankings and comparisons, which are often easier for humans to provide than absolute ratings or demonstrations.

Ways to close The Gap

MollickSmash

  • Drawn extensively from Ethan Mollick who has a Substack called One Useful Thing.
    • AI as a Learning and Teaching Tool:
    • AI, particularly GPT-4, is being effectively used as a tutor and learning aid for students and as a class preparation tool for teachers.
    • It offers adaptive, useful instruction, enhancing learning while reducing busywork.

Decision Transformers based on Preference-Ordering (DPO)

  • Decision models, trained on human preference data, guide the LLM towards preferred outputs, incorporating logic that reflects learned preferences [Decision Transformers: https://arxiv.org/abs/2106.01345].

Ways to close The Gap

MollickSmash

  • Drawn extensively from Ethan Mollick who has a Substack called One Useful Thing.
    • AI as a Learning and Teaching Tool:
    • AI, particularly GPT-4, is being effectively used as a tutor and learning aid for students and as a class preparation tool for teachers.
    • It offers adaptive, useful instruction, enhancing learning while reducing busywork.

Decision Transformers based on Preference-Ordering (DPO)

  • Decision models, trained on human preference data, guide the LLM towards preferred outputs, incorporating logic that reflects learned preferences [Decision Transformers: https://arxiv.org/abs/2106.01345].

The Paradigm Shift: Unsupervised Learning

  • Challenge: The traditional approach faced a major challenge: limited availability of labelled data. Creating these curated datasets is expensive and time-consuming.
  • Concept: The key with unsupervised learning is that the data itself provides the answer, the feedback for the AI to learn. The AI doesn’t need explicit labels; it learns from the inherent structure and patterns within the data.

The Paradigm Shift: Unsupervised Learning

  • Challenge: The traditional approach faced a major challenge: limited availability of labelled data. Creating these curated datasets is expensive and time-consuming.
  • Concept: The key with unsupervised learning is that the data itself provides the answer, the feedback for the AI to learn. The AI doesn’t need explicit labels; it learns from the inherent structure and patterns within the data.

October 2024

Microsoft AI for Science

  • Chris Bishop, is Microsoft Technical Fellow and Director of Microsoft Research AI for Science Microsoft Research Podcast
  • Bishop’s career began in physics, including a PhD in quantum field theory and work on nuclear fusion.
    • He transitioned to machine learning after being inspired by Geoff Hinton’s backpropagation paper, applying neural networks to fusion data at the JET experiment.
    • After a research professorship, he joined Microsoft Research in Cambridge, UK, and later founded the Microsoft Research AI for Science team. AI4Science to Empower the Fifth Paradigm of Scientific Discovery

Secrecy, Espionage, and the Perils of Algorithmic Breakthroughs:

  • The Data Wall and the Next Paradigm Shift: The conversation delves into the technical challenges of overcoming the “data wall” in AI, where simply scaling up existing approaches might not be sufficient. The guests anticipate a need for fundamental algorithmic breakthroughs, potentially involving sophisticated self-play techniques and novel reinforcement learning architectures.
  • Guarding the Secrets - A New Manhattan Project?: They stress the paramount importance of safeguarding these breakthroughs from espionage, suggesting that the stakes for AI might rival or even surpass those of the nuclear age.
  • Tacit Knowledge vs. Explicit Code: While acknowledging the role of talented individuals like Alec Radford, they argue that the most valuable insights might lie not in easily copied code but in the accumulated tacit knowledge and experimental learnings within leading AI labs.
  • Learning from History’s Mistakes: The guests draw lessons from historical cases of parallel invention, such as the development of the atomic bomb, arguing that even small delays can prove decisive in global power dynamics. They highlight the German pursuit of heavy water reactors for their nuclear programme, a technically inferior path chosen due to a crucial scientific insight that the US managed to keep secret, as an example of how even seemingly small advantages can have enormous strategic implications.

Opportunities and Innovations

  • 🔑 AI-Enhanced Pedagogical Techniques:
    • Flipped Classrooms: AI can provide customised content for students to study at home, enabling more interactive and problem-solving activities in class.
    • Personalised Learning: AI’s adaptability can cater to individual student needs, potentially reshaping the one-size-fits-all education model.
  • 🌱 Growth in Creative and Critical Thinking:
    • AI aids in brainstorming and idea generation, fostering creativity in students who might struggle with these skills naturally.
    • By challenging decision biases and encouraging diverse perspectives, AI acts as a catalyst for developing critical thinking skills.

October 2024

Microsoft AI for Science

  • Chris Bishop, is Microsoft Technical Fellow and Director of Microsoft Research AI for Science Microsoft Research Podcast
  • Bishop’s career began in physics, including a PhD in quantum field theory and work on nuclear fusion.
    • He transitioned to machine learning after being inspired by Geoff Hinton’s backpropagation paper, applying neural networks to fusion data at the JET experiment.
    • After a research professorship, he joined Microsoft Research in Cambridge, UK, and later founded the Microsoft Research AI for Science team. AI4Science to Empower the Fifth Paradigm of Scientific Discovery

Secrecy, Espionage, and the Perils of Algorithmic Breakthroughs:

  • The Data Wall and the Next Paradigm Shift: The conversation delves into the technical challenges of overcoming the “data wall” in AI, where simply scaling up existing approaches might not be sufficient. The guests anticipate a need for fundamental algorithmic breakthroughs, potentially involving sophisticated self-play techniques and novel reinforcement learning architectures.
  • Guarding the Secrets - A New Manhattan Project?: They stress the paramount importance of safeguarding these breakthroughs from espionage, suggesting that the stakes for AI might rival or even surpass those of the nuclear age.
  • Tacit Knowledge vs. Explicit Code: While acknowledging the role of talented individuals like Alec Radford, they argue that the most valuable insights might lie not in easily copied code but in the accumulated tacit knowledge and experimental learnings within leading AI labs.
  • Learning from History’s Mistakes: The guests draw lessons from historical cases of parallel invention, such as the development of the atomic bomb, arguing that even small delays can prove decisive in global power dynamics. They highlight the German pursuit of heavy water reactors for their nuclear programme, a technically inferior path chosen due to a crucial scientific insight that the US managed to keep secret, as an example of how even seemingly small advantages can have enormous strategic implications.

Opportunities and Innovations

  • 🔑 AI-Enhanced Pedagogical Techniques:

    • Flipped Classrooms: AI can provide customised content for students to study at home, enabling more interactive and problem-solving activities in class.
    • Personalised Learning: AI’s adaptability can cater to individual student needs, potentially reshaping the one-size-fits-all education model.
  • 🌱 Growth in Creative and Critical Thinking:

    • AI aids in brainstorming and idea generation, fostering creativity in students who might struggle with these skills naturally.

    • By challenging decision biases and encouraging diverse perspectives, AI acts as a catalyst for developing critical thinking skills.

      Key Characteristics

  • Learns from pairwise comparisons

    • More natural for human evaluators

    • Captures relative preferences

    • Basis for reward modeling

    • Enables alignment training

    • Scales better than demonstrations

      Technical Details

      Pairwise Comparison:

      Given: Prompt x, Outputs y₁ and y₂
      Human judgment: y₁ ≻ y₂ (y₁ preferred)
      

      Bradley-Terry Model:

      P(y₁ ≻ y₂ | x) = σ(r(x,y₁) - r(x,y₂))
      
      Where:
      - r: Reward/preference function
      - σ: Sigmoid function
      

      Training Objective:

      Maximize: Σ log P(y_winner ≻ y_loser | x)
      

      Usage in AI/ML

      Preference learning forms the basis of reward model training in RLHF, where human rankings of model outputs are converted into a learned preference function.

      Academic Context

      Preference learning provides the foundation for modern alignment techniques like RLHF and DPO, recognizing that humans are better at comparative judgments than absolute specifications of desired behavior.

      Primary Sources:

    • Ouyang et al., arXiv:2203.02155 (2022) - RLHF

    • Rafailov et al., arXiv:2305.18290 (2023) - DPO

  • RLHF: Major application of preference learning

    • Reward Model: Implements learned preferences

    • Direct Preference Optimisation (DPO): Direct use of preferences

    • Human Feedback: Source of preference data

    • Ranking: Core data format

      Preference Elicitation Methods

      Pairwise Comparison:

    • Compare two outputs

    • Simple binary choice

    • Most common method

    • Clear and efficient

      Ranking:

    • Order multiple outputs

    • More informative per query

    • Can decompose into pairwise

    • Higher cognitive load

      Best-of-N Selection:

    • Choose best from set

    • Efficient information gathering

    • Partial ranking

    • Scales to larger sets

      Cardinal Ratings:

    • Numerical scores (1-5, 1-10)

    • Absolute judgments

    • Harder to calibrate

    • Can convert to pairwise

      Advantages Over Demonstrations

      Easier for Humans:

    • Comparison easier than creation

    • Less cognitive load

    • More consistent

    • Faster to provide

      More Scalable:

    • Can compare model outputs

    • Don’t need expert writers

    • Cheaper per judgment

    • Higher throughput

      Richer Signal:

    • Multiple comparisons per prompt

    • Captures subtle distinctions

    • Better uncertainty estimates

    • Fine-grained preferences

      Challenges

      Transitivity:

    • Human preferences may be intransitive

    • (A ≻ B, B ≻ C, C ≻ A possible)

    • Models assume transitivity

    • Can cause inconsistencies

      Context Dependence:

    • Preferences vary with context

    • Presentation order effects

    • Relative vs. absolute quality

    • Anchor effects

      Aggregation:

    • Multiple annotators may disagree

    • How to aggregate conflicting preferences

    • Majority vote vs. modeling disagreement

    • Meta-preferences

      Implementation Approaches

      Reward Model (RLHF):

      1. Collect preference comparisons
      2. Train reward model on comparisons
      3. Use reward in RL optimization
      4. Two-stage process

      Direct Preference Optimization (DPO):

      1. Collect preference comparisons
      2. Directly optimize policy on preferences
      3. No separate reward model
      4. Single-stage process

      Preference Dataset Construction

      Sampling Strategy:

    • Diverse prompt coverage

    • Multiple outputs per prompt (4-9 typical)

    • Controlled output diversity

    • Include edge cases

      Comparison Assignment:

    • Random pairs from outputs

    • All pairs (exhaustive)

    • Active selection (informative pairs)

    • Balanced difficulty

      Quality Control:

    • Multiple annotators per comparison

    • Agreement metrics

    • Consensus requirements

    • Expert validation

      Mathematical Foundations

      Bradley-Terry-Luce Model:

      P(i chosen from set S) = exp(rᵢ) / Σⱼ∈S exp(rⱼ)
      

      Thurstone Model:

      Utility: uᵢ = rᵢ + εᵢ, εᵢ ~ N(0, σ²)
      P(i ≻ j) = P(uᵢ > uⱼ)
      

      Plackett-Luce Model: Extension to rankings of multiple items

      Applications

      Language Model Alignment:

    • InstructGPT, ChatGPT, Claude

    • Instruction following

    • Safety and helpfulness

    • Reducing harmful outputs

      Recommendation Systems:

    • Learning user preferences

    • Ranking items

    • Personalization

    • A/B testing

      Robotics:

    • Learning from demonstrations

    • Task preferences

    • Safe behaviour learning

    • Human-robot interaction

      Preference Learning vs. Supervised Learning

      Preference Learning:

    • Comparative judgments

    • Easier data collection

    • More natural for humans

    • Captures relative quality

      Supervised Learning:

    • Absolute labels

    • Demonstration-based

    • Higher quality signal

    • More expensive

      Best Practices

      Data Collection:

    • Clear comparison guidelines

    • Sufficient context for judgment

    • Avoid ties when possible

    • Track annotator metadata

      Model Training:

    • Validate on held-out comparisons

    • Monitor transitivity violations

    • Check calibration

    • Test generalization

      Quality Assurance:

    • Inter-annotator agreement

    • Consistency checks

    • Expert review samples

    • Continuous monitoring

      Active Preference Learning

      Uncertainty-Based:

    • Query pairs with uncertain preferences

    • Maximum information gain

    • Reduce labeling cost

    • Faster convergence

      Diversity-Based:

    • Cover preference space

    • Avoid redundant queries

    • Balance with uncertainty

    • Better generalization

      Historical Development

    • 1920s: Thurstone’s law of comparative judgment

    • 1950s: Bradley-Terry model

    • 2000s: Preference learning in ML

    • 2017-2020: Application to RL

    • 2022: RLHF mainstream adoption

    • 2023+: DPO and alternatives

      Significance

      Preference learning recognizes that humans excel at comparative judgments, enabling more natural and scalable feedback collection for aligning AI systems with human values and preferences.

      OWL Functional Syntax

      UK English Notes

    • “Optimisation” (not “optimization”)

    • “Behaviour” (not “behavior”)

    • “Recognised” (not “recognized”)

      Last Updated: 2025-10-27 Verification Status: Verified against RLHF and DPO papers

      Academic Context

  • Preference learning is a specialised machine learning paradigm focused on learning from comparative judgments, such as “A is better than B,” rather than relying on absolute labels or demonstrations.

  • This approach aligns closely with human decision-making processes, which often involve relative preferences rather than fixed scores.

  • The academic foundations of preference learning draw from fields including econometrics, social choice theory, reinforcement learning, and optimisation.

  • Key developments include formal models of preference data, such as pairwise comparisons and ranking distributions, which enable algorithms to infer latent utility functions from observed preferences[2][3][6].

    Current Landscape (2025)

  • Preference learning has seen increasing adoption in personalised recommendation systems, virtual assistants, and human-aligned AI models.

  • Notable implementations include training large language models (LLMs) like GPT-4 and LLaMA 3.2 using reinforcement learning from human preferences to better align outputs with user expectations[2].

  • Industry platforms leverage preference learning to enhance user experience by predicting and adapting to subjective tastes, particularly in e-commerce and content curation.

  • In the UK, and specifically in North England, AI research hubs in Manchester and Leeds have contributed to advancing preference learning techniques, often focusing on human-computer interaction and ethical AI alignment.

  • Technical capabilities:

  • Preference learning algorithms efficiently handle subjective, context-dependent data, often requiring fewer explicit labels but more nuanced comparative feedback.

  • Limitations include the need for high-quality preference data and challenges in scaling to complex, multi-dimensional preference spaces.

  • Standards and frameworks for preference learning are emerging, emphasising transparency, fairness, and interpretability in preference-based models[3][6].

    Research & Literature

  • Key academic papers and sources:

  • Christiano, P. F., Leike, J., Brown, T., et al. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30.
    DOI: 10.5555/3295222.3295349

  • Gopalan, A., Saha, A., Bengio, Y., et al. (2023). Do You Prefer Learning with Preferences? NeurIPS Tutorial.
    Available: https://sites.google.com/view/pref-learning-tutorial-neurips/home[3]

  • Recent survey: Preference learning made easy: Everything should be understood from the sampling distribution of pairwise preference data (2025). arXiv:2502.10505[6]

  • Ongoing research directions include:

  • Improving sample efficiency and robustness of preference-based models.

  • Integrating preference learning with multi-agent systems and social choice frameworks.

  • Enhancing interpretability and ethical alignment in human-AI interaction.

    UK Context

  • The UK has a vibrant AI research ecosystem with significant contributions to preference learning from universities and institutes in North England.

  • Manchester’s AI research groups focus on human-centred machine learning, including preference elicitation methods.

  • Leeds and Newcastle have active projects exploring preference learning in healthcare decision support and personalised education technologies.

  • Sheffield’s AI labs contribute to developing scalable algorithms for preference-based optimisation.

  • Regional case studies include collaborations between academia and industry in Leeds, applying preference learning to improve customer experience in retail and digital services.

  • The UK government and research councils support ethical AI initiatives that incorporate preference learning to ensure AI systems respect human values and societal norms.

    Future Directions

  • Emerging trends:

  • Greater integration of preference learning with large-scale language models and reinforcement learning frameworks.

  • Development of hybrid models combining absolute and relative feedback for richer learning signals.

  • Expansion into new domains such as autonomous systems, personalised medicine, and adaptive education.

  • Anticipated challenges:

  • Collecting reliable and unbiased preference data at scale.

  • Balancing model complexity with interpretability and user trust.

  • Addressing ethical concerns around manipulation and privacy in preference elicitation.

  • Research priorities:

  • Designing frameworks for transparent and fair preference learning.

  • Enhancing cross-disciplinary collaboration to incorporate insights from psychology, economics, and social sciences.

  • Developing regionally relevant applications that reflect UK societal values and regulatory environments.

    References

    1. Christiano, P. F., Leike, J., Brown, T., et al. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30. DOI: 10.5555/3295222.3295349
    2. Gopalan, A., Saha, A., Bengio, Y., et al. (2023). Do You Prefer Learning with Preferences? NeurIPS Tutorial. Available at: https://sites.google.com/view/pref-learning-tutorial-neurips/home
    3. Preference learning made easy: Everything should be understood from the sampling distribution of pairwise preference data (2025). arXiv:2502.10505.
    4. Additional UK-specific AI research reports and government ethical AI frameworks (2024-2025).

    Metadata

  • Last Updated: 2025-11-11

  • Review Status: Comprehensive editorial review

  • Verification: Academic sources verified

  • Regional Context: UK/North England where applicable

Provenance