Weak-to-strong generalisation is an AI-alignment research paradigm investigating whether a more capable model can be reliably supervised and improved using labels or feedback from a weaker supervisor. It serves as an empirical analogue for the superalignment problem, in which humans must oversee superhuman systems they cannot fully evaluate. Findings explore how strong students recover latent capabilities beyond the noisy weak teacher’s own performance.

Content

  • Experiments fine-tune a strong pretrained model on labels produced by a weaker model and measure how much of the strong model’s potential is recovered, often boosted by auxiliary confidence-based or bootstrapping techniques. The paradigm is a proxy for future scenarios where human oversight is the “weak” signal, informing scalable-oversight and elicitation methods.