Crowdsourcing is the practice of obtaining contributions, labour or judgements from a large distributed group of people, typically through an open call mediated by an online platform. In machine learning it is widely used to collect, label and validate training data by decomposing work into microtasks distributed across many contributors. Effective crowdsourcing combines incentive design with quality-control mechanisms to aggregate noisy individual inputs into reliable results.
- Crowdsourcing obtains contributions from a large distributed group via an open call, often to label or validate data. It is a form of Data Collection that draws on Human-in-the-Loop effort and Collective Intelligence to feed Machine Learning.
Overview
- Many tasks remain easier for humans than machines; crowdsourcing harnesses distributed human judgement at scale to perform them economically.
- Work is decomposed into small microtasks dispatched to many contributors, whose individual, noisy outputs are aggregated into reliable results.
- Quality depends as much on incentive and reputation design as on the task itself, since misaligned incentives produce low-quality or adversarial contributions.
Mechanisms
- Task decomposition into microtasks suitable for non-experts.
- Redundant assignment with consensus or majority aggregation.
- Quality control through gold standards, screening and reputation.
- Incentive design to reward accurate, timely contributions.
Applications
- Labelling training datasets for machine learning.
- Content moderation, transcription and translation.
- Human evaluation and preference collection for model alignment.