A system design pattern that dynamically directs inference requests to different AI models based on task complexity, cost, or performance requirements to optimize overall efficiency.
Overview
- [Industry analysis] Harvey’s experiment with a ‘worker advisor’ architecture, where an open-weight GLM 5.1 worker delegates high-stakes tasks to a closed Opus 4.7 advisor, resulted in increased performance at a lower cost than using Opus 4.7 alone. (Source: Harvey (cited by AI Daily Brief host), via AI Daily Brief, 2026-08-24)