Internal, non-public evaluation benchmarks and datasets used by AI developers to measure and optimize model performance on specific, often domain-specific, tasks.
Overview
- Intercom’s Chief Product Officer Paul Adams announced a new model called Apex for Finn, claiming it has a higher resolution rate, fewer hallucinations, and is far cheaper than any other model, enabled by domain-specific proprietary evals from billions of interaction data points. (Source: Paul Adams (Intercom CPO), via AI Daily Brief, 2026-08-24)