A classification performance metric representing the proportion of actual positive instances that an artificial intelligence model correctly identifies, calculated as the ratio of true positives to all actual positives (true positives plus false negatives), measuring the model’s completeness in detecting positive cases, particularly critical in applications where missing positive instances (false negatives) carries significant cost or consequences.
FN (False Negatives): Missed positive instances (Type II errors)
Also known as Sensitivity, True Positive Rate (TPR), or Hit Rate.
Context and Significance
Recall answers the question: “Out of all actual positive cases, how many did the model find?” This metric is essential in scenarios where missing positive cases is particularly costly or dangerous—such as disease screening (missing cancer cases), security threat detection (missing threats), or quality control (missing defects). High recall ensures comprehensive detection of positive instances, though it says nothing about how many negative instances are incorrectly flagged (that’s related to precision and specificity).
Recall trades off with precision: achieving 100% recall is trivial (predict every instance as positive) but results in terrible precision. The challenge lies in maintaining high recall whilst managing false positive rates, with application-specific requirements determining the appropriate balance.
Used in: Model Evaluation, threshold selection, performance monitoring
Reported in: Model Cards, clinical validation reports, audit documentation
Examples and Applications
Cancer Screening Test: Recall of 95% means test identifies 95 out of 100 actual cancer cases, missing 5—high recall critical for early detection despite false positives requiring follow-up
Airport Security Screening: Threat detection with 99.9% recall catches 999 out of 1,000 actual threats—extremely high recall necessary despite inconvenience of false alarms (low precision acceptable)
Email Spam Filter: Spam detection with 85% recall catches 85 out of 100 spam emails, allowing 15 through—lower recall acceptable as users can delete spam, but high precision critical to avoid filtering legitimate mail
Manufacturing Defect Detection: Quality control with 92% recall identifies 92 out of 100 defective products—remaining 8% reach customers, requiring balance with inspection costs (precision)
Calculation and Implementation
Standard Calculation:
from sklearn.metrics import recall_scorerecall = recall_score(y_true, y_pred)# For multi-class: specify average parameter# 'micro', 'macro', 'weighted', or None for per-class
Employ cost-sensitive learning to optimise application-specific objectives
Use calibration to improve reliability of probability estimates for threshold setting
Variants and Related Metrics
Micro-averaged Recall (multi-class): Aggregate TP and FN across classes
Recallmicro=∑i(TPi+FNi)∑iTPiMacro-averaged Recall (multi-class): Average of per-class recalls
Recallmacro=n1∑iRecalliWeighted Recall: Recall averaged across classes weighted by support
Recall@K: Proportion of relevant items in top K recommendations (ranking tasks)
Sensitivity Analysis: In medical contexts, often reported as sensitivity with confidence intervals
ISO/IEC and Standards Alignment
ISO/IEC 25059 (Quality Model for AI Systems):
Recall as metric for functional completeness
Coverage of actual positive cases
ISO/IEC 25024 (Data Quality Metrics):
Recall in context of output completeness measurement
NIST AI RMF Integration
MEASURE Function:
MEASURE-2.2: Appropriate metrics including recall selected based on application risks
MEASURE-2.3: Recall measured across different contexts and subgroups
Recall critical for Safety (detecting hazards) and Reliability trustworthiness characteristics
Medical and Diagnostic Context
In medical and diagnostic testing, recall (sensitivity) is conventionally reported alongside specificity (true negative rate):
Sensitivity (Recall): Ability to correctly identify those with condition
Specificity: Ability to correctly identify those without condition
Specificity=TN+FPTN
Together, sensitivity and specificity provide comprehensive picture of diagnostic test performance.
Recall, also known as sensitivity or true positive rate, is a foundational metric in classification tasks, measuring the proportion of actual positive instances that a model correctly identifies
It is especially relevant in domains where missing positive cases (false negatives) can have serious consequences, such as healthcare or fraud detection
Key developments and current state
Recall remains a core component of model evaluation, often used alongside precision and the F1-score to provide a balanced view of performance
Recent advances in machine learning have led to more nuanced applications, including multi-label and hierarchical classification, where recall is adapted to suit complex data structures
Academic foundations
The concept of recall is rooted in statistical decision theory and has been formalised in the context of information retrieval and pattern recognition since the mid-20th century
Current Landscape (2025)
Industry adoption and implementations
Recall is widely used in sectors such as healthcare, finance, and cybersecurity, where the cost of missing positive instances is high
Notable organisations and platforms
Google Cloud AI and Amazon SageMaker incorporate recall as a standard metric in their model evaluation suites
UK-based companies like eMed Healthcare UK (formerly Babylon Health, which collapsed in 2023) and Revolut use recall to optimise their diagnostic and fraud detection systems
UK and North England examples where relevant
In Manchester, NHS England’s digital teams (formerly NHS Digital, which merged into NHS England in 2023) employ recall to evaluate AI-driven diagnostic tools for early disease detection
Leeds City Council uses recall metrics in its smart city initiatives to identify and respond to public safety incidents
Newcastle University’s Institute for Data Science applies recall in research on predictive maintenance for industrial systems
Sheffield’s Advanced Manufacturing Research Centre (AMRC) leverages recall to ensure the reliability of AI models in manufacturing quality control
Technical capabilities and limitations
Recall is effective in identifying the completeness of positive case detection but can be misleading in imbalanced datasets if used in isolation
High recall often comes at the cost of increased false positives, which can be problematic in applications where precision is also critical
Standards and frameworks
Recall is included in major machine learning evaluation frameworks such as scikit-learn, TensorFlow, and PyTorch
The UK’s National Institute for Health and Care Excellence (NICE) recommends the use of recall in the evaluation of AI models for clinical decision support
Research & Literature
Key academic papers and sources
Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4), 427-437. https://doi.org/10.1016/j.ipm.2009.03.002
Powers, D. M. W. (2011). Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. Journal of Machine Learning Technologies, 2(1), 37-63. https://doi.org/10.5121/jmlt.2011.2103
Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 10(3), e0118432. https://doi.org/10.1371/journal.pone.0118432
Ongoing research directions
Researchers are exploring the integration of recall with other metrics to provide a more comprehensive evaluation of model performance
There is growing interest in developing adaptive recall metrics for dynamic and evolving datasets
UK Context
British contributions and implementations
The UK has been at the forefront of integrating recall into AI-driven healthcare and public sector applications
The Alan Turing Institute has published several studies on the use of recall in evaluating AI models for social good
North England innovation hubs (if relevant)
Manchester’s Digital Health Enterprise Zone is a leader in applying recall to improve the accuracy of AI diagnostics
Leeds’ Data City initiative uses recall to enhance the reliability of data-driven decision-making in urban planning
Newcastle’s Centre for Urban and Regional Development Studies (CURDS) applies recall in research on smart city technologies
Sheffield’s AMRC is pioneering the use of recall in advanced manufacturing and industrial AI
Regional case studies
A recent study by the University of Manchester demonstrated the effectiveness of recall in reducing false negatives in AI-driven cancer screening
Leeds City Council’s use of recall in its smart city platform has led to a significant improvement in the detection of public safety incidents
Future Directions
Emerging trends and developments
The integration of recall with other metrics to provide a more holistic view of model performance
The development of adaptive recall metrics for dynamic and evolving datasets
Anticipated challenges
Balancing recall with precision in imbalanced datasets
Ensuring the interpretability and transparency of recall-based evaluations
Research priorities
Developing new methods to optimise recall in multi-label and hierarchical classification tasks
Exploring the use of recall in real-time and streaming data environments
References
Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4), 427-437. https://doi.org/10.1016/j.ipm.2009.03.002
Powers, D. M. W. (2011). Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation. Journal of Machine Learning Technologies, 2(1), 37-63. https://doi.org/10.5121/jmlt.2011.2103
Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 10(3), e0118432. https://doi.org/10.1371/journal.pone.0118432