Metrics Glossary
This glossary is a cross-domain lookup index for evaluation metrics used across the wiki. Each row names the metric, gives its field of application, and links to the page that owns the fuller definition, formula, examples, and caveats.
Classification and Probability Calibration
Metric Field of application Meaning Accuracy Classification Fraction of examples whose predicted class equals the label. Precision Classification Fraction of predicted positives that are true positives. Recall Classification Fraction of actual positives found by the model. F1 score Classification Harmonic mean of precision and recall for a class or averaging scheme. Balanced accuracy Imbalanced classification Accuracy averaged across classes so a majority class cannot dominate the score. ROC-AUC Binary ranking and classification Threshold-free discrimination score based on the receiver operating characteristic curve. average precision Rare-positive classification Area-style summary of precision-recall ranking quality, often more informative under extreme imbalance. Log loss Probabilistic classification Negative log-likelihood that rewards calibrated probability assigned to the observed label. Brier score Probability calibration Mean squared error between predicted probability and binary outcome. Expected calibration error Probability calibration Weighted gap between predicted confidence and observed accuracy across bins. Sensitivity Clinical and safety classification Positive-class recall, usually emphasizing missed positives. Specificity Clinical and safety classification Fraction of actual negatives correctly rejected.
Regression
Ranking and Retrieval
Metric Field of application Meaning Precision@k Ranked retrieval Fraction of the visible top results that are relevant. Recall@k Ranked retrieval Fraction of known relevant results recovered in the top results. MAP Ranked retrieval Mean of average precision scores across queries. MRR Ranked retrieval Average reciprocal rank of the first relevant result. NDCG Ranked retrieval Rank-discounted graded relevance normalized by the ideal ordering. Rank of first expected source RAG retrieval benchmarks Position of the earliest labelled evidence source in the retrieved list.
Recommendation
Metric Field of application Meaning Top-k precision Recommender ranking Fraction of recommended items in a list that match the user’s held-out relevant set. Top-k recall Recommender ranking Fraction of the user’s held-out relevant items recovered by the recommendation list. Catalog coverage Recommender list health Portion of the item catalog or user space reached by recommendations. Intra-list diversity Recommender list health Degree to which items in the same recommendation list differ from one another. Novelty Recommender list health Degree to which recommended items are not already obvious or popular. Serendipity Recommender list health Degree to which recommendations are both unexpected and useful.
Time-Series Forecasting
Metric Field of application Meaning MAE Point forecasting Average absolute forecast error in the target unit. RMSE Point forecasting Square root of average squared forecast error, emphasizing large misses. MAPE Point forecasting Average absolute percentage error, unstable when actuals are zero or near zero. WAPE Demand and portfolio forecasting Total absolute error divided by total actual volume. MASE Cross-series forecasting Absolute error scaled by a naive or seasonal-naive baseline. Forecast bias Point forecasting Signed average error showing systematic overforecasting or underforecasting. pinball loss Probabilistic forecasting Asymmetric loss for evaluating a forecasted quantile. Empirical interval coverage Prediction intervals Fraction of realized values falling inside predicted intervals. Interval width Prediction intervals Size of the predicted interval, reported alongside coverage. Sharpness Probabilistic forecasting Concentration of a predictive distribution when calibration is acceptable. Forecast calibration Probabilistic forecasting Agreement between predicted quantiles or intervals and observed frequencies.
NLP and Text Generation
Metric Field of application Meaning Macro-F1 NLP classification Class-level F1 averaged equally across labels. Micro-F1 NLP classification F1 computed from aggregated counts across labels. Span F1 Sequence labelling and extraction F1 over matched spans rather than only token labels. Field exact match Information extraction Whether an extracted field value exactly matches the reference value. Field exact accuracy Information extraction Fraction of schema fields whose extracted values exactly match references. Character error rate OCR and transcription Edit distance normalized by reference character count. Word error rate OCR and speech-style transcription Word-level edit distance normalized by reference word count. Perplexity Language modelling Exponentiated average negative log-likelihood per predicted token. BLEU Machine translation and generation Reference-overlap score based on modified n-gram precision.
Computer Vision and Video
Metric Field of application Meaning Intersection over union Detection and segmentation Overlap divided by union between predicted and reference regions. Dice coefficient Segmentation Twice the overlap divided by combined predicted and reference region sizes. Average precision Object detection Precision-recall summary after confidence sorting and overlap-based matching. mAP Object detection Mean detection average precision across classes and often overlap thresholds. Pixel accuracy Semantic segmentation Fraction of pixels assigned the correct class. Mean IoU Semantic segmentation Average class-wise region overlap score. Panoptic Quality Panoptic segmentation Segment overlap penalized by unmatched predicted and reference segments. Boundary F-score Segmentation Boundary precision and recall with a spatial tolerance. Hausdorff distance Medical segmentation Worst nearest-surface error between predicted and reference boundaries. ASSD Medical segmentation Average symmetric surface distance between predicted and reference surfaces. Surface Dice Medical segmentation Fraction of surface points within an acceptable distance tolerance. PCK Pose estimation Fraction of visible keypoints within a normalized distance threshold. Temporal IoU Temporal localization Overlap divided by union for predicted and reference time segments. Temporal mAP Temporal action detection Mean average precision for time segments across temporal overlap thresholds.
Clustering and Representation
Metric Field of application Meaning Silhouette score Clustering Geometry-only score comparing within-cluster distance with nearest other-cluster distance. Adjusted Rand index Clustering with labels Chance-adjusted agreement between a clustering and reference labels.
Generative AI, RAG, and Agents
Metric Field of application Meaning Context recall RAG retrieval Fraction of expected evidence recovered into the model context. Citation precision RAG answers Fraction of cited sources that correspond to expected or supporting evidence. Citation coverage RAG answers Degree to which answer claims or required facts have citations. Answer support RAG answers Degree to which generated claims are backed by retrieved evidence. Claim support rate RAG answers Fraction of checked answer claims judged supported by evidence. Abstention quality RAG and generative systems Whether the system refuses or answers appropriately when evidence is missing. Task success Generative task systems Whether the final output satisfies the task-specific success criteria. Pass predicate Agent evaluation Boolean conjunction of outcome correctness, required actions, forbidden-action absence, and budget compliance. Budget compliance Agent evaluation Whether a trace stays within resource, latency, or call limits. Source coverage RAG and evaluation datasets Degree to which required source documents or evidence categories are exercised.
Experiment and Agreement Statistics
Metric Field of application Meaning p-value Hypothesis testing Probability of a statistic at least as extreme under a specified null model. Confidence interval Statistical estimation Repeated-sampling interval procedure with nominal long-run parameter coverage. Statistical power Experiment planning Probability of detecting a specified effect under the planned test design. Bootstrap interval Evaluation uncertainty Interval estimated by resampling examples and recomputing a statistic. Evaluation coverage Evaluation datasets Fraction of required slices, cases, sources, or paths represented in the evaluation. Raw agreement Human evaluation Fraction of reviewer labels that match before chance adjustment. Cohen’s kappa Human evaluation Reviewer agreement adjusted for expected chance agreement.
How to use this page
Use this glossary when a metric name appears before its full explanation. For study, jump from the metric to the owning subject area: classification and regression metrics usually live in classical machine learning, forecast metrics in time-series forecasting, ranked-list metrics in search or recommendation systems, and judge or trace metrics in generative AI and experimentation.