Submitted:
27 August 2026
Posted:
28 August 2026
You are already at the latest version
Abstract
User behavior, content attributes, and real-time contextual features are deployed continuously and in a large scale search and advertising ranking systems. Current fixed-ratio gray scale rollouts are unable to change the degree of validation as a function of feature quality, service status or model impact and some failures are only identified once traffic volume is higher. This paper analyzes the history of deployment of features on an online major U.S. service. It uses the deployment increment, known as a complete production deployment, as the unit of analysis, and manual or automatic (within 24 hours of deployment) rollbacks as the primary outcome, forming a hierarchical Bayesian discrete-time survival model. There are four types of model inputs: data quality metrics (missed value rate, outlier proportion, distribution gap between online and offline data, default value trigger rate); service performance metrics (read latency, cache hit rate, request failure rate, version discrepancies between data centers); ranking quality metrics (RCE changes, calibration error, candidate score offset); and traffic stability metrics (exposure differences between user segments, regional request fluctuations). The research examined 8, 640 production release records from 1, 286 online features during an 18-month period, of which 612 were identified as rollbacks, and matched these with some 7.4 billion online prediction results. The data was split into training, validation and independent test sets according to the release time of the features, and the versions of neighbouring features were only kept within a single data set. Results of the test demonstrated a AUROC of 0.901, AUPRC of 0.714, Brier Score of 0.064, and calibration slope of 0.96, and a risk ratio for the high-risk release group of 3.47 (95%CI 2.89–4.18) after a 24-hour rollback. This was followed by a shadow check on 164 new releases. With the phased rollout of 1%, 5%, 10% or 25% of traffic based on the predicted risk, the median time to detect anomalies was reduced from 29 min to 12 min, the average time spent in the validation cycle was lowered from 5.8 days to 2.7 days, unnecessary rollback was decreased from 21 to 10, and fluctuations in online RCE were controlled within 0.20 percentage points. Captures data anomalies, service degradation and variations in ranking performance in one overall risk assessment, delivering a calibrated, quantitative basis for decisions on scaling up and rollback of features online.
Keywords:
online feature deployment
; rollback risk
; Bayesian survival model
; phased rollout
; risk calibration
; machine learning infrastructure
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.