Submitted:
15 August 2026
Posted:
17 August 2026
You are already at the latest version
Abstract
Background. Digital glucose technologies include continuous glucose monitoring (CGM), emerging sensing modalities, multi-vendor platforms, forecasting models and artificial-intelligence (AI) briefings. These components are already chained into one patient journey, yet they are still often judged by a single convenient metric, such as MARD, a fused curve, RMSE/R² or fluent text. Visual continuity across brands can look like a medical record while stack safety for care remains unevaluated. Risk can therefore propagate along the product stack.Framework and contributions. We propose a four-layer, stack-wise evaluation framework for digital glucose care comprising sensing (L1), platform interoperability and provenance (L2), forecasting (L3), and human–AI briefing (L4). The framework makes three contributions. First, it organises the glucose product stack as successive clinical gates and identifies the false reassurance associated with each layer, complementing horizontal multidomain AI tools. Second, it specifies the cross-brand longitudinal account as an informatics object and treats L2 provenance as the hinge on which L3 and L4 may inherit trust, with acceptance criteria independent of sensor MARD. Third, it translates the gates into a brand-switching cascade test (Figure 3), a Have-you checklist for procurement and ethics (Table 3), an evidence-bound briefing template (Figure 2), and constrained deployment modes.Findings. Synthesising public evaluation and interoperability literature, we show that silent L2 splicing can carry forward as unchanged L3 score thresholds and categorical L4 prose even when L1 labelled claims are individually acceptable; provenance-visible segments, forecast abstention and evidence-bound briefings interrupt that cascade. Tables 1–3 and Figures 1–3 operationalise the gates.Conclusion. Stack-wise gating, with L2 provenance as the hinge, makes readiness for care an evaluable claim before procurement. A visually continuous multi-vendor glucose curve remains an interface achievement until layer-wise warrants are in place.
Keywords:
continuous glucose monitoring
; digital diabetes
; interoperability
; platform provenance
; multi-vendor platforms
; glucose forecasting
; clinical decision support
; AI safety
; stack-wise gating
1. Introduction
1.1. Evaluation Lag in a Stacked Digital Glucose Journey
Evaluation of digital health technologies often lags deployment. Tools move from pilot to procurement while assessment frameworks remain fragmented, heterogeneous or oriented to single outcomes that are easy to report [8,9]. In digital glucose care this lag is acute. Sensors, multi-vendor applications, forecasting models and AI briefings are already chained into one patient journey, yet assessment still collapses onto whichever metric is locally convenient: a MARD percentage for a sensor, a smooth curve on a phone screen, an RMSE reduction for a predictor, or fluent prose from a language model. Each of these artefacts can be locally true and clinically insufficient. When multi-brand timelines are plotted as one continuous curve, interface continuity is especially easy to mistake for continuity of a medical record.
The clinical stakes are not abstract. Glycaemic readings and alerts inform decisions about carbohydrate intake, insulin adjustment (where clinically authorised), driving, exercise, sleep and escalation to caregivers or clinicians. When evaluation stops at average point disagreement with a reference, it can under-specify performance in hypoglycaemia, in intended-use populations, and under real-world wear conditions. When evaluation stops at interface continuity across brands, it can under-specify whether successive device segments remain medically comparable. When evaluation stops at point forecast error, it can under-specify lead-time to dangerous events and false-alarm burden after the input stream changes. When evaluation stops at linguistic fluency, it can under-specify evidence binding, missed-event rates and human fallback. The problem is therefore not the absence of metrics, but the absence of a stack-wise clinical safety gate.
1.2. from Single Devices to Glucose Gateways
Digital glucose care is no longer a single device class. People living with diabetes and clinicians navigate among CGM sensors, capillary blood glucose meters (BGM), optical or other emerging sensing modalities, vendor-specific companion applications, third-party chronic-disease applications, and general health platforms that ingest glucose streams. A real user need has emerged for continuity when brands change because of cost, supply, insurance formulary, skin reaction, form factor or clinical recommendation. Clinicians likewise prefer not to open multiple vendor portals merely to reconstruct a fortnight of glycaemia. Industry attention is shifting from “who sells the sensor” toward “who owns the daily open frequency, the longitudinal data account, and the downstream advisory layer.”
That shift creates genuine value. It also creates a new class of false reassurance. A multi-brand device picker and a single plotted timeline can look like a medical record. In information-quality terms, however, fitness for clinical purpose depends on more than accessibility of a number on a screen [4]. Provenance, consistency, conformance to local units and statistical conventions, and maintainability of the account over time are part of whether the information remains usable for diagnosis, therapy and prognosis. Diabetes interoperability commentaries similarly distinguish device interoperability from data interoperability and emphasise ownership, privacy, liability and standards for sharing across manufacturers and into electronic records [5]. Empirical scans of regulated diabetes software further show that publicly accessible real-time application programming interfaces (APIs) remain scarce, so “connectable” user interfaces do not imply stable official access [6].
1.2B. from Public Debate to Layer-Specific Questions
Public debate about digital diabetes tools often oscillates between enthusiasm and scepticism. Enthusiasm emphasises access, continuity and patient empowerment. Scepticism emphasises regulatory ambiguity, data opacity and over-promise. Both instincts are incomplete if they remain slogan-level. Continuity is valuable when medically comparable; empowerment is valuable when alerts are warranted and escalatable; scepticism is warranted when fusion destroys provenance or when fluency substitutes for evidence. The four-layer framework converts these instincts into layer-specific questions that can be answered with dossiers, interfaces and protocols rather than with brand loyalty or blanket distrust.
1.3. Beyond Technical Validation: Without Reinventing Evaluation
Failures of clinical AI and digital tools more broadly often arise from weak algorithms and from inadequate evaluation, governance and implementation, including unintended consequences when machine-learning outputs are treated as authoritative without sociotechnical safeguards [1,17,23]. Multidomain clinical-AI tools, sociotechnical checklists, health-IT evaluation traditions and trustworthy-AI guidance already insist that deployability depends on more than offline technical scores [1,2,7,18,19,24]. Building on that inheritance rather than proposing a separate evaluation theory, this paper supplies a glucose-stack gate for sensing, platform accounts, forecasts and briefings. How the gate divides labour with NASSS, dHTA inventories and related stances is taken up in Section 2.6 and Table 1.
1.4. Aim, Innovations and Contributions
Accordingly, the premise of this paper is that a visually continuous multi-vendor glucose curve is an interface achievement until medical spliceability and downstream warrants are shown, and that stack-wise gating across sensing, platform provenance, forecasting and briefing is required before clinical adoption. Three questions organise the contribution:
- Why do favourable single-layer metrics fail to clear a chained glucose product stack?
- Under what informatics conditions is a cross-brand account medically spliceable, and how does that account become the trust hinge for forecasts and briefings?
- How can silent splicing propagate into prediction scores and natural-language claims, and what checklist and template can procurement and ethics review use to interrupt that path?
Contribution 1: Stack-wise gates (evidence:Table 2;Figure 1). We organise sensing, platform accounts, forecasts and briefings as successive clinical gates. Table 2 and Figure 1 specify the clinical question and common false reassurance at each layer, complementing horizontal multidomain tools such as CARES [1].
Contribution 2: Cross-brand account as an informatics hinge (evidence:Section 3.2). We specify the multi-vendor longitudinal account as an informatics object defined by provenance, authorised or standards-based access, and regulatory scope. Its acceptance criteria are independent of sensor MARD, and L2 acts as the hinge that determines what provenance L3 and L4 can inherit [4,5,6,14,15].
Contribution 3: Cascade stress test and governance artefacts (evidence:Figure 3;Table 3;Figure 2;Section 5.4). We apply a brand-switching stress test to forecasting and overnight briefing. The resulting governance artefacts comprise a Have-you checklist, an evidence-bound briefing template and constrained deployment modes specialised beyond generic clinical-AI life-cycle lists [2].
2. Evaluation Gaps and Related Frameworks
2.1. Gap A: Point Metrics as Proxy for Clinical Safety
Point-accuracy indices summarise average disagreement between a candidate measurement system and a reference. In CGM evaluation, MARD has become a lingua franca of marketing and many clinical summaries; RMSE and R² play analogous roles in forecasting papers. These indices are not meaningless. Average relative or absolute disagreement can support comparisons within a defined protocol, and regulators and clinicians routinely ask for them. International consensus statements on CGM use and on time-in-range interpretation nevertheless emphasise clinically oriented metrics and visualisation (including standardised reporting such as the ambulatory glucose profile) beyond a single accuracy number [10,11]. Methodological work further cautions that published MARD values are strongly influenced by study design and should not be treated as precise, freely comparable labels of sensor truth [12,13]. The gap is the leap from “average point disagreement looks acceptable” to “the system is clinically safe for the decisions it will influence.”
Several limitations are well recognised in clinical metrology discussions even when they are under-emphasised in product dashboards. First, averages can mask clinically asymmetric errors: the same MARD can hide different behaviour in hypoglycaemia, euglycaemia and hyperglycaemia. Second, study cohorts and wear conditions may not match the intended-use population or the conditions in which decisions are made (exercise, illness, medications, skin sites, adhesive failure, compression lows). Third, point metrics do not encode actionable timing: a sensor can be “accurate on average” while still being late, unstable during warm-up, or poorly characterised for trend arrows that users treat as decision cues. Fourth, when optical or placement-sensitive modalities are involved, localisation and wear geometry may participate in accuracy; treating the device as a black-box number source then under-specifies the sensing claim.
For evaluation practice, Gap A implies a gate: L1 must ask whether the reading is credible for the intended population, use condition and decision, not only whether a convenient cohort produced a favourable MARD or R². Favourable point metrics remain necessary for many sensing claims; stack-level adoption still requires the downstream gates in this framework.
2.2. Gap B: Multi-Vendor Gateway Apps Mistaken for Medical Records
Consumer diabetes and wellness applications increasingly offer what can be called a universal glucose entry: users authorise or bind devices from several CGM and BGM brands and view one spliced timeline, often alongside diet, medication and activity logs. The user need is authentic. Without aggregation, history fragments across vendor applications, report formats diverge, and brand switches break longitudinal comparison for both patients and clinicians. Unified time-in-range summaries, event definitions and clinician-facing reports are legitimate product goals.
The clinical risk is equally authentic. If manufacturer, model, algorithm or firmware version, warm-up status, calibration state, acquisition channel and quality flags are flattened into a series of timestamps and glucose values, visual continuity can be mistaken for medical comparability. Sampling frequencies differ; units differ; warm-up policies differ; labelled indications and accuracy claims differ; AGP-like statistical conventions may differ; trend arrows and on-device alarms may not transfer cleanly across brands. A visually continuous curve is an interface achievement. It is not automatically a continuous medical record.
Information-quality synthesis for digital health technologies places provenance among dimensions that support fitness for clinical purpose, alongside accuracy, interpretability, plausibility, relevance, completeness, timeliness, security, consistency, conformance and maintainability [4]. In that vocabulary, a gateway that maximises accessibility while destroying provenance improves one IQ category at the expense of informativeness. Diabetes interoperability commentaries add legal and systems constraints: combined products raise questions of liability; permission models may differ across manufacturers; open and transparent data-handling standards for routine care remain incomplete even as CGM-to-EHR data-standards agendas [14] and brief iCoDE steering-committee notes [15] proceed in specific jurisdictions [5,14,15]. Real-time access studies in national markets show that official public APIs are the exception rather than the rule, while do-it-yourself pathways face technical cardinality limits and legal tension with manufacturer terms [6]. Gap B therefore cannot be closed by UI polish alone.
The gateway pattern also interacts with secondary use. Longitudinal accounts are attractive for research and population analytics precisely because they appear continuous. If provenance is absent, secondary analyses may silently mix incompatible measurement systems and then report associations as if they were biologically homogeneous. Primary-care and research uses therefore share an interest in L2 discipline even when their governance pathways differ [5,6]. This paper therefore treats medical spliceability of the cross-brand account as the hinge layer (L2) of stack-wise adoption.
2.3. Gap C: Forecasts Tuned to Point Error Rather than Events
Forecasting models for glucose trajectories are often optimised and reported using point-wise error on relatively stationary series. Lower RMSE can be a useful engineering signal. Clinical safety questions, however, centre on whether dangerous events can be anticipated with usable lead-time, whether false alarms are tolerable in daily life, and whether performance remains meaningful when the input distribution shifts, including after a brand switch that changes sampling cadence, missingness patterns, warm-up segments or quality-flag regimes. A model trained under one vendor’s stream should not silently inherit trust when an undocumented second vendor enters the same longitudinal account. Gap C is thus tightly coupled to Gap B: forecasts consume the L2 account object; if that object is a spliced fiction, the forecast inherits the fiction.
2.4. Gap D: Natural-Language Alerts That Are Hard to Audit
AI briefings and conversational explanations are entering the same product chain. Fluent natural language can accelerate comprehension and action. It can also obscure which evidence, which device stream, which quality flags and which rule or model version justified a claim such as “overnight risk is rising.” Automation bias and alert fatigue are longstanding concerns in clinical decision support [16]; sociotechnical checklists explicitly ask whether alerts are relevant, timely and not overwhelming, and whether continuous monitoring and escalation paths exist after deployment [2]. Multidomain evaluation frameworks likewise emphasise engagement, usability, failure modes and governance [1]. Gap D is the briefing failure mode in which L4 does not inherit L2 platform provenance or expose human fallback.
2.5. Cascading Gaps
The four gaps compound. A reading that is “good enough” on average, a spliced account that is not medically comparable, a forecast tuned to point error, and an unauditable briefing form a cascade rather than four isolated defects. Stack-wise evaluation is required because local optimisation at one layer can increase risk at another.
2.6. Division of Labour with Related Frameworks
Scoping reviews of evaluation frameworks for digital health in chronic disease report heterogeneity of conceptual, results, logical and theory-of-change approaches, with limited evidence that consistent frameworks are applied through monitoring and evaluation cycles [8]. Scoping work on dHTA methodological frameworks likewise finds many tools and substantial terminological heterogeneity, while confirming that domains such as safety, clinical effectiveness, economic aspects, organisational and legal-regulatory issues (and, increasingly, interoperability-related concerns) recur across initiatives [9]. The Non-adoption, Abandonment, Scale-up, Spread and Sustainability (NASSS) framework theorises why technologies fail to become mainstream when complexity accumulates across condition, technology, value proposition, adopters, organisations and wider context [3]; that analysis is invaluable for predicting implementation fate and pairs with layer-wise clinical acceptance questions for care-facing reliance. Digital glucose care already has evaluative vocabulary. Relative to the speed of product chaining, what this paper adds is a stack-specific clinical safety gate (sensing → platform account → forecast → briefing), with L2 provenance as the hinge.
Related stances clarify complementary jobs (Table 1); the third column states what this paper adds beyond each stance.
Single habits (MARD alone, a fused app display alone, RMSE alone, or fluent AI text alone) therefore remain necessary but not sufficient. Our framework inherits HIT evaluation tradition [7] and specialises it to the digital glucose stack rather than a free-floating ethics slogan.
3. A Four-Layer Evaluation Framework
3.0. Overview and Interdependence
This section specifies the four-layer stack and the design choices (vertical gates; L2 as informatics hinge; acceptance cues rather than a single dashboard). Figure 1 summarises the stack as three readable layers of information: gates (clinical questions at L1–L4), false reassurances (the metric or UI artefact that typically substitutes for safety), and the hinge (L2 provenance as the critical fold-point of cascade risk when an upstream gate fails).
Layers are interdependent. A weak L2 turns L3 and L4 into amplifiers of an informatics fiction that L1 MARD cannot repair. Conversely, strong L1 evidence for each device does not license silent splicing at L2, nor does a validated forecast license unauditable prose at L4. The framework is iterative in the same spirit as multidomain evaluation tools [1]: gates should be revisited when populations, protocols, interfaces or models change.
3.1. L1 Sensing
Clinical question. Is the reading credible for the intended population, use condition and decision?
False reassurance. Favourable MARD / R² in a convenient cohort.
Scope. L1 covers CGM, BGM and emerging modalities (including optical approaches) insofar as they generate glucose-related measurements used for care or self-management. L1 states the clinical gate that sensing evidence must answer before downstream layers may treat the stream as trustworthy input; detailed metrological protocols remain with device dossiers and labelling.
Evaluation stance. First, intended population and use condition should be explicit: age bands, diabetes type, pregnancy, dialysis, critical illness, and supervised versus unsupervised use change the meaning of “validated.” Second, hypoglycaemia-relevant performance deserves separate attention because average error can conceal clinically costly misses or overestimates near decision thresholds. Third, wear and localisation assumptions should be stated where they participate in accuracy, skin site, adhesive integrity, compression, motion, optical coupling or vessel targeting for optical methods. Fourth, labelling and instructions for use remain part of the sensing claim: a number presented outside its labelled context is already an L1 failure even if the underlying hardware performed well in a trial.
Published sensing literature may illustrate that placement or visualisation constraints matter for certain modalities; L1 argumentation rests on public clinical-metrology logic. Favourable L1 evidence is a necessary pillar of responsible stacking and becomes clinically actionable when L2–L4 also clear their gates.
L1 acceptance cues (summary). Intended population and condition stated; hypoglycaemia-relevant evidence reviewed; wear/localisation assumptions explicit where they affect accuracy; labelled indications respected in downstream display and alerting.
Decision context matters: a reading used for long-term lifestyle reflection is not the same decision object as a reading used to interpret nocturnal symptoms or to support peri-procedural glycaemic assessment. L1 therefore begins by naming the decision class. Without that naming, point metrics float free of clinical meaning, and downstream layers inherit a number stripped of its warrant.
3.2. L2 Platform Interoperability and Provenance
Clinical question. After a brand switch (or concurrent multi-device use), is the longitudinal account still medically spliceable, and within what regulatory and interface boundaries?
False reassurance. “One-screen fusion”: a visually continuous multi-vendor curve, a single AGP-like report, or a device picker that merely lists brands as connectable.
L2 minimum inspectable set. Provenance fields × disclosed acquisition channels × semantic segment demarcation × regulatory-scope mapping. The subsections below operationalise that set for dossier review; open interchange ratification is a subsequent consensus step (Section 6.5).
3.2.1. Why L2 Is an Informatics Layer
L2 is the informatics layer increasingly treated as a glucose gateway: third-party chronic-care applications, ecosystems that ingest multiple brands, and general health platforms aggregating CGM/BGM streams. The product promise is a patient-owned multi-vendor diabetes data account, unified summaries, event definitions, clinician reports and continuity across device changes. That promise must be evaluated on its own terms, not collapsed into sensor MARD.
Diabetes therapy interoperability literature distinguishes device interoperability (components working together for an intended medical purpose) from data interoperability, exchange, interpretation and storage under common standards, including pathways into electronic medical records [5]. Both matter for glucose gateways. A platform may display numbers from multiple sensors without achieving either form of interoperability in a regulated sense. Conversely, formal interoperability initiatives (CGM-to-EHR data-standards agendas [14] and iCoDE steering-committee progress notes [15]) show that clinical reuse requires agreements on datasets, workflows and jurisdictional privacy regimes [5,14,15].
3.2.2. What Must Be Reconciled Across Brands
A capable L2 platform must reconcile, at minimum: sampling frequency; units; timestamps and time zones; missingness; sensor warm-up; outlier handling; calibration pathway; raw versus algorithmically processed values; trend arrows and device alarms; wear start/stop; device generation and firmware; AGP statistical conventions; and labelled indications and accuracy claims that differ by brand. Without retaining source metadata, a “continuous glucose map” may be only optically continuous. Consistency and conformance (presenting values in stable units and according to local or international conventions) are patient-safety issues, not cosmetic preferences [4].
3.2.3. Minimum Provenance Fields
Each glucose point or session segment should carry, at least:
- device manufacturer and model;
- sensor or session identifier;
- source application;
- acquisition method (manufacturer-official API; user-authorised official cloud export; standards-based health platform; or other disclosed channel);
- raw-or-processed flag;
- algorithm or firmware version;
- timestamp and time zone;
- unit;
- quality flag;
- warm-up status;
- calibration status.
Storing only {time, glucose} is insufficient for clinical reuse. Provenance here is not a bureaucratic extra; it is how clinicians and downstream models know when two segments must not be treated as exchangeable evidence [4].
3.2.4. Access Pathway Matters as Much as the UI
Preferential evaluation should favour manufacturer-official APIs, user-authorised official cloud exports, standards-based health-platform interfaces (including FHIR-oriented exchange patterns discussed in health-interoperability practice [21]), and written data-sharing agreements. Market evidence that public real-time APIs are scarce [6] strengthens rather than weakens this gate: scarcity means many “multi-brand” experiences will be tempted to rely on opaque pipelines. Unofficial decoding or third-party “modifier” algorithms that regenerate glucose from raw sensor signals change medical meaning. The platform may cease to be a data courier and become part of the measurement chain, with corresponding validation and regulatory-boundary obligations. Open-source code availability does not imply manufacturer authorisation, clinical equivalence to the official application, or coverage under a medical-device registration [5,6]. Likewise, a registration that covers a companion application for a specific meter does not automatically cover later multi-vendor CGM ingest, interpretive reports or decision-support features added to the same consumer application.
Evaluation therefore privileges disclosed, authorised or standards-based acquisition channels. Undisclosed or unofficial regeneration of glucose values is treated as an L2 failure mode until validation boundaries are explicit.
3.2.5. Neutrality Without Brand Polemics
Platforms that market device neutrality while remaining tightly coupled (commercially or via supply chain) to one sensor vendor face a sustainability question: will competing manufacturers grant stable, official data access over time? Device neutrality depends on transparent compatibility policies and formal interface agreements, not on a multi-brand picker alone. This is a governance question for procurement and for long-term clinical reliance on the account [5].
3.2.6. L2 Gates
Ask whether: (i) provenance is preserved end-to-end; (ii) brand-switch segments are visually and semantically demarcated; (iii) acquisition channels are disclosed and preferentially official or standards-based; (iv) software functions stay inside stated regulatory scope; and (v) downstream consumers of the account (forecasts, briefings, clinic reports) can see when streams are not comparable.
3.2.7. Clinician-Facing Exports and the “PDF Problem”
Even when on-screen fusion is carefully demarcated, clinician exports can reintroduce Gap B. A PDF or spreadsheet that drops provenance columns while keeping a single glucose column recreates the medically ambiguous series in a format that looks more official precisely because it is printable. L2 evaluation should therefore include export paths: what a physician sees in clinic matters as much as what a patient sees at home. Prefer exports that retain manufacturer, model, segment boundaries and quality flags. Where legacy clinic systems can accept only simple tables, the accompanying cover note should state non-equivalence across brand epochs rather than implying a single measurement system.
3.2.8. Caregivers, Proxies and Shared Accounts
Glucose gateways often support caregiver viewing, parents of children with diabetes, partners, or remote family. Shared viewing multiplies the harm of silent splicing: more people may act on a fused curve that is not medically continuous. L2 and L4 gates should ask whether caregiver interfaces inherit the same provenance visibility and escalation paths as the primary user interface. A briefing that withholds uncertainty from caregivers while showing it to clinicians (or the reverse) creates asymmetric risk. The framework does not prescribe a single sharing model; it requires that sharing not become a channel for carrying forward weak warrants into confident collective action.
3.3. L3 Forecasting
L3 consumes the L2 account object: forecasts and risk scores inherit whatever longitudinal series the platform constructs, including silent splices.
Clinical question. Can dangerous events be anticipated with usable lead-time and acceptable false-alarm burden when inputs change?
False reassurance. Lower point-wise prediction error on a stationary single-brand series.
Evaluation stance. Prefer event-centred endpoints (for example, anticipation of clinically defined hypoglycaemia or hyperglycaemia episodes with stated lead-time distributions) over exclusive reliance on RMSE/R². Report false-alarm burden in terms that matter to users and clinics (interruptions per week, nights affected, overrides). Require external robustness checks and explicit stress tests when L2 splices change sampling regimes or provenance. A model trained under one vendor’s cadence should not silently inherit trust after an undocumented brand switch, an instance of clinically meaningful dataset shift [22].
Forecasting evidence dossiers should meet the following methodological gate. Reviews of CGM sensing algorithms and of data-driven glucose prediction underline both the maturity of sensing pipelines and the heterogeneity of forecasting strategies [25,26]. If a forecast product consumes a multi-vendor account, its evidence dossier must address distribution shift at the account boundary together with in-distribution point error on clean single-brand test sets. Where uncertainty estimates exist, they should be exported to L4 so that briefings can carry calibrated language rather than categorical prose alone.
L3 acceptance cues (summary). Event lead-time and false-alarm metrics reported; robustness after device or provenance change addressed; dependence on L2 segment comparability documented.
3.4. L4 Human–AI Briefing
L4 likewise consumes the L2 account (and L3 outputs where present): natural-language claims must inherit segment, channel and quality-flag warrants rather than overwrite them.
Clinical question. Is the alert actionable, evidence-bound and escalatable?
False reassurance. Fluent natural-language “explanation” of a fused curve.
Evaluation stance. Require claim-to-evidence binding: which segment, which device stream, which quality flags, which rule or model version. Quantify missed events and false alarms for the briefing layer itself, beyond the upstream model alone. Define human fallback and escalation, who is notified, with what urgency, and what happens when provenance is incomplete or the model abstains. Alert burden and automation bias belong in the same gate as fluency [1,2]. L4 must inherit L2 provenance; narrative polish must not overwrite it.
For product evaluation, briefings should make disagreement and uncertainty visible rather than present probabilistic outputs as authoritative statements.
L4 acceptance cues (summary). Evidence binding including stream and flags; missed-event and false-alarm accounting; documented human fallback and escalation; alert burden considered.
3.3.1. Event-Centred Reporting Without Prescribing a Single Endpoint
Different clinical programmes may prioritise different events: Level 1 versus Level 2 hypoglycaemia, nocturnal events, hyperglycaemic emergencies, or composite risk scores used in research. The framework does not legislate one endpoint for all products. It requires that the chosen endpoint be event-like, that lead-time be reported relative to that endpoint, and that false alarms be counted in user-relevant units. A product that only reports RMSE after claiming “hypoglycaemia risk prediction” fails L3 regardless of engineering elegance.
3.3.2. Distribution Shift at the Brand Boundary
Brand switching is a clinically common distribution shift that forecasting papers often omit. Sampling intervals, missingness, warm-up policies and quality-flag semantics differ. Even when both brands are excellent sensors in isolation, the joint series is a new object. L3 evidence should either stay within brand segments or explicitly validate across the boundary. Transfer learning claims, if made, require the same honesty: transferred from which device distribution to which?
3.4.1. Evidence Templates for Briefings
An evidence-bound briefing can be implemented without revealing proprietary model weights. A minimal template includes: alert type; triggering window; device segment identifier; quality flags considered; model or rule version; confidence or abstention state; recommended human action; escalation contact. If any field is unavailable because L2 destroyed provenance, the briefing should degrade gracefully (toward uncertainty language and human review) rather than invent narrative certainty. Figure 2 schematises this template as an auditable interface pattern.
3.4.2. Fluency as a Risk Factor
Fluency is usually treated as a usability benefit. At L4 it is also a risk factor. The more natural the sentence, the easier it is to forget that the warrant may be thin. Evaluation should therefore include adversarial reading: can a sceptical clinician break the sentence back into evidence? If not, fluency has outrun auditability.
4. Brand-Switching Stress Test: Silent Splicing and Cross-Layer Cascade
Under labelled L1 claims that remain individually acceptable, silent L2 splicing is enough to legitimise an unchanged L3 score threshold and categorical L4 overnight prose; provenance-visible demarcation and forecast abstention (or segment-aware operation) interrupt that path (Figure 3). This stress test supplies Contribution 3’s cascade evidence through a weak-path versus strong-path comparison.
4.1. Scenario
Consider a patient who wears Brand A CGM for two weeks, migrates to Brand B for cost or supply reasons, and uses a third-party health application that displays both periods as one curve while logging meals and generating AI risk messages. The scenario is common enough to be clinically realistic and specific enough to stress every layer. Clinics can re-run the same walk-through against local dossiers [1].
4.2. Weak Versus Strong Paths and Cascade Figure
Weak path (Figure 3, top). Silent fusion of Brand A and Brand B; an unchanged forecast threshold after the switch; a fluent overnight claim without disclosed segment, flags or model version.
Strong path (Figure 3, bottom). Visually and semantically demarcated segments with retained provenance; forecast abstention or segment-aware operation across the boundary; uncertainty language plus escalation when warrants are incomplete.
One gating question per layer is annotated on the figure and restated below:
- L1: Were Brands A and B each validated for this patient’s population and use condition (including hypoglycaemia) on their own labelled claims?
- L2: Does the app retain manufacturer, model, algorithm version, warm-up and quality flags across the switch? Was Brand B ingested via an official authorised channel rather than an opaque pipeline? Are the two segments marked as non-equivalent for clinical comparison?
- L3: If a forecast or hypoglycaemia risk score was fitted on Brand A cadence, was it re-validated or gated after Brand B data entered the same account?
- L4: When the briefing says “your overnight risk is rising,” can a clinician or auditor see which segment, flags and rule version supported the sentence, and what happens if provenance is incomplete?
Figure 3.
Brand-switching stress test and cross-layer cascade. Research question on the figure: can silent L2 splicing carry forward as L3 scores and L4 prose when L1 labelled claims remain individually acceptable, and how do strong gates interrupt that cascade? Top row (weak path, vermillion): Panels A–C (silent fusion; unchanged risk threshold; fluent overnight claim). Bottom row (strong path, bluish green): Panels D–F (demarcated segments; forecast abstention or segment-aware operation; uncertainty briefing with escalation). Curves use arbitrary units (a.u.). Colour key (Okabe–Ito): Brand A = blue; Brand B = orange; weak path = vermillion; strong path = bluish green. Abbreviations: L1–L4 as in Figure 1.
Figure 3.
Brand-switching stress test and cross-layer cascade. Research question on the figure: can silent L2 splicing carry forward as L3 scores and L4 prose when L1 labelled claims remain individually acceptable, and how do strong gates interrupt that cascade? Top row (weak path, vermillion): Panels A–C (silent fusion; unchanged risk threshold; fluent overnight claim). Bottom row (strong path, bluish green): Panels D–F (demarcated segments; forecast abstention or segment-aware operation; uncertainty briefing with escalation). Curves use arbitrary units (a.u.). Colour key (Okabe–Ito): Brand A = blue; Brand B = orange; weak path = vermillion; strong path = bluish green. Abbreviations: L1–L4 as in Figure 1.

4.3. Cascade Analysis
Suppose L1 is provisionally acceptable for each device in isolation: both Brand A and Brand B have labelled claims that match the patient’s population. The dangerous failure mode begins at L2. If the application plots a continuous line across the switch without segment demarcation, users and clinicians may compare “week 1 versus week 3” as if the measurement system were constant. If warm-up points from Brand B are not flagged, early-session bias may be interpreted as clinical deterioration or improvement. If acquisition of Brand B occurred through an undisclosed channel, the medical meaning of the second segment is uncertain even when the plot looks perfect.
L3 then consumes the account. A predictor trained on Brand A’s sampling and missingness patterns may emit calibrated scores only under Brand A. After the switch, the same score threshold may not correspond to the same event lead-time. If the product does not gate or re-validate, the forecast silently carries an L2 defect forward as an apparent clinical risk signal.
L4 completes the cascade. A fluent sentence (“your overnight risk is rising”) can be correct relative to a model output and still be clinically misleading if the underlying segment is warm-up, quality-flagged, or drawn from a non-comparable brand without disclosure. Without evidence binding and escalation paths, neither patient nor clinician can distinguish “model believes risk is rising on a comparable series” from “model reacted to a splice artefact.”
If L2 answers are weak, L3 and L4 inherit a spliced fiction: downstream “intelligence” amplifies an informatics failure that no amount of MARD improvement at L1 can repair. Conversely, if L2 demarcates segments, preserves provenance and discloses channels, L3 and L4 can be designed to abstain, down-weight or explicitly conditionalise claims across the boundary. The stress test maps where false reassurance becomes actionable harm and why stack-wise gating is required.
4.4. Secondary Variants
The same skeleton applies to concurrent multi-device use (CGM plus intermittent BGM), to transitions between professional and consumer sensors, and to clinic-facing reports that fuse patient-imported files from multiple vendors. In each variant, the L2 questions (provenance, channel, demarcation, regulatory scope) remain the hinge.
4.5. What “Passing” Looks Like in the Stress Test
The strong path in Section 4.2 can be stated as dossier criteria without naming vendors. At L1, Brand A and Brand B each carry labelled evidence matching the patient’s population. Hypoglycaemia performance is discussed rather than hidden inside a single average. At L2, the application stores provenance for every point. The Brand A and Brand B epochs appear as visually distinct segments with an explicit notice that the segments may not be clinically equivalent. The acquisition channel for Brand B is disclosed in both the interface and audit log, and warm-up and quality flags remain available to downstream services. At L3, the forecast operates within a single-brand segment, uses a segment-aware model validated across the switch, or abstains when comparable history is insufficient. At L4, the briefing exposes the supporting segment, flags, model version and escalation contact. When provenance is incomplete, it uses uncertainty language and recommends human review.
Multi-vendor care remains admissible when provenance is visible; what the framework targets is silent multi-vendor care: fusion without provenance, forecasts without boundary handling, briefings without evidence.
4.6. Stress-Test Epilogue: Overnight Alert After a Switch
Return to the sentence “your overnight risk is rising,” issued three nights after migration from Brand A to Brand B. Under weak L2, the model may be responding to denser sampling, a warm-up tail, a unit conversion artefact, or a true physiological change; the user cannot tell. Under strong L2 and L3 abstention rules, the system might instead say: “Brand B segment is recent; overnight prediction is paused until comparable history accumulates; contact your care team if symptoms occur.” The second message is less “intelligent” in a marketing sense and more clinically honest. Stack-wise gating selects the second message whenever warrants are incomplete.
5. Checklist for Procurement, Governance and Ethics
5.1. Purpose
Multidomain evaluation tools and sociotechnical checklists emphasise actionability [1,2]. Table 3 compresses the stack-wise gates into pre-deployment conversation items for clinical leads, informatics, procurement and ethics review, so that technical elegance in any single layer cannot substitute for the others.
5.2. Layer-Wise Checklist
5.3. Using the Checklist
Hospitals, vendors and ethics boards can treat Table 3 as a gate: technical elegance in any one row does not clear the others. A sensor with strong L1 evidence does not approve an opaque multi-vendor account. A forecast with lower RMSE does not approve unauditable prose. A fluent briefing does not approve missing provenance. Where answers are negative, options include refusing deployment, restricting use to single-brand segments, requiring official APIs, disabling cross-segment forecasts, or forcing human review before patient-facing claims.
The checklist is also a documentation prompt. “Mapped to regulatory scope” forces an explicit sentence about which functions are medical-device functions and which are wellness features. “Disclosed acquisition channels” forces honesty about whether the gateway is an official courier or something else. These sentences matter for liability and for trust [5,6].
5.4. Constrained Deployment Modes When Gates Fail
Layer-wise evaluation supports five constrained modes that preserve user value while withholding unjustified clinical claims. These modes extend the framework beyond binary approve-or-reject decisions:
- Single-brand mode: allow display and coaching only within one manufacturer segment; disable cross-segment statistics.
- Provenance-visible mode: allow multi-brand display only if segment colours, labels and exportable provenance accompany every plot and report.
- Forecast-abstention mode: permit historical viewing but disable predictive scores for a washout period after a brand switch or when quality flags exceed a threshold.
- Human-in-the-loop briefing mode: generate draft language for clinician review before patient delivery.
- Clinic-import mode: accept official vendor reports as PDFs or certified exports without claiming a fused medical record inside the third-party application.
These five modes acknowledge the authentic user need behind glucose gateways while preventing the cascade described in Section 4, and give ethics and procurement committees operational alternatives when a gate fails.
5.5. Documentation Artefacts That Make Gates Auditable
Gates without artefacts become rhetoric. For L2, useful artefacts include a data dictionary of provenance fields; an acquisition-channel matrix (brand × channel × authorisation status); and sample clinician exports that show segment demarcation. For L3, useful artefacts include a validation report that states training-device mix, switch-handling rules and event-level metrics. For L4, useful artefacts include an evidence template for each alert type and an escalation standard operating procedure. For L1, the artefact is simply the labelled evidence dossier keyed to population and condition, already expected in device regulation, but often disconnected from app-level claims. Aligning these artefacts to Table 3 converts the checklist from a meeting agenda into an audit trail.
6. Discussion
6.1. Summary of Findings
Contribution 1. We organise sensing → platform account → forecast → briefing as successive clinical gates and name the false reassurance that typically substitutes for safety at each layer (Table 2; Figure 1). Relative to horizontal multidomain tools such as CARES [1], the increment is a vertical gate for the already-chained glucose product stack rather than another transversal AI checklist. Methodologically, this makes “readiness for care” a stack claim with inspectable acceptance cues. The claim is evaluative and governance-facing; population-specific thresholds remain in product dossiers (Section 6.5).
Contribution 2. We specify the cross-brand longitudinal account as an informatics object with four acceptance elements: provenance fields, disclosed authorised or standards-based channels, semantic segment demarcation, and regulatory-scope mapping. Section 3.2 defines L2 as the hinge that determines what provenance L3 and L4 can inherit. Building on information-quality and interoperability literature [4,5,6,14,15], the framework distinguishes interface connectivity from a medically spliceable and auditable account. Jurisdictional classification and quantitative cut-offs stay with regulators and local protocols.
Contribution 3. The brand-switching stress test shows that silent L2 splicing can carry an unchanged L3 threshold and categorical L4 overnight prose even when L1 labelled claims remain individually acceptable, while provenance-visible demarcation, forecast abstention or segment-aware operation, and uncertainty-plus-escalation interrupt that path (Figure 3; Section 4). Relative to generic clinical-AI life-cycle lists [2], the increment is cascade evidence plus governance artefacts: Table 3, Figure 2 and five constrained deployment modes (Section 5.4). These artefacts make stack readiness operable in procurement and ethics review. Incidence rates and named-product performance are outside this preprint’s scope (Section 6.5).
Taken together, readiness for care is a stack property, not a screenshot property, and L2 is the hinge on which L3 and L4 either amplify or disclose uncertainty. This paper addresses product-stack adoption gates; the relation of point-prediction error to event safety belongs to forecasting-evaluation work.
6.2. Interpretation: Why the Cascade Occurs and How the Design Interrupts It
The cascade is architectural rather than merely algorithmic. Downstream layers consume whatever longitudinal object L2 constructs. If L2 maximises accessibility while destroying provenance (sampling cadence, warm-up state, quality flags, acquisition channel, labelled indication), L3 and L4 inherit an informatics fiction that no incremental MARD improvement at L1 can repair. In information-quality terms, accessibility rises while fitness for diagnosis, therapy and prognosis erodes [4]. In interoperability terms, a UI that “connects” multiple brands without official real-time APIs or standards-based exchange can still look continuous while remaining medically opaque [5,6,14,15]. Forecasting then faces a clinically meaningful dataset shift at the brand boundary [22]; default trajectory thresholds or scores calibrated under one vendor’s missingness pattern need not retain the same event lead-time after splicing. Briefing systems compound the problem because fluent language invites automation bias: the more natural the sentence, the harder it is for patients and clinicians to demand the warrant [16].
The four-layer design interrupts that path by forcing claim–warrant alignment at each layer. L1 asks whether the reading is credible for the intended population, use condition and decision (including hypoglycaemia-relevant evidence) not merely whether an average relative difference looks favourable [10,11,12,13]. L2 asks whether successive brand segments are medically spliceable under disclosed channels and retained provenance. L3 asks for event-centred endpoints and robustness when inputs change, exporting uncertainty rather than discarding it for categorical prose [25,26]. L4 asks whether the sentence can be expanded into segment identifiers, flags, model versions, abstention states and escalation contacts (Figure 2). The five constrained deployment modes in Section 5.4 (single-brand display, provenance-visible multi-brand plots, forecast washout, human-in-the-loop briefing, clinic-import of official reports) are practical consequences of the same interpretation: user demand for continuity is real, but continuity without warrants is the failure mode the framework targets.
Stack-wise gating is therefore claim evaluation. It addresses substitution of a lower-layer warrant for a higher-layer claim. One example is using a favourable MARD to justify a fused AI coaching product whose account does not preserve segment provenance.
6.3. Relation to Prior Work
These basic findings tie well with multidomain arguments that clinical AI fails when evaluation stops at technical metrics and neglects data appropriateness, safety, usability and governance [1,17,18,19], and with cautions about unintended consequences of machine learning in medicine [23]. They are also consistent with sociotechnical models of health IT in complex adaptive systems and with checklist approaches that ask whether alerts are timely, non-overwhelming and backed by escalation paths after deployment [2,24]. Viewpoint literature urging developers not to reinvent health-IT evaluation theory supports specialising inherited technology–user–organisation dimensions to a concrete stack rather than inventing a free-floating ethics slogan [7].
NASSS [3] explains non-adoption, abandonment and scale-up fate under accumulating complexity. Layer-wise clinical acceptance questions for sensing, platform accounts, forecasts and briefings remain a separate instrument. Scoping reviews of chronic-disease digital-health evaluation and of dHTA methodological frameworks confirm both the abundance of tools and the terminological heterogeneity of safety, interoperability and regulatory domains [8,9]. This paper supplies a glucose-stack gate for the chain those inventories leave under-articulated.
CLIQ / information-quality synthesis places provenance among dimensions of fitness for clinical purpose [4]. Diabetes interoperability commentaries distinguish device from data interoperability and emphasise liability, permission models and incomplete open standards for routine care, including CGM-to-EHR data-standards agendas [5,14] and brief iCoDE steering-committee progress notes [15]. Real-time access scans show that public APIs remain scarce [6]. The present framework embeds those insights inside L2 of an adoption gate that also disciplines L3 and L4. International CGM and time-in-range consensus statements already push beyond a single accuracy number toward clinically oriented metrics and standardised visualisation [10,11]; methodological work on MARD further cautions against treating published averages as freely comparable sensor “truth labels” [12,13]. We inherit that caution at L1 and refuse to let a favourable MARD clear the rest of the stack.
SaMD clinical-evaluation guidance provides an internationally discussed vocabulary for software functions that inform clinical management [20]. The four-layer gate leaves jurisdictional classification to regulators and asks whether claimed functions at L2–L4 remain inside the labelled and registered envelope of the sensing components they display or interpret. FHIR-oriented exchange patterns in health-interoperability practice [21] illustrate the standards-based channel preference that L2 privileges. Reviews of CGM sensing algorithms and of data-driven glucose prediction underline both mature sensing pipelines and heterogeneous forecasting strategies [25,26]. These reviews support assessing L3 through event lead-time and boundary robustness, rather than through point error alone on clean single-brand test sets.
Where the literature is thinner is the explicit chaining of gateway fusion → forecast → natural-language briefing as one clinical journey. Product practice has moved faster than evaluative vocabulary. The brand-switching cascade is offered as a stress test that makes that under-articulated chain inspectable without naming commercial gateways: naming without systematic comparative measurement becomes anecdote; systematic measurement of unofficial pipelines risks sliding into operational detail outside the evaluative norms developed here. Clinics can instantiate Table 3 against local shortlists.
6.4. Implications and Refinement of Contribution
These findings suggest that contribution in digital glucose care should be restated as stack contribution. For vendors, a multi-brand plot does not by itself establish medical spliceability. Lower RMSE does not establish event-level robustness after brand switching, and fluent explanations do not establish briefing auditability. Relevant infrastructure includes official APIs, provenance schemas, segment-aware model gates and evidence templates. Device neutrality depends on transparent compatibility policies and formal interface agreements, not on a multi-brand picker alone.
For clinics, procurement and ethics boards, purchase dossiers that contain only MARD, fused-curve screenshots or marketing claims about AI coaching are incomplete. Table 3 supports refusal of false reassurance and supplies alternatives to binary approve/reject decisions via constrained modes. Ethics boards can require plain-language answers to the Have-you items; clinical teams can insist that multi-vendor reports demarcate brand segments before medication review. A thirty-minute rapid screen (identify claimed layers, demand one artefact per layer, walk Figure 3 aloud) prevents meetings from ending at demos (Section 5; Figure 3).
For research reporting, sensing, forecasting and briefing papers can reduce silent cross-layer generalisation by stating devices and firmware, warm-up inclusion, multi-vendor provenance mix and whether briefings may cite quality flags. Even without adopting L1–L4 labels, reporting those assumptions makes claims auditable. Dataset-shift awareness at brand boundaries should become a first-class item in external validation plans [22].
For clinical informatics, digital glucose care is a canonical cyber-physical-clinical stack: physical sensors, software platforms, predictive models and communicative interfaces jointly produce clinical meaning. Evaluation that attends only to user acceptance or algorithmic discrimination misses the distinctive failure mode, conversion of heterogeneous measurement systems into an apparently homogeneous personal data account. Analogous risks appear wherever wearable streams are fused across vendors into longitudinal dashboards; glucose is unusually clear because supply-driven brand switching is common, hypoglycaemia decisions are time-critical, and consumer applications already market multi-brand continuity.
The framework therefore states a portable evaluative norm for digital glucose care: stack readiness is judged layer by layer, with L2 provenance as the hinge, so that a visually continuous multi-vendor curve is treated as an interface object until medical spliceability and downstream warrants are shown. Universal glucose gateways solve a real fragmentation problem; they also elevate provenance, authorised interfaces and regulatory scope to first-class evaluation objects. Stack-wise gating permits multi-vendor care when segments remain provenance-visible, and targets silent fusion without warrants.
6.5. Limitations
Empirical force in this preprint rests on publicly cited literature and on the brand-switching stress test (Figure 3). Priority next steps are multi-vendor cohorts, procurement-site application of Table 3, and measurement of checklist operating characteristics in live review. The cascade finding is a structured clinical–informatics argument under common switching scenarios, ready for those empirical instantiations.
The provenance field list and briefing template are working schemas for open specifications and future consensus; jurisdictions differ, “regulatory scope” must be read under local rules, and SaMD versus wellness boundaries change over time [20]. Quantitative thresholds (labelled MARD claims, locally tolerated false-alarm rates, washout durations after a brand switch) belong in product dossiers and clinic protocols. Gateway descriptions stay generic so that Table 3 maps onto local shortlists without commercial anecdote. Concurrent multi-device use, professional-to-consumer transitions and clinic-facing fused imports remain secondary variants of the same L2 hinge and merit dedicated case studies. Human-factors evidence for evidence-bound briefings under alert burden is still thin relative to language-model productisation [16]; site-specific validation is listed as a deployment step once Figure 1, Figure 2 and Figure 3 have framed the gates.
6.6. Future Research
One important future direction is to operationalise L2 provenance schemas in open specifications (field dictionaries, acquisition-channel matrices and export formats that preserve segment demarcation) so that “medically spliceable” becomes a testable interface property. A second direction is segment-aware forecasting benchmarks that treat brand switches as first-class distribution shifts, reporting event lead-time, false-alarm burden and abstention rates alongside point error [22,25,26]. A third is human-factors and implementation studies of evidence-bound briefings: adversarial readability by clinicians, caregiver interface symmetry, alert-burden trade-offs and escalation protocol adherence. A fourth is empirical application of Table 3 in hospital procurement and ethics review, refining checklist length and wording against real dossiers. A fifth is cross-stack learning: testing whether the same vertical-gate pattern helps other multi-vendor wearable dashboards beyond glucose. Such work can feed vendors building official APIs, clinics that require provenance-visible fusion, and researchers who report layer assumptions as routine methods detail.
From the above discussion, digital glucose care should be judged as a stack: L1 credibility in the intended population and conditions; L2 provenance-preserving, authorised spliceability; L3 event-centred robustness across account boundaries; and L4 evidence-bound, escalatable language. Layered warrants make readiness for care evaluable.
7. Conclusions
This paper proposes a stack-wise evaluation framework for digital glucose care. L2 provenance acts as the hinge linking sensing, platform accounts, forecasting and human–AI briefing as successive clinical gates. The brand-switching stress test shows that silent splicing can carry into forecast thresholds and briefing language even when L1 labelled claims remain individually acceptable. Provenance-visible demarcation, abstention and evidence-bound configurations interrupt that path. Table 2 and Table 3 and Figure 1, Figure 2 and Figure 3 operationalise layer-wise readiness for procurement and governance. A visually continuous multi-vendor glucose curve constitutes a medical record only to the extent that the required warrants are documented.
Use of Artificial Intelligence
The manuscript was written by the authors. Qwen3.7 was used only for wording polish and grammar checks. The authors take full responsibility for the final text and verified that no AI-assisted edits altered scientific claims, citations, or figure content. No generative AI was used to create figures or to invent results.
Conflicts of Interest
The authors declare no conflicts of interest.
Data and Code Availability: All cited sources are publicly available through their original publishers. This literature synthesis releases no accompanying experimental code package.
Acknowledgments
We thank colleagues who commented on earlier drafts of this evaluation framework. This preprint synthesises publicly available literature.
References
- Heslin, S.M. Qualitative framework for evaluating clinical data science systems: beyond technical validation. Healthc. Inform. Res. 2026, 32(2), 196–199. [Google Scholar] [CrossRef]
- Owoyemi, A.; Osuchukwu, J.; Salwei, M.E.; Boyd, A. Checklist approach to developing and implementing AI in clinical settings: instrument development study. JMIRx Med. 2025, 6, e65565. [Google Scholar] [CrossRef] [PubMed]
- Greenhalgh, T.; Wherton, J.; Papoutsi, C.; Lynch, J.; Hughes, G.; A’Court, C.; et al. Beyond adoption: a new framework for theorizing and evaluating nonadoption, abandonment, and challenges to the scale-up, spread, and sustainability of health and care technologies (NASSS). J. Med. Internet Res. 2017, 19(11), e367. [Google Scholar] [CrossRef] [PubMed]
- Fadahunsi, K.P.; O’Connor, S.; Akinlua, J.T.; Wark, P.A.; Gallagher, J.; Carroll, C.; et al. Information quality frameworks for digital health technologies: systematic review (CLIQ). J. Med. Internet Res. 2021, 23(5), e23479. [Google Scholar] [CrossRef] [PubMed]
- Jendle, J.; Adolfsson, P.; Choudhary, P.; Dovc, K.; Fleming, A.; Klonoff, D.C.; et al. A narrative commentary about interoperability in medical devices and data used in diabetes therapy from an academic EU/UK/US perspective. Diabetologia 2024, 67(2), 236–245. [Google Scholar] [CrossRef] [PubMed]
- Randine, P.; Wolff, M.K.; Pocs, M.; Connell, I.R.O.; Cafazzo, J.A.; Årsand, E. Unlocking real-time data access in diabetes management: toward an interoperability model. J. Diabetes Sci. Technol. online ahead of print. 2025. [Google Scholar] [CrossRef] [PubMed]
- Cresswell, K.; de Keizer, N.; Magrabi, F.; Williams, R.; Rigby, M.; Prgomet, M.; et al. Evaluating artificial intelligence in clinical settings, let us not reinvent the wheel. J. Med. Internet Res. 2024, 26, e46407. [Google Scholar] [CrossRef] [PubMed]
- Bashi, N.; Fatehi, F.; Mosadeghi-Nik, M.; Askari, M.S.; Karunanithi, M. Digital health interventions for chronic diseases: a scoping review of evaluation frameworks. BMJ Health Care Inform. 2020, 27(1), e100066. [Google Scholar] [CrossRef] [PubMed]
- Segur-Ferrer, J.; Moltó-Puigmartí, C.; Pastells-Peiró, R.; Vivanco-Hidalgo, R.M. Methodological frameworks and dimensions to be considered in digital health technology assessment: scoping review and thematic analysis. J. Med. Internet Res. 2024, 26, e48694. [Google Scholar] [CrossRef] [PubMed]
- Danne, T.; Nimri, R.; Battelino, T.; Bergenstal, R.M.; Close, K.L.; DeVries, J.H.; et al. International consensus on use of continuous glucose monitoring. Diabetes Care 2017, 40(12), 1631–1640. [Google Scholar] [CrossRef] [PubMed]
- Battelino, T.; Danne, T.; Bergenstal, R.M.; Amiel, S.A.; Beck, R.; Biester, T.; et al. Clinical targets for continuous glucose monitoring data interpretation: recommendations from the international consensus on time in range. Diabetes Care 2019, 42(8), 1593–1603. [Google Scholar] [CrossRef] [PubMed]
- Reiterer, F.; Polterauer, P.; Schoemaker, M.; Schmelzeisen-Redecker, G.; Freckmann, G.; Heinemann, L.; et al. Significance and reliability of MARD for the accuracy of CGM systems. J. Diabetes Sci. Technol. 2017, 11(1), 59–67. [Google Scholar] [CrossRef] [PubMed]
- Kirchsteiger, H.; Heinemann, L.; Freckmann, G.; et al. Performance comparison of CGM systems: MARD values are not always a reliable indicator of CGM system accuracy. J. Diabetes Sci. Technol. 2015, 9(5), 1030–1040. [Google Scholar] [CrossRef] [PubMed]
- Espinoza, J.; Xu, N.Y.; Nguyen, K.T.; Klonoff, D.C. The need for data standards and implementation policies to integrate CGM data into the electronic health record. J. Diabetes Sci. Technol. 2023, 17(2), 495–502. [Google Scholar] [CrossRef] [PubMed]
- Yeung, A.M.; Huang, J.; Klonoff, D.C.; Seigel, R.E.; Goldman, J.M.; Shah, S.N.; et al. iCoDE June 22, 2022 Steering Committee meeting summary report. J. Diabetes Sci. Technol. 2022, 16(6), 1575–1576. [Google Scholar] [CrossRef] [PubMed]
- Lyell, D.; Coiera, E. Automation bias and verification complexity: a systematic review. J. Am. Med. Inform. Assoc. 2017, 24(2), 423–431. [Google Scholar] [CrossRef] [PubMed]
- Kelly, C.J.; Karthikesalingam, A.; Suleyman, M.; Corrado, G.; King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019, 17(1), 195. [Google Scholar] [CrossRef] [PubMed]
- Lekadir, K.; Frangi, A.F.; Porras, A.R.; Glocker, B.; Cintas, C.; Langlotz, C.P.; et al. FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ 2025, 388, e081554. [Google Scholar] [CrossRef] [PubMed]
- World Health Organization. Ethics and governance of artificial intelligence for health: WHO guidance; World Health Organization: Geneva; CC BY-NC-SA 3.0 IGO, 2021; Available online: https://www.who.int/publications/i/item/9789240029200.
- International Medical Device Regulators Forum (IMDRF) SaMD Working Group. Software as a Medical Device (SaMD): Clinical Evaluation. IMDRF/SaMD WG/N41FINAL:2017. 2017. Available online: https://www.imdrf.org/documents/software-medical-device-samd-clinical-evaluation.
- Benson, T.; Grieve, G. Principles of Health Interoperability: FHIR, HL7 and SNOMED CT, 4th ed.; Springer: Cham, 2021; ISBN 978-3-030-56882-5. [Google Scholar] [CrossRef]
- Finlayson, S.G.; Subbaswamy, A.; Singh, K.; Bowers, J.; Kupke, A.; Zittrain, J.; et al. The clinician and dataset shift in artificial intelligence. N. Engl. J. Med. 2021, 385(3), 283–286. [Google Scholar] [CrossRef] [PubMed]
- Cabitza, F.; Rasoini, R.; Gensini, G.F. Unintended consequences of machine learning in medicine. JAMA 2017, 318(6), 517–518. [Google Scholar] [CrossRef] [PubMed]
- Sittig, D.F.; Singh, H. A new sociotechnical model for studying health information technology in complex adaptive healthcare systems. Qual. Saf. Health Care 2010, 19 (Suppl 3), i68–i74. [Google Scholar] [CrossRef] [PubMed]
- Facchinetti, A. Continuous glucose monitoring sensors: past, present and future algorithmic challenges. Sensors 2016, 16(12), 2093. [Google Scholar] [CrossRef] [PubMed]
- Woldaregay, A.Z.; Årsand, E.; Walderhaug, S.; Albers, D.; Mamykina, L.; Botsis, T.; et al. Data-driven modeling and prediction of blood glucose dynamics: machine learning applications in type 1 diabetes. Artif. Intell. Med. 2019, 98, 109–134. [Google Scholar] [CrossRef] [PubMed]
Figure 1.
Four-layer safety evaluation stack for digital glucose care (evaluation architecture; synthesis of Section 2 and Section 3). Research question on the figure: which layer-wise gates, rather than a single dashboard metric, should decide care-facing adoption? Left to right: L1 Sensing; L2 Platform interoperability and provenance; L3 Forecasting; L4 Human–AI briefing (each panel: false reassurance above, clinical question below). The vermillion double arrow marks the cascade of risk when an upstream gate fails. L2 platform provenance is positioned as the hinge between sensing claims and downstream forecast and briefing trust. Bottom bar: stack-wise gates. Abbreviations: AI, artificial intelligence; MARD, mean absolute relative difference; R², coefficient of determination; L1–L4, the four stack layers.
Figure 1.
Four-layer safety evaluation stack for digital glucose care (evaluation architecture; synthesis of Section 2 and Section 3). Research question on the figure: which layer-wise gates, rather than a single dashboard metric, should decide care-facing adoption? Left to right: L1 Sensing; L2 Platform interoperability and provenance; L3 Forecasting; L4 Human–AI briefing (each panel: false reassurance above, clinical question below). The vermillion double arrow marks the cascade of risk when an upstream gate fails. L2 platform provenance is positioned as the hinge between sensing claims and downstream forecast and briefing trust. Bottom bar: stack-wise gates. Abbreviations: AI, artificial intelligence; MARD, mean absolute relative difference; R², coefficient of determination; L1–L4, the four stack layers.

Figure 2.
Minimal evidence template for a human–AI glucose briefing at L4 (interface pattern). Research question on the figure: how must a natural-language claim be bound to machine-readable evidence, and what happens when a required field is missing? Top bar: fields inherited from L2 provenance. Panel A: example claim (“Overnight risk is rising”). Panel B: eight evidence fields with illustrative values. Gate: Panel C releases an evidence-bound briefing when all required fields are present; Panel D degrades to uncertainty language plus human review when any required field is missing. Abbreviations: L2, platform interoperability and provenance; L4, human–AI briefing.
Figure 2.
Minimal evidence template for a human–AI glucose briefing at L4 (interface pattern). Research question on the figure: how must a natural-language claim be bound to machine-readable evidence, and what happens when a required field is missing? Top bar: fields inherited from L2 provenance. Panel A: example claim (“Overnight risk is rising”). Panel B: eight evidence fields with illustrative values. Gate: Panel C releases an evidence-bound briefing when all required fields are present; Panel D degrades to uncertainty language plus human review when any required field is missing. Abbreviations: L2, platform interoperability and provenance; L4, human–AI briefing.

Table 1.
Related frameworks and commentaries versus the four-layer glucose gate.
| Related stance | What it does well | What this paper adds |
|---|---|---|
| Multidomain clinical-AI evaluation beyond technical validation (e.g. Heslin CARES) [1] | Pre-/post-deployment domains; worked clinical scenario; practical evaluation tool framing | A vertical stack gate for glucose products (sensing→platform→forecast→briefing) |
| NASSS [3] | Non-adoption, abandonment, scale-up complexity | Layer-wise clinical acceptance questions for care-facing reliance |
| CLIQ / information quality [4] | Provenance and fitness-for-clinical-use dimensions | Glucose-gateway rules: spliceability, channels, segment demarcation |
| Diabetes device/data interoperability commentaries; real-time API models [5,6] | Device vs data interoperability; official API scarcity; liability/standards | Embedding interoperability into L2 of a four-layer adoption gate |
| Sociotechnical AI checklists (e.g. Owoyemi CASoF) [2] | Actionable “Have you…?” items across the AI life cycle | A short checklist specialised to glucose-stack layers (Table 3) |
| “Do not reinvent HIT evaluation” viewpoints [7] | Inherit technology–user–organisation theory | HIT-informed specialisation to digital glucose care |
Table 1 note. Structured contrast of published evaluation stances with the four-layer glucose gate (literature synthesis for dossier use). Columns: related stance; what it does well; what this paper adds. Rows cite representative frameworks and commentaries [1,2,3,4,5,6,7]. Abbreviations: AI, artificial intelligence; HIT, health information technology; L2, platform interoperability and provenance; NASSS, Non-adoption, Abandonment, Scale-up, Spread and Sustainability.*.
Table 2.
Layer-wise clinical questions, false reassurances and acceptance cues.
| Layer | Clinical question | Common false reassurance | Acceptance cues (this paper) |
|---|---|---|---|
| L1 Sensing | Is the reading credible for the intended population, use condition and decision? | Favourable MARD / R² | Population/condition alignment; hypoglycaemia range; wear/localisation where relevant |
| L2 Platform | After a brand switch, is the longitudinal account still medically spliceable, within which interface and regulatory bounds? | “One-screen fusion” | Official/standards-based channels; provenance fields; segment demarcation; regulatory scope |
| L3 Forecasting | Can dangerous events be anticipated with usable lead-time and acceptable false alarms when inputs change? | Lower point prediction error | Event-centred endpoints; robustness after provenance/sampling change |
| L4 Briefing | Is the alert actionable, evidence-bound and escalatable? | Fluent AI “explanation” of a fused curve | Claim–evidence binding; miss/false-alarm accounting; human fallback |
Table 3.
Checklist for procurement, clinical governance and ethics review.
| Layer | Checklist items (“Have you…?”) |
|---|---|
| L1 | Have you stated the intended population and use condition? Have you reviewed hypoglycaemia-relevant evidence? Are wear/localisation assumptions explicit where they affect accuracy? Are labelled indications respected in display and alerting? |
| L2 | Have you disclosed acquisition channels and preferred official or standards-based interfaces? Are per-point/session provenance fields retained? Are brand-switch segments demarcated visually and semantically? Are application functions mapped to regulatory scope? Have you ruled out silent third-party regeneration of glucose from raw signals without a validation boundary? Can downstream systems see when streams are not comparable? |
| L3 | Have you reported event lead-time and false-alarm metrics, not only point error? Have you addressed robustness after device or provenance change? Is dependence on L2 segment comparability documented? |
| L4 | Have you bound claims to evidence (including stream and flags)? Have you accounted for missed events and false alarms at the briefing layer? Is a human fallback and escalation path documented? Have you assessed alert burden and automation-bias risk? |
Table 3 note. Pre-deployment conversation aid for clinical leads, informatics, procurement and ethics review. Columns: layer; checklist items (“Have you…?”). Rows: L1–L4. Layer definitions as in Figure 1 and Table 2. Negative answers map to refusal, restriction, re-validation or mandatory human review.*.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.