This trial evaluates CITE, a retrieve-and-verify layer that audits an AI-generated care plan against a full-text evidence corpus and flags patient-specific codifiable safety hazards to the clinician. The co-primary outcomes are how accurately CITE flags these hazards (sensitivity and specificity versus blinded clinician adjudication) and its clinician alert burden and acceptance, compared with AI care plans using safety guardrails alone and with unassisted clinician care, in Medicaid primary care.
A Randomized Controlled Trial of CITE (Clinical Inference Tethered to Evidence), an Evidence-Grounding Retrieve-and-Verify Layer That Flags Unsupported and Inappropriate Recommendations in AI-Generated Care Plans, Versus AI With Safety Guardrails Alone and Unassisted Care, in Medicaid Primary Care
Patients are randomized 1:1:1 to (1) unassisted clinician care; (2) AI-generated care plan with safety guardrails; (3) AI-generated care plan with safety guardrails plus CITE. CITE audits the finalized plan against a frozen, versioned evidence corpus and returns physician-facing flags for patient-specific codifiable safety hazards (a recommended drug contraindicated by this patient's diagnosis or laboratory value; a drug-allergy conflict; a dropped high-risk medication; a guideline-indicated therapy omitted for an active diagnosis; a stated quantity refuted by the corpus), each with a verbatim quote and citation; the clinician retains decision authority. Randomization uses a deterministic HMAC permuted-block scheme; outcome assessors are blinded to arm. The co-primary outcomes are (1) the diagnostic accuracy (sensitivity and specificity) of CITE against blinded clinician adjudication, and (2) clinician alert burden (flags per encounter) and acceptance, comparing the CITE arm with the guardrail arm; both are estimable at the enrolled sample size because they do not depend on a rare between-arm event. The unresolved codifiable-hazard rate by arm is reported as a descriptive secondary: codifiable hazards are infrequent, so the trial is not powered for a between-arm efficacy contrast on hazard reduction. A prior trial of a different mechanism (a generic deterministic rule-corpus that surfaced roughly 30 or more flags per encounter and was uninformative) was completed with null results and is registered separately; this trial evaluates a materially different, patient-specific intervention and set of outcomes. Determined exempt by WCG IRB (low risk). Analysis is pre-registered on OSF (https://doi.org/10.17605/OSF.IO/ENXCW).