Duke OCT claim withdrawn
Six reproduction variants failed the prespecified gate. The earlier Duke performance claim was withdrawn and is not presented as valid evidence.
The figures below reproduce the manuscript's dataset roles and scope. They do not establish prospective performance, target-device performance, clinical utility, or regulatory readiness.
| Task | Dataset and role | N | AUC (95% CI) | Fixed-point or scope note |
|---|---|---|---|---|
| Diabetic retinopathy | Messidor-2External | 1,744 | 0.9691 (0.9595-0.9772) | Sensitivity 97.16%; specificity 76.07% at fixed threshold 0.204983 |
| Glaucoma | AIROGSInternal | 8,837 | 0.9792 (0.9721-0.9851) | Internal evaluation; not strictly external |
| Glaucoma | REFUGEStrictly external | 400 (40 positive) | 0.9278 (0.8642-0.9763) | Sensitivity 70.00%; specificity 99.44% at fixed threshold |
| Glaucoma | ORIGAStrictly external | 650 (168 positive) | 0.8729 (0.8399-0.9028) | Sensitivity 58.93%; specificity 92.53% at fixed threshold |
| OCT: Normal / AMD / DME | KermanyInternal | 11,129 | 0.9990 (CI not stored) | Internal B-scan evaluation; no stored confidence interval |
| OCT: AMD vs no AMD | OCTIDLimited external | 261 (55 AMD; 0 DME) | 0.9999 (0.9995-1.0000) | Supports AMD evaluation only; no valid external DME estimate |
AUC values and operating-point figures are reported from the current Version 1 manuscript submitted to medRxiv. Dataset composition, confidence intervals, and task definitions should be read in the full manuscript before reuse. Public DOI assignment is pending screening.
Six reproduction variants failed the prespecified gate. The earlier Duke performance claim was withdrawn and is not presented as valid evidence.
OCTID contained no DME cases. Its result cannot validate external DME performance.
In the 354-image OCTDL confounder set, 347 scans with genuine non-target pathology were labeled Normal for the AMD/DME task. Normal therefore means no AMD/DME detected, not a healthy retina.
Good AUC did not guarantee stable probabilities or fixed operating points across external datasets.
The DR ONNX export failed its prespecified logit-parity tolerance; other portability evidence is incomplete.
There was no prospective intent-to-screen study, target-device validation, human-factors study, or integrated clinical workflow evaluation.
The right next step is independent replication and device-specific evaluation, followed by prospective work if performance, safety, and workflow gates are met.