Skip to content
Retrospective research evidence

The results are promising. The limits are part of the result.

The figures below reproduce the manuscript's dataset roles and scope. They do not establish prospective performance, target-device performance, clinical utility, or regulatory readiness.

RetGuard manuscript performance summary
TaskDataset and roleNAUC (95% CI)Fixed-point or scope note
Diabetic retinopathyMessidor-2External1,7440.9691 (0.9595-0.9772)Sensitivity 97.16%; specificity 76.07% at fixed threshold 0.204983
GlaucomaAIROGSInternal8,8370.9792 (0.9721-0.9851)Internal evaluation; not strictly external
GlaucomaREFUGEStrictly external400 (40 positive)0.9278 (0.8642-0.9763)Sensitivity 70.00%; specificity 99.44% at fixed threshold
GlaucomaORIGAStrictly external650 (168 positive)0.8729 (0.8399-0.9028)Sensitivity 58.93%; specificity 92.53% at fixed threshold
OCT: Normal / AMD / DMEKermanyInternal11,1290.9990 (CI not stored)Internal B-scan evaluation; no stored confidence interval
OCT: AMD vs no AMDOCTIDLimited external261 (55 AMD; 0 DME)0.9999 (0.9995-1.0000)Supports AMD evaluation only; no valid external DME estimate

AUC values and operating-point figures are reported from the current Version 1 manuscript submitted to medRxiv. Dataset composition, confidence intervals, and task definitions should be read in the full manuscript before reuse. Public DOI assignment is pending screening.

Important negative and incomplete evidence

What the manuscript does not hide

Duke OCT claim withdrawn

Six reproduction variants failed the prespecified gate. The earlier Duke performance claim was withdrawn and is not presented as valid evidence.

No external DME estimate

OCTID contained no DME cases. Its result cannot validate external DME performance.

Broad 'Normal' label

In the 354-image OCTDL confounder set, 347 scans with genuine non-target pathology were labeled Normal for the AMD/DME task. Normal therefore means no AMD/DME detected, not a healthy retina.

Calibration and thresholds shifted

Good AUC did not guarantee stable probabilities or fixed operating points across external datasets.

Portability remains open

The DR ONNX export failed its prespecified logit-parity tolerance; other portability evidence is incomplete.

No clinical workflow validation

There was no prospective intent-to-screen study, target-device validation, human-factors study, or integrated clinical workflow evaluation.

The evidence supports a validation partnership - not clinical deployment

The right next step is independent replication and device-specific evaluation, followed by prospective work if performance, safety, and workflow gates are met.