Three AI-assisted CNS pathology atlases for students, residents, and clinicians: created on Replit and audited by Gemini, created by Claude and audited by Codex, and created by Codex and audited by Claude. Explore synthetic teaching images alongside their audit findings, corrections, and limitations. AI review does not establish clinical accuracy.
Evidence review · September 20, 2026
Introduction to Atlas Evaluation
A convincing synthetic slide is not necessarily an accurate one. These three CNS pathology atlases pair AI-assisted creation with review by a different AI system: Replit with Gemini, Claude with Codex, and Codex with Claude. Evaluation asks whether the depicted structures match the stated diagnosis, whether descriptions and classifications are current, and whether errors and subsequent corrections are documented.
This is a review of the live interfaces and publicly available audit evidence—not a new slide-by-slide medical accuracy audit. The atlases differ in scope, image style, and review method, so their published findings cannot be used as a head-to-head accuracy ranking. Cross-model review can identify problems, but models can share errors and agreement is not proof of correctness.
What a rigorous evaluation should check
- Image fidelity: tissue architecture, cellular detail, stain appearance, and disease-defining structures—not just visual realism.
- Text and image agreement: whether described findings are actually visible, with appropriate terminology and molecular context.
- Audit transparency: reviewer, model/version, criteria, references, per-slide findings, corrections, and repeat review of revised images.
- Educational safety: clear synthetic-image labels, limitations, and comparison with verified real slides and specialist-reviewed references.
What the current evidence shows
Replit Atlas · Audited by Gemini
The public audit separates text review from image review: it reports 69 text entries reviewed (61 passed and 8 fixed), and 68 synthetic images initially rated as 53 inaccurate and 15 partially accurate. It documents regeneration actions for all 68 images.
Assessment: Its main strength is traceability: readers can inspect individual image ratings and the stated corrections. These are initial audit results, not an accuracy estimate for the regenerated images. The reviewed audit page does not establish a post-regeneration pass rate or specialist validation.
Read the published auditClaude Atlas · Audited by Codex
The live atlas displays 39 slides with schematic synthetic fields, stain options, microscopic features, and an Audit tab. Its public audit notes provide condition-specific caveats and references, including WHO CNS5 terminology updates and refinement of overly absolute diagnostic claims.
Assessment: Its main strength is detailed feedback on teaching text. Some recommendations remain visibly unresolved: the reviewed glioblastoma description calls pseudopalisading necrosis “pathognomonic,” while the audit recommends “highly characteristic.” The notes are not a standardized image-fidelity score, and schematic fields should not be mistaken for real histology.
Explore the atlas and Audit tabCodex Atlas · Audited by Claude
The site presents 30 generated teaching fields, category filters, key morphology, and pattern-based comparisons. Its Accuracy Audit discusses biological realism and limitations, including blurred fine details, repeated structures, homogeneous chromatin, and artificially merged cell borders.
Assessment: Its main strength is an explicit discussion of image artifacts and educational-use limits. The reviewed audit is qualitative and grouped by lesion category; it does not provide per-slide scores, a shared scoring rubric, or a measured post-correction accuracy rate. The Claude audit attribution is supplied by PathoLanding.
Read the atlas evaluationBottom line: These are educational resources with different levels of documented review, not clinically validated diagnostic tools. A fair comparison would require the same cases, a shared rubric, blinded specialist assessment, and repeat scoring after corrections. Use them alongside verified references and faculty guidance, never to diagnose a patient.