From model score to workflow evidence: what CelestyxAI should prove before a pilot
A validation blueprint connecting modern medical-LLM evaluation research with early-stage clinical AI reporting principles.
A pre-pilot should not ask whether the model is 'good'. It should test whether the complete human-AI workflow behaves safely and usefully for a precisely defined intended use.
Evidence unit
The evaluation target is the workflow: intake completeness, correction rate, escalation handling, failure severity, professional acceptance, time burden and traceability — not a single benchmark score.
Human factors
DECIDE-AI and recent evaluation work both emphasize real users, training, operator behavior and small-scale live evaluation. Human override and disagreement are data to measure, not noise to hide.
Proposed CelestyxAI gate
Before prospective deployment, CelestyxAI should pass deterministic regression tests, simulated edge cases, clinician rubric review, subgroup checks and a documented governance readiness review with the partner institution.
This brief is a CelestyxAI synthesis of the cited external evidence and accumulated project design decisions. It is not a peer-reviewed scientific publication and does not establish clinical efficacy.
