CEL-SYN/02 · September 14, 2026
CelestyxAI scientific brief

From model score to workflow evidence: what CelestyxAI should prove before a pilot

A validation blueprint connecting modern medical-LLM evaluation research with early-stage clinical AI reporting principles.

A pre-pilot should not ask whether the model is 'good'. It should test whether the complete human-AI workflow behaves safely and usefully for a precisely defined intended use.

01

Evidence unit

The evaluation target is the workflow: intake completeness, correction rate, escalation handling, failure severity, professional acceptance, time burden and traceability — not a single benchmark score.

02

Human factors

DECIDE-AI and recent evaluation work both emphasize real users, training, operator behavior and small-scale live evaluation. Human override and disagreement are data to measure, not noise to hide.

03

Proposed CelestyxAI gate

Before prospective deployment, CelestyxAI should pass deterministic regression tests, simulated edge cases, clinician rubric review, subgroup checks and a documented governance readiness review with the partner institution.

This brief is a CelestyxAI synthesis of the cited external evidence and accumulated project design decisions. It is not a peer-reviewed scientific publication and does not establish clinical efficacy.

  1. 01https://www.bmj.com/content/377/bmj-2022-070904
  2. 02https://www.nature.com/articles/s41591-024-03328-5
  3. 03https://www.bmj.com/content/388/bmj-2024-081554
  4. 04https://www.nature.com/articles/s41746-025-01963-x
  5. 05https://www.nature.com/articles/s41746-025-01622-1
Evidence library

Return to the full research ledger.

Research & evidence