Research & evidence

A living evidence base behind the CelestyxAI design direction.

We track external clinical-AI research as an engineering input: what it demonstrates, what it does not demonstrate, and how it changes our architecture, validation strategy and intended use.

Evidence policy

External findings remain attributed to their original sources. CelestyxAI implications and briefs are our own evidence syntheses — not peer-reviewed CelestyxAI efficacy claims.

20 evidence records6 synthesis briefs2025–2026 tracked horizon
Evidence map

From publication to design decision.

Each record is read through the same lens: task validity, safety, human factors, governance, interoperability and clinical workflow fit.

01Evaluation
02Safety
03Human factors
04Interoperability
05Governance
06Clinical workflow
Global evidence ledger

Verified research, chronologically organized.

This ledger prioritizes work that materially informs CelestyxAI's current problem: safe, auditable, professional-facing clinical intake assistance. It is intentionally selective rather than exhaustive.

02 Jan 2025Nature MedicineSTUDY

Clinical LLM evaluation needs task-specific, expert-defined criteria rather than generic accuracy alone

EvaluationHuman factorsClinical workflow

Published finding

Johri and colleagues proposed a multidimensional framework for evaluating LLMs in clinical patient-interaction tasks, emphasizing clinical correctness, relevance, communication quality, safety and task-specific expert review.

CelestyxAI design implication

CelestyxAI should validate each delegated intake task with explicit rubrics and domain reviewers instead of treating a general medical benchmark as a proxy for workflow readiness.

Read original source ↗
08 Jan 2025Nature MedicineSTUDY

Medical LLMs can remain benchmark-strong while carrying poisoned medical knowledge

SafetyGovernanceEvaluation

Published finding

Replacing only 0.001% of training tokens with medical misinformation increased harmful medical outputs while conventional medical benchmarks remained largely unchanged.

CelestyxAI design implication

Knowledge provenance, authoritative external policy controls and downstream validation are safety requirements; benchmark performance cannot certify the integrity of a model's medical knowledge.

Read original source ↗
27 Jan 2025npj Digital MedicineSTUDY

Versatile medical LLM evaluation requires broader clinical task coverage

EvaluationClinical workflow

Published finding

MedS-Bench evaluates medical LLMs across 11 higher-level clinical tasks rather than relying on narrow question-answering performance.

CelestyxAI design implication

Evaluation should mirror the exact task taxonomy used in intake: extraction, normalization, synthesis, escalation support and structured output generation need separate measurements.

Read original source ↗
05 Feb 2025The BMJGUIDELINE

FUTURE-AI turns trustworthy healthcare AI into lifecycle requirements

GovernanceSafetyHuman factorsEvaluation

Published finding

An international consortium of 117 experts from 50 countries defined six principles — fairness, universality, traceability, usability, robustness and explainability — and 30 best practices spanning design through monitoring.

CelestyxAI design implication

Traceability, usability and post-deployment monitoring should be first-class product requirements, not compliance documentation added after the technical build.

Read original source ↗
06 Mar 2025npj Digital MedicineCOMMENT

Healthcare LLM implementation is a control, infrastructure and vendor-dependency problem — not only a model problem

GovernanceSafetyClinical workflow

Published finding

The authors frame LLM deployment around trade-offs between institutional control, collaboration, cost, data security, openness and dependence on external providers.

CelestyxAI design implication

CelestyxAI's orchestration layer should remain model-agnostic, version-aware and replaceable so institutional safety controls do not depend on one vendor's behavior or release cycle.

Read original source ↗
07 Mar 2025npj Digital MedicineSTUDY

Healthcare red-teaming exposes failure modes missed by standard evaluation

SafetyEvaluationHuman factors

Published finding

Clinician-led adversarial testing found 20.1% of 1,504 model responses inappropriate across safety, privacy, hallucination/accuracy and bias categories.

CelestyxAI design implication

Pre-pilot testing should contain deliberately adversarial, ambiguous and edge-case intake scenarios with clinical review and explicit failure classification.

Read original source ↗
17 Mar 2025npj Digital MedicineSTUDY

Task-specific deterministic tools can sharply reduce LLM errors in clinical calculations

SafetyEvaluationClinical workflow

Published finding

Across 10,000 trials, giving LLM agents access to task-specific clinical calculation tools produced the largest reductions in incorrect answers compared with unaided models.

CelestyxAI design implication

When a task can be represented deterministically, CelestyxAI should route it to an explicit tool or rule rather than ask the LLM to reproduce the logic probabilistically.

Read original source ↗
19 Mar 2025Applied SciencesSTUDY

LLMs can transform clinical text into FHIR, but prompting, correction and validation materially affect quality

InteroperabilityEvaluation

Published finding

The study compared multiple LLM approaches for converting clinical reports into FHIR bundles and found that examples and iterative correction influence conversion performance.

CelestyxAI design implication

FHIR-oriented generation should terminate in deterministic schema validation, terminology checks and rejectable outputs rather than treating syntactically plausible JSON as interoperable clinical data.

Read original source ↗
09 May 2025npj Digital MedicineSTUDY

LLM workflows show promise for triage and referral, while clinician judgment still matters in evaluation

Clinical workflowEvaluationHuman factors

Published finding

A 2,000-case MIMIC-IV evaluation tested LLM and RAG workflows for triage, specialty referral and diagnosis, while clinician scoring exposed meaningful inter-rater variation.

CelestyxAI design implication

CelestyxAI should not collapse professional disagreement into a single artificial ground truth; evaluation protocols need adjudication rules and transparent inter-rater measures.

Read original source ↗
12 May 2025OpenAIBENCHMARK

HealthBench pushes health-AI evaluation toward realistic expert-scored scenarios

EvaluationHuman factors

Published finding

HealthBench introduced realistic health conversations with physician-developed, case-specific rubrics and expert grading rather than simple multiple-choice accuracy.

CelestyxAI design implication

Use case-specific rubrics and failure severity should complement deterministic test suites; this is a useful benchmark design pattern, not evidence of CelestyxAI performance.

Read original source ↗
13 Jun 2025npj Digital MedicineSTUDY

Clinical AI evaluation is stronger when human review, automated metrics and simulation are combined

EvaluationHuman factorsClinical workflow

Published finding

The SCRIBE framework for ambient clinical scribing combines simulation, computational metrics, reviewer assessment and intelligent evaluation, including adversarial and fairness simulations.

CelestyxAI design implication

CelestyxAI's pre-pilot protocol should combine deterministic regression tests, scenario simulation, expert review and operational workflow measures instead of relying on one metric family.

Read original source ↗
07 Jul 2025npj Digital MedicineSTUDY

Bias auditing should be calibrated to a specific clinical population and risk tolerance

EvaluationGovernanceHuman factors

Published finding

A five-step framework links stakeholder engagement, local population calibration and clinically relevant scenario testing for auditing LLM accuracy and bias.

CelestyxAI design implication

Fairness evaluation must be local and workflow-specific. CelestyxAI should define subgroup analyses with partner institutions before interpreting aggregate performance.

Read original source ↗
07 Oct 2025npj Digital MedicineCOMMENT

The 'evaluation illusion' warns that medical LLM benchmarks can overstate real-world utility

EvaluationClinical workflowHuman factors

Published finding

The authors describe gaps between benchmark data, tasks, automated metrics and real-world translational impact, and call for context-aware evaluation and experimental transparency.

CelestyxAI design implication

CelestyxAI's evidence plan should measure extrinsic workflow outcomes — correction behavior, completeness, time, escalation handling and usability — in addition to output quality.

Read original source ↗
05 Nov 2025npj Digital MedicineSTUDY

LLM-as-a-judge can scale clinical summary evaluation, but only after anchoring to validated human instruments

EvaluationHuman factors

Published finding

A reasoning model showed strong agreement with human evaluators on a validated clinical summarization instrument and reduced evaluation time, while expert review remained the reference standard.

CelestyxAI design implication

Automated evaluators may help scale regression testing, but CelestyxAI should calibrate them against clinician-scored rubrics before using them as quality gates.

Read original source ↗
26 Dec 2025npj Digital MedicineSTUDY

Safety and effectiveness can diverge — especially in high-risk clinical scenarios

SafetyEvaluationClinical workflow

Published finding

CSEDB uses 30 safety/effectiveness metrics across 2,069 expert-authored cases and found a 13.3% performance drop in high-risk scenarios across tested models.

CelestyxAI design implication

Averages are unsafe as the primary acceptance criterion. CelestyxAI should weight red flags, critical-illness recognition and escalation failures by consequence severity.

Read original source ↗
06 Jun 2026BMC Health Services ResearchREVIEW

No single hallucination mitigation strategy is sufficient in healthcare

SafetyEvaluationGovernance

Published finding

A systematic review of 44 studies identified seven major mitigation families, including RAG, knowledge graphs, human-in-the-loop approaches, specialized evaluation and red teaming.

CelestyxAI design implication

CelestyxAI should use layered controls: authoritative rules, curated knowledge, constrained generation, schema validation, human review and adversarial testing rather than betting safety on prompting alone.

Read original source ↗
25 Jun 2026Cell Reports MedicineSTUDY

S.C.O.R.E. formalizes multidimensional expert evaluation of open-ended healthcare LLM outputs

EvaluationSafetyHuman factors

Published finding

S.C.O.R.E. evaluates Safety, Consensus & Context, Objectivity, Reproducibility and Explainability and showed that generic text-overlap metrics can misclassify clinically appropriate responses.

CelestyxAI design implication

CelestyxAI's expert review rubric should explicitly separate safety, evidence alignment, fairness/objectivity, reproducibility and explainability rather than merge them into one quality score.

Read original source ↗
23 Jul 2026PLOS Digital HealthSTUDY

LLM-based health-data mapping to FHIR is feasible, but systematic mapping errors remain

InteroperabilityEvaluationGovernance

Published finding

Mapping four common health data models to FHIR showed promising F1 scores but exposed semantic confusion, structural misalignment, ambiguous terminology and naming variance; the authors call for human validation and governance.

CelestyxAI design implication

CelestyxAI should treat interoperability as a governed transformation pipeline with terminology services, schema validation, provenance and reject/review paths — not as direct LLM-to-EHR generation.

Read original source ↗
19 Aug 2026NatureREVIEW

Healthcare LLM safety must be engineered across the full system lifecycle

SafetyGovernanceHuman factorsClinical workflow

Published finding

A broad review maps healthcare-LLM hazards across design, data, model, inference and deployment environment and identifies layered mitigations around human and system interactions.

CelestyxAI design implication

Clinical safety is an end-to-end system property. Model selection is one layer among policy governance, data integrity, access control, monitoring, human workflow and incident response.

Read original source ↗
20 Aug 2026Online Journal of Public Health InformaticsREVIEW

Healthcare LLM studies still under-evaluate privacy, security, robustness, explainability and verifiability

EvaluationSafetyGovernance

Published finding

A guideline-informed systematic review of 247 studies found evaluation concentrated on accuracy and fairness, while privacy protection appeared in 0.8% of studies and security assurance in none of the included studies.

CelestyxAI design implication

CelestyxAI's evidence matrix should make privacy, security, robustness, explainability and verifiability mandatory validation domains rather than optional technical appendices.

Read original source ↗
CelestyxAI scientific briefs

Evidence translated into a project-specific scientific direction.

These are original CelestyxAI syntheses grounded in the cited literature and the project's accumulated design decisions. They are clearly separated from external publications and are not presented as peer-reviewed papers.

CEL-SYN/01Evidence synthesis

Why the deterministic policy layer remains authoritative

A CelestyxAI evidence synthesis on separating safety-critical clinical constraints from probabilistic language-model behavior.

Grounded in 4 primary / authoritative sources
Open scientific brief →
CEL-SYN/02Evidence synthesis

From model score to workflow evidence: what CelestyxAI should prove before a pilot

A validation blueprint connecting modern medical-LLM evaluation research with early-stage clinical AI reporting principles.

Grounded in 5 primary / authoritative sources
Open scientific brief →
CEL-SYN/03Evidence synthesis

Human-in-the-loop is a control architecture, not a disclaimer

How professional review, override, calibration and audit should be engineered into clinical intake assistance.

Grounded in 4 primary / authoritative sources
Open scientific brief →
CEL-SYN/04Evidence synthesis

FHIR is an interface contract, not a trust shortcut

A CelestyxAI interoperability note on LLM-assisted clinical structuring, schema validation and governed exchange.

Grounded in 3 primary / authoritative sources
Open scientific brief →
CEL-SYN/05Evidence synthesis

A pre-pilot safety case for professional-facing clinical intake AI

A proposed evidence structure for testing rare, high-consequence failures before a controlled institutional validation.

Grounded in 5 primary / authoritative sources
Open scientific brief →
CEL-SYN/06Evidence synthesis

CelestyxAI scientific direction 2025–2026: from AI triage concept to governed intake infrastructure

A project synthesis connecting CelestyxAI's design evolution with the external evidence that now shapes its architecture and validation plan.

Grounded in 5 primary / authoritative sources
Open scientific brief →
Current scientific direction

The model is not the product boundary.

The emerging evidence increasingly supports the direction CelestyxAI has converged on: constrain probabilistic AI inside a governed clinical system whose policy, validation, audit, interoperability and human-control layers remain independently testable.

01Formalizable risk → deterministic control
02Generative task → bounded orchestration
03Output → schema + clinical validation
04Workflow → professional authority
05Integration → governed FHIR contract
06Deployment → continuous evidence
Research collaboration

Working on clinical AI safety, evaluation, human factors or interoperability?

Start a discussion