This ledger prioritizes work that materially informs CelestyxAI's current problem: safe, auditable, professional-facing clinical intake assistance. It is intentionally selective rather than exhaustive.
01
02 Jan 2025Nature MedicineSTUDY
Clinical LLM evaluation needs task-specific, expert-defined criteria rather than generic accuracy alone
EvaluationHuman factorsClinical workflow
Published finding
Johri and colleagues proposed a multidimensional framework for evaluating LLMs in clinical patient-interaction tasks, emphasizing clinical correctness, relevance, communication quality, safety and task-specific expert review.
CelestyxAI design implication
CelestyxAI should validate each delegated intake task with explicit rubrics and domain reviewers instead of treating a general medical benchmark as a proxy for workflow readiness.
Read original source ↗02
08 Jan 2025Nature MedicineSTUDY
Medical LLMs can remain benchmark-strong while carrying poisoned medical knowledge
SafetyGovernanceEvaluation
Published finding
Replacing only 0.001% of training tokens with medical misinformation increased harmful medical outputs while conventional medical benchmarks remained largely unchanged.
CelestyxAI design implication
Knowledge provenance, authoritative external policy controls and downstream validation are safety requirements; benchmark performance cannot certify the integrity of a model's medical knowledge.
Read original source ↗03
27 Jan 2025npj Digital MedicineSTUDY
Versatile medical LLM evaluation requires broader clinical task coverage
EvaluationClinical workflow
Published finding
MedS-Bench evaluates medical LLMs across 11 higher-level clinical tasks rather than relying on narrow question-answering performance.
CelestyxAI design implication
Evaluation should mirror the exact task taxonomy used in intake: extraction, normalization, synthesis, escalation support and structured output generation need separate measurements.
Read original source ↗04
05 Feb 2025The BMJGUIDELINE
FUTURE-AI turns trustworthy healthcare AI into lifecycle requirements
GovernanceSafetyHuman factorsEvaluation
Published finding
An international consortium of 117 experts from 50 countries defined six principles — fairness, universality, traceability, usability, robustness and explainability — and 30 best practices spanning design through monitoring.
CelestyxAI design implication
Traceability, usability and post-deployment monitoring should be first-class product requirements, not compliance documentation added after the technical build.
Read original source ↗05
06 Mar 2025npj Digital MedicineCOMMENT
Healthcare LLM implementation is a control, infrastructure and vendor-dependency problem — not only a model problem
GovernanceSafetyClinical workflow
Published finding
The authors frame LLM deployment around trade-offs between institutional control, collaboration, cost, data security, openness and dependence on external providers.
CelestyxAI design implication
CelestyxAI's orchestration layer should remain model-agnostic, version-aware and replaceable so institutional safety controls do not depend on one vendor's behavior or release cycle.
Read original source ↗06
07 Mar 2025npj Digital MedicineSTUDY
Healthcare red-teaming exposes failure modes missed by standard evaluation
SafetyEvaluationHuman factors
Published finding
Clinician-led adversarial testing found 20.1% of 1,504 model responses inappropriate across safety, privacy, hallucination/accuracy and bias categories.
CelestyxAI design implication
Pre-pilot testing should contain deliberately adversarial, ambiguous and edge-case intake scenarios with clinical review and explicit failure classification.
Read original source ↗07
17 Mar 2025npj Digital MedicineSTUDY
Task-specific deterministic tools can sharply reduce LLM errors in clinical calculations
SafetyEvaluationClinical workflow
Published finding
Across 10,000 trials, giving LLM agents access to task-specific clinical calculation tools produced the largest reductions in incorrect answers compared with unaided models.
CelestyxAI design implication
When a task can be represented deterministically, CelestyxAI should route it to an explicit tool or rule rather than ask the LLM to reproduce the logic probabilistically.
Read original source ↗08
19 Mar 2025Applied SciencesSTUDY
LLMs can transform clinical text into FHIR, but prompting, correction and validation materially affect quality
InteroperabilityEvaluation
Published finding
The study compared multiple LLM approaches for converting clinical reports into FHIR bundles and found that examples and iterative correction influence conversion performance.
CelestyxAI design implication
FHIR-oriented generation should terminate in deterministic schema validation, terminology checks and rejectable outputs rather than treating syntactically plausible JSON as interoperable clinical data.
Read original source ↗09
09 May 2025npj Digital MedicineSTUDY
LLM workflows show promise for triage and referral, while clinician judgment still matters in evaluation
Clinical workflowEvaluationHuman factors
Published finding
A 2,000-case MIMIC-IV evaluation tested LLM and RAG workflows for triage, specialty referral and diagnosis, while clinician scoring exposed meaningful inter-rater variation.
CelestyxAI design implication
CelestyxAI should not collapse professional disagreement into a single artificial ground truth; evaluation protocols need adjudication rules and transparent inter-rater measures.
Read original source ↗10
12 May 2025OpenAIBENCHMARK
HealthBench pushes health-AI evaluation toward realistic expert-scored scenarios
EvaluationHuman factors
Published finding
HealthBench introduced realistic health conversations with physician-developed, case-specific rubrics and expert grading rather than simple multiple-choice accuracy.
CelestyxAI design implication
Use case-specific rubrics and failure severity should complement deterministic test suites; this is a useful benchmark design pattern, not evidence of CelestyxAI performance.
Read original source ↗11
13 Jun 2025npj Digital MedicineSTUDY
Clinical AI evaluation is stronger when human review, automated metrics and simulation are combined
EvaluationHuman factorsClinical workflow
Published finding
The SCRIBE framework for ambient clinical scribing combines simulation, computational metrics, reviewer assessment and intelligent evaluation, including adversarial and fairness simulations.
CelestyxAI design implication
CelestyxAI's pre-pilot protocol should combine deterministic regression tests, scenario simulation, expert review and operational workflow measures instead of relying on one metric family.
Read original source ↗12
07 Jul 2025npj Digital MedicineSTUDY
Bias auditing should be calibrated to a specific clinical population and risk tolerance
EvaluationGovernanceHuman factors
Published finding
A five-step framework links stakeholder engagement, local population calibration and clinically relevant scenario testing for auditing LLM accuracy and bias.
CelestyxAI design implication
Fairness evaluation must be local and workflow-specific. CelestyxAI should define subgroup analyses with partner institutions before interpreting aggregate performance.
Read original source ↗13
07 Oct 2025npj Digital MedicineCOMMENT
The 'evaluation illusion' warns that medical LLM benchmarks can overstate real-world utility
EvaluationClinical workflowHuman factors
Published finding
The authors describe gaps between benchmark data, tasks, automated metrics and real-world translational impact, and call for context-aware evaluation and experimental transparency.
CelestyxAI design implication
CelestyxAI's evidence plan should measure extrinsic workflow outcomes — correction behavior, completeness, time, escalation handling and usability — in addition to output quality.
Read original source ↗14
05 Nov 2025npj Digital MedicineSTUDY
LLM-as-a-judge can scale clinical summary evaluation, but only after anchoring to validated human instruments
EvaluationHuman factors
Published finding
A reasoning model showed strong agreement with human evaluators on a validated clinical summarization instrument and reduced evaluation time, while expert review remained the reference standard.
CelestyxAI design implication
Automated evaluators may help scale regression testing, but CelestyxAI should calibrate them against clinician-scored rubrics before using them as quality gates.
Read original source ↗15
26 Dec 2025npj Digital MedicineSTUDY
Safety and effectiveness can diverge — especially in high-risk clinical scenarios
SafetyEvaluationClinical workflow
Published finding
CSEDB uses 30 safety/effectiveness metrics across 2,069 expert-authored cases and found a 13.3% performance drop in high-risk scenarios across tested models.
CelestyxAI design implication
Averages are unsafe as the primary acceptance criterion. CelestyxAI should weight red flags, critical-illness recognition and escalation failures by consequence severity.
Read original source ↗16
06 Jun 2026BMC Health Services ResearchREVIEW
No single hallucination mitigation strategy is sufficient in healthcare
SafetyEvaluationGovernance
Published finding
A systematic review of 44 studies identified seven major mitigation families, including RAG, knowledge graphs, human-in-the-loop approaches, specialized evaluation and red teaming.
CelestyxAI design implication
CelestyxAI should use layered controls: authoritative rules, curated knowledge, constrained generation, schema validation, human review and adversarial testing rather than betting safety on prompting alone.
Read original source ↗17
25 Jun 2026Cell Reports MedicineSTUDY
S.C.O.R.E. formalizes multidimensional expert evaluation of open-ended healthcare LLM outputs
EvaluationSafetyHuman factors
Published finding
S.C.O.R.E. evaluates Safety, Consensus & Context, Objectivity, Reproducibility and Explainability and showed that generic text-overlap metrics can misclassify clinically appropriate responses.
CelestyxAI design implication
CelestyxAI's expert review rubric should explicitly separate safety, evidence alignment, fairness/objectivity, reproducibility and explainability rather than merge them into one quality score.
Read original source ↗18
23 Jul 2026PLOS Digital HealthSTUDY
LLM-based health-data mapping to FHIR is feasible, but systematic mapping errors remain
InteroperabilityEvaluationGovernance
Published finding
Mapping four common health data models to FHIR showed promising F1 scores but exposed semantic confusion, structural misalignment, ambiguous terminology and naming variance; the authors call for human validation and governance.
CelestyxAI design implication
CelestyxAI should treat interoperability as a governed transformation pipeline with terminology services, schema validation, provenance and reject/review paths — not as direct LLM-to-EHR generation.
Read original source ↗19
19 Aug 2026NatureREVIEW
Healthcare LLM safety must be engineered across the full system lifecycle
SafetyGovernanceHuman factorsClinical workflow
Published finding
A broad review maps healthcare-LLM hazards across design, data, model, inference and deployment environment and identifies layered mitigations around human and system interactions.
CelestyxAI design implication
Clinical safety is an end-to-end system property. Model selection is one layer among policy governance, data integrity, access control, monitoring, human workflow and incident response.
Read original source ↗20
20 Aug 2026Online Journal of Public Health InformaticsREVIEW
Healthcare LLM studies still under-evaluate privacy, security, robustness, explainability and verifiability
EvaluationSafetyGovernance
Published finding
A guideline-informed systematic review of 247 studies found evaluation concentrated on accuracy and fairness, while privacy protection appeared in 0.8% of studies and security assurance in none of the included studies.
CelestyxAI design implication
CelestyxAI's evidence matrix should make privacy, security, robustness, explainability and verifiability mandatory validation domains rather than optional technical appendices.
Read original source ↗