research.embrace.kcl.ac.uk/ai/evaluation
LIVEStudy Coordinator
Admin indexAI Operations: HAIAEvaluation & experiments

Capability 5.6 · Screen 51

Evaluation & experiments

Offline eval runs, A/B and shadow lanes, MRT variant comparison, reproducible outputs export. 5.6.

SUITE v3.2

Offline evaluation

  • Safety recall 99.2% ≥99% · citation precision 97.4% ≥95%
  • Unsupported advice 0.2% ≤0.5% · escalation accuracy 98.8% ≥98%
EXPERIMENT LANES

Online evidence

  • A/B: A 602 · B 602
  • Shadow delta 1.8% · MRT decision record linked
  • Export reproducible bundle

Group H — AI Operations: HAIA · Screen 51

Evaluation & experiments

Offline eval runs, A/B and shadow lanes, MRT variant comparison, reproducible outputs export. 5.6.

5.6