We publish how well our safety engine works.
Most clinical AI tools claim accuracy. We measured ours against a fixed benchmark, with the right answers written down before testing began — and we publish the result, including what it got wrong.
Why the safety layer is not AI
A language model produces an answer that sounds right. That is acceptable when it drafts a sentence. It is not acceptable when it decides whether a medicine is safe for a particular patient, because the failure mode is not an awkward sentence — it is a patient.
So Carenika splits the work. The AI listens and writes. A separate rule engine — explicit, versioned, written by hand — checks the drugs. Same input, same output, every time. Every alert points back to a named rule a clinician can read, question, and overrule.
What the engine checks
Drug–drug interactions · duplicate therapy and class overlap · contraindications against recorded conditions · eGFR-banded renal dosing · hepatic ceilings · maximum daily dose · paediatric mg/kg review · pregnancy rules · an ADR registry that remembers a patient’s past reactions.
What the doctor sees
A Shield score for the prescription, and for each alert: why it fired, what to do instead, and the evidence behind it. Overrides are permitted and logged — the system advises, the physician decides.
50 scenarios. Reference standard fixed before testing.
Drug interactions, comorbidity contraindications, duplicate therapy and disease-specific rules. Thirty-one scenarios require an alert. Nineteen are deliberately safe — because a safety engine that fires on everything is as useless as one that fires on nothing, and false alarms have to be counted.
50
outpatient scenarios
31 / 19
alerts expected vs deliberate safe controls
1.00
specificity on the held-out run — no false alarms
11 / 12
misses traced to deployment, not rule design
Results have been submitted for peer review to ICPS-2026, North South University, Dhaka. The author is the developer of the system under evaluation; the reference standard was fixed before testing to reduce assessment bias.
What this evidence does not prove.
A benchmark is fifty scenarios; clinical practice is not. The cases were written by us, which is why the held-out block — run once, before any disease-specific rules were added — matters more than the full re-run that followed the fixes. We report both. Carenika is decision support: it documents and it warns. It does not diagnose, and it does not prescribe.
Want the full benchmark?
We share the scenario set and results with clinicians, institutions and researchers on request.
Request the benchmark
Carenika