Carenika AI Carenika
Open the app
Safety & evidence

We publish how well our safety engine works.

Most clinical AI tools claim accuracy. We measured ours against a fixed benchmark, with the right answers written down before testing began — and we publish the result, including what it got wrong.

Why the safety layer is not AI

A language model produces an answer that sounds right. That is acceptable when it drafts a sentence. It is not acceptable when it decides whether a medicine is safe for a particular patient, because the failure mode is not an awkward sentence — it is a patient.

So Carenika splits the work. The AI listens and writes. A separate rule engine — explicit, versioned, written by hand — checks the drugs. Same input, same output, every time. Every alert points back to a named rule a clinician can read, question, and overrule.

What the engine checks

Drug–drug interactions · duplicate therapy and class overlap · contraindications against recorded conditions · eGFR-banded renal dosing · hepatic ceilings · maximum daily dose · paediatric mg/kg review · pregnancy rules · an ADR registry that remembers a patient’s past reactions.

What the doctor sees

A Shield score for the prescription, and for each alert: why it fired, what to do instead, and the evidence behind it. Overrides are permitted and logged — the system advises, the physician decides.

The benchmark

50 scenarios. Reference standard fixed before testing.

Drug interactions, comorbidity contraindications, duplicate therapy and disease-specific rules. Thirty-one scenarios require an alert. Nineteen are deliberately safe — because a safety engine that fires on everything is as useless as one that fires on nothing, and false alarms have to be counted.

50

outpatient scenarios

31 / 19

alerts expected vs deliberate safe controls

1.00

specificity on the held-out run — no false alarms

11 / 12

misses traced to deployment, not rule design

The most useful thing the benchmark told us was uncomfortable: most of what the engine missed was already written as a rule — it simply was not live in the build doctors were using. Writing a rule and running a rule are different things, and only scenario testing shows the difference.

Results have been submitted for peer review to ICPS-2026, North South University, Dhaka. The author is the developer of the system under evaluation; the reference standard was fixed before testing to reduce assessment bias.

Limits

What this evidence does not prove.

A benchmark is fifty scenarios; clinical practice is not. The cases were written by us, which is why the held-out block — run once, before any disease-specific rules were added — matters more than the full re-run that followed the fixes. We report both. Carenika is decision support: it documents and it warns. It does not diagnose, and it does not prescribe.

Want the full benchmark?

We share the scenario set and results with clinicians, institutions and researchers on request.

Request the benchmark