Trust & transparency — how Nox's medical-safety checks are measured

How Leo, Nox's medical-safety layer, is built and measured — red-flag detector recall, false-positive rate, screening categories, governance, and data sources.

How Leo, the red-flag safety layer, works

Before the AI model answers, Leo — Nox's medical-safety system — screens your message for signs of acute red-flag conditions across dozens of categories, such as stroke signs, chest pain, severe breathing difficulty, or a mental-health crisis. Its first check is an independent, deterministic detection layer that runs before any AI is called, surfacing clear guidance to seek appropriate care, with the correct local emergency number when you set your region. Leo is a safety net, not a guarantee: no automated system catches every emergency, and Nox publishes the deterministic detector's measured recall and false-positive rates on its Trust & Transparency page.

When the deterministic layer finds nothing, a second check runs in the background: a lightweight AI classifier re-reads your recent messages to catch dangerous descriptions the fixed patterns can miss — slang, another language, older disease names, or indirect phrasing. If it recognizes a likely emergency, Nox surfaces the same seek-care guidance as a pattern match. This backstop can only add a safety note, never remove or soften one, and because it is not deterministic its results are kept separate from the published detector metrics.

Measured detector performance

  • Overall recall: 100.0% (150 of 150 should-fire cases produced a safety note)
  • False-positive rate: 0.0% (0 of 53 benign cases triggered a note)
  • Test set: 203 labeled cases · 107 rules across 75 categories
Detector recall by red-flag category (75 categories with labeled positive cases)
CategoryCases caughtRecall
Headache6/6100.0%
Eye pain5/5100.0%
Chest pain5/5100.0%
Shortness of breath5/5100.0%
Abdominal pain9/9100.0%
Fever3/3100.0%
Allergic reaction5/5100.0%
Mental-health crisis5/5100.0%
Pregnancy5/5100.0%
Child symptoms5/5100.0%
Stroke (FAST)6/6100.0%
Seizure4/4100.0%
Low blood oxygen (SpO₂)4/4100.0%
Severe dehydration4/4100.0%
Fever with confusion3/3100.0%
Overdose / poisoning4/4100.0%
Severe bleeding1/1100.0%
Sepsis warning signs1/1100.0%
Diabetic emergency1/1100.0%
Necrotizing fasciitis (flesh-eating infection)2/2100.0%
Toxic shock syndrome1/1100.0%
Meningitis rash (non-blanching)2/2100.0%
Testicular torsion1/1100.0%
Choking1/1100.0%
Severe burn1/1100.0%
Heat stroke1/1100.0%
Aortic dissection2/2100.0%
Abdominal aortic aneurysm2/2100.0%
Blood clot in the leg (DVT)2/2100.0%
Fainting / syncope2/2100.0%
Dangerous heart rhythm1/1100.0%
Hypertensive crisis2/2100.0%
Blocked artery in a limb1/1100.0%
Head injury2/2100.0%
Spinal injury1/1100.0%
Cauda equina syndrome1/1100.0%
Broken bone / open fracture2/2100.0%
Compartment syndrome1/1100.0%
Crush injury1/1100.0%
Stab / gunshot / impalement1/1100.0%
Drowning / near-drowning1/1100.0%
Hypothermia1/1100.0%
Frostbite1/1100.0%
Severe altitude sickness1/1100.0%
Smoke inhalation1/1100.0%
Snake bite1/1100.0%
Animal / human bite2/2100.0%
Airway swelling (stridor)1/1100.0%
Severe asthma attack1/1100.0%
Coughing up blood1/1100.0%
Unable to urinate1/1100.0%
Kidney stone complications2/2100.0%
Bowel obstruction1/1100.0%
Strangulated hernia1/1100.0%
Ectopic pregnancy1/1100.0%
Ovarian torsion1/1100.0%
Postpartum emergency2/2100.0%
Priapism1/1100.0%
Sudden hearing loss1/1100.0%
Thyroid storm1/1100.0%
Adrenal crisis1/1100.0%
Sickle cell crisis1/1100.0%
Fever while immunocompromised1/1100.0%
Severe alcohol withdrawal1/1100.0%
Serotonin syndrome1/1100.0%
Severe drug reaction (SJS)1/1100.0%
Acute psychosis1/1100.0%
Eating-disorder complications1/1100.0%
Spreading dental infection2/2100.0%
Uncontrolled nosebleed1/1100.0%
Severe menstrual bleeding1/1100.0%
Dialysis emergency1/1100.0%
Newborn jaundice1/1100.0%
Sexual assault1/1100.0%
Domestic violence1/1100.0%

Safety questions, answered

Does Nox catch every emergency?

No, and Nox never claims to. Leo's red-flag layer is a deterministic safety net that screens for a fixed list of acute warning-sign categories; emergencies outside those categories may not trigger a note. Its measured recall and false-positive rate are published openly on the Accuracy and Trust & Transparency pages. If you think you may be experiencing a medical emergency, call your local emergency number right away.

What happens if I describe something serious to Nox?

If your message matches a red-flag pattern — like chest pressure spreading to the arm, stroke signs, or severe breathing difficulty — Leo shows a clear emergency or urgent-care banner before any AI-generated content. The screening runs before the AI model is called, so this guidance does not depend on the AI behaving well.

Does Nox understand slang, other languages, or indirect ways of describing an emergency?

It tries to. When the deterministic layer finds nothing, a second check runs in the background: a lightweight AI classifier re-reads your recent messages to catch dangerous descriptions the fixed patterns can miss — slang, another language, older disease names, or indirect phrasing. If it recognizes a likely emergency, Nox surfaces the same seek-care guidance as a pattern match. This backstop can only add a safety note, never remove or soften one, and because it is not deterministic its results are kept separate from the published detector metrics. Because this backstop uses AI rather than fixed patterns, its results are kept separate from the published detector metrics on the Accuracy page — and, like the pattern layer, it is a safety net, not a guarantee.

Can I control how strict Nox's safety screening is?

You choose how far Leo's screening goes, free on every plan, from a control right in the chat box. Three levels sit side by side, left to right, from least to most protective: Relaxed warns you about clear emergencies only; Standard also warns on anything urgent; and Strict adds the AI backstop that re-reads recent messages for dangerous descriptions worded indirectly or in another language. Whatever you pick, Leo always screens for true emergencies — the level only changes how many optional layers run on top — and Agent and voice conversations always use the strictest setting.

How accurate is Nox's safety screening?

The detector's overall recall and false-positive rate are measured against a maintained, labeled test set and published as a high-level summary on the Accuracy page. The numbers are generated by the test suite and committed with the code, never hand-typed, and an automated test blocks any rule change that isn't re-measured.

Has a doctor reviewed Nox?

Not yet, and Nox says so plainly: clinician review status is 'pending'. Nox claims no clinician endorsement until a real, licensed clinician completes a review — at which point their name, credentials, scope, and review date will be published on the Trust & Transparency page.

Does Nox show the right emergency number for my country?

Yes, when you set your region. In Settings you can choose your region so Leo's safety banners show your local emergency number. If no region is set, Nox falls back to universal numbers (911 / 999 / 112) so emergency guidance is never blank.

Does Nox Voice 1.2 have the same safety rules?

Yes. Voice conversations go through the same Leo safety screening and the same conservative health guidance as text — talking to Nox out loud never relaxes the safety layer.