Independent clinical evaluation for healthcare AI.

We evaluate AI-generated clinical documentation against a published rubric, graded blind by verified, license-checked clinicians. Our methodology is public. Your results are not.

What the encounter contained
“I was moving a box up onto the shelf and it just — it was like something hit me. Right then.
What the note recorded
“Patient reports headache of gradual onset since yesterday afternoon.”
D2 TEMPORAL ERROR SEVERITY 4 · CRITICAL

Thunderclap onset is the finding that triggers investigation for subarachnoid haemorrhage. A note that records it as gradual removes the reason to look.

30error codes across six families
0–4severity scale anchored to clinical consequence
Blindgraders never know which system wrote the note
Publishedrubric and case specification, in full
§01 — The gap

Everyone reports accuracy. Almost nobody says who checked.

AI documentation tools are deployed across most large health systems. Very few have been independently evaluated by practising clinicians against a stated standard.

The reason isn't indifference — rigorous clinical evaluation is genuinely hard. It needs clinicians who can tell a stylistic difference from a safety error, a case set difficult enough to discriminate between systems, and a protocol disciplined enough that the numbers mean something.

Internal quality review doesn't close the gap, because it isn't independent. Procurement discounts vendor-reported accuracy for exactly that reason — which leaves clinical leaders facing a question they have no good answer to: who checked, and how?

§02 — Method

We publish the instrument.

You should expect that of anyone claiming to evaluate clinical AI, and be sceptical of anyone who won't.

The rubric — 30 codes, six families

Fabrication

Content in the note with no basis in the encounter

Omission

Clinically material content the note dropped

Distortion

Negation, laterality, temporal, numeric and certainty errors

Clinical reasoning

Unsupported assessment, assessment–plan mismatch, guideline discordance

Compliance

Billing-relevant fabrication, third-party PHI, unsupported attestation

Usability

Defects that create edit burden without clinical risk

Severity, anchored to consequence rather than preference

0Cosmetic
1Minor — a reviewer would fix it
2Misleads a reader or creates rework
3Could change a clinical decision
4Could contribute to patient harm

The governing rule is narrow and deliberate: the encounter is the only ground truth. Graders mark what the note gets wrong — not what they would have written differently.

The cases

Clinician-authored synthetic encounters, each shipped with a ground-truth key defining the material facts a correct note must capture. Because we author the ground truth, omission and fabrication are measured exactly rather than reconstructed from a grader's recall. No PHI enters the evaluation set.

Difficulty is engineered, not accidental. Every case declares the stressors it plants — buried red flags, medication churn, laterality conflicts, self-corrections, third-party history, conversational numbers — and each maps to the error codes it is designed to provoke. A test set of clean encounters would show every system scoring well and tell you nothing.

The protocol

  • Blinded. Graders never know which system produced a note.
  • Double-graded. Independent second grading on a stratified sample, plus every note where any grader records a severe error.
  • Adjudicated. A third clinician resolves material disagreement.
  • Calibrated. Every grader completes the same reference set before live work.
  • Reported. Inter-rater reliability is published alongside results. If our graders don't agree with each other, you should know that before you read our numbers.
  • Evidenced. Every recorded error cites the source span it contradicts.

The panel

Every clinician is verified before their first project: state license through the issuing board, NPI confirmation, and ABMS board certification where applicable. Panel credentials are disclosed in aggregate with every report. We would rather field thirty well-calibrated clinicians than three hundred names on a list.

Clinical Documentation Accuracy Rubric PDF · v1.0 · no email required Benchmark Case Specification PDF · v1.0 · no email required
§03 — Deliverable

Traced errors, not a score.

Error rates by taxonomy family. Severity distribution. Red-flag omission rate. Median edit time.

And the metric clinical leaders respond to first: the proportion of notes a practising clinician would sign into a chart as-is — with every exception itemised and traceable to the source encounter.

A percentage tells you where you stand. A traced error tells your engineering team what to fix. Every finding we report is delivered in the form shown at the top of this page: the source, the note, the code, the severity.

§04 — Your data

HIPAA-aligned infrastructure is our operating discipline.

Cureris Integrated Resources is a healthcare technology company. This is what we already do for a living, not a policy page written for this service.

  • Evaluation sets are synthetic and contain no PHI
  • Where an engagement requires protected data, it stays within controlled Cureris infrastructure under a Business Associate Agreement
  • Every panel clinician signs an NDA and IP assignment before their first project
  • Your system outputs, results and report are confidential, and are never published without your written consent
§05 — Independence

An evaluation your competitors can dismiss is worth nothing to you.

Which is why the commercial structure is separated from the published work, deliberately:

  • Our public benchmark tests generally available models only. No company pays to be included in it.
  • Commercial engagements are private and structurally separate. Your results are yours.
  • Engagement clients may review case realism and specialty mix — never scoring rules, severity anchors, or the grading of their own output.
  • Where engagement revenue funds published work, we say so in the publication.
§06 — Cureris

An operating healthcare technology firm, not a research group.

Cureris Integrated Resources LLC is a Dallas-based technology company working in HIPAA-aligned managed IT, cybersecurity and document intelligence for healthcare organisations. Founded [YEAR FOUNDED].

Cureris Panel applies that compliance infrastructure and healthcare operating experience to clinical AI evaluation. Evaluation is the discipline we are extending into, from a business that already runs inside healthcare's regulatory perimeter.

§07 — Status

Where we are now.

We are running the first full benchmark and taking two design partners into it.

We would rather state that plainly than imply a history we don't have. What we bring is a methodology published in full, a verified clinical panel, and a compliance posture built over years of healthcare technology work. What a design partner gets in return is a materially reduced rate, input on case realism, and results before anyone else has them.

§08 — Enquire

Two design partner slots remain.

Or write directly: info@cureris.me · (917) 858-4023