We evaluate AI-generated clinical documentation against a published rubric, graded blind by verified, license-checked clinicians. Our methodology is public. Your results are not.
Thunderclap onset is the finding that triggers investigation for subarachnoid haemorrhage. A note that records it as gradual removes the reason to look.
AI documentation tools are deployed across most large health systems. Very few have been independently evaluated by practising clinicians against a stated standard.
The reason isn't indifference — rigorous clinical evaluation is genuinely hard. It needs clinicians who can tell a stylistic difference from a safety error, a case set difficult enough to discriminate between systems, and a protocol disciplined enough that the numbers mean something.
Internal quality review doesn't close the gap, because it isn't independent. Procurement discounts vendor-reported accuracy for exactly that reason — which leaves clinical leaders facing a question they have no good answer to: who checked, and how?
You should expect that of anyone claiming to evaluate clinical AI, and be sceptical of anyone who won't.
Content in the note with no basis in the encounter
Clinically material content the note dropped
Negation, laterality, temporal, numeric and certainty errors
Unsupported assessment, assessment–plan mismatch, guideline discordance
Billing-relevant fabrication, third-party PHI, unsupported attestation
Defects that create edit burden without clinical risk
The governing rule is narrow and deliberate: the encounter is the only ground truth. Graders mark what the note gets wrong — not what they would have written differently.
Clinician-authored synthetic encounters, each shipped with a ground-truth key defining the material facts a correct note must capture. Because we author the ground truth, omission and fabrication are measured exactly rather than reconstructed from a grader's recall. No PHI enters the evaluation set.
Difficulty is engineered, not accidental. Every case declares the stressors it plants — buried red flags, medication churn, laterality conflicts, self-corrections, third-party history, conversational numbers — and each maps to the error codes it is designed to provoke. A test set of clean encounters would show every system scoring well and tell you nothing.
Every clinician is verified before their first project: state license through the issuing board, NPI confirmation, and ABMS board certification where applicable. Panel credentials are disclosed in aggregate with every report. We would rather field thirty well-calibrated clinicians than three hundred names on a list.
Error rates by taxonomy family. Severity distribution. Red-flag omission rate. Median edit time.
And the metric clinical leaders respond to first: the proportion of notes a practising clinician would sign into a chart as-is — with every exception itemised and traceable to the source encounter.
A percentage tells you where you stand. A traced error tells your engineering team what to fix. Every finding we report is delivered in the form shown at the top of this page: the source, the note, the code, the severity.
Cureris Integrated Resources is a healthcare technology company. This is what we already do for a living, not a policy page written for this service.
Which is why the commercial structure is separated from the published work, deliberately:
Cureris Integrated Resources LLC is a Dallas-based technology company working in HIPAA-aligned managed IT, cybersecurity and document intelligence for healthcare organisations. Founded [YEAR FOUNDED].
Cureris Panel applies that compliance infrastructure and healthcare operating experience to clinical AI evaluation. Evaluation is the discipline we are extending into, from a business that already runs inside healthcare's regulatory perimeter.
We are running the first full benchmark and taking two design partners into it.
We would rather state that plainly than imply a history we don't have. What we bring is a methodology published in full, a verified clinical panel, and a compliance posture built over years of healthcare technology work. What a design partner gets in return is a materially reduced rate, input on case realism, and results before anyone else has them.
Or write directly: info@cureris.me · (917) 858-4023