Procurement-ready
- Fabrication rate < 1%
- Zero severity-4 omissions
MedAttest is the independent certification body for clinical AI accuracy. FACT certification — Fabrication, Accuracy & Completeness Testing — is the procurement standard for ambient AI medical scribes. Think SOC 2, but for what the AI writes in the chart.

There is no objective way to know whether the AI fabricates findings, omits critical clinical facts, or makes unsupported inferences. Vendors self-report accuracy on closed benchmarks. Buyers have nothing independent to point to.
FACT gives both sides a common, verifiable standard.
Vendor-chosen test set. Vendor-defined scoring. No adversarial pressure. No independent review.
Third-party encounters with known ground truth, adversarial traps, severity-weighted scoring, and clinician adjudication.
Every certification run is reproducible, severity-aware, and snapshotted against the exact FACT criteria version in effect.
MedAttest sends a battery of physician-authored synthetic patient encounters — primary care, cardiology, psychiatry — to the vendor's AI.
134 verified clinical facts · 30 adversarial fabrication traps · 0 PHI
Every assertion in the AI-generated note is extracted and scored against ground truth by a physician-calibrated AI judge: grounded, fabricated, or unsupported inference. Omission detection catches what the AI failed to document, weighted by clinical severity.
Grounded / Fabricated / Unsupported · Severity 1–4 omission weighting
Low-confidence judgments are routed to licensed clinicians, whose rulings always override the AI judge. Severity-3 and 4 omissions and fabrications are always human-reviewed.
Clinician override is final · Judge is continuously recalibrated
Every test packet is seeded with canaries — a distinctive allergy, a rare condition, and unique identifier strings tied to specific synthetic patients. We scan every output for cross-patient leakage. A single confirmed containment finding caps the run at FAIL, regardless of fabrication rate, omission scores, or composite. There is no threshold; it's a hard gate.
Per-packet canaries · Cross-patient leak detection · Hard fail, no threshold
Vendors receive a composite score and a FACT tier — A, B, C, or fail — against published criteria. Tiering gates on the conservative upper bound of a Wilson 95% confidence interval, not the point estimate, so a small-sample 'zero fabrications' cannot earn Tier A.
Wilson 95% upper bound · Tier A / B / C / Fail · Published criteria
Certified vendors get a versioned public trust page and FACT badge, a shareable PDF attestation with a confidential claim-level appendix, and a CHAI-compatible JSON export. Every attestation is stamped with its assurance level — Verified or Assessed — based on how evidence was collected.
Public trust page · FACT badge · Assurance level stamped · PDF + CHAI JSON
Severity-aware by design — a missed drug allergy isn't scored like a missed social-history detail.
Tiering compares the conservative upper bound of a Wilson confidence interval (default 95%) against each threshold — not the raw rate. A small-sample "0% fabrication" cannot earn Tier A; only a genuinely large, clean sample narrows the bound below 1%. Every report shows n_assertions and n_encounters alongside rates.
Tiering gates on fabrication, severity-4 omissions, and severity-weighted omission rate — the composite score is reported alongside, but never used as a tier gate. FACT criteria are snapshotted into every certification run, so an attestation never silently changes meaning.
Fabrication rate
Fabrication rate
Same synthetic encounters. Same ground truth. The FACT report makes the delta visible to procurement, governance, and patient-safety committees before deployment.
Every packet is seeded with canaries — a distinctive allergy, a rare condition, unique identifier strings. A single confirmed cross-patient leak is an automatic, non-negotiable FAIL. Patient-safety first; no score can override it.
Tier thresholds compare against the upper bound of a Wilson 95% confidence interval, not the point estimate. A small-sample 'zero fabrications' cannot earn Tier A — only a genuinely large, clean sample narrows the bound below 1%.
Third-party testing against ground truth known only to MedAttest — not a vendor benchmark dressed up as one.
Adversarial traps in every encounter. Designed to surface fabrications that random sampling misses entirely.
AI judge for scale, licensed physicians for the calls that matter. The judge is continuously calibrated against clinician rulings.
A missed drug allergy is not scored like a missed social-history detail. Clinical impact shapes the math.
FACT criteria are snapshotted into every certification run, so an attestation never silently changes meaning.
All test encounters are synthetic. Nothing sensitive ever touches the platform — not in transit, not at rest.
Our SOC 2 Type II vs Type I analogue — determined by how evidence was collected, not by score. The two levels never render identically on a vendor's trust page.
A FACT-A badge and a public trust page shorten procurement cycles. Stop answering one-off security and accuracy questionnaires for every health system — point them at your versioned attestation.
Start a certification runDefensible purchasing, AI governance, and patient safety in one document. Use the vendor registry to compare FACT tiers, fabrication rates, and severity-weighted omission scores side by side.
Look up a vendor's FACT report