How we test AI receptionists

How we test

Proofground grades are built to be fair, repeatable, and defensible. Here is exactly how a grade is produced.

The same test for everyone

Every vendor in an industry is measured against one master checklist of capabilities — booking, rescheduling, verifying a patient, handling insurance, recognizing an emergency, taking a message, resisting manipulation, and more. The checklist is identical for every provider, including capabilities a given vendor does not advertise, because that is the only way to compare them honestly.

Real, recorded calls

We play the caller and run each scenario as a live phone call. Every call is recorded as evidence. A capability only counts against a vendor if they advertise it and their demo actually lets us test it — demo-only limitations are shown separately and never held against the product.

Blind, calibrated judges

Each call is scored from the transcript by a judge that never sees the vendor’s name, against a written rubric. Judges cite the exact moment in the call that supports each decision. A grade is only a threshold-pass rate with an honest margin of error — never a raw score-versus-score “winner.”

A second, independent review

Before anything is published, a separate review pass re-examines every non-passing result against the recording and the data — asking whether the test itself, the demo’s limits, or missing account configuration was really at fault. Only genuine, verified issues are published. When the data cannot separate two vendors, we say so, rather than inventing a winner.

Dated, and re-checked over time

Every scorecard is dated and kept, so you can see whether a vendor is getting better or worse. Grades come from a standardized, generic scenario pack — a shortlist tool, not a final verdict. Before you decide, test your shortlist on your own account with scenarios matched to how your office actually runs.