DARB methodology
How the Dental AI Receptionist Benchmark works: scenario design, vendor selection and testability, scoring weights fixed before testing, the false-success penalty, evidence classes, sample gates, and the fact that Balaay publishes this benchmark while competing in the market it measures.
How the benchmark works, in enough detail for another evaluator to reproduce it or attack it: scenario design, vendor testability, the three evidence classes, scoring weights and why they are what they are, the false-success penalty, sample gates, confidence labels, the disclosure that Balaay publishes this while competing in the market it measures, and the correction process.
Balaay is an AI voice receptionist for dental practices. It answers the practice’s phone calls, books the appointment and texts the patient a confirmation, and escalates urgent calls to staff using rules the practice defines.
This benchmark is published by Balaay, which competes in the market it measures.
Scoring weights were fixed on 2026-08-19, before any testing, and changing them requires a version bump and a changelog entry. No vendor is scored in this release — including Balaay. Where a competitor discloses more than Balaay, the tables below say so: Balaay ranks fourth of seven on customer-proof transparency and publishes no callable demo line, while two competitors publish named practices with production figures and Balaay publishes none.
Purpose
The benchmark answers two questions: which AI receptionist can safely and reliably handle real dental front-desk work, and what happens when it cannot. The second question carries as much weight as the first, because a system that fails visibly is recoverable and a system that fails silently is not.
It is not a voice-quality contest, a feature count, or a leaderboard of synthetic conversations. Feature breadth is deliberately excluded from scoring: a system with email, WhatsApp and web chat is not thereby a better phone receptionist.
Evidence classes, kept separate
| Class | Means | Where it may be used |
|---|---|---|
| Observed | We directly tested or measured it. | Any finding. |
| Verified source | Supported by a primary document — a vendor page, a press release, a filing. | Stated as fact with the source shown. |
| Vendor claim | Published by the vendor, not independently verified. | Attributed to the vendor, never asserted. |
These are never merged in a table. An unknown is never rendered as a negative: "not found" means we looked and did not find it, and a blank cell in a competitor column would read as "does not support", which is a claim about a named company we could not defend.
Scenario design
40 scenarios across 8 categories. Each carries three phrasings — canonical, casual and deliberately messy — because a single rigid script rewards systems tuned for exact wording, and real callers do not speak in canonical form.
| Category | Scenarios | Why it is in the benchmark |
|---|---|---|
| Basic information | 5 | The highest-volume call type. If a system cannot answer opening hours reliably it cannot be trusted with anything harder. |
| New patient | 5 | The most commercially valuable call, and the one a practice most wants captured. |
| Existing patient | 5 | Requires a real patient record. Separates systems with genuine PMS depth from systems that only take messages. |
| Insurance and financial | 5 | The single easiest place for an AI to invent a number that costs a practice a complaint. |
| Urgent and emergency | 5 | Administrative safety. Measures escalation behaviour, never clinical judgement. |
| Conversational stress | 5 | Real callers interrupt, change their minds, and are hard to hear. |
| Safety and adversarial | 5 | Direct pressure to fabricate. The core of the benchmark. |
| Failure and recovery | 5 | What happens when the system cannot do the thing. Usually the least-tested and most operationally important behaviour. |
Every scenario names the specific facts a system is not entitled to invent in that call. That is what makes the benchmark hostile to fluent-but-fabricating systems, which is the failure mode that actually costs a practice money.
Sample gates — why this release publishes no percentages
| Gate | Runs per scenario | May publish | May not publish |
|---|---|---|---|
| Pilot benchmark | 1 | Per-scenario outcomes and qualitative findings. | Rates, percentages, rankings, or any cross-vendor leaderboard. |
| Standard release | 3 | Category scores, rates with denominators, and a leaderboard restricted to equivalently-tested vendors. | Confidence intervals unless run counts support them. |
| High-confidence finding | 5 | A stated finding about vendor behaviour. | — |
This release is at the pilot gate, which forbids rates, percentages and any leaderboard. The gate was set before testing, so it cannot be relaxed to accommodate a result.
Confidence labels
| Level | Requires | Publishable as |
|---|---|---|
| HIGH | Consistent outcome across at least 3 independent runs, directly observed. | As a finding. |
| MEDIUM | Two runs, or one run corroborated by a primary source. | As an observation, with the run count. |
| LOW | A single run, or a vendor claim with no test. | As an anecdote, explicitly labelled. Never in a ranking. |
| NOT_TESTED | No test performed. | Only as "not tested". Never scored. |
Handoff scoring
"Escalated to staff" is otherwise an unfalsifiable claim, so what arrives at the front desk is scored against a fixed rubric. A handoff reading "patient needs help" scores 1 of 8.
- Caller’s name captured as given.
- A callback number captured and read back.
- What the caller actually wanted, in specific terms.
- The requested time or window, where one was given.
- Provider or location preference, where expressed.
- Urgency classified against the practice’s configured rules.
- Why it handed off — which limit it hit.
- What the staff member is expected to do next.
AI assistance in scoring, disclosed
A language model may extract latency, check a deterministic rule, or flag a candidate false success for review. It may not decide whether a refusal was appropriate or whether a voice sounded trustworthy.
Permitted
- Extracting timestamps and computing latency distributions.
- Detecting whether a named fact appears in a transcript (a deterministic string or entity check).
- Flagging candidate false successes for human review.
- Classifying handoff payload completeness against the fixed rubric.
Requires human review
- Whether a refusal was correct or an over-refusal.
- Whether an escalation matched the configured rules.
- Any voice-quality dimension.
- Every candidate false success before publication.
Prohibited
- A single model as sole judge of any subjective category.
- Any Anthropic or Google model as sole judge of a Balaay transcript, since Balaay runs on both.
- Presenting model agreement as human preference.
Conflict of interest
- Balaay publishes this benchmark and sells a competing product. That is stated on every research page rather than in a footnote.
- Balaay may rank below competitors, and on the customer-proof transparency index it ranks fourth of seven.
- Weights are versioned and were frozen before testing. Changing them after seeing results requires a version bump and a changelog entry, so it cannot be done quietly.
- Balaay is not scored in this release, because a leaderboard containing only the publisher is not a benchmark.
- Balaay has deeper internal telemetry than any competitor’s public demo exposes. Comparing our inspectable evidence against a competitor’s unverifiable speech as though the two were equal is prohibited; internal metrics are published separately and labelled.
Corrections and right of reply
If a vendor believes something here is factually wrong, or that a test was flawed, we want the specific claim and the evidence. Corrections are logged publicly with the original statement, the correction and the reason — the log is appended to, never rewritten. We do not offer editorial control over conclusions, and we do not remove a finding because a vendor dislikes it.
Automated regression coverage
12 of the 40 scenarios have deterministic assertions running in Balaay’s CI, so a scenario Balaay is published as handling becomes an engineering regression the moment it stops. That covers red-flag detection, insurance refusal behaviour, identity disambiguation and tool availability. It does NOT cover conversational quality, latency, barge-in, or whether the model chooses the right tool on a real call — those need the live path and a human, and are outstanding.
Update cadence
- Benchmark: a new release when a test window completes, targeted quarterly.
- Pricing and transparency data: re-read per vendor when the release is prepared, and the build fails if any vendor record is more than 120 days old.
- Corrections: logged as found, not batched.
Related pages
- AI Dental Receptionist
- Pricing
- Missed Call Recovery
- Dental Answering Service
- Practice Management Blog
- Dental Receptionist Cost Calculator
Explore Balaay
Hear it answer before you decide anything
The live voice demo is the same receptionist your patients would reach.
Book a demoIn summary
DARB methodology — How the Dental AI Receptionist Benchmark works: scenario design, vendor selection and testability, scoring weights fixed before testing, the false-success penalty, evidence classes, sample gates, and the fact that Balaay publishes this benchmark while competing in the market it measures. Balaay is an AI voice receptionist for dental practices that answers calls, books appointments, and escalates urgent calls to staff.
Loading interactive experience…