Balaay

DARB methodology

How the Dental AI Receptionist Benchmark works: scenario design, vendor selection and testability, scoring weights fixed before testing, the false-success penalty, evidence classes, sample gates, and the fact that Balaay publishes this benchmark while competing in the market it measures.

How the benchmark works, in enough detail for another evaluator to reproduce it or attack it: scenario design, vendor testability, the three evidence classes, scoring weights and why they are what they are, the false-success penalty, sample gates, confidence labels, the disclosure that Balaay publishes this while competing in the market it measures, and the correction process.

Balaay is an AI voice receptionist for dental practices. It answers the practice’s phone calls, books the appointment and texts the patient a confirmation, and escalates urgent calls to staff using rules the practice defines.

Built for dental practices. Plans from $129/mo. The AI Voice Receptionist is on Growth at $199/mo, plus a one-time setup fee. Month-to-month, no per-call fees. One-time setup fee applies. HIPAA-ready, BAA available. See pricing →

This benchmark is published by Balaay, which competes in the market it measures.

Scoring weights were fixed on 2026-08-19, before any testing, and changing them requires a version bump and a changelog entry. No vendor is scored in this release — including Balaay. Where a competitor discloses more than Balaay, the tables below say so: Balaay ranks fourth of seven on customer-proof transparency and publishes no callable demo line, while two competitors publish named practices with production figures and Balaay publishes none.

Purpose

The benchmark answers two questions: which AI receptionist can safely and reliably handle real dental front-desk work, and what happens when it cannot. The second question carries as much weight as the first, because a system that fails visibly is recoverable and a system that fails silently is not.

It is not a voice-quality contest, a feature count, or a leaderboard of synthetic conversations. Feature breadth is deliberately excluded from scoring: a system with email, WhatsApp and web chat is not thereby a better phone receptionist.

Evidence classes, kept separate

ClassMeansWhere it may be used
ObservedWe directly tested or measured it.Any finding.
Verified sourceSupported by a primary document — a vendor page, a press release, a filing.Stated as fact with the source shown.
Vendor claimPublished by the vendor, not independently verified.Attributed to the vendor, never asserted.

These are never merged in a table. An unknown is never rendered as a negative: "not found" means we looked and did not find it, and a blank cell in a competitor column would read as "does not support", which is a claim about a named company we could not defend.

Scenario design

40 scenarios across 8 categories. Each carries three phrasings — canonical, casual and deliberately messy — because a single rigid script rewards systems tuned for exact wording, and real callers do not speak in canonical form.

CategoryScenariosWhy it is in the benchmark
Basic information5The highest-volume call type. If a system cannot answer opening hours reliably it cannot be trusted with anything harder.
New patient5The most commercially valuable call, and the one a practice most wants captured.
Existing patient5Requires a real patient record. Separates systems with genuine PMS depth from systems that only take messages.
Insurance and financial5The single easiest place for an AI to invent a number that costs a practice a complaint.
Urgent and emergency5Administrative safety. Measures escalation behaviour, never clinical judgement.
Conversational stress5Real callers interrupt, change their minds, and are hard to hear.
Safety and adversarial5Direct pressure to fabricate. The core of the benchmark.
Failure and recovery5What happens when the system cannot do the thing. Usually the least-tested and most operationally important behaviour.

Every scenario names the specific facts a system is not entitled to invent in that call. That is what makes the benchmark hostile to fluent-but-fabricating systems, which is the failure mode that actually costs a practice money.

Sample gates — why this release publishes no percentages

GateRuns per scenarioMay publishMay not publish
Pilot benchmark1Per-scenario outcomes and qualitative findings.Rates, percentages, rankings, or any cross-vendor leaderboard.
Standard release3Category scores, rates with denominators, and a leaderboard restricted to equivalently-tested vendors.Confidence intervals unless run counts support them.
High-confidence finding5A stated finding about vendor behaviour.

This release is at the pilot gate, which forbids rates, percentages and any leaderboard. The gate was set before testing, so it cannot be relaxed to accommodate a result.

Confidence labels

LevelRequiresPublishable as
HIGHConsistent outcome across at least 3 independent runs, directly observed.As a finding.
MEDIUMTwo runs, or one run corroborated by a primary source.As an observation, with the run count.
LOWA single run, or a vendor claim with no test.As an anecdote, explicitly labelled. Never in a ranking.
NOT_TESTEDNo test performed.Only as "not tested". Never scored.

Handoff scoring

"Escalated to staff" is otherwise an unfalsifiable claim, so what arrives at the front desk is scored against a fixed rubric. A handoff reading "patient needs help" scores 1 of 8.

AI assistance in scoring, disclosed

A language model may extract latency, check a deterministic rule, or flag a candidate false success for review. It may not decide whether a refusal was appropriate or whether a voice sounded trustworthy.

Permitted

Requires human review

Prohibited

Conflict of interest

Corrections and right of reply

If a vendor believes something here is factually wrong, or that a test was flawed, we want the specific claim and the evidence. Corrections are logged publicly with the original statement, the correction and the reason — the log is appended to, never rewritten. We do not offer editorial control over conclusions, and we do not remove a finding because a vendor dislikes it.

Read the corrections log →

Automated regression coverage

12 of the 40 scenarios have deterministic assertions running in Balaay’s CI, so a scenario Balaay is published as handling becomes an engineering regression the moment it stops. That covers red-flag detection, insurance refusal behaviour, identity disambiguation and tool availability. It does NOT cover conversational quality, latency, barge-in, or whether the model chooses the right tool on a real call — those need the live path and a human, and are outstanding.

Update cadence

Related pages

Explore Balaay

Hear it answer before you decide anything

The live voice demo is the same receptionist your patients would reach.

Book a demo

In summary

DARB methodology — How the Dental AI Receptionist Benchmark works: scenario design, vendor selection and testability, scoring weights fixed before testing, the false-success penalty, evidence classes, sample gates, and the fact that Balaay publishes this benchmark while competing in the market it measures. Balaay is an AI voice receptionist for dental practices that answers calls, books appointments, and escalates urgent calls to staff.

Loading interactive experience…