Hazlo Domain Valuation Accuracy Benchmark
This public benchmark explains how Hazlo measures appraisal calibration, what the June 19, 2026 evaluation set contains, what the aggregate results say, and—equally important—what they do not prove. The current edition is a deterministic calibration smoke test, not a claim of market-wide or cross-vendor superiority.
Dates and source version
Page published and reviewed August 22, 2026. Source benchmark generated June 19, 2026 with engine v3-enhanced. The source date is kept separate from the publication date so readers can tell exactly which engine run the figures describe.
Evaluation set
The focus-class pack contains 25 domains across 9 practical domain classes. 18 have human-approved retail target ranges, 7 also have approved confidence ranges, 24 third-party estimate points are retained for context, and 2 rows have public sale citations.
Confirmed sales have the highest authority. Human-approved ranges are used for calibration when no sale exists. Third-party estimates are informational only and are not treated as truth.
Aggregate results from the dated run
9 of 18 approved-target retail midpoints were inside their approved bands: 50%. 5 were above and 4 were below. Median absolute distance to the nearest band edge was 1%, while mean absolute distance was 72.5%; the large gap shows that a few severe misses remain.
For the 2 confirmed-sale points, 50% were within one-half to two times the reference and median absolute price delta was 68%. Two rows are far too small for a general accuracy conclusion.
For 24 third-party estimate points, 41.7% were within one-half to two times the estimate and median absolute delta was 66%. This is agreement, not accuracy, because competitor or analyst estimates are not ground truth.
How errors are measured
For an approved range, a result counts as within when the retail midpoint falls between the low and high values. If it falls outside, distance is measured to the nearest band edge. The median describes a typical row; the mean remains visible because it is sensitive to severe tail misses.
For a confirmed sale or third-party estimate, the report measures absolute percentage difference and whether the retail midpoint is between 0.5× and 2× of the reference price. These broad ratio bands reflect the uncertainty and skew common in domain sales.
Holdout status and leakage controls
This published edition is a deterministic calibration smoke test. It does not claim an outcome-based sales holdout because only two confirmed-sale reference points are included.
Hazlo's separate historical-sales backtest assigns domains deterministically to a 70% calibration set and a 30% holdout set using a stable hash of the domain. It excludes pending duplicate-review sales, excludes the target domain from its comparable pool, and never sends the sale price into the appraisal engine. No aggregate from that mutable database-backed harness is presented as a result on this page.
Reproduce the aggregate report
Run npx tsx scripts/runFocusClassBenchmark.ts from the project root. The runner disables the optional AI reviewer, startup probe, and thinking mode, then writes the machine-readable and human-readable reports. The public JSON distribution below mirrors the dated aggregate evidence used on this page.
- Machine-readable evidence: /appraisal-accuracy-benchmark/data.json
- Runner: scripts/runFocusClassBenchmark.ts
- Evaluation source: server/services/calibration/benchmarkDataset.ts
- Generated JSON report: docs/appraisal-baseline-report.json
- Generated Markdown report: docs/appraisal-baseline-report.md
Known-sale source citations
The pizza.com reference is linked to contemporary reports from BBC News and Wired. The thunder.io reference is linked to TLD Investors and DN.com reports of its May 2026 Afternic sale. On August 22, 2026, an uncited calm.com value was moved to labeled, unverified informational context and excluded from sale metrics.
Limitations
The approved-target set is small, curated, and partly informed by human judgment; it is not a random sample of the domain market.
Only two confirmed public sales are included, so the sales score is descriptive and not statistically strong.
Third-party estimates are not ground truth and are never used to declare accuracy or superiority.
Several domain classes have no approved target or confirmed sale and therefore contribute only to output-distribution checks.
Historical sale prices can reflect buyer-specific strategy, venue, timing, financing, and private deal terms that an automated appraisal cannot observe.
The report tests one dated engine version. It should not be assumed to describe later versions until the benchmark is rerun and republished.
An uncited calm.com value was moved from the confirmed-sale score to labeled, unverified informational context on August 22, 2026. It is not treated as ground truth.
How comparisons should use this benchmark
Comparison pages may cite this benchmark to describe Hazlo's own published evaluation. They must not convert agreement with competitor estimates into an accuracy score, claim that Hazlo wins a head-to-head test that was not run, or assign unsupported scores to another product. Readers should compare disclosed methodology, output fields, source dates, and limitations using the same domains and reference sales.
Frequently Asked Questions
Does this benchmark prove Hazlo is the most accurate appraiser?
No. The current edition is a small calibration smoke test with only two confirmed-sale points. It is useful for exposing misses and tracking a dated engine, not for proving market-wide or cross-vendor superiority.
Why are competitor estimates separated from actual sales?
An estimate is another model's opinion, not a transaction. Agreement with it can provide context, but it cannot establish accuracy.
Is there a holdout set?
The published focus-class edition does not claim a sales holdout result. Hazlo's separate database-backed sales harness uses a deterministic 70/30 split, but its mutable aggregate is not presented on this page.