lgoyal6 / apparatus-comps

apparatus‑comps.

Fire engines and ambulances trade from five to seven figures, and the tools that appraise them publish no accuracy figure. That is not unusual and it is not a criticism. Almost nobody in appraisal publishes one.

This is that number. A comparable-sales estimator, back-tested on a time-based split, reported with the error bar it actually needs rather than the one that would look better.

half the estimates
36% off or worse
land within 20%
30.5% of the time
to cover 9 in 10
6.8x wide a range
worst single miss
858% and it is in the report
Try it

Price a truck and see what it compared against

Describe a fire engine or an ambulance. The repository's comps model runs in your tab, against the same 903 cached listings the back-test used, and gives you a number, the range it actually needs, and the ten trucks it reached that number from.

starting the engine comps.py and the dataset, unmodified
the estimate
range that covers 9 in 10
how wide that is
comparables used

Read the comparables, not the number

The estimate is a weighted average of those ten trucks in log space, and the range comes from how much they disagree. When the ten are a mixed bag, the range is enormous and the number means very little, which is the honest thing an appraisal tool can tell you and the thing none of them do.

Ask it for something the set barely has, an aerial from the 1980s, and watch the comparables stop resembling what you typed. That is the moment to distrust it, and you can only see it because the comparables are on the page.

the same model
comps.py, unmodified.The one the back-test scored.
the same listings
903 cached rows.Asking prices, not sales.
what to look at
The width, and the ten.A tight number over a mixed bag is a lie.
Figure 1

How wrong is it, exactly

Absolute percentage error on the held-out set, by percentile. The median is the number an appraisal tool would quote if it quoted one. The right tail is the number that costs someone a deal.

Loading
Split
Percentile
at this percentile
-
the estimate is off by
-
median error
-
mean error
-

Why the mean and the median disagree

Mean error is 60.7% and median error is 36.0%. The gap is entirely the right tail: the worst single prediction on the test set put $95,786 on a pumper asking $10,000. Quoting the mean overstates the typical case, and quoting the median hides the case that matters.

The time-based split is the primary one: it trains on older listings and tests on newer, which is the order the world arrives in. Splitting the same data at random cuts mean error sharply for all three fitted models, by 9 to 20 points, and leaves median error slightly worse for two of them. It flatters the tail, not the typical case, which is why both are on the toolbar.

typical miss
36%.Half of estimates are worse than this.
one in ten
Off by 127% or more.And one in twenty by 197%.
beats the obvious baseline
By 11.5 points of median error.Against a global median at 47.5%.
Figure 2

The error bar it actually needs

An estimate with a range attached is only useful if the range is honest. Nominal coverage assumes the errors are log-normal. Actual coverage is measured on the held-out set, and the width is what a seller would be shown.

measured on the held-out set width is the ratio of the top of the range to the bottom

The intervals are honest and they are enormous. Every band covers more than its nominal rate, so the model is not overconfident. Covering nine cases in ten takes a range whose top is 6.8 times its bottom, which on a $50,000 truck is roughly $19,000 to $130,000.

That is the real product question. A range that wide is close to useless to a seller, and it is the correct answer given this data. Narrowing it needs closed sale prices, which is the thing the whole exercise does not have.

not overconfident
Every band over-covers.68% nominal measured at 75.3%.
and barely usable
6.8x wide for 90%.About $19k to $130k on a $50k truck.
what would fix it
Closed sale prices.These are asking prices on live listings.
Figure 3

Where it is worst

Median error by slice, so a buyer can tell which quotes to distrust. Cheap vehicles are the hard ones.

Sliced by
the five worst single predictionstime-based test set
Figure 4

Where it loses