Last month I ran an AI ductwork takeoff on a real hospital set. The manual number, done by a senior mechanical estimator with fifteen years on similar jobs, came back at 3,396 linear feet. The AI came back at 1,096 linear feet. A third of the manual number. On the surface, that’s a disaster : 78% accurate is a passing grade on a chemistry exam, not on a bid you’re going to stake your company’s margin on.
I stared at that gap for an hour before I understood what I was actually looking at. And once I did, I stopped grading AI takeoffs on quantity match. Here’s why.
The diff feels rigorous. It isn’t.
Every estimator I’ve worked with has the same reflex when they see an AI takeoff: pull up their own numbers, subtract, look at the delta. If the delta is small, the AI passes. If it’s big, the AI fails. Simple, mechanical, feels like the disciplined thing to do.
But the manual number carries things the drawings don’t show.
An experienced ductwork estimator doesn’t just quantify what’s on the sheet. They add the twenty linear feet of round duct they know has to run between two mechanical rooms because the equipment schedules make the connection obvious even though the drafter never drew it. They add the fittings that always exist on a transition of that size, even when the plan only shows the straight run. They add the drop-and-rise around the beam the structural set implies but the mechanical set doesn’t call out. That’s not slop. That’s the experience the firm is billing for.
AI can only quantify what’s actually on the drawing. It can’t infer un-drawn geometry. It literally cannot see what isn’t there.
So when I diff my 3,396 LF against the AI’s 1,096 LF, I’m not measuring an accuracy failure. I’m measuring a mixture: some real AI misses on ductwork that was drawn but got missed, plus a large chunk of my own experience-adds for scope that wasn’t drawn and never will be. The AI’s job and my job overlap, but they aren’t the same job.
Diffing quantities to grade AI is like grading a spellchecker by whether it wrote the essay. Wrong instrument for the measurement.

Estimators diff quantities because it’s the only tool they have
I don’t blame anyone for the reflex. For thirty years the only way to spot-check a junior estimator’s work was to diff quantities against a senior’s. It was the only measurement tool the trade had. The tool worked because both estimators were doing the same job. Quantifying scope, plus applying experience-adds. The comparison was apples to apples.
AI changes the job. AI does one half of what the estimator does, extremely fast, at the level of what’s actually drawn. The estimator still owns the other half: catching un-drawn scope, applying the trade knowledge that turns a set of lines into a priceable takeoff. The two halves have always been there. The estimator just used to do both.

So the question isn’t “did the AI match my number?”, that’s asking whether the AI did both halves of the job, which it never will. The question is “how long from drawings-in to a takeoff I’d stake my name on?”
That question has a real answer. And it’s a metric that actually maps to money.
The metric that matters: elapsed hours to a reviewed takeoff
On the same hospital set, here’s what happened.
Manual takeoff, cold, senior estimator: 3 hours 47 minutes. AI takeoff plus my review pass: 22 minutes. I caught the un-drawn transitions the AI couldn’t see. I added the fittings the drawings implied. I bumped the four sections of round duct where the equipment schedule made the connection obvious. The number I sent to procurement wasn’t the AI’s number and it wasn’t a pure manual number either. It was the AI’s number plus my twenty-two minutes of experience.
3 hours 47 minutes to 22 minutes. Same estimator. Same experience. Same drawings. Same priceable output.
On mechanical takeoffs I’ve been running this way for two months, the pattern is consistent: about 4 hours to 1 hour including my review. On structural steel takeoffs, 50 minutes to 5 minutes. On electrical, 6 hours to 45 minutes. The multiplier varies by trade because the drawings vary. Mechanical is dense, steel is sparse, electrical falls in between. But the shape is the same. AI compresses the quantify-what’s-drawn half from hours to seconds. My review pass, the half AI can’t do, stays where it always was.
That’s the metric. Elapsed hours from drawings-in to a reviewed, priceable takeoff. Not quantity match. Not F1 score. Not accuracy percentage. Hours.
What my review pass actually looks for
If I’m not diffing quantities, what am I doing in the 22 minutes?
I’m checking the AI’s counts against the equipment schedule, which is where the un-drawn scope hides. I’m looking at every duct transition and asking whether the size change implies a fitting the AI didn’t count. I’m scanning the drawing for the geometry the drafter left out. The room-to-room connection that has to exist for the mechanical to work at all. I’m spot-checking the AI on drawn scope where I know estimators, human and AI, get it wrong: complex junctions, curved runs, sheets with poor legibility.
Twenty-two minutes of trade experience applied where it matters. Not three hours forty-seven minutes of counting.
The AI didn’t replace my expertise. It gave me back three and a half hours of my week.

Why this reframe matters for every estimator evaluating AI right now
If you’re evaluating AI takeoff software today, the vendors will hand you their accuracy numbers and you’ll do what estimators have always done: pull the same drawings through your own workflow, diff the quantities, and either buy the tool or bounce it. That evaluation will tell you almost nothing.
It won’t tell you how much time the tool actually saves once you factor in your review pass. It won’t tell you whether the AI is missing scope you can’t afford to have missed, versus missing un-drawn scope you were always going to add yourself. It won’t tell you whether your team’s review time is a bottleneck or a rounding error.
Try this instead. Take one bid you already priced manually. Run it through the AI. Time your review pass from the moment you open the AI’s output to the moment you’d sign off on the number. Compare that elapsed time against your original manual hours. That’s the number that maps to margin, to bid capacity, to how many pursuits you can chase per quarter with the same team.

If the AI cuts your elapsed time by 4x, you can bid 4x the work with the same team. If it cuts by 8x, the math on your business changes. That’s the evaluation question. Not “did it match my number.”
The 78% quantity match on the hospital set doesn’t matter. The 22 minutes does.
Cristian Meyer is a Solutions Engineer at Boon. He runs takeoffs on real customer sets weekly. If you want to run the same test on one of your own bids, get in touch, first run is on us.
Technical review: Arjun Bansal, Boon AI.