AITakeoff benchmarkTogal.AI
Togal.AI's takeoff benchmark scores the best general AI model at 34.0 out of 100 on 33 drawing tasks
Togal.AI, which sells AI takeoff software, built Construction's Last Exam from 33 quantity takeoff tasks on real drawings and ran the general AI models through it. OpenAI's GPT-6 Sol led them at 34.0. Togal's own platform, run with a person in the loop, scored 97.7.
- GPT-6 Sol, best general model
- 34.0
- Togal.ai with a person in the loop
- 97.7
- Takeoff tasks on real drawings
- 33
Leaderboard, first page
GPT-6 Sol led the general models at 34.0, and GPT-6 Luna came second at 21.6
Overall score out of 100, the mean of the linear, area and counting family scores
Construction's Last Exam leaderboard, first page, as read on October 8, 2026. The top row is Togal's own platform with a person in the loop, which Togal labels a reference mark. The dashed line carries its 97.7 down the chart. API cost is Togal's scaled estimate for one run through all 33 tasks. Source: Togal.AI, Construction's Last Exam
Togal.AI, a Miami company that sells AI takeoff software, launched Construction's Last Exam on October 6, 2026. It is a benchmark, a fixed set of test tasks scored the same way for every model, with 33 quantity takeoff tasks drawn on real construction drawings. The tasks ask a model to trace net room areas, run wall centerlines for drywall, and locate every receptacle or sprinkler head on a sheet. On Togal's leaderboard the best general-purpose model, OpenAI's GPT-6 Sol, scored 34.0 out of 100. Togal's own platform, run with a person in the loop, scored 97.7.
What it does
Each of the 33 tasks gives a model one drawing and one instruction
Each task gives a model a drawing and one instruction, and the model has to return geometry. Togal sorts the 33 tasks into three families, and the overall score is the mean of the three family scores.
The exam
Togal split the 33 tasks into 12 area, 12 counting and 9 linear takeoffs
Each square is one task. Repeated instructions run on different drawings.
Areas
12 tasks
Traced regions such as net room area, balconies, building footprint, siding and stone on elevations, and paving and sod on site plans.
- Trace every balcony drawn on the plan
- Trace net floor area of every room×3
- Trace every Floratam sod area on landscape
- Trace the threshold area of every door
- Trace the building footprint on the plan
- Trace every siding-finished wall area on elevation
- Trace every stone-clad area on the elevation
- Extract net area of every hatched classroom
- Trace the corridor net area on the plan
- Trace every paved area on the site plan
Counting
12 tasks
The location of every instance of a symbol, such as toilets, sinks, duplex receptacles, smoke detectors, light fixtures and upright sprinkler heads.
- Locate every B- or F-tagged light fixture
- Locate every instance of the target symbol
- Locate every toilet shown on the plan
- Locate every illuminance calculation point on plan
- Locate every sink shown on the plan
- Locate every duplex receptacle on the drawing×2
- Locate every room light on the drawing
- Locate every smoke detector on the drawing
- Locate every Blue Mist plant on the sheet
- Locate every upright sprinkler head on plan
- Locate every toilet including those in stalls
Linear
9 tasks
Polylines along wall centerlines for drywall, door centerlines, and wall perimeters for painters.
- Trace walls with centerlines on the plan×2
- Trace door centerlines for the door contractor
- Trace door centerlines cutting thresholds in half
- Trace wall centerlines for drywallers×4
- Trace wall perimeters for painters
Task prompts as Togal lists them. Sources: Togal.AI, Prompts and Model Answers and overview
The instructions read like the way takeoff work gets divided on a real bid. Four of the nine linear tasks say "Trace wall centerlines for drywallers." Others ask for door centerlines for the door contractor, wall perimeters for painters, siding and stone areas on elevations, and paved and sod areas on a site plan. The counting tasks include duplex receptacles, smoke detectors, upright sprinkler heads, and toilets including those in stalls. The press release describes one task as measuring every balcony in a building and scoring the answer against the correct measurements.
So the output is a set of shapes, the same polygons, polylines and points an estimator draws in takeoff software. A related test of how models read drawings, ARCH-B, asks models to match drawings to photos of the same building.
In a contractor's office, the reader for this leaderboard is the preconstruction manager or technology lead deciding whether estimators should do takeoff in a general chatbot such as ChatGPT or Claude or in a takeoff product.
Evidence so far
Togal wrote the 33 tasks, ran the models and scored its own platform at 97.7
The evidence is Togal's own benchmark site and its press release. Togal wrote the tasks, its estimators validated the answer key, and it ran the models and published the scores, its own platform included. Digital Journal reported that the results have not yet been checked through peer review or by a third party. The rule for grading a traced polygon or a located symbol against the answer key is not described on the overview page.
The table supports these readings.
- The spread among general models is wide. GPT-6 Sol at 34.0 led GPT-6 Luna, in second among general models, by 12.4 points. The rest of the first page sits between 12.3 and 21.6.
- Cost does not line up with score. Togal scales each provider's published API rates, the per-use charges for sending work to a model through its programming interface, so that one full run of GPT-6 Astra costs about $14.0. On that scale GPT-6 Sol cost about $3.5 and scored 34.0, while GPT-6 Astra scored 19.2.
Cost and score
GPT-6 Sol cost about $3.5 a run and scored 34.0, while GPT-6 Astra cost about $14.0 and scored 19.2
API cost for one run through all 33 tasks, in US dollars, against overall score out of 100
Nine general models on the leaderboard's first page. Togal scales each provider's published API rates so that one full run of GPT-6 Astra costs about $14.0. Togal's own platform has no API cost listed and is left out. Source: Togal.AI, Construction's Last Exam
- Counting was the strongest family for the four highest-scoring general models. GPT-6 Sol scored 40.5 on counting and 29.8 on areas. Claude Opus 5.5 did best on areas, at 20.6.
Scores by task family
GPT-6 Sol scored 40.5 on counting, its strongest family, and 29.8 on areas
Family score out of 100, three panels on one scale. The dashed line in each panel is Togal's platform with a person in the loop.
Eight general models with family scores on the leaderboard's first page. DeepSeek Flash 4.1 has an overall score only. Source: Togal.AI, Construction's Last Exam
Togal's 97.7 needs its label read carefully. The progress chart calls it "a reference mark (human-in-the-loop, max. 30 seconds)," and the leaderboard tags the row "Ours." The general-model rows carry no such label. So the 97.7 combines Togal's software with a person's input, capped at 30 seconds, and the distance to 34.0 compares that setup with models working alone.
The progress chart also shows how fast the general models moved. Togal lists 23 general-model runs by release date.
Progress by release date
No listed model scored above 3.9 through February 2026, and GPT-6 Sol reached 34.0 on September 22
Overall score out of 100 for 23 general-model runs, by model release date. The stepped line is the best score to date.
Scores as shown on Togal's progress chart, which can differ from the leaderboard in the last decimal. Source: Togal.AI, Construction's Last Exam
Through February 2026, no listed model scored above 3.9. GPT-5.6 Sol reached 11.1 in July, Gemini 3.8 Flash and GPT-6 Astra both reached 19.2 in early September, and GPT-6 Sol reached 34.0 on September 22. Two scores differ between the chart and the leaderboard (Gemini 3.8 Flash shows 19.2 and 19.1, Claude Opus 5 shows 12.1 and 12.7), so treat the last decimal as approximate.
Togal's press release calls the exam the first AI benchmark built for the construction industry. The release says Togal has trained its own vision models since 2019 on tens of millions of construction plans and serves more than 10,000 users in 30 countries.
How to measure it on your projects
Test a tool on the 10 to 15 costliest line items from two or three recent bids
Anyway, the exam runs on drawings Togal chose, and your own sheets are the better test for your trades. If you pilot any AI takeoff tool, whether a general model or a product such as Quotr or STACK IQ, track quantity variance by line item. It is the difference between the tool's quantity and the estimator's checked quantity, divided by the estimator's quantity, for each line such as LF of drywall partition, SF of flooring, or count of receptacles.
Capture the baseline before the pilot:
- Pick two or three recent bids where the takeoff was checked and the job went ahead, so field quantities can confirm the estimate.
- Record the estimator's hours on takeoff for each bid and the quantity for 10 to 15 line items that carry the most cost.
- Run the same sheets through the tool and compute the variance for each line, plus the estimator's time to review and correct the tool's output.
Split the results the way Togal does, into linear, area, and count items, since the models in the exam scored differently on each. If a tool counts fixtures well and misses wall centerlines, you can use it on count items and keep linear takeoff manual. A first step on one job is to take a single floor plan from a recent bid and have the estimator compare the tool's receptacle count and drywall LF against the numbers that went into the bid.
Availability and cost
One model's run through all 33 tasks cost about $0.2 to $14.0 in API charges
The benchmark site is public at cle.togal.ai, with a page listing all 33 task prompts. Togal says the drawing PDFs, its own baseline output, and the general models' run files are in a public Google Drive folder. On the first page of the leaderboard, one model's run through all 33 tasks ranged from about $0.2 to about $14.0 in API charges. The GitHub repository linked from the site returned a not-found page when this post was written on October 8. Togal's overview page marks this as a geometry-only release and links a preview of a fuller suite covering reasoning and end-to-end tasks.
The data
The leaderboard's first page runs from 97.7 down to 12.3
| Rank | Model | Maker | Linear | Areas | Counting | API cost, all 33 tasks | Overall score |
|---|---|---|---|---|---|---|---|
| 1 | Togal.ai (human in the loop) | Togal.ai | 96.8 | 98.2 | 97.8 | not listed | 97.7 |
| 2 | GPT-6 Sol | OpenAI | 30.9 | 29.8 | 40.5 | $3.5 | 34.0 |
| 3 | GPT-6 Luna | OpenAI | 15.3 | 21.3 | 26.5 | $0.2 | 21.6 |
| 4 | GPT-6 Astra | OpenAI | 21.8 | 12.1 | 24.5 | $14.0 | 19.2 |
| 5 | Gemini 3.8 Flash | 18.7 | 13.6 | 25.0 | $1.2 | 19.1 | |
| 6 | Gemini 3.7 Flash | 14.5 | 18.5 | 16.4 | $1.1 | 16.6 | |
| 7 | Claude Opus 5.5 | Anthropic | 11.8 | 20.6 | 15.0 | $8.0 | 16.2 |
| 8 | Claude Opus 5 | Anthropic | 14.8 | 10.6 | 12.7 | $7.5 | 12.7 |
| 9 | Claude Fable 5.1 | Anthropic | 8.7 | 16.9 | 11.1 | $13.3 | 12.5 |
| 10 | DeepSeek Flash 4.1 | DeepSeek | not listed | not listed | not listed | $0.2 | 12.3 |
The overall score is the mean of the three family scores. Rank 1 is Togal's platform with a person in the loop, capped at 30 seconds. The results have not yet been checked through peer review or by a third party. Source: Togal.AI, Construction's Last Exam. Data: leaderboard, progress by release date, tasks, task families (CSV).