Skip to content

FeatureAI

AITakeoff benchmarkTogal.AI

Togal.AI's takeoff benchmark scores the best general AI model at 34.0 out of 100 on 33 drawing tasks

Togal.AI, which sells AI takeoff software, built Construction's Last Exam from 33 quantity takeoff tasks on real drawings and ran the general AI models through it. OpenAI's GPT-6 Sol led them at 34.0. Togal's own platform, run with a person in the loop, scored 97.7.

GPT-6 Sol, best general model
34.0
Togal.ai with a person in the loop
97.7
Takeoff tasks on real drawings
33

Leaderboard, first page

GPT-6 Sol led the general models at 34.0, and GPT-6 Luna came second at 21.6

Overall score out of 100, the mean of the linear, area and counting family scores

020406080100RANKMODELMAKEROVERALL SCORE, 0 TO 100SCOREAPI COST1Togal.ai, person in the loopTogal.ai97.7n/a2GPT-6 SolOpenAI34.0$3.53GPT-6 LunaOpenAI21.6$0.24GPT-6 AstraOpenAI19.2$14.05Gemini 3.8 FlashGoogle19.1$1.26Gemini 3.7 FlashGoogle16.6$1.17Claude Opus 5.5Anthropic16.2$8.08Claude Opus 5Anthropic12.7$7.59Claude Fable 5.1Anthropic12.5$13.310DeepSeek Flash 4.1DeepSeek12.3$0.2Togal.ai with a personin the loop, 97.7 0255075100OVERALL SCORE, 0 TO 100SCORE1Togal.ai, person in the loop97.72GPT-6 Sol34.03GPT-6 Luna21.64GPT-6 Astra19.25Gemini 3.8 Flash19.16Gemini 3.7 Flash16.67Claude Opus 5.516.28Claude Opus 512.79Claude Fable 5.112.510DeepSeek Flash 4.112.3Togal.aireference, 97.7

Construction's Last Exam leaderboard, first page, as read on October 8, 2026. The top row is Togal's own platform with a person in the loop, which Togal labels a reference mark. The dashed line carries its 97.7 down the chart. API cost is Togal's scaled estimate for one run through all 33 tasks. Source: Togal.AI, Construction's Last Exam

Togal.AI, a Miami company that sells AI takeoff software, launched Construction's Last Exam on October 6, 2026. It is a benchmark, a fixed set of test tasks scored the same way for every model, with 33 quantity takeoff tasks drawn on real construction drawings. The tasks ask a model to trace net room areas, run wall centerlines for drywall, and locate every receptacle or sprinkler head on a sheet. On Togal's leaderboard the best general-purpose model, OpenAI's GPT-6 Sol, scored 34.0 out of 100. Togal's own platform, run with a person in the loop, scored 97.7.

What it does

Each of the 33 tasks gives a model one drawing and one instruction

Each task gives a model a drawing and one instruction, and the model has to return geometry. Togal sorts the 33 tasks into three families, and the overall score is the mean of the three family scores.

The exam

Togal split the 33 tasks into 12 area, 12 counting and 9 linear takeoffs

Each square is one task. Repeated instructions run on different drawings.

Areas

12 tasks

Traced regions such as net room area, balconies, building footprint, siding and stone on elevations, and paving and sod on site plans.

  • Trace every balcony drawn on the plan
  • Trace net floor area of every room×3
  • Trace every Floratam sod area on landscape
  • Trace the threshold area of every door
  • Trace the building footprint on the plan
  • Trace every siding-finished wall area on elevation
  • Trace every stone-clad area on the elevation
  • Extract net area of every hatched classroom
  • Trace the corridor net area on the plan
  • Trace every paved area on the site plan

Counting

12 tasks

The location of every instance of a symbol, such as toilets, sinks, duplex receptacles, smoke detectors, light fixtures and upright sprinkler heads.

  • Locate every B- or F-tagged light fixture
  • Locate every instance of the target symbol
  • Locate every toilet shown on the plan
  • Locate every illuminance calculation point on plan
  • Locate every sink shown on the plan
  • Locate every duplex receptacle on the drawing×2
  • Locate every room light on the drawing
  • Locate every smoke detector on the drawing
  • Locate every Blue Mist plant on the sheet
  • Locate every upright sprinkler head on plan
  • Locate every toilet including those in stalls

Linear

9 tasks

Polylines along wall centerlines for drywall, door centerlines, and wall perimeters for painters.

  • Trace walls with centerlines on the plan×2
  • Trace door centerlines for the door contractor
  • Trace door centerlines cutting thresholds in half
  • Trace wall centerlines for drywallers×4
  • Trace wall perimeters for painters

Task prompts as Togal lists them. Sources: Togal.AI, Prompts and Model Answers and overview

The instructions read like the way takeoff work gets divided on a real bid. Four of the nine linear tasks say "Trace wall centerlines for drywallers." Others ask for door centerlines for the door contractor, wall perimeters for painters, siding and stone areas on elevations, and paved and sod areas on a site plan. The counting tasks include duplex receptacles, smoke detectors, upright sprinkler heads, and toilets including those in stalls. The press release describes one task as measuring every balcony in a building and scoring the answer against the correct measurements.

So the output is a set of shapes, the same polygons, polylines and points an estimator draws in takeoff software. A related test of how models read drawings, ARCH-B, asks models to match drawings to photos of the same building.

In a contractor's office, the reader for this leaderboard is the preconstruction manager or technology lead deciding whether estimators should do takeoff in a general chatbot such as ChatGPT or Claude or in a takeoff product.

Evidence so far

Togal wrote the 33 tasks, ran the models and scored its own platform at 97.7

The evidence is Togal's own benchmark site and its press release. Togal wrote the tasks, its estimators validated the answer key, and it ran the models and published the scores, its own platform included. Digital Journal reported that the results have not yet been checked through peer review or by a third party. The rule for grading a traced polygon or a located symbol against the answer key is not described on the overview page.

The table supports these readings.

  1. The spread among general models is wide. GPT-6 Sol at 34.0 led GPT-6 Luna, in second among general models, by 12.4 points. The rest of the first page sits between 12.3 and 21.6.
  2. Cost does not line up with score. Togal scales each provider's published API rates, the per-use charges for sending work to a model through its programming interface, so that one full run of GPT-6 Astra costs about $14.0. On that scale GPT-6 Sol cost about $3.5 and scored 34.0, while GPT-6 Astra scored 19.2.

Cost and score

GPT-6 Sol cost about $3.5 a run and scored 34.0, while GPT-6 Astra cost about $14.0 and scored 19.2

API cost for one run through all 33 tasks, in US dollars, against overall score out of 100

010203040$0$5$10$15API COST FOR ALL 33 TASKS, USDOVERALL SCOREGPT-6 Sol$3.5, 34.0GPT-6 LunaGPT-6 Astra$14.0, 19.2Gemini 3.8 FlashGemini 3.7 FlashClaude Opus 5.5Claude Opus 5Claude Fable 5.1DeepSeek Flash 4.1 010203040$0$5$10$15API COST FOR ALL 33 TASKS, USDOVERALL SCOREGPT-6 Sol$3.5, 34.0GPT-6 LunaGPT-6 Astra$14.0, 19.2Gemini 3.8 FlashGemini 3.7 FlashClaude Opus 5.5Claude Opus 5Claude Fable 5.1DeepSeek Flash 4.1

Nine general models on the leaderboard's first page. Togal scales each provider's published API rates so that one full run of GPT-6 Astra costs about $14.0. Togal's own platform has no API cost listed and is left out. Source: Togal.AI, Construction's Last Exam

  1. Counting was the strongest family for the four highest-scoring general models. GPT-6 Sol scored 40.5 on counting and 29.8 on areas. Claude Opus 5.5 did best on areas, at 20.6.

Scores by task family

GPT-6 Sol scored 40.5 on counting, its strongest family, and 29.8 on areas

Family score out of 100, three panels on one scale. The dashed line in each panel is Togal's platform with a person in the loop.

GPT-6 SolGPT-6 LunaGPT-6 AstraGemini 3.8 FlashGemini 3.7 FlashClaude Opus 5.5Claude Opus 5Claude Fable 5.1LINEAR9 tasks050100Togal 96.830.9AREAS12 tasks050100Togal 98.229.820.6COUNTING12 tasks050100Togal 97.840.5 LINEAR9 tasks050100Togal 96.8GPT-6 Sol30.9GPT-6 LunaGPT-6 AstraGemini 3.8 FlashGemini 3.7 FlashClaude Opus 5.5Claude Opus 5Claude Fable 5.1AREAS12 tasks050100Togal 98.2GPT-6 Sol29.8GPT-6 LunaGPT-6 AstraGemini 3.8 FlashGemini 3.7 FlashClaude Opus 5.520.6Claude Opus 5Claude Fable 5.1COUNTING12 tasks050100Togal 97.8GPT-6 Sol40.5GPT-6 LunaGPT-6 AstraGemini 3.8 FlashGemini 3.7 FlashClaude Opus 5.5Claude Opus 5Claude Fable 5.1

Eight general models with family scores on the leaderboard's first page. DeepSeek Flash 4.1 has an overall score only. Source: Togal.AI, Construction's Last Exam

Togal's 97.7 needs its label read carefully. The progress chart calls it "a reference mark (human-in-the-loop, max. 30 seconds)," and the leaderboard tags the row "Ours." The general-model rows carry no such label. So the 97.7 combines Togal's software with a person's input, capped at 30 seconds, and the distance to 34.0 compares that setup with models working alone.

The progress chart also shows how fast the general models moved. Togal lists 23 general-model runs by release date.

Progress by release date

No listed model scored above 3.9 through February 2026, and GPT-6 Sol reached 34.0 on September 22

Overall score out of 100 for 23 general-model runs, by model release date. The stepped line is the best score to date.

010203040JAN 2025JUL 2025JAN 2026JUL 2026OVERALL SCORE, GENERAL-PURPOSE MODELSThrough February 2026, no listed model above 3.9GPT-5.6 Sol11.1 in JulyGemini 3.8 Flash and GPT-6 Astra19.2 in early SeptemberGPT-6 Sol34.0 on September 22 010203040JAN 2025JULJAN 2026JULOVERALL SCORE, GENERAL-PURPOSE MODELSThrough Feb 2026,none above 3.9GPT-5.6 Sol11.1, July19.2, earlySeptemberGPT-6 Sol34.0, Sep 22

Scores as shown on Togal's progress chart, which can differ from the leaderboard in the last decimal. Source: Togal.AI, Construction's Last Exam

Through February 2026, no listed model scored above 3.9. GPT-5.6 Sol reached 11.1 in July, Gemini 3.8 Flash and GPT-6 Astra both reached 19.2 in early September, and GPT-6 Sol reached 34.0 on September 22. Two scores differ between the chart and the leaderboard (Gemini 3.8 Flash shows 19.2 and 19.1, Claude Opus 5 shows 12.1 and 12.7), so treat the last decimal as approximate.

Togal's press release calls the exam the first AI benchmark built for the construction industry. The release says Togal has trained its own vision models since 2019 on tens of millions of construction plans and serves more than 10,000 users in 30 countries.

How to measure it on your projects

Test a tool on the 10 to 15 costliest line items from two or three recent bids

Anyway, the exam runs on drawings Togal chose, and your own sheets are the better test for your trades. If you pilot any AI takeoff tool, whether a general model or a product such as Quotr or STACK IQ, track quantity variance by line item. It is the difference between the tool's quantity and the estimator's checked quantity, divided by the estimator's quantity, for each line such as LF of drywall partition, SF of flooring, or count of receptacles.

Capture the baseline before the pilot:

  1. Pick two or three recent bids where the takeoff was checked and the job went ahead, so field quantities can confirm the estimate.
  2. Record the estimator's hours on takeoff for each bid and the quantity for 10 to 15 line items that carry the most cost.
  3. Run the same sheets through the tool and compute the variance for each line, plus the estimator's time to review and correct the tool's output.

Split the results the way Togal does, into linear, area, and count items, since the models in the exam scored differently on each. If a tool counts fixtures well and misses wall centerlines, you can use it on count items and keep linear takeoff manual. A first step on one job is to take a single floor plan from a recent bid and have the estimator compare the tool's receptacle count and drywall LF against the numbers that went into the bid.

Availability and cost

One model's run through all 33 tasks cost about $0.2 to $14.0 in API charges

The benchmark site is public at cle.togal.ai, with a page listing all 33 task prompts. Togal says the drawing PDFs, its own baseline output, and the general models' run files are in a public Google Drive folder. On the first page of the leaderboard, one model's run through all 33 tasks ranged from about $0.2 to about $14.0 in API charges. The GitHub repository linked from the site returned a not-found page when this post was written on October 8. Togal's overview page marks this as a geometry-only release and links a preview of a fuller suite covering reasoning and end-to-end tasks.

The data

The leaderboard's first page runs from 97.7 down to 12.3

Construction's Last Exam leaderboard, first page, scores out of 100
RankModelMakerLinearAreasCountingAPI cost, all 33 tasksOverall score
1Togal.ai (human in the loop)Togal.ai96.898.297.8not listed97.7
2GPT-6 SolOpenAI30.929.840.5$3.534.0
3GPT-6 LunaOpenAI15.321.326.5$0.221.6
4GPT-6 AstraOpenAI21.812.124.5$14.019.2
5Gemini 3.8 FlashGoogle18.713.625.0$1.219.1
6Gemini 3.7 FlashGoogle14.518.516.4$1.116.6
7Claude Opus 5.5Anthropic11.820.615.0$8.016.2
8Claude Opus 5Anthropic14.810.612.7$7.512.7
9Claude Fable 5.1Anthropic8.716.911.1$13.312.5
10DeepSeek Flash 4.1DeepSeeknot listednot listednot listed$0.212.3

The overall score is the mean of the three family scores. Rank 1 is Togal's platform with a person in the loop, capped at 30 seconds. The results have not yet been checked through peer review or by a third party. Source: Togal.AI, Construction's Last Exam. Data: leaderboard, progress by release date, tasks, task families (CSV).