Skip to content

FeatureAI

AI benchmarkARCH-BHarvard and Northeastern

Top AI model matches building drawings to photos 83.9% of the time in the ARCH-B benchmark

Two researchers tested 25 AI models on 354 multiple-choice questions that pair photos, floor plans, elevations and sections of the same building. Gemini 3.1 Pro Preview scored 83.90%, against 35.35% for untrained people recruited online.

Fig. 1 The top model scored 83.90%. Untrained people averaged 35.35%.
020406080100%TOP MODEL83.90%Gemini 3.1 Pro PreviewAVERAGE MODEL, PER QUESTION53.36%UNTRAINED PEOPLE35.35%A GUESS25%EACH SQUARE IS ONE OF 25 MODELS020406080100%TOP MODEL83.90%Gemini 3.1 Pro PreviewAVERAGE MODEL, PER QUESTION53.36%UNTRAINED PEOPLE35.35%A GUESS25%25 MODELS

Accuracy on 354 four-choice questions, percent. Each square is one of the 25 models. Source: ARCH-B, arXiv, posted September 28, 2026.

Title
ARCH-B: Architectural Representation, Comprehension and Hierarchy Benchmark
Questions
354
Task types
11
Models
25
Human answers
5,830
Posted
Sept. 28, 2026
Doc no.
arXiv:2609.34047v1

Two researchers from Harvard and Northeastern University posted ARCH-B to arXiv on September 28, 2026. It is a benchmark, a fixed set of test questions, with 354 multiple-choice items that ask an AI model to match photographs, floor plans, elevations and sections of the same building. Kieran Sagar Parikh and Jose Luis Garcia del Castillo y Lopez ran it on 25 multimodal models, meaning models that read images as well as text, and Gemini 3.1 Pro Preview scored highest at 83.90%. Untrained people recruited online averaged 35.35%.

The skill under test comes up in tools that relate a drawing set to the building, such as software that pins field photos to a sheet or AI takeoff that reads plan PDFs.

Each question offers four images, so a guess scores 25%

Every question has a short instruction, one or more reference images, and four answer images with one correct choice, so a model that guesses scores 25%. The 354 questions fall into 11 task types. Most ask for the image that matches the references, in either direction between drawings and photos. Two task types ask for the one image that belongs to a different building.

“Choose the photograph that corresponds to the building shown in this/these floorplan/s.”

A typical ARCH-B instruction

The images come from real-estate listings on Zillow and Realtor.com and from project pages on ArchDaily and Dezeen, collected daily from November 2024 to October 2025. That gave the authors 3.9 million images, 58,000 of them floor plans, grouped by the listing or project they came from. The grouping is what makes an answer checkable, because every image in a group shows the same building.

The wrong answers are deliberately similar to the right one. For each question, the authors picked the three images from other buildings that an image model rated most similar to the correct one, so each decoy is a similar-looking building or drawing. They then cut the pool down in three steps:

  1. Generate 2,200 candidate questions, 200 per task type.
  2. Run them through four screening models (Claude Sonnet 4, Gemini 2.5 Pro, Pixtral Large and GPT-4.1) and keep the 1,013 that two or fewer of the four answered correctly.
  3. Have two researchers review every remaining question by hand and drop any with duplicate images, mislabeled drawing types, or no single answer a person could defend.
Fig. 2 Screening and review cut 2,200 candidate questions to 354
05001,0001,5002,000CANDIDATE QUESTIONS, 200 PER TASK TYPE2,200KEPT AFTER SCREENING, 2 OR FEWER OF 4 MODELS RIGHT1,013KEPT AFTER REVIEW BY TWO RESEARCHERS35405001,0001,5002,000CANDIDATE QUESTIONS, 200 PER TASK TYPE2,200KEPT AFTER SCREENING, 2 OR FEWER OF 4 MODELS RIGHT1,013KEPT AFTER REVIEW BY TWO RESEARCHERS354

Questions at each stage of curation. The pool came from 3.9 million images in 95,000 real-estate listings and 41,000 architectural projects. Source: Parikh and Garcia del Castillo y Lopez, ARCH-B, arXiv.

In a contractor’s office, the likely user is a VDC manager or technology lead comparing vendors whose products read drawings. The released runner lets you score a given model on the same 354 questions. The source images are listing photos and published design work, and the set contains no construction document sets or progress photos. The authors name construction documents and images from different stages of construction as possible additions in future versions.

Gemini 3.1 Pro Preview led 25 models at 83.90%

The results come from the authors’ own runs of the 25 models and a crowdsourced human study, all reported in the paper. A replication by another group is not yet public.

Fig. 3 Scores ran from 83.90% for Gemini 3.1 Pro Preview down to 10.45% for Gemma 3 4B
020406080100%RANK AND MODELDEVELOPERACCURACYVALID1Gemini 3.1 Pro PreviewGoogle83.9097.462Gemini 3.5 FlashGoogle83.0598.873Claude Fable 5Anthropic80.2398.874GPT-4.1OpenAI77.97100.005GPT-5.5 reasoningOpenAI76.5599.726Gemini 3 Flash PreviewGoogle75.9998.027GPT-5.4 reasoningOpenAI74.2999.728Gemini 2.5 ProGoogle71.47100.009GPT-5.5OpenAI70.9099.7210Claude Opus 4.6Anthropic70.34100.0011Claude Opus 4.5Anthropic65.2599.1512GPT-5.4OpenAI64.4198.8713Grok 4.20 reasoningxAI57.0699.1514Grok 4.3xAI50.2898.8715GPT-5.4 mini reasoningOpenAI50.00100.0016GPT-5.4 miniOpenAI46.6199.7217Gemma 3 27BGoogle38.9899.7218Claude Sonnet 4.6Anthropic37.8560.4519Claude Sonnet 4.5Anthropic34.4668.0820Pixtral LargeMistral32.7797.1821GPT-5.4 nano reasoningOpenAI29.1099.7222Claude Haiku 4.5Anthropic17.5194.6323Gemini 2.5 FlashGoogle17.5124.2924GPT-5.4 nanoOpenAI16.9599.4425Gemma 3 4BGoogle, Ollama10.45100.00A GUESSUNTRAINED PEOPLE0255075100%RANK AND MODELACCURACYVALID1Gemini 3.1 Pro Preview83.9097.462Gemini 3.5 Flash83.0598.873Claude Fable 580.2398.874GPT-4.177.97100.005GPT-5.5 reasoning76.5599.726Gemini 3 Flash Preview75.9998.027GPT-5.4 reasoning74.2999.728Gemini 2.5 Pro71.47100.009GPT-5.570.9099.7210Claude Opus 4.670.34100.0011Claude Opus 4.565.2599.1512GPT-5.464.4198.8713Grok 4.20 reasoning57.0699.1514Grok 4.350.2898.8715GPT-5.4 mini reasoning50.00100.0016GPT-5.4 mini46.6199.7217Gemma 3 27B38.9899.7218Claude Sonnet 4.637.8560.4519Claude Sonnet 4.534.4668.0820Pixtral Large32.7797.1821GPT-5.4 nano reasoning29.1099.7222Claude Haiku 4.517.5194.6323Gemini 2.5 Flash17.5124.2924GPT-5.4 nano16.9599.4425Gemma 3 4B10.45100.00GUESSPEOPLE

Accuracy and valid response rate on the 354 questions, percent. Clouds mark valid response rates under 70%, where scores partly measure formatting failures. Gemma 3 4B was run on Ollama. Source: Parikh and Garcia del Castillo y Lopez, ARCH-B, arXiv.

Models scored 58.78% from photos to plan and 44.14% the other way

The more useful finding is in the task table. Averaged across all 25 models, accuracy was 58.78% when the model had photos and had to pick the floor plan, and 44.14% when it had floor plans and had to pick the photo. People scored about the same in both directions, 34.74% and 36.41%.

Fig. 4 The 25-model average fell from 58.78% to 44.14% when the direction flipped
PHOTOS TOFLOOR PLANFLOOR PLANSTO PHOTOA GUESS, 25%44.14%58.78%AVERAGE OF 25 MODELS36.41%34.74%UNTRAINED PEOPLEPHOTOS TOFLOOR PLANFLOOR PLANSTO PHOTOGUESS44.14%58.78%MODELS36.41%34.74%PEOPLE

Average accuracy on the two floor plan tasks, percent: 23 questions that give 2 to 4 photos and ask for the plan, and 56 that give 1 to 4 plans and ask for the photo. Source: Parikh and Garcia del Castillo y Lopez, ARCH-B, arXiv.

The authors read this as models finding it easier to check a plan’s layout against views they can see than to picture how a plan will look once built.

The best model beat untrained people on every task by 40.38 to 68.00 points

The best model on each task beat the human average by 40.38 to 68.00 percentage points.

Fig. 5 The narrowest lead over people was 40.38 points, on picking the photo from floor plans
020406080100%UNTRAINED PEOPLEAVERAGE OF 25 MODELSBEST MODELA GUESS, 25%Pick the floor plan that matches 2 to 4 photos95.6534.74Pick the photo that matches 1 to 4 floor plans76.7936.4140.38 POINTSFind the photo from a different building84.7821.61Pick the photo that matches three others100.0032.0068.00 POINTSPick the interior that matches 4 exterior photos96.0035.85Pick the exterior that matches 4 interior photos85.0029.70Pick the elevation that matches the floor plans90.6236.95Pick the floor plan that matches the elevations90.0032.25Find the odd one out among a plan, elevation, section and photo100.0033.93Pick the photo that matches a plan, elevation and section96.5542.44Pick the exterior photo that matches 2 elevations96.8856.150255075100%PEOPLE25-MODEL AVERAGEBESTGUESSPick the floor plan that matches 2 to 4 photos95.6534.74Pick the photo that matches 1 to 4 floor plans76.7936.4140.38 POINTSFind the photo from a different building84.7821.61Pick the photo that matches three others100.0032.0068.00 POINTSPick the interior that matches 4 exterior photos96.0035.85Pick the exterior that matches 4 interior photos85.0029.70Pick the elevation that matches the floor plans90.6236.95Pick the floor plan that matches the elevations90.0032.25Find the odd one out among a plan, elevation,section and photo100.0033.93Pick the photo that matches a plan, elevation andsection96.5542.44Pick the exterior photo that matches 2 elevations96.8856.15

Accuracy by task type, percent. The best model differs by task. Dimension lines mark the smallest and largest gaps between the best model and untrained people. Source: Parikh and Garcia del Castillo y Lopez, ARCH-B, arXiv.

That human baseline needs a caveat the authors state themselves. Participants were English-speaking US adults on the Prolific research platform who each answered 10 questions from one task type (they were paid $2 per session whatever their score), and the authors call them a non-expert baseline. They suggest future comparisons with architecture students and practitioners. A superintendent or architect who reads elevations every day would be a different reference point.

Questions no screening model solved held the other models to 28.65%

The authors also checked whether screening with four models had made the benchmark hard only for those four. They grouped questions by how many screening models got them right and scored the models outside the screening set on each group.

Fig. 6 Accuracy rose from 28.65% to 59.02% as screening difficulty fell
020406080100%NO SCREENING MODEL RIGHT33 QUESTIONS28.65%ONE OF FOUR RIGHT77 QUESTIONS41.38%TWO OF FOUR RIGHT244 QUESTIONS59.02%A GUESS0255075100%NO SCREENING MODEL RIGHT, 33 QUESTIONS28.65%ONE OF FOUR RIGHT, 77 QUESTIONS41.38%TWO OF FOUR RIGHT, 244 QUESTIONS59.02%A GUESS

Accuracy of the models outside the screening set, percent, by how many of the four screening models answered each question correctly. Source: Parikh and Garcia del Castillo y Lopez, ARCH-B, arXiv.

Accuracy rose from 28.65% to 59.02% as screening difficulty fell, which the authors take as evidence that the hard questions are hard in general. Per question, average model accuracy was 53.36%, with a range from 4% to 96%. No question was answered correctly by every model, and none was missed by all of them.

Check a photo tool’s location match rate in both directions

So the paper’s split by direction is worth copying. If you pilot a tool that ties photos to drawings, such as a reality capture product or a photo log that places shots on sheets, track the location match rate. It is the share of a fixed sample of items that the tool placed on the correct sheet and in the correct room or grid bay, as checked by the project engineer who knows the job.

Measure it both ways, since the models in ARCH-B scored differently in each direction:

Photo to plan

Take field photos the tool has already placed and check each placement against the sheet.

Plan to photo

Pick rooms or grid bays on a sheet, ask the tool for the matching photos, and check that each returned photo shows that area.

Capture the baseline before the pilot starts. Pull a sample from last month’s photo log, record how many photos carry a correct location tag, and time how long the project engineer spends tagging a week’s photos by hand. AI takeoff tools that read plan PDFs, such as the one Quotr raised a seed round for, can be checked the same way by comparing a sample of extracted quantities to an estimator’s count.

A first step on one job is to pick a single floor and one drawing sheet, and have the project engineer mark the correct location for each photo taken there since the last pay application.

ARCH-B is free, and the runner is under the MIT License

ARCH-B is free. The question set, a JSON schema and the evaluation runner are posted as ancillary files with the arXiv paper, and the README and runner are under the MIT License. The question file links to each image at its original public address with a SHA-256 hash, a fingerprint the runner uses to confirm the image has not changed, and the authors do not redistribute the images themselves.

The runner calls each model provider’s API, the paid connection a program uses to send questions to the model, and the README notes that provider use can incur charges. The authors also host a version of the human-study quiz at archbench.codecolab.org, where anyone can try sample questions.

The data for all 11 task types and 354 questions

ARCH-B accuracy by task type, percent. Source: ARCH-B, arXiv.
TaskQuestionsBest modelBest25-model averagePeoplePeople Average Best
Pick the floor plan that matches 2 to 4 photos23Gemini 3.1 Pro Preview95.6558.7834.74
Pick the photo that matches 1 to 4 floor plans56GPT-4.176.7944.1436.41
Find the photo from a different building46Gemini 3.1 Pro Preview84.7841.2221.61
Pick the photo that matches three others24GPT-4.1100.0060.1732.00
Pick the interior that matches 4 exterior photos25Gemini 3.1 Pro Preview96.0052.6435.85
Pick the exterior that matches 4 interior photos20Claude Opus 4.585.0050.0029.70
Pick the elevation that matches the floor plans32Gemini 3.5 Flash90.6250.2536.95
Pick the floor plan that matches the elevations30Gemini 3.5 Flash90.0057.8732.25
Find the odd one out among a plan, elevation, section and photo37Claude Fable 5100.0064.2233.93
Pick the photo that matches a plan, elevation and section29Gemini 3 Flash Preview96.5559.8642.44
Pick the exterior photo that matches 2 elevations32Gemini 3.5 Flash96.8861.0056.15

Download the data: tasks (CSV), leaderboard (CSV), screening (CSV), curation counts (CSV).