AI benchmarkARCH-BHarvard and Northeastern
Top AI model matches building drawings to photos 83.9% of the time in the ARCH-B benchmark
Two researchers tested 25 AI models on 354 multiple-choice questions that pair photos, floor plans, elevations and sections of the same building. Gemini 3.1 Pro Preview scored 83.90%, against 35.35% for untrained people recruited online.
Accuracy on 354 four-choice questions, percent. Each square is one of the 25 models. Source: ARCH-B, arXiv, posted September 28, 2026.
- Title
- ARCH-B: Architectural Representation, Comprehension and Hierarchy Benchmark
- Questions
- 354
- Task types
- 11
- Models
- 25
- Human answers
- 5,830
- Posted
- Sept. 28, 2026
- Doc no.
- arXiv:2609.34047v1
Two researchers from Harvard and Northeastern University posted ARCH-B to arXiv on September 28, 2026. It is a benchmark, a fixed set of test questions, with 354 multiple-choice items that ask an AI model to match photographs, floor plans, elevations and sections of the same building. Kieran Sagar Parikh and Jose Luis Garcia del Castillo y Lopez ran it on 25 multimodal models, meaning models that read images as well as text, and Gemini 3.1 Pro Preview scored highest at 83.90%. Untrained people recruited online averaged 35.35%.
The skill under test comes up in tools that relate a drawing set to the building, such as software that pins field photos to a sheet or AI takeoff that reads plan PDFs.
Each question offers four images, so a guess scores 25%
Every question has a short instruction, one or more reference images, and four answer images with one correct choice, so a model that guesses scores 25%. The 354 questions fall into 11 task types. Most ask for the image that matches the references, in either direction between drawings and photos. Two task types ask for the one image that belongs to a different building.
“Choose the photograph that corresponds to the building shown in this/these floorplan/s.”
A typical ARCH-B instruction
The images come from real-estate listings on Zillow and Realtor.com and from project pages on ArchDaily and Dezeen, collected daily from November 2024 to October 2025. That gave the authors 3.9 million images, 58,000 of them floor plans, grouped by the listing or project they came from. The grouping is what makes an answer checkable, because every image in a group shows the same building.
The wrong answers are deliberately similar to the right one. For each question, the authors picked the three images from other buildings that an image model rated most similar to the correct one, so each decoy is a similar-looking building or drawing. They then cut the pool down in three steps:
- Generate 2,200 candidate questions, 200 per task type.
- Run them through four screening models (Claude Sonnet 4, Gemini 2.5 Pro, Pixtral Large and GPT-4.1) and keep the 1,013 that two or fewer of the four answered correctly.
- Have two researchers review every remaining question by hand and drop any with duplicate images, mislabeled drawing types, or no single answer a person could defend.
Questions at each stage of curation. The pool came from 3.9 million images in 95,000 real-estate listings and 41,000 architectural projects. Source: Parikh and Garcia del Castillo y Lopez, ARCH-B, arXiv.
In a contractor’s office, the likely user is a VDC manager or technology lead comparing vendors whose products read drawings. The released runner lets you score a given model on the same 354 questions. The source images are listing photos and published design work, and the set contains no construction document sets or progress photos. The authors name construction documents and images from different stages of construction as possible additions in future versions.
Gemini 3.1 Pro Preview led 25 models at 83.90%
The results come from the authors’ own runs of the 25 models and a crowdsourced human study, all reported in the paper. A replication by another group is not yet public.
Accuracy and valid response rate on the 354 questions, percent. Clouds mark valid response rates under 70%, where scores partly measure formatting failures. Gemma 3 4B was run on Ollama. Source: Parikh and Garcia del Castillo y Lopez, ARCH-B, arXiv.
Models scored 58.78% from photos to plan and 44.14% the other way
The more useful finding is in the task table. Averaged across all 25 models, accuracy was 58.78% when the model had photos and had to pick the floor plan, and 44.14% when it had floor plans and had to pick the photo. People scored about the same in both directions, 34.74% and 36.41%.
Average accuracy on the two floor plan tasks, percent: 23 questions that give 2 to 4 photos and ask for the plan, and 56 that give 1 to 4 plans and ask for the photo. Source: Parikh and Garcia del Castillo y Lopez, ARCH-B, arXiv.
The authors read this as models finding it easier to check a plan’s layout against views they can see than to picture how a plan will look once built.
The best model beat untrained people on every task by 40.38 to 68.00 points
The best model on each task beat the human average by 40.38 to 68.00 percentage points.
Accuracy by task type, percent. The best model differs by task. Dimension lines mark the smallest and largest gaps between the best model and untrained people. Source: Parikh and Garcia del Castillo y Lopez, ARCH-B, arXiv.
That human baseline needs a caveat the authors state themselves. Participants were English-speaking US adults on the Prolific research platform who each answered 10 questions from one task type (they were paid $2 per session whatever their score), and the authors call them a non-expert baseline. They suggest future comparisons with architecture students and practitioners. A superintendent or architect who reads elevations every day would be a different reference point.
Questions no screening model solved held the other models to 28.65%
The authors also checked whether screening with four models had made the benchmark hard only for those four. They grouped questions by how many screening models got them right and scored the models outside the screening set on each group.
Accuracy of the models outside the screening set, percent, by how many of the four screening models answered each question correctly. Source: Parikh and Garcia del Castillo y Lopez, ARCH-B, arXiv.
Accuracy rose from 28.65% to 59.02% as screening difficulty fell, which the authors take as evidence that the hard questions are hard in general. Per question, average model accuracy was 53.36%, with a range from 4% to 96%. No question was answered correctly by every model, and none was missed by all of them.
Check a photo tool’s location match rate in both directions
So the paper’s split by direction is worth copying. If you pilot a tool that ties photos to drawings, such as a reality capture product or a photo log that places shots on sheets, track the location match rate. It is the share of a fixed sample of items that the tool placed on the correct sheet and in the correct room or grid bay, as checked by the project engineer who knows the job.
Measure it both ways, since the models in ARCH-B scored differently in each direction:
Photo to plan
Take field photos the tool has already placed and check each placement against the sheet.
Plan to photo
Pick rooms or grid bays on a sheet, ask the tool for the matching photos, and check that each returned photo shows that area.
Capture the baseline before the pilot starts. Pull a sample from last month’s photo log, record how many photos carry a correct location tag, and time how long the project engineer spends tagging a week’s photos by hand. AI takeoff tools that read plan PDFs, such as the one Quotr raised a seed round for, can be checked the same way by comparing a sample of extracted quantities to an estimator’s count.
A first step on one job is to pick a single floor and one drawing sheet, and have the project engineer mark the correct location for each photo taken there since the last pay application.
ARCH-B is free, and the runner is under the MIT License
ARCH-B is free. The question set, a JSON schema and the evaluation runner are posted as ancillary files with the arXiv paper, and the README and runner are under the MIT License. The question file links to each image at its original public address with a SHA-256 hash, a fingerprint the runner uses to confirm the image has not changed, and the authors do not redistribute the images themselves.
The runner calls each model provider’s API, the paid connection a program uses to send questions to the model, and the README notes that provider use can incur charges. The authors also host a version of the human-study quiz at archbench.codecolab.org, where anyone can try sample questions.
The data for all 11 task types and 354 questions
| Task | Questions | Best model | Best | 25-model average | People | People Average Best |
|---|---|---|---|---|---|---|
| Pick the floor plan that matches 2 to 4 photos | 23 | Gemini 3.1 Pro Preview | 95.65 | 58.78 | 34.74 | |
| Pick the photo that matches 1 to 4 floor plans | 56 | GPT-4.1 | 76.79 | 44.14 | 36.41 | |
| Find the photo from a different building | 46 | Gemini 3.1 Pro Preview | 84.78 | 41.22 | 21.61 | |
| Pick the photo that matches three others | 24 | GPT-4.1 | 100.00 | 60.17 | 32.00 | |
| Pick the interior that matches 4 exterior photos | 25 | Gemini 3.1 Pro Preview | 96.00 | 52.64 | 35.85 | |
| Pick the exterior that matches 4 interior photos | 20 | Claude Opus 4.5 | 85.00 | 50.00 | 29.70 | |
| Pick the elevation that matches the floor plans | 32 | Gemini 3.5 Flash | 90.62 | 50.25 | 36.95 | |
| Pick the floor plan that matches the elevations | 30 | Gemini 3.5 Flash | 90.00 | 57.87 | 32.25 | |
| Find the odd one out among a plan, elevation, section and photo | 37 | Claude Fable 5 | 100.00 | 64.22 | 33.93 | |
| Pick the photo that matches a plan, elevation and section | 29 | Gemini 3 Flash Preview | 96.55 | 59.86 | 42.44 | |
| Pick the exterior photo that matches 2 elevations | 32 | Gemini 3.5 Flash | 96.88 | 61.00 | 56.15 |
Download the data: tasks (CSV), leaderboard (CSV), screening (CSV), curation counts (CSV).