How we test
Earlier studies
These studies informed the next benchmark. Each used its own setup and one human reference. Their scores are kept separate from the upcoming versioned releases.
Guided inspection
July 28, 2026 · 23 scored task rows · One human reference
20 unique images plus 4 repeat tasks. One undecidable task excluded, leaving 23 binary rows and 19 unique decidable images. Container labels include 8 positive and 15 negative rows. Repeats are not independent evidence.
Explore this study →Image-labeling pilot
Published pilot summary · 23 binary task rows · One human reference
20 unique images and 4 hidden repeats. One undecidable task excluded. These are the published pilot scores; original response traces have not been verified. The setup differs from the guided inspection run.
Explore this study →Disagreement review
72 selected disagreement images · One human relabel
GPT-5.5 and Grok 4.5 used different evidence workflows. Agreement excludes abstentions. These selected disagreement images are not a representative sample of port traffic.
Explore this study →Tool-use run
July 28, 2026 · 23 binary task rows · Mixed inspection workflows
20 unique images plus 4 repeats. GPT-5.5 used strict autonomous inspection; other models used guided steps. GPT-5.5 completed 14 of 24 tasks. Label agreement alone does not capture protocol errors, runtime failures or stopped turns.
Explore this study →Versioned benchmarks
Human labels are the answer key. Model labels are the answers being tested.
Choose the images
Define a port, task and camera/date split. Check image provenance and duplicates before preparing a review packet.
Independent review
At least two reviewers submit their own judgments, with model answers hidden. Resolve disagreements and retain uncertainty when the image cannot support a clear answer.
Freeze the version
The reference answers, image manifest and scoring protocol are fixed. A correction creates a new version, not an overwritten score.
Run and score
Models answer the same questions under recorded conditions. The platform records the model, prompt, permitted tools and failures, then scores the answers against the frozen reference.
Approve publication
Publish reviewed results with their benchmark version and sample counts. Private images and raw responses stay outside the public report.
What Montréal v1.0 tests
Truck presence and container/chassis attachment in still images. Answer Yes, No or Unsure separately. One qualifying truck anywhere in the image is enough for Yes. This does not measure vehicle counts, movement or waiting time.
Read both kinds of mistake
Found means the model answered Yes where the human reference says Yes. A false alarm means it answered Yes where the reference says No. We also report unanswered, uncertain and failed tasks. Uncertain reference labels are excluded from binary scoring and must remain visible in the evaluation record.
One reference, many model runs
You do not label the dataset again for every model. We reuse the frozen reference when testing additional models. New ports or changed images, labels or rules have their own benchmark versions.