← PortMind results

How we test

Earlier studies

These studies informed the next benchmark. Each used its own setup and one human reference. Their scores are kept separate from the upcoming versioned releases.

Guided inspection

July 28, 2026 · 23 scored task rows · One human reference

20 unique images plus 4 repeat tasks. One undecidable task excluded, leaving 23 binary rows and 19 unique decidable images. Container labels include 8 positive and 15 negative rows. Repeats are not independent evidence.

Explore this study →

Image-labeling pilot

Published pilot summary · 23 binary task rows · One human reference

20 unique images and 4 hidden repeats. One undecidable task excluded. These are the published pilot scores; original response traces have not been verified. The setup differs from the guided inspection run.

Explore this study →

Disagreement review

72 selected disagreement images · One human relabel

GPT-5.5 and Grok 4.5 used different evidence workflows. Agreement excludes abstentions. These selected disagreement images are not a representative sample of port traffic.

Explore this study →

Tool-use run

July 28, 2026 · 23 binary task rows · Mixed inspection workflows

20 unique images plus 4 repeats. GPT-5.5 used strict autonomous inspection; other models used guided steps. GPT-5.5 completed 14 of 24 tasks. Label agreement alone does not capture protocol errors, runtime failures or stopped turns.

Explore this study →

Versioned benchmarks

Human labels are the answer key. Model labels are the answers being tested.

Choose the images

Define a port, task and camera/date split. Check image provenance and duplicates before preparing a review packet.

Independent review

At least two reviewers submit their own judgments, with model answers hidden. Resolve disagreements and retain uncertainty when the image cannot support a clear answer.

Freeze the version

The reference answers, image manifest and scoring protocol are fixed. A correction creates a new version, not an overwritten score.

Run and score

Models answer the same questions under recorded conditions. The platform records the model, prompt, permitted tools and failures, then scores the answers against the frozen reference.

Approve publication

Publish reviewed results with their benchmark version and sample counts. Private images and raw responses stay outside the public report.

What Montréal v1.0 tests

Truck presence and container/chassis attachment in still images. Answer Yes, No or Unsure separately. One qualifying truck anywhere in the image is enough for Yes. This does not measure vehicle counts, movement or waiting time.

Read both kinds of mistake

Found means the model answered Yes where the human reference says Yes. A false alarm means it answered Yes where the reference says No. We also report unanswered, uncertain and failed tasks. Uncertain reference labels are excluded from binary scoring and must remain visible in the evaluation record.

One reference, many model runs

You do not label the dataset again for every model. We reuse the frozen reference when testing additional models. New ports or changed images, labels or rules have their own benchmark versions.

Open the labeling guide →