How well do models understand a port?
PortMind evaluates vision models on port camera images. We ask models to identify trucks and container attachments, then compare their answers with human labels.
Earlier studies
No score available: Mistral Small 3.1, Moondream 3.1.
July 28, 2026 · 23 binary task rows · Mixed inspection workflows
Study setup and limitations
20 unique images plus 4 repeats. GPT-5.5 used strict autonomous inspection; other models used guided steps. GPT-5.5 completed 14 of 24 tasks. Label agreement alone does not capture protocol errors, runtime failures or stopped turns.
Read about this study →Misses and false alarms
In guided inspection, Llama 3.2 Vision found all 8 positive tasks but also flagged 14 of 15 negative tasks.
Select a model to open its results. 8 positive and 15 negative task rows, including repeats.
Tasks
Truck presence
Can you see at least one truck? Pickups and larger road trucks count.
Container attachment
Is any truck carrying a shipping container or towing an empty chassis? A container standing nearby does not count.
Each question has three answers: Yes, No or Unsure. The benchmark tests what is visible in an image, not traffic flow or waiting time.
See labeling examples →How we evaluate models
Label the images
Reviewers independently answer the same questions, with model answers hidden.
Review the answers
Reviewers resolve disagreements before we finalize the reference labels.
Run the models
Each model receives the same images and questions. We record its answers and any failed requests.
Report the results
We compare model answers with the reference labels and report agreement, misses and false alarms.
Contact
For questions about the benchmark or evaluating a model, get in touch.