FieldWorkArena

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

886 tasks over on-site images, videos and documents captured in real factories, warehouses and retail stores (711 perception, 121 decision-making, 54 combination tasks), written from interviews with site workers and managers. An agent must extract information, detect safety or procedural violations and produce reports; answers are scored against ground truth with a weighted mix of exact and near-match scoring on a 0-1 scale (reported here as percent). It tests multimodal agents on physical-world field operations rather than digital environments.

52% · human 74%
GPT-5.2 (2025-12-11)
paper
2026-06-07 → 2026-06-07: 35% → 52%
Released
2025-05
Maintainer
Fujitsu Research (with Carnegie Mellon University)
Status
active
Contamination
low
Metric
accuracy rate (total) (percent, ↑)
Tasks
886
Domains
multimodal general-assistant safety reasoning
human
74% human evaluators on a random sample of perception tasks (paper, score 0.74)

Notes. Factory dataset (V1.0) was released on the Fujitsu site in February 2025, before the May 2025 arXiv paper. Dataset access requires an application form (gated HuggingFace), and the site lists its leaderboard as 'coming soon', so the only public numbers are the authors' own MLLM evaluations in the paper (v4, June 2026); task counts and the human figure are from that revision. The paper's 0-1 scores are multiplied by 100 here.

Full ledger

SystemDeveloperScoreDateSourceConditions
GPT-5.2 (2025-12-11)OpenAI52%paper
Table 2 total score 0.52 in arXiv v4; best model in the paper.
Gemini 2.5 ProGoogle DeepMind46%paper
Table 2 total score 0.46 in arXiv v4.
GPT-4o (2024-08-06)OpenAI35%paper
Table 2 total score 0.35 in arXiv v4 (7 Jun 2026); models run as agents with default parameters.