FieldWorkArena
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
886 tasks over on-site images, videos and documents captured in real factories, warehouses and retail stores (711 perception, 121 decision-making, 54 combination tasks), written from interviews with site workers and managers. An agent must extract information, detect safety or procedural violations and produce reports; answers are scored against ground truth with a weighted mix of exact and near-match scoring on a 0-1 scale (reported here as percent). It tests multimodal agents on physical-world field operations rather than digital environments.
- Released
- 2025-05
- Maintainer
- Fujitsu Research (with Carnegie Mellon University)
- Status
- active
- Contamination
- low
- Metric
- accuracy rate (total) (percent, ↑)
- Tasks
- 886
- Domains
- multimodal general-assistant safety reasoning
- human
- 74% human evaluators on a random sample of perception tasks (paper, score 0.74)
Notes. Factory dataset (V1.0) was released on the Fujitsu site in February 2025, before the May 2025 arXiv paper. Dataset access requires an application form (gated HuggingFace), and the site lists its leaderboard as 'coming soon', so the only public numbers are the authors' own MLLM evaluations in the paper (v4, June 2026); task counts and the human figure are from that revision. The paper's 0-1 scores are multiplied by 100 here.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GPT-5.2 (2025-12-11) | OpenAI | 52% | paper | Table 2 total score 0.52 in arXiv v4; best model in the paper. | |
| Gemini 2.5 Pro | Google DeepMind | 46% | paper | Table 2 total score 0.46 in arXiv v4. | |
| GPT-4o (2024-08-06) | OpenAI | 35% | paper | Table 2 total score 0.35 in arXiv v4 (7 Jun 2026); models run as agents with default parameters. |