FieldWorkArena
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
886 个任务,基于真实工厂、仓库和零售店现场采集的图像、视频和文档(711 个感知任务、121 个决策任务、54 个组合任务),依据对现场工人和管理者的访谈编写。智能体需提取信息、检测安全或流程违规并生成报告;答案对照真值,以精确匹配与近似匹配的加权组合按 0-1 评分(此处以百分比报告)。它测试多模态智能体在物理世界现场作业而非数字环境中的表现。
- 发布
- 2025-05
- 维护者
- Fujitsu Research (with Carnegie Mellon University)
- 状态
- active
- 污染风险
- low
- 指标
- accuracy rate (total) (percent, ↑)
- 题量
- 886
- 领域
- multimodal general-assistant safety reasoning
- 人类
- 74% human evaluators on a random sample of perception tasks (paper, score 0.74)
备注. Factory dataset (V1.0) was released on the Fujitsu site in February 2025, before the May 2025 arXiv paper. Dataset access requires an application form (gated HuggingFace), and the site lists its leaderboard as 'coming soon', so the only public numbers are the authors' own MLLM evaluations in the paper (v4, June 2026); task counts and the human figure are from that revision. The paper's 0-1 scores are multiplied by 100 here.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-5.2 (2025-12-11) | OpenAI | 52% | 论文 | Table 2 total score 0.52 in arXiv v4; best model in the paper. | |
| Gemini 2.5 Pro | Google DeepMind | 46% | 论文 | Table 2 total score 0.46 in arXiv v4. | |
| GPT-4o (2024-08-06) | OpenAI | 35% | 论文 | Table 2 total score 0.35 in arXiv v4 (7 Jun 2026); models run as agents with default parameters. |