微信内可能无法直接打开本站。请点右上角 ··· → 在浏览器打开,或复制链接。
CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning
RSS 官方收录 · 可信分层展示
关键摘要
arXiv:2609.…
- 03526v1 Announce Type: new Abstract: Multimodal language models achiev…
- To probe this distinction, we introduce CulturalMenuBench, a benchmark…
- Evaluating 12 models exposes a substantial knowledge-application gap: …
摘要引擎:抽取
正文提要
arXiv:2609.03526v1 Announce Type: new Abstract: Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attribution. Evaluating 12 models exposes a substantial knowledge-application gap: models exceeding 94% on standard multiple-choice tasks drop to at most 56% when attributing dishes to Chinese regional cuisines, despite an identical four-way format. Diagnostic analyses explain why: error patterns are consistent with random guessing, accuracy tracks visual distinctiveness rather than cultural structure, and models classify cuisines more accurately from dish names alone than from images (+7-18 points). The knowledge is thus present but cannot be activated through visual input. An ablation confirms these tasks genuinely require procedural evidence: removing sequential cooking images selectively degrades process-grounded tasks while others remain stable. Overall, CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context. Code and data are publicly available.