Benchmarking Research and AI AgriBench

AgMMU: A Comprehensive Agricultural Multimodal Understanding and Reasoning Benchmark

Evaluation and training data sets for evaluating and developing vision-language models (VLMs) to produce factually accurate answers for the knowledge-intensive agriculture domain. The AgMMU evaluation data set includes 3390 open-ended questions for factual question answering (OEQs) and 5793 multiple-choice questions (MCQs).

Published in the NeurIPS Data Sets and Benchmarks Track, 2026.

MIRAGE: A Benchmark for Multimodal Information‑Seeking and Reasoning in Agricultural Expert‑Guided Conversations

A new multimodal benchmark designed to evaluate vision-language models in realistic expert consultation settings. MIRAGE incorporates natural user queries, expert responses, and images derived from real interactions between real users and domain experts. The benchmark includes both single-turn (MIRAGE-MMST) and multi-turn (MIRAGE-MMMT) tasks, assessing not just accuracy but also the ability to simulate expert conversational decisions, such as whether to clarify or respond.

Published in the NeurIPS Data Sets and Benchmarks Track, 2026.

AI AgriBench Benchmarking Consortium

The benchmarking research in the CropWizard project led to the formation of the AI AgriBench Benchmarking Consortium in 2025, which is led by CropWizard PI Vikram Adve, and the Center for Digital Agriculture. The consortium was formed in close collaboration with the founding members, including Bayer Crop Sciences, John Deere, Microsoft, the Extension Foundation, and Kissan AI. The broad goal of the consortium is to provide farmers, policymakers, and the public with trustworthy mechanisms for evaluating and confidently using the new generation of AI-driven agricultural advisory tools and services, as well as public chatbots and Large Language Models.

CropWizard: Generative AI for Agriculture
Email: vadve@illinois.edu