{"ID":23475261,"CreatedAt":"2026-09-18T01:09:05.407443952Z","UpdatedAt":"2026-09-20T18:11:56.143995915Z","DeletedAt":null,"paper_url":"https://arxiv.org/abs/2609.20325","arxiv_id":"2609.20325","title":"AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images","abstract":"Agricultural image understanding requires fine-grained recognition of plant diseases, pests, crop structures, and botanical species under complex real-world conditions. Despite recent advances in Multimodal Large Language Models (MLLMs), existing models remain limited to text-only outputs and lack pixel-level visual grounding capabilities. In this work, we introduce AgriScope, a unified pixel-grounded multimodal framework for agricultural image understanding. AgriScope jointly supports image-level, region-level, and pixel-level understanding within a unified framework, enabling tasks such as grounded caption generation, referring expression segmentation, and multi-turn multimodal interaction for agricultural imagery. AgriScope integrates biologically specialized semantic representations with dense spatial grounding through biological-semantic encoding, dense spatial representations, and pixel decoding. To support large-scale grounded learning, we introduce AgriGround, a large-scale pixel-grounded agricultural multimodal instruction-tuning dataset containing over 500K images and 11M instruction-following samples spanning plant disease analysis, crop and weed identification, insect pest recognition, and fine-grained botanical understanding. AgriGround is constructed through a multi-stage automatic annotation pipeline that integrates multimodal caption generation, phrase-level grounding, segmentation mask generation, and task-oriented instruction synthesis to produce densely grounded supervision. Extensive experiments across multiple agricultural vision-language tasks demonstrate the effectiveness of AgriScope in pixel-grounded multimodal understanding, establishing a strong benchmark for agricultural vision-language learning and visual grounding. The dataset and code will be made publicly available at (https://github.com/boudiafA/AgriScope)","short_abstract":"Agricultural image understanding requires fine-grained recognition of plant diseases, pests, crop structures, and botanical species under complex real-world conditions. Despite recent advances in Multimodal Large Language Models (MLLMs), existing models remain limited to text-only outputs and lack pixel-level visual gr...","url_abs":"https://arxiv.org/abs/2609.20325","url_pdf":"https://arxiv.org/pdf/2609.20325v1","authors":"[\"Abderrahmene Boudiaf\",\"Mohamad Alanssari\",\"Irfan Hussain\",\"Sajid Javed\"]","published":"2026-09-17T12:59:11Z","proceeding":"cs.CV","tasks":"[\"cs.CV\"]","methods":"[\"Large Language Model\",\"Language Model\"]","has_code":false,"code_links":[{"ID":639801,"CreatedAt":"2026-09-18T01:09:05.407443952Z","UpdatedAt":"2026-09-18T01:09:05.407443952Z","DeletedAt":null,"paper_id":23475261,"paper_url":"https://arxiv.org/abs/2609.20325","paper_title":"AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images","repo_url":"https://github.com/boudiafA/AgriScope","is_official":false,"mentioned_in_paper":false,"mentioned_in_github":true,"github_stars":0}]}
