[1]

Yu X, Zhao Y, Gao Y, Yuan X, Xiong S. 2022. Benchmark platform for ultra-fine-grained visual categorization beyond human performance. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), October 10–17, 2021, Montreal, QC, Canada. USA: IEEE. pp. 10265–10275 doi: 10.1109/ICCV48922.2021.01012

[2]

Savary S, Willocquet L, Pethybridge SJ, Esker P, McRoberts N, et al. 2019. The global burden of pathogens and pests on major food crops. Nature Ecology & Evolution 3(3):430−439

doi: 10.1038/s41559-018-0793-y
[3]

Gerhards R, Andújar Sanchez D, Hamouz P, Peteinatos GG, Christensen S, et al. 2022. Advances in site-specific weed management in agriculture — a review. Weed Research 62(2):123−133

doi: 10.1111/wre.12526
[4]

Jin X, Liu T, McCullough PE, Chen Y, Yu J. 2023. Evaluation of convolutional neural networks for herbicide susceptibility-based weed detection in turf. Frontiers in Plant Science 14:1096802

doi: 10.3389/fpls.2023.1096802
[5]

Horowitz AR, Ghanim M, Roditakis E, Nauen R, Ishaaya I. 2020. Insecticide resistance and its management in Bemisia tabaci species. Journal of Pest Science 93(3):893−910

doi: 10.1007/s10340-020-01210-0
[6]

Behere GT, Tay WT, Russell DA, Batterham P. 2008. Molecular markers to discriminate among four pest species of Helicoverpa (Lepidoptera: Noctuidae). Bulletin of Entomological Research 98(6):599−603

doi: 10.1017/S0007485308005956
[7]

Kim KC, Byrne LB. 2006. Biodiversity loss and the taxonomic bottleneck: emerging biodiversity science. Ecological Research 21(6):794

doi: 10.1007/s11284-006-0035-7
[8]

Valan M, Makonyi K, Maki A, Vondráček D, Ronquist F. 2019. Automated taxonomic identification of insects with expert-level accuracy using effective feature transfer from convolutional networks. Systematic Biology 68(6):876−895

doi: 10.1093/sysbio/syz014
[9]

Yu X, Zhao Y, Gao Y, Xiong S. 2021. MaskCOV: a random mask covariance network for ultra-fine-grained visual categorization. Pattern Recognition 119:108067

doi: 10.1016/j.patcog.2021.108067
[10]

Alayrac JB, Donahue J, Luc P, Miech A, Barr I, et al. 2022. Flamingo: a visual language model for few-shot learning. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS 2022), 2022, New Orleans, Louisiana, USA. NeurIPS. https://openreview.net/attachment?id=EbMuimAbPbs&name=supplementary_material

[11]

Liu H, Li C, Wu Q, Lee YJ. 2023. Visual instruction tuning. In Advances in Neural Information Processing Systems 36. December 10-16, 2023. New Orleans, Louisiana, USA. NeurIPS. pp. 34892–34916 doi: 10.52202/075280-1516

[12]

Liu Y, Duan H, Zhang Y, Li B, Zhang S, et al. 2025. MMBench: is your multi-modal model an all-around player? In Computer Vision – ECCV 2024. Cham: Springer. pp. 216–233 doi: 10.1007/978-3-031-72658-3_13

[13]

Li B, Ge Y, Ge Y, Wang G, Wang R, et al. 2024. SEED-bench: benchmarking multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 16–22, 2024, Seattle, WA, USA. USA: IEEE. pp. 13299–13308 doi: 10.1109/CVPR52733.2024.01263

[14]

Yue X, Ni Y, Zheng T, Zhang K, Liu R, et al. 2024. MMMU: a massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 16–22, 2024, Seattle, WA, USA. USA: IEEE. pp. 9556–9567 doi: 10.1109/CVPR52733.2024.00913

[15]

Maji S, Kannala J, Rahtu E, Blaschko M, Vedaldi A. 2013. Fine-grained visual classification of aircraft. Technical Report. University of Oxford, Oxford, UK. arXiv:1306.5151. 10.48550/arXiv.1306.5151

[16]

Krause J, Stark M, Jia D, Li FF. 2013. 3D object representations for fine-grained categorization. In 2013 IEEE International Conference on Computer Vision Workshops, December 2–8, 2013, Sydney, NSW, Australia. USA: IEEE. pp. 554–561 doi: 10.1109/ICCVW.2013.77

[17]

Islam A, Biswas MR, Zaghouani W, Belhaouari SB, Shah Z. 2023. Pushing boundaries: Exploring zero shot object classification with large multimodal models. In 2023 Tenth International Conference on Social Networks Analysis, Management and Security (SNAMS), November 21–24, 2023, Abu Dhabi, United Arab Emirates. USA: IEEE. pp. 1–5 doi: 10.1109/SNAMS60348.2023.10375440

[18]

Wu W, Yao H, Zhang M, Song Y, Ouyang W, et al. 2024. GPT4Vis: what can GPT-4 do for zero-shot visual recognition? arXiv Preprint

doi: 10.48550/arXiv.2311.15732
[19]

Tong S, Liu Z, Zhai Y, Ma Y, LeCun Y, et al. 2024. Eyes wide shut? exploring the visual shortcomings of multimodal LLMs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 16–22, 2024, Seattle, WA, USA. USA: IEEE. pp. 9568–9578 doi: 10.1109/CVPR52733.2024.00914

[20]

Kimi Team, Bai T, Bai Y, Bao Y, et al. 2026. Kimi K2.5: visual agentic intelligence. arXiv Preprint

doi: 10.48550/arXiv.2602.02276
[21]

OpenAI. ChatGPT. OpenAI, 2025. https://chatgpt.com

[22]

Google. Gemini. Google, 2025. https://gemini.google.com

[23]

Qwen Team. Qwen3.5: towards native multimodal agents. February 2026. https://qwen.ai/blog?id=qwen3.5

[24]

Wah C, Branson S, Welinder P, Perona P, Belongie S. 2011. The Caltech-UCSD Birds-200-2011 Dataset. Technical Report CNS-TR-2011-001. Pasadena, CA: California Institute of Technology. www.vision.caltech.edu/visipedia/CUB-200-2011.html

[25]

Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, et al. 2021. An image is worth 16×16 words: transformers for image recognition at scale. 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3–7, 2021. OpenReview.net. https://openreview.net/forum?id=YicbFdNTTy

[26]

Demidov D, Sharif M, Abdurahimov A, Cholakkal H, Khan F. 2023. Salient mask-guided vision transformer for fine-grained classification. In Proceedings of the 18th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, February 19–21, 2023. Lisbon, Portugal. Portugal: Science and Technology Publications. pp. 27–38 doi: 10.5220/0011611100003417

[27]

Zhang K, Cao J, Li HS, Cai Q. 2023. Cross-layer feature fusion vision transformer for fine-grained visual classification. In 2023 5th International Conference on Data-driven Optimization of Complex Systems (DOCS), September 22–24, 2023, Tianjin, China. USA: IEEE. pp. 1–6 10.1109/DOCS60977.2023.10294512

[28]

e K, Zhang X, Ren S, Sun J. 2016. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 27-30, 2016, Las Vegas, NV, USA. USA: IEEE. pp. 770–778 doi: 10.1109/CVPR.2016.90

[29]

Krizhevsky A, Sutskever I, Hinton GE. 2012. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems 25 (NIPS 2012), December 3–6, 2012, Lake Tahoe, Nevada, USA. Curran Associates, Inc. pp. 1097–1105 https://proceedings.neurips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html

[30]

Pan Z, Yu X, Zhang M, Gao Y. 2021. Mask-guided feature extraction and augmentation for ultra-fine-grained visual categorization. arXiv Preprint

doi: 10.48550/arXiv.2109.07755
[31]

Yu X, Wang J, Gao Y. 2023. CLE-ViT: contrastive learning encoded transformer for ultra-fine-grained visual categorization. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, August 19–25, 2023, Macau, SAR China. International Joint Conferences on Artificial Intelligence Organization. pp. 4531–4539 doi: 10.24963/ijcai.2023/504

[32]

Zhang P, Yu X, Gu M, Wu Y, Gao Y, et al. 2025. Revisiting continual ultra-fine-grained visual recognition with pre-trained models. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. ACM. pp. 9474–9482 doi: 10.24963/ijcai.2025/1053

[33]

Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, et al. 2021. Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning (ICML 2021), July 18–24, 2021. Proceedings of Machine Learning Research. Vol. 139. PMLR. pp. 8748–8763 https://proceedings.mlr.press/v139/radford21a.html

[34]

Yin S, Fu C, Zhao S, Li K, et al. 2024. A survey on multimodal large language models. National Science Review 11(12):nwae403

doi: 10.1093/nsr/nwae403
[35]

Li J, Li D, Savarese S, Hoi S. 2023. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning 202:814

[36]

Dai W, Li J, Li D, Tiong A, Zhao J, et al. 2023. InstructBLIP: towards general-purpose vision-language models with instruction tuning. In Advances in Neural Information Processing Systems 36. December 10–16, 2023. New Orleans, Louisiana, USA. NeurIPS. pp. 49250–49267 doi: 10.52202/075280-2142

[37]

OpenAI, Achiam J, Adler S, Agarwal S, Ahmad L, et al. 2024. GPT-4 technical report. arXiv Preprint

doi: 10.48550/arXiv.2303.08774
[38]

Gemini Team. 2024. Gemini 1.5: unlocking multimodal understanding across millions of tokens of context. arXiv Preprint

doi: 10.48550/arXiv.2403.05530
[39]

Anthropic. 2024. The Claude 3 Model Family: Opus, Sonnet, Haiku. Technical Report. www.anthropic.com/news/claude-3-family

[40]

Li C, Wong C, Zhang S, Usuyama N, Liu H, et al. 2023. LLaVA-med: training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36 (NeurIPS 2023 Datasets and Benchmarks Track), December 10–16, 2023, New Orleans, LA, USA. Curran Associates, Inc. pp. 28541–28564 https://proceedings.neurips.cc/paper_files/paper/2023/hash/5abcdf8ecdcacba028c6662789194572-Abstract-Datasets_and_Benchmarks.html

[41]

Liu F, Dai W, Zhang C, Zhu J, Yao L, et al. 2025. Co-LLaVA: efficient remote sensing visual question answering via model collaboration. Remote Sensing 17(3):466

doi: 10.3390/rs17030466
[42]

Li Y, Du Y, Zhou K, Wang J, Zhao X, et al. 2023. Evaluating object hallucination in large vision-language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, Singapore. Stroudsburg, PA, USA: ACL. pp. 292–305 doi: 10.18653/v1/2023.emnlp-main.20

[43]

Jiang Y, Irvin J, Wang JH, Chaudhry MA, Chen JH, et al. 2024. Many-shot in-context learning in multimodal foundation models. arXiv Preprint

doi: 10.48550/arXiv.2405.09798
[44]

Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (NeurIPS 2022), November 28 – December 9, 2022, New Orleans, LA, USA. Curran Associates, Inc. pp. 24824–24837 https://proceedings.neurips.cc/paper_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html

[45]

Menon S, Vondrick C. 2023. Visual classification via description from large language models. The Eleventh International Conference on Learning Representations (ICLR 2023), May 1–5, 2023, Kigali, Rwanda. OpenReview.net. https://openreview.net/forum?id=41GFzHFc4zx

[46]

Chia YK, Chen G, Tuan LA, Poria S, Bing L. 2023. Contrastive chain-of-thought prompting. arXiv Preprint

doi: 10.48550/arXiv.2311.09277
[47]

Jiang X, Tang H, Gao J, Du X, He S, et al. 2024. Delving into multimodal prompting for fine-grained visual classification. Proceedings of the AAAI Conference on Artificial Intelligence 38(3):2570−2578

doi: 10.1609/aaai.v38i3.28034
[48]

Zhang Z, Zhang A, Li M, Zhao H, Karypis G, et al. 2024. Multimodal chain-of-thought reasoning in language models. Transactions on Machine Learning Research 2024. https://openreview.net/forum?id=y1pPWFVfvR

[49]

Pratt S, Covert I, Liu R, Farhadi A. 2023. What does a platypus look like? generating customized prompts for zero-shot image classification. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), October 1–6, 2023, Paris, France. pp. 15645–15655 doi: 10.1109/ICCV51070.2023.01438

[50]

Liu M, Roy S, Li W, Zhong Z, Sebe N, et al. 2024. Democratizing fine-grained visual recognition with large language models. The Twelfth International Conference on Learning Representations (ICLR 2024), May 7–11, 2024, Vienna, Austria. OpenReview.net. https://openreview.net/forum?id=c7DND1iIgb

[51]

Pan J, Zhong R, Xia F, Huang J, Zhu L, et al. 2025. ChatLeafDisease: a chain-of-thought prompting approach for crop disease classification using large language models. Plant Phenomics 7(3):100094

doi: 10.1016/j.plaphe.2025.100094
[52]

Lu Y, Bartolo M, Moore A, Riedel S, Stenetorp P. 2021. Fantastically ordered prompts and where to find them: overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 2022, Dublin. Ireland Association for Computational Linguistics. pp. 8086−8098 doi: 10.18653/v1/2022.acl-long.556

[53]

Zhang Y, Zhou K, Liu Z. 2023. What makes good examples for visual in-context learning? In Advances in Neural Information Processing Systems 36. December 10–16, 2023. New Orleans, Louisiana, USA. NeurIPS. pp. 17773–17794 doi: 10.52202/075280-0780

[54]

Min S, Lyu X, Holtzman A, Artetxe M, Lewis M, et al. 2022. Rethinking the role of demonstrations: what makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, Abu Dhabi, United Arab Emirates. Stroudsburg, PA, USA: ACL. pp. 11048–11064 doi: 10.18653/v1/2022.emnlp-main.759

[55]

Ferber D, Wölflein G, Wiest IC, Ligero M, Sainath S, et al. 2024. In-context learning enables multimodal large language models to classify cancer pathology images. Nature Communications 15(1):10104

doi: 10.1038/s41467-024-51465-9
[56]

Yao L, Yang Y. 2026. Large language models are contrastive reasoners. Expert Systems with Applications 301:130407

doi: 10.1016/j.eswa.2025.130407
[57]

Lu X, Yu X, Wang K, Wang Y, Wang P, et al. 2023. SupCon-ViT: supervised contrastive learning for ultra-fine-grained visual categorization. In 2023 International Conference on Digital Image Computing: Techniques and Applications (DICTA), 2023, Port Macquarie, Australia. USA: IEEE. pp. 281–288 doi: 10.1109/DICTA60407.2023.00046

[58]

Chiang WL, Zheng L, Sheng Y, Angelopoulos AN, Li T, et al. 2024. Chatbot arena: an open platform for evaluating LLMs by human preference. Proceedings of the 41st International Conference on Machine Learning (ICML 2024), July 21–27, 2024, Vienna, Austria. PMLR. Vol. 235. pp. 8359–8388 https://proceedings.mlr.press/v235/chiang24b.html

[59]

Simonyan K, Zisserman A. 2015. Very deep convolutional networks for large-scale image recognition. International Conference on Learning Representations (ICLR 2015), May 7–9, 2015, San Diego, CA, USA. arXiv:1409.1556. doi: 10.48550/arXiv.1409.1556

[60]

Chen Y, Bai Y, Zhang W, Mei T. 2019. Destruction and construction learning for fine-grained image recognition. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, Long Beach, CA, USA. USA: IEEE. pp. 5152–5161 doi: 10.1109/CVPR.2019.00530