| [1] |
Yu X, Zhao Y, Gao Y, Yuan X, Xiong S. 2022. Benchmark platform for ultra-fine-grained visual categorization beyond human performance. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), October 10–17, 2021, Montreal, QC, Canada. USA: IEEE. pp. 10265–10275 doi: 10.1109/ICCV48922.2021.01012 |
| [2] |
Savary S, Willocquet L, Pethybridge SJ, Esker P, McRoberts N, et al. 2019. The global burden of pathogens and pests on major food crops. |
| [3] |
Gerhards R, Andújar Sanchez D, Hamouz P, Peteinatos GG, Christensen S, et al. 2022. Advances in site-specific weed management in agriculture — a review. |
| [4] |
Jin X, Liu T, McCullough PE, Chen Y, Yu J. 2023. Evaluation of convolutional neural networks for herbicide susceptibility-based weed detection in turf. |
| [5] |
Horowitz AR, Ghanim M, Roditakis E, Nauen R, Ishaaya I. 2020. Insecticide resistance and its management in Bemisia tabaci species. |
| [6] |
Behere GT, Tay WT, Russell DA, Batterham P. 2008. Molecular markers to discriminate among four pest species of Helicoverpa (Lepidoptera: Noctuidae). |
| [7] |
Kim KC, Byrne LB. 2006. Biodiversity loss and the taxonomic bottleneck: emerging biodiversity science. |
| [8] |
Valan M, Makonyi K, Maki A, Vondráček D, Ronquist F. 2019. Automated taxonomic identification of insects with expert-level accuracy using effective feature transfer from convolutional networks. |
| [9] |
Yu X, Zhao Y, Gao Y, Xiong S. 2021. MaskCOV: a random mask covariance network for ultra-fine-grained visual categorization. |
| [10] |
Alayrac JB, Donahue J, Luc P, Miech A, Barr I, et al. 2022. Flamingo: a visual language model for few-shot learning. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS 2022), 2022, New Orleans, Louisiana, USA. NeurIPS. https://openreview.net/attachment?id=EbMuimAbPbs&name=supplementary_material |
| [11] |
Liu H, Li C, Wu Q, Lee YJ. 2023. Visual instruction tuning. In Advances in Neural Information Processing Systems 36. December 10-16, 2023. New Orleans, Louisiana, USA. NeurIPS. pp. 34892–34916 doi: 10.52202/075280-1516 |
| [12] |
Liu Y, Duan H, Zhang Y, Li B, Zhang S, et al. 2025. MMBench: is your multi-modal model an all-around player? In Computer Vision – ECCV 2024. Cham: Springer. pp. 216–233 doi: 10.1007/978-3-031-72658-3_13 |
| [13] |
Li B, Ge Y, Ge Y, Wang G, Wang R, et al. 2024. SEED-bench: benchmarking multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 16–22, 2024, Seattle, WA, USA. USA: IEEE. pp. 13299–13308 doi: 10.1109/CVPR52733.2024.01263 |
| [14] |
Yue X, Ni Y, Zheng T, Zhang K, Liu R, et al. 2024. MMMU: a massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 16–22, 2024, Seattle, WA, USA. USA: IEEE. pp. 9556–9567 doi: 10.1109/CVPR52733.2024.00913 |
| [15] |
Maji S, Kannala J, Rahtu E, Blaschko M, Vedaldi A. 2013. Fine-grained visual classification of aircraft. Technical Report. University of Oxford, Oxford, UK. arXiv:1306.5151. 10.48550/arXiv.1306.5151 |
| [16] |
Krause J, Stark M, Jia D, Li FF. 2013. 3D object representations for fine-grained categorization. In 2013 IEEE International Conference on Computer Vision Workshops, December 2–8, 2013, Sydney, NSW, Australia. USA: IEEE. pp. 554–561 doi: 10.1109/ICCVW.2013.77 |
| [17] |
Islam A, Biswas MR, Zaghouani W, Belhaouari SB, Shah Z. 2023. Pushing boundaries: Exploring zero shot object classification with large multimodal models. In 2023 Tenth International Conference on Social Networks Analysis, Management and Security (SNAMS), November 21–24, 2023, Abu Dhabi, United Arab Emirates. USA: IEEE. pp. 1–5 doi: 10.1109/SNAMS60348.2023.10375440 |
| [18] |
Wu W, Yao H, Zhang M, Song Y, Ouyang W, et al. 2024. GPT4Vis: what can GPT-4 do for zero-shot visual recognition? |
| [19] |
Tong S, Liu Z, Zhai Y, Ma Y, LeCun Y, et al. 2024. Eyes wide shut? exploring the visual shortcomings of multimodal LLMs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 16–22, 2024, Seattle, WA, USA. USA: IEEE. pp. 9568–9578 doi: 10.1109/CVPR52733.2024.00914 |
| [20] |
Kimi Team, Bai T, Bai Y, Bao Y, et al. 2026. Kimi K2.5: visual agentic intelligence. |
| [21] |
OpenAI. ChatGPT. OpenAI, 2025. https://chatgpt.com |
| [22] |
Google. Gemini. Google, 2025. https://gemini.google.com |
| [23] |
Qwen Team. Qwen3.5: towards native multimodal agents. February 2026. https://qwen.ai/blog?id=qwen3.5 |
| [24] |
Wah C, Branson S, Welinder P, Perona P, Belongie S. 2011. The Caltech-UCSD Birds-200-2011 Dataset. Technical Report CNS-TR-2011-001. Pasadena, CA: California Institute of Technology. www.vision.caltech.edu/visipedia/CUB-200-2011.html |
| [25] |
Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, et al. 2021. An image is worth 16×16 words: transformers for image recognition at scale. 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3–7, 2021. OpenReview.net. https://openreview.net/forum?id=YicbFdNTTy |
| [26] |
Demidov D, Sharif M, Abdurahimov A, Cholakkal H, Khan F. 2023. Salient mask-guided vision transformer for fine-grained classification. In Proceedings of the 18th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, February 19–21, 2023. Lisbon, Portugal. Portugal: Science and Technology Publications. pp. 27–38 doi: 10.5220/0011611100003417 |
| [27] |
Zhang K, Cao J, Li HS, Cai Q. 2023. Cross-layer feature fusion vision transformer for fine-grained visual classification. In 2023 5th International Conference on Data-driven Optimization of Complex Systems (DOCS), September 22–24, 2023, Tianjin, China. USA: IEEE. pp. 1–6 10.1109/DOCS60977.2023.10294512 |
| [28] |
e K, Zhang X, Ren S, Sun J. 2016. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 27-30, 2016, Las Vegas, NV, USA. USA: IEEE. pp. 770–778 doi: 10.1109/CVPR.2016.90 |
| [29] |
Krizhevsky A, Sutskever I, Hinton GE. 2012. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems 25 (NIPS 2012), December 3–6, 2012, Lake Tahoe, Nevada, USA. Curran Associates, Inc. pp. 1097–1105 https://proceedings.neurips.cc/paper/2012/hash/c399862d3b9d6b76c8436e924a68c45b-Abstract.html |
| [30] |
Pan Z, Yu X, Zhang M, Gao Y. 2021. Mask-guided feature extraction and augmentation for ultra-fine-grained visual categorization. |
| [31] |
Yu X, Wang J, Gao Y. 2023. CLE-ViT: contrastive learning encoded transformer for ultra-fine-grained visual categorization. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, August 19–25, 2023, Macau, SAR China. International Joint Conferences on Artificial Intelligence Organization. pp. 4531–4539 doi: 10.24963/ijcai.2023/504 |
| [32] |
Zhang P, Yu X, Gu M, Wu Y, Gao Y, et al. 2025. Revisiting continual ultra-fine-grained visual recognition with pre-trained models. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence. ACM. pp. 9474–9482 doi: 10.24963/ijcai.2025/1053 |
| [33] |
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, et al. 2021. Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning (ICML 2021), July 18–24, 2021. Proceedings of Machine Learning Research. Vol. 139. PMLR. pp. 8748–8763 https://proceedings.mlr.press/v139/radford21a.html |
| [34] |
Yin S, Fu C, Zhao S, Li K, et al. 2024. A survey on multimodal large language models. |
| [35] |
Li J, Li D, Savarese S, Hoi S. 2023. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning 202:814 |
| [36] |
Dai W, Li J, Li D, Tiong A, Zhao J, et al. 2023. InstructBLIP: towards general-purpose vision-language models with instruction tuning. In Advances in Neural Information Processing Systems 36. December 10–16, 2023. New Orleans, Louisiana, USA. NeurIPS. pp. 49250–49267 doi: 10.52202/075280-2142 |
| [37] |
OpenAI, Achiam J, Adler S, Agarwal S, Ahmad L, et al. 2024. GPT-4 technical report. |
| [38] |
Gemini Team. 2024. Gemini 1.5: unlocking multimodal understanding across millions of tokens of context. |
| [39] |
Anthropic. 2024. The Claude 3 Model Family: Opus, Sonnet, Haiku. Technical Report. www.anthropic.com/news/claude-3-family |
| [40] |
Li C, Wong C, Zhang S, Usuyama N, Liu H, et al. 2023. LLaVA-med: training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36 (NeurIPS 2023 Datasets and Benchmarks Track), December 10–16, 2023, New Orleans, LA, USA. Curran Associates, Inc. pp. 28541–28564 https://proceedings.neurips.cc/paper_files/paper/2023/hash/5abcdf8ecdcacba028c6662789194572-Abstract-Datasets_and_Benchmarks.html |
| [41] |
Liu F, Dai W, Zhang C, Zhu J, Yao L, et al. 2025. Co-LLaVA: efficient remote sensing visual question answering via model collaboration. |
| [42] |
Li Y, Du Y, Zhou K, Wang J, Zhao X, et al. 2023. Evaluating object hallucination in large vision-language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, Singapore. Stroudsburg, PA, USA: ACL. pp. 292–305 doi: 10.18653/v1/2023.emnlp-main.20 |
| [43] |
Jiang Y, Irvin J, Wang JH, Chaudhry MA, Chen JH, et al. 2024. Many-shot in-context learning in multimodal foundation models. |
| [44] |
Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35 (NeurIPS 2022), November 28 – December 9, 2022, New Orleans, LA, USA. Curran Associates, Inc. pp. 24824–24837 https://proceedings.neurips.cc/paper_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html |
| [45] |
Menon S, Vondrick C. 2023. Visual classification via description from large language models. The Eleventh International Conference on Learning Representations (ICLR 2023), May 1–5, 2023, Kigali, Rwanda. OpenReview.net. https://openreview.net/forum?id=41GFzHFc4zx |
| [46] |
Chia YK, Chen G, Tuan LA, Poria S, Bing L. 2023. Contrastive chain-of-thought prompting. |
| [47] |
Jiang X, Tang H, Gao J, Du X, He S, et al. 2024. Delving into multimodal prompting for fine-grained visual classification. |
| [48] |
Zhang Z, Zhang A, Li M, Zhao H, Karypis G, et al. 2024. Multimodal chain-of-thought reasoning in language models. Transactions on Machine Learning Research 2024. https://openreview.net/forum?id=y1pPWFVfvR |
| [49] |
Pratt S, Covert I, Liu R, Farhadi A. 2023. What does a platypus look like? generating customized prompts for zero-shot image classification. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), October 1–6, 2023, Paris, France. pp. 15645–15655 doi: 10.1109/ICCV51070.2023.01438 |
| [50] |
Liu M, Roy S, Li W, Zhong Z, Sebe N, et al. 2024. Democratizing fine-grained visual recognition with large language models. The Twelfth International Conference on Learning Representations (ICLR 2024), May 7–11, 2024, Vienna, Austria. OpenReview.net. https://openreview.net/forum?id=c7DND1iIgb |
| [51] |
Pan J, Zhong R, Xia F, Huang J, Zhu L, et al. 2025. ChatLeafDisease: a chain-of-thought prompting approach for crop disease classification using large language models. |
| [52] |
Lu Y, Bartolo M, Moore A, Riedel S, Stenetorp P. 2021. Fantastically ordered prompts and where to find them: overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 2022, Dublin. Ireland Association for Computational Linguistics. pp. 8086−8098 doi: 10.18653/v1/2022.acl-long.556 |
| [53] |
Zhang Y, Zhou K, Liu Z. 2023. What makes good examples for visual in-context learning? In Advances in Neural Information Processing Systems 36. December 10–16, 2023. New Orleans, Louisiana, USA. NeurIPS. pp. 17773–17794 doi: 10.52202/075280-0780 |
| [54] |
Min S, Lyu X, Holtzman A, Artetxe M, Lewis M, et al. 2022. Rethinking the role of demonstrations: what makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, Abu Dhabi, United Arab Emirates. Stroudsburg, PA, USA: ACL. pp. 11048–11064 doi: 10.18653/v1/2022.emnlp-main.759 |
| [55] |
Ferber D, Wölflein G, Wiest IC, Ligero M, Sainath S, et al. 2024. In-context learning enables multimodal large language models to classify cancer pathology images. |
| [56] |
Yao L, Yang Y. 2026. Large language models are contrastive reasoners. |
| [57] |
Lu X, Yu X, Wang K, Wang Y, Wang P, et al. 2023. SupCon-ViT: supervised contrastive learning for ultra-fine-grained visual categorization. In 2023 International Conference on Digital Image Computing: Techniques and Applications (DICTA), 2023, Port Macquarie, Australia. USA: IEEE. pp. 281–288 doi: 10.1109/DICTA60407.2023.00046 |
| [58] |
Chiang WL, Zheng L, Sheng Y, Angelopoulos AN, Li T, et al. 2024. Chatbot arena: an open platform for evaluating LLMs by human preference. Proceedings of the 41st International Conference on Machine Learning (ICML 2024), July 21–27, 2024, Vienna, Austria. PMLR. Vol. 235. pp. 8359–8388 https://proceedings.mlr.press/v235/chiang24b.html |
| [59] |
Simonyan K, Zisserman A. 2015. Very deep convolutional networks for large-scale image recognition. International Conference on Learning Representations (ICLR 2015), May 7–9, 2015, San Diego, CA, USA. arXiv:1409.1556. doi: 10.48550/arXiv.1409.1556 |
| [60] |
Chen Y, Bai Y, Zhang W, Mei T. 2019. Destruction and construction learning for fine-grained image recognition. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, Long Beach, CA, USA. USA: IEEE. pp. 5152–5161 doi: 10.1109/CVPR.2019.00530 |