Search
2026 Volume 41
Article Contents
RESEARCH ARTICLE   Open Access    

GeoTP: parameter-level intelligent planner of GIS tool-use chains based on instruction tuning of large language models

More Information
  • Geospatial task solving typically requires orchestrating multiple tools into a workflow, and automating this process is crucial for enhancing the efficiency of Geographic Information Systems (GIS) applications. Traditional rule-based tool planning systems suffer from limited flexibility and adaptability, while approaches based on large language models (LLMs) for autonomous tool-chain composition and parameter selection face challenges, including high costs associated with closed-source models and scarce training data for open-source alternatives. To address these challenges, this paper proposes an instruction-tuning framework tailored for geospatial tasks using open-source LLMs. First, we introduce a structured tool description and a parameterized instruction-generation framework, leveraging a self-instruct strategy to automatically construct a large-scale instruction-tuning dataset, named GeoITC, which enhances the model's understanding of parameter constraints and tool dependencies. Second, we employ a parameter-efficient fine-tuning technique (LoRA) to train the LLaMA-3-8B model, resulting in a domain-specialized geospatial model, GeoTP. Experiments conducted on our newly constructed multi-level evaluation benchmark, GeoITC-Eval, demonstrate that GeoTP significantly outperforms baseline models in both tool-chain accuracy (CTOA) and parameter accuracy (PA), and exhibits strong generalization capabilities on unseen tools and complex, long-chain tasks. This study shows that fine-tuning open-source LLMs with carefully designed parameterized instructions can effectively enhance their ability to plan complex geospatial workflows and configure precise parameters, offering a practical pathway toward developing low-cost, privately deployable geospatial agents.
  • 加载中
  • Supplementary Table S1 Specification of the geospatial tools in the GeoITC dataset.
  • [1] Arefiev N, Terleev V, Badenko V. 2015. GIS-based fuzzy method for urban planning. Procedia Engineering 117:39−44 doi: 10.1016/j.proeng.2015.08.121

    CrossRef   Google Scholar

    [2] Albert L, Rottensteiner F, Heipke C. 2014. Land use classification using conditional random fields for the verification of geospatial databases. ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences II-4:1−7 doi: 10.5194/isprsannals-ii-4-1-2014

    CrossRef   Google Scholar

    [3] Resch B, Sagl G, Törnros T, Bachmaier A, Eggers JB, et al. 2014. GIS-based planning and modeling for renewable energy: Challenges and future research avenues. ISPRS International Journal of Geo-Information 3(2):662−692 doi: 10.3390/ijgi3020662

    CrossRef   Google Scholar

    [4] Franch-Pardo I, Napoletano BM, Rosete-Verges F, Billa L. 2020. Spatial analysis and GIS in the study of COVID-19. A review. Science of the Total Environment 739:140033 doi: 10.1016/j.scitotenv.2020.140033

    CrossRef   Google Scholar

    [5] Fan Y, Wu W, Wang W, Liu M, Wen Q. 2016. 中国灾害遥感研究进展 [Research progress of disaster remote sensing in China]. 遥感学报 [National Remote Sensing Bulletin] 20(5):1170−1184 (in Chinese) doi: 10.11834/jrs.20166171

    CrossRef   Google Scholar

    [6] He W, Jiang Z. 2023. Uncertainty quantification of deep learning for spatiotemporal data: challenges and opportunities. arXiv Preprint: 2311.02485 doi: 10.48550/arXiv.2311.02485

    CrossRef   Google Scholar

    [7] Zhu AX, Zhao FH, Liang P, Qin CZ. 2021. Next generation of GIS: must be easy. Annals of GIS 27(1):71−86 doi: 10.1080/19475683.2020.1766563

    CrossRef   Google Scholar

    [8] Gao S, Goodchild MF. 2013. Asking spatial questions to identify GIS functionality. Proc. Fourth International Conference on Computing for Geospatial Research and Application, July 22–24, 2013, San Jose, CA, USA. USA: IEEE. pp. 106–110 doi: 10.1109/COMGEO.2013.18
    [9] Allen DW. 2011. Getting to know ArcGIS ModelBuilder. Redlands: Esri Press. 362 pp
    [10] Mericskay B. 2018. Automation of workflows for the installation of a wind farm. In QGIS and Applications in Territorial Planning, eds. Baghdadi N, Mallet C, Zribi M. vol. 3. UK, USA: ISTE Ltd, John Wiley & Sons. pp. 125−168 doi: 10.1002/9781119457121.ch5
    [11] Scheider S, Nyamsuren E, Kruiger H, Xu H. 2021. Geo-analytical question-answering with GIS. International Journal of Digital Earth 14(1):1−14 doi: 10.1080/17538947.2020.1738568

    CrossRef   Google Scholar

    [12] Guo T, Chen X, Wang Y, Chang R, Pei S, et al. 2024. Large language model based multi-agents: a survey of progress and challenges. arXiv Preprint: 2402.01680 doi: 10.48550/arXiv.2402.01680

    CrossRef   Google Scholar

    [13] Wang L, Ma C, Feng X, Zhang Z, Yang H, et al. 2024. A survey on large language model based autonomous agents. Frontiers of Computer Science 18(6):186345 doi: 10.1007/s11704-024-40231-1

    CrossRef   Google Scholar

    [14] Cui Y, Gao Y. 2025. 基于大语言模型的地理信息查询系统 [Geographic information query system based on large language models]. 计算机科学与应用 [Computer Science and Application] 15(8):198−206 (in Chinese) doi: 10.12677/csa.2025.158210

    CrossRef   Google Scholar

    [15] Wang MQ, Li BZ, Wang ZL, Liu SC, Liao C, et al. 2026. 一种集成知识图谱和大语言模型的智能地图制图框架 [An Automatic Cartography Framework Integrating Knowledge Graph and Large Language Model]. 武汉大学学报 (信息科学版) [Geomatics and Information Science of Wuhan University] 51(4):840−851 doi: 10.13203/j.whugis20240266

    CrossRef   Google Scholar

    [16] Akinboyewa T, Ning H, Lessani MN, Li Z. 2024. Automated floodwater depth estimation using large multimodal model for rapid flood mapping. Computational Urban Science 4(1):12 doi: 10.1007/s43762-024-00123-3

    CrossRef   Google Scholar

    [17] Akinboyewa T, Li Z, Ning H, Lessani MN. 2025. GIS Copilot: Towards an autonomous GIS agent for spatial analysis. International Journal of Digital Earth 18(1):2497489 doi: 10.1080/17538947.2025.2497489

    CrossRef   Google Scholar

    [18] Zhu G, Zhang Z, Cao L, Ma K, Xu X, et al. 2025. 基于DeepSeek的文生地图智能体构建方法 [Research on the Construction Method for Text-to-Map Agent Based on DeepSeek]. 地球信息科学学报 [Journal of Geo-information Science] 27(9):2165−2176 (in Chinese) doi: 10.12082/dqxxkx.2025.250207

    CrossRef   Google Scholar

    [19] Wei C, Zhang Y, Zhao X, Zeng Z, Wang Z, et al. 2025. GeoTool-GPT: a trainable method for facilitating Large Language Models to master GIS tools. International Journal of Geographical Information Science 39(4):707−731 doi: 10.1080/13658816.2024.2438937

    CrossRef   Google Scholar

    [20] Zhang Y, Wei C, He Z, Yu W. 2024. GeoGPT: an assistant for understanding and processing geospatial tasks. International Journal of Applied Earth Observation and Geoinformation 131:103976 doi: 10.1016/j.jag.2024.103976

    CrossRef   Google Scholar

    [21] Yao LW, Ren F, Wang SY, Wang Y, Zhang BC, et al. 2025. 地震灾害应急预案的生成式编制方法 [A Generative Method for Earthquake Emergency Plans]. 武汉大学学报(信息科学版) [Geomatics and Information Science of Wuhan University] 50(6):1159−1174 (in Chinese) doi: 10.13203/j.whugis20250127

    CrossRef   Google Scholar

    [22] Vinayan Kozhipuram A, Shailendra S, Kadel R. 2025. Retrieval-augmented generation vs. baseline LLMs: a multi-metric evaluation for knowledge-intensive content. Information 16(9):766 doi: 10.3390/info16090766

    CrossRef   Google Scholar

    [23] Hou S, Jiao H, Liang J, Shen Z, Zhao A, et al. 2026. GeoCogent: an LLM-based agent for geospatial code generation. International Journal of Geographical Information Science 40:1073−1106 doi: 10.1080/13658816.2025.2549460

    CrossRef   Google Scholar

    [24] Hao BW, Liu YF, Li LY, Wang J, Peng Y. 2024. 基于多模态推荐指令的大语言模型指令微调 [The instruction tuning of large language models with multi-modal recommendation instruction]. 北京邮电大学学报 [Journal of Beijing University of Posts and Telecommunications] 47(4):36−43 (in Chinese) doi: 10.13190/j.jbupt.2023-269

    CrossRef   Google Scholar

    [25] Zhang Y, Li J, Wang Z, He Z, Guan Q, et al. 2025. Geospatial large language model trained with a simulated environment for generating tool-use chains autonomously. International Journal of Applied Earth Observation and Geoinformation 136:104312 doi: 10.1016/j.jag.2024.104312

    CrossRef   Google Scholar

    [26] Chen G, Jiao H, Hou S, Liu Z, Xie L, et al. 2025. GeoJSEval: an automated evaluation framework for large language models on JavaScript-based geospatial computation and visualization code generation. ISPRS International Journal of Geo-Information 14(10):382 doi: 10.3390/ijgi14100382

    CrossRef   Google Scholar

    [27] Wang C, Wen L, Jia S, Zhang X, Xu L. 2025. Light-IF: endowing LLMs with generalizable reasoning via preview and self-checking for complex instruction following. arXiv Preprint: 2508.03178 doi: 10.48550/arXiv.2508.03178

    CrossRef   Google Scholar

    [28] Wang Y, Kordi Y, Mishra S, Liu A, Smith NA, et al. 2023. Self-instruct: aligning language models with self-generated instructions. Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers), Toronto, Canada, 2023. USA: ACL. pp. 13484−13508 doi: https://doi.org/10.18653/v1/2023.acl-long.754
    [29] Hao S, Gu Y, Ma H, Hong J, Wang Z, et al. 2023. Reasoning with language model is planning with world model. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Singapore: Association for Computational Linguistics. pp. 8154−8173 doi: 10.18653/v1/2023.emnlp-main.507
    [30] Liu A, Feng B, Xue B, Wang B, Wu B, et al. 2024. Deepseek-v3 technical report. arXiv Preprint: 2412.19437 doi: 10.48550/arXiv.2412.19437

    CrossRef   Google Scholar

    [31] Yang A, Li A, Yang B, Zhang B, Hui B, et al. 2025. Qwen3 technical report. arXiv Preprint: 2505.09388 doi: 10.48550/arXiv.2505.09388

    CrossRef   Google Scholar

  • Cite this article

    Tang J, Li Z, Yang J, Lin D, Zhang B, et al. 2026. GeoTP: parameter-level intelligent planner of GIS tool-use chains based on instruction tuning of large language models. The Knowledge Engineering Review 41: e012 doi: 10.48130/ker-0026-0013
    Tang J, Li Z, Yang J, Lin D, Zhang B, et al. 2026. GeoTP: parameter-level intelligent planner of GIS tool-use chains based on instruction tuning of large language models. The Knowledge Engineering Review 41: e012 doi: 10.48130/ker-0026-0013

Figures(12)  /  Tables(8)

Article Metrics

Article views(364) PDF downloads(69)

RESEARCH ARTICLE   Open Access    

GeoTP: parameter-level intelligent planner of GIS tool-use chains based on instruction tuning of large language models

The Knowledge Engineering Review  41 Article number: e012  (2026)  |  Cite this article

Abstract: Geospatial task solving typically requires orchestrating multiple tools into a workflow, and automating this process is crucial for enhancing the efficiency of Geographic Information Systems (GIS) applications. Traditional rule-based tool planning systems suffer from limited flexibility and adaptability, while approaches based on large language models (LLMs) for autonomous tool-chain composition and parameter selection face challenges, including high costs associated with closed-source models and scarce training data for open-source alternatives. To address these challenges, this paper proposes an instruction-tuning framework tailored for geospatial tasks using open-source LLMs. First, we introduce a structured tool description and a parameterized instruction-generation framework, leveraging a self-instruct strategy to automatically construct a large-scale instruction-tuning dataset, named GeoITC, which enhances the model's understanding of parameter constraints and tool dependencies. Second, we employ a parameter-efficient fine-tuning technique (LoRA) to train the LLaMA-3-8B model, resulting in a domain-specialized geospatial model, GeoTP. Experiments conducted on our newly constructed multi-level evaluation benchmark, GeoITC-Eval, demonstrate that GeoTP significantly outperforms baseline models in both tool-chain accuracy (CTOA) and parameter accuracy (PA), and exhibits strong generalization capabilities on unseen tools and complex, long-chain tasks. This study shows that fine-tuning open-source LLMs with carefully designed parameterized instructions can effectively enhance their ability to plan complex geospatial workflows and configure precise parameters, offering a practical pathway toward developing low-cost, privately deployable geospatial agents.

    • Geographic Information Systems (GIS), a core tool for spatial analysis and decision support, have been widely applied in various fields, including urban planning[1,2], traffic management[3], disaster response[4,5], and resource management[6]. Faced with increasingly complex spatial analysis needs, it is crucial to precisely configure each tool's parameters (e.g., buffer distance, field expression) to solve problems. Automating this process by generating executable workflows with correct tool sequences and compliant parameter settings directly from natural language instructions is crucial for significantly improving the efficiency of GIS applications and lowering the barrier to spatial analysis[7,8].

      Rule-based systems (e.g., ArcGIS ModelBuilder) are classic solutions for formalizing specific workflows[9,10], but they are inflexible and difficult to adapt to dynamic changes in tasks[11]. In recent years, large language models (LLMs) have demonstrated powerful capabilities in semantic understanding, reasoning, and code generation, providing new approaches to solving intelligent geospatial tasks[12,13]. Researchers have begun to explore the deep integration of LLMs and GIS[14,15], forming a range of parallel research methods and technical approaches, aiming to achieve automation at different levels and in different scenarios.

      Deeply integrating LLMs into existing GIS platforms (e.g., QGIS) has become the mainstream practice for achieving copilot-style interaction. Such research aims to overcome the challenges of integrating AI into complex GIS environments, including tool diversity, heterogeneity in data standards, and the accurate conversion of natural language into GIS commands[16]. The GIS Copilot[17] plugin guides LLMs in generating accurate PyQGIS code through a structured library of tool descriptions and a code review module, demonstrating a high success rate for basic and intermediate tasks. This type of work marks the evolution from externally invoked tools to in-platform intelligent agents. Its advantage lies in the close integration of the GIS software's computing ecosystem and visualization capabilities, which enhances analytical transparency and control over results. Utilizing the planning and tool invocation capabilities of commercial LLMs has become a key pathway to achieve high-performance, general-purpose automated systems[18]. The GeoGPT framework centers on GPT-4, which autonomously plans and invokes a predefined toolset through prompt engineering to complete the entire workflow from data collection to visualization[19]. Such systems demonstrate exceptional accuracy when handling complex workflows involving up to eight steps, highlighting the pivotal role of powerful LLMs in complex task planning.

      However, the outstanding performance of LLMs largely relies on closed-source commercial models such as GPT-4, which present significant limitations regarding data privacy, localized deployment, and long-term usage costs. Additionally, LLM-based approaches still face challenges, including tool semantic ambiguity and output hallucinations[20]. To address LLMs' reliance on domain-specific knowledge and their hallucination issues, retrieval-augmented generation (RAG) technology[21] has been systematically integrated into geospatial task automation frameworks. RAG provides real-time, reliable knowledge support for LLMs' decision-making and generation processes by retrieving relevant information from external knowledge bases, such as authoritative scientific reports, professional literature, or structured tool documentation, thereby significantly enhancing the accuracy and reliability of the output[22]. In the field of GIS, RAG is regarded as a potential solution to address the issue of lengthy tool documentation that exceeds the contextual limits of models, enabling LLMs to dynamically retrieve parameter details and usage examples for required tools. GeoCogent systematically constructs a multimodal geospatial knowledge base and deeply integrates the RAG mechanism, effectively improving the accuracy of code generation[23]. However, its performance remains constrained by retrieval quality and real-time responsiveness, and it cannot fundamentally internalize domain knowledge as a stable capability of the model. To fundamentally reduce reliance on closed-source models and develop an open-source LLM capable of private deployment and intelligent, parameter-level planning within GIS tool-use chains, researchers have shifted toward deeply empowering open-source models through instruction tuning[24]. Parameter-efficient fine-tuning techniques, such as LoRA, are favored for their ability to incorporate domain knowledge efficiently. Zhang et al. applied LoRA to fine-tune LLaMA-2-7B, training a specialized model capable of generating GIS tool-use chains[25]. Chen et al. significantly improved the understanding and generative performance of open-source models in the geoscience domain through a combined strategy of continued pre-training and instruction tuning[26]. These efforts preliminarily validate that, through domain-specific data fine-tuning, open-source LLMs can acquire capabilities to handle geospatial tasks. Researchers from GeoGPT and GeoCogent have also clearly indicated that developing specialized local LLMs tailored for the GIS field is a crucial future direction to reduce reliance on commercial APIs and enhance domain adaptability.

      Despite significant progress, current solutions based on open-source LLMs still face several bottlenecks. First, their planning accuracy and parameter compliance in complex tasks remain significantly weaker than those of closed-source models such as GPT-4, which limits their practical deployment. Second, existing datasets largely rely on manual construction, resulting in small scale and low diversity, and particularly lack standardized constraints on tool parameters. Consequently, the models tend to generate syntactically correct but semantically erroneous parameter configurations. Third, there is a lack of systematic evaluation benchmarks to comprehensively assess model performance in tool-use chain planning and parameter configuration. To address these challenges, this study proposes an instruction tuning training framework for open-source LLMs for geospatial tasks. The core contributions are as follows:

      (1) A structured method for geospatial tool description and parameterized instruction data generation.By introducing a self-instruct generation strategy and incorporating domain knowledge, we guide a large language model to automatically generate diverse, multi-complexity 'instruction–tool chain' pairs. This approach places particular emphasis on the structured representation of parameter templates, data types, and value constraints, thereby laying a foundation for the model to learn precise tool invocation logic and parameter adaptation.

      (2) A high-quality instruction tuning dataset, GeoITC, and a systematic evaluation benchmark, GeoITC-Eval. The dataset covers single-tool, dual-tool, and multi-tool tasks, and extends to unseen tools and extra-long workflow scenarios, forming a multi-level evaluation system. Model performance is comprehensively quantified from two dimensions: tool chain order accuracy (CTOA) and parameter accuracy (PA).

      (3) A high-performance geospatial tool orchestration model, GeoTool Planner (GeoTP).Based on the LLaMA-3-8B model, we inject geospatial tool usage capabilities via LoRA. Experimental results show that GeoTP achieves outstanding performance across multiple tasks. It not only substantially outperforms comparable models on known tool tasks, but also demonstrates robust generalization to unseen tools and long-chain workflow planning.

    • The framework of the proposed method, illustrated in Fig. 1, consists of three components: dataset construction, instruction tuning, and GIS tool planning.

      Figure 1. 

      The framework of the GIS tool planning model GeoTP.

    • Two datasets, GeoITC and the evaluation dataset GeoITC-Eval, were constructed for fine-tuning General LLMs to enhance their capability in solving Geospatial Tasks. To overcome the bottlenecks of scarce geospatial instructional data and high manual design costs, a data-efficient construction approach that combines the self-guidance generation strategy with a Simulated Environment was proposed. First, the collected GIS tools are standardized and presented with parameterized descriptions to establish a standard format that includes tool functions, structured parameter templates, data constraints, and examples. Based on this, the high-performance LLMs are guided to automatically generate diverse geospatial task instructions of varying levels of complexity (single-tool, dual-tool, and triple-tool) within a simulated GIS execution environment, while simultaneously producing their corresponding, fully parameterized tool-use chain solutions.

    • Based on the open-source general-purpose LLM LLaMA-3-8B, the model is fine-tuned using the LoRA Fine-tuning technique. This approach freezes the original weights of the base model (LLaMA-3-8B) and introduces trainable Low-rank matrices alongside key modules such as attention for adaptive adjustment. It significantly reduces the number of training parameters, thereby lowering GPU memory consumption and training time. By applying LoRA fine-tuning on the high-quality dataset GeoITC, the model learns to map natural language descriptions to structured tool-use chains and parameter configurations. The resulting GIS tool intelligent planning model, GeoTP, retains the general language capabilities of the Base model while accurately mastering the planning and execution of complex geospatial workflows.

    • Driven by natural-language task instructions, the fine-tuned LLMs autonomously orchestrate GIS tool-use chains and configure parameters. When entering an instruction, the primary objectives of the geospatial analysis task and the corresponding input data requirements must be clearly defined, with explicit annotations of data types and storage paths. This ensures that the task instruction is complete and semantically accurate, without requiring the specification of GIS tools, operational steps, or parameter settings.

    • The core objective of instruction tuning is to align the model's output with the dialogue style and knowledge system of human experts or specific domains, thereby significantly enhancing the model's ability to understand and execute complex tasks based on given instructions. To achieve this goal, constructing a high-quality instruction tuning dataset is crucial, and obtaining accurate 'instruction–response' pairs is particularly important. This process requires not only that the tool-use chain logic in the responses be correct and comprehensive, but also that the parameter configuration of each tool strictly comply with the standards and semantic constraints of the geospatial domain. For example, when the model is tasked with identifying suitable areas for park construction in urban planning—ensuring that the area is within 500 m of a school and avoids water protection zones—it must accurately parse multiple geographic constraints and plan the correct sequence of spatial analysis steps. These steps include selecting appropriate tools and setting corresponding parameters, such as the buffer distance in 'buffer analysis' and the layer processing method in 'overlay analysis', ultimately generating a complete and executable tool-use chain. The ability to understand and configure tool functions and parameters is essential to enabling the model to convert users' natural-language descriptions into a professional and reliable sequence of operations. It also serves as the foundation for deploying practical LLMs in specialized domains.

      Currently, the self-instruct strategy is a widely adopted and efficient approach for constructing Instruction Tuning Datasets. This strategy guides LLMs themselves to participate in data generation, automatically creating 'instruction–solution' pairs and thereby reducing reliance on manual annotation. The process consists of two stages: instruction generation and solution construction, with the latter focusing on producing executable sequences of operations that include specific tool invocations and parameter configurations. In this study, the strategy is employed to automatically generate supervised data using a high-performance LLM, covering diverse geospatial tasks, their corresponding tool-use chains, and parameterized operations. This provides structured and large-scale training resources that enable precise planning and execution of the model in complex geospatial scenarios. The geospatial domain-specific model presented in this study is built on a general LLM and further refined through instruction tuning in the geospatial domain. This process enables the model to acquire professional capabilities such as understanding spatial analysis tasks, autonomously orchestrating tool-use chains, and appropriately configuring parameters.

      Because DeepSeek-V3.1 performs exceptionally well in structured output and tool-calling tasks and can stably generate syntactically correct and logically coherent tool chains, the model has implicitly absorbed some GIS tool knowledge during large-scale pretraining, making it better aligned with the needs of this study for geographic information scenarios. Moreover, the model supports a context window of up to 128 K tokens, which enables it to effectively handle long task descriptions involving multi-tool nesting and complex parameter constraints[27]. Therefore, this paper selects it as a high-performance teacher model and uses structured prompt engineering to explicitly express its implicit spatial inference knowledge, thereby generating the high-quality dataset GeoITC to empower small open-source models that can be deployed locally.

    • This study aims to guide open-source LLMs in developing a deeper understanding of the functional characteristics, parameter settings, and applicable contexts of GIS tools, with a particular focus on enabling them to master multi-tool collaborative methods for specific geospatial tasks. Relevant tool documentation was systematically collected, and textual descriptions of their functions and parameters were extracted. Specifically, the original documentation of 18 commonly used QGIS tools was standardized, and a unified structured description template was designed for each tool, as shown in Eq. (1). Where $ {n}_{i} $ represents the tool name, $ {f}_{i} $ indicates the tool function description, $ {p}_{i} $ denotes the parameter design template, $ {r}_{i} $ specifies the input data requirements, $ {e}_{i} $ provides an example of using the tool, and $ T $ stands for the complete tool combination.

      $ \begin{aligned}{t}_{i}&=\{{n}_{i},{f}_{i},{p}_{i},{r}_{i},{e}_{i}\}\\ T&=\{{t}_{1},{t}_{2},...,{t}_{18}\} \end{aligned} $ (1)

      Table 1 presents the names, functions, and parameters of the three tools, with the complete tool list provided in Supplementary Table S1. The structured description template for the tools is shown in Fig. 2 and primarily includes algorithm functions, parameter explanations, and data constraints. Although in actual engineering deployment the final storage path can be fully distributed automatically by the underlying GIS platform, this study still treats it as a mandatory generated parameter during the experimental evaluation stage. The purpose of this design is to maximize the testing of the large model's instruction-following capability under strong constraints and its alignment accuracy with a specific location. This structured format helps reduce the complexity of the input text and alleviates the understanding burden on the LLMs.

      Table 1.  Specification of sample geospatial tools in GeoITC.

      No. Tool name Function description Parameter name
      1 OSM Downloader This tool allows users to download OpenStreetMap (OSM) data by selecting an area using a rectangle. Download Area, Output Format
      2 OpenTopography DEM Downloader This tool downloads Digital Elevation Models (DEMs) for the extent defined by the user from OpenTopography. DEM Type, Download Extent, Output Raster
      3 Points Along Geometry This algorithm creates a point layer with points distributed along the lines of an input vector layer. Input Layer, Distance, Start Offset, End Offset, Output Layer

      Figure 2. 

      Structured tool description template.

    • To construct diverse and parameter-complete descriptions of geospatial tasks, this study employs the collected QGIS toolset and its structured description template, applying a Self-Instruct strategy[28] to guide a general LLM in generating geospatial task descriptions that encompass a wide range of scenarios with well-defined parameter settings. Given the strong semantic understanding and planning capabilities of DeepSeek-V3.1, this study primarily relies on it to implement geospatial task generation using the self-instruct strategy.

      In practice, given the diversity and complexity of geospatial tasks, three distinct prompt templates were designed for single-tool, dual-tool, and triple-tool tasks. In the single-tool task generation template (Fig. 3), the model is explicitly required to generate instructions that depend solely on the specified tool and include complete parameter settings, covering both the data path and parameter values. To enhance the model's understanding of the correspondence between tasks and parameters, manually constructed example instructions containing parameterized samples were incorporated into the prompt. In the dual-tool (Fig. 4) and triple-tool (Fig. 5) task-generation templates, the structure is similar but emphasizes integrating the tools into a coherent workflow, with parameters between successive steps properly linked and transferable. Through this hierarchical prompting system, geospatial task data spanning diverse scenarios and complexities were successfully collected, providing high-quality supervision signals with rich parameters and rigorous logic for model training.

      Figure 3. 

      The prompt for single-tool geospatial task generation.

      Figure 4. 

      The prompt for dual-tool geospatial task generation.

      Figure 5. 

      The prompt for three-tools geospatial task generation.

      This approach not only alleviates the scarcity of geospatial task instruction data but also ensures the accuracy and executability of generated tasks in tool invocation and parameter configuration through structured parameter descriptions and hierarchical prompt design. It provides a reliable means to enhance the LLMs' capability in planning and parameterized tool use within professional GIS fields.

    • For handling geospatial tasks, the core objective lies in generating a complete solution with a correct sequence of tools and accurate parameter configuration. Traditionally, obtaining such a solution relied on manual operations and process documentation by GIS experts, a method that is inefficient and difficult to scale up. To address this, this study adopts an automatic generation strategy based on an LLM. The core lies in guiding the model, through carefully designed prompt engineering, to produce parameterized and executable workflows. The specific prompt design is illustrated in Fig. 6. First, a structured description of tools is provided to the model, enabling it to accurately understand the tool call interfaces and thereby establish a foundation for generating reasonable parameter values. Second, after understanding the task objective, the model needs to plan the tool execution sequence and perform the critical task of parameter adaptation. This requires the model to parse the task context (for example, 'Calculate the road length of main roads along both banks of the Yangtze River in Wuhan city, with the input data being the regional boundary file wuhan_river_area.shp') and recognize that the input file wuhan_river_area.shp defines the spatial scope of the task. The model then infers that obtaining road data within this boundary is the primary step. It understands that 'both banks of the Yangtze River' represents a filter condition based on a spatial relationship, but also realizes that data directly downloaded from an open-source map such as OSM typically do not contain this attribute. Therefore, it plans to first retrieve all road data, followed by spatial analysis for filtering if necessary. This reasoning guides the model to initially select the OSM Downloader Tool and correctly map the spatial extent of wuhan_river_area.shp to the Download Area parameter of that tool. The model recognizes that 'calculating length' requires precise geometric computation, while the downloaded geographic data are typically in a Geographic Coordinate System. Therefore, a projection transformation must be performed to obtain lengths measured in meters. This leads the model to sequentially include the Reproject Layer and Field Calculator tools in the tool-use chain. It can automatically set the output of the previous step as the input for the next, forming a coherent data flow. More importantly, it can generate the Professional Expression $length for the Field Calculator tool to compute the Geometric Length of each line feature after projection, and name the resulting field road_length_m, demonstrating an accurate understanding of GIS professional operations and semantics. It should be noted that although this paper represents tool combinations in sequential form, this design is capable of expressing complex multi-branch topologies such as tree structures and Directed-Acyclic Graphs (DAGs). In the design of the GeoITC dataset, dependencies between tools are not limited to adjacent upstream and downstream relationships; subsequent tools can explicitly reference parameter variables to accept the outputs of multiple preceding tools as inputs simultaneously.

      Figure 6. 

      The prompt used for parameterized tool-use chain generation.

      The specific parameter values in this dataset are strictly treated as learning examples, rather than fixed rules that the model is expected to memorize. Our design objective is to teach the model, through a large number of examples covering different numerical values and data sources, to understand the semantic typing of parameters, their numerical constraints, and their contextual roles in chained propagation. Whether the model can dynamically generate adaptive parameters based on new task descriptions and data files will be evaluated by its generalization performance on unseen tools and long-chain tasks. This design aims to cultivate the model's ability to abstract general parameter inference capabilities from concrete examples, rather than template matching.

      Since LLMs cannot directly execute GIS operations, a structured output format is designed to simulate this process. Specifically, the model's output must include two parts: first, an assessment of task feasibility; if the task is solvable, the model should then generate a high-level solution framework and produce a workflow that includes a detailed sequence of tools, parameter configuration, and data flow. To enhance the reliability of the generated results, this design incorporates a self-verification mechanism: the sequence and functional intent of the tools planned in the framework must remain consistent with the specific invocations and parameter settings in each step of the workflow. Any inconsistency indicates the presence of hallucinations or logical errors in the output and should therefore be discarded. Studies have shown that this step-by-step inference approach of 'planning first, then refining'[29] can effectively enhance the accuracy and reliability of model responses, thereby enabling the systematic acquisition of high-quality and verifiable geospatial task tool-use chains.

      The GeoITC constructed in this study is a parameterized tool-use chain Instruction Tuning Dataset for geospatial tasks, built based on the self-guidance generation strategy (Eq.[2]). The term instruction denotes the instruction task, and output represents the parameterized tool-use chain. To generate Instruction Tuning data with varying levels of complexity, this study collected three levels of geospatial tasks along with their corresponding tool-use chains, including single-tool, dual-tool, and triple-tool tasks. Specifically, for single-tool tasks, each tool was repeated 35 times to help LLMs fully understand its functions and usage methods (18 × 35 = 630). All dual-tool combinations were repeated 20 times to enable LLMs to learn how different tool combinations can be used to solve geospatial tasks (18 × 17 ÷ 2 × 20 = 3,060). After completing the alignment between the framework and the workflow, unmatched cases were removed, resulting in a total of 2,740 dual-tool instructions. For the triple-tool tasks, all triple-tool combinations were repeated four times, generating 18 × 17 × 16 ÷ 3 ÷ 2 × 4 = 3,264 instructions. After excluding unmatched cases, 2,088 instructions were retained. Detailed statistics of all collected instruction tuning data are presented in Table 2, where the tool numbers in Table 2 indicate the number of tools required to solve the corresponding geospatial tasks.

      $ \begin{aligned}dat{a}_{i}&=\{instructio{n}_{i},inpu{t}_{i},outpu{t}_{i}\}\\ GeoITC&=\{dat{a}_{1},data{2}_{2},...,dat{a}_{5458}\} \end{aligned} $ (2)

      Table 2.  Statistics of the instruction tuning data in GeoITC.

      Number of tools 1 2 3 Total
      Number of instructions 630 2,740 2,088 5,458
    • To assess the capability of language models in solving geospatial tasks, the evaluation dataset GeoITC-Eval was collected. Similar to the instruction tuning data, three levels of Geospatial Tasks were gathered. For tasks that require only a single tool, a total of 270 tasks were collected (18 tools, each repeated 15 times). For tasks requiring two tools, the 18 available tools were paired in all possible combinations to form nine unique pairs; this process was repeated ten times, resulting in 90 tasks. For tasks requiring three tools, the 18 tools were grouped into sets of three to form six unique triples; this process was also repeated ten times, yielding 60 tasks in total. Unlike the instruction tuning data collected solely through Deepseek-V3.1, the evaluation data require professional verification. Specifically, the details and framework of each geospatial task, as well as every step of the workflow, were reviewed by two GIS experts holding master's degrees. This process ensures that the evaluation data can be effectively used to assess model performance. Table 3 presents the statistical results of all collected evaluation data.

      Table 3.  Statistics of the GeoITC-Eval,GeoITC-EvalProlong and GeoITC-EvalExtend.

      Number of tools 1 2 3 4 5 Total
      GeoITC-Eval 270 90 60 0 0 420
      GeoITC-EvalProlong 0 0 0 30 20 50
      GeoITC-EvalExtend 15 15 15 0 0 45

      In addition, two extended datasets were collected to evaluate the performance of GeoTP in scenarios beyond the training data. Specifically, the first extended tool-use chain length dataset, GeoITC-EvalProlong, includes complex tasks that require four to five GIS tools to complete. Using GeoITC-EvalProlong, GeoTP can be tested on tasks that are more complex than those in the training data. Furthermore, three external GIS tools were collected, and the second extended tool dataset, GeoITC-EvalExtend, was constructed based on tools that were not used during GeoITC training. Table 4 presents information on these three tools. Specifically, three levels of geospatial tasks were collected: for geospatial tasks requiring a single tool, 15 instructions were gathered (five repetitions for each of the three tools); for geospatial tasks requiring dual-tool collaboration, 15 instructions were collected, each based on one external tool and one of the 18 training tools. This study collected 15 instructions for geospatial tasks that require three tools to complete. Each instruction is based on one external tool and two of the 18 training tools. In this way, GeoITC-EvalExtend can be used to evaluate the performance of GeoTP when operating with unfamiliar GIS tools, thereby demonstrating its transfer capability.

      Table 4.  The information of three external geospatial tools.

      No. Tool name Function description Parameter name
      1 Clip This algorithm clips a vector layer using the features of an additional polygon layer. Input Layer, Overlay Layer, Output Layer
      2 Difference This algorithm extracts features from the Input layer that fall completely outside or only partially overlap the features from any of the Overlay layer(s). Input Layer, Overlay Layer, Output Layer
      3 Aspect This algorithm calculates the aspect of the Digital Terrain Model in input. Input Layer, Z Factor, Output Layer
    • Given the limited computational capacity of local devices, full-parameter fine-tuning of LLMs is not feasible. Therefore, a parameter-efficient fine-tuning strategy must be adopted to make full use of local computational resources while optimizing the model and minimizing computational and storage demands during training. In this study, the LoRA method is employed within the GeoTP framework due to its outstanding performance demonstrated in recent research on Instruction tuning. Compared with full-parameter fine-tuning of LLMs, LoRA can significantly reduce the number of trainable parameters. For the weight matrix of a large pre-trained language model, the low-rank adaptation (LoRA) method freezes the matrix and constrains its update $ \Delta W $ through a low-rank decomposition, where $ B\in {\mathbb{R}}^{d\times r} $ and $ A\in {\mathbb{R}}^{r\times k} $ are two trainable parameters, and the rank is $ r\ll \min (d,k) $. For the linear layer $ h={W}_{0}x $, its modified forward propagation is expressed as $ h $ (Eq. [3]). It is worth noting that the adopted LoRA method is scalable and can adapt to different system resources. When computational resources are significantly limited, the LoRA configuration can be adjusted by reducing the number of elements in matrix, which are then decomposed into low-rank matrices A and B. In addition, the ranks of matrices $ A $ and $ B $ can also be reduced.

      $ h={W}_{0}x+BAx $ (3)
    • To comprehensively quantify the quality of workflows generated by LLMs, this study evaluates them from two dimensions: tool-use chain accuracy and parameter accuracy. Tool-use chain accuracy (Tool Correctness and Order) assesses whether the model can plan the use of all necessary tools and whether the sequence of tool invocations is reasonable. This dimension focuses on the overall structure and logical flow of the workflow, requiring the generated workflow to be fully consistent with the expected workflow in terms of the number of tools, identifiers, and their order. Parameter accuracy verifies whether the parameters configured by the model for each tool correctly match the tool definitions, including the presence of all required parameters, conformity of parameter value data types, and compliance of value ranges. In the quantitative evaluation, the tool correctness and order accuracy (TCOA) is used to calculate the proportion of tool-use chain samples with correct tool usage order to the total number of samples (Eq. [4]). Here, $ {N}_{s} $ denotes the total number of samples, $ {N}_{t} $ represents the number of samples with a completely correct tool-use chain, and I[·] is the indicator function, which equals 1 when the condition inside the parentheses is satisfied and 0 otherwise. The symbol $\land $ denotes a logical AND operator, requiring that both the preceding and following conditions be satisfied simultaneously. The criterion for a completely correct tool-use chain $ Correc{t}_{toolchain} $ must meet the following two conditions: the number of tools $ m $ in the generated tool-use chain $ G $ must be exactly the same as that in the expected tool-use chain $ E $, and the tool names at each corresponding position must fully match those in the expected tool-use chain.

      $ \begin{aligned}TCOA&=\dfrac{{N}_{t}}{{N}_{s}}\times 100{\text{%}} \\ {N}_{t}&=\sum \limits_{j}^{{N}_{s}}\text{I}[Correc{t}_{toolchain}(j)]\\ Correc{t}_{toolchain}&=(m=n)\wedge \overset{n}{\underset{i}{\wedge }}({g}_{i}={e}_{i})\\ G&=\left[{g}_{1},{g}_{2},...,{g}_{m}\right]\\ E&=\left[{e}_{1},{e}_{2},...,{e}_{n}\right] \end{aligned} $ (4)
      ${ \begin{aligned}PA&=\dfrac{{N}_{p}}{{N}_{s}}\times 100{\text{%}} \\ {N}_{p}&=\sum \limits_{{T}_{j}}^{{N}_{s}}\text{I}[Correc{t}_{params}({T}_{j})]\\ Correc{t}_{params}&=\underset{{T}_{j}}{\wedge }\left\{(PG=PT)\wedge \underset{pg\in PG}{\wedge }\left[(TypeCheck(pg)\wedge RangeCheck(pg)\right]\right\}\\ PG&=\left[p{g}_{1},p{g}_{2},...,p{g}_{m}\right]\\ PT&=\left[p{t}_{1},p{t}_{2},...,p{t}_{n}\right] \end{aligned} }$ (5)

      The calculation of Parameter Accuracy (PA) is shown in Eq. (5), which is used to determine the proportion of tool-use chain samples with fully reasonable parameter settings to the total number of samples. Here, $ {N}_{s} $ denotes the total number of samples, and $ {N}_{p} $ represents the number of tool-use chain samples with fully reasonable parameter settings. Similarly, I[·] is the indicator function, and $ {T}_{j} $ denotes the j-th generated tool-use chain. For each tool within the tool-use chains, the parameter validity criterion $ Correc{t}_{params} $ must simultaneously satisfy the following three conditions: the parameter names $ PG $ in the generated tool-use chains must exactly match the parameter names $ PT $ in the toolset; the data type of each parameter value must comply with the tool's definition requirements; and the parameter value must fall within the tool's specified range of parameter values.

    • Selecting appropriate LLMs is essential for elucidating how instruction tuning enhances the capability of LLMs in geospatial tasks. LLaMA-3-8B has highly robust general-purpose semantic comprehension and logical mapping capabilities, and its open-source ecosystem and tool-chain are the most complete, enabling it to serve as a standard capability anchor for validating the universality of the proposed method frame. The core API documentation and community code for QGIS and most advanced open-source GIS toolchains are primarily in English, and the extremely high representational density of the LLaMA series models on English corpora helps the model quickly internalize structured QGIS tool descriptions during fine-tuning. With 8 billion parameters, it is currently the optimal size for efficient private fine-tuning via LoRA on consumer-grade GPUs, allowing it to maintain extremely high inference speed while ensuring sufficient capacity for domain knowledge. Therefore, this paper adopts the LLaMA-3-8B model. In addition, to compare the performance of different LLMs, DeepSeek-V3.1[30] and Qwen3-max[31] were evaluated on the collected evaluation dataset. This comparison helps determine whether, after applying LoRA fine-tuning using the collected instruction tuning dataset to smaller open-source models, their performance can surpass that of the most advanced open-source LLMs currently available.

      Since all available QGIS tool information has been integrated into the GeoTP model through the collected instruction tuning dataset, GeoITC during the LoRA fine-tuning process, the evaluation only requires inputting the geospatial task into the GeoTP model. However, to ensure an effective and fair comparison, a strategy suitable for general-domain baseline LLMs was designed. Specifically, for the Deepseek-V3.1 and Qwen3-max models, a system prompt was created to provide them with the collected GIS tools and their corresponding descriptions, as shown in Fig. 7.

      Figure 7. 

      Baseline models generate prompts for parameterized tool-use chains.

    • Performance Evaluation Based on the Standard Benchmark (GeoITC-Eval) Table 5 presents the performance comparison between the baseline models and GeoTP on the GeoITC-Eval evaluation dataset. To assess the models' capabilities under varying task complexities, the number of tools required for each task type (ranging from one to three) was recorded. Overall, GeoTP, trained with instruction tuning, significantly outperforms other baseline models in most cases, particularly in tool-use chain accuracy. Specifically, in single-tool tasks, GeoTP achieves a TCOA of 91.5%, far exceeding the baseline models, indicating that it has approached expert-level performance in basic tool invocation and sequential planning. This result further shows that relying solely on prompt engineering makes it difficult for general-purpose large language models to acquire geospatial analysis logic understanding and parameterized decision-making capabilities comparable to those enabled by domain-specific Instruction Tuning. Its parameter accuracy (PA) also reaches 90.0%, demonstrating the model's ability to precisely match parameter types and values, which further validates the effectiveness of parameterized tool descriptions in the instruction data. In multi-tool tasks (involving two to three tools), as task complexity increases, the performance of all models declines to some extent; however, GeoTP still maintains a leading position. In particular, in tool-use chain planning (TCOA: 65.6% for two tools and 58.3% for three tools), it significantly outperforms models of comparable scale such as LLaMA-8B and the general-purpose model Qwen3-max. This indicates that training with structured tool-use chains data enables the model to perform multi-step reasoning and tool integration more effectively. It is noteworthy that although DeepSeek-V3.1 achieves slightly higher parameter accuracy in certain tasks (e.g., three-tool PA) than GeoTP, its tool-use chains accuracy remains consistently lower.

      Table 5.  The results obtained by baselines and GeoTP on the evaluation dataset GeoITC-Eval.

      Number of tools 1 2 3
      Number of instructions 270 90 60
      LLaMA-8B CTOA 7.8%(21) 26.7%(24) 26.7%(16)
      PA 21.5%(58) 42.4%(38) 50.0%(30)
      DeepSeek-v3.1 CTOA 62.6%(169) 63.3%(57) 50.0%(30)
      PA 53.0%(143) 43.3%(39) 65.0%(39)
      Qwen3-max CTOA 54.1%(146) 54.4%(49) 28.3%(17)
      PA 50.4%(136) 43.3%(39) 51.7%(31)
      GeoTP CTOA 91.5%(247) 65.6%(59) 58.3%(35)
      PA 90.0%(243) 48.9%(44) 61.7%(37)
      Best results are in bold.

      This further demonstrates that instruction tuning data specifically designed for geospatial tasks plays a crucial role in enhancing the model's capability in professional tool utilization. In summary, GeoTP's outstanding performance in both tool-use chains planning and parameter configuration validates the effectiveness of the parameterized and structured instruction tuning dataset (GeoITC) proposed in this study for training domain-specific models. It significantly enhances the ability of open-source models to perform tool invocation and workflow generation in complex geospatial tasks.

    • In more complex task scenarios (with four and five tools), the performance of GeoTP and the baseline models was further evaluated on the GeoITC-EvalProlong dataset. It should be noted that during training, GeoTP was exposed only to tasks involving up to three tools; therefore, such tasks impose higher demands on its generalization and compositional reasoning capabilities. As shown in Table 6, GeoTP still maintains a significant advantage over the general baseline models in tool-use chain planning and parameter configuration (PA). This indicates that fine-tuning the model with the structured and parameterized instruction data (GeoITC) proposed in this study enables it to generalize the learned tool combination logic to longer task chains. It is noteworthy that although GeoTP demonstrated excellent parameter accuracy (PA: 76.7%) in the four-tool tasks, all performance metrics declined substantially when the tool-use chain was extended to five tools. This indicates that the current approach still requires improvement in planning and parameter configuration stability when dealing with highly complex workflows. Nevertheless, GeoTP's relatively superior performance, particularly its ability to maintain high parameter configuration consistency in the four-tool tasks, provides preliminary evidence that the proposed data construction method helps the model sustain better logical coherence in complex scenarios. This offers a valuable direction for developing specialized models capable of handling long-chain geospatial tasks.

      Table 6.  The experiment results of baselines and GeoTP on the evaluation dataset GeoITC-EvalProlong.

      Number of tools 4 5
      Number of instructions 30 20
      LLaMA-8B CTOA 10.0%(3) 35.0%(7)
      PA 6.7%(2) 25.0%(5)
      DeepSeek-v3.1 CTOA 20.0%(6) 40.0%(8)
      PA 40.0%(12) 35.0%(7)
      Qwen3-max CTOA 16.7%(5) 45.0%(9)
      PA 36.7%(11) 20.0%(4)
      GeoTP CTOA 66.7%(20) 50.0%(10)
      PA 76.7%(23) 50.0%(10)
      Best results are in bold.
    • Finally, to examine the model's generalization and transferability when encountering external tools unseen during training, GeoTP and the baseline models were evaluated on the GeoITC-EvalExtend dataset, as shown in Table 7. All tasks in this dataset were constructed using tools that were not employed during model training. The experimental results indicate that GeoTP outperformed all baseline models across both dimensions of tool-use chain accuracy and parameter accuracy (PA). In particular, for tasks requiring two or three tools, GeoTP achieved a CTOA of 53.3% in both cases, which is significantly higher than that of the second-best model, Qwen3-max (40.0%). This demonstrates that, after being trained with the structured and parameterized instruction data designed in this study, the model can not only correctly plan the execution order of unseen tools but also exhibits strong reasoning capabilities for tool combination. More importantly, GeoTP also demonstrates stable performance in parameter accuracy (60.0% for single-tool tasks and 40.0% for triple-tool tasks), significantly outperforming other models in multi-tool scenarios. This further indicates that the model does not rely on rote memorization of specific tool parameters learned during training. Instead, it has acquired a generalized capability to infer parameter types, constraints, and appropriate values based on tool descriptions, enabling rapid adaptation to new tools. Compared with its performance on tasks within the training distribution (Table 5), GeoTP maintains a comparable performance level on external tool tasks, with an even more pronounced advantage in multi-tool settings. These results suggest that the instruction tuning data constructed in this study effectively help the model understand the general functional logic and parameter specifications of GIS tools, thereby equipping it with genuine generalization ability in tool usage rather than limiting it to combinations of previously seen tools. This further validates the effectiveness of the data construction method proposed in this paper and the robustness of the model.

      Table 7.  The experiment results obtained by baselines and GeoTP on the evaluation dataset GeoITC-EvalExtend.

      Number of tools 1 2 3
      Number of instructions 270 90 60
      LLaMA-8B CTOA 0.0%(0) 0.0%(0) 0.0%(0)
      PA 0.0%(0) 33.3%(5) 6.7%(1)
      DeepSeek-v3.1 CTOA 20.0%(3) 13.3%(2) 6.7%(1)
      PA 46.7%(7) 13.3%(2) 6.7%(1)
      Qwen3-max CTOA 60.0%(9) 40.0%(6) 40.0%(6)
      PA 33.3%(5) 33.3%(5) 6.7%(1)
      GeoTP CTOA 60.0%(9) 53.3%(8) 53.3%(8)
      PA 60.0%(9) 46.7%(7) 40.0%(6)
      Best results are in bold.
    • This section presents two real-world geospatial task cases to illustrate how GeoTP interprets users' natural language instructions and generates logically coherent and parameter-complete executable workflows. These two cases demonstrate that the model trained with the structured and parameterized instruction data (GeoITC) proposed in this paper can not only perform accurate tool planning but also achieve precise parameter inference and adaptation.

      As shown in Fig. 8, this workflow fully demonstrates the entire process from road centerline discretization and endpoint elevation sampling to slope calculation. GeoTP first identifies the core objective as 'calculating road slope' and, based on geospatial knowledge, understands that road slope is the ratio of the elevation difference between the two ends of a road segment to its horizontal distance. Therefore, to calculate the slope of continuous road segments, the road centerline must first be divided into discrete segments, then the elevations of the start and end points of each segment must be obtained, and finally, the mathematical computation is performed. Specifically, GeoTP first uses the Split Lines by Maximum Length tool to divide the continuous road centerline into equal-length segments. The key parameter inference is reflected in GeoTP's ability to automatically set a reasonable segment length based on the typical requirements of road slope analysis, ensuring the precision of slope calculation. Next, to obtain the elevation of each road segment, GeoTP uses the Extract Specific Vertices tool and precisely sets the vertex indices to extract the start and end vertices of each line segment, generating a point layer that contains all endpoints. GeoTP then applies the Sample Raster Values tool to spatially associate the generated endpoints with the DEM data, sample elevation values, and assign a unified prefix to the elevation fields for differentiation. Finally, GeoTP employs the Field Calculator tool to link the sampled elevation values back to the original road segments and perform slope calculations. The capability of generating key parameters is demonstrated as GeoTP automatically creates the field calculation expression based on the slope formula: abs(('end_elev' - 'start_elev') / ${\$} $length * 100), where $length represents the geometric length of the segment, and 'end_elev' and 'start_elev' are the sampled elevation fields. This demonstrates GeoTP's capability for logical planning and parameter context inference in multi-step tasks.

      Figure 8. 

      Case 1: demonstration of GeoTP using four tools to calculate the road slope.

      As shown in Fig. 9, this case demonstrates the complete workflow from road network data acquisition and geometric preprocessing to turning radius calculation. GeoTP first identifies the core task as 'extracting the turning radius of roads within a specified area'. Based on general principles of geospatial analysis, it plans the steps of obtaining the raw road network data, followed by geometric refinement and specialized computation. Specifically, GeoTP first employs the OSM Downloader Tool and, according to the description of the 'specified area' in the instruction, accurately maps the spatial extent of the input file rectangle_range.shp to the tool's Download Area parameter. Subsequently, GeoTP recognizes that the raw geographic coordinate data downloaded are unsuitable for distance-based geometric calculations. Therefore, it automatically inserts the Reproject Layer step to convert the data into a projected coordinate system suitable for local-scale measurements, ensuring the accuracy of subsequent computations. Next, to calculate the turning radius at each point along the road, it is necessary to generate a dense set of sampling points along the road alignment. GeoTP selects the Points Along Geometry tool and, based on common practices in road geometry analysis, sets an appropriate sampling interval to balance computational efficiency and detail resolution. Finally, GeoTP applies the Turning Radius Calculation tool to compute the turning radius for the generated path points.

      Figure 9. 

      Case 2: demonstration of GeoTP using four tools to calculate the road turning radius.

    • In practical applications, input data may contain errors or outliers. Therefore, this section focuses on examining the robustness of GeoTP when processing data containing exceptions. Specifically, for each geospatial task in GeoTP, this study randomly selects words from the input data for analysis. Characters in correct words are randomly substituted. For example, when the input is word, the system randomly selects one character in the word and substitutes it with another character, ultimately generating ward as the input result. Multiple experiments were conducted to evaluate GeoTP's performance under different levels of input exceptions, including cases in which each task contains 2, 4, 6, 8, or 10 character errors. All experimental results are presented in Fig. 10. The data show that as the number of input exceptions increases, GeoTP's accuracy gradually declines, which is largely consistent with the expectations of this study. Notably, despite the increased number of errors in the input data, GeoTP still achieves higher accuracy than the baseline models, which fully demonstrates the robustness of the proposed model.

      Figure 10. 

      Performance of GeoTP in processing anomalous input data on the GeoITC-Eval dataset.

    • For many geospatial applications, real-time processing capability is a critical requirement. This section evaluates the computational efficiency of GeoTP. Specifically, this paper tests the computational efficiency of GeoTP on the evaluation dataset GeoITC-Eval and records the average processing time required for tasks of different complexity levels. As shown in Table 8, the time required for GeoTP to complete a task increases with task complexity, depending on the number of tools required by each task. Although the computational efficiency of GeoTP is slightly lower than that of DeepSeek-V3.1, its performance remains within an acceptable range.

      Table 8.  Average time (seconds) for GeoTP and DeepSeek-V3.1 to process geospatial tasks in GTChain-Eval.

      Number of tools 1 2 3
      DeepSeek-v3.1 5.818 7.926 11.672
      GeoTP 8.511 12.358 17.378
    • This section discusses the effect of the LoRA rank r on the performance of the GeoTP model. During training, the scale of the tunable parameter module increases with the LoRA rank r, and the dimensionality of the introduced low-rank adaptation matrices becomes larger. Theoretically, this expansion enhances the model's plasticity. To investigate the relationship between parameter scale and downstream task performance, models trained with different ranks r were systematically evaluated using the GeoITC dataset, while keeping other hyperparameters constant. The experimental results are shown in Fig. 11. Overall, under all rank r settings, the performance of GeoTP significantly surpasses that of the original unfine-tuned model and other baseline models listed in Table 5, which fully demonstrates the robustness of the proposed 'data–fine-tuning' framework. These results indicate that fine-tuning with the designed structured instruction data (GeoITC) can effectively enhance the model's core capabilities in geospatial tasks, even under varying degrees of parameter adaptation scale. Furthermore, the model's performance does not increase monotonically with a higher rank r. When r = 16, GeoTP achieved the best results in both tool-use chain and Parameter Accuracy (PA). However, when r increased to 32, the performance slightly declined. This suggests that an excessively large LoRA module size is not the key to improving performance and may even hinder the model's efficient learning of tool-use chain logic and parameter configuration rules from high-quality instruction data due to the introduction of redundant parameters. A moderate and balanced rank r provides sufficient parameter flexibility to accommodate domain-specific knowledge while avoiding overfitting, thereby achieving optimal generalization performance in downstream tasks.

      Figure 11. 

      Performance of GeoTP with different LoRA ranks on GeoITC-Eval.

      This experiment further validates the effectiveness of the data strategy presented in this paper: by carefully constructing instruction data that emphasizes tool sequences and parameter specifications, the model can efficiently learn and master complex geospatial workflow planning and execution capabilities with relatively limited parameter adjustments.

    • This section discusses the impact of the GeoITC training dataset size on model performance. By adjusting the proportion of GeoITC data used for training, the relationship between training data volume and model performance was systematically analyzed. As shown in Fig. 12, the experimental results indicate that as the amount of training data increases, the performance of the GeoTP model exhibits a consistent upward trend in both key metrics: Chain-of-Tool Accuracy (CTOA) and Parameter Accuracy (PA).

      Figure 12. 

      Performance of GeoTP with different proportions of training data on GeoITC-Eval.

      Specifically, when the training data volume increased from 0% to 100%, the CTOA improved from 14.5% to 81.2%, and the PA increased from 30.0% to 81.2%. This significant performance improvement validates the effectiveness of the data collection strategy proposed in this study. It demonstrates that training with structured tool-use chains data enables the model to better grasp the logical sequence and combination patterns among tools, while the parameterized instruction design helps the model learn how to configure parameters appropriately according to the task context, thereby maintaining execution accuracy in complex scenarios. It is noteworthy that even when only 25% of the training data were used, GeoTP achieved 43.6% in CTOA and 62.1% in PA, significantly outperforming the original model without fine-tuning and surpassing several baseline models (see Table 5). This fully demonstrates the robustness of the proposed framework, indicating that the designed instruction fine-tuning data can effectively convey the planning logic and parameter configuration knowledge of geospatial tasks even at a limited scale, thereby providing a feasible solution for model training in resource-constrained scenarios.

    • This paper proposes an effective open-source LLM-based technical framework to address three major challenges in the automatic resolution of Geospatial Tasks using LLMs: the scarcity of high-quality instruction data, the inconsistency in tool parameter configuration, and the limited performance of open-source LLMs. The research contributions are reflected in three aspects. First, a structured method is introduced for GIS tool description and parameterized instruction data generation. By incorporating a self-guidance generation strategy and leveraging domain knowledge, LLMs are guided to automatically generate diverse 'instruction–toolchain' pairs covering a wide range of scenarios and complexities. This approach particularly enhances the structured representation of parameter templates, data types, and value constraints, thereby establishing a solid foundation for the model to learn precise tool invocation logic and achieve effective Parameter Adaptation. Second, by efficiently constructing the GeoITC dataset within a simulated environment to cover scenarios involving multiple tool combinations, this study effectively addresses the scarcity of high-quality training data in the geospatial domain. GeoITC not only includes multi-level tasks such as single-tool, dual-tool, and triple-tool configurations but also emphasizes the structured representation of parameter transfer relationships among tools, providing rich supervision signals for the model to learn complex workflow logic. Furthermore, by developing a systematic evaluation benchmark, GeoITC-Eval, which incorporates multi-level task complexity and external tool generalization scenarios, the model can be comprehensively assessed in a standardized and quantifiable manner across two dimensions: chain tool operation accuracy (CTOA) and parameter accuracy (PA). This benchmark establishes a complete data foundation and evaluation framework for training and performance assessment of intelligent planning models for GIS tools. Finally, based on the LLaMA-3-8B model, the LoRA technique is applied to inject GIS tool usage capabilities, resulting in a specialized GIS tool planning model named GeoTool Planner (GeoTP). Experimental results show that GeoTP significantly outperforms baseline models across all evaluation metrics. For known tool tasks, it achieves a single-tool tool-use chain accuracy (CTOA) of 91.5% and a parameter accuracy (PA) of 90.0%. When handling tasks involving previously unseen external tools or longer tool-use chains (four to five tools), the model still maintains over 53.3% CTOA and up to 76.7% PA, demonstrating reliable generalization and strong planning ability for complex tasks. Ablation studies further confirm the critical role of structured parameter data and appropriately chosen LoRA rank in enhancing model performance.

      In summary, this study not only introduces GeoTP, a high-performance and privately deployable intelligent planning tool for GIS tools, but also establishes a comprehensive methodological framework covering data generation, model training, and performance evaluation. Experimental results demonstrate that fine-tuning with carefully designed parameterized instruction data can significantly enhance open-source LLMs, enabling them to accurately master the planning logic and parameter adaptation capabilities required for complex geospatial workflows. The outstanding performance of GeoTP on known tasks, along with its strong generalization ability when handling unseen tools and extended tool-use chains, provides a solid technical foundation for developing low-cost, highly reliable, and privacy-preserving intelligent planning models for GIS tools. Although GeoTP demonstrated outstanding performance in the experiments, the evaluation results indicate that there is still room for improvement in handling extremely complex tasks, such as long chains involving multiple tools. In addition, reducing the computational overhead of model deployment and enabling dynamic learning from large-scale tool repositories remain key challenges in practical applications. Future work will focus on exploring techniques such as retrieval-augmented generation to allow the model to dynamically invoke larger toolsets without retraining, as well as optimizing the model architecture and quantization strategies to enhance inference efficiency while maintaining performance.

      • This research is supported in part by the National Natural Science Foundation of China (Grant Nos 42130112, 42371479).

      • The data used in this study were provided by public data platforms and collected from publicly available datasets. Therefore, no ethics committee approval was required for this study.

      • The authors confirm contribution to the paper as follows: study conception and design: Tang J, Yang J; data collection: Lin D, Li Z, Yang G; analysis and interpretation of results: Yang J, Tang J; draft manuscript preparation: Tang J, Yang J, Zhang B, Fang L. All authors reviewed the results and approved the final version of the manuscript.

      • The datasets generated during and/or analyzed during the current study are available from the corresponding author upon reasonable request.

      • The authors declare that they have no conflict of interest.

      • Copyright: © 2026 by the author(s). Published by Maximum Academic Press, Fayetteville, GA. This article is an open access article distributed under Creative Commons Attribution License (CC BY 4.0), visit https://creativecommons.org/licenses/by/4.0/.
    Figure (12)  Table (8) References (31)
  • About this article
    Cite this article
    Tang J, Li Z, Yang J, Lin D, Zhang B, et al. 2026. GeoTP: parameter-level intelligent planner of GIS tool-use chains based on instruction tuning of large language models. The Knowledge Engineering Review 41: e012 doi: 10.48130/ker-0026-0013
    Tang J, Li Z, Yang J, Lin D, Zhang B, et al. 2026. GeoTP: parameter-level intelligent planner of GIS tool-use chains based on instruction tuning of large language models. The Knowledge Engineering Review 41: e012 doi: 10.48130/ker-0026-0013

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return