-
Figure 1.
Framework for the identification of direct targets of natural products. Natural products or bioactive compounds are isolated from natural sources and subjected to activity-guided purification, with tiliroside shown as an illustrative example. Candidate targets can first be nominated using omics-based analyses, artificial intelligence-assisted structure prediction, molecular docking, and network pharmacology. Experimental target-fishing strategies are subsequently divided into label-free and label-based approaches. (a) Label-free strategies detect ligand-induced changes in protein stability, protease susceptibility, or local conformation. DARTS identifies proteins protected from proteolysis after ligand binding; CETSA evaluates ligand-induced alterations in thermal stability; TPP extends thermal stability analysis to the proteome scale; and LiP-MS detects changes in local protease accessibility and maps ligand-responsive protein regions. (b) Label-based strategies use activity- or affinity-based chemical probes, photoaffinity labeling, click chemistry, and biotin-streptavidin enrichment to capture compound-interacting proteins. Depending on the target-fishing strategy, candidate proteins are recovered by affinity enrichment or collected from soluble or proteolytic fractions and subsequently identified by immunoblotting or LC-MS/MS. Candidate targets should be further evaluated using orthogonal binding and functional validation assays.
-
Figure 2.
Framework for validating direct binding and functional causality. (a) Biophysical and cellular target-engagement assays offer complementary evidence for compound–protein interactions. SPR, MST, ITC, and NMR assess binding kinetics, affinity, thermodynamics, and structural changes, respectively. CETSA measures target engagement in cells or lysates via thermal stability shifts, but should be combined with purified-protein biophysical assays to confirm direct binding. (b) Orthogonal biochemical validation uses affinity pull-down, competition assays, functional readouts, and site-directed mutagenesis to confirm binding specificity and identify critical residues. Competition by excess unmodified compound supports specific binding, while loss of binding or activity after mutation indicates a defined binding site.
-
Figure 3.
Single-cell and spatial multi-omics platforms for resolving the pharmacological mechanisms of natural products. (a) scRNA-seq identifies drug-responsive cell populations and transcriptional states after quality control, normalization, integration, clustering, and annotation. (b) scATAC-seq and single-cell multiome approaches profile chromatin accessibility, transcription-factor motifs, and regulatory links between accessible elements and gene expression. (c) Spatial transcriptomics and spatial multi-omics retain tissue architecture and map treatment-responsive pathways or cellular interactions to defined pathological niches. These modalities provide contextual and mechanistic evidence but do not independently establish direct compound-target binding.
-
Figure 4.
AI-assisted candidate prioritization as an upstream component of natural product research. (a) Conventional screening uses prior knowledge and large experimental libraries to select compounds, followed by cellular and animal validation. (b) AI-assisted screening integrates chemical, ADMET, target, disease, and multi-omics information using machine-learning or deep-learning models to prioritize candidates before experiments. This workflow reduces the initial search space but does not replace target-engagement or efficacy validation.
-
Figure 5.
AI-assisted drug-target identification and target-based drug screening. (a) Drug-based target prediction, which integrates 1D, 2D, and 3D similarity information of drugs together with multi-source data such as known ligands, protein sequences, binding-pocket features, and knowledge graphs to identify potential targets, followed by molecular docking to validate candidate hits. (b) Target-based drug screening, which uses protein sequence and three-dimensional structural information, combined with SMILES representations, molecular fingerprints, and other features, to build predictive models that score and rank candidate compounds, thereby enabling efficient virtual screening and prioritization of promising drug candidates.
-
Figure 6.
Evidence-informed closed-loop framework for natural product target discovery.
-
Method/readout Label and biological setting Main advantages Main limitations Best use and required validation Affinity or biotin pull-down Requires an immobilized or tagged ligand; lysates or intact-cell-compatible probes Direct enrichment; compatible with competition and quantitative proteomics Probe modification may alter permeability or affinity; matrix and abundant protein background Unbiased capture when a validated probe is available; confirm with free-compound competition and an orthogonal binding assay Photoaffinity/click chemistry Minimal photo-crosslinker and clickable handle; usually intact cells or lysates Captures weak or transient interactions; preserves spatial proximity Photochemical background and crosslinking-radius effects; synthesis and controls are demanding Transient or low-affinity interactions; require inactive-probe, no-UV, and competition controls Degradation-based profiling Ligand incorporated into a degrader or molecular-glue workflow; intact cells Event-driven signal amplification; can reveal low-occupancy binders Depends on ternary-complex geometry, E3 expression, and proteasome competence Functional target nomination when degradation chemistry is feasible; validate direct binding and degradation dependence DARTS No ligand modification; native lysates and limited proteolysis Simple, inexpensive, and compatible with chemically intractable ligands Biased by protein abundance, protease accessibility, and indirect conformational changes Focused or discovery-scale screening; validate by dose-dependent protection and a biophysical assay CETSA/TPP No ligand modification; lysates, intact cells, tissues; immunoblot or MS readout Measures engagement in a biologically relevant environment; proteome-wide with TPP Not all binders shift thermal stability; complexes and downstream effects can produce indirect shifts Cellular target engagement and proteome-wide deconvolution; combine with purified-protein binding and genetics PELSA/LiP-MS No ligand modification; peptide-level proteolysis in native mixtures Detects local structural responses and can suggest responsive protein regions Peptide detectability and protease accessibility limit coverage; responsive regions are not necessarily binding sites Mapping local conformational responses; confirm by mutagenesis or structural analysis SPROX/TRAP No ligand modification; oxidation or residue-accessibility readout in complex proteomes Orthogonal physicochemical evidence; sensitive to local folding or accessibility changes Requires appropriate reactive residues and specialized quantitative proteomics Complementary discovery when thermal or proteolytic shifts are weak; validate direct engagement SIP/pHDPP/
DiffPOPNo ligand modification; solvent-, pH-, or gradient-induced precipitation Scalable and applicable to structurally diverse compounds Solubility changes may be indirect and are influenced by protein physicochemical properties Proteome-wide prioritization; require orthogonal engagement and functional testing SPR-MS or target-immobilized fishing Immobilized protein or ligand; fractions, extracts, or lysates Links real-time binding detection with MS identification; useful for trace constituents or complex mixtures Immobilization can alter conformation; mass transport and nonspecific surface binding require controls Ligand fishing or target fishing in mixtures; confirm affinity, activity, and cellular relevance Table 1.
Decision-oriented comparison of representative natural product target-fishing strategies.
-
Method Core readout Main strengths Main limitations Primary evidential role SPR Surface refractive-index change during binding Real-time Ka, Kd, and KD; low sample use Immobilization, mass transfer, and nonspecific binding Direct binding and kinetics ITC Heat change during solution-phase titration KD, stoichiometry, enthalpy, and entropy High sample demand; weak or low-heat interactions are difficult Direct binding and thermodynamics FP/HTRF Binding-dependent rotation or time-resolved energy transfer Homogeneous, scalable, and suitable for competition Requires tracers or paired reagents; interference risk Screening and displacement evidence MST Binding-dependent thermophoretic movement Low sample use; broad affinity range; complex matrices possible Fluorescence, adsorption, and aggregation artifacts Orthogonal affinity measurement NMR Chemical-shift or relaxation changes Weak-binding detection and interaction-surface mapping Protein size, labeling, solubility, and instrument access Direct binding and residue-level information X-ray/cryo-EM Atomic or near-atomic complex structure Binding pose, pocket geometry, and critical contacts Sample preparation and conformational-state limitations Structural confirmation and mutation design Table 2.
Complementary methods for validating small-molecule-protein interactions.
-
Research question Recommended modality and tools Expected output Key caution Which cell populations respond to treatment? scRNA-seq; Seurat/Scanpy; harmony or scVI when integration is required Cell-type abundance, transcriptional states, and sample-level treatment effects Dissociation and batch effects can mimic cell loss or induction; use biological replicates Does treatment induce a cell-state transition? scRNA-seq time course; monocle or Slingshot Pseudotime ordering, branch points, and state-associated genes Pseudotime is not direct lineage or chronological proof Which cells communicate after treatment? scRNA-seq or spatial data; CellChat/CellPhoneDB; NicheNet for ligand-to-target links Altered ligand-receptor networks and predicted receiver-cell programs Expression-based interactions require protein-level and perturbational validation Is chromatin regulation altered? scATAC-seq or RNA + ATAC multiome; ArchR/Signac Accessible elements, motif activity, peak-to-gene links, and regulatory programs Sparse peak counts and inferred links can reduce robustness How should modalities be integrated? Matched or unmatched multi-omics; Seurat WNN, MOFA+, totalVI or graph-based models Shared latent states, modality-specific factors, and cross-modal regulatory links Integration can obscure modality-specific biology; benchmark against unimodal results Where does the response occur in tissue? Spatial transcriptomics/proteomics; cell2location, Tangram or SPOTlight Spatial niches, cell-state maps, and region-specific interactions Spot resolution, deconvolution assumptions, and histological registration affect inference Table 3.
Question-driven selection of single-cell and spatial multi-omics strategies.
-
AI category Typical inputs and outputs Strength for natural products Main limitation and validation requirement Ligand-based prediction Fingerprints, SMILES, molecular graphs, known compound-target pairs→ranked targets Fast reverse screening and off-target nomination when related ligands are annotated Weak extrapolation beyond known chemical space; use scaffold-aware validation and direct-binding assays Structure-based prediction Protein structures or pockets and ligand conformers→docking poses, scores, or target ranks Can suggest binding sites and rationalize stereochemical interactions Protein flexibility and scoring errors; validate affinity, pose-dependent mutations, and cellular engagement GNN/molecular representation learning Molecular and interaction graphs→learned embeddings and interaction probabilities Captures nonlinear structural and network features Data leakage and opaque features; require external or prospective testing Knowledge-graph/network reasoning Compound-target-disease-pathway relations→mechanistic paths or candidate targets Integrates sparse, heterogeneous evidence and supports polypharmacology hypotheses Database popularity bias and correlation without causality; trace evidence and validate each edge experimentally Perturbation/omics modeling Drug-response, CRISPR, transcriptomic, proteomic, or single-cell signatures→target or pathway ranking Links compounds to context-specific cell states and phenotypes May prioritize downstream effectors rather than binders; combine with target-fishing and engagement assays Multimodal/foundation models Chemical, structural, omics, imaging, and text data→joint representations and multiple predictions Potential to integrate natural-product structure with cell-context and disease knowledge Modality imbalance, interpretability, and domain shift; benchmark each output and perform prospective validation Table 4.
Task-oriented comparison of AI approaches for natural product target research.
Figures
(6)
Tables
(4)