2026 / Conference paper
TargetSage: Identifying Therapeutic Target Genes with Interpretable and Robust LLM Reasoning
NeurIPS 2026
Abstract
Therapeutic target identification—determining which genes or proteins to modulate for a desired clinical effect—is a foundational step in drug discovery. However, existing computational methods that rank candidates at genome scale fail to capture mechanistic knowledge locked in the biomedical literature, lack interpretability, and suffer from the false-negative problem by treating unlabeled but potentially druggable genes as definite negatives. We propose TargetSage, an LLM reasoning agent for therapeutic target identification built on positive-unlabeled (PU) learning. Our approach consists of three modules and provides an evidence-grounded, continuously updatable foundation for interpretable target prioritization at genome scale. Specifically, first, Agentic Profiling enriches each gene with LLM-gathered biomedical evidence; second, Guided Reasoning produces gene-specific rationales and explicit reasoning attributes through a GRPO-trained policy rewarded by downstream Adjusted F1; third, PU-Aware Scoring produces genome-wide druggability scores with an LLM-informed class-prior. Evaluated on 15 benchmark tasks spanning 19,032 human protein-coding genes, TargetSage consistently outperforms ten baselines, surpassing the strongest by 37% in macro-average Adjusted F1 and generalizing to independent held-out benchmarks. Using only the training data collected in 2021, TargetSage outperforms all ten baselines at ranking genes that gain clinical approval by 2025. On an independent validation set of Phase II/III clinical candidates, it further achieves a 51% higher enrichment odds ratio at the top-25 shortlist than the runner-up baseline.