$ Hi there, I'm
Yuhe Wu
$ Sometimes I █
Current Research

LLM in Finance
Fintech Thrust @ HKUST(GZ)
LLM evaluation, LLM self-evolution, and agent cognition for financial decision-making

Fintech
Fintech Lab @ DUFE
Applied AI, time-series forecasting, and explainable models for financial systems
Education

Ph.D. in Fintech (Incoming)
HKUST(GZ) · 2026.09 ~
Supervised by Prof. Guang Zhang

B.S. in Economic Statistics
DUFE · 2022 – 2026
Supervised by Prof. Zhuang Liu and Prof. Xu Qiang
I am an undergraduate student in Economic Statistics at Dongbei University of Finance and Economics (DUFE), supervised by Prof. Zhuang Liu. I will be joining the Hong Kong University of Science and Technology (Guangzhou) as a Ph.D. student in Fall 2026, supervised by Prof. Guang Zhang.
My research focuses on Large Language Models (LLMs) in Finance and Social Sciences, spanning LLM evaluation, LLM self-evolution, and agent cognition. I have published papers at venues including ACL 2026, EMNLP 2026, KDD 2025, and Annals of Operations Research, with additional manuscripts under review at NeurIPS 2026, KDD 2026, AAAI 2027, and Management Science. I also serve as a reviewer for leading venues such as ACL, KDD, NeurIPS, AAAI, and IJOC.
(* equal contribution · † corresponding author)


PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations
Yuhe Wu, Guangyu Wang, Yuran Chen, Jiatong Zhang, Yutong Zhang, Yujie Chen, Jiaming Shang, Guang Zhang†, Zhuang Liu†
In 2025, 2.6% of CS conference papers were found to contain suspected hallucinated citations, nearly 9 times the rate from the previous year. ICLR 2026 now treats hallucinated references as academic ethics violations, where a single instance can lead to rejection. Whether LLMs are a trustworthy tool for research has become an unavoidable question. But fixing hallucinations requires understanding where they come from. When a model gives a wrong answer, the reasons can be entirely different: it may lack the relevant knowledge altogether, it may have memorized incorrect facts, it may have the right knowledge but fail in reasoning, or it may reason correctly but ignore the user's constraints. These four types of errors originate from distinct stages of the generation pipeline, and each requires a different fix, yet existing evaluations mix all queries together and only score final outputs, making them indistinguishable. More critically, fixing one type often worsens another: strengthening instruction following can hurt reasoning, while injecting knowledge can cause forgetting. PRISM decomposes hallucinations along three generation stages into four independently testable dimensions, providing systematic diagnosis across 9,448 instances and 24 LLMs, making both what to fix and how to fix it actionable.


BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications
Jianing Hao*, Yuhe Wu*, Yuanjian Xu*, Shichang Meng, Shuai Yuan, Wei Zeng, Zixuan Zhang, Guang Zhang†
In discussions with over ten leading financial institutions, a recurring question emerged: LLMs are powerful, but which business scenarios can they reliably support, and what foundational capabilities underpin these applications? Existing benchmarks cannot answer this. They either cover narrow tasks or focus only on surface-level accuracy, consistently lacking a causal chain from foundational capability to business performance. BizCompass was built to address this gap: it covers four foundational disciplines at the knowledge level, including finance, economics, statistics, and operations research, and structures tasks around three business roles at the application level, including analyst, trader, and consultant, forming a dual-axis evaluation framework that not only measures how well models perform, but diagnoses why they fall short.


BizSage: A Self-Evolving Multi-Agent Framework for Business Research with Efficient Knowledge Retrieval
Yuhe Wu, Guangyu Wang, Jiaxin Liu, Guang Zhang†
While multi-agent systems based on large language models have shown promise in automating the progressive workflow of academic research, extending them to economics and business research presents two challenges. First, existing methods mostly retrieve at the paper level, yet the evidence needed for research tasks is often distributed across different sections, creating a granularity mismatch that hinders retrieval coverage and precision. Second, these fields demand strict empirical rigor, yet current systems provide limited mechanisms for learning from evaluation feedback. We present BizSage, a multi-agent framework combining corpus-level fine-grained retrieval with quality-driven self-evolution. We build a Lateral Knowledge Graph (LKG) by merging section-level knowledge graphs and apply Personalized PageRank (PPR) to surface semantically relevant and structurally important sections. Seven specialized agents collaborate under a Meta-Review self-evolution mechanism that distills failure modes from evaluation traces into reusable strategies. On a benchmark spanning four domains and three tasks, BizSage ranks first on the majority of metrics, achieves pairwise win-rates above 60% against six baselines, and produces zero hallucinated citations.


Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
Yuhe Wu, Guangyu Wang, Yujie Chen, Jiatong Zhang, Yuran Chen, Yutong Zhang, Xiyin Cheng, Wenpeng Cao, Guang Zhang†, Zhuang Liu†
People increasingly turn to large language models for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this through single-turn judgments or pressure-laden rebuttals, assumptions that poorly match real-world guidance seeking. We introduce narrative captivity, a failure mode in which a model treats an unopposed one-sided account as complete and aligns with the narrator's interpretation without seeking missing perspectives. We build a benchmark of 5,078 interpersonal-conflict scenarios spanning six moral dimensions. Across 17 LLMs, narrative captivity is widespread: end-state judgments under multi-turn narration shift by 25 percentage points on average beyond the matched single-turn baseline. Stage-level analysis identifies preference optimization as a major contributor, while four inference-time strategies provide only partial mitigation.


Beyond When: How Root Content Propagates as Information Cascade Trees
Guangyu Wang*, Yuhe Wu*, Zhaonan Wang†
Simulating information cascades, the traces formed as content propagates across networks, is central to understanding collective behavior online. Existing approaches either only predict aggregate cascade size or synthesize flat cascade sequences via Temporal Point Processes, which model each propagation action as an event. None can generate complete cascade structures or leverage root content. We propose CasT² (Cascade Simulation on Time-ordered Trees), a task that conditions on root content and jointly infers when each event occurs and how the propagation path unfolds, recovering the complete tree structure. We design a framework extending flow matching to tree space with a depth-aware probability path that first constructs the trunk and then expands peripheral branches, while a Transformer backbone iteratively refines cascades through insertions and deletions preserving tree validity. A large language model analyzes root content through established propagation theories, distilling structured semantic profiles that condition generation. We also contribute CasT²-1.4M, a benchmark comprising 1.4M cascades across commenting, reposting, and citation domains.

Ph.D. Student (Incoming)
Starting Ph.D. journey in Fintech at Hong Kong University of Science and Technology (Guangzhou), focusing on LLM Agents, evaluation, and finance applications.

Research Assistant at Fintech Thrust
Research Assistant under Prof. Guang Zhang. Papers accepted at ACL 2026 (Main & Findings) and EMNLP 2026 (Main & Findings), spanning LLM evaluation, multi-agent systems, and socially embedded decision making.

Group Core Member
Core member of the research group supervised by Prof. Xu Qiang at the National Accounting Center.

Research Group Leader
Led the Fintech Lab research group under Prof. Zhuang Liu. Published in Annals of Operations Research and KDD-UMC 2025, with further work accepted at ACL 2026 and EMNLP 2026.

Business Admin → Economic Statistics
Transferred from Business Administration to Economic Statistics. Started exploring data science, statistical modeling, and quantitative methods.
12 awards spanning 4 categories
ACL 2026 Diversity & Inclusion (D&I) Subsidy
ACL 2026
KDD-25 Undergraduate Scholarship
KDD 2025
Dashang Group Academic Research Scholarship (TOP 1%)
DUFE
Excellent Innovation & Entrepreneurship Team Scholarship (TOP 1%)
DUFE
Advanced Individual in Academic Competitions (TOP 1%)
DUFE
National 2nd Prize, 9th Financial Innovation & Entrepreneurship Competition
National
Bronze Award, China International Innovation Competition
National
National 3rd Prize, 10th Student Statistical Modeling Competition
National
Top 100, ICBC Cup National Fintech Innovation Competition
National
First-Class Scholarship (TOP 2%)
DUFE
Merit Student (University Level)
DUFE
Excellent Communist Youth League Member
DUFE




