FinanceHarness: Autonomous Financial Deep Research Framework

By Yijia Xiao, Rujun Han, Yanfei Chen, Zifeng Wang, Ke Jiang, Zhongying CuiZhu, Vishy Tirumalashetty, Wei Wang, Burak Gokturk, Tomas Pfister, Chen-Yu Lee

Rating

1221
Battle Count: 50

Relevance

5/10
The paper is primarily about financial research and report generation rather than direct trading strategies. However, it is highly relevant to quantitative trading in several ways: (1) The PIT sandbox and point-in-time evaluation directly address look-ahead bias, a critical concern in quantitative backtesting. (2) The harness includes valuation tools (DCF, WACC, comparables), risk analytics (VaR, beta, correlation, duration, convexity), and option pricing that are directly applicable to quantitative workflows. (3) The entity graph and situation mining could inform factor construction and event-driven strategies. (4) The benchmark tests forward-looking reasoning (post-cutoff), which is essential for predictive trading models. (5) The paper explicitly distinguishes itself from trading-focused work (TradingAgents, Trading-R1) but provides infrastructure that could support research-driven trading. The relevance is moderate-to-high for the research/analysis side of quantitative trading rather than execution/optimization.

Implementation Complexity

8/10
The system is highly complex with multiple layers: (1) A 100+ million article web corpus with publication-date extraction and FAISS IVF-SQ8 indexing. (2) A finance entity graph with 4.37M edges extracted via LLM. (3) A multi-stage benchmark construction pipeline (situation mining, cutoff selection, question generation, quality filtering, ILP balancing, expert annotation). (4) A layered agent harness with model/serving, runtime, tool surface, modes, skills, grounding, and recovery layers. (5) GRPO training infrastructure. (6) LLM-as-judge evaluation with 5-tier rubric scoring. (7) PIT access control enforcement. The code is available but reproducing the full pipeline requires significant infrastructure (embedding servers, FAISS indices, multiple LLM APIs, expert annotation workflows). The harness itself is modular but the end-to-end system is very complex.

Reproducibility

4/5
Code is publicly available at GitHub (https://github.com/Yijia-Xiao/FinanceHarness). Leaderboard is available at https://financegym.github.io/. The benchmark questions and report submission code are released, but the underlying 100+ million article corpus is NOT publicly released. Grading runs through a private leaderboard with private rubrics to preserve scoring integrity. The PIT sandbox, FAISS index, and evaluation protocol are well-documented. However, the full corpus and some internal infrastructure details are not available, limiting full reproduction.

About this paper

Methodology: FinanceHarness. Problem types: Natural Language Processing, Information Retrieval, Report Generation, Multi-hop Reasoning, Time Series Forecasting, Causal Inference, Risk Management, Portfolio Optimization, Structured Prediction.

The interactive Everscope explorer (charts, battles, favorites) loads below.