FinTradeBench: A Financial Reasoning Benchmark for LLMs

By Yogesh Agrawal, Aniruddha Dutta, Md Mahadi Hasan, Santu Karmaker, Aritra Dutta

Rating

1307
Battle Count: 50

Relevance

6/10
The paper directly evaluates LLM reasoning over trading signals (momentum, volatility, EMA, RSI, MACD, drawdowns) and company fundamentals, which are core inputs to quantitative trading strategies. However, it is a diagnostic benchmark rather than a trading system—it evaluates whether LLMs can reason about these signals correctly, not whether they can generate profitable strategies. The finding that RAG degrades trading-signal reasoning (models cannot parse raw OHLCV tables) is highly relevant for quant teams considering LLM integration. The hybrid reasoning gap (fundamentals vs. market dynamics conflict) mirrors real quant challenges. The paper does not propose trading algorithms, backtesting frameworks, or execution strategies.

Implementation Complexity

7/10
The full pipeline involves: (1) SEC HTML parsing with hierarchical indexing and metadata injection; (2) Dual-track RAG with FAISS dense retrieval, BM25 sparse matching, cross-encoder reranking, and temporal query paths; (3) TELeR-based multi-prompt generation with intra-model self-filtering; (4) Automated numerical auditing against structured financial knowledge bases; (5) Human-LLM judge calibration with iterative prompt engineering; (6) Scaling across 101 companies × 10 years. Requires GPU infrastructure, multiple LLM APIs, financial data pipelines, and domain expertise for seed question authoring. The evaluation framework with paired t-tests and multi-dimensional scoring adds further complexity.

Reproducibility

4/5
The authors commit to releasing the full dataset (questions, gold answers, golden indicators, derived signals), complete codebase for data processing, retrieval, generation, and evaluation under CC BY 4.0 license. Versioned benchmark with frozen releases ensures comparability. TELeR prompt taxonomy, evaluation scripts, and model versions are documented. However, raw SEC filings are provided for only 7 representative companies in the supplement; full reconstruction requires downloader scripts. Compute resources (NVIDIA 4070, H100s, cloud APIs) are specified but exact configurations may vary.

About this paper

Methodology: Calibration-then-Scaling Framework. Problem types: Natural Language Processing, Zero-shot Learning, Classification, Time Series Forecasting.

The interactive Everscope explorer (charts, battles, favorites) loads below.