Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements

By Arianna Miola, Bruno Spaccavento, Lorenzo Silotto, Marco Bianchetti, Luca Cagliero

Rating

1718
Battle Count: 70

Relevance

2/10
The paper focuses on financial document question answering and regulatory compliance rather than quantitative trading strategies. While the extracted financial indicators (CET1, Total Assets, liquidity reserves) could inform fundamental analysis or risk management inputs for trading systems, the paper does not address trading signals, portfolio construction, or market prediction. The primary contribution is a benchmark dataset and RAG pipeline evaluation methodology.

Implementation Complexity

7/10
The multi-stage RAG pipeline involves: PDF processing with Azure Document Intelligence, custom hierarchical chunking, contextual enrichment via GPT-4.1 (with dynamic heading-level adjustment for large files), vector indexing in Qdrant, optional cross-encoder reranking, and LLM generation. Integration of multiple proprietary APIs, handling of very long documents (avg 198k words), and managing the contextual enrichment process for files exceeding model context windows add significant engineering complexity. However, the individual components are well-documented and modular.

Reproducibility

3/5
The benchmark dataset is publicly available on GitHub. However, the pipeline relies on proprietary APIs (OpenAI GPT-4o, GPT-o1-high, VoyageAI, Cohere, Azure Document Intelligence) which may change over time. The stochastic nature of reasoning models (GPT-o1-high) affects result reproducibility. Experiments were conducted in late 2024–early 2025 with models available at that time, limiting temporal reproducibility.

About this paper

Methodology: Multi-stage RAG Pipeline with Contextual Chunk Enrichment. Problem types: Natural Language Processing, Information Retrieval, Question Answering, Ranking.

The interactive Everscope explorer (charts, battles, favorites) loads below.