FinReflectKG - HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems

By Mahesh Kumar, Bhaskarjit Sarmah, Stefano Pasquali

Rating

1183
Battle Count: 54

Relevance

3/10
The paper is primarily focused on hallucination detection in KG-augmented financial QA systems over SEC 10-K filings. While not directly about quantitative trading strategies, it has indirect relevance: (1) reliable extraction of financial metrics from filings is foundational for fundamental analysis and factor models; (2) hallucination detection ensures data integrity for trading signals derived from textual sources; (3) the benchmark could validate AI systems used in investment research pipelines. However, the paper does not address market prediction, portfolio optimization, or trading strategy development directly.

Implementation Complexity

6/10
Moderate complexity. Requires access to multiple LLMs (Qwen-3-235B, GPT-OSS-120B, Lynx-8B), NLI models (DeBERTa-v3), specialized classifiers (Vectara, LettuceDetect), and embedding models (Qwen-0.6B, Stella-400M). The KG triplet extraction pipeline (FinReflectKG) with reflection-driven extraction adds complexity. Threshold optimization across 0.05-0.95 range for multiple metrics requires careful implementation. Statistical tests (Cochran's Q, McNemar with Holm-Bonferroni) and bootstrap CIs add analytical complexity. Hardware requirements range from vLLM endpoints to Apple M4 Pro GPUs.

Reproducibility

4/5
Dataset (755 annotated examples), evaluation scripts, and detector implementations promised to be publicly available on GitHub upon publication. Complete prompts for generation and evaluation provided in appendices. Hardware and hyperparameter specifications detailed in Table 1. Statistical tests (Cochran's Q, McNemar with Holm-Bonferroni) and bootstrap CIs (10,000 iterations) specified. However, the FinReflectKG pipeline details (graph schema, storage backend, ingestion) are not fully specified, and the internal annotation web app is not publicly available.

About this paper

Methodology: FinBench-QA-Hallucination Benchmark with Controlled KG Noise Evaluation. Problem types: Classification, Natural Language Processing, Anomaly Detection, Information Extraction, Graph Learning.

The interactive Everscope explorer (charts, battles, favorites) loads below.