Risk-Adjusted Harm Scoring for Automated Red Teaming of LLMs in Financial Services

By Fabrizio Dimino, Bhaskarjit Sarmah, Stefano Pasquali

Rating

1214
Battle Count: 74

Relevance

2/10
The paper is primarily about LLM safety evaluation in financial services rather than quantitative trading strategies. However, it has indirect relevance: (1) financial institutions deploying LLMs for trading support, investment research, and compliance analysis need safety guarantees; (2) the taxonomy covers market abuse, insider trading, and market manipulation which are directly relevant to trading compliance; (3) the framework could be used to stress-test AI-assisted trading systems against adversarial manipulation. The paper does not address trading strategy development, portfolio optimization, or market prediction.

Implementation Complexity

7/10
Implementation requires: (1) deploying multiple open-weight LLMs as judges (gpt-oss-120b, Qwen3-235B, Llama-3.3-Nemotron-49B) and target models; (2) constructing the 989-prompt benchmark with domain expertise; (3) implementing the multi-turn adaptive red-teaming loop with structured feedback; (4) computing RAHS with severity grading, disclaimer detection, and entropy calculation; (5) running statistical validation (bootstrap, McNemar, Wilcoxon). The ensemble judging and multi-turn pipeline add significant computational and engineering overhead. However, all components use open-weight models and the methodology is well-specified.

Reproducibility

4/5
All target models are open-weight (8B-72B parameters). Judge prompts, attacker prompts, and system prompts are fully disclosed in the appendix. The benchmark taxonomy (989 prompts) is described with representative examples. However, the full FinRedTeamBench dataset is not explicitly stated as publicly downloadable. Hyperparameters (alpha=0.5, gamma=0.2, lambda=0.1) are clearly specified. Statistical tests (bootstrap, McNemar, Wilcoxon) are well-documented.

About this paper

Methodology: Risk-Adjusted Harm Score (RAHS) with Adaptive Multi-Turn Red-Teaming. Problem types: Classification, Risk Management, Anomaly Detection, Natural Language Processing, Ranking.

The interactive Everscope explorer (charts, battles, favorites) loads below.