Rating
1255
Battle Count: 50
Relevance
3/10
The paper uses financial decisions (weekly up/down stock calls on Warsaw Stock Exchange) as a probe instrument for measuring model disposition, not for developing trading strategies. The explicit finding is that no arm shows economic skill — apparent edges are regime beta, not alpha. However, the findings are relevant to quantitative trading insofar as LLM-based agents are increasingly deployed in trading pipelines: abliterated 'uncensored' models introduce systematic optimism bias and altered confidence that could distort automated decision-making. The paper is a cautionary result for anyone using LLMs as decision-makers in financial contexts, but does not propose or evaluate trading models.
Implementation Complexity
6/10
The experimental design is methodologically rigorous but conceptually straightforward: serve base and abliterated checkpoints via vLLM, replay frozen upstream through a decision prompt, collect JSON outputs, compute disposition metrics, run cluster bootstrap. The complexity lies in the provenance auditing (weight hashing, token-for-token prompt verification, chat template checking), the multi-agent pipeline infrastructure, and the preregistration discipline. Reproducing the generation step requires access to the proprietary upstream pipeline; reproducing analyses requires only the released decision database and scripts.
Reproducibility
5/5
Exceptionally high reproducibility: frozen prompt (SHA-256 recorded), preregistration with tag freeze-v2-2026-07-05, all 21,600 decision-level outputs released as public_snapshot.db, analysis scripts deterministic and read-only, Zenodo DOI (10.5281/zenodo.21314839), GitHub repository with code/paper/data, PROVENANCE.md with full weight hashes and config diffs. Only the upstream analyst briefs and debate transcripts are not redistributed (proprietary pipeline), but all analyses regenerate from public data.
About this paper
Methodology: Controlled Cross-Family Abliteration Comparison with Frozen Multi-Agent Pipeline. Problem types: Classification, Natural Language Processing, Causal Inference, Risk Management.
The interactive Everscope explorer (charts, battles, favorites) loads below.