Martingale Doppelgänger-Eval: An Identification Framework for Auditing Candlestick Understanding in Vision-Language Models

By Ziyao Wang

Rating

1436
Battle Count: 50

Relevance

4/10
The paper is highly relevant to quantitative trading as a diagnostic/audit tool for VLM-based financial analysis systems. It demonstrates that current VLMs cannot reliably interpret candlestick evidence and instead rely on trend shortcuts, which has direct implications for any trading system incorporating VLM chart analysis. However, the paper explicitly states it does not target tradability or propose trading strategies. The martingale-null framework and identification theorems provide rigorous foundations for evaluating whether AI-generated technical analysis narratives are grounded in visual evidence. The findings (negative beta_E for commercial APIs, trend dominance) are critical warnings for quantitative practitioners considering VLM integration.

Implementation Complexity

7/10
Implementation requires: (1) OHLCV data pipeline with normalization and multi-style rendering, (2) four distinct benchmark splits with controlled label mechanisms, (3) counterfactual editing at the OHLCV level with matched magnitude and locality constraints, (4) structured prompt engineering and JSON parsing, (5) structural behavioral regression estimation, (6) overlap-weighted propensity scoring for artifact robustness, (7) block-aware sequential testing with e-processes for metered API evaluation, and (8) comprehensive statistical toolkit including MDE calculations and FDR control. The theoretical framework (Theorem 1, Propositions 1-5) requires careful implementation to maintain identification guarantees.

Reproducibility

4/5
The paper provides a comprehensive reproducibility package including source manifests, generation and renderer code, prompt banks, parsers, evaluation scripts, and recomputation scripts. When raw market sources cannot be redistributed, indices, hashes, and scripts are provided. All statistical protocols (block bootstrap, FDR control, sequential testing) are specified. However, commercial API results depend on specific model versions and evaluation windows, and the full benchmark release is described as 'intended' rather than confirmed available.

About this paper

Methodology: Martingale Doppelgänger-Eval. Problem types: Causal Inference, Classification, Computer Vision, Time Series Forecasting, Anomaly Detection.

The interactive Everscope explorer (charts, battles, favorites) loads below.