MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code

By Varstern (Yifan) Wang, Siyu Wang, Yuecheng He

Rating

1778
Battle Count: 50

Relevance

10/10
Highly relevant. It addresses a critical failure mode in automated trading: 'silent failures' where code runs but implements the wrong logic. It provides a rigorous method to verify that LLM-generated trading code behaves as intended, which is crucial for operational risk management.

Implementation Complexity

7/10
Requires setting up a custom backtesting engine with strict causality enforcement, state logging, and sandboxing. The benchmark generation pipeline involves complex combinatorial logic and LLM-based back-translation with verification gates.

Reproducibility

5/5
The benchmark suite, prompt templates, and evaluation harness are available on GitHub. Market data is public (Binance). The methodology is fully described, including the generator logic and metric definitions.

About this paper

Methodology: MintEval. Problem types: Code Generation, Natural Language Processing, Algorithmic Execution, Risk Management.

The interactive Everscope explorer (charts, battles, favorites) loads below.