Rating
1194
Battle Count: 166
Relevance
6/10
Directly relevant to DeFi quantitative trading and stablecoin-based trading infrastructure. The benchmark addresses peg-risk prediction for USDT, USDC, and DAI which are core trading pairs and collateral assets. Findings show current agentic systems cannot reliably detect rare depeg events, which is critical for automated trading, liquidation management, and settlement in stablecoin-based markets. The cost-aware evaluation framework is relevant for production trading systems where inference cost and latency matter. However, the paper focuses on risk classification rather than price-path forecasting or direct trading signal generation.
Implementation Complexity
5/10
Moderate complexity. The benchmark framework involves: (1) constructing historical-replay cases from hourly OHLCV data with feature engineering (peg deviation, volatility, volume signals, market context), (2) running six heterogeneous agentic platforms through a shared OpenRouter backend, (3) strict JSON parsing and validation, (4) frozen calibration rules for label mapping, (5) multi-metric evaluation (classification, probabilistic, continuous, cost, reliability), and (6) two-block experimental design. The core evaluation logic is straightforward, but orchestrating six different agent frameworks with consistent protocol compliance adds engineering overhead. No model training required; all agents use the same frozen LLM backend.
Reproducibility
5/5
Exceptional reproducibility: dataset released on Hugging Face, source code on GitHub, frozen protocol with fixed backend (Qwen 72B, temperature 0), strict JSON schema, versioned leaderboards, case lists, prompt packets, calibration rules, execution metadata, raw/calibrated predictions, cost logs, evaluation scripts, freeze manifests, and no-API recomputation scripts. Both 120-case and 507-case blocks use identical protocol, agent runner, and logging schema. Anonymization ablation and leakage-safety audits included.
About this paper
Methodology: StableEval Arena. Problem types: Classification, Risk Management, Anomaly Detection, Time Series Forecasting, Imbalanced Learning.
The interactive Everscope explorer (charts, battles, favorites) loads below.