STEER: Assessing the Economic Rationality of Large Language Models
By Narun Raman, Taylor Lundy, Samuel Amouyal, Yoav Levine, Kevin Leyton-Brown, Moshe Tennenholtz
Rating
1294
Battle Count: 28
Relevance
7/10
The benchmark could be useful for evaluating LLMs' potential in financial decision-making and strategy development, though it's not directly focused on quantitative trading.
Implementation Complexity
6/10
The benchmark implementation is complex due to the wide range of economic concepts covered, but the authors provide tools and a web interface to facilitate its use.
Reproducibility
4/5
The paper provides detailed methodology and a web interface for generating benchmark questions, validating them, and visualizing experimental results.
About this paper
Methodology: STEER. Problem types: Classification, Decision Making, Multi-task Learning.
The interactive Everscope explorer (charts, battles, favorites) loads below.