STEER: Assessing the Economic Rationality of Large Language Models

By Narun Raman, Taylor Lundy, Samuel Amouyal, Yoav Levine, Kevin Leyton-Brown, Moshe Tennenholtz

Rating

1294
Battle Count: 28

Relevance

7/10
The benchmark could be useful for evaluating LLMs' potential in financial decision-making and strategy development, though it's not directly focused on quantitative trading.

Implementation Complexity

6/10
The benchmark implementation is complex due to the wide range of economic concepts covered, but the authors provide tools and a web interface to facilitate its use.

Reproducibility

4/5
The paper provides detailed methodology and a web interface for generating benchmark questions, validating them, and visualizing experimental results.

About this paper

Methodology: STEER. Problem types: Classification, Decision Making, Multi-task Learning.

The interactive Everscope explorer (charts, battles, favorites) loads below.