Research Questions
- Can LLMs imitate human agents in behavioural economics experiments?
- Can these experiments be reproduced using LLMs and expanded with new parameters?
- When “endowed” with different attributes, can LLMs represent diverse, human-like perspectives?
Results
- GPT-3 qualitatively reproduced the findings of behavioural economics experiments run with human participants.
- More advanced models (e.g., davinci-003) performed better, while smaller models did not respond reliably to the attributes they were endowed with.
- LLM-based experiments are significantly faster and more cost-efficient than human subject studies.
- LLMs can behave like diverse agents when endowed with different personalities or viewpoints.
Findings
- Experiment Replication:
- Classic behavioural findings, such as social preferences, fairness judgments, and status quo bias, were successfully reproduced using LLMs.
- Diversity Through Endowment:
- When models were given different political views or preferences, their responses shifted predictably, showing that viewpoint diversity can be controlled.
- Limitations of Smaller Models:
- Smaller GPT-3 variants (ada, babbage, curie) often failed to respond to the endowed attributes or to capture nuanced behavioural patterns.
- Memorisation Concerns:
- Because LLMs may have been exposed to descriptions of these experiments during training, questions remain regarding the originality of the reproduced behaviours.
- Ethical Considerations:
- Running experiments without human participants has advantages, but the authenticity of AI-generated responses and the risk of misrepresentation remain open concerns.
Scores
- LLM Models: 5
- Synthetic Data: 4
- Method: 4
- Speed: 4
- Ethics: 3
- Accuracy: 3
- Demographics: 2
If you would like to explore this research in more detail, click here to read the full paper.