Research Questions
- How does LLM-based search affect efficiency, accuracy, user perception, and error handling compared to traditional search methods?
- How can users’ overconfidence in incorrect LLM outputs be reduced?
- How do confidence signals (e.g. colour-coded confidence indicators) influence user behaviour?
Results
- LLM-based search cut task completion time in half (1.6 min vs. 3.4 min).
- Users submitted fewer but more complex queries.
- Accuracy was similar for routine tasks, but on the task that contained an error, LLM users’ accuracy dropped sharply (47% vs. 93%).
- Participants found the LLM experience more satisfying.
- Confidence cues (colour coding) helped users detect errors and doubled accuracy.
Findings
- Efficiency:
- LLM search was 50% faster, and queries were fewer in number but more complex.
- Accuracy & Overconfidence:
- Users performed well on standard tasks, but when the dataset contained an error, LLM users showed strong overconfidence in incorrect answers.
- User Perception:
- Although the LLM experience felt more satisfying, users often failed to notice mistakes.
- Mitigation (Experiment 2):
- Confidence cues (green = high confidence, red = low confidence) increased accuracy from 26% to 53–58%. Users also ran more follow-up queries to test uncertain outputs.
Scores
- LLM Models: 4
- Synthetic Data: 0
- Method: 5
- Speed: 5
- Ethics: 3
- Accuracy: 5
- Demographics: 0
If you would like to explore this research in more detail, click here to read the full paper.