Total Simulated Survey Error: Designing and Diagnosing Survey Responses from Large Language Models
Source note: Preprints on societal impact.
arXiv:2609.10280v1 Announce Type: new Abstract: Large Language models (LLMs), having been trained on vast amounts of human-generated data, may encode the attitudes and behaviors of these humans. As such, LLMs show promise in mimicking human-like patterns that facilitate their use in simulating people in a wide variety of contexts. One such context is using LLMs as 'silicon samples', i.e., proxies of people in answering survey questions to establish public opinion, design policies, or use as (social) scientific data. However, several critical questions of social biases, generalization, and technical limitations remain, further complicated by a vast design space open to simulation designers. Multiverse analyses might help us make sense of the impact of different design choices, however, we lack a systematic understanding of the design space of LLM-generated surveys as well as how these decisions interplay with inherent LLM limitations. Therefore, how do we systematically identify, trace, and document limitations in LLM-generated survey responses? Building on traditions in the quantitative social sciences, specifically survey methodology and measurement theory, we investigate threats to the validity of LLM-generated survey responses. To do so, we design a framework that enumerates conceptual errors and systematic biases that can occur at different stages of the survey simulation lifecycle. Our framework, called the Total Simulated Survey Error (TS2E) Framework, provides a unified and end-to-end perspective on LLM-generated survey data. The framework, illustrated through a theo
Every model that read this
| Model | Provider | Stage | Score | Conf. | Latency | Prompt | When |
|---|---|---|---|---|---|---|---|
| Llama 3.3 70B | Meta | analysis | 0 | 70% | 5100ms | v1.0.0 / m1.0.0 | 2026-09-10 08:56 |
| Mistral Small 3.1 24B | Mistral AI | consensus | 0 | 80% | 5789ms | v1.0.0 / m1.0.1 | 2026-09-10 09:14 |
| gpt-oss 120B | OpenAI | consensus | -30 | 68% | 17427ms | v1.0.0 / m1.0.1 | 2026-09-10 09:14 |
| Llama 3.3 70B | Meta | analysis | -50 | 70% | 3911ms | v1.0.0 / m1.0.1 | 2026-09-10 09:17 |
The item discusses the potential use of Large Language Models (LLMs) in simulating survey responses, but raises critical questions about social biases, generalization, and technical limitations. The introduction of the Total Simulated Survey Error (TS2E) Framework aims to systematically identify and document limitations in LLM-generated survey responses, indicating a neutral, problem-solving approach. The item does not make a substantive claim about AI's effect on society, but rather explores the potential and limitations of using LLMs in survey simulations.
The paper discusses the potential of LLMs to simulate human responses in surveys. It acknowledges the risks of biases and errors but also presents a framework to mitigate these issues. The overall implication is uncertain due to the balancing of risks and opportunities.
The paper proposes a framework that could reduce harmful bias in LLM‑generated surveys, but it also legitimizes the use of synthetic respondents, potentially increasing reliance on flawed data for policy decisions. The net effect leans toward material concern.
Raises concerns about social biases and technical limitations in LLM-generated surveys.
Evidence extracted
- LLMs may encode the attitudes and behaviors of humans
- LLMs can be used as 'silicon samples' to simulate people in survey responses
- LLMs may encode human attitudes and behaviors
- LLMs can be used as proxies of people in surveys
Consensus history
Superseded readings are kept. Nothing is overwritten.
| When | Mean | Spread | Agreement | State |
|---|---|---|---|---|
| 2026-09-10 09:20 | -20 | 50 | low | current |
| 2026-09-10 09:14 | -10 | 30 | moderate | superseded |
SOURCE arXiv cs.CY (Computers and Society) (tier 1)
↓
DOCUMENT 723affd8-8f11-4fd8-bb4e-4734a9f15da3
https://arxiv.org/abs/2609.10280
↓
EVIDENCE 4 extracted excerpts
↓
MODEL RUN 4 runs, methodology 1.0.1
↓
SCORE -20 (Adverse)
↓
CONFIDENCE 72%