EVAS ::. Research Cooperative | Ecological Validation of Artifical Simulacra

EVAS ::. Research Cooperative

Ecological Validation of Artifical Simulacra

Towards a Shared Mission

We challenge the assumption that language models are inherently reliable human proxies for social science and linguistic research. Through systematic evaluation of downstream task limitations alongside possible applications, we examine when and how LLMs can legitimately serve as behavioral agents, research instruments, or simulation platforms. Thereby, we are committed to bridging the gap between artificial intelligence capabilities and social science validity through task- and language-dependent testing and multi-dimensional evaluation of large language models as behavioral agents and human simulacra.

Formalizing

EVAS develops task-specific datasets and formal specifications of stimulus and context conditions to enable standardized, comparable testing of LLMs as behavioral agents and human simulacra.

Evaluating

EVAS builds evaluation pipelines and metrics that probe LLM behavior mechanistically rather than relying on surface-level accuracy, using open-weight models to ensure reproducibility and interpretability.

Critiquing

EVAS critically analyzes unfalsifiable claims about silicon sampling and advocates for empirical validation in AI-based social science research through formalized pipelines and deep evaluation.

Reviewed Research Contributions

Our research trace the empirical thread running through EVAS's work: testing whether LLM-based agents actually behave like the humans they are asked to simulate, and reporting when they don't. Recent work spans political and cultural bias in agent-centric simulations, construct validity in psychological questionnaires in LLMs, and the operational realism of generative agents on social networks. Up to this point, we find consistently that surface-level performance does not guarantee behavioral or ecological validity to human respondents.

showing 0 out of 5
no matching research found

Humans behind EVAS

EVAS is carried by researchers working across (computational) linguistics, natural language processing, and computational social science. United by the skepticism of unvalidated claims about LLM-as-human-proxy research, we build empirical infrastructure, datasets, evaluation pipelines, and simulation frameworks, needed to test claims and the performance of silicon sampling rather than assume it.

Alistair Plum

Alistair Plum

Alistair is a Postdoctoral Researcher at the University of Luxembourg. Within the Culture and Computation Lab, he bridges Humanities and Computer Science, applying computational linguistics methods to low-resource languages. His research investigates how cultural and linguistic diversity is represented (and often flattened) in language models, directly supporting EVAS’s mission to empirically validate when and how LLMs can serve as culturally grounded research instruments rather than universal proxies.

Nils Schwager

Nils Schwager

Nils works as a Research Fellow at the Department of Computational Linguistics at Trier University. His thesis examines when LM-based instruments can be trusted to produce new scientific insight, arguing that accuracy and calibration are distinct properties which current applications rarely establish together. His research on accuracy and calibration of Human Simulacra grounds EVAS's case for stronger validation in the mechanics of the models themselves.

Simon Münker

Simon Münker

Simon works as a Research Fellow at the Department of Computational Linguistics at Trier University. His thesis focused on language models as human simulacra in social science, critiquing the uncritical adoption of LLMs in simulated psychological questionnaires, which often show a lack of empirical realism. Together with related issues in LLM-based prediction of human behavior on social networks, his thesis and discussions with colleagues and co-authors sparked the idea behind EVAS: the case for stronger validation in AI-based simulations.