Towards a Shared Mission
We challenge the assumption that language models are inherently reliable human proxies for social science and linguistic research. Through systematic evaluation of downstream task limitations alongside possible applications, we examine when and how LLMs can legitimately serve as behavioral agents, research instruments, or simulation platforms. Thereby, we are committed to bridging the gap between artificial intelligence capabilities and social science validity through task- and language-dependent testing and multi-dimensional evaluation of large language models as behavioral agents and human simulacra.
Formalizing
EVAS develops task-specific datasets and formal specifications of stimulus and context conditions to enable standardized, comparable testing of LLMs as behavioral agents and human simulacra.
Evaluating
EVAS builds evaluation pipelines and metrics that probe LLM behavior mechanistically rather than relying on surface-level accuracy, using open-weight models to ensure reproducibility and interpretability.
Critiquing
EVAS critically analyzes unfalsifiable claims about silicon sampling and advocates for empirical validation in AI-based social science research through formalized pipelines and deep evaluation.




