SLALOM: Simulation Lifecycle Analysis via Longitudinal Observation Metrics for Social Simulation
AI 摘要
SLALOM框架通过纵向观测指标评估LLM社会模拟的合理性,提升仿真验证标准。
主要贡献
- 提出SLALOM框架,用于评估社会模拟的真实性
- 使用动态时间规整(DTW)对齐模拟轨迹和实际数据
- 关注模拟过程而非最终结果,解决“停表问题”
方法论
SLALOM利用模式导向建模(POM)将社会现象视为时间序列,通过DTW比对模拟轨迹与经验数据,评估结构真实性。
原文摘要
Large Language Model (LLM) agents offer a potentially-transformative path forward for generative social science but face a critical crisis of validity. Current simulation evaluation methodologies suffer from the "stopped clock" problem: they confirm that a simulation reached the correct final outcome while ignoring whether the trajectory leading to it was sociologically plausible. Because the internal reasoning of LLMs is opaque, verifying the "black box" of social mechanisms remains a persistent challenge. In this paper, we introduce SLALOM (Simulation Lifecycle Analysis via Longitudinal Observation Metrics), a framework that shifts validation from outcome verification to process fidelity. Drawing on Pattern-Oriented Modeling (POM), SLALOM treats social phenomena as multivariate time series that must traverse specific SLALOM gates, or intermediate waypoint constraints representing distinct phases. By utilizing Dynamic Time Warping (DTW) to align simulated trajectories with empirical ground truth, SLALOM offers a quantitative metric to assess structural realism, helping to differentiate plausible social dynamics from stochastic noise and contributing to more robust policy simulation standards.