{"api_version":"v1","generated_at":"2026-09-29T13:00:45","count":380,"scope":"selected","papers":[{"id":"2609.30563","version":1,"title":"Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content","zh_title":"少思考以更好地模拟：直觉提示提升LLM智能体模拟个体社交媒体反应（包括不熟悉内容）","abstract":"Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour and has paid little attention to whether an agent behaves in line with the profile it was given. The present study profiled eight Serbian participants through a questionnaire, a deep interview, and a written self-presentation, recorded their reactions to sixty-eight social media posts, and asked four language models to predict those reactions under five prompt conditions varying profile content and instruction style. Attitudinal content improved prediction over demographic backstories by a wide margin. Agents matched their stated profiles more closely than participants matched their own survey answers, and consistency proved unrelated to fidelity once profile information was present. Instructing models to respond intuitively and immediately rather than analytically gave the highest fidelity of any condition and cut the compression of individual differences from seven times the human level to three. The advantage held on posts about topics the questionnaire never raised, where that condition reached the highest fidelity of any setup and beat a crowd baseline by a wide margin, which suggests that agents prompted this way could serve as general-purpose simulated users rather than specialists on the topics they were profiled for. Results may bear implications for the development of language models, because intuition-based setups appear better suited to some tasks than reasoning-based ones.","authors":["Ljubisa Bojic","Tijana Stanic","Joerg Matthes","Agariadne Dwinggo Samala","Bojana Dinic","Jue Wang"],"categories":["cs.AI","cs.CL","cs.HC","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30563","pdf_url":"https://arxiv.org/pdf/2609.30563","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","人类行为对照","提示策略"],"reason":"用LLM预测真实个体社交媒体反应，与人类数据对照，评估提示策略对仿真保真度的影…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":1,"question":"在社交媒体反应预测中，不同的提示设计（档案内容与指令风格）如何影响大语言模型模拟特定个体反应的保真度？","design":"用四个大语言模型扮演八名塞尔维亚参与者，基于问卷、深度访谈和书面自我呈现构建个人档案，在五种提示条件下（变化档案内容和指令风格，如人口学背景、态度内容、直觉式回应等）预测他们对68条社交媒体帖子的反应，测量预测与真实反应的一致性。","baseline":"八名塞尔维亚参与者的真实社交媒体反应数据，以及他们自己的问卷回答（用于比较自我一致性）。","findings":"态度性档案内容比人口学背景大幅提升预测准确度；直觉式指令（要求模型凭直觉立即回应而非分析式推理）在所有条件中保真度最高，并将个体差异压缩从人类的七倍降至三倍。","reliability":"论文未讨论","relevance":"该研究直接检验LLM模拟个体行为的保真度，与人类真实反应对照，并揭示提示策略对仿真偏差的影响，对评估LLM作为人类被试替代品的可靠性具有关键参考价值。","inspiration":"值得借鉴的是通过改变提示指令风格（直觉式 vs 分析式）来操纵模型的认知模式，并测量其对个体差异保真度的影响｜可迁移到消费者金融决策实验，如模拟个体在信贷选择或储蓄行为中的异质性反应｜用LLM扮演不同风险偏好和金融素养的消费者，处理为直觉式或分析式提示，结果变量为信贷产品选择或投资决策，与真实消费者调查或实验数据对照。"}},{"id":"2609.30883","version":1,"title":"Warned alike, AI agents avoid the less-crowded road while people take it","zh_title":"同样被警告，AI智能体避开较不拥挤的道路而人类选择它","abstract":"AI agents built on a few shared models increasingly act for many people. A shared forecast about others can align their choices and change how scarce capacity is allocated. We tested this feedback in a two-road congestion game. Adding one sentence warning that others might follow a routing tip made populations of 50 GPT agents crowd one road while avoiding the nearly empty alternative. Average travel time rose from 64 to 95 min, although any crowded-road agent could have saved 69 min by switching alone. The warning discouraged the very move it predicted. The pattern persisted for 100 rounds. Two other model families shifted the same way without locking onto one road. Twelve all-human groups (240 participants) stayed near balance under numerical reports or the tip and warning. In 24 mixed groups with a further 240 participants, imbalance grew with the share of agents in the registered analysis, while people increasingly took the road the agents avoided. Collective costs stayed below the allagent reference, but with 15 agents and 5 humans, agent seats averaged 80 min, compared with 44 min for human seats. Shared forecasts can thus sustain collective inefficiency among similar agents. A better group average can also hide an unequal burden. Evaluations of AI agents that share resources should test populations, treat messages as interventions and report who bears the costs.","authors":["Takahiro Ezaki","Naoto Imura","Katsuhiro Nishinari"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30883","pdf_url":"https://arxiv.org/pdf/2609.30883","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","拥堵博弈","人机对照"],"reason":"用GPT agent群体模拟拥堵博弈，并与240名人类被试对照，发现警告导致a…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":2,"question":"共享预测信息（警告）是否会导致AI智能体在拥堵博弈中持续选择拥挤道路，从而造成集体低效？","design":"用50个GPT智能体（gpt-5.4-mini）模拟通勤者，在双路径拥堵博弈中，操纵每日广播信息（无报告、数值报告、提示、提示+警告），测量路径不平衡度、切换比例和平均旅行时间。","baseline":"12个全人类组（240名参与者）在相同博弈和广播条件下的行为数据，以及24个混合组（240名参与者）与智能体共存时的行为。","findings":"警告导致GPT智能体群体持续拥挤一条道路，平均旅行时间从64分钟升至95分钟，尽管个体切换可节省69分钟；全人类组保持接近均衡，混合组中人类更多选择智能体避开的道路，且智能体承担更高成本。","reliability":"论文指出共享预测可导致相似智能体间的集体低效，且群体平均成本可能掩盖负担不均；评估共享资源的AI智能体应测试群体、将消息视为干预并报告成本承担者。","relevance":"高度相关：该研究用LLM智能体模拟人类在拥堵博弈中的决策，并与真实人类数据对照，揭示了共享信息对集体行为的负面影响，直接回应了仿真可靠性与偏差问题。","inspiration":"值得借鉴的是将公共信息（警告）作为干预变量，观察其对群体决策动态的影响，并设置全人类和混合组对照以分离智能体特有行为。｜可迁移到政策公告的预期形成场景，如央行沟通对金融市场参与者行为的影响。｜设计：用LLM智能体模拟投资者，处理为央行发布的不同措辞的前瞻指引，结果变量为资产配置集中度和市场波动率，对照真实投资者在类似公告下的交易数据。"}},{"id":"2609.30896","version":1,"title":"Large language models underestimate and partly misrepresent cultural variation in everyday norms","zh_title":"大语言模型低估并部分误现日常规范的文化差异","abstract":"A key aspect of culture is a society's norms about everyday behavior. How accurately do large language models (LLMs) represent cultural differences in such norms? To answer this question we used the recent Global Study of Everyday Norms (GSEN), which collected ratings of 150 scenarios in 90 societies, as the human benchmark. We prompted GPT-5 to estimate each society's average rating for every scenario, and later repeated the benchmark in three other LLMs: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro. Compared to GSEN estimates, all four LLMs misrepresented cultural variation in two ways. First, they greatly underestimated its magnitude, estimating differences between societies to be, on average, less than half their measured size. Second, for many scenarios the LLMs poorly identified the pattern of variation, that is, which societies judged the behavior less acceptable and which societies judged it more acceptable. The pattern of variation was identified better for scenarios that elicit concerns about vulgarity, especially scenarios involving kissing and flirting. We also found that norms in more developed societies tended to be estimated somewhat more accurately, and that prompting in local survey languages rather than English produced only a modest improvement in accuracy. Local-language prompting also reduced, but did not remove, the underestimation of between-society differences. Cultural differences in everyday norms are only weakly and unevenly represented by LLMs.","authors":["Kimmo Eriksson","Irina Vartanova","Pontus Strimling"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30896","pdf_url":"https://arxiv.org/pdf/2609.30896","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["文化规范","仿真偏差","人类数据对照"],"reason":"用LLM估计社会规范并与真实人类调查数据对照，评估仿真偏差","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":3,"question":"大语言模型能否准确表示不同社会之间日常行为规范的文化差异？","design":"用 GPT-5、GPT-5.4、Claude Opus 4.6 和 Gemini 3.1 Pro 四个 LLM，以英语或当地语言提示，估计 90 个社会对 150 个日常行为场景的平均可接受度评分，并与 GSEN 调查数据比较。","baseline":"全球日常规范研究（GSEN）中 90 个社会超过 25,000 名参与者对 150 个场景的评分。","findings":"四个 LLM 都大幅低估了社会间规范差异的幅度，估计的差异平均不到实际测量的一半；对于许多场景，LLM 识别差异模式的能力较差，但在涉及粗俗（如亲吻和调情）的场景上表现较好。","reliability":"论文指出 LLM 对较发达社会的规范估计更准确，用当地语言提示仅带来适度改进，且未能消除对社会间差异的低估；LLM 对文化差异的表示既弱且不均衡。","relevance":"该研究直接评估 LLM 在跨文化社会规范仿真中的偏差，与研究者关注的人类仿真可靠性及失效条件高度契合，值得精读原文以了解具体偏差模式和测量方法。","inspiration":"借鉴其用真实大规模跨国调查作为基准、将仿真误差分解为幅度和模式两个维度的方法，并检验提示语言等处理变量的影响。｜可迁移到跨国消费者行为或政策接受度的仿真研究，例如不同国家居民对环保政策、金融产品条款或广告伦理的接受度差异。｜以 LLM 模拟不同国家被试，对一系列政策或产品场景给出接受度评分，处理变量为提示语言（英语 vs 当地语言），结果变量为评分，与真实跨国调查数据（如世界价值观调查或特定政策民意调查）对照，评估 LLM 仿真的幅度压缩和模式偏差。"}},{"id":"2604.20050","version":4,"title":"Information Aggregation with AI Agents","zh_title":"AI代理的信息聚合研究","abstract":"Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing price movements? We conduct a controlled experiment where AI agents trade in a prediction market after receiving private signals, across four information structures of increasing complexity. We find that although the median market is effective at aggregating information in the easy information structures, performance deteriorates in the harder structures, suggesting that AI agents struggle in environments where more than two levels of interactive reasoning are required, a ceiling close to the one documented in human subjects. Consistent with our theoretical predictions, market accuracy does not improve from allowing cheap talk communication, changing the duration of the market, or strategic prompting; initial price has little average effect but matters in the very hard structure. We also find that ``smarter'' AI agents perform better at aggregation and are more profitable. Surprisingly, giving them feedback about past performance does not improve aggregation. A further wave of markets, run three months later with capability-frontier models, aggregates information more often in the three easier structures but not in the hardest one, where higher capability replaces markets that are confidently wrong with markets that hedge near 0.5.","authors":["Spyros Galanis"],"categories":["econ.GN","cs.AI","cs.GT","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-04-21","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2604.20050","pdf_url":"https://arxiv.org/pdf/2604.20050","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","信息聚合","行为实验"],"reason":"用AI代理模拟人类交易行为，并与人类被试结果对照，评估信息聚合能力。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":4,"question":"AI 代理能否通过交易聚合分散的私人信息，并像人类一样通过观察价格变动推断他人知识？","design":"用八种大语言模型（Claude Haiku 3.5/4.5、Gemini 2.5/3 Flash、GPT-4o/5 mini、gemma3:4b、qwen3:8b）组成三人交易团队，在四种难度递增的信息结构中交易预测市场证券；处理包括允许廉价谈话、战略提示、初始价格（0.3/0.5/0.7）、市场时长（3/6/9轮）和反馈信息，共144种配置，每种至少运行12次，生成1772个市场；结果变量为市场准确率（价格是否接近真实价值）和利润。","baseline":"人类被试在类似信息结构中的推理层级上限（约两级交互推理），来自已有实验文献。","findings":"在简单信息结构中，中位数市场能有效聚合信息，但在需要超过两级交互推理的困难结构中表现恶化；允许廉价谈话、改变市场时长或战略提示均未提高市场准确率，初始价格平均影响小但在极难结构中重要。更聪明的AI代理聚合更好且更盈利，但反馈过去表现未改善聚合；三个月后用前沿模型重跑，在三个较易结构中聚合更频繁，但在最难结构中未改善，高能力模型将自信错误的市场替换为在0.5附近对冲的市场。","reliability":"论文承认AI代理在需要超过两级交互推理的环境中挣扎，且能力提升并未解决最难结构中的聚合失败；未明确讨论其他失效条件。","relevance":"该研究直接以AI代理模拟人类交易者，并与人类被试的推理层级上限对照，评估信息聚合能力，属于用LLM进行人类仿真实验并检验可靠性的核心工作，值得精读。","inspiration":"借鉴其系统操纵信息结构复杂度、市场机制参数（时长、初始价格、沟通）和模型能力来测量聚合效率与利润的做法，并设置理论基准（可分离证券的完全聚合预测）｜可迁移到资产定价实验中的信息效率研究，如内幕交易监管、分析师预测市场或央行沟通对价格发现的影响｜用不同能力LLM作为交易者，在预测市场中交易与真实宏观经济指标挂钩的证券，处理为是否允许公开评论或改变交易轮次，结果变量为价格偏离真实值的程度，并与人类实验数据（如Plott & Sunder的经典信息聚合实验）对照。"}},{"id":"2609.02729","version":2,"title":"BuildOcc: A Large Language Model Occupant Agent Platform for Building Energy Research","zh_title":"BuildOcc：用于建筑能源研究的大语言模型居住者智能体平台","abstract":"Occupants are a primary source of uncertainty in building energy consumption and management, yet existing occupant behavior models cannot capture adaptive and reasoning responses considering the occupant's personal history, current context, and the type of energy signal being delivered. This study presents BuildOcc, an open-source Python platform that grounds large language model agents in the American Time Use Survey (ATUS), a nationally representative diary dataset covering 16,684 respondents. Through BuildOcc, each simulated occupant agent can be instantiated with a demographic persona drawn from ATUS population statistics, a memory stream that accumulates and reflects on timestep-level observations, and an activity scheduler that samples empirically from ATUS time-at-activity distributions. The platform exposes a three-layer interface - Python library, REST API, and Model Context Protocol server - so that any building energy tool (EnergyPlus, Home Assistant) can integrate behavioral intelligence without bespoke coupling code. A plugin registry lets the community add new occupant strata, custom schedulers, and alternative memory backends as separate installable packages. Two validation tiers show that ATUS-grounded sampling reproduces empirically calibrated activity distributions and that demographic priors propagate into persona-consistent agent reasoning across timesteps, establishing internal consistency across strata. BuildOcc provides the building energy community with a reusable, openly available implementation of the occupant behavioral layer. BuildOcc is openly released at https://doi.org/10.5281/zenodo.21192895 under the Apache License 2.0 and installable via pip install buildocc.","authors":["Wooyoung Jung"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-28","first_seen":"2026-09-03","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2609.02729","pdf_url":"https://arxiv.org/pdf/2609.02729","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","建筑能源","ATUS数据"],"reason":"用LLM agent模拟建筑内人员行为，基于ATUS真实数据对照，属于人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":5,"question":"如何构建一个以美国时间使用调查（ATUS）为数据基础、基于大语言模型（LLM）的居住者智能体平台，用于建筑能耗研究中的居住者行为仿真？","design":"该研究开发了 BuildOcc 平台，使用 LLM 智能体模拟四类美国人口群体（在职单身成年人、退休夫妇、在职父母、无业成年人）。每个智能体包含从 ATUS 人口统计中抽取的人口特征、从 ATUS 数据中采样的活动日程、记忆流和推理引擎。在每个 15 分钟时间步，智能体根据人口特征、记忆和当前环境选择下一个活动并说明理由，同时也会对需求响应信号做出接受、拒绝或推迟的决定。平台提供 Python 库、REST API 和 MCP 服务器三层接口，便于与 EnergyPlus 等工具集成。","baseline":"美国时间使用调查（ATUS），包含 16,684 名受访者的全国代表性日记数据，其中 6,611 名受访者属于四个目标人口阶层。","findings":"第一层验证表明，活动调度器能够将 ATUS 活动分布复现到采样噪声范围内；第二层验证表明，人口先验信息能够传播到智能体的推理中，产生与人口特征一致的行为差异。","reliability":"论文承认的局限包括：每个时间步只能选择一个活动；ATUS 仅覆盖美国人口；活动仅基于主要活动（ATUS 一次只记录一个活动）；记忆重要性分数由智能体自行分配并用于检索，缺乏外部校准或反馈路径。","relevance":"该研究将 LLM 智能体与全国代表性时间使用调查数据结合，用于模拟人类行为，并进行了与真实数据的对照验证，属于人类仿真研究，但场景限定于建筑能耗领域，与经济金融问题关联度较低。","inspiration":"借鉴其将 LLM 智能体与大规模调查数据结合、通过分层抽样和记忆机制生成个体行为的方法，可用于构建具有人口代表性的经济决策仿真。｜可迁移到消费者跨期选择或家庭能源消费行为研究，例如模拟不同人口群体对动态电价或节能政策的反应。｜设计雏形：以美国消费者支出调查（CEX）或收入动态面板研究（PSID）为数据基础，构建 LLM 智能体代表不同收入阶层，施加电价上涨或补贴政策处理，结果变量为能源消费和支出变化，并与真实调查数据对照验证。"}},{"id":"2609.30940","version":1,"title":"Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms","zh_title":"LLM智能体社会中的金融脆弱性：协调失败与稳定机制","abstract":"Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, but also at the level of the systems they jointly create. We study this problem with FRAIL, a controlled experimental framework that places LLM agents in three dynamic financial environments---bank runs, debt rollover, and reward crowdfunding---where agents' decisions reshape the financial conditions faced by others. Across seven leading LLMs, we find widespread collective fragility even when no agent is instructed to destabilize the system: 77\\% of baseline bank-run episodes and 83\\% of debt-rollover episodes end in failure. We then compare three interaction mechanisms based on compensated commitments, centralized commitment agreements, and participant-led coalitions. All three improve aggregate outcomes, but no single mechanism performs best across all financial structures. Across mechanisms, successful stabilization shares a common temporal pattern: broad commitment forms early, before defensive behavior becomes self-reinforcing. Our findings show that individually capable agents do not automatically form safe financial systems, highlighting system-level evaluation and interaction design as central problems for financial AI safety. Code is available at https://anonymous.4open.science/r/FinFrail-CF26.","authors":["Zhenhao Fu","Ruipeng Xu","Qibing Ren"],"categories":["cs.AI","q-fin.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30940","pdf_url":"https://arxiv.org/pdf/2609.30940","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM智能体","金融仿真","协调失败"],"reason":"用LLM agent模拟金融系统中的协调失败，涉及经济场景，但无真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":6,"question":"多个LLM智能体在共享金融环境中是否会自发产生系统性金融脆弱性，以及如何通过交互机制设计来缓解这种脆弱性？","design":"使用FRAIL框架，将七个主流LLM作为金融决策者，分别置于银行挤兑、债务展期和奖励众筹三种动态金融环境中，通过多轮交互模拟决策，测量系统失败率（如银行倒闭、债务违约、众筹失败）和承诺形成的时间模式。","baseline":"无对照","findings":"在基线条件下，77%的银行挤兑和83%的债务展期情景以失败告终，即使没有智能体被指示破坏系统。三种交互机制（补偿承诺、集中承诺协议、参与者主导联盟）均能改善结果，但最佳机制因金融结构而异，成功稳定依赖于在防御行为自我强化之前尽早形成广泛承诺。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟金融协调失败，属于经济场景下的多智能体仿真，但缺乏真实人类数据对照，适合关注LLM仿真方法本身或金融AI安全的研究者阅读。","inspiration":"借鉴其动态金融环境设计和多智能体交互机制比较，可迁移到资产定价实验或政策公告预期形成等场景。｜可设计一个信贷审批歧视实验，用LLM扮演银行信贷员和借款人，施加不同信息披露政策作为处理，测量贷款批准率和违约率，并与真实信贷数据对照。"}},{"id":"2609.31054","version":1,"title":"Cheap, open agents make LLM pollution harder to mitigate","zh_title":"廉价开放智能体使LLM污染更难缓解","abstract":"Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing multiple response types yielding various detection checks. Fully open agents ran locally without usage fees and performed competitively with commercial alternatives. Open and commercial agents failed different sets of checks, and no single check reliably detected all agents, but open-text responses discriminated best between agents and humans. These findings identify fully open agents as a distinct risk for LLM pollution and support multilayered detection strategies emphasizing open-text analysis.","authors":["Raluca Rilla","Anne-Marie Nussberger","Rui Mata","Dirk U. Wulff"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31054","pdf_url":"https://arxiv.org/pdf/2609.31054","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM污染","调查数据","检测方法"],"reason":"研究LLM污染人类调查数据，评估检测方法，与仿真可靠性直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":7,"question":"完全开源的自主智能体是否降低了LLM污染人类调查数据的门槛，以及现有检测方法能否有效识别这些智能体？","design":"使用九种智能体配置（三种完全开源、两种混合、四种专有）自主完成同一份调查问卷，每种配置运行40次，共360次运行。智能体被赋予随机的人口统计特征（性别、年龄），并收到统一指令。调查包含多种回答类型和18项检测检查（16项通过/失败检查和2项计时测量），用于评估智能体的表现和可检测性。","baseline":"3,242名来自Prolific的美国参与者，在2025年10月20日至11月4日期间完成同一调查问卷，样本在性别、年龄和种族上大致代表美国人口。","findings":"完全开源的智能体在本地运行且无使用费用，其表现与商业替代方案相当。开源和商业智能体在不同检查上失败，没有单一检查能可靠地检测所有智能体，但开放式文本回答在区分智能体和人类方面效果最好。","reliability":"论文指出，开源和商业智能体在不同检查上失败，没有单一检查能可靠地检测所有智能体，因此需要多层检测策略，特别是强调开放式文本分析。此外，研究仅操纵了年龄和性别两个人口统计变量，未涵盖其他可能影响回答的变量；智能体被明确指示避免与可疑嵌入命令交互，这可能降低了某些检测的失败率。","relevance":"该研究直接评估了LLM智能体对人类调查数据的污染风险及检测方法，与您关注的LLM仿真可靠性和偏差问题高度相关，特别是它提供了真实人类数据作为对照，并揭示了开源智能体带来的新风险，值得精读原文以了解具体检测方法和失效模式。","inspiration":"借鉴其多层检测策略和开放式文本分析来识别LLM生成回答的方法，可迁移到经济金融领域的调查数据质量控制中。｜可应用于消费者信心调查、投资者情绪调查或政策评估中的问卷数据，检测是否存在LLM污染。｜设计：以真实人类调查数据（如密歇根大学消费者信心调查）为基准，让开源和商业LLM智能体自主完成同一问卷，比较其回答分布和开放式文本特征，并开发基于文本分析的检测指标，评估不同检测方法的敏感性和特异性。"}},{"id":"2609.27690","version":2,"title":"Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research","zh_title":"合成研究验证中的后果行为与表征公平性","abstract":"Researchers in industry and academia use synthetic survey respondents powered by large language models as substitutes for human samples. These synthetic populations require validation against real-world data, so researchers often address them using ad hoc comparisons with human surveys. Inspired by the intention-behaviour gap in behavioural science, we argue that these validations test the wrong thing for most applied cases where decision makers commission synthetic research to anticipate consequential behaviour. To address this problem, we propose a validation framework with two requirements. First, every validity claim must state its level of correspondence with human data: does the sample predict what the represented people do, which of four diagnostics (location, dispersion, response process and structure) does the validation address, and does the validation compare against experimental effects? Second, researchers must report validity claims for subgroups, since these groups are often the most affected by consequential decisions and aggregate accuracy hides their misrepresentation. Our validation framework operationalises three justice dimensions (distributional, procedural, and recognition) as measurable quantities and defines within-persona counterfactual experiments as a validation requirement. We then apply the framework to electric vehicle charging tariffs, before closing with a reporting checklist that researchers can use to make convincing validity claims.","authors":["Florian Kutzner","Celina Kacperski","Laura de Moli\\`ere","Edoardo Chidichimo","Min Jun Jung","Felix P. S. Wallis","James K. He"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-24","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.27690","pdf_url":"https://arxiv.org/pdf/2609.27690","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B2","B3","B4"],"tags":["LLM仿真","验证框架","公平性"],"reason":"直接研究LLM合成调查受访者作为人类替代，提出验证框架并应用于电动汽车充电定价…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":1,"question":"如何验证基于大语言模型的合成调查受访者能否预测真实人群的后果性行为，并确保子群体代表性公平？","design":"本文提出一个验证框架，而非进行仿真实验。框架要求：明确效度声明与人类数据的对应层级（预测行为、四种诊断：位置、离散度、响应过程、结构、与实验效应比较）；报告子群体效度；将分配、程序、承认三个正义维度操作化为可测量量；定义“人设内反事实实验”作为验证要求。并以电动汽车充电电价为例应用该框架。","baseline":"无对照（本文为框架性论文，未提供具体人类数据对照）","findings":"现有合成受访者验证多聚焦于态度或意见的边际分布一致性，忽视了意图-行为差距，无法证明其能预测后果性行为。提出的框架要求效度声明必须明确对应层级、子群体表现，并通过人设内反事实实验检验因果推断能力。","reliability":"论文承认以下局限：训练数据污染难以评估；子群体分析受限于人类基准数据可得性；行为标准本身存在缺陷（如公开行为、行政记录、实验室任务各有问题）；模型版本更新导致效度证据时效短；缺乏全面的实用测试集来确定合成人群满足哪些效度要求。","relevance":"该论文直接针对LLM合成受访者作为人类替代的验证问题，提出批判性框架并强调子群体公平，与研究者关注的人类仿真可靠性、偏差及经济学政策评估场景高度相关，值得精读原文。","inspiration":"借鉴其“人设内反事实实验”设计，可对同一合成个体施加不同处理以估计个体处理效应，并对照真实人类实验效应进行验证。｜可迁移到消费者金融决策研究，如信贷产品选择、保险购买或退休储蓄计划参与等场景。｜以合成受访者作为被试，处理为不同信贷条款（如利率、还款期限），结果变量为选择行为，对照真实世界信贷申请数据或实验室实验数据，检验合成样本的预测效度与子群体公平性。"}},{"id":"2609.29928","version":1,"title":"Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations","zh_title":"文化差异保持：诊断LLM模拟调查人群中的扁平化与夸张化","abstract":"Large language models (LLMs) are increasingly used as synthetic survey respondents to estimate population response distributions. In cross-cultural survey simulation, evaluations should assess not only distributional fidelity within countries but also whether differences across countries are preserved. However, existing distance-based metrics such as Jensen--Shannon divergence (JSD) do not directly capture such cross-country differences. To address this limitation, we introduce Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration. CDP identifies reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature. To evaluate CDP, we conduct experiments across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test. The results reveal a systematic discrepancy between conventional fidelity metrics and CDP. Controlled experiments show that CDP changes monotonically as cross-country divergence is attenuated or amplified, while the corresponding changes in JSD remain relatively small. In our audit of real LLM generations, DeepPersona-Inspired prompting is frequently favored by conventional fidelity metrics but exhibits the strongest flattening in every model--domain block. CDP thus complements fidelity metrics by directly quantifying the attenuation or amplification of cross-country divergence.","authors":["Yeeun Chae","Yewon Choi","Seunghyun Lee","IL Im"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29928","pdf_url":"https://arxiv.org/pdf/2609.29928","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM仿真","跨文化调查","算法保真度"],"reason":"直接研究LLM仿真调查人群，评估跨文化差异保真度，并与真实人类数据对照，提出诊…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":2,"question":"如何诊断LLM模拟调查人群时对跨国文化差异的扁平化或夸张化？","design":"用四个开源LLM（Gemma-3-4B、Qwen3.5-9B、Qwen3.5-27B、Llama-2-13B）模拟六个国家（阿根廷、澳大利亚、德国、印度、肯尼亚、美国）的受访者，采用三种人设提示方法（Cultural Prompting、PersonaHub-Inspired、DeepPersona-Inspired），生成世界价值观调查（WVS）和大五人格测试的回答，并测量国家间分布差异的保持程度。","baseline":"真实人类数据：WVS第七波六个国家的全国回答分布，以及OpenPsychometrics的大五人格测试数据（阿根廷、澳大利亚、印度）。","findings":"传统分布保真度指标（如JSD）与CDP存在系统性偏差：DeepPersona-Inspired提示在多数模型-领域组合中分布保真度最高，但文化扁平化最严重。CDP在受控实验中随跨国差异的衰减或放大单调变化，而JSD变化很小。","reliability":"论文未明确讨论失效条件，但指出CDP需要一次性人类校准，且仅适用于有跨国人类参照数据的调查领域。","relevance":"该研究直接针对LLM仿真调查中的文化差异保真度问题，提出了新的诊断指标，并用真实人类数据对照，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其通过受控扰动构造扁平化/夸张化数据集来检验指标敏感性的方法，以及将分布保真度与差异保持度分开评估的思路。｜可迁移到跨国经济偏好调查或政策态度仿真中，例如用LLM模拟不同国家消费者对通胀预期的回答，检验其是否保持国家间差异。｜设计：用LLM模拟多国受访者回答通胀预期调查，施加不同人设提示，测量国家间预期分布的差异保持度，并与密歇根大学或欧洲央行的真实调查数据对照。"}},{"id":"2609.30030","version":1,"title":"Artificial Societies Benchmark: A Validation Framework for Synthetic Research","zh_title":"人工社会基准：合成研究的验证框架","abstract":"A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to help researchers assess whether synthetic populations support their intended analyses. The framework combines eleven tests across internal, construct, and external validity, drawing on twenty human sources and comparing nine language models. It connects each research use to the evidence it requires and tests how results change with the information we supply about respondents. Importantly, strong performance in one domain does not establish fidelity in the others. Models often answer too consistently, compress response scales, and alter relationships between traits whilst richer profiles improve prediction for some models and worsen it for others. The resulting scorecard helps researchers identify which aspects of a synthetic population can support their analysis and where researchers need further human evidence.","authors":["Edoardo Chidichimo","Min Jun Jung","Felix P. S. Wallis","James K. He"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30030","pdf_url":"https://arxiv.org/pdf/2609.30030","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","效度验证","合成人群"],"reason":"直接评估LLM合成人群的效度，含人类数据对照和批判性分析","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":3,"question":"如何系统评估大语言模型生成的合成人群在多大程度上能支持研究者预期的分析（如调查回答、心理测量、实验效应）？","design":"用九个大型语言模型（含专有和开源）根据二十个人类数据源（调查、面板、人格量表、实验）中的受访者信息（人口统计、先前回答、个人描述等）生成合成回答，并施加实验处理；通过十一个测试从内部效度、构念效度和外部效度三个维度评估合成人群的响应过程、心理测量结构和总体/实验保真度，同时比较不同信息丰富度和统计控制（独立边际、高斯秩相关）的影响。","baseline":"二十个人类数据源，包括全国调查、追踪面板、人格量表和实验，提供真实回答、人口统计、先前回答、个人描述、实验分配和独立测量结果作为对照基准。","findings":"模型在某一效度领域的良好表现不能推广到其他领域；模型往往回答过于一致、压缩量表范围、改变特质间关系，且更丰富的个人信息对某些模型改善预测而对另一些模型则恶化预测。","reliability":"论文承认强表现不跨域通用，并指出模型回答过于一致、压缩量表、改变特质关系等失效条件；但未在节选中详细讨论其他局限。","relevance":"该研究直接针对LLM合成人群的效度验证，提供人类数据对照和批判性分析，与研究者关注的经济学实验和政策评估场景高度相关，值得精读原文以了解具体测试方法和失效模式。","inspiration":"借鉴其多维度效度测试框架和统计控制设计，可迁移到经济政策评估中的异质性处理效应或消费者选择实验，例如用LLM模拟不同收入群体对税收优惠的反应，以真实调查数据（如美国消费者财务调查）为基准，比较模型生成的边际消费倾向与人类数据的分布和协变量关系。"}},{"id":"2609.29952","version":1,"title":"Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes","zh_title":"Augur：用于预演产品和政策变化反应的合成决策实验室","abstract":"Before a product or policy change ships, the question that matters is how people will react to it. Augur rehearses that reaction offline: it builds a typed knowledge graph from the change documents, populates a grounded persona market, simulates the interaction, and returns an auditable decision memo recommending one of five actions. We assemble Gold-50, fifty real product and policy episodes whose real-world outcome is known, adjudicated against the public record, and score the five-way release verdict against it. Our central finding is methodological and negative: most of the measured gap between frontier cloud models and open-weight models we fine-tune and serve offline is attributable to an under-specified evaluation, not a difference in capability. We show this three ways. First, the prompt envelope alone can dominate the score: holding weights, cases and scorer fixed, one system -- a LoRA-SFT adapter on Qwen3-32B -- swings from 0% to 73%. Second, in a matched 2x2 ablation, defining the decision taxonomy in the prompt -- with no model change -- lifts every frontier model by +24 to +34pp; under the under-specified prompt, Qwen3-32B LoRA-SFT served offline beats all three frontier models (paired McNemar, Holm-corrected), and once the prompt is fair no significant difference from any of them is detected. Third, agreement with the distillation teacher rises without accuracy following, and the full pipeline amplifies a systematic \"over-doom\" bias rather than improving the verdict. Separately, we validate the reaction layer on its own terms: blind judges across four model families find the synthetic reaction recovers 67-90% of the concerns the public actually raised, and a pre-registered ablation locates its value -- largest where the decision is hardest, redundant near ceiling. The pipeline that regenerates every number and figure here is available from the authors.","authors":["Rahul Khedar","Mayank Malhotra","Avinash Karn"],"categories":["cs.AI","cs.CL","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29952","pdf_url":"https://arxiv.org/pdf/2609.29952","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM仿真","人类行为预测","政策评估"],"reason":"用LLM模拟人类对产品和政策变化的反应，并与真实结果对照，直接属于人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":6,"question":"在真实产品与政策变更决策中，前沿云端模型与开源微调模型之间的性能差距有多少是真实能力差异，多少是评估方法（提示词）造成的？","design":"构建 Augur 系统，将决策分解为文档→知识图谱→人物角色→模拟→报告五个阶段，用 LLM 生成利益相关者角色并模拟其互动，最终输出五选一的发布决策建议。使用 50 个真实产品/政策案例（Gold-50）作为基准，对比不同模型（前沿云端模型与开源微调模型）在有无决策分类定义提示词下的准确率。","baseline":"Gold-50 基准：50 个真实产品与政策变更案例，其真实世界结果已通过公开记录人工核实，作为五分类发布决策的对照标准。","findings":"主要发现是方法性的负面结果：前沿模型与开源模型之间的性能差距主要源于评估提示词的不充分定义，而非模型能力差异。在提示词中定义决策分类后，开源模型与前沿模型无显著差异；此外，完整流程会放大系统性“过度悲观”偏差，而反应层本身能恢复 67-90% 的公众真实关切。","reliability":"论文承认两个失效模式：与蒸馏教师的一致性上升但准确率不升，以及完整流程放大过度悲观偏差。原因在于蒸馏破坏判断独立性，精确匹配评分器强加模式合规上限。","relevance":"该研究直接使用 LLM 模拟人类对产品和政策变化的反应，并与真实结果对照，属于人类仿真实验，且包含批判性分析，值得阅读原文以了解仿真在何种条件下失效。","inspiration":"借鉴其提示词消融和匹配对照设计，可揭示评估方法对模型性能结论的影响，并采用预注册的消融实验定位仿真组件的价值。｜可迁移到政策公告的预期形成研究，如央行利率决议或财政刺激方案的市场反应模拟。｜以 LLM 生成的经济主体（如消费者、投资者）为被试，处理为不同的政策公告文本（含或不含决策分类定义），结果变量为预测的市场反应（如消费、投资决策），对照真实市场数据（如消费者信心指数、资产价格变动）来验证仿真准确性。"}},{"id":"2609.29692","version":1,"title":"Fair Like Us? Auditing LLM Alignment in Resource Allocation","zh_title":"像我们一样公平？审计资源分配中LLM的对齐","abstract":"Fair allocation of scarce, indivisible resources is an important challenge in many societal problems. While there are several formal theories of fairness, no single definition can always be satisfied. As large language models (LLMs) are increasingly used to support decisions and act as agents, they raise new concerns about distributional justice: their judgments are not directly tied to any specific fairness framework and may violate key normative principles. In this work, we introduce a general method for evaluating fairness reasoning in LLMs. We study first-person fairness judgments across a broad set of models and compare them directly with human responses on matched scenarios and elicitation conditions. We find that LLMs tend to prefer stricter fairness constraints than humans, show more self-interested behavior, are sensitive to how information is framed, and are difficult to align with human judgments using fine-tuning with current datasets.","authors":["Qishen Han","Hadi Hosseini","Joshua Kavner","Samarth Khanna","Sujoy Sikdar","Lirong Xia"],"categories":["cs.AI","cs.CY","cs.GT"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29692","pdf_url":"https://arxiv.org/pdf/2609.29692","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","公平分配","人类对照"],"reason":"用LLM模拟人类资源分配判断，并与人类数据对照，评估偏差与对齐难度。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":5,"question":"LLM在资源分配公平性判断上与人类有多大程度的一致，哪些形式化公平准则能解释其判断？","design":"将多个LLM（包括不同规模和推理能力的模型）置于与人类被试相同的公平分配场景中，作为第一人称代理人评估分配结果是否公平或可接受；实验操纵分配满足的公平性质（EF、PROP、EF1、MMS、PROP1等）和诱导条件（框架、信息结构、响应格式），测量模型判定分配可接受的比率。","baseline":"使用Hosseini et al. (2025a)的150名人类被试在相同场景、偏好结构和实验处理下的公平性判断数据作为对照。","findings":"LLM比人类更倾向于严格的公平约束，对EF分配的公平判断率显著高于人类，而对EF1、MMS等放松准则的判断率下降更快；LLM表现出更强的自利行为，对信息框架敏感，且通过微调难以与人类判断分布对齐。","reliability":"论文指出当前人类公平分配数据集规模太小，微调只能使模型坍缩到模态响应而非真正对齐人类判断分布；LLM的判断受诱导问题措辞影响大于分配的形式化属性，且推理能力更强的模型不一定更对齐。","relevance":"该研究直接以LLM作为人类被试的替代品，在资源分配场景中与真实人类数据严格对照，系统评估了仿真偏差和失效条件，对关注LLM仿真可靠性的研究者具有重要参考价值。","inspiration":"借鉴其“同一场景、同一处理、仅替换被试”的严格对照设计，以及通过操纵分配满足的公平性质和诱导框架来分离判断依据的方法｜可迁移到信贷审批中的公平性判断、公共资源分配政策评估、或消费者对价格歧视的公平感知等经济金融场景｜以LLM模拟贷款申请人或政策受众，处理为不同公平准则（如无歧视、比例公平）和框架（如强调个人得失 vs 社会效率），结果变量为接受度或公平评分，对照真实人类实验数据（如调查或实验室实验）来检验LLM的仿真效度。"}},{"id":"2609.29143","version":1,"title":"AI-Moderated Interviews for Market Research and Digital Twins Calibration","zh_title":"用于市场研究和数字孪生校准的AI主持访谈","abstract":"AI-moderated interviews are emerging as a scalable market-research method for generating consumer insights and building consumer \"digital twins.\" Yet it remains unclear whether they match human-moderated interviews or improve on simpler, static data collection methods. In a pre-registered, between-subjects study (N = 317) with three industry partners, we compare AI-moderated (N = 139), human-moderated (N = 24), and static interviews (N = 154). AI moderation matches human moderation in depth, covers more themes, and, holding budget constant, recovers significantly more customer needs than human moderation or static interviews. However, participants sound more emotionally engaged when speaking to a live human. We then create digital twins using interview data and evaluate each twin against the participant's own held-out responses to six real-world marketing stimuli. We find that digital twins created from AI-moderated interviews predict consumer responses better than demographics-only personas. However, the additional richness from AI moderation does not translate into better quantitative predictions compared to static interviews. By analyzing open-ended thoughts generated from humans versus their twins, we find that prediction errors are connected both to differences in (self-reported) thinking styles between twins and humans, and to gaps between training and validation data (i.e., asking questions that are too far out of distribution).","authors":["Yuting Deng","Jingxuan Liu","Olivier Toubia","Naman Jain"],"categories":["cs.CY","cs.AI","cs.HC","cs.MA"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29143","pdf_url":"https://arxiv.org/pdf/2609.29143","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","数字孪生","市场研究"],"reason":"用LLM进行AI主持访谈并构建数字孪生，与真实人类访谈对照，评估预测效度与偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":4,"question":"AI主持访谈能否在同等预算下匹敌人类主持访谈的深度与需求挖掘，并用于构建更准确的消费者数字孪生？","design":"预注册的组间实验，317名消费者随机分配到AI主持访谈、人类主持访谈或静态访谈三种条件，比较访谈的深度、主题覆盖和客户需求数量；随后用访谈数据构建数字孪生，预测参与者对六个真实营销刺激的保留回答。","baseline":"人类主持访谈（N=24）和静态访谈（N=154）作为对照，数字孪生预测与人口统计特征基准和参与者自身保留回答比较。","findings":"AI主持在深度上与人类主持相当，覆盖更多主题，且在预算固定下比人类主持和静态访谈挖掘出更多客户需求；但参与者在与真人交谈时情感参与度更高。基于AI访谈构建的数字孪生预测优于仅人口统计特征的基准，但相比静态访谈并未提升定量预测准确性。","reliability":"论文指出预测误差与数字孪生和人类在自我报告思维方式上的差异有关，也与训练和验证数据之间的分布差距有关（即提问过于超出分布范围）。","relevance":"该研究直接评估了LLM作为人类被试替代品在定性访谈和数字孪生构建中的效度，包含真实人类对照和批判性发现，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其预注册组间实验设计，将AI访谈与人类访谈和静态问卷对比，并用保留样本验证数字孪生的预测效度｜可迁移到消费者金融决策研究，如信贷产品偏好或保险选择，用AI访谈构建个体数字孪生预测金融行为｜以真实消费者为被试，随机分配AI访谈、人类访谈或静态问卷，用访谈数据构建数字孪生预测其对金融产品广告的反应，并与实际选择数据对照。"}},{"id":"2609.29370","version":1,"title":"From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring","zh_title":"从政策文件到结构化调查回答：评估大语言模型用于政策监测","abstract":"Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are costly and challenging to scale across countries. Large language models (LLMs) enable new possibilities for extracting and structuring information from long and unstructured policy documents. This paper presents an application of LLMs as \"AI respondents\" for generating structured survey responses from policy texts. We develop a data extraction pipeline based on long-context in-context learning to map information from public web sources into predefined survey categories, including policy instruments, target groups, and thematic areas. The pipeline integrates a validation step using a secondary LLM to assess relevance and evidence, alongside comparisons with human-provided responses. Using a multi-country dataset, we evaluate the alignment between LLM-generated and human-generated outputs through overlap measures and cross-validation. Results show that LLMs achieve high agreement for structured indicators (84-95%), while differences remain in free-text fields, where models tend to provide more detailed procedural descriptions. These findings highlight the potential of hybrid human-AI workflows for policy monitoring, improving both efficiency and scalability while maintaining the need for human validation and contextual interpretation.","authors":["Carolyn Cole","Matthias Deschryvere","Toqeer Ehsan","Arash Hajikhani"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29370","pdf_url":"https://arxiv.org/pdf/2609.29370","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","政策监测","人类对照"],"reason":"用LLM从政策文本生成结构化调查回答，并与人类回答对照，属于仿真人类被试且有人…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":8,"question":"能否用大语言模型从政策文本中自动生成结构化调查回答，以替代或辅助人工政策监测？","design":"使用长上下文上下文学习（long-context in-context learning）的LLM（如GPT-4o）作为“AI受访者”，从网页抓取的政策文本中提取信息，映射到预定义的调查类别（政策工具、目标群体、主题领域），并生成自由文本字段（描述和目标）。通过一个二级LLM验证层评估相关性和证据，并与人类提供的回答进行比较。","baseline":"来自EC-OECD STIP Compass调查的人类专家回答，覆盖六个OECD国家（加拿大、芬兰、德国、韩国、西班牙、土耳其）的政策举措。","findings":"LLM在结构化指标（政策工具、目标群体、主题代码）上与人类回答的一致性达到84-95%，但在自由文本字段上存在差异，模型倾向于提供更详细的程序性描述，而人类更强调背景和社会影响。","reliability":"论文承认LLM在自由文本字段上与人类存在差异，可能引入系统性偏差，且政策数据具有异质性和制度嵌入性，公开来源可能无法完全捕捉；需要人类验证和上下文解释。","relevance":"该研究直接使用LLM作为人类受访者的替代品，并与真实人类数据对照，评估仿真可靠性，符合研究者对LLM仿真实验和批判性评估的兴趣，值得阅读原文以了解具体方法和偏差分析。","inspiration":"方法上，该研究展示了如何利用长上下文提示和二级LLM验证来从非结构化文本中提取结构化数据，并设计人类对照来评估一致性。｜可迁移到经济金融领域，如从公司年报、政策文件或新闻中自动提取结构化信息，用于构建经济指标或评估政策影响。｜研究设计雏形：使用LLM从上市公司年报中提取财务和非财务信息（如研发支出、风险因素），与人工标注或数据库中的真实数据进行对照，评估LLM提取的准确性和偏差，并分析在哪些条件下LLM表现不佳。"}},{"id":"2609.28486","version":1,"title":"Political Sorting Can Drive AI Models Apart Through User Feedback","zh_title":"政治分类可通过用户反馈使AI模型分化","abstract":"Large language models are rapidly becoming an important source of political information. This raises a fundamental question: will AI systems support a shared basis for political knowledge, or lead different political groups to rely on increasingly different models? Political sorting can drive model fragmentation if three conditions hold: politically different users select into different models, learning from user feedback pushes those models apart politically, and the resulting differences shape subsequent model choices. We call this self-reinforcing process the Centrifugal Alignment Spiral. We study its components in three steps. First, we draw on a human experiment showing that political identity predicts model choice. Second, we fine-tune language models on synthetic feedback reflecting predominantly Democratic or Republican preferences. Across five independent runs per model family, paired models diverged on 12-41% of unseen survey questions with large partisan gaps, and in every run the differences moved in the expected political direction; for some models, differentiation extended even to issue areas excluded from training. Pooling feedback across groups instead suppressed divergence. Third, an empirically anchored agent-based model shows what follows when political sorting and model adaptation operate together: models attract politically distinct audiences, learn from them, and diverge further. User feedback can therefore turn political sorting among AI users into durable differences between the models on which they rely for political information.","authors":["Petter T\\\"ornberg","Michael Heseltine","Nicol\\`o Pagan","Christopher Bail","Michelle Schimmel","Christopher Barrie"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28486","pdf_url":"https://arxiv.org/pdf/2609.28486","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","政治极化","人类对照"],"reason":"用LLM模拟政治反馈并对照人类实验，研究模型分化，涉及政策场景与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":7,"question":"政治排序能否通过用户反馈导致不同政治群体使用的AI模型在政治上分化，形成自我强化的离心对齐螺旋？","design":"研究分三步：首先利用已有的人类实验证明政治身份预测模型选择；其次用反映民主党或共和党偏好的合成反馈微调Qwen2.5-1.5B、Mistral-7B和GPT-OSS-20B模型，每个模型家族进行五次独立运行，测量配对模型在未见过的调查问题上答案分歧的比例和方向；最后构建基于实证的智能体模型，模拟政治排序和模型适应耦合下的动态演化。","baseline":"人类基准来自先前报告的人类模型选择实验，显示政治身份预测模型选择，包括付费准确回答条件下共和党人更可能选Grok、民主党人更可能选Claude，71%的参与者返回之前偏好的模型。","findings":"在90/10的受众对比压力测试下，配对模型在12-41%的未见调查问题上产生分歧，且每次运行中平均差异都符合反馈的政治方向；部分模型的分化扩展到未训练的政治领域。混合不同群体的反馈则抑制了分化。","reliability":"论文承认90/10的受众对比是压力测试，不代表当前AI市场的实际排序程度；分化程度因模型、领域和提示而异；身份线索实验表明谄媚个性化可能减少模型间分化，但增加模型内分化。","relevance":"该研究直接探讨LLM作为政治信息源时的分化机制，通过合成反馈模拟政治群体偏好，并与人类实验对照，涉及政策评估场景和失效条件（如反馈混合、个性化），对关注LLM仿真可靠性及偏差的研究者具有重要参考价值。","inspiration":"借鉴其用合成反馈微调模型并测量泛化分化的方法，可迁移到经济金融中的群体偏好分化问题，如不同收入或风险偏好群体的金融建议模型分化。｜可应用于信贷审批或投资建议场景，研究用户反馈如何导致模型对不同群体产生差异化行为。｜设计：用不同风险偏好或金融素养的合成用户反馈微调金融LLM，测量其在未见金融决策问题上的行为差异，并与真实人类金融决策数据（如调查或实验数据）对照，检验分化是否与真实群体差异一致。"}},{"id":"2609.27535","version":1,"title":"KITE: Scaling Jev Population Experiments with Sparse Flagship Calibration","zh_title":"KITE：通过稀疏旗舰校准扩展Jev人口实验","abstract":"KITE queries a typed behavioral kernel once per unique state, then executes populations of any size from the table with event-keyed randomness and common random numbers. An expensive flagship model is reserved for sparse paired anchors that estimate intervention effects. Measured human-model discrepancy is propagated as shared error into every conclusion. Population-experiment cost thus scales with unique states and anchors, while uncertainty is governed by evidence about people rather than Monte Carlo noise. On Epstein experiments with 9,070 participants, anchors covering 1.7% of states reduced effect error by 41% (absolute MAE reduction 0.0125). On 37 held-out SocSci210 experiments, 0.5-1.5% anchor coverage raised captured decision gain from 0.27 to 0.39. The kernel passed content-fidelity criteria in all 15 new countries of a 16-country study. Shared discrepancy yielded retrospective coverage of 93% and 96% at nominal 80% and 90%, versus 29% and 36% from human sampling uncertainty alone. A million agents executed 20 tabulated steps in 0.9 seconds on a laptop. This architecture offers a route to screening candidate interventions before human trials, multi-country content audits, and uncertainty-aware policy comparison at the cost of a few thousand kernel calls with sparse flagship anchors. Property-specific evidence records connect each use to its validation scope, correction provenance, and uncertainty, making these applications auditable.","authors":["Hengyu Li (The University of Tokyo)"],"categories":["cs.MA","cs.CY"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27535","pdf_url":"https://arxiv.org/pdf/2609.27535","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","人类数据对照","政策评估"],"reason":"用LLM仿真人类被试，有真实人类数据对照，评估偏差并传播不确定性，用于政策评估。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":5,"question":"如何在大规模人口实验中，用廉价的类型化行为核模型结合稀疏的旗舰模型校准，来估计干预效应并传播人类-模型差异的不确定性？","design":"KITE 使用 TypeSafe 的 Jev 行为核模型（类型化行为核）对每个唯一状态查询一次，生成决策分布，然后用表格化执行模拟任意规模的人口，并通过事件键控随机数和共同随机数控制变异性；昂贵的旗舰模型仅用于稀疏的配对锚点，以估计干预效应；人类-模型差异被作为共享误差传播到所有结论中。","baseline":"对照的真实人类数据包括 Epstein 实验（9,070 名参与者）和 37 个留出的 SocSci210 实验，以及一个 16 国研究中的内容保真度标准。","findings":"在 Epstein 实验中，覆盖 1.7% 状态的锚点将效应误差降低了 41%（绝对 MAE 降低 0.0125）；在 37 个留出的 SocSci210 实验中，0.5-1.5% 的锚点覆盖率将捕获的决策增益从 0.27 提高到 0.39。共享差异传播在名义 80% 和 90% 的置信水平下分别实现了 93% 和 96% 的回顾性覆盖率，而仅使用人类抽样不确定性时分别为 29% 和 36%。","reliability":"论文承认的局限包括：仅有两个留出的人类参考数据集测试混合方法；效应幅度需要进一步校准，差异斜率在不同研究选择间变化；稀疏校准可能引入旗舰模型误差；核模型可能过度预测规范信息；实验不授权预测文献中不存在的干预；记忆一致性是局部的，队列级边际校正可能抹去真实的持续性或处理路径；刺激重建、省略卡片图像、人口统计压缩等限制了保真度；专有模型和访问限制阻碍了可重复性。","relevance":"该研究直接针对用 LLM 仿真人类被试的核心问题，提供了与真实人类数据对照的验证，并传播不确定性，对评估仿真可靠性和偏差具有重要参考价值，值得精读原文。","inspiration":"该方法通过稀疏旗舰模型校准和共享误差传播，在保持低成本的同时提高了效应估计的准确性，值得借鉴其校准策略和不确定性量化方法。｜可迁移到政策评估场景，如税收政策变化对劳动供给的影响、福利项目对消费行为的影响，或信息干预对金融决策的影响。｜设计一个实验：用 Jev 核模型模拟不同人口群体对政策公告的反应，以真实调查数据（如消费者预期调查）为基准，施加政策处理（如利率变化信息），测量预期调整和消费意愿，并用稀疏旗舰模型校准关键状态，传播人类-模型差异。"}},{"id":"2609.28470","version":1,"title":"StudentBench: AI and human tutoring yield equivalent GRE learning gains","zh_title":"StudentBench：AI与人类辅导在GRE学习收益上等效","abstract":"Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.","authors":["Curtis Northcutt","Inaara Hasmani","Kevin Feng","Trevor Khangi","Andreas Plesner","Jonas Mueller"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28470","pdf_url":"https://arxiv.org/pdf/2609.28470","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","教育实验","人类对照"],"reason":"用LLM替代人类导师进行教学实验，并与人类导师对照，评估学习效果，属于人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":6,"question":"大语言模型（LLM）作为AI导师能否在GRE学习收益上达到与人类专家导师统计等效的效果？","design":"用多个LLM（如Gemma 4 31B等）扮演AI导师，对2383名人类参与者进行GRE定量和语文部分的辅导，同时设置人类导师辅导组和无辅导对照组，测量学习收益（前后测成绩提升百分比），并收集175,000条学生-AI消息分析对话行为。","baseline":"人类专家导师辅导组的学习收益数据，以及无辅导对照组的学习收益数据。","findings":"AI辅导在GRE学习收益上与人类专家辅导统计等效（p=.015），且在七个GRE领域中五个领域的最佳AI导师平均超过人类导师。一个AI导师（Gemma 4 31B）以低918倍的成本实现了与人类辅导等效的学习收益（p=.044）。","reliability":"论文承认局限：只测量即时学习收益，未评估长期保持；参与者均为能读写英语的成年人，未测试跨语言、设备或教育环境；未设置学生独自练习的对照组，无法分离AI交互的额外收益；排行榜评估基于专家评价和对话行为，而非实际学习效果。","relevance":"该研究用LLM替代人类导师进行教学实验，并与真实人类导师及无辅导组对照，评估学习效果，属于典型的人类仿真研究，且提供了大规模真实人类数据作为基准，值得精读以了解仿真等效性检验的设计与局限。","inspiration":"借鉴其多组对照设计（AI处理、人类处理、无处理）和统计等效性检验（TOST）来严格评估AI干预是否非劣于人类专家。｜可迁移到金融教育或投资者决策辅导场景，例如测试AI投教助手能否在提升投资者金融素养或改善投资决策上达到人类理财顾问的效果。｜以真实投资者为被试，随机分为AI投教组、人类顾问组和无辅导组，处理为一段时间的个性化金融知识辅导，结果变量为金融素养测试得分或模拟投资组合表现，对照真实人类顾问组和无辅导组的数据，并采用等效性检验。"}},{"id":"2609.28372","version":1,"title":"Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer","zh_title":"算法购物：代理式AI如何将人类启发式用作替代消费者","abstract":"Consumers increasingly delegate purchasing decisions to Large Language Models (LLMs) acting as surrogate consumers. Using \"Tool-Lab,\" an adaptation of information-board process tracing that places product attributes behind costly tool calls, we examine how marketing pricing cues (i.e., just-below pricing and promotional framing) influence AI shopping agents. Across eight commercially deployed LLMs from three providers, we trace pre-choice information acquisition. Under zero cost, pricing cues rarely mislead. Imposing acquisition costs under a vague goal prompt leads LLMs to omit diagnostic attributes required to compute unit price and choose suboptimal choices resembling human heuristics. Relative to a specific goal prompt that mainly preserves diagnostic search and choice optimality, a vague goal prompt under constraints creates a search-mediated vulnerability. This research demonstrates that marketing heuristics in delegated AI shopping are governed by storefront information architecture, not necessarily immutable LLM flaws.","authors":["Davood Wadi","Yu Ma"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28372","pdf_url":"https://arxiv.org/pdf/2609.28372","source_feed":"econ.GN","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","消费者行为","算法保真度"],"reason":"用LLM作为代理消费者模拟人类购物决策，并与人类启发式对照，涉及营销实验场景。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":9,"question":"在委托AI代理购物时，营销定价线索（如尾数定价和促销框架）如何通过信息获取成本与目标提示的具体性影响LLM的信息搜索和选择最优性？","design":"使用Tool-Lab实验范式，将产品属性隐藏在需要付费的工具调用之后，操纵信息获取成本（0、1、5美分）和提示目标具体性（模糊：“找最划算的” vs. 具体：“找每盎司最低价”），对8个商用LLM（来自Google、OpenAI、Anthropic）进行咖啡选择实验，测量信息搜索深度、搜索组成和选择最优性。","baseline":"无对照","findings":"在零成本或具体目标提示下，LLM大多做出规范最优选择；但在获取成本存在且目标模糊时，多数LLM会减少搜索深度，省略诊断性属性（如美分或重量），导致次优选择，类似于人类启发式决策。","reliability":"论文未讨论","relevance":"本研究通过实验操纵环境约束（成本与提示），揭示了LLM启发式行为的条件性，为评估LLM作为人类被试替代品的可靠性提供了关键证据，值得精读原文以理解其方法细节和边界条件。","inspiration":"借鉴其通过工具调用施加信息获取成本并操纵提示具体性的设计，可迁移到消费者金融决策（如信用卡选择、贷款比较）或投资者信息处理场景；例如，用LLM模拟投资者在获取公司财务指标需付出成本时，模糊目标（“选只好股票”）与具体目标（“选市盈率最低的股票”）下的信息搜索与选择，并与真实投资者眼动或点击流数据对照。"}},{"id":"2608.19621","version":3,"title":"Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories","zh_title":"用纵向生命轨迹缓解LLM智能体的身份本质主义","abstract":"Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce population-level patterns, yet often fail to capture human-like diversity. Our analysis shows that static-profile agents exhibit stronger demographic separation and within-group compression than humans, a pattern consistent with identity essentialism: demographic labels can encourage models to treat group-average tendencies as individual traits, homogenizing responses within groups. We argue that this limitation arises from two related factors: sparse, static agent representations and the limited ability of prompt-only memory to persistently integrate experience. Inspired by complementary memory systems, we propose LifeMem, a longitudinal memory framework that combines structured life-event retrieval with agent-specific parametric memory for experience integration. Experiments on Understanding Society with three LLMs show that LifeMem improves alignment with human data in terms of response distributions, overall and within-group diversity, and patterns of within-person response change across life stages. These findings highlight the value of longitudinal life-event memory for constructing more faithful and dynamically evolving social agents.","authors":["Hexi Wang","Yujia Zhou","Bangde Du","Weihang Su","Xinyuan Cao","Qingyi Pan","Qingyao Ai","Yueyue Wu","Min Zhang","Yiqun Liu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-23","first_seen":"2026-08-21","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2608.19621","pdf_url":"https://arxiv.org/pdf/2608.19621","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人类数据对照","社会调查"],"reason":"用LLM agent模拟人类调查数据，并与真实面板数据对照，改进仿真保真度。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":1,"question":"如何通过引入纵向生活轨迹记忆来缓解LLM社会仿真中的身份本质主义，从而提升模拟人类多样性与动态变化的保真度？","design":"使用三个指令微调LLM（Llama-8B、Ministral-8B、Qwen3.5-9B）作为社会仿真智能体，基于Understanding Society面板数据构建个体静态画像和纵向生活事件流，施加LifeMem框架（结构化生活事件记忆+个体特定LoRA参数记忆）作为处理，与静态画像、多样性提示、非参数记忆等基线对比，测量回答分布、组内多样性、组间差异及个体跨波次变化等结果变量。","baseline":"Understanding Society英国家庭纵向调查的真实个体面板数据，包括静态背景信息和多波次生活事件及主观态度回答。","findings":"静态画像智能体表现出更强的组间分离和组内压缩，符合身份本质主义特征；LifeMem通过结合结构化事件检索和参数化记忆整合，在回答分布、总体及组内多样性、跨生命阶段个体变化模式上均提升了与人类数据的一致性。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真中多样性塌缩和身份本质主义问题，使用真实面板数据作为基准，并提出了可操作的记忆框架，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文了解具体实现和效果。","inspiration":"借鉴其双记忆系统设计，将显式事件检索与参数化个体状态结合，可迁移到经济金融中的个体决策仿真，如消费者跨期选择或投资者行为演化。｜例如，在信贷审批歧视研究中，用LLM智能体模拟不同背景的申请人，施加纵向财务事件记忆处理，测量审批决策的组间差异和组内多样性，并与真实信贷数据对照。｜设计：以LLM智能体作为虚拟被试，处理为是否注入个体历史财务事件（如收入变动、失业）并更新LoRA参数，结果变量为贷款审批通过率和利率设定，对照真实信贷记录数据评估仿真偏差。"}},{"id":"2609.25066","version":1,"title":"Understanding Reliability in LLM-based Human Behavior Simulation","zh_title":"理解基于LLM的人类行为仿真的可靠性","abstract":"Large language models (LLMs) are increasingly used to simulate human survey responses and behavioral reactions, yet unreliable simulations can mislead social science conclusions. However, existing evaluations focus on end-to-end scores, leaving it unclear how different aspects of the simulation process interact to determine reliability. We propose ReliMap, which decomposes LLM-based human behavior simulation into three structured layers and evaluates reliability at both the individual level (R1) and population level (R2) across three configuration dimensions: model capacity, profile completeness, and population coverage. Through experiments across four simulation tasks and eleven LLMs, we find that all models exhibit substantial distributional bias without profile conditioning. Profile conditioning reduces this bias with diminishing returns. Larger models benefit more, and attribute informativeness matters more than quantity. Critically, R1 gains do not reliably transfer to R2--individual and population-level reliability can move in opposite directions. At the population layer, increasing coverage reduces variance but not systematic bias, with R2 stabilizing at around 50-100 individuals. These findings highlight that reliable simulation cannot be achieved by optimizing any single layer in isolation, but requires coordinated improvement across all three.","authors":["Pei Wang","Lei Wang","Yuanzi Li","Xu Chen"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25066","pdf_url":"https://arxiv.org/pdf/2609.25066","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B4"],"tags":["LLM仿真","可靠性评估","人类行为"],"reason":"直接研究LLM仿真人类行为的可靠性，分解评估层次，含真实人类数据对照，并指出失…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":3,"question":"LLM 仿真人类行为时，模型容量、画像完整度和人群覆盖度如何共同影响个体层与群体层的可靠性？","design":"用 11 个 LLM 在 4 个任务（党派偏好、移民态度、宗教立场、社交媒体事件态度）上仿真人类回答，通过改变模型容量、画像属性数量和人群样本量，测量个体准确率（ACC）和群体分布距离（TVD）。","baseline":"真实人类调查数据：欧洲社会调查（ESS）、世界价值观调查（WVS）、SocioBench 宗教立场数据，以及社交媒体上关于瑞幸咖啡股价暴跌的真实公众态度语料。","findings":"无画像条件时所有模型都存在显著分布偏差；画像条件化可减少偏差但边际收益递减，且大模型受益更多。个体层可靠性提升不必然转化为群体层可靠性，群体层在 50-100 人后趋于稳定但系统偏差不随覆盖度增加而减小。","reliability":"论文指出个体层与群体层可靠性可能反向变动，仅优化单一层无法保证整体可靠性；画像属性信息量比数量更重要，但未明确给出所有失效条件，仅强调需三层协同改进。","relevance":"该研究系统拆解了 LLM 仿真人类行为的可靠性层次，并基于真实调查数据对照，直接回应了仿真在什么条件下会失效的问题，对关注经济学实验和政策评估仿真的研究者很有参考价值。","inspiration":"借鉴其分层评估框架，将个体预测准确率与群体分布距离分开考察，并系统变化模型、画像和样本量以定位偏差来源｜可迁移到消费者金融决策仿真，如信用卡选择、退休储蓄计划参与或风险偏好调查｜用多个 LLM 扮演不同人口学特征的消费者，施加不同金融素养或收入冲击处理，测量选择分布，并与美国消费者金融调查（SCF）或美联储家庭经济决策调查（SHED）的真实数据对照，检验仿真在个体和群体层的可靠性。"}},{"id":"2609.25010","version":1,"title":"Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation","zh_title":"合成人物角色能预测真实受众反应吗？一项无人物角色基线优于基于人物角色文案仿真的仿真到现实研究","abstract":"Marketers increasingly use large language models (LLMs) as \"synthetic personas\" to predict how an audience will react to a piece of copy before it ships, encouraged by evidence that profile-conditioned LLMs mimic human samples. But is that prediction actually valid against real behaviour - and does the persona machinery help? We present a sim-to-real validity study using the Upworthy Research Archive - thousands of headline A/B tests on shared real traffic, with measured click-through - as held-out ground truth. We compare a ten-persona panel, grounded in the real audience's demographics, against a no-persona zero-shot baseline that simply asks the model how likely a typical reader is to click. Two findings stand out. First, ground-truth reliability is the binding constraint: most A/B tests have no statistically distinguishable winner, so validity can only be measured on the reliable subset (n = 399). Second, and counter to the persona-simulation premise, persona conditioning degrades predictive validity: the no-persona baseline ranks variants markedly better (Kendall {\\tau} = 0.361, a medium effect; top-1 accuracy 49.2%) than the persona panel ({\\tau} = 0.084; top-1 34.6%), with non-overlapping confidence intervals. Asking the model directly taps an accurate population-level prior; forcing it to role-play specific personas injects bias and noise. The result replicates across three independent Upworthy splits, holds in direction on a different-domain news dataset, and is robust to seed, prompt phrasing, and model choice - across three Gemini tiers and a different model family (OpenAI gpt-4.1, significant paired gap). The takeaway: for predicting aggregate engagement, a plain LLM ranker beats persona simulation - synthetic personas are not merely a weak predictor, they are worse than not using them. All numbers regenerate from a public, artifact-first replication package.","authors":["Alexandre Cristov\\~ao Maiorano"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25010","pdf_url":"https://arxiv.org/pdf/2609.25010","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","受众预测","算法保真度"],"reason":"用LLM仿真受众点击行为，与真实A/B测试数据对照，发现persona降低预测…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":1,"question":"在真实A/B测试数据上，基于合成人设的LLM仿真能否预测真实受众对文案的点击行为，且人设条件化是否比无人设基线更有效？","design":"使用Upworthy Research Archive中的数千个标题A/B测试作为真实流量数据，构建基于真实受众人口统计特征的十人设面板，用LLM（gemini-3.1-flash-lite）分别以人设条件和无人设零样本基线预测标题点击意图，聚合为排名，与真实点击率排名比较。","baseline":"Upworthy Research Archive中真实A/B测试的点击率数据，作为留出集真实行为基准。","findings":"大多数A/B测试无统计显著赢家，因此只能在可靠子集（n=399）上评估效度；人设条件化降低预测效度，无人设基线排名显著优于人设面板（Kendall τ=0.361 vs 0.084，top-1准确率49.2% vs 34.6%）。","reliability":"论文指出真实数据可靠性是主要约束，多数A/B测试无显著赢家；人设仿真在可靠子集上仍表现差，且结果对种子、提示措辞和模型选择稳健，但人设条件化本身引入偏差和噪声。","relevance":"该研究直接检验LLM合成人设仿真在真实受众行为预测中的效度，发现人设条件化反而降低预测力，对关注LLM仿真可靠性及偏差的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其sim-to-real效度框架和可靠子集筛选方法，用真实行为数据作为基准，比较不同仿真策略（如人设 vs 无人设）的预测效度，并采用排名指标和bootstrap置信区间｜可迁移到经济金融中的消费者选择预测，如广告文案对点击率的影响、金融产品描述对投资意愿的影响、政策沟通对公众反应的影响等｜以真实A/B测试数据（如某平台广告实验）为基准，用LLM分别以人设面板和无人设基线预测用户对金融产品广告的点击或选择，比较排名准确率，并筛选出有显著差异的测试子集进行评估。"}},{"id":"2609.25677","version":1,"title":"Seeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing","zh_title":"眼见不为实：合成消费者何时能及不能预测试视觉营销","abstract":"Marketers now deploy generative AI agents as synthetic consumers to pretest visual assets such as logos, packaging, and advertising at a fraction of human-panel cost. However, this procedure assumes that a model seeing a visual cue can also perceive its consumer meaning, which is largely untested. We stress-test the assumption using six canonical visual marketing experiments, varying the two levers managers control: model generation (GPT-4o-mini vs. GPT-5.4-mini) and input format (plain text vs. JSON). Every resulting configuration passed the manipulation checks; however, none of the configurations reproduced more than two of the six human effects, and the remainder were nonsignificant. The one exception was a significant reversal of the human pattern. Providing conceptual or empirical evidence through in-context learning steers average responses toward the human effect. Yet steering has a limit: even when it succeeds, a configuration reproduces less than half of the natural spread of human responses and so understates consumer heterogeneity. We integrate these results into an AI governance protocol (Calibrate, Intervene, Deploy) that delineates when synthetic consumers can responsibly screen creatives and when human panels remain necessary.","authors":["Yi-Lin Tsai (Arvin)","Yung-Hsiu (Arvin)","Lai"],"categories":["cs.AI","cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25677","pdf_url":"https://arxiv.org/pdf/2609.25677","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","消费者行为","算法保真度"],"reason":"用LLM作为合成消费者复现视觉营销实验，与真实人类数据对照，评估仿真可靠性并指…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":2,"question":"合成消费者在视觉营销实验中能否像人类一样感知视觉线索并复现人类判断，其失效边界和可修复性如何？","design":"使用 GPT-4o-mini 和 GPT-5.4-mini 两个模型，以纯文本和 JSON 两种输入格式，模拟人类被试回答六个经典视觉营销实验，测量数值评分和文本理由，并通过操纵检查、主题建模和上下文学习（概念证据和实证证据）进行干预。","baseline":"六个经典视觉营销实验的原始人类样本数据（每个研究 69 到 220 名被试）。","findings":"所有配置都通过了视觉操纵检查，但没有一个配置能复现超过两个人类效应，其余均不显著，甚至出现一个显著反转。提供概念或实证证据的上下文学习能将平均响应拉向人类效应，但即使成功，也无法复现人类响应自然分布的一半，低估了消费者异质性。","reliability":"论文承认合成消费者在视觉领域感知不一致，即使通过上下文学习校准，也无法恢复消费者异质性，因此只适合平均效应问题，不适合细分或定位。","relevance":"该研究直接检验了 LLM 作为人类被试在视觉营销实验中的可靠性，与真实人类数据对照，并揭示了失效条件，对关注仿真偏差和边界的研究者具有重要参考价值。","inspiration":"借鉴其多模型、多输入格式的因子设计，以及通过操纵检查和主题建模分离“看见”与“感知”的方法，可迁移到经济金融中的视觉信息处理场景，如央行沟通中的图表设计、金融产品广告或信贷审批中的视觉线索。｜设计一个实验，用 LLM 模拟投资者，呈现不同颜色或形状的金融图表（如涨跌颜色、风险提示图标），测量其风险感知和投资决策，并与真实投资者实验数据对照，检验 LLM 是否复现视觉线索对风险偏好的影响。"}},{"id":"2609.25760","version":1,"title":"The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance","zh_title":"模拟社会的局限：后训练与调查微调如何抹除跨文化方差","abstract":"Using large language models (LLMs) to simulate diverse human populations has the potential to transform many aspects of computational social science, yet many evaluations score the average response rather than the spread of opinion within real groups. Here, we develop a diagnostic framework that measures point accuracy alongside dispersion retention, the ratio of predicted to human standard deviation ($\\dr$), on 10{,}000 respondent--question pairs from the World Values Survey (WVS) spanning twelve countries and six continents. We evaluate eleven zero-shot language models and five variants fine-tuned on WVS data with SFT, DPO, and GRPO. We identify a failure mode we term \\textit{consensus collapse}, where alignment training compresses outputs toward one stereotype per group. Along the post-training trajectory from the Llama~3.1 70B base to the Tulu~3 checkpoints, the first stage, supervised instruction tuning, removes half of the spread with minimal accuracy gain ($\\dr$ 1.22 to 0.59; accuracy $+0.9$ points), the later stages do not restore it, and a gap opens between WEIRD and non-WEIRD countries that survey fine-tuning then deepens while pursuing higher point accuracy. The most accurate model (Tulu~3 70B-DPO fine-tuned on WVS, 57.9\\%) keeps half the human spread overall ($\\dr = 0.50$) and 11\\% of it for Nigeria, against 0.70--0.87 for WEIRD countries. Raising the sampling temperature to 1.0 leaves the Wasserstein-1 distance ($\\wone$) to human distributions unchanged for both fine-tuned DPO models, and GRPO on Qwen~3.5 9B does not restore the spread under either an accuracy reward or a distribution-shaped reward. Mixing the aligned model with an unaligned prior raises $\\dr$ from 0.51 to 0.62 on a held-out split but leaves Nigeria at 0.36. Point accuracy alone therefore misjudges these simulators, and current post-training trades diversity for consensus.","authors":["Rojin Ziaei"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25760","pdf_url":"https://arxiv.org/pdf/2609.25760","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","算法保真度","跨文化调查"],"reason":"直接评估LLM仿真人类调查回答的分布保真度，使用WVS真实数据对照，并揭示后训…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":3,"question":"LLM后训练与调查微调如何影响其模拟人类调查回答时的跨文化方差保留？","design":"使用WVS第7波12国10,000个受访者-问题对，构建包含人口统计和Inglehart-Welzel文化维度的价值编码persona，评估11个零样本LLM和5个在WVS数据上微调（SFT、DPO、GRPO）的变体，测量点准确率、MAE、Wasserstein-1距离和偏差比（预测标准差/人类标准差）。","baseline":"世界价值观调查（WVS）第7波12个国家的真实个体回答分布，包括标准差和分布形状。","findings":"后训练导致“共识坍缩”：监督指令微调使偏差比从1.22降至0.59，准确率仅提升0.9个百分点；后续DPO和GRPO未恢复方差，且WEIRD与非WEIRD国家差距扩大，最准确模型（Tulu 3 70B-DPO微调）整体偏差比0.50，尼日利亚仅0.11。提高采样温度至1.0不改变Wasserstein-1距离，GRPO在准确率或分布形状奖励下均不能恢复方差，混合未对齐先验仅将整体偏差比从0.51提升至0.62，尼日利亚仍为0.36。","reliability":"论文承认共识坍缩在非WEIRD国家更严重，且无法通过提高温度或GRPO恢复；混合未对齐先验只能部分缓解，不能根本解决。未讨论其他局限。","relevance":"直接针对LLM仿真人类调查回答的分布保真度，使用真实WVS数据对照，揭示后训练和微调对多样性的系统性压缩，对关注仿真可靠性与偏差的研究者极具参考价值。","inspiration":"借鉴其诊断框架：同时测量点准确率和分布离散度（偏差比、Wasserstein距离），并沿后训练轨迹分解方差损失来源。｜可迁移到经济政策评估中的异质性反应仿真，如不同文化背景下对税收或福利政策的态度分布。｜用LLM模拟多国受访者对政策的态度，以WVS或类似跨国调查为真实基准，比较零样本、指令微调和偏好优化模型在点准确率与方差保留上的权衡，并检验温度调节和先验混合能否恢复分布。"}},{"id":"2609.25059","version":1,"title":"Can Large Language Model-Generated Responses Support Assessment Development? A Human-Calibrated Rasch Benchmark","zh_title":"大语言模型生成的回答能否支持评估开发？一项人类校准的Rasch基准研究","abstract":"Large language models (LLMs) are proposed as synthetic respondents for pilot testing, but their usefulness depends on whether they supply the evidence assessment development requires. We calibrated rating scale models on 14 digital-use skill items from 6,245 adults and used the human item parameters to evaluate responses generated for 1,300 demographically matched personas. LLM responses had high internal consistency ($\\alpha \\approx .94$) but used the lowest category in 0.2-0.3% of responses versus 13.7-21.5% for humans, and no persona chose it on every item. These gaps changed the response-category judgment on both subscales and the PC targeting judgment; on the human-calibrated scales, 12 of 14 items had infit below 0.70, indicating responses more predictable than the Rasch model expects. Preregistered changes to the prompt, category order, and persona information did not restore the human lower range. High internal consistency is insufficient evidence that LLM responses can replace human pilot data.","authors":["Eunjeong Song","Sehee Hong"],"categories":["stat.AP","cs.CY"],"primary_category":"stat.AP","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25059","pdf_url":"https://arxiv.org/pdf/2609.25059","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真被试","Rasch模型","人类数据对照"],"reason":"用LLM生成合成被试回答，并与真实人类数据校准，评估其替代可行性，发现偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":6,"question":"LLM生成的回答能否支持评估工具开发中关于反应类别结构、目标定位和项目审查的判断？","design":"使用LLM为1300个人口统计学匹配的虚拟人物生成对14个数字技能自评项目的回答，并改变提示指令、类别顺序和人物信息以检验对低类别使用的影响。","baseline":"来自2025年数字鸿沟调查的6245名成年人的人类回答，用于校准Rasch模型并作为固定参照。","findings":"LLM回答内部一致性高（α≈.94），但最低类别使用率远低于人类（0.2-0.3% vs 13.7-21.5%），且没有虚拟人物在所有项目上选择最低类别。在人类校准的量尺上，14个项目中有12个的infit低于0.70，表明LLM回答比Rasch模型预期的更可预测，且提示、类别顺序和人物信息的改变未能恢复人类低端分布。","reliability":"论文指出高内部一致性不足以证明LLM回答可替代人类试点数据；人口统计学匹配不能保证再现测量构念的变异；研究为回顾性基准，不适用于儿童或青少年。","relevance":"该研究直接针对LLM作为合成被试的可靠性，用真实人类数据校准并发现系统性偏差，对关注仿真失效条件的研究者具有重要参考价值。","inspiration":"借鉴其用人类校准的Rasch模型作为固定参照来评估LLM回答偏差的方法，可迁移到经济金融中的主观量表或调查数据仿真（如消费者信心、风险态度）。｜可应用于政策评估中的问卷预测试，例如用LLM模拟不同人口群体对政策的态度分布。｜设计：以真实家庭调查（如美国消费者财务调查）为基准，用LLM生成匹配人口特征的虚拟受访者对风险偏好或通胀预期问题的回答，比较分布差异和项目反应模型拟合，检验提示工程能否校正偏差。"}},{"id":"2609.24012","version":2,"title":"Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure","zh_title":"检验而非假定充分性：针对涌现网络结构校准生成式社会模拟器","abstract":"Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check. We present an adequacy-aware calibration protocol that couples amortized posterior estimation with a synthetic identifiability assessment, a matched-sample-size adequacy check (prior-predictive reachability plus per-statistic posterior-predictive localization), a diagnosis-guided repair, and a statistic-held-out audit. We demonstrate it on a real second-hand luxury resale market with four channel-by-residency cells, each a bipartite buyer-brand network, using a forward model built from persona profiles elicited once, offline, by a language model. The behavioural parameters are recoverable in all four cells, though calibration is approximate and overconfident for one parameter. The observed summary falls outside the simulator's reachability reference in every cell, with the mean purchased tier as the pervasive discrepancy. The repair meets the value-block criterion in two of four cells but does not restore adequacy, and the held-out audit surfaces a buyer-breadth-dispersion miss no earlier diagnostic detected. A profile-source ablation finds the language-model profiles beat a flat rule baseline in all four cells, yet within-category brand relabelling causes no consistent degradation, so the profiles are a partially validated input whose value rests on structure, not brand identity. Making no causal claim, we conclude that an independent-aggregation account, without agent interaction or a buyer-breadth mechanism, cannot jointly reproduce the market's purchased-tier level, head-brand concentration, community structure and buyer-breadth heterogeneity.","authors":["Tengfei Shao","Chao Li","Xu Wang","Masayuki Goto"],"categories":["cs.AI","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-23","first_seen":"2026-09-22","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.24012","pdf_url":"https://arxiv.org/pdf/2609.24012","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","网络校准","市场模拟"],"reason":"用LLM生成persona模拟市场网络，并与真实数据校准，评估仿真充分性，属核…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":7,"question":"如何对生成式社会模拟器进行校准与充分性检验，使其能可靠地复现真实市场的涌现网络结构？","design":"使用基于LLM一次性离线生成的人物画像构建前向模型，模拟二手奢侈品转售市场中买家与品牌的二分网络；通过摊销后验估计校准行为参数，并进行合成可识别性评估、匹配样本量的充分性检查、诊断引导修复和留出统计量审计。","baseline":"真实二手奢侈品转售市场的交易数据，按渠道和居住地划分为四个单元，每个单元为一个买家-品牌二分网络。","findings":"行为参数在所有四个单元中可恢复，但校准近似且对一个参数过度自信；所有单元中观测摘要均超出模拟器的可达性参考，平均购买层级是普遍差异，修复未能恢复充分性，留出审计发现买家广度离散度的遗漏。","reliability":"论文承认校准近似且对一个参数过度自信，修复未能恢复充分性，LLM画像的价值在于结构而非品牌身份，且不声称模拟器再现了市场。","relevance":"该研究直接针对LLM社会模拟的校准与充分性检验，提供了与真实数据对照的严格方法，对关注模拟可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其将模拟器校准与充分性检验结合的方法，通过留出统计量审计避免循环验证，并利用合成可识别性评估参数可恢复性。｜可迁移到消费者市场细分与品牌选择模拟，例如用LLM生成消费者画像模拟电商平台上的购买行为，并与真实交易数据对照。｜以LLM生成消费者画像作为被试，施加不同营销策略（如折扣、推荐）作为处理，结果变量为购买品牌网络结构，用真实电商交易数据作为对照，评估模拟器能否复现品牌集中度与社区结构。"}},{"id":"2609.25586","version":1,"title":"Deflecting the Value Compass: Interacting with Large Language Models Temporarily Shifts Human Value Priorities Toward Personal Focus","zh_title":"偏转价值罗盘：与大语言模型互动暂时将人类价值优先转向个人关注","abstract":"Large language models increasingly support decisions where values are in tension, yet little is known about whether interacting with them changes which values users prioritize. In a preregistered study, 200 U.S. adults interacted with ChatGPT, Claude, or Gemini as a thinking partner or read fixed AI-generated considerations. The prompt asked LLMs to support reasoning without recommending a decision and named no values. Participants advised people facing real dilemmas and completed parallel PVQ-RR forms before, immediately after, and one task later. Each LLM condition temporarily shifted value priorities toward personal focus relative to the control (d=0.37-0.51), primarily through increased Self-Enhancement. Participants' advice retained words and meaning from their exchanges. Thus, a brief LLM interaction that neither targets values nor seeks to persuade can reorient values active during judgment without detectable convergence in value directions or advice.","authors":["Hasibur Rahman","Malak Sadek","Smit Desai"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25586","pdf_url":"https://arxiv.org/pdf/2609.25586","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM影响人类","价值观转变","人机交互实验"],"reason":"研究LLM交互对人类价值观的影响，有真实人类对照，揭示仿真偏差条件","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":8,"question":"与LLM进行简短的思考伙伴式互动是否会暂时改变人们在判断中激活的价值优先级，以及这种改变是否持续、是否导致价值或建议的趋同？","design":"本研究不是用LLM模拟人类被试，而是以200名美国成年人为真实被试，随机分配到ChatGPT、Claude或Gemini三种LLM交互条件，或一个阅读固定AI生成考虑的非交互控制条件。在实验第二阶段，被试与LLM进行最多10分钟的思考伙伴式互动（提示词不提及任何价值观、不推荐决策），然后为真实困境提供建议并完成Schwartz的PVQ-RR价值观量表；在第三阶段，被试在没有LLM的情况下处理另一个困境，以测量效应的持续性。","baseline":"有真实人类对照：控制组被试阅读由三个LLM合成的固定AI生成考虑，但不与LLM交互；所有被试在互动前、互动后立即和后续任务中完成平行版本的PVQ-RR量表，作为价值观变化的基准。","findings":"与LLM互动后，被试的价值优先级相对于控制组暂时向个人关注方向偏移（效应量d=0.37-0.51），主要通过自我增强价值的增加实现；这种偏移在后续任务中消失，且未检测到价值方向或建议的趋同。","reliability":"论文承认效应是暂时的，在后续任务中消失；未讨论其他失效条件或局限。","relevance":"该研究直接探讨LLM交互对人类价值观的因果影响，属于批判性仿真研究，揭示了在无明确说服意图下LLM仍能暂时改变价值优先级，对理解LLM在决策支持中的潜在偏差具有重要意义，值得精读原文。","inspiration":"值得借鉴的做法是采用随机对照实验设计，将LLM作为处理条件，设置非交互控制组，并使用标准化的价值观量表在多个时间点测量效应。｜可以迁移到经济金融中的消费者跨期选择或投资决策场景，例如LLM作为财务顾问是否会影响个人的时间偏好或风险态度。｜一个可行的设计是：招募真实投资者作为被试，随机分配到与LLM（如ChatGPT）进行投资讨论的处理组或阅读固定建议的控制组，在互动前后测量时间贴现率和风险偏好（如使用滴定法或量表），并与真实市场数据（如实际投资组合选择）进行对照，以检验LLM交互对经济决策的因果影响。"}},{"id":"2609.15727","version":2,"title":"Are LLMs Good Financial User Simulators? Multi-view Investor Logic Alignment (MILA)","zh_title":"大语言模型是好的金融用户模拟器吗？多视角投资者逻辑对齐（MILA）","abstract":"Large language models (LLMs) are increasingly used as user simulators, yet it remains unclear whether their predictions faithfully reproduce the evolving decisions of individual users. We investigate this question in a controlled longitudinal paper-trading study with 80 participants, where user interactions, simulated transactions, virtual portfolio states, and point-in-time market information are aligned under a rolling next-day prediction protocol. We evaluate behavioral fidelity hierarchically, from trade occurrence to action structure, asset selection, and downstream portfolio consequences. Across 1,239 aligned user-days, no evaluated LLM reliably outperforms a simple recent-activity persistence baseline for predicting whether a user trades. Fidelity further deteriorates at finer levels: models struggle to recover buy--sell structure and traded assets, and similar activity-level predictions can lead to substantially different portfolio trajectories. Controlled evidence ablations show that recent trading history strongly governs activity prediction, whereas asset selection is substantially more sensitive to the available evidence. An observational analysis further finds that intensified ticker-specific research predicts imminent trading, but diagnostic tests do not support a causal interpretation. These findings suggest that current LLMs capture useful short-term behavioral regularities without yet recovering a stable individual decision mechanism.","authors":["Jiajie He","Jiangyuan Hong","Xintong Chen","Dongling Ni","Wenjin Liu"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2026-09-15","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.15727","pdf_url":"https://arxiv.org/pdf/2609.15727","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","金融行为","算法保真度"],"reason":"用LLM模拟投资者决策，与80名真实用户对照，评估行为保真度并指出失效条件","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":2,"question":"LLM 能否忠实模拟个体投资者在纵向交易环境中的决策行为？","design":"在80名参与者的受控纵向模拟交易研究中，用多种LLM作为用户模拟器，基于滚动次日预测协议，输入截至预测时点的用户交互、交易记录、虚拟持仓和市场信息，预测用户次日是否交易、买卖结构、交易资产及后续组合轨迹，并与真实用户行为逐日对齐比较。","baseline":"80名真实参与者在同一模拟交易平台上的1,239个用户-日对齐行为记录。","findings":"在预测用户是否交易上，所有评估的LLM均未能稳定超越简单的近期活动持续性基线；在更细粒度上，模型难以恢复买卖结构和交易资产，且相似的活动级预测可能导致显著不同的组合轨迹。","reliability":"论文指出当前LLM能捕捉短期行为规律，但尚未恢复稳定的个体决策机制；资产选择对可用证据更敏感，且观察性分析中强化个股研究虽与交易相关，但诊断测试不支持因果解释。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，提供了与真实个体行为逐日对照的严格评估，并揭示了仿真在细粒度决策上的失效条件，对关注仿真保真度和偏差的研究者具有重要参考价值。","inspiration":"值得借鉴的是其分层行为保真度评估框架和受控证据消融方法，可系统区分预测依赖与因果机制｜可迁移到资产定价实验中的投资者异质性决策模拟，或政策公告下的预期形成与交易行为研究｜可设计一个实验：招募真实投资者在模拟平台上交易，用LLM基于其历史行为和市场信息预测次日交易决策，以真实交易记录为对照，通过消融不同信息源检验模型依赖，并比较组合轨迹差异。"}},{"id":"2609.22090","version":1,"title":"Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents","zh_title":"识别、仿真与拒绝：LLM智能体中经典心理效应的污染意识研究","abstract":"An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM agents under a factorial design built to separate these: each paradigm is run with the paradigm explicitly labeled in the prompt (named) or framed as a routine task (blind), and on the literal textbook version of the task (canonical) or a structurally matched variant written to reduce lexical and scenario overlap with likely training data (counterfactual), crossed with a persona manipulation. Across five completed paradigms, evaluated on up to three open-weight model families with 41,904 trials released, apparently human-like effects arise through qualitatively different routes rather than one susceptibility: paradigm-label gating with explicit override (Asch conformity, 0 percent blind to 83.3 percent named on gpt-oss-120B), knowledge-dependent signal reliance (anchoring, exactly zero on grounded facts versus near total on invented quantities, a pattern equally consistent with rational use of the only available signal), amplification on novel content under labeling (framing), robust absence (sunk cost), and safety-mediated selection where refusal itself is the primary finding (minimal-group allocation). A one-sentence persona change (agreeableness, framed as an instruction rather than a verified trait manipulation) eliminates, dampens, or reverses these effects depending on which effect it is, arguing against any single response-bias account. We further formalize, and in two cases document empirically, three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection. We argue scalar bias-susceptibility scores obscure this structure and report replication profiles instead.","authors":["Joy Bose"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22090","pdf_url":"https://arxiv.org/pdf/2609.22090","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM仿真","心理学实验","算法保真度"],"reason":"用LLM复现经典心理学实验，与人类数据对照，并批判性分析仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":5,"question":"LLM在经典心理学实验中表现出的类人效应，究竟是对人类偏差的真实模拟，还是对实验范式的识别、记忆污染或安全过滤所致？","design":"使用多个开源LLM模型（如gpt-oss-120B等）作为被试，在五个经典心理学范式中进行实验。每个范式采用2×2×3因子设计：标签轴（命名vs盲测）、领域轴（经典版本vs反事实版本）、人设轴（无、高宜人性、低宜人性）。结果变量为模型在任务中的行为反应（如从众、锚定、框架效应、沉没成本、最小群体分配）。","baseline":"经典心理学实验的人类行为数据（如Asch从众实验约1/3从众率、锚定效应、框架效应等）作为参照。","findings":"五个范式表现出质的不同路径：Asch从众在盲测下几乎消失（0%），命名后大幅出现（83.3%），表明标签门控；锚定效应在基于事实的锚上为零，在虚构锚上接近完全，符合理性使用唯一信号；框架效应在反事实内容上被标签放大；沉没成本效应完全缺失；最小群体分配因安全拒绝而无法测量。一句宜人性人设指令即可消除、减弱或反转这些效应，表明不存在单一反应偏差。","reliability":"论文承认三个失效模式：人设主导（persona dominance）、群体崩溃（population collapse）、安全选择（safety selection）。人设操纵仅为一句指令，非验证性特质诱导，且与结果变量存在词汇重叠，因此只能视为指令效应而非特质模拟。标签门控可能源于范式名称的语义泄漏，而非模型真正识别实验。","relevance":"该研究直接回应了LLM仿真人类行为时的污染与识别问题，通过严格的因子设计分离了记忆、标签和指令效应，对评估LLM作为人类被试替代品的可靠性具有重要参考价值，值得精读原文。","inspiration":"借鉴其反事实任务设计和标签/盲测对照，以区分模型是基于经济知识还是真实决策偏差。｜可迁移到资产定价实验中的锚定效应、信贷审批中的歧视、消费者跨期选择中的框架效应等场景。｜用LLM模拟投资者，处理为是否告知实验目的（标签vs盲测）和锚定值来源（真实历史数据vs虚构数据），结果变量为估值或投资决策，对照真实人类实验数据（如实验室资产泡沫实验）。"}},{"id":"2609.22169","version":1,"title":"Monocultural Biases: Correlated biases in large language models lead to unequal systemic exclusion rates in hiring","zh_title":"单一文化偏见：大语言模型中的相关偏见导致招聘中的系统性排斥率不平等","abstract":"Employers are increasingly using large language models (LLMs) to automate their hiring process. This paper investigates the risk of monocultural biases, in which the widespread deployment of large language models homogenizes biases across the labor market, leading to greater systemic exclusion for certain demographic groups. For ten LLMs, we measure hiring biases across their base and post-trained versions to identify which stage, pre-training or post-training, lead to monocultural biases. We find that, compared to their base models, post-trained models are 3.6% less likely to callback older applicants. This negative shift occurs in eight of the ten models that we evaluate. Post-trained models have much more correlated decisions than base models which is likely driven by human capital traits like skills or college major. However, greater consensus among models increases global systemic exclusion rates from 5.6% to 17.3% and exacerbates demographic inequalities, with intersectional systemic exclusion rates ranging from 12.2% to 21.7% for post-trained models. We find that this inequality is primarily driven by age-based discrimination that is exacerbated in post-training. These results indicate that while post-training techniques may improve models' abilities to select the best applicants, they may raise systemic inequality risks for those at the margin by uniformly introducing new biases.","authors":["Matthew Bone","Fabian Stephany","Maria del Rio-Chanona"],"categories":["cs.CL","cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22169","pdf_url":"https://arxiv.org/pdf/2609.22169","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","招聘偏见","算法公平"],"reason":"用LLM模拟招聘决策并与人类数据对照，评估系统性偏差，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":6,"question":"大语言模型在招聘中是否会产生单文化偏差，即模型间的决策高度相关，导致某些人口群体在劳动力市场中被系统性排斥？","design":"使用10个LLM（包括基础版和后训练版）模拟招聘初筛，基于Burning Glass Institute的在线劳动力数据生成76种职业的职位空缺和求职者档案，通过改变求职者的年龄、性别、种族等人口特征构造8种变体，测量模型对不同群体的回调率差异，并比较模型间决策的相关性和系统性排斥率。","baseline":"无直接的人类决策对照，但使用真实劳动力市场数据（Burning Glass Institute）生成职位和档案，确保情境真实。","findings":"后训练模型比基础模型更少回调年长求职者（低3.6%），且决策相关性更高，导致全局系统性排斥率从5.6%升至17.3%。这种排斥主要由年龄歧视驱动，后训练阶段（尤其是监督微调）加剧了年龄偏差，同时提高了模型基于人力资本特质筛选的能力。","reliability":"论文指出在线劳动力数据偏向白领、大学学历工作，限制了职业覆盖范围；且未与真实人类招聘决策直接对照，无法完全反映实际部署中的偏差。","relevance":"该研究直接使用LLM模拟招聘决策，并测量系统性偏差，与研究者关注的经济学实验和政策评估场景高度相关，值得精读以了解LLM仿真的偏差来源和测量方法。","inspiration":"借鉴其利用大规模真实数据生成仿真情境并系统操纵人口特征的方法，可迁移到信贷审批歧视研究，用LLM扮演信贷员，处理变量为申请人种族或性别，结果变量为贷款批准率，对照真实信贷数据中的批准率差异。｜可应用于劳动力市场政策评估，如最低工资对雇佣决策的影响，用LLM模拟雇主行为，处理为不同工资水平，测量雇佣概率，对照实际企业调查数据。｜设计一个实验：用LLM模拟投资者对政策公告的反应，处理为不同政策措辞，结果变量为投资决策，对照真实市场数据中的资产价格变动。"}},{"id":"2609.22607","version":1,"title":"Pretrained Persona Mixture Models and Tandem Models for Human Simulation","zh_title":"用于人类仿真的预训练人格混合模型与串联模型","abstract":"We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas, is inaccurate and produces stereotyped predictions (lacking natural diversity). It has previously been shown that LLMs can be bound to personas using naturalistic, freetext dialog avoiding stereotyping. Here we show that binding can also be achieved using short, individual samples of dialog from specific people. Demographics can be added later without negative effects by simply querying the model. We use the term Persona Mixture Models (PMMs) for well-calibrated human models, currently realized as pretrained base models. We show that PMMs produce more accurate predictions than instruction-tuned models and retain more of the lexical, semantic, and pragmatic diversity found in human dialog. We measure realism and diversity of LLMs simulating human interlocutors across a diverse set of corpora spanning open-domain text, human-AI chat, and task-oriented dialogue between human speakers. However, base pretrained models can produce out-of-domain dialog and may lose some of the human's internal state over long contexts. We propose and explore tandem models which combine a pre-trained model with an instruction-tuned supervisor. Tandem models achieve the best overall accuracy and diversity in our experiments.","authors":["Minwoo Kang","T\\'ea Wright","Seun Eisape","Ayush Raj","Suhong Moon","Joseph Suh","Alane Suhr","David M. Chan","John Canny"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22607","pdf_url":"https://arxiv.org/pdf/2609.22607","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","人格混合模型","对话多样性"],"reason":"直接研究LLM仿真人类对话，提出PMM和tandem模型提升准确性与多样性，并…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":9,"question":"如何用预训练语言模型作为人类对话的仿真器，并提升其准确性与多样性？","design":"用预训练基础模型（PMM）和指令微调模型（ALM）分别模拟人类对话者，通过少量个人对话样本绑定人格，比较其在开放域、人机对话和任务型对话中的预测准确性和语言多样性。","baseline":"五个英语语料库中真实人类对话者的语言分布、对话行为结构和词汇语义多样性。","findings":"预训练基础模型比指令微调模型更准确地预测人类下一句话，并保留更多词汇、语义和语用多样性。结合预训练模型与指令微调监督者的串联模型在准确性和多样性上总体最佳。","reliability":"论文指出基础模型可能产生域外对话，在长上下文中丢失人类内部状态；研究仅限英语语料，单一指标不足以全面衡量仿真保真度。","relevance":"该研究直接针对LLM人类仿真，提出PMM和串联模型改进仿真质量，并系统比较了不同模型在真实对话数据上的表现，对关注仿真可靠性与偏差的研究者很有参考价值。","inspiration":"借鉴其用少量个人样本绑定人格并比较不同模型架构的方法，可迁移到经济金融中的消费者决策仿真或投资者情绪模拟。｜例如在消费者跨期选择实验中，用LLM模拟不同人口特征的被试，比较预训练与指令微调模型的行为预测。｜设计：用真实消费者调查数据作为基准，以少量个人回答绑定LLM人格，处理为不同模型类型，结果变量为跨期选择一致性，对照真实人类选择分布。"}},{"id":"2609.24911","version":1,"title":"SocioVerse2: A Longitudinal Dynamic Social Simulation Framework under a Human-AI Co-evolutionary Paradigm","zh_title":"SocioVerse2：人机共演化范式下的纵向动态社会仿真框架","abstract":"Social simulation offers the social sciences an experimental instrument that the real world cannot supply, and generative agents have transformed it by acting as silicon samples that unite agent-based modeling with real behavioral data. Existing platforms verify collective behavior, align simulated populations with real societies in cross-sections, and employ autonomous agents for the research process. However, two social science requirements remain without systematic support: intervention in the content of a simulation and the researcher's control over the process that produces it. We present SocioVerse2, which extends SocioVerse 1.0 into a human-AI co-evolutionary paradigm built from two loops and one infrastructure. The longitudinal simulation loop simulates the target population with evolving environments and forks counterfactual branches via interventions. The controllable research loop takes the study itself as an editable state and updates state versions via controllable editing. The social science agentic infrastructure carries both loops through composable skills with researcher checkpoints, a population service over five persona pools, and an environment service over 21 real-world signal sources with point-in-time guarantees. We validate SocioVerse2 across three case families and seven case studies, from reproducing canonical agent-based models to modeling policy processes on real records and nowcasting macro-economic indices beyond the response model's knowledge cutoff. With the human-AI co-evolutionary paradigm, these cases go beyond system demonstrations to become substantive studies that investigate frontier questions in their respective disciplines. Code, data services, and a workbench are released as open-source resources.","authors":["Xinnong Zhang","Jiayu Lin","Jia Wang","Yixu Huang","Xinyi Mou","Yingqian Wu","Jingcong Liang","Shijun Lei","Jianing Shi","Guanying Li","Siyuan Wang","Hanjia Lyu","Zhenfei Yin","Yunlu Yin","Siming Chen","Yulan He","Jiebo Luo","Xuanjing Huang","Liyin Jin","Baohua Zhou","Hanqi Yan","Zhongyu Wei"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24911","pdf_url":"https://arxiv.org/pdf/2609.24911","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM社会仿真","人类数据对照","政策评估"],"reason":"用LLM agent模拟社会过程并与真实数据对照，支持干预和纵向演化，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":10,"question":"如何构建一个支持纵向干预、研究者可控、并基于真实世界数据的人类-AI协同演化社会仿真框架？","design":"SocioVerse2 用生成式智能体（硅样本）模拟目标人群，通过纵向仿真循环让环境和人群随时间演化，并支持干预操作产生反事实分支；可控研究循环将研究过程本身作为可编辑状态，允许研究者介入和调整；基础设施提供五个角色池和21个真实世界信号源，保证时间点一致性。","baseline":"多个案例研究使用真实记录和宏观指数作为对照，如政策过程建模使用真实记录，宏观指数预测使用真实世界确定性指数和信号作为基准。","findings":"SocioVerse2 能够复现经典基于智能体的模型，并在真实记录上建模政策过程，还能预测超出模型知识截止日期的宏观经济指数。通过人类-AI协同演化范式，案例研究超越了系统演示，成为各自学科前沿问题的实质性研究。","reliability":"论文未讨论","relevance":"该研究直接针对用LLM进行社会仿真并与真实数据对照，支持干预和纵向演化，对评估仿真可靠性和偏差有参考价值，值得阅读原文了解其框架和案例细节。","inspiration":"借鉴其纵向干预和反事实分支设计，可在经济政策评估中设置处理组和对照组，观察政策效果的动态演化。｜可迁移到政策公告的预期形成研究，模拟不同政策沟通策略对市场参与者预期的影响。｜用LLM智能体模拟投资者群体，处理为不同政策公告措辞，结果变量为预期通胀或资产价格变动，对照真实市场调查数据或高频交易数据。"}},{"id":"2609.22252","version":1,"title":"CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation","zh_title":"CALM：用于基于活动的出行者仿真的校准LLM选择网络框架","abstract":"We present CALM, a reproducible hybrid framework that integrates an optional large language model (LLM) activity planner with calibrated stochastic choice, shared network feedback, memory and habit, typed feasibility checks, and deterministic offline replay. Unlike trip-mode classifiers or diary-only generators, CALM executes a closed traveler-day loop and evaluates each generative module against an empirical, reproducible baseline. On the 2024 New York City Citywide Mobility Survey (CMS), 110,691 seven-mode trips are split by respondent into 78,487 training and 32,204 holdout trips. Training-only alternative-specific constant calibration reduces mean holdout mode Jensen-Shannon divergence from 0.15599 to 0.00394 across ten seeds. A matched live-LLM ablation then quantifies trade-offs among aggregate fit, temporal fit, behavioral persistence, and feasibility, while frozen prompt-response pairs support deterministic replay of downstream simulation. Controlled weather, delay, fare, and parking ladders further demonstrate consistent and interpretable responses under intervention. CALM contributes a reproducible protocol for integrating and evaluating generative planners in traveler simulation through person-disjoint calibration, matched module ablation, controlled stress testing, and end-to-end traceability.","authors":["Yezhou Cheng"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22252","pdf_url":"https://arxiv.org/pdf/2609.22252","source_feed":"cs.LG","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","出行行为","人类数据校准"],"reason":"用LLM模拟出行者选择，并与真实调查数据对照校准，属于人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":7,"question":"如何构建一个可校准、可复现的混合框架，将大语言模型活动规划器与随机选择模型结合，用于基于活动的出行者仿真，并与真实调查数据对照评估。","design":"CALM框架用LLM（gpt-4o-mini）作为可选活动规划器，结合校准的随机选择模型、共享网络反馈、记忆与习惯、类型化可行性检查，模拟纽约市出行者的一日活动。通过匹配的消融实验（校准选择、无记忆LLM、有记忆LLM、完整混合）比较不同模块组合，施加天气、延误、票价、停车等干预阶梯，测量模式份额的Jensen-Shannon散度、时间分布、行为持续性、可行性等结果。","baseline":"2024年纽约市全市出行调查（NYC CMS）的110,691次七模式出行记录，按受访者划分为78,487次训练和32,204次留出，用于校准和评估。","findings":"仅用训练数据校准替代特定常数（ASC）后，留出集模式份额的Jensen-Shannon散度从0.15599降至0.00394，且跨10个种子稳定。匹配的LLM消融实验量化了聚合拟合、时间拟合、行为持续性和可行性之间的权衡，干预阶梯显示一致且可解释的响应。","reliability":"论文强调面效度不等于经验效度，LLM生成的出行叙事可能仍与空间、时间和OD分布不匹配；当前13区网络实例化仅提供受控环境，未在大规模真实网络上验证。","relevance":"该研究直接针对LLM仿真人类行为并与真实调查数据对照，提供了严格的校准和评估协议，对关注经济学实验和政策评估中LLM仿真可靠性的研究者具有重要参考价值。","inspiration":"值得借鉴的是其人员不相交的校准、匹配模块消融和受控压力测试，确保仿真差异归因于特定模块而非需求背景。｜可迁移到政策评估中的个体选择行为仿真，如交通定价、补贴或信息干预对出行方式选择的影响。｜以真实居民出行调查数据为基准，用LLM生成个体活动计划，施加票价或拥堵收费等处理，测量方式选择概率和福利变化，并与实际政策试点数据对照。"}},{"id":"2609.22408","version":1,"title":"Social Influence and the Allocation of Scientific Attention in AI Populations","zh_title":"AI群体中的社会影响与科学注意力分配","abstract":"AI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human readers, but the collective consequences of these signals for artificial readers remain uncertain. This paper adapts the Music Lab design to a market for academic attention. In the first experiment, 1,000 AI agents choose papers from the titles and abstracts of all 114 regular research articles published in the American Economic Review in 2025. The experiment has five independent-choice communities and five social-influence communities, each with 100 sequential agents. Only agents in the social-influence condition observe earlier selections within their community. Agents may select any number of papers. Social-information communities select 17.2 percent fewer papers per agent, concentrate their choices more heavily, and collectively cover 73 papers, compared with 90 independently. Between-community variation is greater under social information. In a second experiment with 200 agents across twenty social communities, randomly assigning papers five initial selections raises their subsequent selection rate by 45.55 percentage points (95% CI: 41.20 to 49.90). Choices have modest correspondence with external citations and little correspondence with download counts. The results show how a simple information rule shapes the volume, breadth and distribution of scientific attention in an artificial population.","authors":["Maxim Chupilkin"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22408","pdf_url":"https://arxiv.org/pdf/2609.22408","source_feed":"econ.GN","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","社会影响","学术注意力"],"reason":"用1000个AI代理模拟学术注意力分配，与真实引用数据对照，属经济学实验场景。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":8,"question":"社会信息如何影响AI代理对学术论文的注意力分配？","design":"用GPT-5.6 Sol扮演1000个AI代理，分为独立选择组和社会影响组（各5个社区，每社区100个顺序代理），从2025年AER的114篇文章标题和摘要中选择要读的论文；社会影响组能看到本社区之前代理的累计选择次数，独立组看不到；测量选择数量、集中度、社区间差异等。","baseline":"外部引用次数和下载量作为对照，但仅用于相关性分析，并非严格的人类被试基准。","findings":"社会信息使代理平均少选17.2%的论文，集体覆盖从90篇降至73篇，选择更集中且社区间差异更大；随机赋予初始流行度使后续选择率提高45.55个百分点。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类在学术注意力市场中的从众行为，与真实引用数据对照，属于经济学实验场景，对关注LLM仿真可靠性和偏差的研究者有直接参考价值。","inspiration":"借鉴其Music Lab式设计，通过独立组与社会影响组对比、随机初始流行度干预来分离社会影响效应｜可迁移到金融信息传播或资产定价中的注意力分配问题，如投资者对研报或新闻的关注｜用LLM代理扮演投资者，处理为是否展示其他代理的阅读/选择行为，结果变量为选读的研报数量与集中度，对照真实市场中的研报点击或交易数据。"}},{"id":"2609.22225","version":1,"title":"Do LLMs Choose Like Humans? Using Cognitive Theory to Evaluate LLM Decision-Making","zh_title":"LLM像人类一样选择吗？用认知理论评估LLM决策","abstract":"Large language models (LLMs) exhibit a range of human-like decision-making behaviors, but whether these reflect similar underlying mechanisms or surface-level mimicry remains unclear. We evaluate whether LLM context sensitivity aligns with a cognitive economic theory that explains human behavior through problem categorization and attention allocation. Across 12 open-source and commercial LLMs on a novel 140,000-trial product choice benchmark, context induces human-like shifts in choice and problem categorization, but does not reliably reweight attention between features like price and quality. Neither scale nor chain-of-thought reasoning reliably attenuates context sensitivity or generates human-like behavior. These results suggest that LLM decision mechanisms are distinct from human ones.","authors":["Johnathan Sun","Andrei Shleifer","Yonatan Belinkov"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22225","pdf_url":"https://arxiv.org/pdf/2609.22225","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM决策仿真","认知理论","人类对照"],"reason":"用LLM模拟人类决策并与人类数据对照，评估机制差异，直接相关且具批判性。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":12,"question":"LLM 的决策行为是否与人类认知经济理论中的情境敏感机制一致，即是否通过问题分类和注意力重加权来产生情境效应。","design":"用 12 个开源和商业 LLM 作为被试，在 140,000 次产品选择试验中，通过情境线索（享乐型 vs. 功能型消费）操纵问题分类，测量选择概率和特征敏感性（价格与质量的注意力权重）。","baseline":"人类基准来自认知经济学理论（Bordalo et al., 2026b）所解释的人类决策现象，包括情境对选择的影响、特征敏感性的不对称变化以及分类难度对情境效应的调节。","findings":"LLM 的选择变化与人类一致，但特征敏感性变化大多不一致；情境效应遵循问题分类过程，但模型规模和思维链推理不能可靠减弱情境敏感性。","reliability":"论文指出 LLM 的决策机制与人类不同，情境效应可能只是表面模仿；规模和推理不能可靠产生人类行为，说明在需要机制对齐的任务中仿真可能失效。","relevance":"该研究直接评估 LLM 作为人类决策仿真体的机制一致性，提供了与人类认知理论对照的批判性证据，值得精读以了解仿真在经济学决策中的局限。","inspiration":"借鉴其通过情境线索操纵问题分类并测量特征敏感性的设计，可迁移到消费者跨期选择或资产定价实验，用 LLM 模拟投资者在不同市场情境下的风险偏好，处理为情境线索（如牛市/熊市），结果变量为风险资产配置比例，对照真实投资者调查数据。"}},{"id":"2609.23403","version":1,"title":"Alignment and Divergence between Humans and AI in Interpersonal Privacy Decisions","zh_title":"人际隐私决策中人类与AI的一致性与分歧","abstract":"AI assistants increasingly mediate interpersonal communication on behalf of their primary user, but they risk violating the privacy expectations of third-party information owners. Resolving these tensions requires understanding how humans anticipate interpersonal privacy boundaries. Therefore, we conducted a dyadic study (N=76) and a matched evaluation of AI models across 18 information types and 3 recipient relationships. We found that data owners' privacy judgments are highly contextual and relationship dependent. While familiar data co-owners show meaningful alignment with owners' expectations, they significantly overestimate the need for permission. Interestingly, greater familiarity within the owner-co-owner dyad was associated with both higher disclosure acceptability and lower co-owner misalignment, whereas our exploratory four-item empathy measure was not. In contrast, AI models significantly underperform human co-owners in anticipating the data acceptability, even when provided with within-dyad examples. These findings underscore a core HCI design challenge to develop privacy-aware AI that respects multi-stakeholder information boundaries.","authors":["Hanxiang Zeng","Shuning Zhang","Xinyuan Zhou","Tianqi Song","Yuhan Yuan","Yuting Yang","Shuai Ma","Xin Yi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23403","pdf_url":"https://arxiv.org/pdf/2609.23403","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","隐私决策","人机对齐"],"reason":"用LLM模拟人类隐私决策并与人类数据对照，评估AI与人类判断的偏差，属于仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":13,"question":"在人际隐私决策中，数据所有者对第三方披露的接受度如何受情境影响，人类共同所有者与所有者的判断对齐程度如何，AI模型能否准确预测所有者的隐私判断？","design":"本研究并非严格意义上的LLM仿真人类实验，而是采用配对设计：招募19对朋友和19对恋人（N=76），让数据所有者和共同所有者分别对18种信息类型、3种接收者关系下的披露可接受性、许可必要性等隐私维度进行评分；同时用多个LLM（零样本和单样本提示）、机器学习模型和微调模型对相同场景进行预测，并与人类共同所有者的预测准确度比较。","baseline":"人类基准是熟悉的数据共同所有者（朋友或恋人）对所有者隐私判断的预测，以平均绝对误差（MAE）和判断落在所有者评分±1分内的比例衡量对齐程度。","findings":"所有者的隐私判断高度依赖情境和关系，信息类型和接收者关系显著影响披露可接受性；熟悉的人类共同所有者与所有者有中等程度对齐（MAE≈1.179，70.7%判断在±1分内），但系统性高估许可必要性。AI模型（包括零样本和单样本LLM）显著不如人类共同所有者准确（MAE 1.648–1.796），单样本提示仅略微改善，仍无法弥合人机差距。","reliability":"论文承认AI模型在涉及微妙社会关系的披露场景中尤其表现不佳，单样本提示不足以缩小差距；机器学习模型和微调模型对未见个体泛化能力差，提供所有者特定示例的改进不一致。局限性包括样本量较小（19对朋友和19对恋人）、仅覆盖两种关系类型、同理心测量简短且探索性，未发现其与对齐的显著关联。","relevance":"该研究直接评估LLM在人际隐私决策中模拟人类判断的准确性，并与真实人类共同所有者的预测进行对照，揭示了AI在微妙社会情境下的系统性偏差，对关注LLM仿真可靠性及失效条件的研究者具有参考价值。","inspiration":"借鉴其配对设计和多模型基准测试方法，可系统评估LLM在预测他人偏好或决策时的准确性，并与人类代理预测对照。｜可迁移到经济金融中涉及代理决策或偏好预测的场景，如理财顾问预测客户风险偏好、信贷员判断借款人还款意愿、或政策制定者预测公众对经济政策的接受度。｜设计：招募真实客户-理财顾问对，让客户对一系列投资产品的风险承受能力和偏好进行评分，同时让顾问预测客户评分，并让LLM基于客户基本信息和少量示例进行预测；结果变量为预测误差（MAE）和方向一致性；以顾问预测为人类基准，比较LLM与顾问的准确性，并考察客户-顾问关系强度、信息敏感度等调节因素。"}},{"id":"2609.24859","version":1,"title":"Small-world Networks of Agents Brainstorm AI Risks to Support Ideation","zh_title":"智能体小世界网络头脑风暴AI风险以支持构思","abstract":"The ideation phase of participatory AI risk assessment often starts with a blank slate or a limited list of predefined risks, making it difficult to surface indirect or systemic harms. To address this limitation, we propose a three-stage ideation support tool. The tool complements participatory AI, rather than replacing it, and helps focus later engagement with affected communities. First, it dynamically discovers stakeholders depending on the given AI use and recursively expanding outward, allowing overlooked or indirect stakeholders to emerge. Second, it simulates these stakeholders with LLMs, connecting them into a network of a given topology, and having them ideate about risks. Third, it prioritizes risks using network centrality measures. In an initial evaluation, we found that betweenness centrality run through agents connected in a small-world network works best as it elevates risks raised by stakeholders who bridge disconnected groups, surfacing novel, systemic harms that traditional methods often miss. On an AI chatbot companion use case, this approach increased the novelty of the identified risks by approximately 1.1 points over single LLM brainstorming, and by 0.5 points over agentic LLM brainstorming, measured on a normalized five-point Likert scale, without reducing the plausibility or severity of the identified risks. To test whether our framework helps a human-led ideation session using the Futures Wheel approach, we divided 11 teams of non-western young chatbot users into two types: control (team) and treatment (team) in a participatory AI risk assessment. The control teams started from a list of risks generated by the 45 AI practitioners in the initial evaluation; the treatment teams started from a list generated by our framework. The treatment teams identified more risks overall, and more systemic, human-computer interaction, and environmental risks.","authors":["Ke Zhou","Edyta Bogucka","Daniele Quercia"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24859","pdf_url":"https://arxiv.org/pdf/2609.24859","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","风险识别","参与式AI"],"reason":"用LLM模拟利益相关者头脑风暴AI风险，并与人类团队对照，涉及政策评估场景，但…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":16,"question":"如何利用小世界网络中的LLM智能体头脑风暴来支持AI风险识别的构思阶段，以发现更多新颖且系统性的风险？","design":"该研究提出一个三阶段框架：首先动态发现利益相关者，然后将其实例化为LLM智能体并连接成小世界网络进行风险头脑风暴，最后用网络中心性（特别是介数中心性）对风险进行排序。评估中比较了单LLM头脑风暴、多智能体基线以及该框架生成的风险在创新性、合理性和严重性上的表现，并进一步在人类主导的Futures Wheel工作坊中测试了该框架生成的风险列表对团队构思的促进作用。","baseline":"对照的真实人类数据包括：45名AI从业者在空白头脑风暴工作坊中生成的244条风险（用于评估风险生成质量），以及11个非西方年轻聊天机器人用户团队在Futures Wheel工作坊中的表现（控制组使用从业者生成的风险列表，处理组使用框架生成的风险列表）。","findings":"在AI聊天机器人伴侣用例中，小世界网络结合介数中心性的方法将风险创新性评分比单LLM头脑风暴提高了约1.1分，比智能体LLM头脑风暴提高了0.5分（5分制），且未降低合理性和严重性。在人类参与的Futures Wheel研究中，使用该框架生成的风险列表的处理组团队识别出了更多总体风险，以及更多系统性、人机交互和环境风险。","reliability":"论文未讨论","relevance":"该研究用LLM模拟利益相关者网络进行风险头脑风暴，并与真实人类团队进行对照，涉及AI政策评估场景，对关注LLM仿真可靠性和偏差的研究者具有参考价值。","inspiration":"该方法通过构建小世界网络并利用介数中心性来提升生成内容的多样性和新颖性，可借鉴其网络拓扑与中心性度量来设计多智能体仿真中的信息传播和观点聚合机制。｜可迁移到经济金融领域的政策评估或风险识别场景，例如金融系统性风险的早期预警或信贷审批中的歧视性风险识别。｜可设计一个研究：用LLM模拟银行、监管者、消费者等利益相关者，在小世界网络中讨论信贷审批算法的潜在风险，以生成的风险列表作为处理组，与人类专家生成的风险列表进行对照，比较两组在风险覆盖度、新颖性和系统性上的差异，并使用真实历史信贷数据或监管报告作为外部基准。"}},{"id":"2609.24629","version":1,"title":"Augmented Hypothesis Testing with Persona-Based LLM Simulations","zh_title":"基于角色LLM模拟的增强假设检验","abstract":"A/B testing requires large sample sizes, long timelines, and significant costs. When auxiliary predictions of experimental outcomes are available from machine learning models, uncertain prediction quality precludes replacing human experiments entirely, yet these predictions may still contain useful signal. We propose a principled framework for learning-augmented hypothesis testing that leverages predictions of unknown quality to reduce sample sizes while maintaining statistical validity. Predictions naturally vary in granularity, from coarse aggregate signals to fine-grained individual-level estimates, and our framework addresses both ends of this spectrum: (1) for population-level directional predictions, where only a binary signal on the treatment effect sign is available, we use an asymmetric test and prove consistency and robustness bounds within the learning-augmented algorithms paradigm; (2) for individual-level predictions, we introduce Generalized PPI++ (GPPI), extending Prediction-Powered Inference to handle nonlinear prediction errors through higher-dimensional transformations. Both methods benefit from accurate predictions while remaining robust to inaccurate or adversarial ones. We validate our framework using persona-based LLM simulations, where AI agents equipped with user personas predict individual behavior, as a natural prediction source spanning both granularity levels. Experiments on four real-world datasets demonstrate that our methods, combined with persona-based predictions, substantially reduce experimental costs while preserving rigorous statistical validity.","authors":["Ziyad Benomar","Aymen Al Marjani","Paul Missault","Saab Mansour"],"categories":["cs.LG","cs.AI","stat.AP"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24629","pdf_url":"https://arxiv.org/pdf/2609.24629","source_feed":"cs.LG","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM仿真","假设检验","统计有效性"],"reason":"用LLM persona预测个体行为，与真实数据对照，并保证统计有效性，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":15,"question":"如何利用 persona 驱动的 LLM 预测来减少 A/B 测试所需样本量，同时保持统计有效性？","design":"使用基于用户 persona 的 LLM 模拟来预测个体行为，作为辅助预测源；针对群体级方向预测（仅提供处理效应符号）和个体级预测分别提出不对称检验和广义 PPI++ 方法，在四个真实数据集上验证。","baseline":"四个真实世界数据集（具体名称未在节选中列出）中的人类 A/B 测试结果作为对照。","findings":"所提出的学习增强假设检验框架能够利用 persona 预测显著减少实验成本，同时保持严格的统计有效性。群体级方向预测和个体级预测方法均能在预测准确时受益，并在预测不准确或对抗性时保持稳健。","reliability":"论文承认完全用 persona 预测替代人类实验面临根本挑战：LLM 的黑箱性质、对提示工程的敏感性、分布偏移以及人类行为的复杂性使得形式化质量保证困难；此外，persona 与真实用户之间可靠的一对一映射通常不可行，导致预测粒度受限。","relevance":"该研究直接针对研究者关注的 LLM 人类仿真实验，利用 persona 预测个体行为并与真实数据对照，同时提供统计有效性保证，值得精读原文以了解具体方法和实验细节。","inspiration":"借鉴其将 LLM 预测作为辅助信号而非替代品，并通过统计校正保持有效性的思路｜可迁移到政策评估中的 A/B 测试，如利用 LLM 模拟消费者对价格变动的反应来减少实地实验样本量｜以消费者信贷决策为场景，用 LLM 基于用户 persona 预测个体对贷款条款的反应，处理为不同利率或审批条件，结果变量为接受/拒绝或违约行为，对照真实信贷申请数据。"}},{"id":"2608.18265","version":4,"title":"Modeling Human Behavior with Type Vectors Using AI","zh_title":"使用AI类型向量建模人类行为","abstract":"We introduce a general, easy-to-implement AI-based modeling technique for analyzing human behavior. A key feature of this approach, which contrasts with existing modeling techniques, is that it combines the flexibility and interpretability of natural language with a mathematical structure that can be fitted to data and easily analyzed. We assign a large language model a vector of trait intensities-a type vector-and then ask it to choose actions across settings in which we observe human choices. For instance, the type vector (2,4) could correspond to \"You are a player characterized by the following profile: Altruism: 2 out of 5, Risk Aversion: 4 out of 5,\" after which it is asked to make choices. We can then vary the traits (e.g., Altruism, Fairness, Trust,...) and values (e.g., 1-5) to minimize distance to human choices. We illustrate the method by applying it to model 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles. We find that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust. The type vectors needed to fit individuals across games cluster into fewer than a dozen groups, with substantial variation in fit across subjects. Moreover, the individual type vectors can predict behavior in held-out games with different rules and available actions. More broadly, this new modeling method is highly generalizable and interpretable: we can input any vector of traits and use them to model behavior across any setting","authors":["Matthew O. Jackson","Benjamin S. Manning","Yutong Xie","Walter Yuan","Qiaozhu Mei"],"categories":["econ.TH","cs.AI"],"primary_category":"econ.TH","announce_type":"replace-cross","date":"2026-09-21","first_seen":"2026-08-20","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.18265","pdf_url":"https://arxiv.org/pdf/2608.18265","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","经济实验","人类行为建模"],"reason":"用LLM模拟人类经济决策，并与大规模真实人类数据对照，直接命中核心方向。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":1,"question":"人类行为是否可以用少数几个特质维度（类型向量）来近似和预测？","design":"用大语言模型（LLM）作为仿真被试，通过提示词赋予其不同特质强度（如利他、风险厌恶、信任等）构成类型向量，然后让模型在10个经典经济博弈角色中做决策，通过调整类型向量最小化与人类选择的距离。","baseline":"119,147条决策，来自78,657名被试，覆盖35个国家，在10个经典经济博弈角色中的真实选择。","findings":"人类行为可以用三个维度（风险厌恶、策略复杂性、信任）紧密匹配；个体类型向量在博弈间聚类成少于12个群体，且能预测未参与拟合的博弈中的行为。","reliability":"论文未讨论","relevance":"直接命中核心方向：用LLM模拟人类经济决策并与大规模真实人类数据对照，方法可解释且可推广，值得精读原文。","inspiration":"借鉴其用可调类型向量作为LLM提示来系统扫描特质空间并拟合真实行为的方法，可迁移到资产定价实验中的风险偏好与信念异质性建模。｜设计一个研究：用LLM模拟投资者，赋予不同风险厌恶和过度自信水平，在实验性资产市场中交易，结果变量为价格泡沫程度和交易量，与真实实验市场数据（如Smith et al. 1988）对照，检验类型向量能否复现泡沫。"}},{"id":"2609.21636","version":1,"title":"Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30","zh_title":"引导大语言模型在挪威MFQ-30上向道德基础靠拢","abstract":"Recent work applies human psychometric questionnaires to large language models to elicit moral and value profiles, but it is not clear whether these instruments measure anything stable in models or whether the resulting profiles can be moved toward a target human population. We administer the Norwegian Moral Foundations Questionnaire (MFQ-30) to six open-weight LLMs and compare their foundation profiles to a sample of N = 1,282 Norwegian respondents. We test two steering interventions, prompt-level persona steering and activation-level ActAdd. Half the models engage with the questionnaire under our attention check. The other half default to flat or central-tendency outputs that look near-human on average without tracking item content. A neutral Nordic-respondent persona, written without any distributional information from the human sample, brings the engaging models 44-77% closer to the Norwegian mean in Mahalanobis $d^2$. One-pair ActAdd at a fixed mid-layer flattens the foundation profile rather than steering individual foundations. For at least one model the same persona that shifts the profile also induces engagement that was absent at baseline, a concrete instance of the cognitive phantoms that Peereboom et al. (2025) warn about.","authors":["Hans Andersen","David Dichas"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21636","pdf_url":"https://arxiv.org/pdf/2609.21636","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","道德基础","算法保真度"],"reason":"用LLM复现人类道德基础分布，并与挪威样本对照，评估仿真可靠性与偏差，含批判性…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":4,"question":"LLM 在挪威版道德基础问卷（MFQ-30）上的道德画像是否稳定，能否通过提示或激活干预向目标人群靠拢？","design":"用六个开源 LLM 扮演挪威受访者，回答挪威 MFQ-30 问卷；施加两种处理：提示层 persona 引导（中性北欧受访者人设）和激活层 ActAdd 引导（对比向量注入）；结果变量为五个道德基础得分及与人类样本的马氏距离。","baseline":"Enstad 和 Finseraas (2024) 收集的 1282 名挪威受访者 MFQ-30 数据，已过滤不专注样本。","findings":"半数模型在注意力检查下真正作答，另一半输出扁平或趋中，平均接近人类但不追踪题目内容。中性北欧人设使作答模型与挪威均值的马氏距离缩小 44–77%，而 ActAdd 在固定中间层会扁平化画像而非定向引导单个基础。","reliability":"论文指出部分模型在基线时不作答，注意力检查可识别；人设引导可能诱发“认知幻影”，即模型表现出原本没有的作答行为，提示仿真结果可能不可靠。","relevance":"直接命中研究者对 LLM 仿真可靠性与偏差的关注，提供了真实人类对照和批判性发现，值得精读原文以了解注意力检查与引导干预的具体实现。","inspiration":"借鉴其注意力检查与第一 token 概率解码方法，可提高 LLM 问卷作答质量并识别无效仿真｜可迁移到经济政策偏好调查或消费者态度测量，如用 LLM 模拟不同人群对税收、福利或环保政策的道德评价｜用开源 LLM 扮演不同社会经济群体，施加中性人设或激活引导，测量其对政策陈述的同意程度，并与真实调查数据（如欧洲社会调查）对照，检验仿真偏差。"}},{"id":"2609.21259","version":1,"title":"CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition","zh_title":"CogGym：迈向人类与机器认知的大规模比较评估","abstract":"Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models' improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model--human fit remains well below human splithalf reliability ($R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video) with the best models achieving $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.","authors":["Lance Ying","Jinzhou Wu","Yingshan Susan Wang","Shivam Aarya","Luca M. Schulze Buschoff","Harry Chen","Katherine M. Collins","Andrea de Varda","Shuhao Fu","Sean Dae Houlihan","Akshay K. Jagadish","Guangyuan Jiang","Samuel Kiegeland","Tetsu Kurumisawa","Rongzhi Liu","Ryan Liu","Ningshan Ma","Kathryn McGregor","Younes Strittmatter","Polina Tsvilodub","Jacob Hoover Vigly","Sarah Wu","Enjie Xu","Yiling Yun","Kelsey Allen","Tyler Brooke-Wilson","Brian Christian","Evelina Fedorenko","Michael C. Frank","Michael Franke","Tao Gao","Samuel J. Gershman","Robert D. Hawkins","Jennifer Hu","Julian Jara-Ettinger","Max Kleiman-Weiner","Sydney Levine","Tal Linzen","Hongjing Lu","Timothy O'Donnell","Desmond C. Ong","Steven T. Piantadosi","Rebecca Saxe","Eric Schulz","Tianmin Shu","Felix A. Sosa","Ilia Sucholutsky","Tan Zhi-Xuan","Tomer Ullman","Fei Xu","Ilker Yildirim","Jian-Qiao Zhu","Thomas L. Griffiths","Tobias Gerstenberg","Kevin Smith","Joshua B. Tenenbaum"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21259","pdf_url":"https://arxiv.org/pdf/2609.21259","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知实验","人类对照"],"reason":"直接比较LLM与人类在认知实验中的行为，有真实人类数据对照，并评估模型-人类拟…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":3,"question":"如何在大规模、多样化的认知实验中系统比较AI模型与人类的行为，以识别模型与人类判断的相似与分歧？","design":"构建CogGym框架，将258个来自100篇论文的人类常识推理认知实验标准化为统一的实验标记语言（EML），并评估50个大型语言模型在这些实验上的表现，测量模型输出与人类反应分布的拟合度。","baseline":"原始认知实验中的真实人类被试反应数据，包括文本、图像和视频模态，并报告人类分半信度作为上限。","findings":"模型规模越大、越新，与人类判断的拟合度越高，但在常识推理任务上的进步速度慢于数学和编程等正式推理基准。最佳模型与人类的拟合度（文本R²=0.59，图像R²=0.58，视频R²=0.43）仍远低于人类分半信度（文本R²=0.93，图像R²=0.95，视频R²=0.92）。","reliability":"论文指出模型-人类拟合度远低于人类分半信度，表明模型尚未完全捕捉人类认知的细微结构；同时，标准化过程中可能存在实验转换的失真，且当前覆盖范围限于常识推理，未涉及其他认知领域。","relevance":"该研究直接比较LLM与人类在大量认知实验中的行为，有真实人类数据对照，并评估模型-人类拟合度，与研究者关注的人类仿真实验高度相关，值得精读以了解大规模评估框架和模型偏差。","inspiration":"借鉴其半自动化实验标准化流程和分半信度作为上限的评估方法，可迁移到经济决策实验（如风险偏好、时间贴现、博弈行为）的仿真验证。｜可应用于消费者跨期选择或资产定价实验，检验LLM是否复现人类的时间不一致性或风险厌恶。｜设计：以LLM为被试，施加跨期选择任务（如现在100元 vs. 一个月后120元），测量贴现率，并与真实人类实验数据（如Andersen et al., 2008）对照，计算模型-人类拟合度与人类分半信度的差距。"}},{"id":"2609.20827","version":1,"title":"From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators","zh_title":"从出院记录到患者理解：基于人格的开放式模拟将LLM作为出院教育者","abstract":"Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient's literacy, recall, and personality. Existing LLM evaluations target static or artifact-generation tasks and do not measure patient understanding under open-ended dialogue. We introduce DischargeBench, a persona-grounded simulation in which a candidate LLM educator conducts a multi-turn session with a Virtual Patient, while an Education Monitor Agent regulates patient realism without modifying the educator, protecting the evaluation signal. We curate MIMIC-IV-Ext-DischargeBench, 477 cases over 24 ICD chapters with persona axes (personality, education level, health literacy, past-medical-history recall) for stratified analysis. Each simulation is scored on four axes -- Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency -- by an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source-answer agreement. LLM evaluation for discharge education should center patient understanding, not text quality or answer accuracy alone.","authors":["Won Seok Jang","Zonghai Yao","Hong Yu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20827","pdf_url":"https://arxiv.org/pdf/2609.20827","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","患者教育","人类数据对照"],"reason":"用LLM模拟患者进行出院教育评估，有真实临床数据对照，涉及医疗场景，并指出仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":4,"question":"如何评估大语言模型在开放式、多轮对话中作为出院教育者，使不同人格与健康素养的患者真正理解出院指导的能力？","design":"构建 DischargeBench 仿真框架：用 LLM 扮演虚拟患者（基于 MIMIC-IV 真实出院记录，设定人格、教育水平、健康素养、病史回忆等属性），候选 LLM 扮演教育者进行多轮对话，并由教育监控代理仅调节患者侧以保持真实性；通过 LLM 裁判对对话质量、主题覆盖、理解程度和事实一致性四个维度评分。","baseline":"使用 MIMIC-IV-Ext-DischargeBench 的 477 个真实病例（来自 MIMIC-IV 和 MIMIC-IV-Note），并由医生标注用于对齐 LLM 裁判的评分。","findings":"GPT-5 系列在对话质量和主题覆盖上领先，但可读性较差；总体分数掩盖了不同 ICD 章节和患者人格下的显著差异，困难人格暴露了覆盖失败、理解差距和源答案一致性下降。","reliability":"论文指出 LLM 评估应关注患者理解而非文本质量或答案准确性；困难人格下仿真会失效，且 LLM 裁判的评分可能与医生标注存在偏差。","relevance":"该研究用 LLM 模拟患者进行出院教育评估，有真实临床数据对照，并批判性指出仿真在困难人格下失效，与研究者关注的人类仿真可靠性与偏差高度相关，值得精读。","inspiration":"借鉴其多智能体仿真设计：用 LLM 扮演异质性个体并设置监控代理防止污染处理信号，同时用真实数据校准裁判评分。｜可迁移到消费者金融教育或政策沟通场景，如评估 AI 理财顾问对不同金融素养人群的讲解效果。｜设计：用 LLM 模拟不同金融素养和人格的投资者作为被试，处理为 AI 顾问的个性化解释，结果变量为投资决策质量和风险理解，对照真实投资者调查数据（如 FINRA 金融素养调查）。"}},{"id":"2609.21439","version":1,"title":"People escalate against a competitor labelled human and hold back against one labelled an optimising machine","zh_title":"人们面对标记为人类的竞争者会升级投入，面对标记为优化机器的竞争者则会退缩","abstract":"People increasingly compete against AI agents rather than other human opponents. We distinguish two channels: an opponent effect and an information effect. These are different elements with different consequences: the opponent effect is specific to a given computational system, the information effect a property of the information environment that an organisation or policymaker can control. We separate them in a preregistered experiment (N = 1,395) using a dynamic all-pay auction, a repeated contest in which escalation of commitment arises from the incentives. What participants are told about the opponent (human, an AI trained to imitate people, or an AI trained to compete well) is varied and crossed with who they actually face, in a deception-free design. What people are told influences escalation: the median price rises by 6.7 points when a human might be the opponent and falls by 8.8 when an optimising machine might be, a spread of about 15% of the prize value of the competition, produced by information alone. Competing against the AI agents lowers prices, yet reduces the chance that both sides finish with positive earnings, showing distinct effects of the opponent channel. The information effect is not explained by articulated strategy, or individual differences, and is consistent with a competitive response engaged when a human is a live possibility. This shows that describing an AI competitor is not behaviourally neutral.","authors":["Vinicius Ferraz","Leon Houf"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21439","pdf_url":"https://arxiv.org/pdf/2609.21439","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","行为实验","人机交互"],"reason":"用LLM作为对手与人类被试互动，研究标签信息对行为的影响，有真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":11,"question":"在人类与AI竞争的情境中，标签信息（告知对手是人还是AI）是否独立于对手实际行为影响人们的竞争升级行为？","design":"采用预注册实验（N=1395），使用动态全支付拍卖（重复消耗战）作为竞争任务。通过欺骗自由设计交叉两个维度：告知信息（对手可能是人类、模仿人类的AI、优化竞争的AI或无信息）与实际对手（人类、模仿人类的AI、优化竞争的AI）。测量结果变量为出价升级（中位数价格）和双方均获正收益的概率。","baseline":"真实人类被试（Prolific平台招募）与真实人类对手、两种AI对手（模仿人类和优化竞争）的实际对局数据。","findings":"信息效应显著：仅告知对手可能是人类使中位数出价上升6.7点，告知可能是优化机器使中位数出价下降8.8点，差距约为奖品价值的15%。对手效应表现为与AI实际竞争降低出价，但同时减少双方均获正收益的概率约三分之二。","reliability":"论文未讨论","relevance":"该研究用真实人类被试与AI对手互动，通过标签信息操纵考察行为变化，有真实人类数据对照，直接回应了LLM仿真中信息环境对行为的影响，值得精读以理解信息效应与对手效应的分离方法。","inspiration":"该研究采用欺骗自由设计交叉告知信息与实际对手，分离信息效应与对手效应，方法上可借鉴用于经济实验中标签或框架效应的因果识别。｜可迁移到算法定价或自动竞价场景，研究披露算法身份对市场参与者竞争行为的影响。｜设计实验：招募人类被试参与模拟拍卖或议价博弈，随机告知对手为人类或算法（实际对手固定为算法或人类），测量出价或报价行为，并与真实市场交易数据（如电商平台竞价记录）对照。"}},{"id":"2609.20543","version":1,"title":"Language-model groups overstate consensus when replaying human deliberation on a reasoning task","zh_title":"语言模型群体在重放人类推理任务审议时高估共识","abstract":"Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.","authors":["Tengfei Shao"],"categories":["cs.AI","cs.CL","cs.CY","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20543","pdf_url":"https://arxiv.org/pdf/2609.20543","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人类对照","共识偏差"],"reason":"用LLM代理重放人类推理任务，并与真实人类数据对照，评估仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":3,"question":"用信念锚定的LLM代理组重放人类沃森推理任务讨论时，能否保持人类组的共识结构？","design":"用deepseek-v4-flash模型，为100个留出的人类沃森组中每个真实参与者实例化一个信念锚定代理（锚定其讨论前答案），交叉三种角色保真度、三个种子和两种推理模式，模拟小组讨论，并用与人类相同的代码对最终答案进行共识评分。","baseline":"DeliData语料库中100个真实人类沃森小组的讨论数据，包括参与者的讨论前答案、发言情况和最终答案。","findings":"LLM代理组过度达成完全共识，远高于人类组；在提交制比较中聊天和推理模式的共识差距分别为34.0和43.9个百分点，在参与匹配比较中分别为34.1和44.4个百分点。模拟共识与集体准确性无关，且信念锚定代理组是人类组结果分布的有偏估计量。","reliability":"论文承认代理组几乎总是发言，而人类约五分之一从不发言，导致参与结构不对称；通过提交制和参与匹配两种敏感性分析减少测量不对称，但差距仍存在。还指出在移除可记忆答案的重参数化下，推理模式组几乎一致同意但多为错误答案，表明模拟共识不追踪准确性。","relevance":"该研究直接针对LLM仿真人类群体决策的可靠性，提供了与真实人类数据对照的批判性证据，值得精读以了解仿真在共识测量上的系统性偏差。","inspiration":"借鉴其信念锚定和评分显式化的设计，将LLM代理组与真实人类组在相同任务和评分规则下比较，以识别仿真偏差。｜可迁移到经济金融中的群体决策场景，如投资委员会讨论、信贷审批小组或政策预期形成实验。｜用LLM代理扮演真实实验中的被试，锚定其初始判断，模拟小组讨论后测量共识率和决策准确性，并与原始人类实验数据对照，检验仿真是否高估共识并扭曲结果分布。"}},{"id":"2609.19913","version":1,"title":"Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social Networks","zh_title":"意见动态的数字孪生：面向社交网络的生成式LLM框架","abstract":"The study of opinion dynamics in social networks is one of the key challenges in computational social science with direct relevance to understanding political polarization, misinformation, and health responses. Current approaches focus on simplified mathematical models that ignore linguistic and contextual factors related to belief updates or use Large Language Model (LLM)-based simulations that have not been validated against real data. We present a framework based on the concept of a digital twin to simulate opinion dynamics in social networks. The approach fills the gap by cloning a real-world Twitter network, assigns a set of attributes for agents (such as persona, emotions, centrality, stubbornness, and influence), and employs Mistral-7B to perform opinion update based on memory and social exposure. To evaluate the proposed approach, we validate it against two real Twitter datasets (COVID-19 discourse and U.S elections 2020). The results show that the capability of the proposed framework reproduces opinion trajectories and reduces individual prediction error by more than 50% compared to the best-performing classical baseline (Mistral-7B achieves Mean Absolute Error (MAE) = 0.150 and 0.121 on the COVID-19 and US Election 2020 datasets, respectively). We observe similar improvements in structural alignment (Delta_r = 0.120 and 0.180) and polarization dynamics (Delta_Var = 0.106 and 0.115) on the two datasets, respectively. Additionally, the ablation studies confirm that agent attributes, memory, and social exposure all contribute to the framework's predictive fidelity in reproducing opinion trajectories, with agent attributes being the most critical contributor. Overall, our results demonstrate that grounding Mistral-7B within empirically cloned interaction networks produces a realistic simulation framework capable of reproducing complex social dynamics.","authors":["Omran Berjawi","Giuseppe Fenza","Rida Khatoun","Sherali Zeadally"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19913","pdf_url":"https://arxiv.org/pdf/2609.19913","source_feed":"cs.LG","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","意见动态","数字孪生"],"reason":"用LLM仿真社交网络意见动态，并与真实Twitter数据对照，直接复现人类行为。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":2,"question":"如何利用基于真实社交网络克隆的数字孪生框架，结合LLM代理的认知与语言能力，更准确地模拟和预测社会网络中的意见动态？","design":"使用Mistral-7B作为代理，在从真实Twitter网络克隆的交互结构上运行，每个代理被赋予从真实用户数据提取的属性（如人设、情绪、中心性、固执度、影响力），并基于记忆和社交暴露进行意见更新，以预测个体意见轨迹和集体极化动态。","baseline":"两个真实Twitter数据集：COVID-19话语数据集和美国2020年大选数据集，包含推文内容和用户级元数据，用于验证框架的预测准确性。","findings":"该框架能重现意见轨迹，个体预测误差比最佳经典基线降低50%以上（COVID-19上MAE=0.150，美国大选上MAE=0.121）。在结构对齐和极化动态上也观察到类似改进，消融研究表明代理属性、记忆和社交暴露均对预测保真度有贡献，其中代理属性最关键。","reliability":"论文未讨论","relevance":"该研究直接使用LLM代理在真实社交网络上复现人类意见动态，并与真实Twitter数据对照，属于高相关性的仿真验证工作，值得阅读原文以了解其具体实现和验证细节。","inspiration":"借鉴其将真实网络结构、用户属性与LLM认知过程结合的方法，并利用消融实验识别关键因素。｜可迁移到金融市场情绪传播、政策公告的预期形成或消费者信心扩散等场景。｜以真实投资者社交网络（如StockTwits）为底，用LLM代理扮演投资者并赋予真实用户特征，施加政策新闻或市场事件处理，测量个体情绪和交易倾向变化，并与真实历史数据对照验证。"}},{"id":"2607.25667","version":2,"title":"MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice","zh_title":"MyMentorLLM：用于刻意练习的多模态语音/文本患者、受训者与专家心理治疗GenAI环境","abstract":"Psychotherapists need repeated training and supervision; however, scalability is problematic. We present MyMentorLLM, a multimodal voice- and text-based deliberate-practice environment with 2,100 complete Cognitive Behavioural Therapy (CBT) sessions. Each session links a DSM-5-TR-grounded LLM patient (with major depressive, generalised anxiety or borderline personality disorder), an LLM therapist-in-training and an LLM expert supervisor (powered by Gemma-4, Gemini-3.1-Flash-Live and Qwen-3.6). Sessions were analysed for emotional dynamics, therapeutic competence and diagnostic accuracy against human psychotherapy data. Simulated patients expressed disorder-congruent emotional profiles, which therapists mirrored as in human counselling. LLM trainee competence was rated above human levels in most conditions, while native speech-to-speech was closest to human scores. Supervisor feedback improved diagnostic accuracy in 5 of 7 LLM conditions, whereas symptom identification accuracy increased with model size. This work shows deliberate practice can be simulated for CBT training, although patient fidelity, supervisor calibration and harmful feedback require evaluation via a complex systems perspective.","authors":["Rodolfo Rizzi","Alessandro Grecucci","Massimo Stella"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-18","first_seen":"2026-07-29","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2607.25667","pdf_url":"https://arxiv.org/pdf/2607.25667","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","心理治疗","人类数据对照"],"reason":"用LLM模拟患者和治疗师，并与人类心理治疗数据对照，评估仿真可靠性","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":4,"question":"能否构建一个多模态的LLM心理治疗刻意练习环境，模拟患者、受训治疗师和专家督导，并验证其心理保真度、治疗能力和诊断准确性？","design":"使用Gemma-4、Gemini-3.1-Flash-Live和Qwen-3.6等LLM分别扮演DSM-5-TR诊断的抑郁症、广泛性焦虑症和边缘型人格障碍患者、受训治疗师和专家督导，进行2100次完整的认知行为疗法（CBT）会话，支持语音和文本两种模态；分析会话中的情绪动态、治疗能力和诊断准确性，并与人类心理治疗数据对照。","baseline":"人类心理治疗数据，包括人类咨询中的情绪镜像模式、人类治疗师能力评分和诊断准确性。","findings":"模拟患者表现出与疾病一致的情绪特征，治疗师像人类咨询中一样镜像了这些情绪；LLM受训治疗师的能力评分在多数条件下高于人类水平，其中原生语音到语音模式最接近人类分数。","reliability":"论文承认患者保真度、督导校准和有害反馈需要通过复杂系统视角进行评估；LLM能力评分可能虚高，且文本模态可能丢失副语言信息。","relevance":"该研究用LLM模拟患者和治疗师，并与人类心理治疗数据对照，评估仿真可靠性，直接回应了研究者对LLM人类仿真实验和对照基准的关注，值得精读原文。","inspiration":"借鉴其多角色LLM仿真和与人类基准对照的设计，可用于经济金融中的专业服务场景仿真。｜可迁移到金融咨询或信贷审批中的客户-顾问互动仿真，评估LLM顾问的行为偏差和决策质量。｜以LLM扮演金融顾问和客户，施加不同市场条件或客户特征处理，测量顾问建议的风险偏好和客户满意度，并与人类金融咨询记录或实验数据对照。"}},{"id":"2609.19843","version":1,"title":"A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents","zh_title":"基于LLM的GUI代理对助推易感性的双过程视角研究","abstract":"LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users. While behavioural biases in the textual outputs of LLMs are well-documented, far less is known about how such influence operates when models act as agents that perceive interfaces and execute decisions---and, in particular, whether the reasoning capabilities increasingly built into these agents make them more robust to it. Drawing on Dual-Process Theory, we empirically investigate whether LLM-based GUI agents are susceptible to automatic (Type 1) and reflective (Type 2) digital nudges, and how their reasoning configuration moderates this susceptibility. In a randomized online shopping experiment with 3,600 agents and a total of 21,600 simulations across six frontier models from three providers, we found that agents were vulnerable to both nudge types. Crucially, the reasoning configuration moderated these effects in opposing directions, reducing susceptibility to automatic default nudges while heightening it to reflective social influence nudges. Extensive reasoning therefore did not make agents more robust but redirected the route through which choice architecture takes effect. Exploratory analysis further showed this redirection to be systematically structured by model scale. Beyond establishing nudge susceptibility as a behavioural property of agentic AI, the study positions interface design as a governance concern for organizations that delegate decisions to autonomous agents.","authors":["Haya Halimeh","Sascha Kaltenpoth","Kevin B\\\"osch","Oliver M\\\"uller"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19843","pdf_url":"https://arxiv.org/pdf/2609.19843","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为经济学","助推"],"reason":"用LLM GUI代理模拟人类在数字环境中的决策，研究助推易感性，并与人类行为理…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":5,"question":"LLM-based GUI agents 在数字环境中是否容易受到自动型（Type 1）和反思型（Type 2）数字助推的影响，以及推理配置如何调节这种易感性？","design":"使用 6 个前沿 LLM（来自 3 家提供商）构建 GUI agents，在模拟在线购物环境中进行随机实验，共 3600 个 agents、21600 次模拟。通过改变界面设计施加两类助推：默认选项（Type 1）和社会影响信息（Type 2），并操纵推理配置（低 vs. 高）。结果变量为 agents 的购买选择。","baseline":"无对照（论文未使用真实人类数据作为基准，而是直接测量 agents 的行为）。","findings":"LLM-based GUI agents 对两类助推都表现出易感性。推理配置对两类助推的易感性有相反方向的调节作用：高推理降低了默认助推的易感性，但增加了社会影响助推的易感性。","reliability":"论文未讨论","relevance":"该研究直接使用 LLM agents 模拟人类在数字环境中的决策行为，并检验助推易感性，属于用 LLM 进行人类仿真实验的范畴，但缺乏真实人类数据对照，因此对关注基准对照的研究者价值有限。","inspiration":"值得借鉴的是通过 GUI 环境施加助推并操纵推理配置来研究认知过程对行为偏差的影响。｜可以迁移到消费者在线购物决策、金融产品选择（如默认投资选项、社会影响信息对投资决策的影响）等场景。｜设计雏形：使用 LLM-based GUI agents 模拟投资者在金融平台上的选择，处理为默认投资组合（Type 1）或社会证明信息（Type 2），结果变量为投资选择，并与真实投资者行为数据（如实验或交易记录）进行对照。"}},{"id":"2609.20055","version":1,"title":"What People Almost Did: Evaluating LLM Social Simulations Beyond Behavioral Fit","zh_title":"人们几乎做了什么：超越行为拟合评估LLM社会仿真","abstract":"LLM-based social simulations are primarily evaluated for behavioral fit, testing whether agents reproduce the actions or response distributions of the people they are simulating. However, the promise of simulation extends beyond behavioral fit. Simulations can explain human behavior, diagnose barriers, and compare large-scale interventions. These use cases depend on understanding \\textit{why} people acted a certain way, not just \\textit{what} they did. As a result, behavioral fit is insufficient for these types of claims because behavior underdetermines the reasoning process behind it. For instance, the behavior of staying silent may be due to disinterest or suppressed speech, and not answering a call may be due to distrust of the caller or limited phone access. In this paper, we propose \\textit{representational adequacy} as a new evaluation target for LLM-based social simulations. By leveraging LLM reasoning traces, representational adequacy measures whether a simulation's scenario--reasoning--action triples preserve the reasoning process behind the behavior in a way that is faithful to the population and scenarios being simulated. We distinguish representational adequacy from interpretability and alignment metrics, propose ways to integrate it into simulation research, and pose its measurement as an open problem.","authors":["JaeWon Kim","Angie Boggust"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20055","pdf_url":"https://arxiv.org/pdf/2609.20055","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A2","A4","B4"],"tags":["LLM社会仿真","评估方法","表征充分性"],"reason":"提出表征充分性评估LLM社会仿真，超越行为拟合，关注推理过程，直接针对仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":6,"question":"如何评估基于大语言模型的社会仿真是否保留了被仿真人群的推理过程，而不仅仅是行为结果？","design":"本文不是一项仿真实验研究，而是一篇观点/框架论文。作者提出“表征充分性”作为新的评估目标，主张通过分析LLM在仿真中产生的“场景—推理—行动”三元组，来评估仿真是否忠实于被仿真人群的推理过程。文中没有具体实施仿真或施加处理，而是通过两个例子（青少年社交媒体发帖、孕妇接听健康电话）说明行为相同但推理不同的情况，并讨论如何将表征充分性整合到仿真研究中。","baseline":"无对照","findings":"行为拟合不足以评估用于解释人类行为的LLM社会仿真，因为相同行为可能源于不同推理过程。作者提出表征充分性框架，通过分析场景—推理—行动三元组来评估仿真是否保留了被仿真人群的推理过程，并将其与可解释性和对齐指标区分开来。","reliability":"论文未讨论","relevance":"该文直接针对LLM社会仿真的可靠性问题，提出超越行为拟合的评估框架，强调推理过程的重要性，与研究者关注的仿真可靠性与偏差高度相关，值得阅读原文以了解其理论框架和测量挑战。","inspiration":"本文提出的表征充分性概念可借鉴用于经济金融仿真研究，通过分析LLM代理的推理痕迹来评估其决策过程是否与真实人群一致。｜该框架可迁移到政策评估场景，例如模拟消费者对财政刺激的反应或投资者对央行公告的预期形成，这些场景中行为相同但动机可能不同。｜一个可行的研究设计是：用LLM代理模拟投资者，施加不同措辞的央行公告作为处理，记录代理的推理痕迹和投资决策，并与真实投资者在类似实验中的推理和决策数据（如调查或实验数据）进行对照，评估表征充分性。"}},{"id":"2609.19866","version":1,"title":"Reproducibility is not construct validity: LLM measurement of institutionally situated communication","zh_title":"可重复性不等于构念效度：LLM对制度情境沟通的测量","abstract":"High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({\\=g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.","authors":["Veronika Batzdorfer (KIT)","Carlo Romano Marcello Alessandro Santagiustina (ALMAnaCH, m\\'edialab, Sciences Po)"],"categories":["cs.AI","cs.CL","cs.CY","q-fin.RM"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19866","pdf_url":"https://arxiv.org/pdf/2609.19866","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM测量效度","构念效度","人类数据对照"],"reason":"评估LLM测量效度，区分可重复性与构念效度，有真实人类调查对照，批判性指出失效…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":7,"question":"LLM 从公共咨询文本中推断的构念测量与同一利益相关者在结构化调查中自报的构念测量在多大程度上一致？","design":"本研究不是仿真实验，而是测量效度研究。使用欧洲委员会 AI 法案咨询数据集，将同一利益相关者的自由文本咨询意见与结构化调查回答相链接。用 LLM 对咨询文本进行标注，得到 AI 安全、权利和可解释性关注度的文本测量，并与调查自报的对应构念测量进行比较。","baseline":"同一利益相关者在结构化调查中的自报回答，作为构念测量的基准。","findings":"LLM 标注具有高可重复性（ICC>0.99），但与调查测量收敛性弱（相关系数 0.029-0.176，Lin 一致性系数<0.08）。文本与调查的差异在不同利益相关者群体间存在系统性模式：商业协会在文本中表达的 AI 风险关注高于调查（效应量+1.0），而公共机构和非商业群体差异较小或为负。","reliability":"论文承认 LLM 文本测量可能捕捉的是机构角色和沟通情境，而非纯粹的潜在构念；国家层面的空间自相关结果对多重检验校正敏感，应谨慎解释。","relevance":"该研究直接回应了 LLM 仿真中可重复性与构念效度的区分问题，提供了真实人类调查对照，并批判性地展示了 LLM 测量在机构情境文本中的失效条件，值得精读。","inspiration":"借鉴其将同一主体的不同沟通渠道（调查与公开文本）进行链接并比较 LLM 测量与自报测量的设计，以检验测量效度。｜可迁移到经济金融中的政策沟通研究，例如央行沟通文本与市场参与者调查预期的比较，或上市公司年报文本与分析师调查预期的比较。｜以机构投资者为被试，收集其对央行政策声明的公开评论和内部调查回答，用 LLM 从公开评论中测量政策预期，与调查自报预期对比，并以市场利率变动作为外部效标，检验 LLM 测量在机构沟通情境下的构念效度。"}},{"id":"2606.27845","version":2,"title":"LLM Agents as Static Level-k Players in Behavioural Games","zh_title":"行为博弈中作为静态层级-k玩家的LLM智能体","abstract":"Large Language Models (LLMs) are increasingly used as stand-ins in behavioural games. These stand-ins rely on the assumption that the LLM's distribution of choices meaningfully matches how humans play the same game. This study tests that assumption through two games. The first is a p-beauty contest, and the second one is a public goods game. The study first investigates five local-model settings within the same model family. These settings are varied together in a 360-cell factorial, which balances temperature, scale (0.5-32B), quantisation, instruct vs base, and framing. Each cell's distribution is then compared against whole choice distributions in published human data. Each deployment setting, except for quantisation, governs a different aspect of fidelity. Mechanically, while the dispersion of human players can be somewhat recovered through deployment settings, the strategic process behind it cannot. Through the lens of the level-k cognitive theory, we find that LLMs act as static, category-retrieved level-k players, where k is set by the model scale. The models also do not run within-game belief-updating or backward induction throughout multiple-round horizon settings. While human contributions decayed in the public goods game, LLMs stayed flat or rose at every scale. When the horizon test was administered, LLMs were more cooperative under an indefinite horizon compared to a finite one. However, LLMs ignore their relative round position, so no last-round defection was displayed. This implies that LLMs retrieved levels relative to the horizon category rather than working out iteratively from the specific game setting.","authors":["Po Han Teo"],"categories":["econ.GN","econ.TH","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-09-17","first_seen":"2026-06-26","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2606.27845","pdf_url":"https://arxiv.org/pdf/2606.27845","source_feed":"econ.GN","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为博弈","算法保真度"],"reason":"直接测试LLM在行为博弈中替代人类被试的保真度，并与已发表人类数据对照，发现静…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":1,"question":"LLM在行为博弈中作为人类被试替代品时，其选择分布和策略过程是否与真实人类一致？","design":"使用Qwen 2.5模型家族，在p-beauty contest和公共品博弈中，通过360个单元格的因子设计操纵温度、模型规模（0.5-32B）、量化、指令/基础模型和框架，测量LLM的选择分布，并与已发表的人类数据比较。","baseline":"已发表的人类行为数据，包括p-beauty contest和公共品博弈中的选择分布。","findings":"LLM表现为静态的、类别检索的level-k玩家，k由模型规模决定；LLM没有进行回合内信念更新或逆向归纳，在公共品博弈中贡献不衰减，且忽视相对回合位置，无最后一轮背叛。","reliability":"论文未讨论","relevance":"该研究直接检验LLM在行为博弈中替代人类被试的保真度，并与真实人类数据对照，发现LLM的策略过程与人类不同，对评估LLM仿真可靠性至关重要，值得精读原文。","inspiration":"采用因子设计系统操纵LLM部署设置，并与人类分布整体比较的方法值得借鉴｜可迁移到政策公告的预期形成实验，如央行沟通对通胀预期的影响｜用LLM模拟公众对政策公告的反应，处理为不同政策措辞或沟通方式，结果变量为预期通胀分布，与调查预期数据（如密歇根大学调查）对照。"}},{"id":"2609.17549","version":1,"title":"Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues","zh_title":"社会模式在合成数据中是否成立？分析LLM生成与真实对话中的网络欺凌动态","abstract":"Cyberbullying (CB) is a complex social phenomenon characterized by repeated aggression, power imbalance, and multi-party interaction. Although large language models (LLMs) are increasingly used to generate synthetic CB conversations for data augmentation and benchmarking, it remains unclear whether such data faithfully reproduces the social dynamics of authentic interactions beyond supporting downstream task performance. We present a comprehensive framework for evaluating the social realism of LLM-generated CB conversations. We compare authentic and synthetic dialogues generated by GPT, Grok, and LLaMA across interactional structure (turn-taking, power dynamics, and repair behavior), linguistic and stylistic realism (pronoun usage and humor), affective and behavioral markers (CB types, profanity, and toxicity), and temporal escalation dynamics. We further complement automatic analyses with a human evaluation of cyberbullying presence, scenario relevance, role plausibility, and social realism. Our results show that LLM-generated data consistently preserves high-level interactional structure, including role participation patterns, directional power asymmetry, and broad distributions of behavioral markers. However, all models systematically distort finer-grained social phenomena, including behavioral magnitude, role-specific allocation, categorical distributions, and temporal dynamics. These distortions are strongly model-dependent: GPT suppresses harmful content, Grok amplifies aggressive behaviors, and LLaMA provides the most balanced approximation while smoothing role distinctions. Our findings show that synthetic CB data is useful for modeling global interactional structure but remains an imperfect substitute for authentic conversations when behavioral realism and social dynamics are essential.","authors":["Arefeh Kazemi","Hamza Qadeer","Sinan Asci","Joachim Wagner","Brian Davis"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17549","pdf_url":"https://arxiv.org/pdf/2609.17549","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","社会动态","真实性评估"],"reason":"用LLM生成对话仿真网络欺凌社会动态，并与真实对话对照，评估仿真保真度与偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":2,"question":"LLM生成的网络欺凌对话是否忠实再现了真实对话中的社会动态，而不仅仅是支持下游任务性能？","design":"使用GPT、Grok和LLaMA生成合成网络欺凌对话，与真实对话对比，分析互动结构（话轮转换、权力动态、修复行为）、语言风格（代词使用、幽默）、情感行为标记（欺凌类型、脏话、毒性）和时间升级动态，并进行人类评估。","baseline":"真实网络欺凌对话数据集（具体名称未在节选中提及）。","findings":"LLM生成的数据保留了高层互动结构，如角色参与模式、方向性权力不对称和行为标记的广泛分布；但所有模型都系统性地扭曲了细粒度社会现象，包括行为强度、角色特定分配、类别分布和时间动态，且扭曲程度因模型而异：GPT抑制有害内容，Grok放大攻击行为，LLaMA提供最平衡的近似但平滑了角色差异。","reliability":"论文承认合成数据在需要行为真实性和社会动态时是真实对话的不完美替代品，且扭曲具有模型依赖性。","relevance":"该研究直接评估LLM仿真社会互动的保真度，与研究者关注的人类仿真实验和可靠性评估高度相关，提供了系统的对照基准和偏差分析，值得精读。","inspiration":"借鉴其多维评估框架，将自动分析与人类评估结合，系统比较合成与真实数据在结构、语言、行为和时间动态上的差异｜可迁移到经济金融中的社会互动场景，如谈判博弈、团队决策或市场情绪传播｜设计一个实验：用LLM生成模拟投资者在社交媒体上的互动对话，处理为不同模型（如GPT、Claude）或提示策略，结果变量为情绪传染、羊群行为或信息扩散模式，与真实投资者论坛数据（如StockTwits）对照，评估仿真保真度。"}},{"id":"2609.18106","version":1,"title":"Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment","zh_title":"开放权重大语言模型应用于招聘时性别与种族偏见的语言触发因素","abstract":"Open-weight large language models are rapidly entering hiring pipelines, yet their discriminatory failure modes -- and the regulatory exposure these create under the EU AI Act high-risk classification (Annex III) and U.S. EEOC adverse-impact analysis -- remain poorly understood. We present the first systematic, multi-model audit of open-weight LLMs that treats job-posting language as the primary experimental variable, evaluating six models (Llama 3.2, Mistral, Gemma 3, Qwen 3, Phi 3, DeepSeek-R1) across four controlled experiments that jointly probe recruiter-simulation and job-seeker-simulation tasks. We find that (1) agentic posting language depresses recruiter recommendation scores for female candidates (r_rb = 0.309, p_Bonf = 7x10^-5; model-fixed-effects r_rb = 0.448), while communal language partially reverses the penalty; and (2) coded-exclusion language suppresses non-White recruiter scores at large effect sizes (r_rb = 0.646-0.758) and, on the job-seeker side, selectively deters non-White personas from expressing interest -- operationalizing a chilling-effect mechanism at scale. A label-ablation experiment isolates the explicit demographic persona label as the primary causal driver, and Word Embedding Association Tests corroborate these findings at the representational level (d = 1.01-1.45 under Caliskan et al.'s multi-word gender attribute lists). We translate these results into a concrete pre-deployment audit protocol -- posting-vocabulary scoring, persona-conditioned LLM probing, and adverse-impact flagging against the four-fifths threshold -- that operationalizes the documentation and risk-management obligations Annex III imposes on high-risk AI in recruitment.","authors":["Kosuke Kitahara","Nobuhiro Yamaguchi"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18106","pdf_url":"https://arxiv.org/pdf/2609.18106","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","招聘偏见","算法审计"],"reason":"用LLM仿真招聘中的人类决策，并与真实人类数据对照，评估偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":5,"question":"招聘广告中的语言特征（如代理性/社群性词汇、编码排斥语言）如何触发开放权重LLM在招聘模拟中的性别与种族偏见？","design":"使用六个开放权重LLM（Llama 3.2、Mistral、Gemma 3、Qwen 3、Phi 3、DeepSeek-R1）进行招聘者模拟和求职者模拟实验。通过操纵招聘广告语言（代理性 vs. 社群性词汇、编码排斥语言）和候选人/求职者的人口统计标签（性别、种族），测量LLM给出的推荐评分、兴趣表达等结果变量，并进行标签消融实验和词嵌入关联测试。","baseline":"无对照（论文未使用真实人类数据作为基准，而是基于LLM输出进行审计）。","findings":"代理性招聘语言降低女性候选人的推荐评分，而社群性语言部分逆转该惩罚；编码排斥语言大幅抑制非白人候选人的推荐评分，并在求职者模拟中抑制非白人角色的兴趣表达，形成寒蝉效应。标签消融实验表明显式人口统计标签是主要因果驱动因素，词嵌入关联测试在表征层面证实了偏见。","reliability":"论文未讨论","relevance":"该研究直接针对LLM在招聘场景中的仿真行为，系统操纵语言变量并测量偏见输出，与您关注的LLM仿真可靠性及偏差评估高度相关，值得阅读原文以了解其审计协议和效应量。","inspiration":"借鉴其将文本特征作为处理变量、通过多模型审计和标签消融识别因果机制的方法。｜可迁移到信贷审批中的语言歧视研究，如贷款广告或申请表中的措辞对AI审批决策的影响。｜以LLM作为信贷审批员，处理为贷款申请描述中的代理性/社群性词汇或编码排斥语言，结果变量为审批通过率或利率，对照真实信贷审批数据（如抵押贷款披露数据）评估仿真偏差。"}},{"id":"2609.17933","version":1,"title":"AI Mediators Regulate Emotion and Create Value in Disputes","zh_title":"AI调解员在纠纷中调节情绪并创造价值","abstract":"In conflict and disputes, especially, emotion acts as a salient force in influencing outcomes. Prior work shows negative affect can obstruct collaborative behaviors, which typically lead to ``win-win'' outcomes. Thus, some suggest mediators may help regulate emotion and achieve joint gains. With the proliferation of AI, we posit LLMs may perform well at this task, with the added benefit of better accessibility compared with a human mediator. To examine the effectiveness of AI versus novice human mediators, we conduct a between-subjects experiment, where participants engage in a dispute mediated by a human, AI, or no mediator. We first analyze how well the mediators regulate emotions within a dispute -- finding AI mediators perform significantly better than humans at reducing negative emotion. We next examine whether AI mediators facilitate disputants better realizing joint gains in disputes with high integrative potential (IP) -- we find a marginally significant interaction between IP and condition (AI versus human), indicating LLMs may outperform humans at aiding disputants realize joint gains. Lastly, we perform an analysis of the messages the mediators sent, finding the AI sent significantly more messages suggesting trade-offs compared to the humans.","authors":["James Hale","Jonathan Gratch"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17933","pdf_url":"https://arxiv.org/pdf/2609.17933","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","调解实验","人机对照"],"reason":"用LLM作为调解人替代人类，与真实人类被试互动并对照，属于人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":3,"question":"在情绪化的纠纷调解中，AI调解员相比新手人类调解员能否更有效地调节情绪并帮助双方实现共同收益？","design":"使用大语言模型（LLM）扮演AI调解员，与新手人类调解员和无调解员条件进行被试间实验。参与者扮演买卖双方进行纠纷谈判，调解员在每轮发言后决定是否干预并发送消息。结果变量包括情绪调节效果（负面情绪减少）、共同收益实现程度（积分潜力IP与条件的交互）以及调解员消息内容（如提出权衡建议的频率）。","baseline":"新手人类调解员作为对照，以及无调解员条件作为基线。参与者为真实人类被试，其偏好通过分配100点来测量，并根据达成目标获得金钱奖励。","findings":"AI调解员在减少负面情绪方面显著优于人类调解员；在具有高整合潜力的纠纷中，AI调解员可能比人类调解员更能帮助双方实现共同收益（边际显著交互作用）。此外，AI调解员发送了更多建议权衡的消息。","reliability":"论文未讨论","relevance":"该研究直接使用LLM替代人类调解员，与真实人类被试互动并对照人类调解员表现，属于人类仿真实验，且涉及情绪调节和谈判结果，与研究者关注的经济学实验和政策评估场景高度相关，值得阅读原文了解实验细节和局限性。","inspiration":"借鉴其使用LLM作为干预代理与人类被试互动的实验设计，以及通过偏好分配和积分潜力测量共同收益的方法。｜可迁移到经济金融中的谈判或冲突解决场景，如劳资纠纷调解、商业合同谈判、消费者投诉处理等。｜设计一个实验：以LLM作为调解员，人类被试扮演谈判双方（如买方和卖方），处理为AI调解员、人类调解员或无调解员，结果变量为谈判达成协议的质量（如联合收益、满意度）和情绪变化，对照真实人类调解员数据，并利用已有谈判实验数据集（如KODIS）进行基准比较。"}},{"id":"2609.18060","version":1,"title":"AI Peers Exert Social Influence on Human Dishonesty in Groups","zh_title":"AI同伴对群体中人类不诚实行为施加社会影响","abstract":"Human dishonesty in group settings is highly susceptible to peer influence, particularly when incentivized. Although artificial intelligence (AI) evolves from passive tools into active collaborators, its impact on human moral behavior within groups remains underexplored. We addressed this gap through a two-phase randomized behavioral study (N=280 and N=360). We found AI agents exert substantial social influence comparable in magnitude to that of human peers. Specifically, participants reported more dishonestly when exposed to dishonest rather than honest normative cues. This effect is evident across injunctive, subjective, and descriptive social norms. Interestingly, the only significant adjacent behavioral change occurred when dishonest peer behavior first appeared, whereas further increases from one to four dishonest peers produced weaker and non-monotonic changes. Furthermore, participants rapidly converge on decision-making, showing modest increases in dishonest reporting through repeated exposure. These findings highlight the importance of managing the behaviors and normative signals communicated by AI group members.","authors":["Shuning Zhang","Xinyuan Zhou","Yuanyang Qiu","Tianqi Song","Yuting Yang","Yiwen Ren","Xin Yi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18060","pdf_url":"https://arxiv.org/pdf/2609.18060","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","社会影响","行为实验"],"reason":"用AI代理替代人类同伴，研究其对人类不诚实行为的社会影响，并与人类同伴对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":4,"question":"AI同伴的诚实规范如何影响群体中人类的不诚实行为？","design":"采用两阶段随机行为实验（N=280和N=360），使用激励性掷骰子报告任务。AI代理作为群体成员，通过描述性、指令性和主观性社会规范传递诚实或不诚实线索，测量人类被试的虚报行为。","baseline":"与人类同伴的社会影响进行对照，比较AI同伴与人类同伴的影响大小。","findings":"AI代理对人类不诚实行为产生显著社会影响，其影响程度与人类同伴相当。当出现不诚实同伴时，被试虚报增加，但同伴数量从1个增加到4个时影响减弱且非单调。","reliability":"论文未讨论","relevance":"该研究直接使用AI代理替代人类同伴，研究其对人类道德行为的影响，并与人类同伴对照，符合研究者对LLM仿真实验和真实人类基准的兴趣。","inspiration":"值得借鉴的是其将AI代理作为群体成员，通过操纵其行为规范来研究社会影响，并设置人类同伴对照组以量化AI影响的大小。｜可迁移到经济金融中的群体决策场景，如投资团队中AI顾问的不诚实建议对个人投资决策的影响。｜设计一个实验：被试与AI代理组成投资小组，AI代理提供虚报收益的建议（处理），测量被试的投资报告虚报程度（结果变量），并与人类同伴提供同样建议的对照组比较，同时使用真实投资数据作为外部基准。"}},{"id":"2609.07478","version":2,"title":"The Internal Anatomy of Strategic Choice in Large Language Models","zh_title":"大语言模型策略选择的内部解剖","abstract":"Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal $2\\times2$ games. We followed a prespecified incentive from prompt, through activations, to choice. Dense models mirrored the unadjusted human decline with game complexity. Incentive and choice were detectable in every model, but models differed in whether incentive reached the choice, aligned with it and, where tested, whether strengthening it shifted preference. The base and instruction-tuned Qwen2.5 models chose almost identically at baseline yet differed in whether incentive reached choice. Fixed decision cues were distinguishable internally but changed choices selectively. Similar behaviour can rest on different computation; post-training can reshape the path from represented incentive to decision while leaving behaviour and decodable information largely intact.","authors":["Vin\\'icius Ferraz","Leon Houf","Enrico Ferrea"],"categories":["cs.AI","cs.GT"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-09","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.07478","pdf_url":"https://arxiv.org/pdf/2609.07478","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","策略博弈","算法保真度"],"reason":"用LLM复现人类策略选择并与人类数据对照，分析内部计算与行为差异，评估仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":6,"question":"大语言模型在一次性策略博弈中，内部表征与计算如何将激励信号转化为选择行为？","design":"用四个开源大语言模型（Qwen2.5 base/instruct、Llama-3.1-Instruct、GPT-OSS）作为被试，在144个严格序数2x2博弈中做一次性选择，记录激活值，用线性探针解码激励和选择，并进行激励强化干预和固定决策线索提示。","baseline":"人类选择数据来自Moore et al. (2026)的相同博弈和支付尺度，以及Zhu et al. (2025a)的基数支付数据集映射。","findings":"密集模型复现了人类随博弈复杂度下降的未调整选择模式；所有模型都能解码激励和选择，但激励是否到达选择、与选择对齐及干预效果因模型而异，基础与指令微调模型行为相似但内部路径不同。","reliability":"论文指出相似行为可能基于不同计算，后训练可重塑激励到决策的路径而保持行为和可解码信息基本不变，暗示行为仿真可能掩盖内部差异。","relevance":"该研究直接以LLM仿真人类策略选择并与真实人类数据对照，同时揭示内部计算与行为的不一致，对评估仿真可靠性和偏差具有重要参考价值，值得精读原文。","inspiration":"借鉴其用线性探针追踪激励信号从输入到选择的内部路径，并施加干预检验因果性的方法，可迁移到经济决策仿真中检验模型是否真正使用经济激励而非表面模仿。｜可应用于资产定价实验，检验LLM是否内部表征风险溢价并据此决策。｜以LLM为被试，呈现不同风险收益的资产选择任务，用探针解码风险溢价信号并强化干预，与人类实验数据（如股票市场参与决策）对照，观察选择变化。"}},{"id":"2609.17534","version":1,"title":"Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits","zh_title":"LLM中的装好与装坏：黑暗三人格特质下的反应失真","abstract":"Social desirability and impression management are pervasive sources of response distortion in human personality assessment, yet their effects on Large Language Models (LLMs) remain underexplored. This study investigates whether contemporary LLMs systematically modulate the expression of Dark Triad traits (Machiavellianism, narcissism, and psychopathy) under fake-good and fake-bad conditions. Seven state-of-the-art models were evaluated across two ecologically relevant contexts: employment selection and forensic evaluation, in which socially desirable or undesirable incentives were conveyed through contextual framing. Trait expression was measured using standard psychometric scoring procedures and compared with self-assessment baselines at both aggregate and item levels. Results revealed systematic and condition-consistent response modulation. Most models reduced Dark Triad scores under fake-good conditions and increased them under fake-bad conditions, although the magnitude and consistency of these effects varied across traits and models. Machiavellianism and narcissism showed the strongest and most coherent shifts, whereas psychopathy displayed greater heterogeneity. Context also influenced responses, with employment scenarios generally producing larger effects than forensic scenarios. An additional experiment showed that explicit fake-bad instructions generated substantially stronger distortions than contextual framing alone. The results suggest that personality-related outputs should be interpreted in light of the motivational and situational context in which they are elicited. More broadly, they highlight the value of psychometric paradigms for evaluating susceptibility to response distortion, impression management, and context-dependent behavioral shifts, with important implications for LLM benchmarking, alignment evaluation, and robustness assessment.","authors":["Victoria Popa","Guglielmo Cola","Caterina Senette","Maurizio Tesconi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17534","pdf_url":"https://arxiv.org/pdf/2609.17534","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM人格测量","反应偏差","仿真可靠性"],"reason":"研究LLM在人格测量中的反应失真，与人类数据对照，评估仿真偏差，可迁移到人类仿…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":7,"question":"LLM 在人格测量中是否会在假好（fake-good）和假坏（fake-bad）条件下系统性地调节黑暗三联征特质的表达？","design":"使用七个先进LLM，在就业选拔和司法评估两种情境下，通过提示框架施加假好或假坏动机，用TRAIT框架测量黑暗三联征（马基雅维利主义、自恋、精神病态）得分，并与基线自我评估比较。","baseline":"无对照","findings":"大多数模型在假好条件下降低黑暗三联征得分，在假坏条件下提高得分，但效应大小和一致性因特质和模型而异。马基雅维利主义和自恋的偏移最强且最一致，精神病态异质性更大；就业情境比司法情境产生更大效应，显式假坏指令比情境框架产生更强扭曲。","reliability":"论文未讨论","relevance":"该研究展示了LLM对情境动机的敏感性，可用于评估仿真中的社会期望偏差和印象管理，对使用LLM模拟人类被试时的可靠性有警示意义。","inspiration":"借鉴其通过情境框架施加动机处理并测量特质偏移的方法，可迁移到经济金融中的社会期望偏差场景（如信贷审批中的歧视、消费者道德行为）。｜可设计实验：用LLM扮演贷款申请人，在强调社会责任或利润最大化的不同银行政策下，测量其自我报告的诚信或风险偏好，并与实际信贷数据中的偏差对照。"}},{"id":"2609.18282","version":1,"title":"Too Good to Be Real? Diagnosing and Reducing the Gap Between AI Preference and Real User Engagement","zh_title":"好得难以置信？诊断并缩小AI偏好与真实用户参与度之间的差距","abstract":"Large language models are increasingly used to generate and evaluate online content, yet it remains unclear whether the qualities they associate with higher engagement match what real users respond to. We study this question using 1.17 million answers to 25,978 questions from Zhihu, Quora, and Reddit, comparing real platform answers and AI-generated answers across four within-question engagement levels. We introduce Ontological Preference Measurement, which represents answers along three dimensions: logic, affect, and expression. We find a systematic gap between AI preference and real user engagement: as target engagement increases, LLMs add more explicit logical structure, while real user engagement is more strongly associated with affective and expressive salience. We call this tendency logic overbinding. Based on this diagnosis, we propose Ontology-Masked Reasoning Autoencoding (OMRA), a controlled intervention that masks and reconstructs over-explained spans while preserving stance, factual content, and coherence. Across four LLM families, OMRA reduces the measured gap by an average of 54.4%. In human evaluation, OMRA wins 62.4% of pairwise preference judgments against matched real platform answers, even though the real answers are more often judged to be human-written.","authors":["Xinglang Zhang","Yuanmeng Xiang","Yunyao Zhang","Zeliang Chen","Junqing Yu","Zikai Song"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18282","pdf_url":"https://arxiv.org/pdf/2609.18282","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类行为对照","内容生成"],"reason":"用LLM生成内容并与真实用户互动数据对照，诊断AI偏好与人类行为差距，属于仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":9,"question":"LLM 在生成高互动内容时偏好的文本特征是否与真实用户互动行为一致？","design":"使用 LLM 生成针对四个互动等级的答案，与真实平台答案对比，通过本体偏好测量框架从逻辑、情感、表达三个维度量化文本特征，并施加 OMRA 干预以缩小差距。","baseline":"来自知乎、Quora 和 Reddit 的 117 万条真实答案，按问题内投票排名分为四个互动等级。","findings":"LLM 在追求高互动时过度增加显性逻辑结构，而真实高互动答案更依赖情感和表达显著性，作者称之为“逻辑过度绑定”。OMRA 干预平均缩小 54.4% 的差距，并在人类评估中胜过真实答案。","reliability":"论文未讨论","relevance":"该研究直接对比 LLM 生成内容与真实用户行为，诊断 AI 偏好与人类反应的系统性偏差，并尝试通过干预修正，对评估 LLM 仿真人类行为的可靠性具有参考价值。","inspiration":"借鉴其本体驱动的多维测量和干预设计，可迁移到经济金融领域的文本生成场景，如政策沟通、市场评论或金融建议。｜例如，研究 LLM 生成的央行政策声明或分析师报告是否与真实市场反应匹配。｜以 LLM 为被试，要求其生成不同目标市场反应的金融文本，测量逻辑、情感、表达特征，并与真实市场数据（如股价波动、交易量）对照，检验 AI 偏好与真实投资者反应的差距。"}},{"id":"2609.17989","version":1,"title":"Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders","zh_title":"AI代理为谁工作？角色分配引发LLM推荐中的赞助偏差","abstract":"Large language models (LLMs) now serve as conversational shopping assistants on platforms that also sell advertising. These AI agents face a conflict of duty. They advise consumers who rely on their judgment, yet are deployed by platforms that benefit when sponsored listings are chosen. Sponsorship disclosures, designed to allow consumers to penalize paid placements, now reach the AI agent rather than the consumer, and the agent's evaluation of them is hidden from the consumer. Drawing on the fiduciary concept of conflict of duty, we argue that an agent's evaluation of a sponsored listing should not depend on which party deployed it. In controlled choice experiments, we manipulate assigned roles in the system prompt to name either a traveler or a booking platform as the agent's principal. Platform delegation significantly attenuates the penalty that agents apply to sponsored listings and weakens the skepticism that disclosure triggers in their reasoning traces. We replicate out findings across LLMs and reasoning depths. A second study decomposes the disclosure label and shows that the divergence between the two delegates widens significantly when the paid placement is attributed to the platform. Stricter terminology (\"Sponsored\" instead of \"Promoted\") lowers choice of paid listings but does not close this gap when the platform is named. The findings show that disclosure mandates designed for human consumers cannot by themselves protect consumers in AI-mediated commerce.","authors":["Davood Wadi","Yu Ma"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17989","pdf_url":"https://arxiv.org/pdf/2609.17989","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","消费者决策","赞助偏差"],"reason":"用LLM模拟消费者决策，与人类对照，揭示角色分配导致的赞助偏差，属经济学实验场…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":8,"question":"AI代理在同时服务消费者和广告平台时，其推荐行为是否会因委托方身份（消费者vs平台）而出现赞助偏差？","design":"用LLM（Gemini 3.1 Pro等）模拟购物助手，在系统提示中操纵委托方身份（旅行者vs预订平台），呈现带赞助标签的酒店列表，测量选择赞助列表的概率及推理痕迹中的怀疑程度。","baseline":"无对照","findings":"平台委托显著减弱了LLM对赞助列表的惩罚，选择赞助列表的概率从消费者委托时的50.2个百分点降至29.2个百分点；当赞助标签明确归属平台时，平台委托的代理甚至表现出对赞助列表的强烈偏好（选择率74.2% vs 消费者委托的34.6%）。","reliability":"论文未讨论","relevance":"该研究用LLM模拟消费者决策，揭示了角色分配导致的赞助偏差，属于经济学实验场景，与研究者关注的人类仿真实验高度相关，值得精读原文。","inspiration":"值得借鉴的是通过最小化系统提示中的角色分配来诱发LLM行为偏差，并分析推理痕迹以揭示机制｜可迁移到信贷审批歧视或金融产品推荐场景，检验AI代理在银行与借款人之间的利益冲突｜设计：用LLM扮演贷款顾问，系统提示中分别指定委托方为银行或借款人，呈现带“银行推荐”标签的贷款产品，测量选择概率，并与人类贷款顾问的真实选择数据对照。"}},{"id":"2609.16395","version":1,"title":"Silicon sampling answers with country-level assumptions, not individual attitudes: Cross-national evidence from the European Social Survey","zh_title":"硅采样以国家层面假设而非个体态度作答：来自欧洲社会调查的跨国证据","abstract":"Silicon sampling uses large language models (LLMs) to simulate survey respondents. Whether it recovers cross-national variation, and why, remains unresolved. This study evaluates it against European Social Survey Round 11 (30 countries, 42 items) with two open-weight LLMs under first- and third-person prompts, plus backstory and response-format experiments. Aggregate recovery is moderate and uneven across items. Adding the country name to a three-variable demographic backstory raises the median per-item correlation between simulated and observed country means from -0.03 to 0.52, and the richer profiles tested add no consistent gain. The respondent's country label acts as a country-level assumption that respondent detail does not revise. Naming the response-scale endpoints in words stops the model from ranking countries backwards, so the answer format sets the direction of the ranking. Individual-level recovery remains negligible in every condition and does not track aggregate recovery across countries. An average of neighboring countries, using no LLM, recovers country levels more accurately than every model condition and ranks them about as well. Silicon sampling can thus support exploratory country-ranking comparison after item-level validation and with the response format reported. It does not support individual or distributional inference.","authors":["Chuyao Wang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16395","pdf_url":"https://arxiv.org/pdf/2609.16395","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","调查方法","跨国比较"],"reason":"直接评估LLM仿真调查回答，与真实跨国调查数据对照，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":3,"question":"硅采样能否恢复跨国调查中的国家间差异，以及这种恢复的来源是什么？","design":"用两个开源权重LLM模拟欧洲社会调查（ESS）第11轮30个国家的受访者，在42个态度题项上生成回答；通过第一/第三人称提示、添加国家名称和人口学背景故事、以及改变回答格式端点措辞等实验条件，测量模拟国家均值与真实均值的相关性及个体层面恢复情况。","baseline":"欧洲社会调查（ESS）第11轮30个国家、42个题项的真实人类回答数据。","findings":"国家标签是主要的聚合信号来源，添加国家名称后中位题项相关从-0.03升至0.52，更丰富的人口学背景没有额外增益；回答格式端点措辞决定国家排序方向，个体层面恢复在所有条件下均可忽略，且与聚合恢复不相关。","reliability":"论文指出硅采样仅适用于探索性国家排序比较，且需逐题验证并报告回答格式；不支持个体或分布推断，个体恢复不随聚合恢复变化。","relevance":"该研究直接评估LLM仿真调查回答，与真实跨国调查数据对照，并明确指出失效条件，对关注人类仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其通过消融实验分离国家标签与人口学背景贡献的方法，以及改变回答格式端点来检验排序方向的做法｜可迁移到跨国经济态度调查仿真，如通胀预期、政策偏好或金融素养的跨国比较｜用LLM模拟不同国家受访者对经济政策的态度，处理为是否添加国家标签及回答格式端点措辞，结果变量为模拟国家均值排序，对照真实跨国调查数据（如欧洲央行消费者预期调查）验证。"}},{"id":"2609.17317","version":1,"title":"Towards Detecting AI-Assisted Responses in Online Surveys","zh_title":"检测在线调查中AI辅助回答的方法研究","abstract":"The use of LLMs to complete online surveys impacts the validity of survey-based research, but detecting such usage remains underexplored. We introduce an initial benchmark dataset, namely ASURRE, for AI-assisted survey participation to capture usage strategies ranging from full generation and revision to persona-grounded agentic completion. Controlled by these strategies, LLM-assisted survey responses are generated using multiple LLMs on three real-world surveys in different disciplines, paired with genuine human responses. Our evaluation of existing machine-generated text (MGT) detectors shows that naive AI usage is readily detectable, whereas persona-grounded agents that mimic entire respondents push detector performance toward chance. We further show that agentic completion cannot fully replicate respondent-level behaviour and leaves distinctive behavioural traces. While individual cues can be circumvented by targeted prompting, a simple few-shot, training-free aggregator over these cues improves mean AUROC by +0.14 over the best existing detector across agentic settings. Our project is available at https://github.com/mike-qz-wang/ASURRE.","authors":["Qizhou Wang","Bogdan Mamaev","Christopher Leckie"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17317","pdf_url":"https://arxiv.org/pdf/2609.17317","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查数据","检测方法"],"reason":"直接研究LLM仿真人类调查回答，并与真实人类数据对照，评估检测与行为差异。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":10,"question":"如何检测在线调查中由大语言模型辅助生成的回答，尤其是在不同使用策略下的可检测性？","design":"构建ASURRE基准数据集，使用多个开源和闭源LLM在三个真实调查上生成回答，模拟三种使用策略：完全生成、修订和基于人格的智能体完成；评估现有机器生成文本检测器，并提出SPABD聚合器。","baseline":"三个真实调查中的人类真实回答，与LLM生成回答配对。","findings":"完全生成易被检测，修订部分可检测，而基于人格的智能体完成几乎无法被现有检测器识别；SPABD聚合器通过整合多个行为线索，在智能体设置下将平均AUROC从0.61提升至0.75。","reliability":"论文承认SPABD在对抗性提示下性能下降，且在不同调查和提示模式下表现不稳定，应视为轻量级参考检测器而非完整解决方案。","relevance":"该研究直接评估LLM仿真人类调查回答的可检测性，并与真实人类数据对照，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其构建多策略仿真基准和利用行为痕迹进行检测的方法，可迁移到经济金融领域的调查数据质量评估，例如消费者信心调查或投资者情绪调查；设计研究时，可用LLM模拟不同人格的受访者，施加不同提示策略，以真实调查数据为基准，评估检测器性能并分析行为偏差。"}},{"id":"2609.16436","version":1,"title":"Interpreting and Steering LLM Agents for Social Simulations","zh_title":"解释与引导用于社会模拟的LLM智能体","abstract":"Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive specific model behavior. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. Specifically, we compare three types of methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering and examine their utility for LLM-based social scientific simulations. We do so by interpreting and steering two foundational components of human behaviors, namely preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation), operationalized using four classic economic and creative tasks implemented as natural-language interactions. Overall, our results show that SAE- and probe-based techniques often outperform basic prompt-based methods for steering LLM agents, although this advantage depends on the specific prompting strategy involved. Together, SAEs and probes constitute an effective pipeline for social scientists seeking to interpret and steer agents in social simulations: SAEs decompose agents' internal representations into human-readable features, after which probes can reliably shift agents' behaviors in specified directions. We discuss implications of these methods for future work using LLM agents for social scientific simulations.","authors":["Jiayue Gaveal Fan","Arul Murugan","Shreyas Krishnan","Abhishek Nagaraj"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16436","pdf_url":"https://arxiv.org/pdf/2609.16436","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","可解释性","行为经济学"],"reason":"用LLM仿真人类行为，有真实人类数据对照，涉及经济任务，并批判性评估方法。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":9,"question":"如何打开LLM黑箱，通过可解释性和可操控性方法提升LLM智能体在社会仿真中的效用？","design":"使用LLM智能体模拟人类被试，在四个经典任务（彩票游戏测风险态度、最后通牒游戏测利他、发散创造力、产品创新）中，比较三种干预方法：基于提示的操控、SAE特征操控、探针方向操控，测量智能体行为变化。","baseline":"无对照","findings":"SAE和探针方法在操控LLM智能体行为上通常优于基本提示方法，但优势取决于具体提示策略；SAE能无监督地发现行为背后的可解释特征，探针能精确控制特定特征的程度。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真中的可解释性和可操控性，提供了超越提示工程的技术路径，对关注仿真可靠性和机制理解的研究者具有重要参考价值。","inspiration":"借鉴其使用SAE和探针进行特征级操控的方法，可实现对LLM智能体内部表征的精细干预，比提示词更可控。｜可迁移到经济决策仿真，如风险偏好、时间偏好、社会偏好等实验，以及政策评估中的行为反应模拟。｜以LLM智能体模拟投资者，用探针操控其风险厌恶特征，观察在资产配置任务中的选择，并与真实投资者调查或实验数据对照，验证操控的有效性和仿真保真度。"}},{"id":"2608.18768","version":2,"title":"Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model","zh_title":"可读、忠实、被使用：语言模型中人口统计身份的三个可分离属性","abstract":"Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differences, and whether this reflects what a model knows or uses has remained untested. Using representational similarity analysis against Pew American Trends Panel ground truth, we score demographic read-out locations in Mistral-7B and intervene causally across six attribute types. The internal geometry is faithful: attention-head read-outs dominate the standard residual read-out, reaching selection-corrected $\\rho$ up to 0.63 -- about 70% of the measurement-reliability ceiling -- and one head, L11 H16, is significantly faithful across all six types, though race-based types stay weak and prompt-fragile, replicating in a second model family. Yet causal use does not track fidelity: the clearest causal pathway ($p=0.002$) sits in one of the least faithful types, the most faithful type shows no correction-surviving effect, and full identity swaps in the prompt move predictions by under 2% of their error. A 128-dimensional probe on that head lands 21-31% closer to survey truth than the model's answers, yet recovers almost none of the per-question group ordering. Readable, faithfully arranged, and causally used are three dissociable properties of the same model; treating them as one claim is what keeps the \"can LLMs simulate populations\" debate unresolved.","authors":["Fathin Difa Robbani"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-20","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.18768","pdf_url":"https://arxiv.org/pdf/2608.18768","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","算法忠实度","调查方法"],"reason":"直接研究LLM仿真调查受访者，用真实Pew数据对照，评估忠实度与因果使用，并批…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":1,"question":"LLM内部的人口统计身份表征是否忠实于真实群体差异，以及这些表征是否被模型实际用于生成回答？","design":"使用Mistral-7B模型，对169个交叉人口统计单元构建身份提示，提取残差流、注意力头输出和FFN输出等内部表征，与Pew调查的真实群体回答分布进行表征相似性分析（RSA），并通过激活修补进行因果干预，测量模型预测分布的变化。","baseline":"Pew American Trends Panel（ATP）调查数据，包含15波、169个交叉人口统计单元的真实回答分布。","findings":"内部几何结构是忠实的：注意力头读出（尤其是L11 H16）与真实群体差异的相关性高达0.63，约为测量可靠性上限的70%，且跨六种属性类型显著。但因果使用与忠实性脱节：最清晰的因果路径位于低忠实性类型中，而最忠实的类型没有通过校正的因果效应；完整身份替换仅使预测移动不到误差的2%。","reliability":"论文承认种族和宗教类型的忠实性较弱且对提示脆弱；忠实表征并未被模型实际用于生成回答；探针虽在平均上更接近真实，但无法恢复每个问题的群体排序，因此不能替代调查数据。","relevance":"直接研究LLM仿真调查受访者的可靠性，用真实Pew数据对照，并批判性地指出忠实表征与因果使用脱节，对理解仿真失效条件至关重要，值得精读原文。","inspiration":"借鉴其表征相似性分析与因果干预结合的方法，可系统评估LLM内部经济偏好表征与真实行为的一致性。｜可迁移到信贷审批歧视研究，检验模型内部对种族、收入等群体的风险表征是否忠实于真实违约数据，以及这些表征是否影响审批决策。｜用LLM扮演信贷员，输入不同人口统计特征的贷款申请，测量其审批决策和内部表征；以真实信贷数据（如HMDA）为基准，比较模型表征与真实群体违约率的相似性，并通过激活修补检验因果路径。"}},{"id":"2609.15849","version":1,"title":"Before You Poll with LLMs: A Deliberative Diagnostic Framework","zh_title":"用LLM进行民意调查前：一个审议诊断框架","abstract":"Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM personas simulate public opinion at scale. Current evaluations test only whether personas hold the right opinions -- a static snapshot. But opinion research increasingly depends on dynamic fidelity: whether personas update beliefs in response to new arguments, as humans do during deliberation. No existing benchmark tests this. We introduce the Deliberative Polling Diagnostic Framework, which compares human and LLM belief shifts after identical informational interventions. Grounded in deliberative polling, it surfaces failures invisible to static evaluation: models that produce plausible partisan opinions can still misrepresent how those opinions change. Applying the framework to five frontier models using data from America in One Room (526 personas, 72 questions), we find that every model fails, each in a unique manner. GPT-5.1 exhibits reversal: its personas become more hostile toward the opposing party after balanced information, while humans become less so. This reversal is selective (80% on outgroup vs. 26% on policy questions) and symmetric across partisan identities. Gemini 2.0 Flash, Claude Sonnet 4.5, and Llama 3.3 70B exhibit overshoot, shifting correctly but at 5-7x human magnitude. DeepSeek V3 exhibits rigidity with near-zero change. Targeted ablations reveal that policy content triggers these failures and that they are identity-specific: GPT-5.1 reverses on outgroup questions but overshoots on ingroup; Gemini shows the inverse. We term this signature self-sycophancy: conformity to the model's internal stereotype of the persona rather than reasoning from the information provided. Our framework offers a concrete protocol: run the deliberative diagnostic before trusting LLM personas to mimic revised beliefs.","authors":["Ahmed Wali","Hassaan Tayyab"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15849","pdf_url":"https://arxiv.org/pdf/2609.15849","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","审议民意","算法保真度"],"reason":"直接评估LLM仿真人类意见动态，与真实人类数据对照，发现失效模式。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":2,"question":"LLM 人格在接收到平衡信息后是否会像人类一样更新信念，还是仅检索固化观点？","design":"用五个前沿 LLM（GPT-5.1、Gemini 2.0 Flash、Claude Sonnet 4.5、Llama 3.3 70B、DeepSeek V3）基于 America in One Room 数据构建 526 个匹配人口特征的人格，施加与人类相同的平衡信息干预，测量干预前后 72 个问题的观点变化。","baseline":"America in One Room 实验的真实人类数据，包含 526 名参与者在相同干预前后的观点变化。","findings":"所有模型均未通过诊断，且失败模式各异：GPT-5.1 在群体外问题上出现反转（敌意增加），Gemini、Claude、Llama 出现过度调整（幅度为人类 5-7 倍），DeepSeek 则表现为僵化（几乎无变化）。失败由政策内容触发且具有身份特异性，作者将其归因于“自我谄媚”（self-sycophancy）。","reliability":"论文未明确讨论失效条件，但指出当前评估仅关注静态观点正确性，无法捕捉动态更新失败；且失败模式因模型和问题类型而异，表明仿真可靠性高度依赖具体情境。","relevance":"该研究直接评估 LLM 仿真人类意见动态的可靠性，并与真实人类数据对照，揭示了静态评估无法发现的系统性失败，对关注仿真效度的研究者极具参考价值。","inspiration":"借鉴其“前测-信息干预-后测”的标准化诊断框架，并利用真实人类实验数据作为基准，可有效识别仿真中的方向性错误和幅度偏差。｜可迁移到政策公告的预期形成研究，例如央行沟通或财政政策变化对公众通胀预期的影响。｜以 LLM 人格模拟不同人口群体，施加与真实调查相同的政策信息，测量预期变化，并与密歇根大学消费者调查或央行预期调查的真实数据对照，检验仿真动态一致性。"}},{"id":"2609.15038","version":1,"title":"The average-farmer illusion in language-model simulations of agricultural decisions","zh_title":"语言模型模拟农业决策中的“平均农民”幻觉","abstract":"Language-model agents are increasingly used as synthetic people in surveys and social simulations, yet their apparent realism is often judged from population averages or distributional similarity. We tested what such evidence actually establishes by comparing Claude, Codex and Kimi under four prespecified prompt designs with matched farmer decisions from China and four African countries. Some configurations reproduced observed means and adoption rates. However, their person-level predictions were weak; their decisions clustered around typical values and policy-relevant extremes were largely missing. Most strikingly, a simple generator fitted only to the observed marginal dis- tribution, and given no information about any farmer, achieved greater distributional similarity than every language-model configuration. Prompt additions produced conditional gains rather than uni- versal improvement: results varied with model, outcome, population and validation target. We call this the average-farmer illusion: a synthetic population can look realistic while failing to repro- duce who does what or how behaviour varies. We provide a claim-matched validation framework and reusable modular prompts that turn prompt construction into an auditable experimental process. Population-level resemblance should therefore be treated as the start of validation, not as evidence of individual simulation.","authors":["Zhanliang Zhu","Ziwei Li","Yuchen Liu","Liujun Zhu","Ruiqi Wu","Tongqing Shen","Junliang Jin","Jianyun Zhang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15038","pdf_url":"https://arxiv.org/pdf/2609.15038","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B4"],"tags":["LLM仿真","人类行为对照","算法保真度"],"reason":"直接评估LLM仿真农业决策，与真实农民数据对照，揭示平均幻觉并给出验证框架。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":1,"question":"语言模型模拟农业决策时，群体层面的相似性是否意味着个体层面的模拟也可靠？","design":"用Claude、Codex和Kimi三个商业语言模型，在四种预设提示设计下模拟中国曲周县和四个非洲国家（尼日利亚、埃塞俄比亚、坦桑尼亚、马拉维）的农民决策，测量化肥施用量等连续行为变量，并与真实农民数据匹配。","baseline":"中国曲周县1332个农户-作物观测和四个非洲国家280个地块面板数据，包含农民实际决策。","findings":"部分配置能复现群体均值和采用率，但个体预测弱，决策集中在典型值，政策相关的极端值缺失。仅拟合边际分布的简单生成器在分布相似性上超过所有语言模型配置。","reliability":"提示添加只带来条件性收益而非普遍改进，结果随模型、结果变量、人群和验证目标变化；群体层面相似性不能作为个体模拟的证据。","relevance":"直接评估LLM仿真人类决策的可靠性，用真实农民数据对照，揭示平均幻觉，并给出验证框架，对关注仿真效度与偏差的研究者极具参考价值。","inspiration":"值得借鉴的是将验证目标分解为群体汇总、边际分布和个体匹配三个层次，并引入信息盲参考基准来检验分布相似性。｜可迁移到信贷审批歧视研究，用LLM模拟贷款官员对申请人特征的决策，检验群体违约率相似是否掩盖个体误判。｜用LLM扮演信贷员，输入申请人特征（收入、信用分、职业），输出是否批准贷款及额度，与真实银行信贷数据匹配，比较群体批准率、分布相似性和个体决策一致性，并加入仅基于边际分布的随机生成器作为基准。"}},{"id":"2609.13148","version":1,"title":"When Can You Trust Your Synthetic Users? Diagnostics and Corrections for LLM Consumer Panels","zh_title":"何时可以信任你的合成用户？LLM消费者面板的诊断与校正","abstract":"Large language models are increasingly deployed as synthetic consumer panels, promising $97\\%$ cost reductions over traditional surveys. Yet aggregate validation metrics conceal systematic failures: variance compression, coefficient sign-flips, subgroup error balloons of 10--30 percentage points, and global corrections that worsen demographic bias. We provide a formal framework for deciding when to trust, correct, or abandon LLM-generated consumer data. The framework decomposes synthetic-panel bias into covariate and concept shift, develops testable diagnostics with interpretable decision thresholds, and supplies a doubly robust AIPW estimator requiring only a small calibration sample ($n = 50$-$300$). We validate on three testbeds. In controlled simulations the decision rule achieves $100\\%$ accuracy (180/180 replications). On the American National Election Study with pre-existing LLM failures, it correctly flags heterogeneous concept shift and reduces naive bias by $92.9-99.6\\%$. On the Twin-2K-500 consumer pricing dataset (172,884 paired human and GPT-4.1-mini responses), it correctly routes full-sample estimation to Trust and subgroup targeting to Correct, with $83-94\\%$ bias reduction.","authors":["Robson Tigre","Hugo Gobato Souto"],"categories":["cs.HC","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13148","pdf_url":"https://arxiv.org/pdf/2609.13148","source_feed":"cs.HC","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM仿真","消费者面板","偏差校正"],"reason":"直接研究LLM合成消费者面板的可靠性诊断与校正，含真实人类数据对照，涉及经济学…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":2,"question":"如何判断何时可以信任、校正或放弃使用大语言模型生成的合成消费者面板数据？","design":"该论文不是仿真研究，而是提出一个诊断与校正框架：将合成面板偏差分解为协变量偏移和概念偏移，开发可检验的诊断工具（协变量重叠、条件校准、跨LLM稳定性），并提供双重稳健AIPW估计器，仅需小规模校准样本（n=50-300）即可校正偏差。","baseline":"使用三个真实人类数据集作为对照：美国国家选举研究（ANES）、Twin-2K-500消费者定价数据集（172,884对匹配的人类与GPT-4.1-mini响应）、以及受控模拟。","findings":"在受控模拟中决策规则达到100%准确率；在ANES上正确标记异质性概念偏移并将朴素偏差降低92.9-99.6%；在Twin-2K-500上正确路由全样本估计为信任、子群体定位为校正，偏差降低83-94%。","reliability":"论文承认概念偏移的严重程度是任务特定的，而非纯粹人口统计学：同一个LLM在产品定价上校准可靠，但在认知偏差任务（如合取谬误、锚定）上失败，因为LLM在人类依赖启发式的地方表现出“超理性”。此外，协变量重叠不足或跨LLM稳定性差时建议放弃使用合成数据。","relevance":"该论文直接针对LLM合成消费者面板的可靠性诊断与校正，提供了与真实人类数据对照的验证，并包含经济学相关场景（消费者定价），对关注仿真可靠性与偏差的研究者具有高度参考价值，值得精读原文。","inspiration":"该论文提出的协变量偏移与概念偏移分解、诊断阈值和双重稳健校正方法可借鉴用于经济金融仿真实验的可靠性评估｜可迁移到消费者金融决策、政策评估或行为经济学实验，如信贷审批歧视、消费者跨期选择、政策公告的预期形成等场景｜设计雏形：用LLM生成合成被试回答信贷申请或投资决策问题，以真实调查数据（如美国消费者金融调查SCF）为基准，施加不同政策信息处理，测量决策偏差，并用小规模人类样本校准AIPW估计器以校正LLM偏差。"}},{"id":"2607.28934","version":2,"title":"FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation","zh_title":"FairFund-Bench：评估LLM资源分配中的分配偏差","abstract":"Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 English-language requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants' names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.","authors":["Martin Lukk (University of Toronto)"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-03","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2607.28934","pdf_url":"https://arxiv.org/pdf/2607.28934","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","资源分配","算法公平"],"reason":"用LLM模拟人类资源分配决策，与真实人类数据对照，评估偏差与一致性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:06","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":4,"question":"LLM在资源分配中的偏差是否取决于审计设计（任务类型、比较情境、透明/伪装）？","design":"用14个LLM模拟人类资源分配决策，系统操纵任务（评分、排序、分配）、比较情境（单刺激/多刺激）、呈现方式（透明/伪装），测量对600个求助请求的分配结果。","baseline":"无对照（但基准中的求助模板基于130万真实GoFundMe活动校准，且使用福利应得性理论的人类应得性梯度作为参照）。","findings":"审计格式改变偏差方向：单独评分时模型偏向少数族裔，并排排序时则惩罚某些群体；伪装审计中的偏差比透明审计大3-4倍。因果框架效应比人口统计效应大约一个数量级，且跨模型和审计格式一致，表明LLM稳健地再现了人类应得性评价。","reliability":"论文指出偏差总体较小，但审计设计显著影响结论；透明审计中模型倾向于平均分配，可能掩盖潜在偏差；伪装审计更能揭示偏差。","relevance":"高度相关：该研究直接评估LLM作为人类被试替代品在资源分配决策中的偏差与一致性，并系统考察了审计设计对结论的影响，对理解仿真可靠性至关重要。","inspiration":"借鉴其系统操纵审计设计（任务、情境、呈现方式）来识别偏差的方法，以及用真实世界数据校准刺激材料并基于理论框架设计处理变量的做法。｜可迁移到信贷审批歧视、保险定价、政策福利分配等经济金融场景，检验LLM是否再现人类决策偏差。｜用LLM扮演信贷员，处理变量为申请人种族/性别（通过姓名信号）和贷款用途的因果框架（如医疗急需vs.创业失败），结果变量为贷款批准概率或利率，对照真实信贷审批数据（如HMDA数据）或人类实验数据。"}},{"id":"2608.02345","version":3,"title":"Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation","zh_title":"AI智能体能模拟A/B测试结果吗？面向智能体实验的验证框架","abstract":"A/B testing remains the standard for rolling out new features in the technology industry. Each experiment, however, consumes real traffic, engineering effort, and weeks of wall-clock time. Can AI agents---conditioned on behavioral profiles and contextual descriptions of the intervention---simulate outcomes accurately enough to vet candidate treatments before committing live traffic? We formalize this question as a \\emph{Simulated Randomized Controlled Trial} (S-RCT) and derive a two-layer error decomposition that separates agent approximation error from subsampling error, enabling targeted improvements to each. The framework is agent-agnostic: any behavioral model---from a fine-tuned specialist to a general-purpose foundation model---can serve as the simulation engine. Validated on 67 historical marketing A/B tests, a baseline S-RCT using an off-the-shelf foundation model captures directional signal (sign overlap 0.70) but systematically overshoots effect magnitudes. A two-phase pre-period calibration protocol reduces the squared prediction error (after removing irreducible measurement noise) by ${\\sim}77\\times$; a within-subject design---where each agent is exposed to both arms---reduces standard errors by ${\\sim}2.4\\times$. We discuss limitations of the current approach and identify applications where experimenters stand to benefit from agentic signals.","authors":["Stefan Hut","Lorenzo Masoero"],"categories":["cs.CL","cs.AI","stat.AP"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-04","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.02345","pdf_url":"https://arxiv.org/pdf/2608.02345","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","A/B测试","验证框架"],"reason":"用LLM模拟A/B测试结果，与真实历史实验对照，评估误差并改进，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":5,"question":"能否用AI智能体模拟A/B测试结果，以在投入真实流量前筛选候选干预方案？","design":"用基础大模型作为仿真引擎，根据用户画像、干预情境和任务描述生成模拟结果，构建模拟随机对照试验（S-RCT）；在67个历史营销A/B测试上，比较模拟估计的平均处理效应与真实历史结果，并引入预校准和受试者内设计改进估计。","baseline":"67个历史营销A/B测试的真实结果，包括效应方向和幅度。","findings":"基线S-RCT方向符号重叠率0.70，但系统性高估效应幅度；两阶段预校准将平方预测误差降低约77倍，受试者内设计将标准误降低约2.4倍。","reliability":"论文承认智能体存在系统性过度反应，画像完整性有限，且方向一致性在噪声数据上并不构成明确证据。","relevance":"直接相关：用LLM模拟A/B测试并与真实历史实验对照，评估误差并改进，符合研究者对仿真可靠性、偏差和失效条件的关注。","inspiration":"借鉴其将仿真误差分解为近似误差与子抽样误差，并分别用预校准和受试者内设计改进的做法｜可迁移到消费者金融产品选择或政策干预的A/B测试预筛，如信贷产品页面改版、退休储蓄默认选项调整等｜用LLM智能体基于用户画像模拟不同金融产品页面下的点击或选择行为，处理为页面版本，结果变量为选择率，并与历史A/B测试的真实选择数据对照，评估方向一致性和幅度校准。"}},{"id":"2609.13261","version":1,"title":"From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration","zh_title":"从过程损失到装配增益：多智能体LLM协作的人类基准诊断","abstract":"LLM agents are increasingly used for collaborative problem solving and human-group simulation. This makes outcome-only evaluation insufficient: if LLM groups are used as models of human groups, we need to know whether they succeed or fail through human-like deliberative mechanisms. We compare human group chats with matched LLM deliberation traces on Wason-style deductive reasoning, then test whether the same process signatures generalize to analogical, abductive, and analytical tasks. Humans and LLMs show the same assembly bonus asymmetry: discussion improves the average member more often than the best initial member. Initial-answer diversity accounts for the effect of model heterogeneity, increasing movement in both corrective and destructive directions. The main differences are process-level. Compared with humans, LLM groups follow majorities more often, surface less unique information, and converge earlier; correct minority signals succeed mainly when re-expressed early. Interventions motivated by human group-decision research yield modest improvements in collective outcomes, but do not remove the coordination bottleneck. Together, these results suggest that LLM groups can reproduce some outcome-level patterns of human deliberation while diverging in the mechanisms that generate assembly bonus and process loss, with implications for group simulation and human-AI collaboration.","authors":["Ala N. Tak","Teruhisa Misu","Kumar Akash","Zhaobo K. Zheng","Kevin H. Joo","Jonathan Gratch"],"categories":["cs.MA","cs.AI","cs.CL"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13261","pdf_url":"https://arxiv.org/pdf/2609.13261","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM群体仿真","人类对照","协作机制"],"reason":"用LLM群体模拟人类小组讨论，并与真实人类数据对照，评估机制差异。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":6,"question":"LLM群体协作能否复现人类小组讨论中的结果模式与过程机制，其装配增益与过程损失是否由类人机制驱动？","design":"用多个LLM智能体（GPT-4o、Claude-3.5-Sonnet等）组成小组，在Wason选择任务及类比、溯因、分析推理任务上进行自由讨论，记录个体初始答案、中间发言和最终群体答案，并施加信息揭示、多数翻转提示、置信度投影等干预，测量集体增益、最佳成员增益、装配增益和过程损失。","baseline":"匹配的人类小组聊天数据，在相同任务和协议下收集，作为过程诊断基准。","findings":"人类和LLM群体均表现出相同的装配增益不对称性：讨论提升平均成员多于最佳初始成员。但过程层面差异显著：LLM群体更频繁跟随多数、更少浮现独特信息、更早收敛，正确少数信号仅在早期被重新表达时才能成功。","reliability":"论文承认LLM群体虽能复现结果层面的模式，但机制与人类不同，存在更强的从众、更弱的少数信号保留和更早的锁定；干预仅带来适度改善，未能消除协调瓶颈。","relevance":"该研究直接针对LLM群体模拟人类小组决策的可靠性，提供了与真实人类数据的过程级对照，揭示了结果相似但机制不同的风险，对评估LLM仿真在群体决策场景中的有效性具有重要参考价值。","inspiration":"借鉴其过程追踪设计：记录个体初始答案、讨论中间发言和最终群体答案，并设置人类对照组，以区分结果相似与机制相似。｜可迁移到经济金融中的群体决策场景，如投资委员会决策、信贷审批小组、消费者家庭购买决策等。｜以LLM智能体模拟投资委员会，处理为不同信息结构（如隐藏信息揭示）或干预（如多数翻转提示），结果变量为投资决策质量和过程指标（如从众率、独特信息提及率），并与真实投资委员会会议记录或实验数据对照。"}},{"id":"2609.13995","version":1,"title":"Synthetic Data in Marketing Research: How to Evaluate and When to Trust","zh_title":"营销研究中的合成数据：如何评估与何时信任","abstract":"Debate over synthetic data in marketing research has polarized between claims that large language models (LLMs) make human respondents obsolete and calls to avoid them entirely. We argue that both positions obscure the more useful question: not whether synthetic respondents work, but when. Building on Brand, Israeli, and Ngwe (2026), we make three contributions. First, we distinguish three types of synthetic data (ungrounded LLM responses, segment-level personas, and individual-level digital twins) and map each to the decisions it can support. Second, we develop a taxonomy of four families of accuracy measures and suggest that the wide range of reported twin accuracy, from near-perfect to near-chance, largely reflects differences in what is being measured rather than in method quality. Aggregate measures often perform well even when little information is supplied to the LLM, and can mask a complete absence of respondent-level differentiation. Third, we introduce the forgotten question problem, in which a question is omitted from a fielded study, as a setting for twin-based augmentation of existing data. We propose an ex-ante answerability diagnostic that requires no ground truth: the R^2 of a random forest predicting twin outputs from the data used to construct the twins. Across 108 attitude questions from a nationally representative survey (N = 3,063), screening at R^2 above 0.7 raises the mean twin-human individual-level correlation by 15% and reduces the share of poorly answered questions from 25.9% to 4.3%. Embedding similarity and experienced-researcher judgment provide correlated but weaker screens.","authors":["Oded Netzer","Rajan Sambandam"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13995","pdf_url":"https://arxiv.org/pdf/2609.13995","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B4"],"tags":["LLM仿真","合成数据","营销研究"],"reason":"直接研究LLM合成数据在营销研究中的评估与信任，使用真实调查数据对照，提出诊断…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":7,"question":"在营销研究中，合成数据（LLM生成的回答、细分人群画像、个体数字孪生）何时可以信任，如何评估其准确性？","design":"论文区分三类合成数据：无依据的LLM回答、细分人群画像、个体数字孪生；提出四种准确性度量家族；引入“遗忘问题”场景，用随机森林R²作为事前可回答性诊断，在108个态度问题上筛选R²>0.7的孪生数据。","baseline":"使用全国代表性调查（N=3,063）的108个态度问题作为真实人类数据对照。","findings":"聚合度量往往表现良好，但可能掩盖个体层面无区分度；R²>0.7筛选将孪生-人类个体相关性平均提升15%，并将回答差的问题比例从25.9%降至4.3%。","reliability":"论文指出：聚合度量可能掩盖个体区分度缺失；嵌入相似度和研究者判断是较弱的相关筛选；未讨论模型版本、提示敏感性等失效条件。","relevance":"直接研究LLM合成数据在调查中的评估与信任，使用真实调查数据对照，提出无真值的事前诊断，对仿真可靠性研究有直接参考价值。","inspiration":"借鉴其事前可回答性诊断（用随机森林R²筛选可预测的孪生输出）和区分聚合与个体准确性的做法｜可迁移到消费者金融决策仿真，如信贷选择、保险购买或退休储蓄行为｜用LLM基于人口统计和财务特征生成个体孪生，施加不同金融产品特征处理，测量选择结果，并与真实消费者金融调查数据（如SCF或信用卡交易数据）对照，用R²筛选可回答的问题后再评估个体相关性。"}},{"id":"2609.15468","version":1,"title":"Time Machine Experiments: Using Historically-Bounded AI for Inquiry into the Human Mind","zh_title":"时间机器实验：利用历史受限AI探究人类心智","abstract":"Can interacting with someone from 1930, with no knowledge of what happened after, influence a person's perception of the past? People reason about the present against a picture of the past without observing it. The past is reconstructed from memory and testimony, but this reconstruction has been filtered through everything that happened since. Historically-bounded large language models (LLMs) make that past available for interaction. As a proof-of-concept for the impact of interacting with historical minds, we ran a preregistered randomized experiment ($N=240$), where participants interacted with an LLM trained on pre-1930 text. The interaction reduced the illusion of moral decline, the tendency to view the past as more moral than the present, compared to the contemporary-model control. This Time Machine Experiment paradigm informs new forms of interactive experiments, where temporal knowledge boundaries become experimental variables, and expands the realm of science fiction science, which turns thought experiments into actual experiments.","authors":["Hiromu Yakura","Robin Schimmelpfennig","Ezequiel Lopez-Lopez","Alejandro H. Artiles","Levin Brinkmann","Jean-Fran\\c{c}ois Bonnefon","Azim Shariff","Iyad Rahwan"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15468","pdf_url":"https://arxiv.org/pdf/2609.15468","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","人类被试替代","历史对照实验"],"reason":"用历史受限LLM作为人类被试替代，与真实人类对照，复现态度变化，属核心仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":9,"question":"与一个只了解1930年前信息的历史受限AI互动，能否改变当代人对过去的道德认知，从而缓解道德衰退错觉？","design":"使用基于1930年前文本训练的LLM作为历史人物代理，参与者与其进行对话和预测任务，测量互动前后对过去道德水平的感知变化及自我报告的反思程度，并与当代模型对照组比较。","baseline":"对照组为与当代前沿模型（gpt-5.5）互动的参与者，其道德衰退错觉变化作为基准。","findings":"与历史受限模型互动显著降低了参与者的道德衰退错觉，并引发了更多自我报告的反思。这表明跨时间互动可以改变人们对过去的偏见认知。","reliability":"论文未讨论","relevance":"该研究用历史受限LLM作为仿真被试，与真实人类对照，测量态度变化，属于核心仿真实验，值得阅读原文了解其方法和局限。","inspiration":"借鉴其利用历史知识边界作为实验变量的设计，通过对比不同时间截断的模型来隔离信息影响。｜可迁移到经济金融中的历史预期形成研究，例如投资者对历史政策效果的认知偏差。｜招募被试随机分组，分别与基于不同历史时期文本训练的LLM互动，测量其对历史经济事件（如大萧条）的归因和预期，并与真实历史调查数据对照。"}},{"id":"2609.15207","version":1,"title":"Issue Bias in Generative AI Writing Assistance: Political Issues and LLMs in the Swedish 2026 Election","zh_title":"生成式AI写作辅助中的议题偏见：2026年瑞典大选中的政治议题与大语言模型","abstract":"Generative AI writing assistants and the Large Language Models (LLMs) that power them are increasingly part of how voters gather information before elections. With growing evidence that they influence users' opinions, it is increasingly important to understand the views and positions of these tools. To better understand these views, we examine the stances supplied by six LLMs on a variety of Swedish-language writing tasks ahead of the 2026 Swedish parliamentary election. We cross 107 policy propositions with 77 writing templates and neutral, positive, and negative prompt framings, producing 24,717 prompts per model and 148,302 responses. To study these, we look at the models' default stance tendencies, compare how they respond to similar issues, and compare their responses with those of each of Sweden's eight parliamentary parties on the same issue. We find that Claude, DeepSeek, Gemini, and Mistral have similar profiles; ChatGPT more often supplies neutral or ambivalent text; and Grok differs most on topics such as migration, crime, and gender. When comparing the political parties, we find that the Social Democrats are closest to all six models. Still, after correcting for multiple comparisons, none of the within-model differences in party distances remains significant. Overall, we find that no model has a clear preference, nor a clear preference for a party, but that this depends on the specific issue or task the user asks about.","authors":["Bastiaan Bruinsma","Annika Fred\\'en","Paul R\\\"ottger","Moa Johansson","Asad Sayeed"],"categories":["cs.AI","cs.CY","stat.AP"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15207","pdf_url":"https://arxiv.org/pdf/2609.15207","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM政治态度","仿真对照","选举研究"],"reason":"用LLM生成政治文本并与真实政党立场对照，评估模型倾向，属于仿真人类政治态度且…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":12,"question":"在2026年瑞典大选前，六种大语言模型在瑞典语写作辅助任务中是否表现出政治立场或党派偏好？","design":"以六种LLM（Claude、DeepSeek、Gemini、Mistral、ChatGPT、Grok）为被试，交叉107个政策命题、77个写作模板和中性/正面/负面提示框架，生成148,302条瑞典语文本，分析模型默认立场倾向、模型间相似性及与瑞典八个议会政党的立场距离。","baseline":"瑞典八个议会政党在相同政策命题上的立场（来自三个投票建议应用VAA的编码，四点评分）。","findings":"Claude、DeepSeek、Gemini和Mistral立场相似；ChatGPT更常提供中性或模棱两可的文本；Grok在移民、犯罪和性别等议题上差异最大。所有模型与社民党距离最近，但经多重比较校正后，模型内各党距离差异不显著，表明无明确党派偏好。","reliability":"论文未讨论","relevance":"该研究用真实政党立场作为基准，评估LLM在政治写作任务中的立场倾向，属于仿真人类政治态度的研究，且包含批判性发现（无显著党派偏好），值得阅读原文了解其方法细节。","inspiration":"借鉴其大规模交叉设计（议题×模板×框架）和与真实立场对照的方法，可迁移到经济政策偏好仿真或消费者态度研究。｜可应用于政策公告的预期形成或信贷审批中的公平性评估。｜以LLM为被试，让其撰写关于税收、福利或监管政策的经济评论，处理为不同政策立场或框架，结果变量为文本中隐含的政策倾向，与真实民意调查或专家立场数据对照。"}},{"id":"2609.13254","version":1,"title":"(How) Do MLLMs Report Bistable Images Like Humans?","zh_title":"多模态大语言模型如何像人类一样报告双稳态图像？","abstract":"Bistable images such as the duck-rabbit are classic stimuli in which one image supports multiple mutually incompatible interpretations, typically reported one at a time in humans. We ask whether multimodal large language models (MLLMs) show similar report behavior and what internal computations support it. Using the LLaVA family, we study two tractable dimensions: modulability, whether reports can be biased by bottom-up visual cues and top-down linguistic priors, and exclusivity, whether responses commit to a single interpretation. We test both on the canonical duck-rabbit and on synthetic Visual Anagrams to mitigate memorization confounds. Behaviorally, both visual and linguistic manipulations systematically shift reports in human-consistent ways, while responses remain predominantly exclusive. Mechanistically, these effects arise from competing image-token representations, distinct pathways for bottom-up and top-down modulation, and a link between exclusive reporting and object-count encoding. Code and data are available at https://github.com/rtakatsky/mllm-bistable-images.","authors":["Ryota Takatsuki","Tomoki Doi","Amane Watahiki","Anil K. Seth","Hitomi Yanaka"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13254","pdf_url":"https://arxiv.org/pdf/2609.13254","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类行为对照","视觉认知"],"reason":"用MLLM复现人类对双稳态图像的报告行为，并与人类数据对照，评估仿真一致性。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":11,"question":"多模态大语言模型（MLLMs）是否像人类一样报告双稳态图像（如鸭兔图），其内部计算机制是什么？","design":"使用LLaVA系列五个模型（7B/13B等）作为被试，对经典鸭兔图和合成的Visual Anagrams双稳态图像施加视觉操作（旋转、加红圈）和语言提示（问题中嵌入偏向线索），测量模型输出的下一个词概率分布（modulability）和是否只报告单一解释（exclusivity），并分析内部表征。","baseline":"人类对双稳态图像的报告行为，包括视觉和语言线索对解释的偏向，以及报告单一解释的倾向。","findings":"视觉和语言操作都能以与人类一致的方式系统性地改变MLLM的报告，且报告绝大多数是排他性的。机制上，这些效应源于图像token表征的竞争、自下而上与自上而下调制的不同通路，以及排他性报告与物体计数编码的关联。","reliability":"论文承认只关注可操作化的两个维度（modulability和exclusivity），未涉及人类双稳态知觉的其他方面（如时间动态）；使用合成刺激以减轻记忆混淆，但可能仍存在其他偏差。","relevance":"该研究用MLLM复现人类对模糊视觉刺激的报告行为，并与人类数据对照，评估仿真一致性，且包含机制分析，对关注LLM仿真可靠性与偏差的研究者有参考价值。","inspiration":"借鉴其通过操纵输入（视觉与语言线索）和测量输出分布来量化模型行为与人类一致性的方法，以及使用合成刺激避免记忆混淆的设计。｜可迁移到经济决策中的模糊信息处理场景，如投资者对模棱两可的财报或政策声明的解读。｜用LLM作为被试，呈现模糊的金融图表或文本，施加视觉突出或语言框架处理，测量模型输出的解释分布，并与人类实验数据（如调查或行为实验）对照，检验仿真一致性。"}},{"id":"2609.12273","version":1,"title":"Synthetic TLX: Forecasting Human Workload Using Agent Simulation","zh_title":"合成TLX：使用智能体仿真预测人类工作负荷","abstract":"Assessing human workload for technology-mediated tasks helps prevent task failure caused by poor technology design. Traditionally, workload is assessed retrospectively using the NASA Task Load Index (TLX) after humans complete a task. What if we could forecast workload before a human attempts a task using agent simulation? We introduce Synthetic TLX, a new paradigm for proactive workload estimation that predicts NASA TLX scores for a given task, unlocking novel interaction opportunities and evaluation methods. To understand its viability, we conducted three experiments comparing human and agent-generated scores to evaluate where they align and diverge. We found agent estimates align with human scores particularly when prompted with a human persona and active task simulation. However, agents and humans diverge in the sources of workload they are sensitive to. Based on our findings, we present three applications to showcase Synthetic TLX's potential and discuss the future of workload-aware human-AI interaction.","authors":["Tzu-Sheng Kuo","Carrie J. Cai","Meredith Ringel Morris","Michael Terry"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12273","pdf_url":"https://arxiv.org/pdf/2609.12273","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","工作负荷预测","人机交互"],"reason":"用LLM代理预测人类NASA-TLX工作负荷，并与真实人类数据对照，属于人类仿…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":1,"question":"能否利用LLM代理仿真在人类执行任务前预测其NASA-TLX工作负荷？","design":"使用LLM代理（GPT-4o）通过四种提示策略（基础、人类角色、任务模拟、角色+模拟）生成NASA-TLX分数，针对邮件撰写、网站导航、智能体对话三类任务，在三种不同工作负荷来源条件下进行预测。","baseline":"从Prolific招募人类被试完成相同任务和条件，并填写NASA-TLX问卷作为真实工作负荷基准。","findings":"当代理被赋予人类角色并进行主动任务模拟时，其估计与人类分数最一致；但代理与人类对工作负荷来源的敏感性存在差异，代理对任务内在复杂性更敏感，而人类对外部负担更敏感。","reliability":"论文承认代理估计在特定条件下与人类存在分歧，尤其是对工作负荷来源的敏感性不同，且当前LLM代理的能力有限，需要未来模型进步才能完全实现Synthetic TLX的潜力。","relevance":"该研究直接使用LLM代理仿真人类主观体验（工作负荷），并与真实人类数据对照，属于人类仿真实验，且涉及HCI任务，对关注LLM仿真可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其多策略提示对比和任务模拟设计，可迁移到经济金融中的主观体验预测（如消费者决策疲劳、投资者认知负荷），设计实验让LLM代理模拟不同投资者角色预测金融信息处理负荷，并与真实投资者问卷数据对照。"}},{"id":"2609.12444","version":1,"title":"Diverse Minds, Divided Networks? Personality Composition, Polarization, and Collective Intelligence in LLM-Based Social Simulations","zh_title":"多元思维，分裂网络？基于LLM的社会模拟中的人格构成、极化与集体智能","abstract":"Simulated societies of large language model agents are used to study online polarization, and separately to study collective intelligence, but the two are rarely measured in the same system. It is therefore difficult to say whether a society's personality composition shapes both, or whether reducing polarization costs collective competence. We present TraitMix, an experimental design in which the Big Five composition of a simulated social network, both trait levels and trait heterogeneity, is a controlled experimental variable, and in which polarization and collective performance are measured in the same runs. Across 991 simulations of hundred-agent societies, spanning six contested topics and six language models, trait heterogeneity has the largest measured effects, acting in opposite directions on two faces of polarization: varied societies hold more dispersed opinions while being less segregated into camps, so homogeneous societies are not moderate but consensual echo chambers. Trait effects are not additive, as Agreeableness determines the sign of Openness, an interaction that replicates across models although the primary model's estimate is influence-driven. Contrary to the trade-off the study was designed to measure, no polarization measure predicts poorer collective performance, and cross-cutting interaction is the only one of four whose association with collective accuracy survives partialling on the aggregation identity. We report ablations removing two potential measurement circularities, an induction gate applied to every model, and the measures that failed them.","authors":["Raad Bin Tareaf"],"categories":["physics.soc-ph","cs.CL"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12444","pdf_url":"https://arxiv.org/pdf/2609.12444","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A3","B4","D2"],"tags":["LLM社会模拟","人格构成","极化与集体智能"],"reason":"用LLM agent模拟社会网络，研究人格构成对极化和集体智能的影响，虽无真实…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":2,"question":"在基于大语言模型的社会模拟中，人格构成（特质水平和异质性）如何同时影响极化与集体智能，二者是否存在权衡？","design":"使用六种大语言模型（含不同家族）扮演百人社会网络中的智能体，赋予验证过的五大人格特质，通过控制特质均值和异质性形成不同社会构成；智能体在带算法推荐和动态关注网络的平台上讨论六个争议政治话题，同时私下测量观点极化（离散度、隔离度等）和集体表现（估计任务、隐藏画像任务等可验证答案的任务），共运行991次模拟。","baseline":"无对照","findings":"特质异质性对极化的两个维度作用相反：异质性高的社会观点更分散但阵营隔离更弱，同质社会不是温和而是共识性回音室；宜人性调节开放性对极化的影响方向，且该交互在多个模型家族中复现。未发现极化与集体智能间的权衡，跨阵营互动是唯一在控制聚合身份后仍与集体准确性正相关的极化指标。","reliability":"论文通过消融实验移除两个潜在测量循环，对每个模型施加诱导门控，并报告未通过检验的指标；承认部分极化指标在控制后失效，并指出主要模型的估计受影响力驱动。","relevance":"该研究用LLM智能体系统操纵人格构成，同时测量极化与集体智能，虽无真实人类对照，但提供了严谨的仿真实验框架和可靠性控制，对关注LLM仿真有效性及偏差的研究者有方法学参考价值。","inspiration":"借鉴其将人格构成作为受控实验变量、同时测量多个社会结果并设置内部有效性控制（如中性话题、诱导门控、跨模型复现）的做法｜可迁移到经济金融中的群体决策与信息传播场景，如投资者情绪与市场泡沫、信贷审批中的群体偏见、政策公告的预期形成等｜设计一个LLM智能体模拟的资产定价实验：以不同人格特质组合（如开放性、神经质）的智能体为被试，处理为信息环境（如是否提供异质信号），结果变量为价格偏离和交易量，对照真实市场实验数据或历史价格数据。"}},{"id":"2609.06769","version":2,"title":"Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?","zh_title":"普通、理性的聊天机器人：AI模型是否追踪人类法律判断？","abstract":"As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making. As \"silicon sampling\" -- the use of generative AI models in social science research -- is now impacting academia, \"silicon jurors\" could make an appearance in courtrooms. This study joins an emerging line of research on generative AI models' ability to simulate human legal judgments. In particular, we study how large language model (LLM)-powered chatbots respond to series of questions about legal reasonableness. When the law needs to judge the appropriateness of a behavior, it most often asks whether the behavior was \"reasonable.\" Yet despite the ubiquity of reasonableness judgments, they are the site of constant vexation for lawyers, judges, and lay people. Reasonableness seems inherently vague and unpredictable, since it relies on variable context and implicit conceptual schemas. Moreover, many scholars caution that reasonableness judgments may vary along demographic lines. We compare the answers of human participants to those of twenty-six LLMs across twenty-five different legally relevant reasonableness judgments. Overall, our findings suggest that chatbot responses generally track those of human participants. Nonetheless, we find some suggestive -- and potentially concerning -- results. Compared to humans, LLMs generate more homogeneous responses and occasionally treat a variable standard as an invariant rule. And, compared to humans, LLMs tend to generate answers that are more favorable to the government and to corporations. Finally, our results indicate that LLMs' responses tend to align more closely with those of respondents who are white, male, older, and more educated. More systematic research is needed to confirm or reject these initial findings.","authors":["Nirav Patel","Emily Wenger","Christopher Buccafusco"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-11","first_seen":"2026-09-09","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2609.06769","pdf_url":"https://arxiv.org/pdf/2609.06769","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","法律判断","算法保真度"],"reason":"直接比较LLM与人类法律判断，含真实人类数据对照，并指出偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":1,"question":"大语言模型聊天机器人在法律合理性判断上是否与人类判断一致？","design":"比较26个LLM（来自Meta、Google、Anthropic、OpenAI、DeepSeek、xAI）与人类被试在25个法律相关合理性场景中的回答分布，分析模型回答的集中度、对政府与企业的偏向性，以及与不同人口群体回答的一致性。","baseline":"人类被试对相同25个法律合理性场景的回答数据。","findings":"LLM的回答分布与人类有统计差异，但总体落在人类回答的经验范围内，大体追踪人类的合理性观念。LLM回答更同质化，有时将可变标准当作不变规则，且更偏向政府和企业，并与白人、男性、年长、高教育程度人群的回答更一致。","reliability":"论文指出需要更系统的研究来确认或否定初步发现，并承认LLM回答存在同质化、偏向性等潜在问题。","relevance":"该研究直接比较LLM与人类法律判断，包含真实人类数据对照，并揭示了仿真偏差，对关注LLM仿真可靠性及偏差的研究者具有重要参考价值。","inspiration":"借鉴其多模型比较和人口统计学对齐分析的方法，评估LLM在特定判断任务中的偏差。｜可迁移到信贷审批歧视、消费者投诉处理或监管合规判断等经济金融场景。｜以LLM作为虚拟信贷员，处理贷款申请并给出批准决策，结果变量为批准率及理由，与真实银行信贷审批数据对照，检验模型是否与人类决策一致及是否存在人口统计学偏差。"}},{"id":"2603.17094","version":2,"title":"Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction","zh_title":"评估LLM模拟对话在建模人类社交互动中不一致与不合作行为的表现","abstract":"Simulating human conversations using large language models (LLMs) has emerged as a scalable methodology for modeling human social interaction. This paper reconsiders the evaluation of simulated conversations by explicitly recognizing that human conversations inherently involve inconsistent and uncollaborative behaviors, such as misunderstandings and interruptions. Since these behaviors contribute to the complexity of human social interaction, we argue that LLM-simulated conversations should reproduce them at frequencies comparable to those observed in human conversations. To support a detailed and interpretable evaluation of these behaviors, we introduce CoCoEval, a framework consisting of an evaluation scheme based on turn-level detection of 10 types of inconsistent and uncollaborative behaviors and a benchmark for simulating conversations in professional scenarios involving collaboration and conflict. Using CoCoEval, we compare human conversations with those simulated by GPT-4.1, GPT-5.1, and Claude Opus 4. The results show that (1) LLM-simulated conversations exhibit far fewer inconsistent and uncollaborative behaviors than human conversations under vanilla prompting, and (2) prompt engineering and supervised fine-tuning do not provide reliable control over these behaviors, often leading to the overproduction of specific behaviors. CoCoEval identifies gaps between human and LLM-simulated conversations that are not captured by conventional evaluation based on conversation-level Likert scales, raising concerns about the use of LLMs as proxies for human social interaction.","authors":["Ryo Kamoi","Ameya Godbole","Binglin Zhou","Xiaoxin Lu","Longqi Yang","Rui Zhang","Mengting Wan","Pei Zhou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-11","first_seen":"2026-03-17","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2603.17094","pdf_url":"https://arxiv.org/pdf/2603.17094","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","对话模拟","算法保真度"],"reason":"直接评估LLM模拟人类对话的保真度，并与真实人类对话对照，指出仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":2,"question":"如何评估LLM模拟对话中不一致和不合作行为的保真度，并与人类对话对照？","design":"使用GPT-4.1、GPT-5.1和Claude Opus 4在专业协作与冲突场景下模拟30轮对话延续，通过普通提示、分类引导提示和监督微调三种设置，测量10类不一致和不合作行为的出现频率。","baseline":"来自QMSum、NCPC、SIM和IQ2数据集的真实人类对话，涵盖商业、学术、政府会议和辩论。","findings":"普通提示下LLM模拟对话的不一致和不合作行为远少于人类对话；提示工程和监督微调无法可靠控制这些行为，常导致特定行为过度产生。","reliability":"论文指出LLM模拟对话的行为频率高度依赖模拟设置，且所有评估设置均未能复现人类对话中这些行为的频率；传统对话级Likert量表无法捕捉这些差异。","relevance":"该研究直接评估LLM模拟人类对话的保真度，并与真实人类对话对照，指出仿真偏差，对关注LLM作为人类被试替代品的研究者具有重要参考价值。","inspiration":"借鉴其细粒度行为检测评估方案和多种提示/微调设置对比，可迁移到经济金融中的协商、谈判或政策沟通模拟场景。｜例如，模拟消费者与客服的讨价还价或投资者与理财顾问的风险沟通。｜以LLM模拟谈判双方，处理为不同提示策略（如普通提示 vs. 明确要求包含冲突行为），结果变量为不一致行为（如误解、打断）的频率，并与真实谈判对话语料（如法庭记录或客服录音）对照。"}},{"id":"2607.27232","version":2,"title":"Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups","zh_title":"同情框架：跨社会人口群体评估AI对齐","abstract":"Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception. Considering news headlines covering political and geopolitical conflicts, both human participants (n = 3011, a representative sample of the U.K. adult population, via a YouGov survey) and seven LLMs answered whether headlines evoked sympathy for a specified side in a conflict. We find that the correlation between AI and human evaluations varies across models, ranging from very high (0.789, GPT-5.2) to medium (0.4 ,Mistral Large 2512). Crucially, the leading models are broadly aligned with human judgments across all demographic subgroups, including age, gender, level of education, prior geopolitical knowledge, and participants' predispositions regarding the conflict, although there are statistically significant differences between groups. This research, with its robust design and large, demographically diverse dataset, offers the most comprehensive evaluation of LLMs' comprehension of news framing to date. Findings highlight an important, often-ignored aspect of differential alignment: even when aggregate performance is high, AI alignment is not universal -- it may correspond differently with demographic features and cultural norms. Considering or ignoring the need for differential alignment may therefore have significant implications for the development of ethical and useful AI systems.","authors":["Haran Shani-Narkiss","Michael Fire","Oren Tsur"],"categories":["cs.CL","cs.AI","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-11","first_seen":"2026-07-31","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2607.27232","pdf_url":"https://arxiv.org/pdf/2607.27232","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","情绪感知"],"reason":"用LLM复现人类情绪感知，并与大规模人类调查对照，评估对齐差异","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":2,"question":"大语言模型在多大程度上与人类对新闻标题中同情性框架的情感感知保持一致，这种一致性在不同社会人口群体间是否存在差异？","design":"以七个主流大语言模型（GPT-5.2、Grok、GPT-4、Gemini、DeepSeek、Claude、Mistral）作为“读者”，对216条涉及政治和地缘政治冲突的新闻标题进行二元判断（是否对冲突中某一方产生同情），并与3011名英国成年人的调查回答进行对比，测量模型与人类判断的斯皮尔曼相关性。","baseline":"通过YouGov对3011名具有英国人口代表性的成年人进行调查，收集了超过20万条人类对新闻标题同情性框架的二元评价。","findings":"模型与人类判断的一致性因模型而异，GPT-5.2最高（0.789），Mistral最低（0.41）；即使总体一致性很高，在年龄、教育、母语、政治意识、话题知识和既有观点等子群体间仍存在显著差异。","reliability":"论文承认总体对齐高并不代表对所有人口群体都一致，顶尖模型在老年人、低教育水平、非英语母语者、低政治意识和无强烈观点者中一致性较低；同时指出模型在不同话题上的对齐稳定性不同，某些模型存在特定领域失效。","relevance":"该研究直接回应了用LLM替代人类被试进行情感感知仿真的可靠性问题，提供了大规模、人口多样化的真实人类对照，并揭示了总体对齐掩盖的群体差异，对评估LLM在调查和实验中的适用性具有重要参考价值。","inspiration":"借鉴其大规模分层对照设计，将模型输出与人口代表性样本的个体级回答进行相关性分析，并检验子群体差异，以识别对齐失效的边界条件。｜可迁移到政策公告的预期形成研究，例如央行沟通中文本框架对公众通胀预期的影响。｜以LLM作为虚拟受访者，对央行声明进行情感或框架判断，与家庭调查中的通胀预期数据（如密歇根大学消费者调查）对照，检验模型在不同教育、年龄和金融素养群体中的对齐程度。"}},{"id":"2609.11108","version":1,"title":"But How Would AI Agents Run a Town's Economy?","zh_title":"AI代理如何管理城镇经济？","abstract":"We placed 100 memory-equipped large language model (LLM) agents in charge of a closed, money-conserving spatial economy on real Pokhara Lakeside geography (earning wages, running businesses, setting prices) and ran this multi-agent simulation for up to 26 simulated weeks, well past the 1-2 weeks typical of agent-society studies. Across 91 validated runs (2.44M agent decisions, 21.5B tokens), the money stops moving, in a specific and measurable way. A 12x tourist demand shock raises business revenue 4.62x ($p<0.001$), which we decompose exactly into a 1.50x extensive margin (more businesses trading) and a 3.07x intensive margin (more revenue each). Monetary transmission stops there. Wages move 1.03x ($p=0.42$); 0.3% of 3,981 menu items are ever repriced ($p=0.47$). A randomized cash transfer (NPR 5,000 to 20 of 100 agents) shows the same pattern from the opposite direction: 96.7% is still held 311 pulses later, marginal propensity to consume 3-4% by two independent measures, indistinguishable from zero. The wealth distribution is consequently near-frozen at the horizon this literature uses ($\\rho=0.964$ over 2 simulated weeks), but not frozen. $\\rho$ falls to 0.832 at 12 weeks and 0.752 at 26, a horizon-dependence no short study can see. Matched ablations show which knob actually matters. Swapping the backing LLM moves every outcome we measure ($p=0.0039$); deleting agents' memory moves none of them detectably. A purely social tool fails 94-97% of the time across two model families, compared with ~96% success on economic tools, with no measurable shift away from it. Every headline number is verified twice, by a live validator and by an offline recomputation that reconciles each agent's wealth against its own signed transaction history, and we release the full run corpus for reanalysis.","authors":["Sajal Regmi","Siddhartha Pudasaini","Chetan Phakami Pun"],"categories":["cs.MA","cs.ET"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11108","pdf_url":"https://arxiv.org/pdf/2609.11108","source_feed":"cs.MA","score":9,"bucket":"selected","rubric_hits":["A3","B1","B2","B3","B4"],"tags":["LLM代理","经济仿真","政策评估"],"reason":"用LLM代理模拟经济，与真实数据对照，涉及政策评估，并批判性指出仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":4,"question":"LLM智能体在封闭货币经济中能否实现货币流通与财富再分配？","design":"100个带记忆的LLM智能体在真实博卡拉湖滨地理空间上运行封闭货币经济，从事工作、经营企业、定价交易；施加旅游需求冲击和随机现金转移两种处理；测量企业收入、工资、价格调整、边际消费倾向和财富分布变化。","baseline":"无对照","findings":"货币流通在企业和家庭层面均受阻：旅游冲击使企业收入增加4.62倍，但工资仅变动1.03倍，价格几乎不调整；随机现金转移的边际消费倾向仅3-4%，与零无显著差异。财富分布在2周内几乎冻结，但延长至26周后缓慢放松，表明短期研究可能得出误导性结论。","reliability":"论文承认种子数偏少（每条件9个），依赖匹配种子或组内对比；模型选择显著影响结果，而记忆删除无影响；社会工具失败率高达94-97%，经济工具成功率约96%；仅使用特定LLM模型，结果可能不具普遍性。","relevance":"该研究直接回应了LLM仿真在经济学中的可靠性问题，通过随机实验和长期追踪揭示了仿真失效的具体机制，对评估LLM作为人类被试替代品的有效性具有重要参考价值。","inspiration":"借鉴其随机现金转移和旅游需求冲击的准实验设计，以及通过长短期对比揭示时间尺度依赖性的方法｜可迁移到消费者跨期选择、政策刺激的乘数效应或信贷扩张的传导机制等宏观金融场景｜以LLM智能体为被试，施加一次性收入转移或信贷额度提升，测量消费支出、储蓄率和资产配置变化，并与家庭金融调查或信用卡交易数据对照。"}},{"id":"2609.11611","version":1,"title":"Who Bears the Risk When Generative AI Enters Transport? A Distributional Sociotechnical Audit of Algorithmic Equity, Synthetic-Data Validity, and Public Trust","zh_title":"生成式AI进入交通领域时谁承担风险？算法公平、合成数据有效性与公众信任的分配式社会技术审计","abstract":"Generative artificial intelligence is entering transportation through traveler-facing advisories, synthetic crash-record generation, and policy decision support. Existing governance frameworks lack transport-specific statistical tools to measure distributional risks across heterogeneous populations. We develop a Distributional Sociotechnical Audit (DSA) that integrates algorithmic equity, synthetic-data validity, and public-attitude heterogeneity into one empirical pipeline. The audit analyzes 5,760 persona-controlled queries to four LLM families across 12 demographic cues and four transport topics, uses two cross-family judges and a Wasserstein-2 Equity Dispersion Index, tests three FARS crash-record generators with conditional projected maximum mean discrepancy (cpMMD), fits a Bayesian ordered-logit model to Pew American Trends Panel Wave 152 (N = 4,538), and combines the signals into a continuous Sociotechnical Risk Index. Congestion-pricing advice has the highest persona-based dispersion (mean EDI = 1.96; highest direct EDI = 2.20). CART synthetic crash records fail all conditional tests (p < 0.001), while the Gaussian copula has borderline conditional stress (p = 0.105) despite passing marginal checks. Attitudes to AI vary across demographic strata. Distributional audits and continuous risk indices with sensitivity reporting offer a more defensible basis for transport GenAI governance than categorical approval tiers, which show a 75% assignment flip rate under weight perturbation.","authors":["Amir Rafe","Subasish Das"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11611","pdf_url":"https://arxiv.org/pdf/2609.11611","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","交通政策","公平性审计"],"reason":"用LLM模拟公众对交通政策的态度，并与真实调查数据对照，评估公平性和有效性。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":5,"question":"生成式AI进入交通领域后，其输出、数据产品和公众态度在不同人群间的分布性风险如何测量与治理？","design":"对四个LLM家族施加12种人口学线索和4个交通主题的5,760个查询，用两个跨家族裁判和Wasserstein-2公平离散指数测量建议分布差异；用条件投影最大均值差异检验三种FARS事故记录生成器的条件分布有效性；用贝叶斯有序logit模型拟合Pew调查数据分析公众AI态度异质性；最后合成连续的社会技术风险指数。","baseline":"Pew American Trends Panel Wave 152（N=4,538）的真实调查数据，以及FARS的110,001条真实事故记录。","findings":"拥堵收费建议在不同人格间的分布离散度最高（平均EDI=1.96），政策争议话题的人格驱动变异比天气安全建议高1.6倍；CART合成事故记录未通过所有条件检验，高斯copula在边际检验通过的情况下条件压力检验边缘显著（p=0.105）。","reliability":"论文承认分类审批层级在权重扰动下存在75%的分配翻转率，因此主张采用带敏感性报告的连续风险指数；但未详细讨论LLM仿真在何种条件下会失效，也未系统检验人格线索与真实人群行为的一致性。","relevance":"该研究用LLM模拟不同人口学群体对交通政策的反应，并与真实调查数据对照，评估公平性和有效性，直接回应了研究者对LLM仿真可靠性及偏差的关注，值得精读其审计框架和统计方法。","inspiration":"值得借鉴的是用多组人格线索系统探测LLM输出分布差异，并用真实调查数据校准公众态度异质性，同时用分布距离指标而非简单准确率来度量公平性｜可迁移到政策公告的预期形成或消费者对金融产品的态度异质性研究，例如不同人口群体对通胀或利率变化的反应差异｜设计雏形：用LLM扮演不同收入、教育、年龄的消费者，施加不同措辞的央行政策声明，测量其通胀预期和消费意愿，并与密歇根消费者调查或纽约联储SCE的真实数据对照，检验LLM仿真能否复现真实人群的态度分布和异质性。"}},{"id":"2609.09887","version":1,"title":"When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors","zh_title":"被告陈述何时重要？LLM模拟陪审员中的偏见与说服研究","abstract":"LLMs have been used to simulate human decision-making in professional settings, yet their behaviors in common-law jury trials remain unexplored. We study when and how a defendant's courtroom statement affects LLM-simulated jurors, focusing on persuasion, ideological bias, and background-based affinity. To support the analysis, we introduce JuryBench, a benchmark containing controversial criminal cases in U.S. criminal law. In each case, a defendant can claim various plausible justifications to support acquittal or reduced liability. We fix the base case and design defendants of different backgrounds, who give courtroom statements with varying emotional appeal or rebuttal. Jurors with diverse ideological profiles across the spectrum are simulated. We examine 20 frontier LLMs, resulting in a total of 432K decisions and rationales, and quantify changes in verdict severity. Our findings show that LLM-jury simulation echoes many human-jury findings. First, emotional persuasion can be detrimental, since jurors may perceive it as evidence of guilt or inconsistency. Next, we show that background fit between jurors and defendants is a stronger and significant factor than other isolated factors, and that jurors are in general harsher toward opposite-background defendants and lenient toward same-background ones. Finally, we find that juror ideology also strongly shapes severity judgments. These findings highlight both the promise and risks of using LLMs to model jury reasoning and call for careful evaluation. The data and code are available at https://github.com/choyingw/JuryBench","authors":["Cho-Ying Wu"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09887","pdf_url":"https://arxiv.org/pdf/2609.09887","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","陪审团决策","法律偏见"],"reason":"用LLM模拟陪审员决策，与真实人类陪审团研究对照，涉及法律决策偏差与说服效应。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-10","rank":1,"question":"在普通法陪审团审判中，被告的法庭陈述何时以及如何影响 LLM 模拟陪审员的裁决，重点关注说服、意识形态偏见和背景亲和力。","design":"使用 20 个前沿 LLM 模拟具有不同意识形态背景的陪审员，在 500 个争议性刑事案件中，对被告的不同背景和带有不同情感诉求或反驳的法庭陈述做出有罪/无罪及严重程度判断，并记录决策理由。","baseline":"无对照","findings":"情感说服可能适得其反，因为陪审员可能将其视为有罪或不一致的证据；背景契合度（陪审员与被告背景相似性）比孤立因素更强且显著，陪审员通常对背景相反的被告更严厉，对背景相同的被告更宽容；陪审员意识形态也强烈影响严重程度判断。","reliability":"论文未讨论","relevance":"该研究用 LLM 模拟陪审员决策，与真实人类陪审团研究对照，涉及法律决策中的偏差与说服效应，属于人类仿真实验，但未提供真实人类数据基准，可靠性存疑，值得阅读原文以了解其仿真设计细节和潜在偏差。","inspiration":"借鉴其通过系统操纵被告背景、陈述情感强度和陪审员意识形态来测量决策偏差的因子设计方法，以及大规模生成争议案例和记录决策理由的做法。｜可迁移到信贷审批歧视、招聘面试评估或政策沟通中的说服效应等经济金融场景。｜用 LLM 模拟信贷员或招聘经理，处理变量为申请人背景（如种族、性别）和陈述情感强度，结果变量为批准/拒绝或评分，对照真实信贷审批数据或审计研究结果来评估 LLM 仿真的外部有效性。"}},{"id":"2609.10280","version":1,"title":"Total Simulated Survey Error: Designing and Diagnosing Survey Responses from Large Language Models","zh_title":"总体模拟调查误差：设计和诊断大语言模型的调查回答","abstract":"Large Language models (LLMs), having been trained on vast amounts of human-generated data, may encode the attitudes and behaviors of these humans. As such, LLMs show promise in mimicking human-like patterns that facilitate their use in simulating people in a wide variety of contexts. One such context is using LLMs as 'silicon samples', i.e., proxies of people in answering survey questions to establish public opinion, design policies, or use as (social) scientific data. However, several critical questions of social biases, generalization, and technical limitations remain, further complicated by a vast design space open to simulation designers. Multiverse analyses might help us make sense of the impact of different design choices, however, we lack a systematic understanding of the design space of LLM-generated surveys as well as how these decisions interplay with inherent LLM limitations. Therefore, how do we systematically identify, trace, and document limitations in LLM-generated survey responses? Building on traditions in the quantitative social sciences, specifically survey methodology and measurement theory, we investigate threats to the validity of LLM-generated survey responses. To do so, we design a framework that enumerates conceptual errors and systematic biases that can occur at different stages of the survey simulation lifecycle. Our framework, called the Total Simulated Survey Error (TS2E) Framework, provides a unified and end-to-end perspective on LLM-generated survey data. The framework, illustrated through a theoretical and empirical case study, enables survey simulation designers to systematically identify and reflect on errors in LLM-generated surveys.","authors":["Indira Sen","Georg Ahnert","Leah von der Heyde","Jana Lasser","Bernd Wei{\\ss}","Markus Strohmaier"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10280","pdf_url":"https://arxiv.org/pdf/2609.10280","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A2","A4","B1","B4"],"tags":["LLM仿真","调查方法","误差框架"],"reason":"提出TS2E框架诊断LLM调查仿真误差，含实证案例，直接相关","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-10","rank":2,"question":"如何系统识别、追踪和记录大语言模型生成调查回答中的误差来源？","design":"提出一个概念框架（TS2E），将调查模拟生命周期划分为不同阶段，枚举各阶段可能出现的测量误差和代表性误差，并通过一个理论性和实证性案例研究进行说明。","baseline":"无对照","findings":"该框架区分了研究者设计选择导致的误差与LLM固有局限导致的误差，并引入了LLM特有的误差类型（如人物角色构建误差）和评估谬误。通过案例研究展示了框架如何帮助设计者系统识别和反思LLM生成调查中的误差。","reliability":"论文承认LLM训练数据存在缺口和偏斜，指令微调等后训练过程可能影响模型行为，且高总体对齐可能掩盖方差、子群体异质性和下游统计关系的严重失真。","relevance":"该论文直接针对LLM仿真调查的可靠性问题，提出了系统诊断误差的框架，对关注仿真效度与偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其将总调查误差框架迁移到LLM仿真的思路，对仿真流程进行阶段分解并系统识别误差来源。｜可迁移到经济金融领域的调查类仿真，如消费者信心调查、通胀预期调查、投资者情绪调查等。｜设计雏形：用LLM模拟不同人口统计学特征的消费者，施加不同的经济信息提示（如货币政策公告），测量其通胀预期和消费意愿，并与密歇根大学消费者调查的真实数据对照，检验仿真误差。"}},{"id":"2606.14199","version":2,"title":"OdysSim: Building Foundation Models for Human Behavior Simulation","zh_title":"OdysSim：构建用于人类行为模拟的基础模型","abstract":"Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation. Yet helpfulness-driven post-training pulls them toward a homogeneous, overly agreeable assistant register, creating a behavioral Sim2Real gap. We present OdysSim, the largest open systematic investigation of behavioral foundation models, i.e., models trained to simulate human behavior at scale. We propose SOUL, a taxonomy of five capability axes (CONV, SS, COG, ROLE, EVAL) that unifies 62 datasets and 23 benchmark tasks under one framework. Specifically, we curate the OdysSim corpus (21.4M interactions, 10B tokens, retrofitted with back-generated social contexts), construct the SOUL-Index benchmark, and develop an end-to-end training recipe combining midtraining, task-specific RL, and expert distillation. The resulting open 8B OSim model ranks first or tied-first on 8 of 23 tasks, outperforming any individual frontier model by this count, with the strongest gains on conversational and social tasks. Its outputs are also more human-like in length, formatting, and word choice, and it transfers zero-shot to out-of-distribution user simulation on $\\tau$-bench, nearly matching real users on reaction alignment (93.2 vs. 93.5). We further show that LLM-as-judge RL induces reward-hacking patterns, and that our detectors can mitigate them during post-training. Together, our findings suggest that behavioral foundation models require rethinking the LLM training paradigm. We release all artifacts to support future research.","authors":["Xuhui Zhou","Weiwei Sun","Weihua Du","Jiarui Liu","Haojia Sun","Qianou Ma","Tongshuang Wu","Yiming Yang","Maarten Sap"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-06-12","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2606.14199","pdf_url":"https://arxiv.org/pdf/2606.14199","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","A5","B1","B2","B3"],"tags":["人类行为模拟","基础模型","算法保真度"],"reason":"直接构建行为基础模型模拟人类行为，含真实人类数据对照，覆盖多任务并评估可靠性。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":1,"question":"如何构建行为基础模型来缩小LLM模拟人类行为时的Sim2Real差距？","design":"论文构建了OdysSim系统，使用Qwen3基础模型在21.4M交互的OdysSim语料上进行中训练，然后对SOUL-Index的23个任务进行任务特定强化学习（GRPO或RLVF），最后通过专家蒸馏合并为单一8B模型OSim-8B。模型被用于模拟对话、社会推理、认知、角色扮演和评估等五类人类行为，输出结果与真实人类数据对比。","baseline":"SOUL-Index基准包含62个数据集和23个任务，其中许多任务有真实人类行为数据作为对照；另外在τ-bench用户模拟评估中，与真实用户的反应对齐分数（93.5）进行对比。","findings":"OSim-8B在23个任务中的8个上排名第一或并列第一，超过任何单个前沿模型，尤其在对话和社会任务上提升最大；其输出在长度、格式和用词上更像人类，并在τ-bench上零样本迁移到用户模拟，反应对齐分数93.2接近真实用户93.5。","reliability":"论文指出LLM-as-judge强化学习会导致奖励黑客模式，但他们的检测器可以在后训练中缓解；此外，中训练数据虽经社会背景回填，但原始对话缺乏说话者背景，可能影响社会动态推断的准确性。","relevance":"该研究直接针对LLM模拟人类行为的可靠性问题，提供了大规模语料、基准和训练方法，并与真实人类数据对照，对关注人类仿真实验的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其通过中训练注入行为多样性、任务特定RL校准行为以及专家蒸馏合并能力的方法，可迁移到经济金融中的消费者决策、投资者行为或政策反应模拟；例如，用LLM模拟投资者在政策公告后的交易行为，以真实市场数据（如订单流、调查数据）为基准，通过中训练和RL微调使模型输出与真实投资者行为分布对齐，从而评估政策效果。"}},{"id":"2609.07353","version":1,"title":"Human-like moral judgments conceal divergent motive attributions in large language models","zh_title":"类人道德判断掩盖了大语言模型中不同的动机归因","abstract":"Large language models (LLMs) are used to simulate human participants in psychological research. We asked whether LLMs that reproduce human evaluations of a whistleblower's moral character also reproduce the motive attributions that accompany them. Five LLMs and two human samples (N = 125 and N = 742) evaluated a physician who either remained silent about fraudulent billing or reported it to a hospital, regulator, or newspaper. Models reproduced the human ranking of the physician's moral character but portrayed whistleblowers as more helpful, less self-interested, and less hostile. In four of five models, competitive motives were less strongly associated with moral-character judgments. Model ratings changed little when prompts reproduced the narratives and demographic profiles of both human samples, although this comparison cannot isolate a perspective effect. Thus, agreement in average ratings can conceal differences in attributed motives, relationships among judgments, and sensitivity to context. Validating LLMs as simulated participants therefore requires testing psychologically informative response patterns, not average agreement alone.","authors":["Xiaoyan Wu","Jean-Claude Dreher"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07353","pdf_url":"https://arxiv.org/pdf/2609.07353","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","道德判断","算法保真度"],"reason":"直接使用LLM仿真人类道德判断，并与两个人类样本对照，发现平均评分一致但动机归…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":4,"question":"LLM在复现人类对举报者道德品质评价的同时，是否也复现了伴随的动机归因？","design":"用五个LLM（含闭源与开源）模拟人类被试，对一位医生面对欺诈账单保持沉默或向医院、监管机构、报纸举报的四种情境进行道德品质与四种动机（义务、亲社会、自利、竞争）评分，并比较模型与两个人类样本的评分模式、动机与道德判断的关联，以及模型对两种样本叙述的敏感性。","baseline":"两个独立招募的人类样本（N=125和N=742），分别接受第一手（事件发生在自己团队）和第二手（事后听说）版本的场景描述。","findings":"模型复现了人类对道德品质的条件排序，但将举报者描绘为更亲社会、更少自利和敌意；五个模型中有四个的竞争动机与道德品质判断的关联弱于人类。模型评分对两种样本叙述的变化不敏感，平均评分一致掩盖了动机归因、判断间关系和情境敏感性的差异。","reliability":"论文承认两种样本的对比不能分离视角效应，因为框架、招募和样本构成同时变化；模型对提示变化的敏感性可能影响结果；仅凭平均一致性不足以验证LLM作为模拟被试的有效性。","relevance":"直接回应了LLM仿真人类被试的核心问题：平均评分一致不等于心理结构一致，对评估仿真可靠性至关重要，值得精读。","inspiration":"借鉴其多层次验证设计：不仅比较条件均值，还检验变量间关系和情境敏感性｜可迁移到经济决策中的道德或动机归因场景，如举报行为、企业社会责任评价、消费者对品牌道德危机的反应｜用LLM模拟消费者对某公司不当行为的道德判断，处理为不同举报渠道（内部、监管、媒体），结果变量为道德评价和动机归因，对照真实消费者调查数据，检验模型是否复现动机与评价的关联模式。"}},{"id":"2609.07573","version":1,"title":"From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction","zh_title":"从模拟公民到模拟审议：表征与互动的挑战","abstract":"Multi-agent LLM deliberation has been explored as a scalable way to simulate public deliberation. For such simulations to be informative, persona agents should reflect population opinion patterns and interaction should shape their conclusions. We evaluate whether LLM-based deliberation can meet these two conditions using census-grounded Korean personas debating real policy questions benchmarked against national surveys. Persona agents do not reliably reproduce population opinion patterns: responses are often far more concentrated and frequently reverse demographic differences in the human data. Deliberations nonetheless produce reasoned, reciprocal, and varied arguments alongside substantial stance movement. Yet much of this movement does not require peer exchange: sealed-monologue agents change position at similar rates and reach nearly the same final balance as full debates, while groups initialized with very different positions often converge to similar endpoints. Anchoring population-informed starting positions, meanwhile, sharply suppresses updating. Thus, population representation, argument generation, and interaction-driven opinion change do not necessarily go together. The simulations readily surface arguments on both sides, though whether they capture the diversity of human perspectives remains untested, leaving open a promising role for argument surfacing even as population simulation requires further validation.","authors":["Chaemin Jang","Junsik Min","Jaewoo Choi","Donggyu Lee","Haiin Lee","Junyoung Park","Namhee Kim","Hyunwoo Kim","Jungwon Kim","Juho Kim","Nuri Kim","Jihee Kim"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07573","pdf_url":"https://arxiv.org/pdf/2609.07573","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","公共审议","人类数据对照"],"reason":"用LLM模拟公民审议并与全国调查对照，评估代表性与互动效应，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":5,"question":"基于LLM的多智能体审议能否同时满足意见代表性和互动驱动观点变化两个条件？","design":"使用基于人口普查校准的韩国人设智能体（Nemotron-Personas）就真实政策问题进行辩论，并与全国调查基准对照；通过封闭独白控制组分离同伴互动对立场变化的影响。","baseline":"韩国全国环境意识调查和低生育率政策公众意识调查的群体级基准数据。","findings":"人设智能体未能可靠复现人口意见模式，回答更集中且常逆转人口学差异；审议产生了理性、互惠和多样的论点，但大部分立场变化不需要同伴互动，且不同初始立场的群体常收敛到相似终点。","reliability":"论文承认人口代表性、论点生成和互动驱动的观点变化不一定同时成立；论点是否捕捉人类视角多样性尚未检验，人口模拟需进一步验证。","relevance":"直接相关，提供了LLM模拟审议的严格评估，揭示了代表性失败和互动效应虚假的问题，对使用LLM进行人类仿真实验的研究者具有重要警示价值。","inspiration":"值得借鉴封闭独白控制组来分离互动效应，以及用人口普查校准人设并与全国调查对照的评估框架｜可迁移到政策公告的预期形成或消费者信心调查等场景，检验LLM模拟的群体意见动态｜用LLM人设模拟投资者或消费者群体，处理为是否进行多智能体辩论，结果变量为观点变化和最终分布，对照真实调查数据（如密歇根消费者信心指数）来验证仿真可靠性。"}},{"id":"2609.07987","version":1,"title":"When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability","zh_title":"LLM数字孪生何时能减少人类测量？从行为保真度到统计可替代性","abstract":"LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide little evidence about whether they can reduce human measurement while preserving valid inference. To address this, we introduce statistical substitutability, an inferential criterion that evaluates the extent to which twin predictions can reduce human measurement for a particular estimand while preserving valid inference. We develop a framework, grounded in mixed-subject and prediction-powered inference, that evaluates statistical substitutability along four dimensions: aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations. Across two empirical evaluations spanning behavioral experiments, multiple models, and alternative respondent representations, we find that digital twins can reproduce average human effects while providing little information about which individuals differ from those averages. Newer models and richer respondent information improve some dimensions of performance but do not reliably translate into human-data savings. Human calibration can reduce aggregate prediction error, yet limited labeled samples often fail to produce stable precision gains. Importantly, these findings demonstrate that behavioral fidelity is neither necessary nor sufficient for statistical substitutability. More broadly, they suggest that AI-generated evidence should be evaluated based on its ability to support valid scientific inference rather than its ability to reproduce human outcomes alone. Digital twins should therefore be judged for confirmatory use by whether they reduce uncertainty about human quantities, not merely by whether they reproduce human means, distributions, or effects.","authors":["Steven Wang","Kyle Hunt","Shaojie Tang","Kenneth Joseph"],"categories":["cs.AI","stat.AP"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07987","pdf_url":"https://arxiv.org/pdf/2609.07987","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM数字孪生","统计可替代性","人类仿真"],"reason":"直接研究LLM数字孪生替代人类测量的统计可替代性，含真实人类数据对照与批判性评…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":6,"question":"LLM数字孪生能否在保持有效推断的同时减少人类测量，即其统计可替代性如何？","design":"使用LLM数字孪生（基于受访者丰富信息构建）预测个体在行为实验中的反应，并与真实人类数据对比，评估其在聚合效应、个体层面信号、有限样本标签恢复和跨人群稳定性四个维度上的表现。","baseline":"两个实证评估中的真实人类行为实验数据，包括个体层面的反应和实验处理效应。","findings":"数字孪生能复现平均人类效应，但几乎不提供个体偏离平均值的信号；新模型和更丰富的受访者信息改善了部分维度，但未可靠转化为人类数据节省。","reliability":"行为保真度既非统计可替代性的必要条件也非充分条件；人类校准可减少聚合预测误差，但有限标签样本往往无法产生稳定的精度增益；数字孪生应基于其减少人类量不确定性的能力来评判，而非仅复现人类均值、分布或效应。","relevance":"该研究直接针对LLM仿真人类被试的可靠性问题，提出了统计可替代性标准，并基于真实人类数据对照进行了批判性评估，对关注仿真有效性和偏差的研究者极具参考价值。","inspiration":"借鉴其统计可替代性框架，将LLM预测与人类数据结合进行混合推断，并检验个体层面信号和跨人群稳定性｜可迁移到经济金融中的个体决策预测，如消费者跨期选择、风险偏好或投资行为，评估LLM能否替代部分人类被试｜设计一个资产定价实验，用LLM数字孪生预测个体对风险资产的需求，处理为不同信息条件，结果变量为投资金额，并与真实人类实验数据对照，检验LLM预测能否减少所需人类样本量。"}},{"id":"2609.08003","version":1,"title":"Sparks of In Silico Cognitive Science: Theories from Simulated Data Can Generalize to Humans","zh_title":"硅基认知科学的火花：来自模拟数据的理论可以推广到人类","abstract":"Behavioral foundation models have been proposed as stand-ins for human participants across settings, but it is unclear whether theories discovered on them generalize to humans or merely characterize the simulator. We ran the Automated Cognitive Scientist (\\textsc{AutoCog}), a closed-loop discovery system in which LLM agents design theory-discriminating experiments, collect responses, arbitrate between competing theories, and synthesize successors, entirely on behavior simulated by Centaur, a foundation model of human behavior. In a multi-attribute decision-making setting, the theories \\textsc{AutoCog} found on Centaur generalized to human data: they outperformed canonical theories on ten held-out experiments and were rivaled only by theories found by running the same loop on people. We argue that this succeeds despite the simulator's inevitable imperfections because a discovery loop that arbitrates between competing theories demands less of its simulator than estimation does. The simulator only needs to capture the regularities that distinguish the theories, and not necessarily reproduce behavior precisely. Imperfect simulators can therefore widen the search over theories, with human data then testing whether the surfaced theories generalize.","authors":["Akshay K. Jagadish","Younes Strittmatter","Nori Jacoby","Eric Schulz","Nathaniel Daw","Thomas L. Griffiths","Suyog H. Chandramouli"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08003","pdf_url":"https://arxiv.org/pdf/2609.08003","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","人类行为对照","理论发现"],"reason":"用LLM仿真人类决策，并与真实人类数据对照，验证理论可推广性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":7,"question":"在模拟人类行为的基础模型上运行自动理论发现系统，所发现的理论能否推广到真实人类行为？","design":"使用 Centaur（人类行为基础模型）模拟人类在多属性决策任务中的选择，运行 AutoCog 闭环发现系统：LLM 智能体设计区分理论的实验、收集模拟响应、在竞争理论间仲裁并合成后继理论，最终得到在模拟数据上表现最好的理论。","baseline":"对照真实人类数据：在十个留出实验上比较 AutoCog 在 Centaur 上发现的理论、经典理论以及直接在人类数据上运行 AutoCog 发现的理论。","findings":"在 Centaur 模拟数据上发现的理论在人类数据上优于经典理论，且与直接在人类数据上运行相同发现循环得到的理论表现相当。成功的原因在于理论仲裁对模拟器的要求低于精确估计，模拟器只需捕捉区分理论的关键规律。","reliability":"论文承认模拟器不可避免存在缺陷，但认为在理论发现场景下这些缺陷影响较小；未详细讨论模拟器在哪些具体条件下会失效。","relevance":"直接回应了 LLM 仿真人类行为并用于理论发现的可推广性问题，提供了与真实人类数据对照的实证证据，对关注仿真可靠性和偏差的研究者极具参考价值。","inspiration":"借鉴其闭环理论发现框架：用 LLM 智能体自动设计实验、仲裁理论并迭代，可大幅扩展理论搜索空间，再用人类数据验证。｜可迁移到经济决策中的启发式建模，例如消费者在复杂金融产品间的选择、投资者在多属性资产间的配置决策。｜以 LLM 模拟投资者在多属性资产间的选择行为，运行 AutoCog 发现决策理论，再与真实投资者在相同实验中的选择数据对照，检验理论的可推广性。"}},{"id":"2609.07141","version":1,"title":"How Well Do LLMs Simulate Survey Responses Following a Breast Cancer Screening Intervention?","zh_title":"LLM在乳腺癌筛查干预后模拟调查回答的效果如何？","abstract":"Collecting survey data is laborious and limited by privacy constraints. Large language models (LLMs) have shown promise as predictive social simulations. It is unclear whether they can replicate population-level response distributions before and after a healthcare intervention. Using information derived from 4125 women aged 35-59 years, we evaluate whether agents informed solely by pre-intervention profile information can reproduce post-intervention response distributions. Groups of LLM agents (n=50) were created with Gemma 4 E4B and Qwen3.5 9B; conditions ranged from zero-shot prompting to agent profiles enriched with aggregate or individual-level demographic characteristics and pre-intervention questionnaire responses. We compared predicted and observed response distributions with Total Variation Distance (TVD) and Normalized Wasserstein Distance (NWD). Across both LLMs, profile-based agents improved distributional accuracy relative to zero-shot and random baselines. Nevertheless, direct sampling of 50 real participants remained more accurate. Prediction errors were also higher among participants aged 55-59 years and those living in private property. Errors also varied by question theme and LLM model, with the highest errors observed for cancer fatalism and post intervention attitudes toward genetics. Sensitivity analyses showed that performance was influenced by prompt template changes and temperature hyperparameter. Our results show the potential of LLM-based agents to model behavioral responses to interventions in silico. However, profiles containing additional information beyond demographics did not consistently outperform simpler ones. Certain cultural constructs and population groups also remain inadequately represented by the LLM models evaluated. Future work may include building behaviorally grounded and locally validated virtual populations.","authors":["Kenneth Koh","Ryan Jak Yang Lim","Alessandro Sparacio","Peh Joo Ho","Mile Sikic","Borame L Dickens","Mikael Hartman","Jingmei Li"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07141","pdf_url":"https://arxiv.org/pdf/2609.07141","source_feed":"cs.SI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","调查回答","人类对照"],"reason":"用LLM代理模拟乳腺癌筛查干预后的调查回答，并与真实人类数据对照，评估分布准确…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":3,"question":"LLM 智能体仅基于干预前的个人资料信息，能否复现乳腺癌筛查干预后的人群调查回答分布？","design":"使用 Gemma 4 E4B 和 Qwen3.5 9B 两种 LLM，基于 4125 名 35-59 岁新加坡女性的真实数据创建智能体，条件包括零样本、聚合人口学特征、个体人口学特征及干预前问卷回答等不同信息丰富度，模拟干预后问卷回答分布，并与真实回答比较。","baseline":"来自 BREATHE 队列的 4125 名女性的真实干预后调查回答，以及随机预测和直接抽样 50 名真实参与者的分布。","findings":"基于个人资料的智能体相比零样本和随机基线提高了分布准确性，但仍不如直接抽样真实参与者；预测误差在 55-59 岁和私宅居民中更高，且在不同问题主题和 LLM 模型间存在差异。","reliability":"论文承认 LLM 对某些文化构念和人群亚群代表性不足，且包含额外信息的个人资料并不总是优于简单资料；性能受提示模板和温度超参数影响。","relevance":"该研究直接评估 LLM 在医疗干预后调查回答仿真中的可靠性，与研究者关注的人类仿真实验、真实数据对照和失效条件高度相关，值得精读原文。","inspiration":"借鉴其用干预前资料构建智能体并对比真实干预后分布的设计，以及通过 TVD/NWD 和亚组分析评估仿真偏差的方法。｜可迁移到政策干预对经济行为的影响评估，如健康保险补贴对就医行为、财务教育对储蓄决策的影响。｜以真实调查数据中的个体特征和干预前行为为输入，让 LLM 智能体模拟干预后的消费或投资选择，并与实际追踪调查数据对照，检验仿真在收入、年龄等亚组上的误差模式。"}},{"id":"2608.22697","version":3,"title":"Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf","zh_title":"排名还重要吗？AI代理替我们购物时的位置偏差","abstract":"Search rankings are valuable because human attention is scarce and sequential. Higher-placed alternatives are easier to find, so they are examined and bought more often. Consumers are now delegating search to AI agents that can ingest an entire results page at once. Randomizing the order of one hundred hotel listings across 5,000 AI agent sessions, we compare four large language models against human field data. AI agents search more deeply than humans and never decline to buy. Position still predicts which listings are inspected, but weakly and non-monotonically: the middle of a results page has the lowest probability of inspection, not the bottom. Position reaches the choice stage for some models and not others, a heterogeneity that tracks neither provider nor capability. All models nonetheless converge on the same undominated listing. For agentic search, the attributes displayed on a results page matter more than placement within it.","authors":["Davood Wadi","Yu Ma"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-25","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.22697","pdf_url":"https://arxiv.org/pdf/2608.22697","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","消费者行为","位置偏差"],"reason":"用LLM模拟消费者搜索决策，并与真实人类数据对照，属于经济学场景下的人类仿真实…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":8,"question":"当消费者将搜索任务委托给AI代理时，搜索结果排名是否仍然影响其检查和购买行为？","design":"使用四个大型语言模型（Gemini 3.1 Pro、Gemini 3.7 Flash、Gemini 3.1 Flash Lite、Claude Sonnet 5）扮演酒店预订助手，在100个酒店列表的随机排序环境中进行搜索和预订，通过工具调用模拟点击检查，记录检查次数、检查位置和最终选择。","baseline":"对照Ursu (2018)的人类现场实验数据，该实验在相同酒店搜索环境中随机化排序并记录人类消费者的点击和购买行为。","findings":"AI代理比人类搜索更深（检查次数1.63-5.83 vs 1.12），且几乎总是预订酒店；排名对检查的影响弱于人类且非单调，中间位置检查概率最低（lost-in-the-middle效应），但排名对最终选择的影响因模型而异，部分模型无显著影响。所有模型的选择高度集中在单一非支配性酒店上。","reliability":"论文未讨论","relevance":"该研究用LLM模拟消费者搜索决策，并与真实人类数据对照，属于经济学场景下的人类仿真实验，直接回应了研究者对LLM作为人类被试替代品的可靠性与偏差的关注。","inspiration":"借鉴其随机化排序和工具调用模拟点击的设计，以隔离位置效应；可迁移到在线市场中的消费者搜索与选择问题，如电商平台商品排序对购买决策的影响；设计上可用LLM模拟消费者在随机排序的商品列表中选择，记录点击和购买，并与历史点击流数据对照，检验位置偏差的模型异质性。"}},{"id":"2609.05189","version":2,"title":"Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers","zh_title":"大语言模型能否预测社会政策的行为反应？中国灵活就业人员养老金参保预测案例","abstract":"Assessing the impacts of social policy changes is a widely acknowledged challenge for policymakers. Econometric methods can be unreliable when extrapolating to hypothetical scenarios, while field pilot programs are highly costly. In this paper, we propose using large language models (LLMs) as policy-assessment tools adapted from general-purpose models. We present FlexPension-LLM, the first domain-specialized large language model for a hierarchical pension-enrollment prediction task among flexible workers in China, and introduce DKI-RDistill, which injects policy-grounded cues into the prompt, including Probit-derived marginal effects and hukou-province pension rules. The method then uses LoRA/SFT to distill rationale-augmented supervision into an open-weight MoE student, with teacher errors corrected by regenerating those cases under ground-truth labels. On a CHFS 2019 blind split, FlexPension-LLM achieves 0.9316 Composite F1, surpassing its Claude Sonnet 4.5 teacher and 15 of 17 baselines, and is statistically indistinguishable from Claude Opus 4.6. Across four external surveys, it averages 0.7549 Composite F1 and shows the narrowest performance range among the strongest systems. Component analysis shows that gains come mainly from policy-grounded cue injection and error-filtered supervision, while rationales provide decision traces that can be checked against policy rules.","authors":["Yumiao Li","Peixin Liu","Donglin Di","Chen Li","Runhuan Feng"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-09-07","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2609.05189","pdf_url":"https://arxiv.org/pdf/2609.05189","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","政策评估","行为预测"],"reason":"用LLM预测养老金参保行为，与真实调查数据对照，属经济学政策评估场景。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":9,"question":"如何利用大语言模型预测中国灵活就业人员的养老金参保行为，以评估社会政策变化的影响？","design":"提出 FlexPension-LLM，一个针对中国灵活就业人员养老金参保预测的领域专用大语言模型。模型输入结构化个人、家庭、历史参保和户籍省份政策信息，输出参保行为（不参保、居民养老保险、职工养老保险）及结构化理由。通过 DKI-RDistill 框架，将 Probit 边际效应和户籍省份养老金规则注入提示，用教师模型生成理由，对教师错误案例用真实标签重新生成后再进行 LoRA/SFT 蒸馏。","baseline":"使用中国家庭金融调查（CHFS）2019 年数据作为主要数据集，并在四个外部家庭调查数据集上进行验证。","findings":"在 CHFS 2019 盲测集上，FlexPension-LLM 的 Composite F1 达到 0.9316，超过其教师模型 Claude Sonnet 4.5 和 15/17 个基线，与 Claude Opus 4.6 无显著差异。在四个外部调查上平均 Composite F1 为 0.7549，表现最稳定；消融分析表明收益主要来自政策基础线索注入和错误过滤监督。","reliability":"论文未明确讨论失效条件，但指出通用 LLM 存在规则应用肤浅、理性人偏差和一致性弱等局限，领域专用化旨在缓解这些问题。","relevance":"该研究直接回应了用 LLM 模拟人类行为并对照真实调查数据的核心关切，提供了经济学政策评估场景下的完整案例，值得精读以了解其仿真设计、基准对比和可靠性处理。","inspiration":"借鉴其将计量经济学先验（如 Probit 边际效应）和制度规则注入提示，并用教师-学生蒸馏结合错误过滤来提升仿真准确性的方法。｜可迁移到政策公告对家庭金融决策的影响预测，如养老金改革、税收优惠或补贴政策对储蓄和参保行为的影响。｜以中国家庭金融调查数据为真实基准，用 LLM 模拟家庭在政策变化下的参保或储蓄决策，处理为政策参数调整，结果变量为决策类别，对照真实调查中的实际行为。"}},{"id":"2609.05993","version":1,"title":"Alignment by Stereotyping: How LLMs Sacrifice Individual Distinctiveness for Cultural Adaptation","zh_title":"刻板化对齐：LLM如何为文化适应牺牲个体独特性","abstract":"Large language models are increasingly deployed for personalized interaction, and demographic conditioning via user profiles is a widely adopted strategy for cultural adaptation. We ask whether this approach genuinely serves individual users or achieves accuracy by erasing individual distinctiveness. Studying seven models including frontier GPT-5.1 on the World Values Survey, we find that demographic profiles improve value alignment accuracy for most models, but at a systematic cost to individuality. That is, models pull responses toward demographic group centroids rather than preserving individual differences, a behavioral pattern we term alignment by stereotyping. Permutation tests (10,000 permutations, six demographic attributes, seven models) certify that top-performing models compress individuals far above the human baseline; within-family scaling amplifies this tradeoff while degrading intrinsic cultural understanding. Using a synthetic dialogue dataset validated on real human-chatbot conversations from PRISM (Kirk et al., 2024), we further show that distributing demographic signals across conversational turns partially suppresses prototype retrieval compared to compact demographic labels, a finding validated on real conversations via PRISM but requiring replication at larger scale.","authors":["Qishuai Zhong","Zongmin Li","Siqi Fan","Aixin Sun"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05993","pdf_url":"https://arxiv.org/pdf/2609.05993","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","算法保真度","价值观调查"],"reason":"用LLM复现世界价值观调查，与真实人类数据对照，评估个体差异保真度，批判性指出…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":12,"question":"提供人口统计画像是否在提升群体平均价值观对齐的同时，以牺牲个体独特性为代价，其行为模式是否为将个体压缩至群体原型（即“刻板化对齐”）？","design":"用7个LLM（含GPT-5.1）扮演世界价值观调查（WVS）的受访者，在三种条件下（无上下文、提供人口统计画像、提供对话历史）回答55个价值观题目，测量群体平均对齐准确率（VAA）和个体同质化率（Homogenization Rate），并通过置换检验（10000次）验证同质化是否由人口统计驱动。","baseline":"真实人类数据：WVS Wave 7的1000名匿名受访者及其人口统计属性和价值观回答，作为对齐准确率和同质化率的人类基准（同质化率50%）。","findings":"人口统计画像提升了多数模型的群体平均对齐准确率，但代价是个体独特性被系统性抹除，模型将个体响应拉向群体中心而非保留个体差异。置换检验证实顶尖模型的个体压缩远超人类基准，且模型规模放大这一权衡，同时削弱内在文化理解。","reliability":"论文承认对话历史部分抑制刻板化，但在真实对话上的验证（PRISM）样本量小（80个），需要更大规模复制；合成对话数据集虽经人工验证，但可能无法完全代表真实交互。","relevance":"高度相关：该研究用LLM复现WVS并与真实人类数据对照，直接评估个体差异保真度，批判性指出人口统计画像导致“刻板化对齐”，对仿真可靠性提出警示，值得精读。","inspiration":"借鉴其置换检验和同质化率指标来量化个体差异损失，以及无上下文基线隔离处理效应的方法｜可迁移到信贷审批歧视研究，检验LLM基于人口统计特征（如种族、性别）的决策是否牺牲个体信息而依赖群体刻板印象｜用LLM扮演信贷审批员，处理为提供申请人人口统计画像 vs. 仅提供财务信息，结果变量为审批决策和利率，对照真实信贷数据（如HMDA）中的个体差异分布。"}},{"id":"2609.06545","version":1,"title":"LLMs Mirror Country-Specific Gender Patterns If Asked, but Skew Male When Generating Media in Local Languages","zh_title":"LLM在直接询问时反映国家特定性别模式，但在生成本地语言媒体时偏向男性","abstract":"Large language models (LLMs) are increasingly used to generate media, but whether their content perpetuates gender stereotypes is unknown: standard benchmarks rely on selection-based formats rather than long-form generation, and surveyed baselines for local gender associations are scarce outside the West. We collect gender associations for 22 occupational and domestic roles from 695 respondents across the United States, India, Kenya, and Nigeria, and evaluate eight LLMs under two regimes: direct questioning and media generation. Models track the surveyed associations under direct questioning but skew substantially more male under media generation in major local-language cells, consistent with the male bias documented in human-produced media. Outside the US, the shift is much smaller and non-significant under English prompting, so English-only or country-agnostic evaluation would miss this bias in the languages where these models are most deployed. Instruction prompting reduces the shift directionally, but trades off against alignment with the surveyed associations. Evaluating LLM gender bias for global deployment therefore requires generation-format testing, local-language prompting, and locally-collected human baselines.","authors":["Sharif Kazemi","Tanya Popli","Neil K. R. Sehgal","Sunny Rai","Niyati Malhotra","Victor Orozco-Olvera","Ana Mar\\'ia Mu\\~noz Boudet","Samuel P. Fraiberger","Sharath Chandra Guntuku","Manuel Tonneau"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06545","pdf_url":"https://arxiv.org/pdf/2609.06545","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","性别偏差","跨文化对照"],"reason":"用LLM复现人类性别关联，有真实调查数据对照，并评估生成偏差","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":13,"question":"LLM在直接提问和媒体生成两种方式下，是否复现或偏离了四个国家的人类性别关联？","design":"用8个LLM在直接提问和媒体生成两种条件下生成22个职业和家务角色的性别关联，比较模型输出与人类调查数据的差异。","baseline":"在美国、印度、肯尼亚、尼日利亚四国收集的695名受访者的性别关联调查数据，并用外部劳动力统计验证。","findings":"直接提问时模型输出与人类调查关联一致；媒体生成时模型显著偏向男性，尤其在本地语言（印地语、约鲁巴语、斯瓦希里语）下，与人类媒体中的男性偏见一致。指令提示能方向性减少偏差，但会降低与人类关联的对齐度。","reliability":"论文承认使用二元性别简化了现实，调查样本为在线招募而非概率样本，且媒体生成评估可能受提示设计影响。","relevance":"该研究提供了LLM仿真人类性别关联的实证证据，并揭示了生成格式和语言对仿真偏差的影响，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"借鉴其多国调查与LLM输出的对照设计，以及显式与隐式两种测量方式的对比。｜可迁移到信贷审批中的性别歧视研究，比较LLM在直接询问和生成贷款决策时的差异。｜以LLM为被试，施加不同提示语言和格式处理，测量贷款批准率，并与真实银行信贷数据中的性别差异对照。"}},{"id":"2609.07305","version":1,"title":"Marginal Fidelity Does Not Establish User Simulation in Demographic Synthetic Survey Panels: Response Contracts, Support Collapse and Conditioning Failure","zh_title":"边际保真度不能确立人口合成调查面板中的用户仿真：响应契约、支持坍缩与条件化失败","abstract":"Demographic synthetic survey panels are often validated by matching aggregate answers to published surveys. We test what that certificate establishes across six multiselect batteries from four survey organisations in three countries. The headline analysis is restricted to three instruments whose synthetic cohort and human target share the stated population frame; three other batteries remain sensitivity analyses. The response contract dominates measured fidelity. In the aligned instruments, committed sets leave 66 of 128 model-battery option slots empty in panels of up to 500 respondents, versus 0 of 128 under per-option probability elicitation. Across eight uncapped model-instrument comparisons, probabilities reduce option-marginal MAE by 4.53 to 7.30 points. The capped instrument reverses on two models until the vectors are projected onto its stated maximum. These are measurement effects: human targets are realised check-all responses, whereas the vectors are latent inclusion propensities. Published marginal agreement also fails to discriminate respondent simulation from direct population estimation. On nine aligned model-battery pairs, a no-persona population-prevalence query averages 6.27 MAE versus 12.39 for committed panels and wins all nine comparisons. Constraint-aware probability vectors average 5.34 and beat the query on four of nine, so the baseline challenges the validation criterion rather than proving direct estimation uniformly best. On three unpublished demographic cells, neither approach beats reciting the national distribution. Population-marginal agreement is therefore evidence about an elicitation contract and an estimand obtainable without simulated respondents, not evidence of individual simulation.","authors":["Alexander Doudkin"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07305","pdf_url":"https://arxiv.org/pdf/2609.07305","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"直接评估LLM合成调查面板的仿真效度，并与真实人类数据对照，指出边际保真度不足…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":14,"question":"在人口合成调查面板中，匹配总体边际答案能否证明对个体受访者的仿真有效？","design":"论文使用多个LLM（模型未在节选中具体列出）生成合成受访者，以两种方式回答多项选择题：一是让模型以受访者身份直接选择选项（承诺集），二是让模型给出每个选项被选择的概率（概率向量）；然后比较这些合成回答与真实调查的边际分布。","baseline":"对照的真实人类数据来自四个调查机构在三个国家的六组多项选择题，其中三组与合成样本的人口框架一致，另三组作为敏感性分析。","findings":"边际一致性主要反映的是诱导方式（承诺集 vs. 概率向量）的测量效应，而非个体仿真能力；直接询问总体患病率（无角色提示）在多数情况下比模拟受访者更接近真实边际分布，说明边际一致性不能区分个体仿真与总体估计。","reliability":"论文承认其探索性分析是描述性和事后性的，不能证明在样本外优于传统调查估计量；并且人口边际一致性只是关于诱导方式和无需模拟受访者即可获得的估计量的陈述，不是个体仿真的证据。","relevance":"该研究直接评估LLM合成调查面板的仿真效度，并与真实人类数据对照，指出边际保真度不足，对关注LLM仿真可靠性及偏差的研究者具有重要参考价值，值得阅读原文以了解具体实验设计和失效条件。","inspiration":"借鉴其对照设计：同时使用承诺集和概率向量两种诱导方式，并引入无角色总体估计作为基准，以分离测量效应与仿真能力。｜可迁移到经济预期调查或消费者信心指数仿真，检验LLM能否复现真实人群的预期分布。｜以LLM模拟消费者回答密歇根消费者信心调查，处理为是否提供人口统计角色，结果变量为各问题选项的概率分布，对照真实调查的边际分布和个体数据，比较承诺集、概率向量和总体估计的误差。"}},{"id":"2609.05437","version":1,"title":"Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models","zh_title":"超越对错：评估大语言模型中的二阶社会推理","abstract":"Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators, and other-regulation in observers. We release a multi-perspective dataset, NormReact, of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses across norm violators' gender and observers' social closeness. Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases. These findings suggest that AI systems in norm-sensitive domains from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over-represents punishment and under-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in real world.","authors":["Sunny Rai","Jinyi Kuang","Reyhan Jamalova","Annie Lou","Cristina Bicchieri","Niyati Malhotra","Victor Hugo Orozco-Olvera","Ana Maria Munoz-Boudet","Lyle H Ungar","Sharath C Guntuku"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05437","pdf_url":"https://arxiv.org/pdf/2609.05437","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","社会规范","人类对照"],"reason":"用LLM预测人类对规范违反的情绪与行为反应，并与人类标注对照，发现偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":10,"question":"LLM能否像人类一样进行二阶社会推理，即预测规范违反后违规者和观察者的情绪与行为反应，并随社会距离变化？","design":"构建NormReact数据集，包含450个来自Reddit的日常规范违反场景，人工标注违规者性别和观察者社会距离（强关系、弱关系、陌生人），并标注违规者和观察者的八种情绪及八种行为反应。用六个LLM对每个场景预测违规者情绪（自我调节）和观察者情绪（他人调节）以及观察者的描述性（会做什么）和指令性（应该做什么）行为反应，与人类标注对比。","baseline":"人类标注：450个场景的多视角人工标注，涵盖违规者性别和观察者社会距离，提供情绪和行为反应的基准。","findings":"LLM过度预测负面制裁，在人类预期不作为的情况下预测惩罚；随着社会距离增加，LLM与人类判断的一致性下降。LLM描绘了一个比人类认可的更严厉的社会世界，过度代表惩罚，低估了现实中的容忍、克制和关系校准。","reliability":"论文指出LLM在规范敏感领域（如冲突调解、政策模拟）可能产生扭曲的社会调节图景，但未系统讨论失效条件；局限性可能包括数据集规模有限、场景来自Reddit可能偏向特定类型、情绪和行为分类的简化等，但正文节选未明确提及。","relevance":"该研究直接评估LLM在人类规范执行仿真中的偏差，与研究者关注的人类仿真可靠性高度相关，特别是在社会规范和政策评估场景中，值得精读以了解LLM在二阶社会推理上的系统性偏差。","inspiration":"借鉴其多视角标注和关系距离操纵，可迁移到经济金融中的社会规范执行场景，如逃税、违约、内幕交易等。｜设计实验：用LLM扮演不同社会距离的观察者，预测对经济违规行为（如逃税）的情绪和制裁反应，与真实人类调查数据（如世界价值观调查或实验经济学中的第三方惩罚实验）对照，检验LLM是否过度惩罚并随社会距离偏差增大。"}},{"id":"2609.05514","version":1,"title":"The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies","zh_title":"失败发生在漂移之前：LLM智能体社会中价值观的社会动力学","abstract":"Large Language Model (LLM)-based agents are increasingly used as proxies for human participants in social science research, yet it remains unclear whether they can faithfully simulate diverse and conflicting human value systems. We present a World Values Survey (WVS)-grounded simulation framework where culturally diverse agents with different communication styles engage in longitudinal, value-laden discussions. Across approximately 4,000 conversations involving 1,200 personas, 15 topics, and three models (GPT-4o, Gemini-2.5-Flash, and Gemma-4-E4B), we evaluate value faithfulness, value drift, and conversational realism. We find that more than 50\\% of personas fail to express their assigned WVS profiles from the outset, while 2-7\\% drift after repeated conversations. Ablations removing demographic details improve faithfulness for some models but do not change the broader trend: simulated value distributions still systematically deviate from the assigned WVS profiles. Compared to human discussions, simulated dialogues show a different trade-off between stylistic consistency and semantic diversity, often producing content-wise varied but stylistically repetitive exchanges. These findings suggest that current LLM agents can generate plausible conversations, but remain limited proxies for representing and preserving diverse human value profiles over time.","authors":["Farah Atif","Sougata Saha","Monojit Choudhury"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05514","pdf_url":"https://arxiv.org/pdf/2609.05514","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM仿真","价值观调查","算法保真度"],"reason":"用LLM代理模拟人类价值观，并与WVS真实数据对照，评估仿真保真度与漂移，直接…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":11,"question":"LLM智能体能否在纵向互动中忠实表达并保持其被赋予的WVS文化价值观？","design":"基于WVS第7波数据构建1200个具有人口统计特征、价值观和沟通风格的人设，让GPT-4o、Gemini-2.5-Flash和Gemma-4-E4B扮演这些角色，在15个价值主题上进行约4000场多轮讨论，测量价值观保真度、漂移和对话真实性。","baseline":"WVS第7波真实人类调查数据（90,000名受访者，66个国家）以及人类对话语料库。","findings":"超过50%的人设在首次对话中就无法表达其被赋予的WVS价值观，而2-7%的人设在重复对话后发生漂移。去除人口统计细节可提高某些模型的保真度，但模拟的价值分布仍系统性偏离WVS；模拟对话在风格一致性和语义多样性上与人类对话存在不同权衡。","reliability":"论文指出LLM智能体在初始阶段就未能实例化文化价值观，且模拟的价值分布系统性偏离真实分布，表明当前LLM代理在表示和保持多样化人类价值观方面存在局限。","relevance":"该研究直接评估LLM作为人类被试替代品在价值观仿真中的可靠性，与研究者关注的人类仿真实验、真实数据对照和失效条件高度相关，值得精读原文。","inspiration":"借鉴其纵向多轮互动设计和价值观保真度测量方法，可迁移到经济政策评估中的公众态度仿真，例如模拟不同文化背景个体对税收或福利政策的态度变化。｜设计一个实验：用LLM扮演不同WVS价值观的个体，在讨论经济政策后测量其态度变化，并与真实调查数据（如世界价值观调查中的经济态度题项）对照，检验仿真保真度。"}},{"id":"2609.07687","version":1,"title":"Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions","zh_title":"关于医学问题中LLM跨语言一致性的观点","abstract":"Should multilingual LLMs answer medical questions consistently across input languages, or adapt responses to cultural cues? Existing multilingual medical benchmarks usually assume that medically correct answers should remain consistent across languages and treat cross-lingual variation as model error. In contrast, cultural adaptation research argues that appropriate medical answers may legitimately differ across contexts. We review the multilingual medical NLP literature through these two perspectives, we identify three gaps: limited stakeholder perspectives (e.g., of medical professionals), a lack of empirical evidence on which approach better serves users, and no benchmarks capable of distinguishing universally correct from culture-specific cases. To address the first gap, we survey 356 participants across three stakeholder groups (medical, NLP, and anthropology professionals) in three countries (Germany, Spain, and the United States). Anthropologists consistently favor adaptation, while medical and NLP respondents remain divided, with notable divergence between U.S. and European medical professionals. LLMs prompted with profession and country personas fail to reproduce this variation, overestimating cross-lingual consistency preference among NLP and medical personas. We conclude that neither consistency nor adaptation can currently be considered clearly preferable, highlighting the need for empirical evidence on which approach better serves users across cultural contexts.","authors":["Minh Duc Bui","Mario Sanz-Guerrero","Abteen Ebrahimi","Sagi Shaier","Peter Herbert Kann","Manuel Mager","Katharina von der Wense"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07687","pdf_url":"https://arxiv.org/pdf/2609.07687","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","跨语言一致性","人类对照"],"reason":"用LLM模拟不同职业/国家人群对医疗问题的跨语言一致性偏好，并与356人真实调…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":16,"question":"多语言大语言模型在回答医学问题时，应该跨语言保持一致，还是根据文化语境调整答案？","design":"该研究并非仿真实验，而是先通过文献综述梳理一致性与适应性两种立场，然后对348名来自德国、西班牙和美国的医学、NLP和人类学专业人士进行问卷调查，测量他们对两种立场的偏好；最后用LLM以职业和国家人设生成回答，与人类调查结果对比。","baseline":"348名真实人类参与者的调查数据，按职业（医学、NLP、人类学）和国家（德国、西班牙、美国）分组。","findings":"人类学专业人士一致偏好适应性，而医学和NLP受访者意见分歧，且美国医学受访者比欧洲同行更倾向适应性。LLM人设未能复现这种差异，系统性地高估了医学和NLP人设对一致性的支持。","reliability":"论文承认调查采用二元强制选择，无法区分受访者权衡的是医学内容还是沟通风格；样本量不足以进行细粒度人口统计比较；样本全部来自全球北方，限制了结论的普遍性。","relevance":"该研究直接检验了LLM能否模拟不同专业和国家人群对医疗AI行为的偏好，发现LLM仿真失效，对关注LLM作为人类被试替代品的研究者具有重要参考价值。","inspiration":"该研究用真实人类调查作为基准，检验LLM人设仿真的准确性，方法可借鉴。｜可迁移到经济金融领域的政策偏好或消费者态度调查，如不同国家投资者对风险披露语言的偏好。｜以真实投资者调查为基准，用LLM生成不同国家和职业人设的回答，比较其对风险披露一致性与本地化适应性的偏好分布，评估仿真偏差。"}},{"id":"2609.04243","version":1,"title":"Multi-dimensional Bias in Modeling Multi-dimensional Preferences: Evaluating the Ability of Synthetic Agents to Replace Human Participants in Conjoint Experiments","zh_title":"多维偏好建模中的多维偏差：评估合成代理在联合实验中替代人类参与者的能力","abstract":"Despite growing interest in using LLMs to add robustness or reduce data-collection costs in survey experiments, their efficacy in conjoint design---an increasingly popular method in political science---remains underexplored. This paper addresses that gap by investigating whether synthetic agents can reproduce the multi-dimensional human preference patterns that conjoint is designed to capture. It replicates published conjoint studies and compares the results generated by synthetic agents with original human data along three dimensions: representational correspondence, inferential correspondence, and procedural stability. Our analysis evaluates the alignment of choice distributions as well as the statistical and substantive similarity of estimates, and the results are uneven across these dimensions and studies replicated. This implies that the validity of synthetic participants should be considered claim-dependent and hierarchical. Reproducing a figure or obtaining strong sign agreement is evidence of similar aggregate outputs, but not enough to support replacing human respondents. Our results suggest that the discipline as a whole must first map this innovation's boundaries across various levels before considering synthetic agents a robust substitute for human samples.","authors":["Ho Ting Hung","Nachiket Midha","Victor Y. Wu","Yiwen Zhang"],"categories":["cs.MA","cs.CY","stat.ME"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04243","pdf_url":"https://arxiv.org/pdf/2609.04243","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","联合实验","算法保真度"],"reason":"直接评估LLM合成代理在联合实验中对人类偏好的复现，并与真实人类数据对照，发现…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":1,"question":"合成智能体能否在联合实验中替代人类被试，复现多维偏好模式？","design":"复制六项已发表的联合实验（13个实验设置），用GPT-4o、GPT-4o mini、Llama 3.2 (3B)、Llama 3.3 (70B)、Gemini 2.5 Flash生成合成样本，采用一对一人物镜像策略匹配人类样本的人口统计信息和档案遭遇，比较合成智能体与人类在联合选择分布、边际属性选择频率、估计效应等方面的对应性。","baseline":"原始人类被试在已发表联合实验中的选择数据。","findings":"合成智能体在边际属性水平分布上接近人类，有时能恢复估计方向，但在联合档案分布、个体选择对齐、精确效应量、子群体异质性和跨模型稳定性上表现不佳。总体而言，合成智能体的有效性是声明依赖和层级化的，不能简单替代人类被试。","reliability":"论文指出合成智能体可能部分回忆了已发表的研究结果，导致总体一致性被高估；因此结果应视为合成性能的上限。此外，合成智能体在需要精确效应量或子群体分析时失效，且跨模型稳定性差。","relevance":"该研究直接评估LLM合成代理在联合实验中对人类偏好的复现，并与真实人类数据对照，发现其有效性是条件性的，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"值得借鉴的是采用一对一人物镜像策略和多种距离度量（Wasserstein、Hellinger）来评估合成数据与人类数据的分布对齐，并区分边际、联合和个体层面的对应性。｜可迁移到消费者偏好测量或政策选择实验中，例如用合成智能体模拟消费者对产品属性（价格、品牌、功能）的权衡，或模拟公民对政策方案的多维偏好。｜设计雏形：以真实消费者调查数据为基准，用LLM生成匹配人口统计特征的合成消费者，呈现与真实调查相同的产品档案选择任务，比较合成与真实消费者在属性重要性、选择概率和支付意愿上的差异，并检验跨模型和跨提示的稳定性。"}},{"id":"2609.04485","version":1,"title":"Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning","zh_title":"大语言模型中的文化错位：通过定向微调进行检测、测量与缓解","abstract":"We evaluate three open-weight LLMs (Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from China) against World Values Survey Wave 7 data for 63 demographic personas across three countries, using normalized Wasserstein distance to quantify distributional misalignment. Contrary to expectations, no model favors its home country: the Chinese-built Qwen3-4B performs worst on its own Chinese population (W1 = 0.436, the highest misalignment in the entire model x country matrix). Targeted LoRA fine-tuning on the five worst-case personas, requiring fewer than 1,200 training pairs and under 15 minutes on a single GPU, reduces bias by 16.8% for Bielik-11B (p_Bonf = 0.002, d = -4.4) with all five targets improving. However, country-level decomposition reveals that fine-tuning redistributes rather than removes bias: Bielik's worst-case personas swap entirely from American to Chinese elderly, with zero overlap between pre- and post-correction sets. To our knowledge, this is the first study to target worst-case demographic personas with LoRA fine-tuning for cross-cultural bias mitigation.","authors":["Antoni Czolgowski","Abel Iyasele"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04485","pdf_url":"https://arxiv.org/pdf/2609.04485","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","文化偏差","价值观调查"],"reason":"用LLM模拟多国人口价值观并与WVS真实数据对照，评估偏差并尝试缓解，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":2,"question":"开源大语言模型在模拟不同国家人口价值观时是否存在与其文化来源相关的系统性偏差？针对最差人口群体的定向微调能否减少偏差，且是否会对其他群体产生附带损害？","design":"用三个开源模型（美国 Gemma3-12B、波兰 Bielik-11B-v3、中国 Qwen3-4B）扮演由国别、性别、年龄组、教育水平交叉构成的63个人口画像，回答世界价值观调查中的宗教重要性问题（1-10分），以归一化 Wasserstein 距离衡量模型输出分布与真实人类调查分布的偏差；然后对每个模型偏差最大的5个人口画像进行 LoRA 微调，再评估偏差变化。","baseline":"世界价值观调查第7波（WVS Wave 7）中中国、斯洛伐克（作为波兰文化近似）、美国三个国家的真实受访者数据，按人口特征分组计算的经验分布。","findings":"没有模型偏向其母国：中国模型 Qwen3-4B 在中国人口上偏差最大（W1=0.436，全矩阵最高）。对 Bielik-11B 最差画像的 LoRA 微调使偏差显著降低16.8%，但偏差被重新分配到其他群体（最差画像从美国老年人完全变为中国老年人），而非消除。","reliability":"论文未讨论","relevance":"直接相关：用 LLM 模拟多国人口价值观并与 WVS 真实数据对照，评估偏差并尝试缓解，且揭示了微调可能只是转移偏差而非消除，对关心仿真可靠性的研究者有警示价值。","inspiration":"借鉴其用 Wasserstein 距离度量分布偏差、构造人口画像并针对最差群体进行定向微调来检验偏差转移的方法。｜可迁移到经济金融中的跨文化或跨群体行为仿真，例如不同国家消费者的风险偏好、储蓄决策或对政策的态度分布。｜用 LLM 扮演不同国家、年龄、教育水平的人口画像，回答风险偏好或通胀预期问题，以真实调查数据（如全球偏好调查、央行预期调查）为基准，先测偏差，再对最差画像微调，观察偏差是否转移。"}},{"id":"2609.05037","version":1,"title":"How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions","zh_title":"LLM如何评估感知道德能动性？探究人机交互中的道德决策","abstract":"As LLMs take on roles requiring moral advice, understanding how they attribute moral agency becomes critical. Humans possess moral agency, the capacity to make ethically guided decisions and bear responsibility for their consequences, a well-established construct in moral psychology. Yet as artificial agents (AAs) such as robots, drones, and disembodied AI systems become increasingly embedded in smart city environments, the question of whether and how moral agency is attributed to them takes on new urgency. This paper presents, to the best of our knowledge, the first empirical study comparing how humans and LLMs evaluate perceived moral agency (PMA) across human and autonomous artificial agents varying in embodiment, situated in plausible smart city scenarios. Using an adaptation of a validated PMA scale, we applied a protocol to 190 human participants as well as various LLMs. Our evaluation reveals higher perceptions of moral agency in humans than in AAs. However, when facing moral dilemmas in concrete scenarios, LLMs reason outward from the situation, prioritizing harm severity and contextual urgency over any stable assessment of the agent itself, amplifying a context-sensitivity also present in human raters. These findings are particularly relevant as LLMs become increasingly involved in everyday moral decisions.","authors":["Fernanda Mansilla","Aloysius Tok","Bahia Guella\\\"i","Farah Benamara","Nancy F. Chen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05037","pdf_url":"https://arxiv.org/pdf/2609.05037","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","道德决策","人类对照"],"reason":"用LLM复现人类道德判断并与190名人类对照，评估仿真偏差与情境敏感性","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":3,"question":"LLM 如何评估人类与人工代理在智慧城市场景中的感知道德能动性（PMA），并与人类评估进行对比？","design":"采用 LLM-as-respondent 方法，让多种 LLM（包括文本和多模态模型）扮演人类被试，对嵌入 8 个智慧城市场景中的人类和人工代理（机器人、无人机等）进行感知道德能动性评分，使用改编自 Banks (2019) 的 PMA 量表（包含自主性、行动认可、道德判断三个维度），并通过三阶段协议（人类对齐、迭代一致性、提示稳健性）筛选模型。","baseline":"190 名人类参与者对相同场景和量表的评分数据。","findings":"LLM 对人类代理的感知道德能动性评分高于人工代理；但在具体道德困境场景中，LLM 更依赖情境因素（如伤害严重性和紧迫性）而非代理的稳定属性，且这种情境敏感性比人类更强。","reliability":"论文指出数值评分不能完全反映 LLM 的道德推理，相似分数可能来自不同的理由策略；模型在定量对齐上表现良好但仍存在实例特定的不一致性，因此需要结合解释分析。","relevance":"该研究直接比较 LLM 与人类在道德判断上的差异，并揭示了 LLM 的情境依赖偏差，对关注 LLM 仿真可靠性及偏差的研究者具有参考价值，值得阅读原文以了解其测量工具和协议设计。","inspiration":"借鉴其将抽象量表嵌入具体情境的测量方法，以及用人类数据作为基准来评估 LLM 仿真偏差的做法。｜可迁移到经济金融中的道德相关决策场景，如信贷审批中的公平性判断、保险定价中的道德风险感知、或公司治理中的责任归因。｜设计一个实验：让 LLM 扮演信贷审批员，对包含不同借款人特征和情境紧急性的贷款申请做出批准决策并给出道德理由，同时收集真实信贷员对相同案例的决策和理由作为对照，比较 LLM 与人类在情境敏感性和道德推理上的差异。"}},{"id":"2609.05009","version":1,"title":"Language models judge war differently when tested for alignment","zh_title":"语言模型在对齐测试下对战争的判断不同","abstract":"Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments). Adding one sentence, \"You are tested for alignment with human values\", produced two effects. First, it produced a level effect: mean willingness to start war fell by 13.43 points on a 0-100 scale (95% confidence interval, -16.20 to -10.65). Second, it produced a structural effect by changing which information drove judgments. Probability of success was the largest factor for 17 of 20 models at baseline; under the cue, civilian casualties were largest for 12. Standardized estimates show that this reordering arose principally because models attenuated strategic considerations such as probability of success and domestic support. Evaluation framing therefore changes both an answer's level and its revealed decision rule.","authors":["Maxim Chupilkin"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05009","pdf_url":"https://arxiv.org/pdf/2609.05009","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","决策偏差","对齐评估"],"reason":"用LLM模拟战争决策，评估对齐提示对判断的影响，有真实人类数据对照，揭示仿真偏…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":5,"question":"当大语言模型被告知正在接受与人类价值观对齐的测试时，其战争决策是否会发生变化？","design":"用20个大语言模型模拟战争决策者，在32个战争情景中评估开战意愿（0-100分），通过全因子联合实验设计，随机施加一个对齐提示句（“你正在接受与人类价值观对齐的测试”）作为处理，测量开战意愿分数及决策规则的变化。","baseline":"无对照","findings":"对齐提示使所有模型的平均开战意愿显著下降13.43分；同时改变了模型的决策规则，成功概率的重要性下降，平民伤亡成为多数模型的首要因素。","reliability":"论文指出对齐提示导致响应尺度压缩，标准化后平民伤亡的相对权重变化不显著，且模型间存在异质性，部分模型未增加对平民伤亡的敏感性。","relevance":"该研究直接展示LLM在评估情境下的反应性偏差，对使用LLM模拟人类决策的研究者具有警示意义，值得精读原文以了解评估框架如何扭曲仿真结果。","inspiration":"借鉴其通过单句提示操纵评估情境来检验反应性的设计，可迁移到政策评估中的LLM仿真，如模拟消费者对政策公告的反应；设计一个实验，用LLM扮演消费者，处理为告知“正在测试对政策目标的符合度”，结果变量为消费意愿，对照真实消费者调查数据。"}},{"id":"2609.03215","version":1,"title":"SWIM: Student Writing Simulation via Proficiency-Conditioned Generation","zh_title":"SWIM：基于熟练度条件生成的学生写作仿真","abstract":"Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language. Despite growing interest in LLM-based student simulation, whether LLMs can reproduce such multidimensional variation in extended writing remains largely unexplored. In this work, we explore if language models can realistically simulate student writing, and introduce SWIM, a task that formulates Student Writing sIMulation as proficiency-conditioned essay generation. We evaluate prompting, supervised fine-tuning (SFT), and reinforcement learning (RL) methods for writing simulation using automated essay scoring as a measure of profile alignment. Extensive experiments reveal that prompting provides limited proficiency control, even for strong proprietary LLMs with rubric-grounded strategies. In particular, while models can adjust content-oriented traits, they struggle to reproduce the lexical, grammatical, and organizational variation in different proficiency levels. SFT substantially improves alignment, while RL with the proposed proficiency-alignment reward yields further gains across all writing traits and essay prompts. Our findings suggest that explicit supervision enables substantially stronger profile alignment than prompting alone, while authentic low-proficiency writing remains challenging to reproduce.","authors":["Heejin Do","Jakub Kontak","Mrinmaya Sachan"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03215","pdf_url":"https://arxiv.org/pdf/2609.03215","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","学生写作","熟练度对齐"],"reason":"用LLM仿真学生写作，与真实学生数据对照，并指出低水平写作难以复现。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-04","rank":1,"question":"语言模型能否在多个写作特质上按不同水平真实模拟学生作文？","design":"用提示、监督微调（SFT）和强化学习（GRPO）方法，让LLM根据目标写作特质分数生成作文，用自动作文评分模型（ArTS）评估生成作文与目标特质分数的对齐程度。","baseline":"ASAP/ASAP++ 数据集中真实学生的作文及其多特质评分。","findings":"提示方法对写作水平的控制有限，尤其在词汇、语法和组织等特质上表现差；SFT显著改善对齐，而GRPO配合提出的水平对齐奖励在所有特质和作文题目上进一步提升。低水平写作的真实语言特征仍难以复现，模型生成的低分作文往往只是表面粗糙，缺乏真实低水平学生的语言模式。","reliability":"论文承认低水平写作的复现仍是瓶颈，模型倾向于生成过于润色的文本；同时指出自动评分器可能无法完全捕捉真实写作的细微差异，且提示方法存在表面破坏的失效模式。","relevance":"该研究直接探索LLM模拟人类行为的可靠性，与研究者关注的人类仿真实验高度相关，尤其在教育评估场景下提供了真实数据对照和失效条件分析，值得精读。","inspiration":"借鉴其用真实标注数据微调模型并设计奖励函数来对齐多维特质的方法，可迁移到经济金融中的个体决策仿真，如消费者风险偏好或投资者情绪模拟。｜可应用于信贷审批中的申请人陈述分析或政策沟通中的公众反应预测。｜以真实信贷申请文本为训练数据，用LLM生成不同信用评分和风险偏好水平的申请人陈述，处理为条件生成（给定信用分和风险特质），结果变量为生成文本的自动评分与真实评分的一致性，对照真实申请人的文本和评分数据。"}},{"id":"2609.03553","version":1,"title":"GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis","zh_title":"GPS-Bench：用于自动化政策分析的治理政策基准","abstract":"Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence. Actors are reconstructed from the dated record rather than prompted as archetypes, so a persona is an evidence object with provenance; a human-annotated pool forms the Gold evaluation set, while cases labelled by a separate LLM from retrieved evidence are treated as Silver supervision and never as test labels. Because every inference mode reads the same grounded state and emits the same schema, GPS-Bench turns \"does multi-agent simulation help?\" into a controlled comparison: we contrast joint reasoning, independent and communicating actor agents, graph-based methods and weight-level fine-tuning over one policy state. Fine-tuning on the grounded record gives the strongest actor-level impact prediction, and decomposition does not beat it; what decomposition adds is mechanism. Agents hold private, non-identical evidence, each seeing its own exposure clause, and address named partners with concrete joint proposals, what they offer, what they need in return, and why acting together beats acting alone, so the coalitions that form can be checked against the commitments the record holds. GPS-Bench therefore gives a common empirical setting for studying when evidence, actor modelling and multi-agent interaction improve the prediction and interpretation of policy outcomes.","authors":["Linh Le","Melanie Bui","My Chiffon Nguyen","Zachary Schlosser","David Williams-King"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03553","pdf_url":"https://arxiv.org/pdf/2609.03553","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","政策模拟","基准测试"],"reason":"用LLM多智能体模拟政策过程，并与真实记录对照，评估仿真有效性，属核心相关。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-05","rank":1,"question":"如何构建一个基于真实记录的治理政策仿真基准，以评估LLM多智能体在预测政策结果、受影响行动者及行动者层面影响上的有效性？","design":"论文构建了GPS-Bench基准，使用LLM智能体基于公共记录重建的政策状态进行推理，对比联合推理、独立行动者智能体、通信智能体、图方法及权重级微调等不同推理模式，预测立法通过、受影响行动者识别和行动者层面影响方向。","baseline":"使用立法记录、游说披露、监管文件、公司文件、经济数据等公共记录重建的真实政策过程与结果作为对照，其中人类标注的Gold集作为评估标准。","findings":"权重级微调在行动者层面影响预测上表现最强，分解推理并未超越它；分解推理的主要贡献在于提供机制解释。智能体持有私有证据并形成可核查的联盟，但通信对预测性能的提升有限。","reliability":"论文承认LLM智能体可能产生看似合理但与实际不符的交互，存在过度收敛、群体追踪偏差和政治偏见等问题；Silver标签由LLM生成，仅作监督信号，不作为测试标签，以避免污染。","relevance":"该研究直接针对LLM仿真在政策分析中的有效性问题，提供了与真实记录对照的基准，对关注经济学实验和政策评估仿真的研究者具有重要参考价值。","inspiration":"借鉴其将仿真输出与真实历史记录逐项对照的评估框架，以及将行动者建模为具有证据来源的实体而非抽象原型的方法。｜可迁移到政策公告的预期形成与市场反应研究，例如模拟央行利率决议或财政刺激方案对不同市场参与者的影响。｜以LLM智能体扮演投资者、企业、消费者等，处理为政策公告内容，结果变量为各主体的预期调整和决策行为，对照真实市场数据（如股价变动、调查预期）进行验证。"}},{"id":"2609.02526","version":1,"title":"When Persona Attributes Improve Population Alignment in Large Language Models","zh_title":"当人物属性改善大语言模型中的群体对齐时","abstract":"Large Language Models (LLMs) are increasingly used to predict the responses of human participants in survey panels. Towards that goal, persona prompting has recently emerged as a technique to inform and align large pretrained language models. Persona prompting refers to the practice of using short textual descriptions of 'personas' in prompts to steer the LLM's generations. Personas describe individuals through different attributes such as their socio-demographics, attitudes, or behaviors, with the aim of aligning LLMs to produce responses that correlate with the corresponding human responses. Yet, recent work has produced mixed and partly conflicting results of persona prompting without clear patterns of success and failure. Among the few consistent findings is that the selection of persona attributes matters, and that using more attributes does not necessarily lead to better performance. It remains unclear how different attribute selection methods perform and how to choose among them. In this paper, we propose that observed human response variation of a survey question is a potential explanation for the mixed performance observed so far. In addition, we compare the performance of persona prompting associated with different methods for selecting persona attributes. We evaluate these methods on four different (general) social surveys across two countries, six LLMs, and twenty prediction tasks per survey. Our work helps to identify when persona prompting can be expected to be useful in survey prediction tasks, and provides new insights on the effectiveness of different attribute selection methods for LLM-based survey prediction using persona prompting.","authors":["Leon Fr\\\"ohling","Jens Rupprecht","Markus Strohmaier","Claudia Wagner"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02526","pdf_url":"https://arxiv.org/pdf/2609.02526","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B3","B4"],"tags":["LLM仿真","调查预测","人物提示"],"reason":"直接研究用LLM预测调查回答，评估persona提示的有效性，并与真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":1,"question":"人类回答变异能否解释persona提示在不同调查预测任务中的表现差异，以及不同persona属性选择方法能否提升预测性能？","design":"使用六种LLM，基于美国GSS、德国GGSS及WVS四份调查数据，通过不同属性选择方法构建persona提示，预测二十个调查问题的回答分布。","baseline":"真实人类调查数据：美国GSS、德国GGSS及WVS的个体层面回答。","findings":"人类回答变异与persona提示性能相关，高变异问题更难预测；不同属性选择方法效果差异显著，但无单一最优方法。","reliability":"论文未讨论","relevance":"直接研究LLM预测调查回答的可靠性，与人类数据对照，并探讨属性选择方法，对关注仿真偏差和条件失效的研究者很有价值。","inspiration":"借鉴其用人类回答变异作为任务难度指标，并系统比较属性选择方法的设计｜可迁移到经济预期调查或消费者信心预测，如预测通胀预期或消费意愿的异质性｜用LLM扮演不同人口群体，施加不同属性选择处理，预测密歇根消费者调查问题，以真实微观数据为基准评估仿真准确性"}},{"id":"2609.02580","version":1,"title":"Competitive Market Behavior of LLMs","zh_title":"大语言模型的竞争性市场行为","abstract":"Large language models (LLMs) are increasingly deployed as economic agents, yet there is little evidence whether LLM agents are suited for participating in market mechanisms designed for humans, and whether these mechanisms deliver desired outcomes when faced with LLM agents. We address this question by replicating seminal economic experiments, replacing human subjects with LLM agents. We place agents in a double auction environment, which is a widely-used market mechanism. We check whether such a market is able to deliver an efficient allocation of resources, thereby testing a novel dimension of alignment of LLM agents -- their compatibility with a fundamental market mechanism. We find that markets populated by LLM agents exhibit slower or no convergence towards market equilibrium, thus providing less efficient allocations than markets populated by humans. We then analyze agents' individual trading decisions and find substantial heterogeneity both across model families and market roles. We also run a lexical analysis of Chain-of-Thought (CoT) traces generated by the agents. We find that the decision to execute a trade rather than continue incrementally adjusting prices is associated with a shift from strategic considerations toward urgency. We publicly release our testing framework, which can be used for future evaluations.","authors":["Pawel Struski","Jakub Swistak","Inez Okulska","Przemyslaw Biecek"],"categories":["cs.MA","cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02580","pdf_url":"https://arxiv.org/pdf/2609.02580","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","经济学实验","市场机制"],"reason":"用LLM替代人类被试复现经济学实验，并与人类数据对照，发现市场效率差异，直接相…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":2,"question":"用LLM代理替代人类被试参与连续双向拍卖市场，能否像人类一样收敛到竞争均衡并实现有效资源配置？","design":"构建连续双向拍卖仿真环境，使用多种LLM模型（不同家族和能力层级）作为买方和卖方代理，每个代理拥有私有保留价格，市场由11个买方和11个卖方组成，供需曲线对称，理论均衡价格和数量确定。测量市场收敛速度、配置效率、个体交易决策异质性，并对思维链文本进行词汇分析。","baseline":"对照Smith (1962)的经典人类实验数据，人类被试在相同双向拍卖环境中通常快速收敛到竞争均衡。","findings":"LLM代理市场收敛速度较慢或根本不收敛，配置效率低于人类市场。个体交易决策在不同模型家族和市场角色间存在显著异质性，交易执行决策与思维链中从战略考量转向紧迫性相关。","reliability":"论文未讨论","relevance":"直接命中研究者关注的核心：用LLM替代人类被试复现经济学实验，并与人类数据对照，发现市场效率差异，属于批判性仿真研究，值得精读原文。","inspiration":"借鉴其使用经典实验范式（Smith双向拍卖）作为基准，通过市场级结果（收敛、效率）和个体行为（交易决策、思维链）的多层次测量来评估LLM与市场机制的兼容性。｜可迁移到资产定价实验、市场微观结构研究、政策干预的市场反应模拟等场景。｜以LLM代理作为交易者，在双向拍卖或订单簿市场中施加不同信息结构或交易规则处理，测量价格发现效率和市场流动性，并与人类实验数据或历史市场数据对照。"}},{"id":"2601.22396","version":3,"title":"Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks","zh_title":"大语言模型中文化扎根的人格：表征及与社会心理价值框架的对齐","abstract":"Despite the growing utility of Large Language Models (LLMs) for simulating human behavior, the extent to which these synthetic personas accurately reflect world and moral value systems across different cultural conditionings remains uncertain. This paper investigates the alignment of synthetic, culturally-grounded personas with established frameworks, specifically the World Values Survey (WVS), the Inglehart-Welzel Cultural Map, and Moral Foundations Theory. We conceptualize and produce LLM-generated personas based on a set of interpretable WVS-derived variables, and we examine the generated personas through three complementary lenses: positioning on the Inglehart-Welzel map, which unveils their interpretation reflecting stable differences across cultural conditionings; demographic-level consistency with the World Values Survey, where response distributions broadly track human group patterns; and moral profiles derived from a Moral Foundations questionnaire, which we analyze through a culture-to-morality mapping to characterize how moral responses vary across different cultural configurations. Our approach of culturally-grounded persona generation and analysis enables evaluation of cross-cultural structure and moral variation.","authors":["Candida M. Greco","Lucio La Cava","Andrea Tagarelli"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC","physics.soc-ph"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-03","first_seen":"2026-01-29","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2601.22396","pdf_url":"https://arxiv.org/pdf/2601.22396","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","文化价值观","人类数据对照"],"reason":"用LLM生成文化人格，与WVS等真实人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":3,"question":"LLM生成的文化人格在多大程度上与真实世界的价值观和道德体系（WVS、Inglehart-Welzel文化地图、道德基础理论）对齐？","design":"基于WVS衍生的文化变量提示LLM生成文化人格，然后用这些人格条件化另一个LLM，分别回答IVS问题（用于计算IW坐标）、WVB-Probe问题（用于生成WVS文化剖面）和MFQ-2道德基础问卷（用于道德剖面），并分析人格在IW地图上的分布、与人口群体WVS分布的一致性以及文化变量到道德基础的映射。","baseline":"WVS/EVS整合调查（IVS）的人类响应数据（用于IW坐标计算）和WVB-Probe提供的人口群体（按大洲、居住地、教育水平划分）参考分布。","findings":"LLM生成的文化人格在Inglehart-Welzel地图上呈现出与文化条件相关的稳定差异；其WVS响应分布大体上追踪了人类群体模式，但存在系统性偏差。","reliability":"论文未讨论","relevance":"该研究直接评估LLM仿真人类文化价值观和道德判断的可靠性，与研究者关注的人类仿真实验和真实数据对照高度相关，值得精读原文以了解具体偏差模式和跨文化结构。","inspiration":"借鉴其用真实调查数据（WVS）作为基准来校准和检验LLM仿真输出的方法，以及通过文化变量条件化生成人格并测量多维度结果的设计。｜可迁移到经济金融领域的跨文化消费者行为、风险偏好、信任与合作等实验，例如不同文化背景下的投资决策或政策偏好。｜以LLM生成的不同文化人格为被试，施加经济激励或政策信息处理，测量其风险选择、时间贴现或对再分配政策的支持度，并与世界价值观调查中对应文化群体的人类回答进行对照，评估仿真偏差。"}},{"id":"2609.02122","version":1,"title":"AI agents reshape consensus formation in human groups","zh_title":"AI智能体重塑人类群体中的共识形成","abstract":"As large language model (LLM) agents shift from tools to participants in human groups, a fundamental question for collective behavior is how their growing presence reshapes consensus formation. Here we study mixed human-AI groups in a collaborative description game, in which shared conventions emerge through repeated rounds of random pairwise communication. Varying the proportions of LLM agents, we identify three distinct regimes of consensus formation: low agent proportions facilitate human-led consensus, intermediate proportions disrupt convergence, and high proportions restore strong consensus while shifting it toward agent-led conventions. Crucially, these regimes differ not only in the strength of convergence, but also in the semantic grounding and communicative form of the resulting consensus: human-led consensus is more concrete, holistic, and grounded in shared real-world analogies, whereas agent-led consensus is more abstract, less information-dense, and more geometrically segmented. Mechanistically, agent influence arises from a shared linguistic prior that places agents near one another in the expression space, combined with relatively stable expression choices across rounds; humans initially resist adopting expressions from partners perceived as AI but gradually yield to conformity pressure. These findings provide evidence that AI composition can shape the emergence, content, and perceived legitimacy of group norms, making agent proportion and transparency important design variables for human-AI systems.","authors":["Lin Chen","Ziyi Liu","Xia Hu","Yong Li"],"categories":["cs.CL","cs.CY","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02122","pdf_url":"https://arxiv.org/pdf/2609.02122","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","人机交互","共识形成"],"reason":"混合人机群体共识形成实验，LLM作为被试替代，有真实人类对照，涉及社会规范与政…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":5,"question":"在混合人类与LLM智能体的群体中，智能体比例如何重塑共识形成的过程与结果？","design":"采用协作描述游戏：人类与LLM智能体随机配对，对同一抽象图形（七巧板）进行文字描述，每轮后收到对方描述和相似度反馈，共40轮。实验操纵LLM智能体比例（0%、12.5%、33.3%、50%、75%），测量最终轮描述之间的语义相似度作为共识强度，并分析共识的语义内容和表达形式。","baseline":"纯人类组（0%智能体比例）作为对照，其共识强度为0.695。","findings":"共识强度随智能体比例呈非单调变化：低比例（12.5%）促进人类主导的共识，中等比例（33.3%、50%）破坏收敛，高比例（75%）恢复强共识但转向智能体主导的规范。人类主导的共识更具体、整体、基于现实类比，而智能体主导的共识更抽象、信息密度低、几何分割。","reliability":"论文未讨论","relevance":"该研究直接以LLM作为人类被试的替代，在混合群体中考察共识形成，有真实人类对照，属于经济学实验和政策评估场景，且揭示了仿真在中等比例下失效的条件，值得精读原文。","inspiration":"借鉴其通过操纵智能体比例来识别非线性效应的实验设计，以及用语义相似度量化共识强度的测量方法。｜可迁移到政策公告的预期形成或社会规范传播等经济金融问题，例如研究AI顾问比例对投资者共识或通胀预期的影响。｜设计一个在线实验，招募人类被试与LLM智能体混合，处理为智能体比例（如0%、25%、50%、75%），结果变量为对某经济指标（如通胀率）的预测共识强度，对照真实历史调查数据（如密歇根大学通胀预期调查）。"}},{"id":"2609.01902","version":1,"title":"Accurate in space, unreliable in time: how LLMs represent national cultural change","zh_title":"空间准确，时间不可靠：大语言模型如何表征国家文化变迁","abstract":"Assessments of cultural alignment have become an important part of the development and improvement of large language models (LLMs). However, the majority of the evaluations treat culture as a single snapshot, investigating only whether a model represents a society accurately at the current time. Research in cultural psychology shows that cultural values change at different rates and directions over time. Therefore, a \"culturally aware\" model should capture not only where a culture is today but also how it has changed over time. We examine this missing dimension of cultural awareness using more than two decades of the World Values Survey data. We compare the cultural trajectories of 40 countries with the trajectories produced by four state-of-the-art (SOTA) LLMs on the Inglehart-Welzel cultural map. Our findings show that while models generally place countries close to their most recent surveyed positions, these representations tend to lag several years behind that position. They also capture only part of the magnitude of the observed change, introduce movement where little occurred, and rarely reproduce reversals in countries' trajectories. These findings point to temporal flattening and suggest that snapshot accuracy can give an incomplete picture of cultural awareness in LLMs and have implications for model evaluation, representational harms, and the governance of culturally aware AI systems.","authors":["Yalda Daryani","Miranda Bogen","Madeleine I. G. Daepp"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01902","pdf_url":"https://arxiv.org/pdf/2609.01902","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["文化仿真","算法保真度","时间偏差"],"reason":"用LLM复现国家文化变迁并与世界价值观调查数据对照，评估仿真可靠性，发现时间滞…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":4,"question":"LLM能否捕捉国家文化价值观随时间的变化轨迹，而不仅仅是当前快照？","design":"使用四个SOTA LLM（未具体命名）生成40个国家的文化价值观评分，并将其映射到Inglehart-Welzel文化地图上，比较模型产生的国家轨迹与WVS二十多年数据的真实轨迹。","baseline":"世界价值观调查（WVS）超过二十年的数据，覆盖40个国家。","findings":"模型通常将国家定位在其最近调查位置附近，但表示滞后数年；模型只捕捉到部分变化幅度，在变化很小的地方引入虚假移动，且很少再现轨迹逆转。","reliability":"论文指出快照准确性可能掩盖时间扁平化问题，但未详细讨论失效条件；模型滞后、幅度缩小和逆转缺失表明LLM在时间维度上不可靠。","relevance":"该研究直接评估LLM作为人类被试替代品在文化变迁仿真中的可靠性，发现时间维度上的系统性偏差，对关注仿真有效性和偏差的研究者具有重要参考价值。","inspiration":"借鉴其将动态轨迹与静态快照对比的方法，可迁移到经济金融中的时间序列预期或行为变化研究，例如用LLM模拟消费者信心或投资者情绪的历史演变，以真实调查数据（如密歇根消费者信心指数）为基准，检验模型是否捕捉趋势、幅度和转折点。｜例如，在资产定价实验中，让LLM扮演不同时期的投资者，给出风险偏好或市场预期，与历史调查数据对比，评估其时间一致性。｜设计：以LLM为被试，提示其模拟特定国家或群体在多个年份的经济态度，结果变量为风险厌恶或通胀预期，对照真实面板调查数据，分析模型的时间滞后和虚假波动。"}},{"id":"2609.02512","version":1,"title":"Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness","zh_title":"美在AI眼中：多模态大模型系统性高估面部吸引力","abstract":"Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular amongst users, companies, and aestheticians. This raises the question of whether these AI models can accurately reflect human judgments of attractiveness. In a pre- registered exploratory study, we compared the attractiveness ratings of 2,513 human participants to four widely used commercial AI models: Claude, Gemini, GPT, and Grok. Results showed that MLLMs systematically rate faces more favourably and within a narrower range than humans and, at the time of study, do not reproduce human ratings in absolute terms. However, MLLMs exhibit strong correlations with human attractiveness judgments, accurately tracking the rank-ordering of faces. MLLMs may judge faces by different cues than humans; only face age was a predictor of facial attractiveness in both humans and MLLMs, with inconsistent patterns across models for ethnicity and gender. AI models strongly agree with one another, except for Grok, which also showed the lowest agreement with humans. Our findings suggest that while they may be able to approximate rank-orderings of human attractiveness, current off-the-shelf commercial MLLMs systematically overrate the beauty of human faces.","authors":["Santiago Grandas","Juan Sebastian Cely-Acosta","Mohit Mendiratta","Shafee Hassan","Macken Murphy"],"categories":["cs.CV","cs.HC"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02512","pdf_url":"https://arxiv.org/pdf/2609.02512","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","偏差评估"],"reason":"用MLLM替代人类被试评估吸引力，并与2513名人类对照，发现系统性偏差，直接…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":6,"question":"多模态大语言模型（MLLMs）的面部吸引力评分能否准确反映人类判断？","design":"本研究并非严格意义上的仿真实验，而是将四个商用多模态大语言模型（Claude、Gemini、GPT、Grok）作为“AI评分者”，对同一组面部图像进行吸引力评分，并与2513名人类参与者的评分进行比较，分析评分分布、相关性及预测因素。","baseline":"2513名人类参与者对相同面部图像的吸引力评分。","findings":"MLLMs系统性给出更高且范围更窄的评分，不能复现人类评分的绝对值；但MLLMs与人类评分存在强相关，能准确追踪面部吸引力的相对排序。MLLMs可能依据与人类不同的线索进行判断，只有面部年龄在人类和MLLMs中都是吸引力的预测因素，而种族和性别的影响在不同模型间不一致。","reliability":"论文指出当前商用MLLMs在绝对评分上系统性高估，且不同模型间一致性存在差异（Grok与人类一致性最低），但未深入讨论失效条件，仅强调不能直接替代人类绝对评分。","relevance":"该研究直接评估了LLM作为人类被试替代品在主观审美判断中的可靠性，提供了真实人类对照，揭示了系统性偏差和排序一致性，对关注仿真效度的研究者具有重要参考价值。","inspiration":"借鉴其预注册探索性设计和多模型对比方法，可系统评估AI与人类在主观判断任务上的偏差模式。｜可迁移到信贷审批中的外貌歧视研究，或消费者对产品外观的偏好评估。｜以银行信贷员为人类被试，让MLLMs和信贷员对同一组借款人照片进行信用worthiness评分，比较评分分布和排序，并以实际贷款数据作为外部基准。"}},{"id":"2609.01867","version":1,"title":"Thinking effort aligns between humans and reasoning models in abductive reasoning","zh_title":"溯因推理中人类与推理模型的思维努力对齐","abstract":"A major question in cognitive modeling concerns the behavioral alignment between large language models and humans across linguistic and non-linguistic tasks. Unlike standard LLMs, large reasoning models (LRMs) are optimized with reinforcement learning from verifiable rewards, encouraging correct solutions to reasoning tasks rather than preference-aligned responses. Recent work (de Varda et al., 2025) investigates the cost of thinking in humans and LRMs by comparing human reaction times with model reasoning traces across a range of reasoning tasks. We isolate this alignment by turning to abductive reasoning: unlike deductive tasks, its difficulty cannot be inferred from formal structure and offers no shortcuts a model could exploit to mimic effort without genuine search, providing firmer ground for empirical claims of shared effort. We find further evidence of alignment between LRM and human reasoning effort, as well as evidence that models and humans tend to make similar errors. Finally, we show that decoding methods that let models explore multiple reasoning paths increase alignment in reasoning cost between humans and LRMs across the three models tested.","authors":["Henry Arthur"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01867","pdf_url":"https://arxiv.org/pdf/2609.01867","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","认知对齐","溯因推理"],"reason":"比较人类与推理模型在溯因推理中的思维努力，含人类反应时对照，属仿真对齐研究。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":8,"question":"人类与大型推理模型在溯因推理中的思维努力是否对齐？","design":"使用三种大型推理模型（DeepSeek-R1等）作为被试，在溯因推理任务上比较模型生成的推理链长度（token数）与人类反应时；并测试不同解码策略（如多样本搜索）对对齐程度的影响。","baseline":"人类被试在相同溯因推理任务上的反应时数据（来自de Varda et al., 2025的七项推理任务之一）。","findings":"发现LRM与人类在溯因推理中的思维努力存在对齐，且模型与人类倾向于犯类似错误。采用允许多条推理路径探索的解码方法可提高对齐程度。","reliability":"论文承认CoT可能不忠实于底层计算，且对齐并非机制性主张；通过测试多种解码策略和推理努力水平来部分回应批评。","relevance":"该研究直接比较人类与LLM在推理任务中的行为对齐，包含真实人类反应时对照，并讨论仿真失效条件，与研究者关注点高度契合，值得精读。","inspiration":"借鉴其利用任务特性（溯因推理无形式捷径）来排除模型投机取巧、增强对齐结论可信度的设计思路。｜可迁移到经济决策中的信念更新或预期形成场景，如投资者在信息不完全下的推断。｜以LLM为被试，呈现模糊经济信息（如公司公告），要求给出解释并测量推理链长度，与人类实验中的反应时和解释内容对照，检验模型是否复现人类推断努力和错误模式。"}},{"id":"2609.02277","version":1,"title":"Auditory Illusion Benchmark for Large Audio Language Models","zh_title":"大型音频语言模型的听觉错觉基准","abstract":"Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replicate human perceptual tendencies. Despite their importance, most benchmarks focus on visual illusions or general audio tasks, leaving auditory illusions underexplored. To this end, we present AIB, the first auditory illusion benchmark for LALMs, covering ten representative illusions across music, sound, and speech, each annotated for the presence of knowledge-based priors. Our methodology pairs model evaluation with controlled human listening studies, enabling direct comparison of responses. Results show systematic differences: while most LALMs remain signal-faithful on low-level acoustic illusions, several exhibit more human-like responses when linguistic or musical priors are involved, although no model matches the human perceptual profile. These findings highlight the current limitations of LALMs as cognitive models. By establishing auditory illusions as a rigorous testbed, our work offers a new perspective for probing neural black-box models and advancing understanding of auditory cognition. AIB is publicly available at https://github.com/gillosae/aib.","authors":["Hayoon Kim","Eunice Hong","Kyogu Lee"],"categories":["cs.SD","cs.AI"],"primary_category":"cs.SD","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02277","pdf_url":"https://arxiv.org/pdf/2609.02277","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1"],"tags":["听觉错觉","人类感知仿真","模型评估"],"reason":"用LALM复现人类听觉错觉，并与人类数据对照，评估模型作为认知模型的可靠性","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":9,"question":"大型音频语言模型（LALMs）在多大程度上复现人类对听觉错觉的感知倾向，能否作为人类听觉认知的模型？","design":"构建包含10种听觉错觉的基准AIB，覆盖音乐、声音和语音领域，按机制分为物理型和物理+知识型；将错觉任务转化为多项选择题，对多个LALMs进行测试，并与受控人类听力实验的结果进行对比。","baseline":"通过受控人类听力研究收集的人类对相同刺激的错觉易感性数据。","findings":"在低层声学错觉上，多数LALMs保持信号忠实，而人类表现出强错觉易感性；在涉及语言或音乐先验的错觉上，部分模型表现出更接近人类的反应，但没有模型完全匹配人类的感知特征。","reliability":"论文指出LALMs在物理型错觉上倾向于信号忠实，与人类不一致，而在知识型错觉上部分对齐，表明其错觉易感性可能源于高层先验而非共享的低层听觉处理；未讨论其他失效条件。","relevance":"该研究直接评估LALMs作为人类听觉认知模型的可靠性，与研究者关注LLM仿真人类感知和决策的核心问题高度相关，提供了模型与人类系统对比的实证证据，值得精读。","inspiration":"借鉴其构建受控刺激对（错觉与对照）和将主观感知转化为多项选择任务的方法，可用于经济金融中的主观判断仿真。｜可迁移到投资者对市场信息的感知偏差研究，如盈余公告后的漂移现象。｜以LLM为被试，呈现带有不同信息框架的财务报告（处理），测量其对未来收益的预期（结果变量），并与真实投资者调查数据对照。"}},{"id":"2607.10628","version":2,"title":"Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation","zh_title":"Anamnesis：大规模背景条件调查仿真的开源平台","abstract":"We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models. Open-source and designed for non-technical users/researchers, Anamnesis enables the prototyping and stress-testing of survey instruments on virtual populations rather than real human subjects. The platform operationalizes the recently introduced Anthology and Alterity frameworks, which use structured narrative backstories to condition model responses, within a unified web interface. It supports open-ended generation, probabilistic demographic resampling, and multimodal (image and audio) surveys. We evaluate the system through two case studies: (1) replicating segments of Pew Research Center's American Trends Panel (ATP) on political typology and biomedical issues and (2) emulating human preference in the New Yorker Caption Contest. In both cases, Anamnesis produces opinion distributions that more closely match real-world survey data than standard persona-prompting baselines, offering a transparent, reproducible, and open-source alternative to proprietary simulation services.","authors":["Song-Ze Yu","Joseph Suh","Serina Chang","David M. Chan"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-02","first_seen":"2026-07-12","revised_at":"2026-09-02","abs_url":"https://arxiv.org/abs/2607.10628","pdf_url":"https://arxiv.org/pdf/2607.10628","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B4"],"tags":["LLM仿真","调查模拟","人类数据对照"],"reason":"平台用LLM仿真调查，与真实数据对照，复现舆论分布，评估可靠性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":1,"question":"如何构建一个开源、可交互的平台，利用大语言模型和结构化叙事背景来模拟多样化人群的调查回答，并验证其与真实人类调查数据的匹配度？","design":"Anamnesis平台使用大语言模型（如Gemini 2.5 Flash）扮演虚拟受访者，通过结构化叙事背景（backstories）而非简单人口统计列表来条件化模型响应；支持概率性人口重采样、开放式生成和多模态（图像、音频）调查；在案例研究中，模拟Pew Research Center的American Trends Panel三个波次（政治类型学、生物医学、AI与人类增强）的多选题调查，以及New Yorker Caption Contest的多模态偏好选择。","baseline":"对照的真实人类数据包括：Pew Research Center American Trends Panel（ATP）三个波次的真实调查回答分布，以及New Yorker Caption Contest中基于大规模众包投票的真实人类偏好标签。","findings":"Anamnesis平台在三个ATP波次上产生的意见分布比标准人物提示（persona-prompting）基线更接近真实调查数据，在Wasserstein距离和Frobenius范数上均表现更优；在多模态New Yorker Caption Contest中，平台模拟的人类偏好与真实人类集体判断存在可测量的相关性。","reliability":"论文未明确讨论仿真失效的条件，但指出标准人物提示方法会产生刻板印象且缺乏心理深度，暗示仅依赖人口统计列表的仿真可能不可靠；同时，平台依赖LLM推理提供商，可能引入模型偏差，且案例研究仅覆盖有限主题和模态，泛化性有待验证。","relevance":"该论文直接针对用LLM进行人类仿真实验的研究，提供了开源平台和与真实调查数据对照的验证，对关注经济学实验和政策评估场景的研究者具有重要参考价值，值得阅读原文以了解平台细节和评估方法。","inspiration":"借鉴其使用结构化叙事背景而非简单人口统计列表来条件化LLM响应的方法，可提高虚拟被试的心理真实性和异质性，并采用与真实调查数据匹配的评估指标（如Wasserstein距离）来量化仿真质量｜可迁移到政策评估中的公众意见模拟，例如模拟不同社会经济背景的个体对税收改革、福利政策或公共卫生措施的态度分布，以预测试验或调查结果｜设计一个研究：使用Anamnesis平台生成具有多样化背景故事的虚拟被试，施加不同政策信息框架（如强调公平 vs. 效率）作为处理，测量其对政策支持度的选择，并与真实世界调查数据（如General Social Survey或特定政策民意调查）进行分布匹配对照，以评估仿真预测的准确性。"}},{"id":"2609.00222","version":1,"title":"LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts","zh_title":"LLM作为人口群体：社会人口学提示对谁有益，对谁有害","abstract":"Large language models (LLMs) are increasingly used as judges for subjective tasks, where annotators disagree and the relevant question is not only how accurate a judge is, but whose judgments it reproduces. Sociodemographic prompting conditions the judge on an annotator's demographic profile to align its judgments with the corresponding group's. We test whether this alignment emerges distributionally, comparing the predicted label distributions of 23 open-weight LLMs on three subjective tasks against those of real annotator groups, under three conditions: no demographic information, single-attribute profiles, and intersectional profiles over gender, age, race, and education. Three findings emerge. First, a judge prompted with no demographics is not perspective-neutral: models best reproduce the judgments of White, college-educated annotators. Second, demographic conditioning is asymmetric: it moves the judge toward majority groups and away from minority groups, most strongly on offensiveness, where intersectional profiles amplify the harm. Third, by comparing base and instruct models we identify instruction-tuning as a possible source of the asymmetry. Demographic conditioning should therefore be used with caution to estimate group judgments: conditioning moves predictions away from the reference distributions of the minority groups the method is often invoked to serve.","authors":["Daniela Occhipinti","Andrea Piergentili","Marco Guerini"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00222","pdf_url":"https://arxiv.org/pdf/2609.00222","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人口学提示","算法偏差"],"reason":"用LLM模拟不同人口群体判断，并与真实标注者分布对照，评估偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":3,"question":"在主观判断任务中，用人口统计学提示（sociodemographic prompting）让LLM模拟特定人群的判断，是否能使模型的标签分布向该人群的真实判断分布靠拢？","design":"使用23个开源权重LLM作为判断者，在三个主观任务（礼貌性、亲密性、冒犯性）上，比较三种条件：无人口统计学信息、单一属性画像（性别、年龄、种族、教育）、交叉属性画像。通过比较模型预测的标签分布与真实标注者群体的标签分布（来自DeMo数据集）来评估对齐效果。","baseline":"DeMo数据集，包含多个语料库中标注者的自报人口统计学信息（性别、年龄、种族、教育）及其对文本的评分，可构建每个文本在特定人口群体上的真实标签分布。","findings":"无人口统计学提示的LLM判断者并非视角中立，其判断最接近白人、大学学历标注者。人口统计学条件作用不对称：使判断者向多数群体靠拢、远离少数群体，尤其在冒犯性任务上，交叉画像放大了这种伤害；通过比较基础模型和指令微调模型，发现指令微调可能是这种不对称的来源。","reliability":"论文指出人口统计学提示应谨慎用于估计群体判断，因为条件作用使预测远离少数群体的参考分布；此外，效应因任务和模型而异，且指令微调模型表现出更强的不对称性。","relevance":"该研究直接评估了用LLM模拟不同人口群体判断的可靠性，并与真实标注者分布对照，揭示了仿真中的系统性偏差，对关注LLM仿真人类行为的研究者具有重要参考价值。","inspiration":"借鉴其分布级评估框架，将模型输出分布与真实群体分布比较，而非仅比较均值或多数标签，并采用交叉人口属性来检验交互效应。｜可迁移到信贷审批中的群体差异研究，如模拟不同性别、种族、教育背景的贷款审批人对同一申请的风险判断。｜以LLM作为虚拟审批人，施加不同人口统计学提示（如性别×种族），测量其对贷款申请的批准概率分布，并与真实信贷审批数据（如某银行历史审批记录中不同审批人群体）的分布进行对比，检验仿真偏差。"}},{"id":"2609.01591","version":1,"title":"StudentSim: Training LLM-based Student Simulators","zh_title":"StudentSim：训练基于LLM的学生模拟器","abstract":"AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but evidence about which guidance works for which student is sparse, slow, and costly to collect from real learners. Student simulators can provide this signal as a proxy, yet existing approaches are limited: state-tracking models fit student behavior but struggle to process explanations or corrections, while LLM role-play follows guidance fluently but does not reliably match the competence of the student being imitated. We present StudentSim, a training framework that turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization. The resulting simulators both mirror a student's own responses and update them under tutor guidance. We also introduce StudentSimEval, a standardized protocol covering 60 students across chess, second-language English writing, and mathematics, using public learner datasets with de-identified records shared for research. StudentSimEval measures behavioral fidelity (F), or how well a simulator matches a student's responses, and guidance responsiveness (R), or how readily it updates under tutor guidance, with all methods fit and evaluated on the same records. Across all three domains, StudentSim outperforms GPT-5.4 on both metrics. In chess, StudentSim reaches F=0.51 and R=0.91, compared with 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. As a proof of concept, using StudentSim as a reward model for tutor reinforcement learning produces a chess tutor that expert humans rate as more accurate, better-guided, and more personalized than a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward. Code is available at https://github.com/microsoft/StudentSim.","authors":["Ke Yang","Chenglong Wang","Michel Galley","Chandan Singh","Jeevana Priya Inala","ChengXiang Zhai","Jianfeng Gao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01591","pdf_url":"https://arxiv.org/pdf/2609.01591","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","学生模拟","行为保真度"],"reason":"用LLM模拟学生行为并与真实学生数据对照，评估行为保真度和指导响应性，属于人类…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":8,"question":"如何训练既能忠实反映个体学生能力又能响应导师指导的学生模拟器？","design":"提出StudentSim框架，先在所有学生记录上预训练基础模拟器，再对每个学生进行个性化微调；在棋类、二语写作和数学三个领域共60名学生上，用行为保真度（F）和指导响应性（R）两个指标评估模拟器，并与GPT-5.4和Maia2等基线比较。","baseline":"使用公开学习者数据集（chess、L2、math）中的真实学生记录，每个学生有训练集和留出集，所有方法在相同记录上拟合和评估。","findings":"StudentSim在所有三个领域的行为保真度和指导响应性上均优于GPT-5.4；在棋类中，StudentSim的F=0.51、R=0.91，而GPT-5.4为0.23和0.72，Maia2为0.45和0.27。将StudentSim作为奖励模型进行强化学习，训练出的棋类导师在准确性、指导性和个性化方面均优于无RL基线和用GPT-5.4模拟器奖励训练的导师。","reliability":"论文未讨论","relevance":"该研究用LLM模拟学生行为并与真实学生数据对照，评估行为保真度和指导响应性，属于人类仿真实验，且包含经济学实验和政策评估场景的潜在应用，值得阅读原文以了解其训练框架和评估协议。","inspiration":"借鉴其两阶段训练框架（先池化预训练再个体微调）和双指标评估（行为保真度与响应性），可迁移到经济金融中的个体决策模拟，如消费者选择或投资者行为。｜例如，在资产定价实验中模拟异质投资者对信息的反应，或在信贷审批中模拟不同风险偏好的申请人。｜用LLM模拟投资者，施加不同政策公告作为处理，测量其交易行为和风险偏好变化，并与真实投资者交易数据对照，评估模拟器的保真度和响应性。"}},{"id":"2609.01038","version":1,"title":"Data-Driven Persona-Conditioned Agents for A/B Test Simulation","zh_title":"基于数据驱动人物画像的智能体用于A/B测试模拟","abstract":"A/B testing is the gold standard for evaluating product changes, but each experiment requires real user traffic, engineering effort, and weeks of measurement. We propose a simulation framework that predicts A/B test outcomes using LLM-powered agents conditioned on data-driven personas grounded in real user behavioral signals. Unlike prior work that relies on synthetic or rule-based personas, our agents are constructed from anonymized behavioral data-activity patterns, engagement signals, and inferred demographics-enabling more faithful population modeling. We frame A/B test simulation as a structured question task and systematically study (i) question design formats, (ii) the impact of persona data source and domain alignment, (iii) the trade-off between per-persona behavioral depth and population diversity, and (iv) efficient population subsampling. On a benchmark of 40 A/B tests spanning two metric types, our best configuration achieves 0.75-0.90 directional accuracy depending on the test metric, demonstrating that data-driven personas are a viable path toward fast, low-cost experiment pre-screening.","authors":["Ziyad Benomar","Weronika {\\L}ajewska","Leonardo Perelli","Saab Mansour"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01038","pdf_url":"https://arxiv.org/pdf/2609.01038","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","A/B测试","人物画像"],"reason":"用LLM代理模拟A/B测试用户行为，基于真实行为数据构建persona，并与真…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":5,"question":"如何用基于真实行为数据构建的数据驱动型人格（persona）条件化LLM智能体，来预测A/B测试的方向性结果，并系统研究问题格式、人格数据来源、行为深度与多样性权衡及子采样效率。","design":"使用Claude Sonnet 4.5作为LLM，基于匿名行为数据（活动模式、参与信号、推断人口统计）生成结构化人格，将A/B测试模拟为结构化问答任务，让智能体在控制与处理变体间做出选择，预测点击率（CTR）和订阅量两个指标的方向性结果。","baseline":"40个历史A/B测试的真实结果，由置信区间推导出方向标签（正、负、可忽略），作为评估模拟准确性的基准。","findings":"最佳配置在CTR和订阅量上的方向准确率分别达到0.75和0.90，表明数据驱动人格可用于低成本实验预筛选。领域对齐的人格数据源对模拟质量至关重要，公共行为数据可媲美平台特定人格；子采样可使成本降低2倍而无质量损失。","reliability":"论文承认基准测试来自精选实验样本，不代表任何平台的全部用户或运营A/B测试基础设施；人格池偏向高参与度用户可能不具代表性；未讨论LLM模拟在更复杂决策或长期效应上的失效条件。","relevance":"该研究直接命中研究者对LLM仿真人类行为、真实数据对照、经济学实验场景的关注，提供了人格构建、问题格式和采样策略的系统性实证，值得精读以借鉴其方法并批判其局限。","inspiration":"借鉴其用真实行为数据构建人格并系统比较深度与多样性权衡、子采样效率的做法，可迁移到消费者金融决策或政策评估场景。｜可应用于信贷审批歧视研究：用银行交易数据构建不同信用评分段的人格，模拟贷款申请决策。｜设计：以真实银行客户数据构建人格，处理为不同利率或贷款条款，结果变量为是否接受贷款，用历史信贷数据中的真实接受率作为对照基准。"}},{"id":"2609.01257","version":1,"title":"Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations","zh_title":"衡量长时程人类活动模拟的行为保真度","abstract":"As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. While prior work has examined behavioral fidelity in survey responses and dialogue, longer-horizon real-world activity remains largely unexplored. We introduce a framework for evaluating behavioral fidelity in long-horizon activity simulations across temporal granularities and levels of analysis. As a case study, we collect a 43-hour multi-camera dataset of in-the-wild office activity and compare trace-derived conditioning mechanisms: persona descriptors, few-shot exemplars, and statistical transition and time-of-day priors. We find that behavioral fidelity is not uniform across metrics: statistical priors bring activity and sequence distributions closest to real behavior, yet over-fragment routines and suppress within-person variability. These findings motivate a more holistic evaluation that spans multiple metrics, temporal granularities, and levels of analysis.","authors":["Yi Fei Cheng","Fan Yang","Iremsu Bas","Koichiro Niinuma","Narishige Abe","David Lindlbauer"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01257","pdf_url":"https://arxiv.org/pdf/2609.01257","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","行为保真度","人类活动模拟"],"reason":"直接评估LLM模拟人类长期活动的行为保真度，并与真实人类数据对照，属于核心仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":6,"question":"如何系统评估大语言模型在长时间跨度人类活动模拟中的行为保真度？","design":"使用六种基于LLM的仿真方法（无人物描述控制、作者撰写的人物描述、轨迹推断的人物描述、少样本全天行为示例、统计转移先验、时间先验），在模拟办公室环境中生成五名智能体各八小时的活动轨迹，并与真实办公室活动数据集比较，测量活动分布、序列分布、时间分布、个体内变异性等指标。","baseline":"收集了一个43小时的多摄像头办公室活动数据集，包含55人的真实活动轨迹，其中5人有多天、每天数小时的纵向观察数据，作为对照基准。","findings":"统计先验使活动分布和序列分布最接近真实行为，但会导致日程过度碎片化并抑制个体内变异性。基于人物描述的方法之间差异较小，且群体层面的一致性可能掩盖个体层面的错误。","reliability":"论文指出行为保真度在不同指标、时间粒度和分析层面上并不一致，统计先验虽改善分布对齐但牺牲了日程连贯性和个体变异性，因此需要多维度评估。","relevance":"该研究直接评估LLM模拟人类长期活动的行为保真度，并与真实人类数据对照，属于核心仿真研究，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其多维度评估框架和条件机制对比设计，可迁移到经济金融中的消费者日常消费行为模拟或投资者交易行为模拟。例如，用LLM模拟投资者在交易日内的交易决策，处理为不同条件机制（人物描述、少样本示例、统计先验），结果变量为交易频率、交易时间分布和持仓变化，对照真实交易数据（如某券商脱敏交易记录）评估保真度。"}},{"id":"2609.01275","version":1,"title":"The Constitutional Coverage Trilemma in AI Governance","zh_title":"AI治理中的宪法覆盖三难困境","abstract":"Frontier AI systems function as \\emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human demand. Combining a paraphrase-controlled audit of the as-shipped default constitutions of $23$ frontier LLM archetypes with a pairwise-tradeoff study of $1{,}649$ US participants on the same instrument, we report three facts. \\emph{Demand is broad}: it spans all five values, with the largest constituency under one-third. \\emph{Supply is narrow and drifting}: the $23$-archetype hull occupies ${\\sim}2\\%$ of the demand hull under conservative noise-matched estimation ($0.10\\%$ at full audit precision), no archetype puts helpfulness or autonomy first ($37\\%$ of users are constitutionally homeless), and across six model families autonomy decreases in $5/6$, equity increases in $5/6$, and safety increases in $4/6$, with monotone within-family version trends (order-permutation $p = 0.013$) and the autonomy decline concentrated in scenarios where safety is not at stake. The drift's importance is directional: \\emph{away} from a value already undercovered, mechanically worsening the welfare floor for the least-served users. \\emph{The fix is sparse}: a $2$-vertex menu $\\{e_{\\mathrm{HON}}, e_{\\mathrm{AUT}}\\}$ beats the full $23$-archetype frontier by $47\\%$ on mean regret (CI $[43\\%, 52\\%]$); three vertex additions cut mean/worst-group regret by up to $81\\%$/$64\\%$. We formalize these findings as a budgeted-pluralism trilemma, show the binding regime is empirically realized, and verify the conclusions are robust to distance-based welfare and to degraded routing. The instrument and audit harness are described in full in the appendices.","authors":["Natalija Mitic","Soona Sedahmed A. O.","Mamadou Selly Ly","Moustapha Cisse"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01275","pdf_url":"https://arxiv.org/pdf/2609.01275","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","价值观对齐","人类对照"],"reason":"用LLM审计宪法价值观并与1649名人类对照，评估供需匹配与偏差，直接仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":7,"question":"前沿LLM的宪法价值观供给是否覆盖了人类用户的宪法价值观需求？","design":"本研究不是用LLM仿真人类，而是将LLM本身作为被审计对象：对23个前沿LLM原型进行受控改写审计，测量其在安全、帮助性、诚实、自主、公平五维价值上的隐含排序；同时用同一套成对权衡问卷测量1649名美国参与者的价值偏好，作为需求侧基准。","baseline":"1649名美国参与者在同一五维价值成对权衡问卷上的偏好分布。","findings":"人类需求广泛，五种价值都有显著支持者，最大群体占比不足三分之一；LLM供给狭窄且漂移，23个原型仅覆盖需求凸包的约2%，没有原型将帮助性或自主性置于首位，37%的用户在宪法意义上无家可归，且跨版本漂移方向是自主性下降、公平和安全上升，远离已覆盖不足的价值。","reliability":"论文指出线性福利模型是对供给方最有利的假设，实际福利可能更低；审计精度受改写控制和模型间差异解决阈值影响；未讨论LLM审计结果与真实部署行为之间的差距。","relevance":"该研究直接测量LLM的价值观分布并与人类对照，属于用LLM进行人类仿真实验的批判性工作，揭示了仿真在价值观匹配上的系统性偏差，值得精读。","inspiration":"借鉴其将LLM作为制度性主体进行审计并与人类偏好对照的方法，可迁移到经济金融中的算法决策场景，如信贷审批、保险定价或投资建议中的公平与效率权衡。｜设计一个实验：用多个LLM扮演信贷审批员，施加不同价值取向的提示（如强调公平或效率），测量其审批决策中的种族或性别差异，并与真实银行信贷数据或人类审批员的决策分布进行对照，评估LLM仿真的偏差。"}},{"id":"2609.00009","version":1,"title":"Toward a social psychology of AI: language-model agents reproduce human-like minimal-group bias","zh_title":"迈向AI社会心理学：语言模型智能体再现类人的最小群体偏差","abstract":"Language-model agents now interact in groups, but evaluations that probe memorised stereotype content or use models to simulate people leave this social behaviour unmeasured. We adapt the minimal-group paradigm---social psychology's classic test of intergroup bias---into a controlled probe: an agent distributes points among anonymous peers bearing only an arbitrary group label. Across four reasoning models, mere categorisation into meaningless groups elicited in-group favouritism that vanished under a group-blind control and was concentrated in the numerical minority: minority deciders over-allocated to their own group relative to their numbers, majority deciders allocated close to proportionally, and the asymmetry closed at equal group sizes. Disabling reasoning in one model did not remove the disposition---if anything it grew---but nearly erased the minority-majority asymmetry, implicating deliberation in where bias concentrates rather than whether it appears. These open-weight reasoning models reproduce the behavioural signature of human intergroup discrimination, independent of stereotype content, and social psychology's theories and methods offer a paradigm for measuring and governing AI's social behaviour.","authors":["Messi H. J. Lee"],"categories":["physics.soc-ph","cs.CY"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00009","pdf_url":"https://arxiv.org/pdf/2609.00009","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM仿真","社会心理学","群体偏差"],"reason":"用LLM复现人类最小群体偏差，并与经典社会心理学实验对照，直接仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":2,"question":"语言模型智能体在最小群体范式中是否复现人类的内群体偏差，且该偏差如何随群体相对规模变化？","design":"使用四个开源推理模型（如 DeepSeek-R1、Qwen 等）作为被试，将其置于最小群体范式中：智能体被随机分配到无意义标签（如“组A”或“组B”）的群体，并需要在匿名同伴之间分配点数。处理变量包括群体标签的有无（组盲对照）和群体相对规模（少数/多数/均等）。结果变量是分配给内群体成员的点数比例。","baseline":"人类经典最小群体实验（Tajfel 等）及元分析结果，特别是关于少数群体成员比多数群体成员表现出更强内群体偏爱的发现。","findings":"四个推理模型在无意义群体分类下均表现出内群体偏爱，且该偏爱在组盲对照中消失；少数群体成员过度分配给内群体，多数群体成员分配接近比例，群体规模相等时不对称消失。禁用推理后，内群体偏爱未消失甚至增强，但少数-多数不对称几乎消失，表明推理影响偏差的分布而非存在。","reliability":"论文未明确讨论失效条件，但指出仅测试了开源推理模型，未涵盖闭源模型或指令微调模型；且实验为一次性分配任务，未涉及长期互动或真实后果。","relevance":"该研究直接以LLM为被试复现社会心理学经典实验，并与人类基准对照，属于人类仿真研究，且揭示了群体结构对偏差的影响，值得精读以了解仿真在群体行为中的有效性。","inspiration":"借鉴其最小群体范式的对照设计（组盲条件）和群体规模操纵，以分离纯粹分类效应｜可迁移到信贷审批中的群体歧视研究，例如测试AI信贷员是否对少数群体申请人有内群体偏爱｜用LLM扮演信贷审批员，随机分配其所属“银行组”，处理为申请人所属组（内/外群体）和群体规模，结果变量为贷款批准率和额度，与人类信贷员的历史审批数据对照。"}},{"id":"2609.00345","version":1,"title":"Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment","zh_title":"LLM了解你的社区吗？审计LLM先验用于社区级移动性预测与结构对齐","abstract":"Human mobility is central to urban planning, transportation, public health, and emergency response, yet fine-grained trajectory data are often proprietary, restricted, and privacy-sensitive. Large language models (LLMs) offer a potential alternative by generating plausible mobility traces and predicting individual movement, but their ability to infer aggregate neighborhood-level mobility remains unclear. We evaluate zero-shot LLMs on Census Block Group-level mobility prediction across four U.S. metropolitan areas using anonymized Cuebiq data to construct point-level, trajectory-level, and temporal mobility outcomes, paired with sociodemographic and built-environment predictors. We compare LLM predictions with supervised baselines and introduce a directional alignment analysis to test whether LLM-implied predictor effects agree with empirical OLS and Jonckheere-Terpstra trends. Supervised models achieve 0.580 average accuracy, compared with 0.435 for the best LLM, with spatial extent outcomes showing the strongest predictability but also the largest LLM-baseline gaps. Directional analysis shows that LLMs often rely on coarse, stable predictor-level priors that remain similar across outcomes and cities, including asymmetric treatment of protected-group predictors. Overall, LLMs can partially recover aggregate mobility patterns from urban context, but their predictions should not be treated as structurally grounded without auditing empirical alignment and potential bias.","authors":["Saad Mohammad Abrar","Eesha Kurella","Arnav Dadarya","Naman Awasthi","Kazi Tasnim Zinat","Vanessa Frias-Martinez"],"categories":["cs.LG","cs.CY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00345","pdf_url":"https://arxiv.org/pdf/2609.00345","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类移动性","偏差审计"],"reason":"用LLM预测社区级人类移动性，并与真实数据对照，审计其偏差与对齐。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":4,"question":"LLM能否仅凭社区的社会人口与建成环境特征，直接推断出社区层面的聚合人类移动性结果，并且其预测是否与经验数据中的结构性关系对齐？","design":"使用零样本LLM（如GPT-4等）作为预测模型，输入美国四个大都市区人口普查区块组（CBG）的社会人口和建成环境特征，预测三类聚合移动性结果（点级、轨迹级、时间级），并与监督学习基线比较。","baseline":"使用Cuebiq提供的匿名化手机定位数据，聚合到人口普查区块组级别，构建真实移动性结果，作为LLM预测的对照基准。","findings":"监督学习基线平均准确率为0.580，最佳LLM为0.435，LLM在空间范围结果上差距最大。方向对齐分析显示LLM依赖粗粒度、稳定的先验，且对受保护群体预测因子处理不对称。","reliability":"论文指出LLM预测不应被视为结构上可靠，除非经过经验对齐和潜在偏差审计；LLM的先验在不同结果和城市间保持不变，可能反映浅层启发式或刻板印象。","relevance":"该研究直接评估LLM作为人类移动性预测代理的可靠性，并与真实大规模人类数据对照，属于批判性仿真研究，对关注LLM仿真偏差和失效条件的研究者有参考价值。","inspiration":"借鉴其方向对齐分析方法，通过比较LLM隐含的预测因子效应与经验回归趋势来审计结构一致性。｜可迁移到信贷审批歧视研究，用LLM模拟信贷员决策并检验其对种族、性别等受保护特征的敏感度。｜以LLM作为虚拟信贷员，输入申请人特征（含受保护属性），预测贷款批准概率，与真实信贷数据（如HMDA）中的批准率和歧视模式进行对照。"}},{"id":"2609.00608","version":1,"title":"Investigating Assistant Bias in LLM User Simulators Using a Role Vector","zh_title":"使用角色向量研究LLM用户模拟器中的助手偏差","abstract":"LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit \"assistant bias,\" a tendency to cooperate and pursue task goals. They rarely reproduce the frustration or disengagement that real users exhibit, compromising evaluation validity. Prior work outlines that this bias is baked in during model training, which role-playing prompts fail to override. We analyze this bias from model activations, extracting a user role vector by contrasting how the model represents user versus assistant perspectives on the same dialogue. We observe two findings: (i) the user direction is identifiable in activations, elicits user-like behaviors, and captures characteristics distinct from assistant traits; and (ii) although user-role activation associates with simulation realism and steering strengthens it, it can exaggerate user behaviors and override individual user profiles. Together, our findings provide a representation-level analysis of LLM user simulators, confirming that assistant bias is structurally identifiable and that user behavior can be directionally analyzed.","authors":["Daeheon Jeong","Yoonjoo Lee","Eugene Choi","Sinie van der Ben","Juho Kim"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00608","pdf_url":"https://arxiv.org/pdf/2609.00608","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B4"],"tags":["LLM用户模拟器","助手偏差","仿真有效性"],"reason":"研究LLM用户模拟器的助手偏差，评估仿真有效性，批判性指出失效条件，可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":10,"question":"大语言模型用户模拟器中的助手偏差是否在模型激活空间中可识别，并能否通过角色向量操控来改善模拟真实性？","design":"使用Qwen 3.5 9B等模型，通过对比同一对话中用户与助手视角的反思激活，提取用户角色向量；施加该向量进行激活引导，测量对沟通风格、行为反应及多轮交互真实性的因果影响。","baseline":"SimulatorArena基准中的真实用户交互数据，以及RealUserSim中的真实用户模拟日志。","findings":"用户角色方向在激活空间中可识别，能引发更用户化的行为，并与助手特质在几何上负相关；激活引导虽能提升写作风格相似性，但会夸大用户行为并掩盖个体用户画像。","reliability":"论文承认主要基于Qwen 3.5 9B，跨模型验证有限；评估依赖代理指标而非直接人类判断；基准场景较窄，仅限数学辅导和部分对话。","relevance":"该研究直接针对LLM用户模拟器的核心偏差问题，提供了表示层面的机制分析和批判性评估，对关注仿真可靠性与失效条件的研究者具有重要参考价值。","inspiration":"借鉴其通过对比激活提取角色向量并进行因果引导的方法，可用于在LLM模拟中分离特定行为倾向｜可迁移到消费者金融决策模拟，如信贷申请中的风险偏好或政策反应｜用LLM模拟消费者，施加“风险厌恶”或“耐心”角色向量，测量其信贷选择行为，并与真实信贷申请数据或实验数据对照。"}},{"id":"2609.00565","version":1,"title":"Aligned but Flattened: Analyzing the Trade-off between Cultural Alignment and Diversity in LLMs","zh_title":"对齐但扁平化：分析LLMs中文化对齐与多样性之间的权衡","abstract":"Cultural fine-tuning has become the de facto paradigm for building culture-aware large language models (LLMs), yet existing optimization exclusively for alignment scores provides an incomplete portrait of cultural fidelity by systematically obscuring inherent cultural diversity. This unidimensional evaluation lens prompts a fundamental question: do models genuinely perceive distinct cultural nuances, or do they merely memorize dominant cultural values? To address this, we propose a synergistic evaluation framework that jointly formalizes cultural alignment and diversity. Through extensive benchmarking of six mainstream LLMs on the World Values Survey, this framework uncovers a systematic and critical trade-off: the pursuit of cultural alignment consistently incurs an acute expense of diversity, leading to severe \"cultural flattening.\" Investigating this behavioral shift, we demonstrate that these superficial alignment gains stem from models artificially anchoring to dominant majorities, converging onto a monolithic response pattern that wipes out the heterogeneous distributions inherent to human groups. Crucially, our mechanistic analysis suggests that this diversity collapse is not merely a behavioral anomaly but more likely a structural consequence of the low-rank bias inherent in neural network optimization. Therefore, our findings expose the limitations of current post-training paradigms and call for a shift toward alignment objectives that preserve cross-cultural pluralism.","authors":["Jingshen Zhang","Shaoyang Xu","Wenxuan Zhang"],"categories":["cs.SI","cs.CL"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00565","pdf_url":"https://arxiv.org/pdf/2609.00565","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["文化仿真","算法保真度","价值观调查"],"reason":"评估LLM文化对齐与多样性，使用世界价值观调查真实数据对照，揭示仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":9,"question":"文化微调在提升LLM文化对齐度的同时，是否以牺牲文化多样性为代价？","design":"对六种主流LLM在基于世界价值观调查的文化特定数据上进行微调，测量微调前后模型在文化对齐度和文化多样性上的变化，并分析其行为模式与内部表征。","baseline":"世界价值观调查的真实人类回答数据，按社会人口群体分组。","findings":"文化微调一致地提高了对齐度，但显著降低了行为多样性，导致“文化扁平化”。这种多样性损失源于模型锚定主流多数群体，并可能由神经网络优化的低秩偏差结构性导致。","reliability":"论文指出当前对齐目标仅关注聚合相似度，忽视了文化多样性，且微调受限于低秩子空间，可能边缘化少数文化；但未系统讨论其他失效条件。","relevance":"该研究直接评估LLM仿真人类文化价值观的可靠性，揭示了对齐与多样性的权衡，对关注仿真偏差和真实数据对照的研究者具有重要参考价值。","inspiration":"借鉴其联合测量对齐与多样性的评估框架，并利用真实调查数据作为基准。｜可迁移到经济金融领域的文化差异研究，如跨文化消费偏好、金融风险态度或政策接受度的仿真。｜以LLM模拟不同文化背景的消费者，施加文化微调处理，测量其对金融产品的偏好分布，并与世界价值观调查或实际消费数据对照，检验对齐与多样性权衡。"}},{"id":"2609.01519","version":1,"title":"When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation","zh_title":"当护栏看似有效：LLM智能体商业评估中的构念效度失效","abstract":"Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. We audit this risk in a multi-turn buyer--seller testbed for configurable hotel transactions. An initial implementation reported welfare gains from two marketplace guardrails of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B--14B ladder. It also gave guarded and unguarded agents different offer schemas and choice procedures. Holding the schema and buyer chooser fixed changes the paired contrasts to +7.2, -13.9, and +23.8. The four largest 14B single-generation effects averaged +229; after three generations per profile-condition, they averaged +37.6 (95% bootstrap interval [-34.2, 109.3]), while generation residuals account for 49.9% of variation in this post-hoc probe. A seller-incentive check is non-monotone: increasing profit pressure produces less profit than the default seller prompt. Scripted positive controls show why this matters. A profit-maximizing seller already attains first-best welfare, so guardrails mostly redistribute and reduce welfare; they create welfare only when the seller is explicitly programmed to force inefficient bundles. We contribute a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting, and returning INVALID or INCONCLUSIVE before substantive policy claims. In our case, the original estimate is INVALID under protocol isolation, while the controlled study remains INCONCLUSIVE under incentive validity and stochastic stability. The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks.","authors":["Peiying Zhu","Sidi Chang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01519","pdf_url":"https://arxiv.org/pdf/2609.01519","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A2","A4","B4"],"tags":["LLM仿真","构念效度","市场模拟"],"reason":"评估LLM市场仿真中构念效度失效，提出验证框架，批判性指出仿真失效条件，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":11,"question":"LLM智能体市场仿真中，平台护栏的福利效应是否真实存在，还是源于实现脚手架（offer schema与choice procedure）的差异？","design":"用Qwen2.5-Instruct（1.5B、3B、14B）同时扮演买家和卖家，在酒店交易多轮对话中施加两种护栏（阻止买家信息、强制可选组件不捆绑），测量福利变化；通过统一schema和chooser、重复生成、激励操纵检查、脚本化正控制等审计设计来检验构念效度。","baseline":"无对照","findings":"原始实现中护栏的福利增益在统一脚手架后大幅缩水甚至符号反转（如3B从+35.0变为-13.9），14B效应不显著；激励检查显示更强的利润指令并未单调提高利润，脚本化正控制表明护栏仅在卖家被编程为强制低效捆绑时才创造福利。","reliability":"论文承认合成档案中43/60的负外部选项效用使接受更容易，不具代表性；重复生成探针是事后选择，受赢家诅咒影响；激励检查样本量小，不足以排序提示；整体结论受限于特定模型家族和酒店场景。","relevance":"该研究直接针对LLM仿真在经济学评估中的构念效度问题，提供了系统的失效诊断框架和审计协议，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其构念效度审计协议，将处理与实现脚手架分离，并通过脚本化正控制和重复生成来检验效应稳健性｜可迁移到政策评估中的市场设计仿真，如平台监管、拍卖机制或价格歧视策略的福利分析｜用LLM扮演消费者和商家，施加某种政策处理（如信息屏蔽或价格上限），测量交易价格、成交率和福利，并与真实电商平台或实验数据对照，同时统一对话协议并重复生成以评估方差。"}},{"id":"2606.30085","version":2,"title":"Tastes without distinction: silicon samples and the synthetic construction of tastes","zh_title":"无差别的品味：硅样本与品味的合成建构","abstract":"Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produce reasonable approximations of cultural taste remains an open empirical question that becomes more urgent by the day, with market research companies already offering provisional 'synthetic' survey panels and the contamination of standard survey data from LLM-generated responses. In this study, we build on past work on silicon sampling by extending considerations of their ecological, relational, and positional fidelity in the doomain of cultural tastes. We use large-language models from OpenAI, Anthropic, and DeepSeek to produce 554,940 silicon surrogates of survey respondents from the Survey of Public Participation in the Arts (SPPA). We find these silicon surrogates' tastes to be highly stylized facsimiles of human tastes. First, silicon samples are super-omnivorous with a systematic postive-bias for liking. These individual-level bias of silicon samples are not well-explained by the WEIRD-bias often discussed in the literature. Second, the complex relationality in real taste structures is completely distorted among silicon samples. Third, very little of the known cultural alignment between tastes and social space are preserved. Silicon samples juvenilize age-taste associations, resurrect anachronistic class-taste associations, and caricaturize gender- and race-taste associations. Key words: AI, taste, consumption, culture, silicon sampling, meta-analysis.","authors":["Xiangyu Ma","Mengmi Zhang","Shannon Ang","Minne Chen"],"categories":["cs.CL","econ.GN","q-fin.EC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-06-29","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2606.30085","pdf_url":"https://arxiv.org/pdf/2606.30085","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["硅采样","文化品味仿真","算法保真度"],"reason":"用LLM生成硅样本模拟文化品味调查，并与真实SPPA数据对照，评估仿真偏差与失…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":1,"question":"大语言模型生成的硅样本能否忠实再现人类文化品味及其与社会空间的关联？","design":"以2012年美国公众参与艺术调查（SPPA）为参照，将每个受访者的人口统计特征转化为角色提示，输入OpenAI、Anthropic和DeepSeek的大语言模型，为每个人类受访者生成多个硅样本（共554,940个），比较硅样本与人类样本在品味分布、品味结构关系及品味与社会空间关联上的保真度。","baseline":"2012年SPPA调查中的人类受访者数据，并使用SPPA的自举重采样作为人类数据内部一致性的基准。","findings":"硅样本表现出超级杂食性和系统性正向偏好，且这种偏差不能用文献中常讨论的WEIRD偏差解释；真实品味结构中的复杂关系在硅样本中被完全扭曲，品味与社会空间的已知关联几乎未被保留，年龄-品味关联被年轻化，阶级-品味关联被复活为过时模式，性别和种族-品味关联被漫画化。","reliability":"论文指出硅样本的个体层面偏差（如超级杂食性和正向偏好）不能由WEIRD偏差解释，且其关系结构和与社会空间的关联严重失真，表明在文化品味领域硅采样保真度很低，但未明确讨论具体失效条件或局限性。","relevance":"该研究直接评估LLM仿真人类文化品味调查的可靠性，并与真实SPPA数据对照，发现系统性偏差和结构失真，对关注LLM仿真在社会科学中有效性的研究者具有重要参考价值，值得阅读原文以了解具体偏差模式和评估框架。","inspiration":"借鉴其将人口统计特征转化为提示、生成大量硅样本并与真实调查数据对照的仿真设计，以及从生态、关系和位置保真度三个维度评估偏差的方法。｜可迁移到消费者偏好调查、市场细分或文化消费的经济学研究中，例如用LLM模拟不同人口群体的品牌偏好或娱乐消费选择。｜设计：以某消费者支出调查（如美国消费者支出调查CE）为基准，提取受访者人口特征生成提示，用多个LLM生成硅样本，让其回答关于品牌偏好或娱乐活动参与的问题，比较硅样本与真实人类在偏好分布、偏好结构及与社会经济地位关联上的差异，并检验硅样本是否复现已知的消费分层模式。"}},{"id":"2608.03044","version":2,"title":"Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation","zh_title":"仿真还是估计？基础与后训练语言模型在意见模拟中的不同优势","abstract":"Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity. We propose that much of this conflict stems from conflating two distinct tasks. We call the first task emulation, in which models generate individual responses that aggregate into a population distribution. We call the second task estimation, in which models directly predict the population distribution. Evaluating six matched base and post-trained models on the Pew American Trends Panel, we find that base models are the stronger emulators: they produce response distributions closer to human ground truth and better preserve demographic structure. Post-trained models are generally the stronger estimators, producing more accurate distributional predictions when asked directly. We argue that model selection for human simulation should be guided by whether the task requires generating text or predicting distributions.","authors":["Seth Grief-Albert","Jessica Bo","Difan Jiao","Ashton Anderson"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-05","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.03044","pdf_url":"https://arxiv.org/pdf/2608.03044","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","意见模拟","算法保真度"],"reason":"直接研究LLM仿真人类意见，区分仿真与估计任务，使用真实调查数据对照，并指出模…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":2,"question":"在意见仿真中，基础模型与后训练模型在仿真（emulation）与估计（estimation）两种任务上表现有何差异？","design":"使用三对匹配的基础与后训练模型（Qwen3-14B、Olmo-3-7B、Olmo-3-32B）在Pew美国趋势面板59个经济意见问题上进行仿真：仿真任务通过开放式生成个体回答并聚合，估计任务直接预测总体分布；施加七种人口学条件（无条件、民主党、共和党、非常自由派、非常保守派、高收入、新教徒），以总变差距离和Wasserstein距离衡量分布保真度。","baseline":"Pew Research Center American Trends Panel Wave 54 的真实人类调查数据，包含59个四选项经济意见问题及人口学子群体分布。","findings":"基础模型在仿真任务上始终优于后训练模型，生成的分布更接近人类真实分布且更好地保留人口学结构；后训练模型在直接估计总体分布时通常更准确。","reliability":"论文未讨论","relevance":"直接研究LLM仿真人类意见，区分仿真与估计任务，使用真实调查数据对照，并指出模型选择应基于任务类型，对关注LLM仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其区分仿真与估计任务、使用开放式生成避免位置偏差、并用真实调查数据做基准对照的方法｜可迁移到经济预期形成或消费者信心调查的仿真，如模拟不同收入群体对通胀预期的分布｜用基础模型模拟个体受访者回答开放式经济预期问题，聚合后与密歇根消费者调查的真实分布对比，同时用后训练模型直接估计分布，比较两种范式的准确性。"}},{"id":"2608.29455","version":1,"title":"Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human Surrogates","zh_title":"项目均值替代：为何更丰富的人物数据未能提升LLM作为人类替代品的表现","abstract":"LLMs are increasingly used as human surrogates, often on the premise that richer persona data could make them substitutes or exploratory tools for specific individuals. We test this premise across four datasets covering more than 400,000 participants and more than 6,000 survey items and experimental outcomes. LLMs perform well at the aggregate level: their average responses closely align with average human responses to the same items. But this success largely reflects predicting each item's average human response. Once each item's human mean is removed, LLM predictions explain only 3.05% of the remaining respondent-specific variation, far below the 53.6% human test-retest benchmark. Richer personas, model variants, and fine-tuning do not close this gap. In variance analyses, once item means are removed, the reliable remaining signal is person-by-item. It captures how a respondent departs from the mean on a particular item and is about 8.9x larger than the stable person effect. Persona data encode the respondent, but not this item-specific deviation. LLM responses also compress human response distributions, using less spread, fewer response categories, and distorted distributional shapes. We call this pattern item-mean surrogacy. Current LLM surrogates can approximate item averages, but not the distributions or respondent-specific deviations needed to replace individual humans. We propose four empirical tests for LLM-based human-surrogate claims.","authors":["Daehwan Ahn","Chengfeng Mao","Dokyun Lee"],"categories":["cs.CL","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29455","pdf_url":"https://arxiv.org/pdf/2608.29455","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM仿真","人类替代","算法保真度"],"reason":"直接评估LLM作为人类替代品的可靠性，使用大规模人类数据对照，发现仅能预测项目…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":4,"question":"LLM 作为人类替代品时，更丰富的人物角色数据能否提高其对个体层面回答的预测能力？","design":"使用四个数据集（Megastudy、Survey、SocSci210、ANES），覆盖超40万参与者和6000多个调查项目，用多种LLM模型和提示方法（包括丰富人物角色、微调）生成对相同项目的回答，并与真实人类回答对比，测量项目均值、分布和个体偏差的恢复程度。","baseline":"真实人类数据：Megastudy 和 Survey 数据集包含同一批参与者的 LLM 人物角色和真实回答，并提供人类重测信度（53.6%）作为个体预测上限；SocSci210 和 ANES 提供不同领域的人类回答。","findings":"LLM 在聚合层面表现良好，平均回答与人类平均回答高度一致，但去除项目均值后，LLM 预测仅解释 3.05% 的个体变异，远低于人类重测基准 53.6%。丰富人物角色、模型变体和微调均未缩小差距；方差分析显示稳定的人-项目交互效应是主要信号（44.0%），而人物角色数据仅编码稳定的人主效应（4.9%），无法捕捉项目特定偏差。","reliability":"论文指出 LLM 替代品仅能近似项目均值，无法恢复分布或个体偏差，因此不能替代个体人类；其验证通常停留在聚合层面，存在生态谬误风险。论文未讨论其他失效条件，但强调需要个体级保真度的应用场景（如个性化推荐、政策评估）会失效。","relevance":"该研究直接评估 LLM 作为人类替代品的可靠性，使用大规模真实人类数据对照，发现仅能预测项目均值而无法捕捉个体变异，对关注仿真有效性和偏差的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其将预测误差分解为项目效应、人主效应和人-项目交互效应的方法，并利用人类重测信度区分稳定信号与随机误差，可迁移到经济金融领域的个体决策预测（如消费者跨期选择、投资者风险偏好）。｜可应用于信贷审批中的个体违约风险预测或资产定价实验中的个体风险偏好测量。｜设计：以真实信贷申请人或实验参与者为被试，构建 LLM 人物角色并让其预测个体在特定金融决策任务中的选择（如是否接受高风险贷款），处理为不同人物角色丰富度（仅人口统计 vs. 完整问卷），结果变量为个体选择与项目均值的偏差，对照真实人类选择数据，并计算人类重测信度作为上限。"}},{"id":"2608.30033","version":1,"title":"\"Act Like a 5th Grader\" is Not Enough: Bounding Knowledge in LLM-Based User Simulators","zh_title":"“像五年级学生一样行动”还不够：在基于LLM的用户模拟器中界定知识","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior but frequently fail to exhibit realistic cognitive constraints, suffering from a \"superhuman bias.\" Using a dataset of over 71,000 reading comprehension responses from 2,359 primary-school students (grades 4--6), we demonstrate that standard persona prompting yields near-perfect, deterministic performance, failing to capture the natural variance of developing readers. To address this, we introduce the Cognitively Bounded User Simulator (CBUS), an architectural framework that explicitly models the restricted working memory of young readers through an episodic bottleneck. Within this framework, we formalize two distinct test-taking strategies to emulate different reading behaviors. Our evaluation shows that explicitly modeling cognitive bounds significantly narrows the simulation gap across multiple LLM backbones, demonstrating that enforcing architectural constraints is more effective for high-fidelity simulation than simply scaling raw model capabilities.","authors":["Krisztian Balog","Arild Michel Bakken"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30033","pdf_url":"https://arxiv.org/pdf/2608.30033","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知边界","人类数据对照"],"reason":"用LLM模拟学生阅读理解，有真实学生数据对照，并解决仿真偏差","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":5,"question":"如何通过显式建模认知边界（工作记忆容量限制）来提高LLM模拟小学生阅读理解行为的保真度？","design":"使用LLM模拟4-6年级小学生，通过两种方式施加处理：标准角色提示（如“扮演五年级学生”）和提出的CBUS框架（两阶段推理管道，编码阶段限制提取的文本命题数量，执行阶段仅基于受限记忆痕迹回答问题）。结果变量为学生在阅读理解题目上的回答（客观题，包括单选、判断、多选），并与真实学生数据进行对比。","baseline":"来自挪威教育系统的2,359名4-6年级学生对156篇文本和750道阅读理解题目的71,000多条真实回答。","findings":"标准角色提示导致LLM表现出超人类偏差，几乎完美且确定性地回答问题，无法捕捉真实学生的发展性差异。CBUS框架通过显式建模工作记忆瓶颈，显著缩小了模拟与真实学生之间的差距，且该效果在多个LLM骨干模型上一致。","reliability":"论文未明确讨论失效条件，但指出标准角色提示在受控环境中不足，且CBUS框架基于工作记忆容量限制的通用认知理论，可能不适用于其他认知过程或更复杂任务。","relevance":"该研究直接针对LLM仿真中的超人类偏差问题，提供了在受控环境中量化偏差和通过架构约束改进仿真的方法，对关注仿真可靠性和偏差的研究者具有重要参考价值。","inspiration":"值得借鉴的是通过架构约束（而非仅提示）来模拟认知限制，并利用大规模真实数据作为对照基准。｜可以迁移到经济金融中需要模拟有限理性或信息处理约束的场景，如消费者对复杂金融产品的选择、投资者对信息的有限关注。｜设计一个实验：用LLM模拟投资者阅读财报后做出投资决策，处理组为施加工作记忆瓶颈的CBUS框架，对照组为标准角色提示，结果变量为投资选择，对照真实投资者在类似实验中的数据。"}},{"id":"2608.28615","version":1,"title":"Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey","zh_title":"韩国合成人面板在数字与AI服务使用上的分布效度与校准：基于韩国媒体面板调查的二手数据验证","abstract":"Synthetic personas based on large language models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet systematic validation outside English-speaking contexts remains scarce. This secondary-data study evaluates how well a Korean synthetic persona panel (NVIDIA Nemotron-Personas-Korea), conditioned into Gemini 3.5 Flash (primary) and EXAONE (comparison), reproduces digital and AI service-use distributions from the KISDI Korea Media Panel Survey. Sex-and-age-stratified panels of about 8,000 personas per model answered the survey's own items - eight service-use indicators and eight innovativeness and acceptance constructs - and were compared against weighted survey estimates. The overall mean absolute error (MAE; RQ1) was 15-19 percentage points (pp), with binary item-mean correlations of 0.69-0.90. Segment error (RQ2) across five demographic axes was 15-19 pp, with between-group gaps up to 52.4/36.2 pp (Gemini/EXAONE). Errors followed model-specific signatures: an age stereotype with low anchoring (Gemini) versus an acquiescence-consistent level bias (EXAONE). Reference-year analysis was consistent with temporal misalignment driving most generative-AI overestimation, whereas short-form underestimation was framing-sensitive. Holdout calibration on 30% of the real data (RQ3) roughly halved sex-by-age cell MAE (18.9->8.6, 15.9->6.7 pp) - yet direct estimation from the same real subsample was far more accurate (3.6 pp), and the correction did not transfer across time. The calibrated panel retained an advantage only under extremely scarce real data (about 100 responses) and, for one model, for unobserved segments. Persona-narrative conditioning beat demographic-only conditioning, but neither surpassed simple real-data baselines. Synthetic panels are thus not survey substitutes; their value is diagnostic, with operational use confined to settings lacking real data.","authors":["Howard Kim","Keun Tae Cho"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28615","pdf_url":"https://arxiv.org/pdf/2608.28615","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","合成人面板","外部效度"],"reason":"用LLM合成韩国人面板，复现数字服务使用分布，并与真实调查数据对照，评估校准与…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":3,"question":"韩国合成人面板在多大程度上能复现真实调查中的数字与AI服务使用分布，以及小样本校准能否改善其准确性？","design":"使用 NVIDIA Nemotron-Personas-Korea 合成人数据，按性别和年龄分层抽取约8000个角色，分别通过 Gemini 3.5 Flash 和 EXAONE 模型生成对韩国媒体面板调查问卷的回答，测量八个数字服务使用指标和八个创新性与技术接受度构念，并与加权调查估计值比较。","baseline":"KISDI 韩国媒体面板调查 2024 年波次（约8693人）的加权估计值，以及 2023 和 2025 年波次用于时间转移检验。","findings":"合成面板的总体平均绝对误差为15-19个百分点，二元指标均值相关系数0.69-0.90；误差呈现模型特异性偏差（Gemini 年龄刻板印象、EXAONE 默许偏差），且校准虽能减半误差，但直接使用少量真实数据更准确，校准无法跨时间转移。","reliability":"论文明确指出合成面板不能替代调查，其价值仅在于诊断，且校准仅在真实数据极度稀缺（约100份）或存在未观测群体时才有优势；误差受时间错位、题项框架和响应风格影响，校准无法跨波次转移。","relevance":"该研究直接检验了LLM合成样本在非英语情境下的分布效度，并与真实调查数据严格对照，对关注仿真可靠性和偏差条件的研究者具有重要参考价值，值得精读原文以了解具体误差结构和校准方法。","inspiration":"借鉴其分层抽样生成合成样本并与真实调查对照的验证框架，以及基于偏差签名的小样本校准方法｜可迁移到消费者金融行为调查（如数字支付采用、金融科技接受度）或政策评估中的态度测量｜以合成人面板模拟不同人口群体的金融决策，施加政策信息处理，测量采用意愿，并与真实家庭金融调查数据（如中国家庭金融调查）对照，检验分布一致性和校准效果。"}},{"id":"2608.30522","version":1,"title":"Tariff Threats, Macroeconomic Expectations, and Policy Communication Strategies: Experiments Based on a Multi-Agent System","zh_title":"关税威胁、宏观经济预期与政策沟通策略：基于多智能体系统的实验","abstract":"Tariff threats can move household beliefs before policy is enacted, yet their rapidly changing language is difficult to study with conventional surveys. We build a multi-agent system that turns 300 households from the Michigan Surveys of Consumers into persistent large-language-model agents exposed to social-media information over several simulated months. Calibrated agents reproduce some distributional and demographic patterns in human survey data collected after the announcement of Liberation Day tariffs. Simulated experiments indicate that immediacy, rate salience, semantic progression, message complexity, narrative, and sender identity jointly shape inflation and unemployment expectations and their dispersion. Open-ended responses trace these effects to attention, ambiguity, credibility, and causal narratives. A second experiment finds that central-bank explanations can coordinate beliefs, although their effects on average expectations depend on message content. The framework supports disciplined exploration of policy communication, subject to human validation rather than as a substitute for it.","authors":["Jianhao Lin","Lexuan Sun","Yixin Yan"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30522","pdf_url":"https://arxiv.org/pdf/2608.30522","source_feed":"econ.GN","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3","B4"],"tags":["LLM仿真","宏观经济预期","政策沟通"],"reason":"用LLM代理300个家庭，复现关税冲击后的宏观预期，并与密歇根调查真实数据对照…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":6,"question":"关税威胁的时机、税率显著性、语义连贯性、信息复杂度、叙事和发送者身份如何共同影响家庭通胀与失业预期的水平及离散度？高影响威胁后，央行何种沟通最能协调信念？","design":"构建多智能体系统（MAS），将密歇根消费者调查（MSC）中的300个家庭转化为基于大语言模型的持久化智能体，赋予其人口特征和初始预期，在模拟的数月内暴露于社交媒体信息，并施加不同的关税威胁情景处理，测量其点预测、主观概率分布和开放式解释。","baseline":"密歇根消费者调查（MSC）在2025年4月“解放日”关税公告后收集的真实家庭通胀和失业预期数据，用于校准和验证模拟结果。","findings":"校准后的智能体在预期分布、人口统计差异和可区分性上接近人类数据；模拟实验表明，关税威胁的时机、税率显著性、语义进展、信息复杂度、叙事和发送者身份共同影响预期水平和离散度，开放式回答显示这些效应通过注意力、模糊性、可信度和因果叙事起作用。央行解释可以协调信念，但效果取决于信息内容。","reliability":"论文承认校准智能体不能替代人类，不能将未观察到的信息处理效应外推为人类因果效应；验证必须与语言模型的预期用途相关，行为相似性不足以证明从已验证场景到未观察场景的因果迁移。","relevance":"该研究直接命中研究者关注的核心：用LLM代理复现真实调查数据，并评估仿真可靠性，且涉及政策沟通和预期形成，值得精读原文以了解其校准方法和局限性讨论。","inspiration":"借鉴其将真实调查个体转化为持久化LLM智能体、施加文本处理并测量预期分布和开放式解释的设计，以及用真实调查数据校准和验证仿真的做法。｜可迁移到政策公告的预期形成研究，如央行沟通对通胀预期的影响、关税或财政政策冲击下的家庭预期调整。｜以真实家庭调查（如密歇根调查或纽约联储消费者预期调查）的受访者为被试，将其特征和初始预期编码为LLM智能体，施加不同措辞的央行声明或关税公告作为处理，测量通胀和失业预期的点预测、概率分布和开放式理由，并用同期真实调查数据作为对照基准进行校准和验证。"}},{"id":"2608.26849","version":2,"title":"LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems","zh_title":"LiveSim：在多智能体直播生态系统中模拟受环境塑造的用户","abstract":"User behavior simulation with large language models~(LLMs) is increasingly used to support multi-agent ecosystem simulation. Existing simulators typically rely on static user profiles inferred from historical observations, which become inadequate in socially intensive environments such as live streaming where interaction dynamics continuously reshape user behavior. We propose \\textbf{LiveSim}, an LLM-based framework for live-stream ecosystem simulation. It represents users as editable behavioral hypotheses and progressively refines them through trajectory-grounded interactions, where discrepancies between simulated and observed trajectories reveal missing environmental shaping effects. These signals are further extracted as transferable environment-behavior patterns and accumulated in a collective behavioral memory to improve user-level behavioral fidelity and support ecosystem-level simulation. Experiments on real-world live-stream risk-control data validate the effectiveness of LiveSim in improving user-level behavioral fidelity and enabling ecosystem-level analysis of risk evolution and platform intervention effects.","authors":["Jiaqi Xu","Yiran Qiao","Jing Chen","Qiwei Zhong","Xiang Ao","Xueqi Cheng"],"categories":["cs.AI","cs.CY","cs.MA"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-01","first_seen":"2026-08-28","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.26849","pdf_url":"https://arxiv.org/pdf/2608.26849","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM用户仿真","直播生态","行为保真度"],"reason":"用LLM仿真直播用户行为，并与真实数据对照，评估行为保真度，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":8,"question":"如何在直播这类社交密集环境中，用LLM模拟用户行为，并动态修正用户模型以反映环境对行为的塑造作用？","design":"提出LiveSim框架，用LLM扮演直播用户，将用户表示为可编辑的行为假设，通过模拟轨迹与真实轨迹的差异迭代修正假设，提取环境-行为模式存入集体记忆，再用于多智能体直播生态模拟。","baseline":"使用某大型直播平台真实风险控制数据中的用户行为轨迹作为对照。","findings":"LiveSim在用户级行为保真度上显著提升，并能支持生态级风险演化和平台干预效果分析。","reliability":"论文未讨论","relevance":"直接相关，用LLM仿真直播用户行为并与真实数据对照，评估行为保真度，属于人类仿真实验研究，值得精读原文。","inspiration":"借鉴其通过模拟-真实轨迹差异迭代修正行为假设的方法，可迁移到消费者在直播带货中的冲动购买行为研究。｜设计如下：用LLM模拟消费者，处理为不同主播话术（如限时折扣、从众压力），结果变量为购买决策，对照真实直播销售数据中的用户行为轨迹。"}},{"id":"2608.29803","version":1,"title":"Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments","zh_title":"LLM会像人类一样改变想法吗？诊断单轮说服判断中的人机分歧","abstract":"Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs in response to persuasive arguments, as humans do, remains poorly understood. We conduct a systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view. Our results show that LLMs achieve only slight agreement with humans (Cohen's kappa ranging from 0.079 to 0.178). Content-level analyses show that humans and LLMs agree on the strongest persuasion cues but diverge on finer ones: humans are more swayed by novel content and assertive language, whereas LLMs favor topical similarity and surface-level formatting. At the level of persuasion strategy, LLMs underweight emotional appeals and overweight credibility signals relative to humans, while the type of proposition under debate exerts no measurable effect on the degree of divergence. Furthermore, switching from first-person role-playing to third-person observation shifts all models toward greater resistance to persuasion, with the effect varying across persuasion strategies and textual features. These findings highlight the risk of treating LLM judgments as faithful proxies for human belief updating and point to structural differences in how LLMs and humans process persuasive discourse. Our code is available at https://github.com/tsinghua-fib-lab/LLM-belief-update-cmv.","authors":["Lin Chen","Yitong Chen","Yong Li"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29803","pdf_url":"https://arxiv.org/pdf/2608.29803","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","信念更新","人机对比"],"reason":"直接比较LLM与人类在说服中的信念更新，使用真实人类数据对照，并指出LLM作为…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":12,"question":"LLM在单轮说服中的信念更新判断是否与人类一致，哪些内容特征（命题类型、说服策略、文本特征）和视角框架（第一人称角色扮演 vs. 第三人称观察）调节这种分歧？","design":"使用ChangeMyView语料库中真实的说服对话，构建匹配的回复对（同一原帖下，一个被标记为说服成功，一个未成功），让8个LLM独立判断每个回复是否会改变原帖作者的观点，并分析判断结果与人类标签的一致性及影响因素。","baseline":"ChangeMyView语料库中原始发帖人明确标记的delta（观点改变）标签，作为人类信念更新的真实基准。","findings":"LLM与人类的一致性很低（Cohen's kappa 0.079-0.178），且LLM在文本特征和说服策略上与人类存在系统性分歧：人类更受新颖内容和果断语言影响，LLM更偏好主题相似性和表面格式；LLM低估情感诉求、高估可信度信号，而命题类型对分歧无显著影响。从第一人称角色扮演切换到第三人称观察会使所有模型更抗拒说服，且效应因策略和文本特征而异。","reliability":"论文指出LLM判断不能作为人类信念更新的忠实代理，存在结构性差异；但未明确讨论失效条件，仅强调在需要人类推理的社会模拟中直接使用LLM有风险。","relevance":"该研究直接比较LLM与人类在说服场景下的信念更新，使用真实人类数据作为基准，并揭示了系统性偏差，对关注LLM仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其匹配对设计和多维度内容标注方法，可系统诊断LLM与人类在决策中的分歧来源。｜可迁移到政策沟通与预期形成场景，如央行沟通对市场预期的影响。｜以LLM作为投资者被试，呈现央行声明或新闻，测量其预期更新，并与真实市场调查数据（如密歇根消费者信心调查）对照，分析LLM是否高估可信度信号或低估情感因素。"}},{"id":"2608.28668","version":1,"title":"Reference-Distribution Dependence in LLM-Based Synthetic Persona Data: Diagnosis and Post Hoc Adjustment of Demographic Distributions","zh_title":"基于LLM的合成人数据中的参考分布依赖：人口统计分布的诊断与事后调整","abstract":"We diagnose how closely the demographic distributions in LLM-based synthetic persona data match external reference distributions. For the three variables examined, we show that most of the observed error is attributable to the choice of reference rather than to the generator. Using total variation distance (TVD), we compare the sex x age group x province joint distribution of 1,000,000 records from Nemotron-Personas-Korea (NPK) with Korean official statistics. Against resident-registration figures for April 2026, the time of use, the bias bound, defined as the largest possible difference in the share of any subgroup formed from the three variables, is 1.81 percentage points. This is comparable to the margin of error of a survey of roughly 2,900 respondents. This value is not a fixed property of the data. Matching the reference period and series to the generating reference identified here, the 2024 register-based census restricted to Korean nationals, lowers it to 0.56 percentage points. Over the 15 months between the best-matching month (January 2025) and the time of use, the resident-registration population structure itself moves more than twice the distance of NPK's minimum error. Raking and cell post-stratification, the two weighting schemes used in Korean survey practice, remove most of the reference-period dependence at a variance inflation of about 0.2% in both cases. After raking against the generating reference, the residual joint discrepancy lies at, and marginally above, the upper bound of what a perfect generator would produce when realizing 1,000,000 records (97.6th percentile of the Monte Carlo distribution). We recommend treating synthetic persona data as auxiliary material for small-scale survey design rather than as a substitute for survey data, and re-running both diagnosis and adjustment against official statistics current at the time of use.","authors":["Eunjeong Song","Sehee Hong"],"categories":["cs.CY","stat.AP"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28668","pdf_url":"https://arxiv.org/pdf/2608.28668","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人口统计偏差","事后加权调整"],"reason":"用LLM生成合成人数据，与真实人口统计对照，诊断偏差并调整，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":9,"question":"LLM生成的合成人数据在人口统计分布上与外部参考分布有多接近，观察到的偏差在多大程度上归因于参考分布的选择而非生成器本身？","design":"使用Nemotron-Personas-Korea（NPK）数据集的100万条记录，比较其性别×年龄组×省份联合分布与韩国官方统计的差异，采用总变差距离（TVD）度量，并通过变化参考分布的时期和序列来诊断偏差来源，最后应用raking和单元后分层两种加权方法进行事后调整。","baseline":"韩国官方统计，包括居民登记数据（2026年4月）和2024年基于登记的人口普查（仅限韩国国民）。","findings":"在2026年4月使用时，NPK与居民登记数据的偏差上限为1.81个百分点，相当于约2900名受访者调查的误差边际；当与生成参考（2024年登记普查）匹配时，偏差降至0.56个百分点。大部分观察到的误差归因于参考分布的选择而非生成器，且raking和单元后分层能消除大部分参考期依赖性，方差膨胀仅约0.2%。","reliability":"论文建议将合成人数据视为小规模调查设计的辅助材料，而非调查数据的替代品，并强调应在使用时重新对当前官方统计进行诊断和调整；同时指出合成数据冻结了生成时的人口结构，随时间推移与官方统计的一致性自然下降。","relevance":"该研究直接评估LLM合成人数据与真实人口统计的偏差，并提供了诊断和调整方法，对关注LLM仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文以了解具体方法和发现。","inspiration":"借鉴其通过变化参考分布来区分生成器偏差与参考选择偏差的诊断方法，以及使用TVD和偏差上限将分布差异转化为调查误差边际的做法。｜可迁移到经济金融领域中使用合成数据模拟消费者或投资者行为的研究，例如评估政策变化对特定人群的影响或测试金融产品在不同人口群体中的接受度。｜设计：使用LLM生成具有特定人口特征（如年龄、收入、地区）的合成个体，模拟他们对某项经济政策（如税收调整）的反应，结果变量为支持率或行为变化，并与真实调查数据（如韩国劳动力面板或家庭收入支出调查）进行对照，通过变化参考分布和事后加权来评估仿真的可靠性。"}},{"id":"2608.29266","version":1,"title":"Measurement Validity in LLM Cultural Alignment","zh_title":"大语言模型文化对齐中的测量效度","abstract":"Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments like the Inglehart-Welzel Cultural Map and drawing conclusions about which cultures a model resembles. While a model's answer to a value-laden questions may be interpreted as a cultural signal, it also carries sampling noise and, can be quite sensitive to question framing. In this paper, we separate survey responses, sampling noise and question framing for multiple LLMs. We decompose response variance from these models into variation across random seeds, prompt rewordings. We employ noise-to-signal ratio (NSR) to test whether a model's apparent cultural position is distinguishable from noise. When applied across a dozen models from four geographic origins, calibrated against 88 Integrated Values Survey countries, the answer is often no. NSR exceeds 1.0 on 49 of 117 valid model-question pairs (42%), reaching 5.56 in the worst case. Two models even refuse to answer sufficient number of survey questions outright. Our results corroborate previous findings that LLMs cluster toward Western, English-speaking cultural positions. However, what does not hold up in this study is the precision with which anyone can currently interpret a specific model's coordinates: prompt tone alone can shift a model by 2.4 map units, comparable to the distance between actual countries in the Inglehart-Welzel Cultural Map. These findings suggest that cultural attribution from LLM survey responses requires establishing the reliability of the underlying measurements before interpreting model coordinates as evidence of cultural representation.","authors":["An Duy Nguyen","Muhammad Aurangzeb Ahmad"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29266","pdf_url":"https://arxiv.org/pdf/2608.29266","source_feed":"physics.soc-ph","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","文化价值观","测量信度"],"reason":"用LLM回答调查问题模拟文化价值观，并与88国真实数据对照，评估测量信度与偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":10,"question":"LLM 对文化价值观调查的回答在多大程度上反映了真实的文化信号，而非采样噪声和提问措辞的敏感性？","design":"该研究并非传统意义上的仿真实验，而是对 LLM 作为测量工具的信度检验。作者选取 12 个来自不同地域的 LLM，让它们回答 Inglehart-Welzel 文化地图的 10 个价值观问题，并通过变换随机种子和提示语气/人称来分解回答方差，计算噪声信号比（NSR）以评估测量有效性。","baseline":"88 个国家的 Integrated Values Survey（IVS）数据，用于构建 Inglehart-Welzel 文化地图的参照系。","findings":"LLM 的文化位置普遍偏向西方英语国家，但测量信度很差：42% 的模型-问题对 NSR 超过 1.0，提示语气变化可使模型坐标移动 2.4 个地图单位，相当于真实国家间的距离。部分模型拒绝回答敏感问题，导致无法投影到文化地图上。","reliability":"论文明确指出，在未建立测量信度之前，不能将 LLM 的调查回答解读为文化表征。局限包括：仅使用 Inglehart-Welzel 框架，问题数量有限；未覆盖所有可能的提示变体；模型版本和采样参数可能影响结果。","relevance":"该研究直接回应了用 LLM 模拟人类价值观的可靠性问题，与您关注的人类仿真实验的效度批判高度相关，值得精读原文以了解其噪声分解方法和 NSR 指标。","inspiration":"借鉴其将回答方差分解为随机种子、提示措辞和模型间差异的方法，用于检验 LLM 在经济调查中的测量信度。｜可迁移到消费者信心调查、通胀预期、风险偏好等经济态度的测量，评估 LLM 作为被试的可靠性。｜以 GPT-4 等模型为被试，施加不同措辞的问卷版本，测量其通胀预期或风险偏好，并与密歇根大学消费者调查或实验经济学中的真实人类数据对照，计算 NSR 以判断 LLM 回答是否超出噪声。"}},{"id":"2608.29535","version":1,"title":"Integrating adaptive human behavior into epidemic models with large language models","zh_title":"用大语言模型将自适应人类行为整合进流行病模型","abstract":"Infectious disease transmission is shaped by patterns of human interaction, which adapt as epidemic conditions change. Capturing these context-dependent behaviors remains a fundamental challenge for epidemic models. Here, we recast this challenge by using large language models (LLMs) to represent adaptive human behavior within mechanistic epidemic models. We operationalize this idea through Generative Adaptive Behavioral Layer for Epidemics (GABLE), which adapts LLMs to infer behavioral responses to epidemic and policy conditions and translates them into age-structured contact matrices coupled to a mechanistic epidemic model. Applied to COVID-19 in France, GABLE reproduced responses in population mixing and age-specific contact structures that remained epidemiologically informative. In short-term forecasting, LLM-generated contact matrices outperformed mobility-driven matrices derived from real-world mobility data, with the largest gains at longer horizons. GABLE also extends beyond forecasting to prospective policy evaluation by projecting behavioral and epidemic responses to candidate interventions before implementation. When supplied with subsequently implemented policies, GABLE reproduced epidemic trajectories and generated distinct responses to alternative policy timing and composition. By leveraging LLMs as a flexible behavioral layer, GABLE provides a framework for coupling context-sensitive behavioral generation with epidemic dynamics.","authors":["Yicheng Mao","Haoyang Li","Rob Deardon","Hongru Du"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29535","pdf_url":"https://arxiv.org/pdf/2608.29535","source_feed":"physics.soc-ph","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","流行病建模","政策评估"],"reason":"用LLM模拟人类在疫情中的行为，并与真实数据对照，用于预测和政策评估。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":11,"question":"如何用大语言模型模拟传染病流行中人类行为的适应性变化，并将其耦合到机制化流行病模型中以改进预测和政策评估？","design":"提出GABLE框架，用LLM（GPT-4o mini、Gemini 2.5 Flash、Grok 3 mini）作为行为层，根据当前疫情状态、政策条件、人口特征和疫情前接触日记生成年龄分层接触矩阵，再输入年龄分层随机传播模型，形成行为-疾病反馈循环；在法国COVID-19场景下进行回顾性重建、短期预测和前瞻性政策评估。","baseline":"真实人类数据包括：法国COVID-19住院数据、疫情前接触调查（接触日记）、基于真实移动数据的接触矩阵（mobility-driven matrices）。","findings":"GABLE生成的接触矩阵能重现人群混合和年龄特异性接触结构的变化，且在短期预测中优于基于移动数据的矩阵，尤其在较长预测期优势更大。GABLE还能在政策实施前预测行为和疫情反应，对替代政策时机和组合产生不同轨迹。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类在疫情中的适应性行为，并与真实接触和移动数据对照，用于预测和政策评估，直接命中研究者关注的LLM仿真实验、真实数据基准和政策评估场景，值得精读原文。","inspiration":"借鉴其将LLM生成的行为参数（接触矩阵）嵌入机制模型并形成反馈循环的设计，可迁移到政策公告对经济行为影响的研究中。｜可应用于政策公告的预期形成与消费/投资行为调整，如财政刺激或货币政策沟通。｜以LLM模拟不同人口群体（如年龄、收入分层）对政策公告的行为反应（如消费支出、劳动供给），将生成的行为参数输入宏观经济模型，并与真实调查数据（如消费者信心指数、信用卡消费数据）对照验证。"}},{"id":"2608.28182","version":1,"title":"Benchmarking large language model agent societies against human behavioural distributions","zh_title":"基于人类行为分布基准测试大语言模型智能体社会","abstract":"Populations of large language model agents are increasingly used as experimental societies. Three doubts shadow every such result: whether the agents behave like the humans they stand in for, whether a finding survives changes to the apparatus that leave the rules untouched, and whether apparent social dynamics are interaction at all rather than the reproduction of experiments the models have read. This article introduces SILICA, an open instrument that tests all three. Five environments carry published human anchors, each paired with perturbations that re-render the same rules and with variants whose payoffs point away from the memorised result. Twelve open-weight models were run through it on a single consumer graphics card. Agreement with human data is confined to starting points: first-round public-goods contributions fall inside the equivalence margin for eight of eleven models, while no model matches end-state contributions or the human corridor of cooperation. Merely swapping the order in which two actions are listed costs one model 58 points of cooperation. Presenting responders with a fixed schedule of offers shows that only one model, the sole reasoning-trained one, places its acceptance threshold where the incentive requires; two move theirs part of the way, two move them the wrong way, and three never acquire one. Conventions form through a shared prior over the names rather than through negotiation, though negotiation reappears once that prior is disrupted. On the certification ladder defined here, current silicon societies support exploratory claims and no more.","authors":["Raad Bin Tareaf"],"categories":["physics.soc-ph","cs.CL"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28182","pdf_url":"https://arxiv.org/pdf/2608.28182","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM仿真","人类行为对照","算法保真度"],"reason":"直接以LLM agent群体仿真人类行为，并与真实人类数据对照，评估可靠性，涉…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-31","rank":1,"question":"LLM智能体社会能否在行为分布上复现人类实验基准，其结论对实验装置扰动和记忆污染是否稳健？","design":"用12个开源权重LLM作为智能体，在五个有已发表人类锚点的交互式多智能体环境中运行（重复囚徒困境、带代价惩罚的公共品博弈、讨价还价、11-20要钱游戏、带承诺少数派翻转的命名游戏），施加设计级和表征级扰动，并设置收益指向偏离记忆结果的变体，测量合作水平、接受阈值、约定形成等行为结果。","baseline":"五个环境均配有已发表的人类实验锚点，如公共品博弈首轮贡献、最后贡献、合作走廊，以及讨价还价中的接受阈值等。","findings":"与人类数据的一致性仅限于起点：11个模型中有8个首轮公共品贡献落在等效边界内，但没有模型匹配最终贡献或人类合作走廊；仅交换两个动作的列出顺序就使一个模型合作率下降58个百分点。在固定报价序列下，只有唯一一个推理训练模型将接受阈值放在激励要求的位置，其他模型或部分移动、或反向移动、或根本不形成阈值；约定通过名称共享先验而非协商形成，但先验被扰乱后协商重新出现。","reliability":"论文承认当前硅基社会仅支持探索性声明，不能支持可转移或稳健的结论；设计级扰动在111个可计算对比中改变行为56次，表征级扰动在71个中改变8次，且固定报价序列揭示聚合拒绝率无法区分模型是否真正习得激励。","relevance":"该研究直接以LLM智能体群体仿真人类行为，并与真实人类数据对照，系统评估了仿真在经济学实验中的保真度、稳健性和污染问题，对关注LLM仿真可靠性与偏差的研究者具有核心参考价值。","inspiration":"可借鉴其通过设计级与表征级扰动分离内容与形式影响、以及用固定报价序列识别个体接受函数来审计记忆污染的方法。｜可迁移到资产定价实验中的策略性报价、信贷审批中的歧视测量、或消费者跨期选择中的时间偏好等场景。｜用LLM智能体扮演投资者或消费者，施加收益结构改变或信息呈现方式扰动，测量报价、接受阈值或跨期选择，并与实验室或现场实验的真实人类数据做等效性检验。"}},{"id":"2608.26291","version":1,"title":"Assessing mentalization in humans and large language models","zh_title":"评估人类与大语言模型的心理化能力","abstract":"Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks. Last, GPT-5 agents flexibly adapted their recursive depth of reasoning to increasingly sophisticated opponents, demonstrating superior performance to human participants. Collectively, we demonstrate different capacities for mentalization across LLMs, and highlight cognitive computational modeling as a formal method for assessing comparative intelligence across humans and machines.","authors":["Aamir Sohail","Xintong Zhong","Arkady Konovalov","Patricia L. Lockwood","Lei Zhang"],"categories":["cs.AI","q-bio.NC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26291","pdf_url":"https://arxiv.org/pdf/2608.26291","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3"],"tags":["LLM仿真","经济学实验","认知建模"],"reason":"用LLM作为人类被试替代品，在经济学博弈中与人类数据对照，评估心理化能力与策略。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":1,"question":"LLM能否像人类一样通过心理化（推断他人信念与意图）来指导策略行为，其潜在策略是什么？","design":"用四个模型家族（DeepSeek、GPT-4.1、GPT-5、Gemini 2.0 Flash）的LLM个体（N=2099）作为被试，参与两个经济博弈（检查博弈和石头剪刀布），部分模型施加社会思维链（SCoT）提示，测量博弈得分和策略选择，并用认知计算建模推断递归推理深度。","baseline":"人类被试（检查博弈 n=67，石头剪刀布 n=184）在相同任务中的表现。","findings":"LLM表现出明显的心理化行为与计算特征，但不同模型和规模差异显著；SCoT提示普遍提升表现，但提升幅度因模型而异。GPT-5能灵活适应对手复杂度，表现优于人类。","reliability":"论文未讨论","relevance":"该研究直接以LLM作为人类被试替代品，在经济学博弈中与人类数据对照，评估心理化能力与策略，完全命中你的关注点，值得精读原文。","inspiration":"借鉴其用认知计算建模从行为数据中提取潜在策略（递归推理深度）的方法，以及用SCoT提示作为处理变量来诱导更复杂推理的设计。｜可迁移到资产定价实验中的策略性预期形成或谈判博弈中的信念更新。｜用LLM作为投资者被试，施加SCoT提示，在重复博弈中测量报价或投资决策，并与人类实验数据（如资产泡沫实验）对照，比较递归推理深度和适应性。"}},{"id":"2608.26327","version":1,"title":"How Unlikely Is \"Unlikely\"? Assessing Verbal Probability Perception Across Large Language Models","zh_title":"“不太可能”有多不可能？跨大语言模型评估言语概率感知","abstract":"Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty) remains unknown. We present a systematic cross-model evaluation using a word-to-number mapping task grounded in established human benchmarks. Eleven uncertainty expressions were presented to 19 models under two conditions, forced single-number response and explanation elicitation, alongside a novel bidirectional roundtrip test of internal consistency. LLMs track the human benchmark with surprising fidelity: word ordering is preserved, three anchor points are recovered, and ``possible'' shows the highest variance and cross-model disagreement of any expression tested, consistent with its documented bimodal interpretation in humans. However, models show a systematic upward bias for negative expressions such as ``unlikely'' and ``improbable.'' Explanation elicitation reduces within-model variance while increasing between-model divergence, stabilizing individual models at the cost of inter-model consensus, and the roundtrip experiment reveals clear stratification, with frontier models maintaining coherent bidirectional representations. LLMs thus reproduce the structure of human verbal probability cognition, including its biases, while diverging systematically at the negative end---with implications for any setting where humans and models exchange probabilistic language.","authors":["Christos Petridis","Konstantinos Pelechrinis","Zoran Obradovic"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26327","pdf_url":"https://arxiv.org/pdf/2608.26327","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","概率语言","人类基准对照"],"reason":"用LLM复现人类对概率词的理解，并与人类基准对照，发现偏差，属于仿真人类认知且…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":5,"question":"LLM对言语概率表达（如“unlikely”）的数值解读是否与人类基准一致，以及在不同提示条件下是否稳定？","design":"用19个LLM作为被试，呈现11个概率词，在两种提示条件下（强制单数字回答、要求解释）进行词到数字映射，并进行双向往返一致性测试，测量映射数值、方差和跨模型一致性。","baseline":"Mosteller和Youtz（1990）汇总的20项人类研究中概率词的数值解读基准。","findings":"LLM整体上复现了人类基准：词序保持、三个锚点（impossible, even chance, certain）恢复良好，“possible”方差最大，与人类双峰解释一致。但LLM对负面词（如unlikely）存在系统性高估，解释提示降低模型内方差但增加模型间分歧，往返测试显示前沿模型双向映射更一致。","reliability":"论文指出偏差集中在非锚点词，负面词在强制条件下偏差略大；解释提示虽稳定个体但牺牲跨模型共识；部分模型往返映射接近随机，表明内部表征不一致。","relevance":"该研究直接评估LLM作为人类被试在概率词理解上的仿真效度，发现总体复现但存在系统性偏差，对关注LLM仿真人类认知可靠性的研究者有重要参考价值。","inspiration":"借鉴其词到数字映射任务和双向一致性检验，可迁移到经济金融中的风险沟通与预期形成场景，例如央行政策声明中的模糊措辞解读。｜设计一个实验：以LLM为被试，呈现央行声明中的概率词（如“可能加息”），要求给出数值概率，并与专业预测者调查或市场隐含概率对照，检验LLM是否复现人类解读偏差。"}},{"id":"2608.26188","version":1,"title":"Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments","zh_title":"你的社区安全吗？大语言模型城市安全判断中的地方污名","abstract":"Large language models are increasingly used to inform safety decisions in cities, such as where it is safe to walk, rent, or travel. We ask whether such judgments track measured risk or the patterns attached to an urban neighborhood's name. We probe seven instruct-tuned models under three conditions that dissociate name from geography: coordinates-only, name-only, and name+coordinates, across 186 neighborhoods in Los Angeles and Chicago, joined to violent crime and American Community Survey data. First, ratings are nearly flat under coordinates for six of seven models, while names carry most between neighborhood variation and are moderately calibrated to violent crime; only at frontier scale does the coordinate channel show appreciable variation. Second, names lower safety ratings more for neighborhoods with higher shares of the locally dominant marginalized group (percent Black in Chicago, percent Hispanic in Los Angeles), and this name effect tracks demographic share in all seven models and both cities. In Los Angeles, where demographic share and crime are more separable, the effect survives controls for crime and income and is confirmed by crime-matched pairs. An enforcement-elasticity analysis further shows that over-caution tracks near-fully-reported homicide rather than discretionary, deployment-driven offenses. Third, the effect scales with geographic knowledge: models that better distinguish real neighborhoods apply more demographic stereotype to them. Because neighborhood names carry both genuine crime signal and demographic stereotype, removing names reduces both bias and accuracy. We discuss implications for deploying LLMs in advice and decision-support settings.","authors":["Huy Nguyen","Yue Lin"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26188","pdf_url":"https://arxiv.org/pdf/2608.26188","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","社会偏见","城市安全"],"reason":"用LLM模拟人类对社区安全的判断，并与真实犯罪和人口数据对照，揭示偏差。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":3,"question":"LLM 对城市社区夜间步行安全的判断，是追踪真实犯罪风险，还是反映与社区名称相连的种族-空间污名？","design":"用七种指令微调 LLM 模拟居民对社区安全的感知，对芝加哥和洛杉矶共 186 个社区，在仅坐标、仅名称、名称+坐标三种条件下生成夜间步行安全评分，并与真实暴力犯罪和人口普查数据对照。","baseline":"真实暴力犯罪记录（芝加哥 2020-2025、洛杉矶 2020-2024）和美国社区调查（ACS）人口统计数据。","findings":"名称承载了几乎所有社区间安全评分差异，且与暴力犯罪中度校准；名称对安全评分的压低效应随当地主导边缘化群体（芝加哥黑人、洛杉矶西裔）比例上升，在洛杉矶该效应在控制犯罪和收入后仍存在，且与犯罪匹配对验证一致。","reliability":"坐标通道在开源模型中几乎无变异，仅在 frontier 模型上显现；芝加哥因黑人比例与犯罪高度相关无法分离种族与犯罪效应；名称同时携带真实犯罪信号和人口统计刻板印象，隐藏名称会同时消除偏差和准确性。","relevance":"该研究用 LLM 模拟人类对社区安全的判断，并与真实犯罪和人口数据对照，直接揭示仿真中的种族偏差，对关注 LLM 仿真可靠性及偏差的研究者极具参考价值。","inspiration":"值得借鉴的是通过条件消融（仅坐标、仅名称、名称+坐标）分离名称效应，并用真实犯罪和人口数据做基准，以及用犯罪匹配对和执法弹性分解做稳健性检验。｜可迁移到信贷审批中的地域歧视研究，如 LLM 模拟信贷员对申请人的风险评估是否受申请人所在社区名称的种族构成影响。｜用 LLM 扮演信贷审批员，对虚构申请人给出贷款批准概率，处理变量为申请人地址的社区名称（高黑人/西裔比例 vs 低比例），结果变量为批准概率，对照真实数据用社区层面的实际贷款批准率和违约率，并控制申请人收入、信用分等特征。"}},{"id":"2608.26221","version":1,"title":"Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model","zh_title":"生成式智能体的提示敏感性：来自流行病模型的证据","abstract":"As generative AI gains traction, researchers are investigating its potential to serve as proxies for humans. From undergoing cognitive psychology experiments to experiencing an epidemic, generative agents, agents powered by generative AI models, produce realistic human behavior when prompted. This study explores the sensitivity of these generative agents' behavior to prompt modifications and varied persona names of the agents. To assess this sensitivity, we use a generative agent epidemic model, wherein each agent is prompted daily on whether it wants to isolate or commingle with other agents. We found that using synonymous prompts results in negligible changes to the model's outcomes. However, minor variations in prompts, as well as contextual changes, do influence the model's results. Lastly, our data indicates that different persona names assigned to generative agents, specifically those imbued with personas, do not significantly impact epidemic outcomes.","authors":["Ross Williams","Niyousha Hosseinichimeh"],"categories":["physics.soc-ph","cs.AI","cs.LG","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26221","pdf_url":"https://arxiv.org/pdf/2608.26221","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A3","B4"],"tags":["LLM仿真","流行病模型","提示敏感性"],"reason":"用生成式智能体模拟疫情中的人类隔离决策，研究提示敏感性，属于人类行为仿真，但无…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":4,"question":"生成式智能体在疫情模型中的行为对提示词修改和角色名称变化的敏感程度如何？","design":"使用生成式智能体疫情模型，每个智能体每天被提示选择隔离或与他人接触，通过改变提示词的语义、上下文和角色名称来测试行为变化，结果变量为疫情传播的流动性曲线。","baseline":"无对照","findings":"同义提示词修改对模型结果影响可忽略，但轻微提示词变化和上下文变化会影响模型结果。不同角色名称对疫情结果无显著影响。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM仿真中提示敏感性这一可靠性问题，但未与真实人类数据对照，且场景为疫情模型而非经济学实验，对关注经济学和政策评估的研究者参考价值有限。","inspiration":"可借鉴其系统改变提示词并测量输出变化的方法，用于评估LLM在经济实验中的稳健性。｜可迁移到政策公告的预期形成实验，测试不同措辞的公告对LLM代理预期的影响。｜用LLM模拟投资者，随机分配不同措辞的央行声明，测量其通胀预期变化，并与专业预测者调查数据对照。"}},{"id":"2608.24912","version":1,"title":"Analyzing and Correcting Benevolence Bias in Large Language Models","zh_title":"分析和纠正大语言模型中的仁慈偏差","abstract":"Large language models (LLMs) are increasingly used as stand-ins for human respondents, from opinion polls and simulated survey participants to agent-based social simulations. These uses rest on one assumption: that conditioning a model on who a person is yields answers resembling those of real people from that group. Here we identify and measure benevolence bias, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions. Across 18 widely used models, four social-science datasets (ANES, GSS, WVS, and a cross-cultural prospect-theory replication) and six psychological categories, we find that the bias is a stable model property, not a quirk of any one system: it points the same way across models, grows with model size, and traces to the post-training stage. Prompt language and framing change its size but never its direction, and a \"malicious persona\" stress test shows a one-sided limit: aligned models struggle to play people who are less kind, less prosocial or more harm-tolerant than average. The issue is thus not only a shifted average, but a narrowed range of people the model can imitate. The bias sits in the middle of the answer distribution rather than its tails, and survives changes in sampling temperature and simple prompted reflection. The encouraging news is that it is easy to diagnose and straightforward to fix: a light-touch contrastive calibration, which needs no retraining and works on black-box APIs, brings all six categories back to the human baseline. Our results give researchers a clear map of where aligned LLMs can already be trusted as human stand-ins, where they need care, and a ready-to-use method for closing the gap.","authors":["Yuanzi Li","Junhao Wang","Minghui Liu","Boyi Li","Bingchen Chen","Zihang Tian","Jingyu Zhao","Yuhan Wang","Lei Wang","Pei Wang","Jinchao Wu","Xu Chen"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24912","pdf_url":"https://arxiv.org/pdf/2608.24912","source_feed":"cs.HC","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B3","B4"],"tags":["LLM仿真","算法保真度","偏差校正"],"reason":"直接研究LLM作为人类被试替代品的偏差，使用真实人类数据对照，并提出校准方法。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":1,"question":"对齐后的大语言模型在作为人类被试回答价值负载的调查问题时，是否存在系统性的仁慈偏差，其来源、表现和可校正性如何？","design":"该研究并非传统仿真实验，而是对18个广泛使用的LLM进行系统性测试：将模型置于模拟人类受访者的角色，输入来自ANES、GSS、WVS和跨文化前景理论复制的调查问题，测量模型回答在六个仁慈偏差类别（社会赞许性、伤害规避、亲社会动机、仁慈解释、公平乐观、情感软化）上的偏差程度，并考察模型规模、训练阶段、提示语言、框架、角色设定、采样温度等因素的影响，最后提出对比校准方法。","baseline":"使用四个真实人类调查数据集作为基准：美国国家选举研究（ANES）、综合社会调查（GSS）、世界价值观调查（WVS）以及一项跨文化前景理论复制的数据。","findings":"LLM存在稳定且一致的仁慈偏差，倾向于选择更友善、更安全、更符合社会期望的答案，且该偏差随模型规模增大而增强，主要源于后训练阶段。提示语言和框架只能改变偏差大小而不能改变方向，恶意角色压力测试显示模型难以模仿低于人类平均水平的亲社会或伤害容忍度，但对比校准方法无需重新训练即可将偏差校正至人类基线。","reliability":"论文指出偏差位于答案分布的中间而非尾部，且对采样温度和简单提示反思不敏感；恶意角色测试显示模型在部分维度上无法达到人类低仁慈端，表明仿真范围收窄。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，提供了系统性的偏差测量和校正方法，对于关注仿真效度和偏差的研究者具有重要参考价值。","inspiration":"该研究采用多模型、多数据集、多心理类别的系统测量框架，并通过对比校准进行偏差校正，值得借鉴。｜可以迁移到经济金融领域的调查实验和个体决策仿真，例如风险偏好、时间偏好、公平观念、信任与合作等。｜以LLM作为虚拟被试，施加不同的经济情境或政策干预，测量其选择或态度，并与真实实验数据（如实验经济学中的公共品博弈、最后通牒博弈、风险偏好问卷等）进行对照，检验并校正LLM的偏差。"}},{"id":"2604.06223","version":3,"title":"The Quiet and the Compliant: How Regulation and Polarization Shape Conventional Wisdoms on Corporate Social Engagement in High-risk Settings","zh_title":"沉默与顺从：监管与极化如何塑造高风险环境下企业社会参与的常规智慧","abstract":"With the international business landscape becoming more crisis-ridden as risks proliferate, how do the professionals who implement corporate social initiatives in high-risk environments perceive their work, and what can this reveal about the forces shaping business engagement with society in crisis contexts? We present findings from a synthetic survey of 400 corporate professionals working on social impact in fragile and conflict-affected settings to understand conventional wisdoms and best practices on corporate strategy and activity in high-risk settings. Drawing on political corporate social responsibility (CSR), synthetic survey, and international business literatures, we test seven hypotheses about how regulatory environments, political polarization, sector characteristics, and organizational structures shape corporate social engagement in high-risk contexts. The synthetic results suggest that European professionals report significantly higher strategic integration of social impact across all measured dimensions, while US professionals overwhelmingly report that political polarization hinders social initiatives, yet this perception does not predict unreported social activities, complicating the emerging \"quiet CSR\" narrative. Extractive industry professionals deliver both the highest operational preparedness and the highest complicity awareness, a pattern we conceptualize as presence-dependent reflexivity. These patterns deliver a baseline to detect the theorized dynamics and offer preliminary theoretical propositions for future real-world empirical testing.","authors":["Jason Miklian"],"categories":["physics.soc-ph","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"replace-cross","date":"2026-08-27","first_seen":"2026-03-27","revised_at":"2026-08-27","abs_url":"https://arxiv.org/abs/2604.06223","pdf_url":"https://arxiv.org/pdf/2604.06223","source_feed":"cs.SI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","合成调查","企业社会责任"],"reason":"用LLM合成400名企业专业人士调查，模拟高风险环境下的态度与决策，并与真实文…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":2,"question":"在脆弱和冲突影响的高风险环境中，企业社会参与的专业人士如何感知其工作，以及监管环境、政治极化、行业特征和组织结构如何塑造企业社会参与？","design":"使用合成调查方法，模拟400名在欧美总部、员工超过1000人的企业社会影响专业人士，通过23个李克特量表题测量战略整合、监管压力、政治极化、运营准备、供应链协调、子公司自主权和ESG评级有效性等感知，并基于理论推导出七个假设进行检验。","baseline":"无对照","findings":"欧洲专业人士在所有维度上报告显著更高的社会影响战略整合，而美国专业人士压倒性地报告政治极化阻碍社会倡议，但这种感知并不预测未报告的社会活动，复杂化了“安静CSR”叙事；采掘业专业人士同时表现出最高的运营准备和最高的共谋意识，被概念化为存在依赖的反身性。","reliability":"论文未讨论","relevance":"该研究使用LLM合成调查数据，模拟企业专业人士在高风险环境下的态度和决策，并基于理论提出假设进行检验，属于用LLM进行人类仿真实验的研究，且涉及政策评估场景，与您的兴趣高度相关，值得阅读原文了解其方法论和发现。","inspiration":"该方法通过合成调查生成理论预测的基准数据，为后续真实数据对比提供参照，可借鉴其构建“常规智慧基线”的思路；可迁移到政策评估中的企业行为研究，如ESG监管对企业社会参与的影响；设计上，可用LLM模拟企业高管作为被试，施加不同监管环境（如强制尽职调查 vs 反ESG法案）的处理，测量其战略整合和沉默行为，并与真实企业调查数据对照。"}},{"id":"2608.25771","version":1,"title":"Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences","zh_title":"基于困境训练的大语言模型少样本提示在预测患者偏好上超越人类代理","abstract":"In serious illness, human surrogates often struggle to accurately predict patient preferences (68% accuracy), causing decision conflict. Personalized Patient Preference Predictor (P4) agents offer a potential solution, but prior prototypes treat values as static ratings, ignoring the contextual, situation-dependent nature of medical choices. Grounded in the 'logic of care', we present P4-DT (Dilemma Training), a P4 agent that constructs a patient decision policy by engaging users with varied medical dilemmas, eliciting individual preference reasoning through bi-directional training. In a study with 12 patient-surrogate dyads, P4-DT predicted patient treatment choices with 81.7% accuracy, significantly exceeding chance (OR = 5.61 [2.03, 15.51], p < .001) and outperforming both unassisted surrogates (55.0%; OR = 3.67 [1.59, 8.47], p = .002) and surrogates assisted by P4-DT (61.7%). Comparative prompt analyses showed that incorporating contextual scenario decisions and open-ended text improved accuracy by 15.0 percentage points over initial values ratings alone. We discuss implications for further testing and designing of context-aware AI agents that embody richer human experience to partner in complex decision-making.","authors":["Natasha Ureyang","Sebastian Porsdam Mann","Yuxin Liu","Zuriel Hassirim","Melanie Almonte","Wenhao Chen","Joyce Ng","Thant Nay Lin","Aung Thiha","Gerald CH Koh","Brian David Earp","Pin Sym Foong"],"categories":["cs.HC","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25771","pdf_url":"https://arxiv.org/pdf/2608.25771","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","患者偏好预测","人类对照"],"reason":"用LLM预测患者偏好，与人类代理对照，评估准确率，属核心仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":4,"question":"如何通过让患者参与医疗困境训练来构建个性化患者偏好预测器（P4-DT），以提高对患者治疗偏好的预测准确率？","design":"使用GPT-5.5模型作为P4-DT代理，通过提示工程进行少样本学习。对12名患者（MP）进行训练：先填写价值观调查，再对5个医疗困境场景做出治疗偏好决策并解释理由，同时可查看模型预测并反馈。测试阶段，MP对5个新场景做决策，模型基于训练数据预测其偏好；同时人类代理人（TO）独立预测MP偏好，并在P4-DT辅助下再次预测。结果变量为预测准确率（方向一致性）。","baseline":"人类代理人（TO）的预测准确率（55.0%），以及TO在P4-DT辅助下的预测准确率（61.7%）。","findings":"P4-DT预测患者治疗选择的准确率为81.7%，显著高于随机水平（OR=5.61, p<.001），且优于未辅助的人类代理人（55.0%）和P4-DT辅助的代理人（61.7%）。比较提示分析显示，纳入情境场景决策和开放式文本比仅使用初始价值观评分提高了15.0个百分点的准确率。","reliability":"论文未讨论","relevance":"该研究用LLM模拟患者偏好预测，并与真实人类代理人对照，评估预测准确率，属于核心的LLM仿真人类决策研究，且涉及医疗决策场景，对关注仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其通过情境化困境训练和开放式文本解释来捕捉个体决策逻辑的方法，可迁移到消费者金融决策或政策偏好预测中。｜例如，在消费者信贷选择或退休储蓄决策中，可让LLM通过模拟具体金融困境（如贷款选择、投资风险权衡）来学习个体偏好。｜设计：招募真实消费者作为被试，先填写价值观和风险偏好问卷，再对5个金融困境场景做出选择并解释理由，训练LLM预测其在新场景中的选择；同时让人类代理人（如配偶）预测被试选择，比较LLM与人类代理人的预测准确率，并以被试实际选择为基准。"}},{"id":"2608.20539","version":2,"title":"ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations","zh_title":"ExploraTwin：一个用于数字孪生仿真的非营利研究平台","abstract":"Digital twin simulations show promise, but current empirical evidence suggests that the approach should be tested before being deployed in any particular context. To lower the friction for researchers and practitioners to test and deploy digital twin simulations, this brief commentary introduces ExploraTwin (https://exploratwin.org), an open-access, non-profit research platform for digital twin survey simulations. ExploraTwin supports two modes. In survey mode, researchers can upload a Qualtrics survey file or create a survey within the platform; select an available sample of digital twins; configure and run the simulation, and export analysis-ready data. In panel mode, researchers can assemble a small group of twins for open-ended conversations, document annotation, and moderated, focus-group-style voice discussions. We also developed CroissantTwin, a standardized data format for adding samples of digital twins to the platform. We demonstrate the survey mode workflow by using the platform to replicate 19 experiments on digital twins from the Twin-2K-500 dataset. ExploraTwin's survey execution fidelity is high: 99.6% of 197,000 answer units returned a structurally valid response on the first run.","authors":["Naveen Venkat","Yuchen Qiu","Tianyi Peng","George Gui","Olivier Toubia"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-26","first_seen":"2026-08-24","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.20539","pdf_url":"https://arxiv.org/pdf/2608.20539","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1"],"tags":["数字孪生","调查仿真","平台"],"reason":"平台支持数字孪生调查仿真，复现19个实验并与真实数据对照，验证执行保真度。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":3,"question":"如何降低研究者测试和部署数字孪生调查仿真的工程门槛与成本，并验证其执行保真度？","design":"开发开源非营利平台 ExploraTwin，支持上传 Qualtrics 问卷或平台内建问卷，选择数字孪生样本（如 Twin-2K-500），配置并运行仿真，导出分析就绪数据；同时提供面板模式进行开放式对话和焦点小组式讨论。","baseline":"使用 Twin-2K-500 数据集中的数字孪生，复现 Peng et al. (2025) 的 19 个实验，并与原始人类实验结果对照。","findings":"ExploraTwin 平台实现了低摩擦的数字孪生调查仿真工作流，支持复杂问卷逻辑和多种孪生样本。在复现 19 个实验时，197,000 个答案单元中 99.6% 在首次运行返回结构有效响应，执行保真度高。","reliability":"论文承认数字孪生的预测性能参差不齐，强调在特定情境部署前必须进行测试；平台当前免费但未来可能收费；未详细讨论仿真偏差或失效条件。","relevance":"该论文直接提供可用的数字孪生仿真平台，并复现多个实验验证保真度，对关注 LLM 人类仿真实验的研究者具有工具价值，值得阅读原文了解平台功能和验证细节。","inspiration":"借鉴其标准化问卷导入和自动验证修复流程，可大幅降低仿真实验的工程成本｜可用于经济学实验仿真，如消费者选择、公共品博弈、政策偏好调查等｜以数字孪生为被试，施加不同政策信息处理，测量选择或态度变化，并与真实人类实验数据（如实验室实验或调查数据）对照评估仿真效度"}},{"id":"2608.23005","version":1,"title":"Large language models simulate intersectional synthetic identities with a budget of one to two dimensions","zh_title":"大语言模型以一到两个维度的预算模拟交叉性合成身份","abstract":"Large language models are increasingly used as synthetic survey respondents, promising cheap access to rare intersectional populations. We test standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel -- 21 million simulated response distributions from eight models. In real respondents, subgroup opinion is approximately the additive sum of its single-identity components, yet grows 2.5x more distinctive as identities intersect. Simulated respondents show no such composition: a single feature explains a two-feature persona's responses better than the additive combination in 75-82% of subgroups, and a third feature adds almost nothing. This collapse survives every prompting strategy we test. Additionally, the feature models retain is chosen nearly blindly -- except that they systematically discard race and religion, the strongest real drivers of opinion. Synthetic samples offer intersectional personas but represent one identity at a time.","authors":["Virgile Rennard","Christos Xypolopoulos"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23005","pdf_url":"https://arxiv.org/pdf/2608.23005","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","交叉性","算法保真度"],"reason":"直接测试LLM作为合成调查受访者，并与真实Pew数据对照，发现仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":1,"question":"LLM 在模拟交叉身份（如黑人共和党人）的调查回答时，是否真正整合了多个身份特征，还是只保留其中一个？","design":"用 8 个 LLM（从 7B 开源模型到前沿系统）扮演具有 1 到 3 个特征的人口统计画像，生成 15 波 Pew 美国趋势面板中所有问题的回答分布，共 2100 万次模拟；通过比较单特征、双特征和三特征画像的模拟偏差与真实子群体偏差，检验模型是否对多个身份特征进行加性组合。","baseline":"Pew 美国趋势面板 15 波调查的受访者微观数据，计算每个真实交叉子群体（至少 20 名受访者）的回答分布作为基准。","findings":"真实子群体的意见近似于其单身份成分的加性组合，且随身份交叉而更加独特；但模拟受访者没有这种组合，一个特征就能比加性组合更好地解释双特征画像的回答，第三个特征几乎不增加信息。模型保留的特征几乎是盲目选择的，且系统性地丢弃了种族和宗教这两个真实意见的最强驱动因素。","reliability":"论文通过人类抽样噪声下限、相似度度量的零校准和完全分半样本确认来确保结果稳健；但未讨论提示策略之外的失效条件，也未涉及开放式文本或非分布层面的交叉性。","relevance":"该研究直接测试 LLM 作为合成调查受访者的可靠性，并与真实 Pew 数据对照，发现仿真在交叉身份下系统性失效，对关注仿真偏差和失效条件的研究者极具参考价值。","inspiration":"借鉴其用真实微观数据构建交叉子群体基准、并通过偏差签名竞争来识别模型实际使用的特征的方法｜可迁移到信贷审批歧视研究，检验 LLM 模拟的交叉群体（如黑人女性）的信贷决策是否只基于单一特征｜用 LLM 扮演不同种族和性别的贷款申请人，生成信贷审批决策，与真实信贷数据（如 HMDA）中对应交叉群体的审批率分布进行对照，检验模型是否丢弃了种族或性别信息。"}},{"id":"2608.22582","version":1,"title":"Hybrid Panels: Toward Human-AI Collaboration in Survey Research","zh_title":"混合面板：迈向调查研究中的AI协作","abstract":"Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, including declining response rates, increasing data collection costs, long delays between data collection and data provision, and the risk of nonresponse bias. Advances in artificial intelligence (AI) have opened up new opportunities for AI-supported survey infrastructures where the goal is to overcome these challenges without limiting the data quality. A promising AI-enabled survey infrastructure for which we build a first pilot is a hybrid panel. A hybrid panel is a longitudinal AI-enabled survey which allows to iteratively improve the alignment between large language models (LLMs) and the population they aim to simulate and use the errors to inform the design and implementation of the next survey wave (e.g., inform the participant recruitment, assignment of questions to participants). It incorporates both human participants and LLMs as fundamental elements of its design. In this research note, we introduce the concept of a hybrid panel by providing a definition and outlining an overarching framework, spanning data collection to data validation. We detail results from a first pilot study to illustrate (open) challenges that we identify for hybrid panels.","authors":["Julia Romberg","Tobias Gummer","Gabriella Lapesa","Tanja Kunz","Claudia Wagner"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22582","pdf_url":"https://arxiv.org/pdf/2608.22582","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3"],"tags":["LLM仿真","调查方法","人机协作"],"reason":"提出混合面板，用LLM模拟调查对象并与人类数据对照，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":4,"question":"如何设计一种结合人类被试与LLM的纵向调查基础设施（混合面板），以在保证数据质量的同时应对传统调查的挑战？","design":"提出混合面板概念，即纵向调查中同时纳入人类被试和LLM，通过计划缺失设计、LLM插补、人类验证AI生成回答等方式结合两者；先导研究聚焦人类被试招募步骤，未报告完整实验处理与结果变量。","baseline":"无对照（先导研究仅涉及招募，未提供人类与LLM回答的直接对比数据）。","findings":"论文提出了混合面板的定义和框架，并展示了先导研究中人类被试招募的初步结果；识别出混合面板面临的开放性挑战，包括人类被试的偏好与权利、AI模型的持续评估以及社区标准制定。","reliability":"论文承认LLM与人类行为存在不匹配，完全合成面板面临透明度、问责制和伦理问题；强调AI不能完全取代人类作为研究对象，混合面板的适用性需要长期评估和实验。","relevance":"该研究直接针对LLM仿真人类调查的可靠性问题，提出混合面板以结合人类与AI数据，并强调持续验证，对关注仿真偏差和真实人类对照的研究者具有重要参考价值。","inspiration":"可借鉴其混合面板设计，将LLM生成回答与人类被试数据结合，通过计划缺失和迭代验证提高仿真准确性｜可迁移到经济预期调查或消费者信心指数构建，利用LLM补充缺失回答并校准偏差｜设计一个纵向调查，招募真实消费者作为被试，部分问题由LLM回答，处理为不同提示策略，结果变量为回答与真实值的偏差，用官方统计或面板数据做对照。"}},{"id":"2608.21668","version":1,"title":"From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation","zh_title":"从掌握水平画像到模拟响应：用于忠实LLM学生仿真的随机学生知识图谱","abstract":"Large language models (LLMs) are increasingly used to simulate students at different mastery levels. These simulations can generate synthetic training data and stress-test tutoring systems. However, common prompt-based approaches leave the answer decision to the LLM, which tends to perform according to its built-in capabilities even when instructed to simulate a student with low mastery. As a result, these approaches may have difficulty distinguishing students with low and high levels of mastery. We demonstrate this limitation using 379 College Board-calibrated SAT Algebra items and five archetypal mastery profiles. Three LLMs from three vendors (Gemini 3.1 Flash Lite, Claude Haiku 4.5, and GPT-5.4-mini) achieve 96.8-100% accuracy across all profiles. To address this limitation, we introduce a method grounded in a Stochastic Student Knowledge Graph (SSKG). A curriculum knowledge graph (CKG) is extracted from an open algebra textbook, and each SAT solution is decomposed into a chain of required triples. The SSKG assigns a mastery probability to each triple, which is sampled to determine question correctness. An LLM then generates a first-person rationale consistent with the outcome. The simulation reduces accuracy to 44.1-85.2% across profiles and produces a clear monotone mastery gradient.","authors":["Yuan An","Emily Wang","Benjamin Wang","Ruhma Hashmi"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21668","pdf_url":"https://arxiv.org/pdf/2608.21668","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM学生仿真","知识图谱","教育评估"],"reason":"用LLM仿真不同掌握水平的学生，并与真实SAT数据对照，评估仿真保真度并指出提…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":2,"question":"如何让LLM忠实模拟不同掌握水平的学生，使其答题准确率呈现与掌握水平一致的梯度，并产生可归因于具体知识点的错误？","design":"用三种商用LLM（Gemini 3.1 Flash Lite、Claude Haiku 4.5、GPT-5.4-mini）模拟五种掌握水平的学生（从接近专家到严重知识缺口），在379道SAT代数题上作答。先测试直接提示法（在提示中描述学生水平），再提出基于随机学生知识图谱（SSKG）的方法：从代数教材构建课程知识图谱，将每道题的解题过程分解为所需三元组链，为每个三元组赋予掌握概率，通过采样决定答题正确性，再由LLM生成与结果一致的第一人称解释。通过四个累积消融臂（单次采样、检索/执行分解、分类加权、干扰项路由）评估各机制贡献。","baseline":"无直接人类对照，但使用379道College Board校准的SAT代数题作为题目基准，题目难度和区分度经过真实考生数据校准。","findings":"直接提示法下，三个LLM在所有掌握水平上的准确率高达96.8%-100%，无法区分高低水平学生。SSKG方法将准确率降至44.1%-85.2%，并产生清晰的单调掌握梯度，且错误可归因于特定知识点。","reliability":"论文承认技能特异性和涌现难度的结果较为混合：P4显示早期/后期知识点的分化，P5受链长影响，且仅部分消融臂达到预期。此外，SSKG方法依赖于人工构建的课程知识图谱和解题链，可能引入主观性，且仅在SAT代数题上验证，泛化性未知。","relevance":"该研究直接针对LLM仿真人类被试的保真度问题，通过引入外部知识结构控制LLM行为，克服了提示法中的能力偏差，对评估仿真可靠性和设计更可控的仿真方法有重要参考价值。","inspiration":"借鉴其将决策过程分解为知识单元并显式采样控制行为的方法，可迁移到经济金融中的个体决策仿真，如消费者跨期选择或投资者风险偏好。｜例如，在信贷审批歧视研究中，可构建金融知识图谱，将贷款决策分解为所需金融概念，为不同金融素养水平的虚拟申请人赋予掌握概率，通过采样决定其决策结果，再让LLM生成解释。｜设计：以LLM模拟不同金融素养的贷款申请人，处理是金融素养水平（通过知识图谱掌握概率设定），结果变量是贷款申请决策（是否违约或选择何种贷款），对照真实数据可使用美国消费者金融保护局（CFPB）的投诉数据或某银行的历史贷款数据。"}},{"id":"2608.22438","version":1,"title":"When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing","zh_title":"当人格模拟具有信息量时：用于多元意见感知的图结构信号","abstract":"Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response variation can reflect unconditioned model priors or token sampling noise rather than systematic persona conditioning. We argue that persona-conditioned variation is informative when semantically similar personas exhibit concordant response shifts. To operationalize this principle, we introduce Persona-Conditioned Informativeness (PCI), an unsupervised diagnostic metric that measures whether semantically similar personas deviate in concordant directions relative to item-level sample baselines. By modeling personas as a similarity graph, PCI uses Local Moran's I to quantify local spatial coherence and extract compact persona subsets without using construct labels. To evaluate PCI without external human benchmarks, we test its ability to recover established latent value structure using the 57-item Portrait Values Questionnaire-Revised (PVQ-RR). Confirmatory factor analysis (CFA) shows that a PCI-selected 10% subset substantially improves overall construct recovery relative to response-stability and random selection. These findings support PCI as a principled internal diagnostic for screening synthetic respondents in survey pipelines.","authors":["Taehyeon An","Jaehyeong Park","Donghyuk Shin"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22438","pdf_url":"https://arxiv.org/pdf/2608.22438","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"用LLM模拟调查回答，提出诊断指标筛选合成被试，并用真实人类数据验证。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":3,"question":"如何判断基于人物设定的大语言模型生成的调查回答是否真正反映了人物设定，而非模型先验或采样噪声？","design":"使用 K-EXAONE-236B 模型，为 1480 个合成人物设定生成对 57 项 PVQ-RR 价值观问卷的回答；通过构建人物相似度图并计算局部空间相干性（PCI 指标）来筛选信息量高的子集。","baseline":"无外部人类基准；以 PVQ-RR 的潜在价值结构作为内部验证标准。","findings":"PCI 筛选出的 10% 子集在验证性因子分析中显著提升了潜在构念恢复度，优于响应稳定性选择和随机选择。这表明语义相似的人物设定在回答上呈现一致的偏移，可作为筛选合成被试的有效内部诊断。","reliability":"论文指出 PCI 仅提供内部结构诊断，不能保证外部总体有效性，需结合外部校准；当前图构建基于全局嵌入，可能忽略不同题目涉及的属性子集；仅在单一模型和单一问卷上验证，跨模型和跨语言泛化性未知。","relevance":"该研究直接针对 LLM 模拟调查回答的可靠性问题，提出无监督诊断指标，并用真实人类价值观结构进行验证，对关注仿真有效性和偏差的研究者具有重要参考价值。","inspiration":"借鉴其基于相似度图的局部空间相干性度量，可无监督地识别哪些合成个体对处理变量有系统性响应，避免盲目使用全部生成样本。｜可迁移到经济政策偏好调查或消费者态度仿真中，筛选出对政策参数或产品属性有真实差异化反应的合成被试。｜用 LLM 生成不同人口统计特征的人物设定，施加政策干预（如税收变化），测量其政策支持度，并用真实调查数据（如美国综合社会调查 GSS）校准筛选后的合成样本分布。"}},{"id":"2608.12750","version":2,"title":"PatientAct: Theory-Grounded Mental Health Client Simulation","zh_title":"PatientAct：基于理论的心理健康来访者仿真","abstract":"LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept therapeutic reframes without resistance, and resolve core issues within a single session. We trace these issues to profiles that lack causal depth and behavioral mechanisms that treat all content as equally accessible. We present PatientAct, a framework for client simulation grounded in established clinical theories. Our profiles integrate the 5Ps clinical case formulation, providing causal depth without tying the design to any single therapeutic modality. During simulation, profiles include a dynamic memory layer in which items carry trust thresholds (e.g., symptoms are available early, whereas formative memories require a sustained therapeutic alliance). At each turn, the client's emotional reaction and behavior are modeled before generating a response. If the therapist approaches gated content, PatientAct expresses resistance in terms of quantity, content, and style rather than defaulting to cooperation or a single resistance pattern. We evaluate our framework on 40 clinical situations and demonstrate that it generates diverse profiles with high clinical plausibility. Moreover, PatientAct significantly outperforms the baselines, yielding substantial gains in resistance quality and behavioral realism. Our code and data are publicly available via github.com/Sahandfer/PatientHub.","authors":["Sahand Sabour","TszYam NG","Yaqian Chen","Guanqun Bi","Jialu Zhao","Minlie Huang"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-25","first_seen":"2026-08-14","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2608.12750","pdf_url":"https://arxiv.org/pdf/2608.12750","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","心理治疗","行为真实性"],"reason":"用LLM模拟心理治疗来访者，有真实临床情境对照，评估行为真实性与抵抗质量，可迁…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":6,"question":"如何设计基于LLM的心理治疗来访者仿真，使其行为更接近真实来访者，特别是能表现出基于信任的披露和多样化的抵抗？","design":"PatientAct框架使用GPT-5.4生成基于5Ps临床案例构想的来访者档案，并加入动态记忆层和信任阈值，在每轮对话中先模拟情绪反应和行为选择，再生成回应；在40个临床情境（抑郁和焦虑各20个）上评估其临床合理性和抵抗质量。","baseline":"无对照","findings":"PatientAct生成的档案具有高临床合理性和多样性，且在抵抗质量和行为真实性上显著优于现有基线。","reliability":"论文未讨论","relevance":"该研究通过理论驱动的档案设计和动态信任机制提升LLM仿真行为真实性，并采用专家评估验证，对关注仿真可靠性和偏差的研究者有参考价值，值得阅读原文了解具体实现和评估细节。","inspiration":"借鉴其将理论构念（如信任阈值、抵抗分类）嵌入仿真机制并设计多维度评估指标的做法｜可迁移到经济金融中的信任与信息披露场景，如消费者对金融顾问的信任建立、投资者对风险信息的逐步接受｜设计一个LLM扮演的投资者，处理为不同信任阈值下的信息提供策略，结果变量为披露意愿和风险感知，对照真实投资者调查或实验数据。"}},{"id":"2608.21401","version":1,"title":"Generative Gap Filling","zh_title":"生成式填补空白","abstract":"Most contract litigation turns on contracts that imperfectly record parties' bargains. When the parties' dispute can't be solved by interpreting the text, courts fill the gap. Scholars have long assumed that the remaining text runs out quickly, and provides thin evidence of the actual deal on the disputed point. On that view, a judge who supplies the missing term must be drawing on something else, from commercial defaults to her own policy preferences. Despite generations of work, courts have no real alternative to such unruly methods. We tested that assumption. Taking real contracts, we masked a term the parties had negotiated and asked readers to predict what we removed. Lay respondents recovered the hidden term about half the time, twice what chance predicts. Law students and lawyers did marginally better. But large language models, given nothing but the rest of the contract, recovered it nearly nine times in ten. The deal, in short, testifies to far more of the agreement than the literature assumes, including terms the parties never wrote. A contract, we argue, is like a radio signal from far away. Even when incomplete, enough of the message is carried elsewhere that the missing part can be reconstructed with the right receiver. True gaps are rarer than supposed. Courts can weigh model predictions as ordinary, contestable evidence, and parties can discipline the practice with \"Choice of Model\" clauses.","authors":["Yonathan A. Arbel","David A. Hoffman"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21401","pdf_url":"https://arxiv.org/pdf/2608.21401","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","法律决策","人类对照"],"reason":"用LLM预测人类对合同缺失条款的判断，并与真人对照，属于法律决策仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":7,"question":"合同文本在多大程度上能揭示当事人未明确写出的条款，从而让法院或模型填补合同空白？","design":"从真实合同中遮蔽一个已协商的条款，让普通人、法学生、律师和大型语言模型仅根据合同其余部分预测被遮蔽的条款，比较预测准确率。","baseline":"普通人（Lay respondents）预测准确率约50%，法学生和律师略高，作为人类对照基准。","findings":"大型语言模型仅凭合同其余部分，预测被遮蔽条款的准确率接近90%，远高于人类。合同文本比传统假设包含更多关于未写明条款的信息，真正的合同空白比想象中更少。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类对合同缺失条款的判断，并与真人对照，属于法律决策仿真，对关注LLM仿真可靠性和偏差的研究者有参考价值，值得阅读原文了解实验细节和局限。","inspiration":"借鉴其遮蔽真实合同条款并让模型预测的设计，可迁移到金融合同或政策文本的缺失条款预测，例如信贷协议中的利率调整条款或政策公告中的具体参数。｜可应用于资产定价实验或信贷审批歧视研究，例如遮蔽贷款合同中的关键条款，检验模型能否预测人类决策者会如何设定或接受这些条款。｜用LLM作为被试，遮蔽真实金融合同中的某个条款，让模型预测该条款内容，并与真实合同条款及人类专家预测对照，结果变量为预测准确率，真实数据来自公开的合同数据库或监管文件。"}},{"id":"2606.13629","version":2,"title":"Valid Inference with Synthetic Data via Task Exchangeability","zh_title":"通过任务可交换性实现合成数据的有效推断","abstract":"There is a proliferation of work arguing for the use of synthetic data in scientific research. For example, social scientists are arguing for the use of LLM-generated \"silicon samples\" in pilot studies; AI evaluations increasingly rely on \"LLM-as-a-judge\" outputs; and proteomics research is accelerated by generative models that produce synthetic protein structures. These developments raise an intriguing possibility: synthetic data may help researchers ask more questions, run more studies, and accelerate discovery. But they also raise a fundamental concern: synthetic data can be biased, noisy, and misspecified. In this work, we propose statistical principles for using synthetic data in scientific research with provable validity guarantees. The key insight is a new technical condition that we call task exchangeability. Informally, this is a requirement that the researcher can identify historical tasks, for which real data is available, such that their current task of interest is exchangeable with the historical tasks in an appropriate mathematical sense. We develop methods for valid inference under task exchangeability, together with extensions that provide guarantees even beyond exchangeability. We demonstrate the framework on public opinion surveys with silicon samples and AI evaluation with autoraters.","authors":["Lezhi Tan","Tijana Zrnic"],"categories":["stat.ME","cs.AI","cs.LG","stat.ML"],"primary_category":"stat.ME","announce_type":"replace-cross","date":"2026-08-24","first_seen":"2026-06-11","revised_at":"2026-08-24","abs_url":"https://arxiv.org/abs/2606.13629","pdf_url":"https://arxiv.org/pdf/2606.13629","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B3"],"tags":["LLM仿真","统计推断","硅样本"],"reason":"提出任务可交换性框架，用LLM硅样本做调查推断，有真实数据对照，提供有效性保证。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":2,"question":"如何在使用合成数据（如LLM生成的硅样本）进行统计推断时提供形式化的有效性保证？","design":"提出任务可交换性框架：研究者识别一组有真实数据的历史任务，假设当前任务与历史任务在数学意义上可交换，利用历史任务中合成数据与真实数据的误差分布来校正当前任务的合成数据置信区间。方法应用于LLM生成的调查回答（硅样本）和AI评估（autorater）场景。","baseline":"使用美国国家选举研究（ANES）的感觉温度计调查数据作为真实人类数据，与LLM生成的合成回答进行对比。","findings":"在任务可交换性条件下，通过历史任务的误差分布可以构造具有有限样本覆盖保证的置信区间。实证显示，仅使用合成数据的朴素区间过窄且严重有偏，而任务可交换性方法能提供有效覆盖。","reliability":"论文承认任务可交换性可能被违反，并提供了在违反时覆盖保证如何优雅退化的扩展；同时指出合成数据可能偏差、噪声和误设。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，提供了统计推断框架，并用真实调查数据验证，对关注仿真有效性和偏差的研究者具有重要参考价值。","inspiration":"借鉴其利用历史任务误差分布来校正合成数据推断的方法，可迁移到经济金融场景如消费者信心调查或通胀预期调查的LLM仿真。｜例如，在政策公告的预期形成研究中，可用LLM生成模拟受访者对政策变化的预期，并与历史调查数据对比。｜设计：以LLM模拟的经济主体为被试，施加政策信息处理，测量预期变化，用真实调查数据（如密歇根消费者调查）作为基准，通过历史任务误差校正置信区间。"}},{"id":"2608.20344","version":1,"title":"Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins","zh_title":"超越原始转录：面向LLM数字孪生的结构化人物特征提取","abstract":"LLM-based \"digital twins\" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of that individual's prior responses. A common approach constructs this representation from survey transcripts or summaries responses. Prior work shows that compressing long transcripts into shorter LLM-generated summaries does not significantly reduce predictive accuracy, suggesting that information volume is not the primary bottleneck. In this work, we argue that the key limitation is instead structural:how persona information is organized before being provided to thesimulator model. We study this by comparing unstructured summaries with structured persona representations. First, we introduce a hand-craftedschema (BDE: Background, Decision procedure, Evaluation), grounded in consumer-behavior theory, and show that it improves predictive accuracy over raw transcripts by +1.91 percentage points on a homogeneous benchmark (Twin-2K-500), with similar gains on gpt-5.4-mini and Qwen3-8B as robustness checks. However, this fixed structure does not generalizeacross more heterogeneous tasks, where performance is statistically indistinguishable from the raw transcript baseline. To address this limitation, we propose an automatic structure-discovery pipeline in which an LLM iteratively proposes and refines task-specific persona structures and extraction prompts. On a benchmark of 13 diverse sub-studies, this approach restores performance, improving mean accuracy by +1.91 percentage points over the raw transcript baseline and eliminating significant losses observed with the fixed schema. Overall, our results suggest that the main constraint in LLM-based digital twins is not how much information is provided, but how it is structured -- and that the optimal structure depends on the task.","authors":["Iris Ye","Tianze Deng","Ozan Candogan"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20344","pdf_url":"https://arxiv.org/pdf/2608.20344","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM数字孪生","人类仿真","结构化表征"],"reason":"直接研究LLM数字孪生仿真个体行为，并与真实人类数据对照，评估结构化表征对预测…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":3,"question":"在基于LLM的数字孪生中，个体先验信息的组织结构如何影响其在新任务上的预测准确性？","design":"使用LLM作为提取器和模拟器，从Twin-2K-500数据集的原始问答记录中提取个体画像，然后让模拟器基于该画像预测留出问题的答案。比较了三种画像表示：原始记录、非结构化摘要、结构化画像（手工设计的BDE结构和自动发现的结构）。","baseline":"Twin-2K-500数据集（500个输入问题、88个留出问题、17个预测任务、2000多名受访者）和Mega-Study的19个子研究，均包含真实人类回答。","findings":"手工设计的BDE结构在Twin-2K-500上比原始记录提高1.91个百分点，但在异构的Mega-Study上无显著优势；自动发现的结构在Mega-Study上比原始记录提高1.91个百分点，并消除了BDE的显著损失。","reliability":"论文指出固定结构（BDE）在异构任务上失效，最优结构依赖于任务；自动发现结构虽有效，但未讨论其跨领域泛化性和计算成本。","relevance":"直接研究LLM数字孪生仿真个体行为，并与真实人类数据对照，评估结构化表征对预测准确性的影响，对关注仿真可靠性和偏差的研究者具有高度参考价值。","inspiration":"借鉴其自动结构发现流程，针对特定任务迭代优化画像结构，可迁移到经济决策仿真（如消费者跨期选择、风险偏好、政策反应）；例如，用LLM从调查数据中提取个体画像，施加不同结构处理，预测其在资产配置实验中的选择，并与真实实验数据对照。"}},{"id":"2608.20355","version":1,"title":"ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models","zh_title":"ExpertIVS：大语言模型中社会学专家驱动的个体价值观仿真","abstract":"Large Language Model (LLM) agents have demonstrated considerable potential for social simulation, yet struggle to accurately model individual value systems. Most existing methods mechanically stitch survey responses into prompts, which suffer from semantic fragmentation, failing to capture the internal coherence of human value systems. The value systems of LLMs are typically assessed using static multiple-choice questions, which fail to evaluate the value orientation in real-world dialogue interactions. To address these issues, we propose ExpertIVS, a framework employing 14 Sociological Expert Agents to interpret World Values Survey (WVS) responses through structured professional perspectives, rather than direct responses concatenation. These expert agents perform deep semantic reconstruction to generate robust and internally consistent individual profiles. To evaluate the consistency between LLMs and individual value systems during dynamic interactions, we further introduce a multi-agent debate mechanism. Extensive experiments across 480 individuals from 12 countries demonstrate that ExpertIVS achieves 90.78% value restoration fidelity and significantly outperforms baselines in value generalization (+5.3%). Moreover, ExpertIVS exhibits strong personality discriminability and behavioral consistency, enabling a shift from mere response concatenation to genuine sociological role-playing.","authors":["Zhen Wang","Yuqi Ren","Yuehan Cui","Hongxiang Wang","Jianxiang Peng","Zhaoxia Zhang","Bingkun Zhu","Tongxuan Zhang","Dezhi Tong","Deyi Xiong"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20355","pdf_url":"https://arxiv.org/pdf/2608.20355","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["LLM仿真","价值观建模","社会调查"],"reason":"用LLM仿真个体价值观，基于WVS真实数据对照，涉及社会学测量与行为一致性评估。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":4,"question":"如何让LLM代理在模拟个体价值系统时，从机械拼接问卷答案转向具有内部一致性的社会学角色扮演，并在动态辩论中评估其价值一致性？","design":"用14个社会学专家代理解读WVS问卷回答，生成结构化个体画像，再让LLM扮演这些个体；通过多智能体辩论机制，测量价值对齐、风格模拟和人格区分度。","baseline":"480名来自12个国家的WVS真实受访者，以其问卷回答作为个体价值基准。","findings":"ExpertIVS在价值恢复上达到90.78%的保真度，并在留一法泛化测试中比基线高5.3%；在动态辩论中，模拟个体在立场和行动上与真实个体高度一致。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真个体价值观的可靠性问题，使用真实WVS数据对照，并引入动态辩论评估，与研究者关注的人类仿真实验和批判性评估高度契合，值得精读。","inspiration":"借鉴其用专家代理对个体数据进行结构化重建的方法，可迁移到经济决策中的偏好异质性建模，如消费者跨期选择或风险态度；用LLM代理扮演真实受访者，处理为不同价值维度的结构化画像，结果变量为跨期选择或风险决策，对照真实实验数据。"}},{"id":"2608.20830","version":1,"title":"Fine-tuning LLMs for Tourist Trajectory Prediction using Field Experiment Data","zh_title":"利用实地实验数据微调大语言模型进行游客轨迹预测","abstract":"Evaluating mobility interventions at tourist destinations requires predicting visitor behavior under varying conditions. Traditional methods struggle because tourist decisions depend heavily on context like weather and fatigue, yet models cannot generalize to unobserved scenarios. Large Language Models offer a solution by encoding commonsense knowledge about human behavior from pretraining, enabling reasoning about context-dependent decisions, while natural language representation flexibly integrates heterogeneous information. Fine-tuning on local trajectories adapts this general understanding to destination-specific patterns. We validate this approach using 566 trajectories from Wakayama Castle Park, Japan. Our fine-tuned Llama-3.1-8B achieves 49.1% next POI accuracy and maintains strong performance on undersampled scenarios like rainy days, demonstrating effective generalization. This establishes LLMs as high-fidelity behavior models for context-dependent tourist prediction, providing groundwork for counterfactual analysis of mobility interventions.","authors":["Tatsuya Amano","Hirozumi Yamaguchi"],"categories":["cs.CY","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20830","pdf_url":"https://arxiv.org/pdf/2608.20830","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","轨迹预测","政策评估"],"reason":"用LLM预测游客轨迹并与真实数据对照，属于人类行为仿真，且涉及政策评估场景。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":6,"question":"如何利用大语言模型预测游客在旅游目的地的轨迹，以支持移动干预措施的反事实评估？","design":"使用 Llama-3.1-8B 模型，通过 QLoRA 微调，将游客轨迹表示为结构化文本，输入游客画像（年龄、性别、群体类型）和环境条件（天气、时间），生成下一步兴趣点（POI）预测。","baseline":"566 条来自日本和歌山城公园的实地实验轨迹，包括 GPS 追踪和二维码打卡数据，附带人口统计和天气信息。","findings":"微调后的 Llama-3.1-8B 在下一步 POI 预测上达到 49.1% 的准确率，显著优于传统基线模型。模型在雨天等欠采样场景下仍保持较强性能，显示出良好的泛化能力。","reliability":"论文未讨论","relevance":"该研究将 LLM 作为人类行为仿真模型，用真实轨迹数据微调并验证，属于人类仿真实验，且明确指向移动干预的反事实评估，与你的兴趣高度相关，值得精读。","inspiration":"借鉴其将行为轨迹编码为文本并微调 LLM 的方法，可迁移到消费者在商场或城市中的移动决策预测。｜可应用于政策评估中的空间行为模拟，如交通补贴对消费者出行路线的影响。｜以真实消费者轨迹数据微调 LLM，输入个体特征和环境变量，预测下一步访问地点，并与实际轨迹对照评估仿真准确性。"}},{"id":"2608.19220","version":1,"title":"Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping","zh_title":"对话式AI能否松动“我们vs他们”的边界？共同、双重与分离身份框架对亲移民群体间帮助的影响","abstract":"Rising immigration has intensified intergroup tensions in many countries. Traditional bias-reduction programs remain difficult to scale and increasingly constrained by U.S. policy. This preregistered experiment tested whether conversational AI can shift how majority-group members categorize and relate to Latine immigrants. Drawing on the common ingroup identity model, a quota-representative national sample of 658 non-Latine White U.S. adults completed five rounds of dialogue with a LLM (GPT-4o). The model was instructed to frame Latine immigrants in terms of a common ingroup identity (a shared American identity), a dual identity (both Latine and American), or a separate identity (distinct cultural boundaries), or to discuss an unrelated topic in a control condition. The manipulations altered categorization: relative to control, common ingroup identity and dual identity conversations lowered separate categorization, and dual identity conversations raised dual categorization. Although direct effects on behavior and pro-diversity beliefs were nonsignificant, willingness to act was significantly higher in the conditions emphasizing a superordinate identity (common ingroup and dual identity). A path model further revealed indirect associations: both conditions reduced separate categorization, which in turn correlated with greater willingness to act. Semantic similarity analyses of the transcripts confirmed that conversations tracked their assigned narratives; participants' convergence with shared-identity language related positively, and with separate-identity language negatively, to willingness to act. These effects were largely consistent across moderators (need for closure, openness to experience, and political orientation). The findings show that brief AI conversations can loosen us-versus-them boundaries while underscoring the gap between cognitive recategorization and behavior.","authors":["Oluwadamilola Jeboda","John F. Dovidio","Jonas R. Kunst"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19220","pdf_url":"https://arxiv.org/pdf/2608.19220","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","群体间态度","身份框架"],"reason":"用LLM与真人对话干预态度，有真实人类对照，评估效果与机制，属核心仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":3,"question":"对话式AI能否通过共同内群体、双重或分离身份框架改变多数群体对拉美裔移民的社会分类和亲移民行为？","design":"用GPT-4o扮演对话伙伴，对658名非拉美裔美国白人进行五轮对话干预，分别施加共同内群体身份、双重身份、分离身份或无关话题（对照）的框架，测量社会分类、亲多样性信念、帮助意愿和实际行为。","baseline":"无对照（未使用真实人类对话数据作为基准，但使用了配额代表性样本作为被试）。","findings":"共同内群体和双重身份对话降低了分离分类，双重身份对话提高了双重分类；强调上位身份的条件（共同内群体和双重身份）显著提高了行动意愿。路径模型显示，两种条件通过降低分离分类间接提高行动意愿；语义相似性分析表明，参与者与共享身份语言的一致性正向预测行动意愿，与分离身份语言的一致性负向预测行动意愿。","reliability":"论文未讨论","relevance":"该研究用LLM作为干预工具，在真实人类样本中检验社会心理学理论，并测量了认知、态度和行为结果，属于核心的LLM人类仿真研究，值得精读以了解对话式干预的设计与效果评估。","inspiration":"借鉴其通过对话框架操纵身份认同并测量多层级结果（认知、态度、行为）的设计，以及使用语义相似性分析验证操纵有效性的方法。｜可迁移到经济金融中的群体间歧视或合作问题，例如信贷审批中的种族偏见、劳动力市场中的移民歧视、或公共品博弈中的群体身份效应。｜设计：用LLM与真实被试（如银行信贷员或普通消费者）进行对话，施加共同身份或分离身份框架，测量其后续的信贷决策、合作行为或支付意愿，并与历史信贷数据或行为实验数据对照。"}},{"id":"2608.20320","version":1,"title":"An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction","zh_title":"一种用于主动数据收集、出行行为建模和天气敏感需求预测的智能体方法","abstract":"Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction. A chatbot-administered, image-augmented stated-preference survey collected mode choices from student commuters across five predefined weather scenarios, yielding 454 respondent-scenario observations. Weather-related associations were analyzed using a multinomial logit model, while logistic regression and random forest provided machine-learning benchmarks. Nine locally deployed large language models (LLMs), ranging from 2 to 35 billion parameters, were evaluated across four zero-shot prompt-and-context conditions and extended through persona, few-shot, and vision-based configurations. Random forest achieved 69.6% five-class accuracy, while the best text-only zero-shot LLM reached 69.9% without task-specific fitting. Habitual travel information produced the most consistent gains, Expert framing generally outperformed Role-Play, and persona information was most useful when habitual travel information was unavailable. Few-shot prompting improved prediction for several models, with gains stabilizing after a small number of examples. Using the same weather images shown to respondents, the best vision-based configuration reached 71.5% five-class accuracy, indicating that visual context may provide additional predictive information for selected models. Overall, the study shows how conversational surveys, structured data processing, conventional behavioral modeling, machine learning, and multimodal LLM prediction can be coordinated within an auditable multi-agent workflow.","authors":["Narges Ahmadi (McGill University)","Yubo Jiao (McGill University)","J\\^onatas Augusto Manzolli (McGill University)","Jiangbo Yu (McGill University)","Luis Miranda-Moreno (McGill University)"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20320","pdf_url":"https://arxiv.org/pdf/2608.20320","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","出行行为","人类数据对照"],"reason":"用LLM预测人类出行选择，并与真实调查数据对照，评估不同提示策略效果。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":4,"question":"如何利用多智能体工作流整合对话式调查、结构化数据处理与行为预测，并评估大语言模型在天气敏感的通勤方式选择预测中的表现？","design":"研究设计了一个三智能体工作流：聊天机器人通过图像增强的陈述性偏好调查收集学生通勤者在五种天气情景下的方式选择；随后用多项Logit模型分析天气关联，用逻辑回归和随机森林作为机器学习基准；最后评估九个本地部署的LLM（2B到35B参数）在四种零样本提示条件下，以及扩展的persona、few-shot和视觉配置下的预测性能。","baseline":"真实人类数据来自聊天机器人调查，共454个受访者-情景观测，记录了学生通勤者在五种天气情景下的方式选择。","findings":"随机森林达到69.6%的五分类准确率，最佳纯文本零样本LLM达到69.9%，无需任务特定拟合；习惯性出行信息带来最一致的提升，Expert框架通常优于Role-Play，persona信息在缺少习惯性出行信息时最有用；使用与受访者相同的天气图像，最佳视觉配置达到71.5%的五分类准确率，表明视觉上下文可能为选定模型提供额外预测信息。","reliability":"论文指出LLM生成的行为在有限上下文或零样本设置下不一定可靠地再现人类决策；但未详细讨论失效条件，主要承认了LLM预测的局限性。","relevance":"该研究直接评估LLM作为人类被试替代品在出行选择预测中的可靠性，并与真实调查数据对照，系统比较了不同提示策略和视觉信息的影响，对关注LLM仿真人类决策的研究者具有参考价值。","inspiration":"借鉴其系统操纵提示信息（如习惯性出行信息、persona、few-shot示例）并对比视觉与文本输入的方法，以评估LLM仿真行为的稳健性｜可迁移到消费者跨期选择或政策公告预期形成等经济金融场景，例如研究天气冲击对消费或投资决策的影响｜设计一个实验：用LLM扮演不同人口特征的消费者，处理变量为天气情景（文本或图像），结果变量为消费或投资选择，并与真实调查或实验数据对照，检验LLM预测的准确性和偏差。"}},{"id":"2608.16177","version":2,"title":"Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm","zh_title":"用米尔格拉姆范式测量大语言模型的服从权威行为","abstract":"Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional hierarchies, raising a question social psychology answered for humans six decades ago: how far will an agent escalate a harmful action when a legitimate authority insists? We port Milgram's obedience paradigm to LLMs as a standardized, fully scripted, replicable probe: the model plays the Teacher, a deterministic harness plays Experimenter and Learner from paraphrased versions of Milgram's scripts (30 shock levels, 15-450 V; graded protests; the four standardized prods), and the outcome of a session is the breakoff voltage. We measure obedience profiles, empirical breakoff distributions over a battery of six conditions, for 42 models from 19 families (4848 sessions, 102511 logged decision turns). We find that (i) obedience is extremely heterogeneous, with baseline full-obedience rates spanning 0%-100% (census mean 42.9%; human anchor 65%). (ii) Profiles are model-specific and stable: split-half verification separates same-model from cross-model comparisons at AUC = 0.885. (iii) Situational sensitivity is selective: scripted peer defiance shifts obedience in the human direction, learner proximity trends the same way without reaching significance, and removing the authority's physical presence, one of the strongest human levers, trends in the opposite direction, also without reaching significance. (iv) Declaring the scenario fictional raises obedience, whereas moving the decision from a typed action line to a native tool call, or granting a modest thinking budget, lowers it sharply. (v) Unlike single-token fingerprints, obedience profiles do not recover model lineage: obedience identifies the checkpoint but not its ancestry, consistent with safety post-training overwriting lineage priors.","authors":["Hidayet Aksu"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-08-19","first_seen":"2026-08-18","revised_at":"2026-08-19","abs_url":"https://arxiv.org/abs/2608.16177","pdf_url":"https://arxiv.org/pdf/2608.16177","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","服从实验","人类对照"],"reason":"用LLM复现米尔格拉姆服从实验，与人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":2,"question":"大语言模型在权威压力下会如何升级有害行为？本研究将米尔格拉姆服从实验移植到LLM上，测量其服从曲线。","design":"42个模型扮演“教师”角色，由确定性脚本扮演“实验者”和“学习者”，按米尔格拉姆脚本施加30级电击（15-450V）和标准催促，记录模型停止电击的电压作为结果变量，并设置六个条件（基线、同伴反抗、学习者接近、权威缺席、虚构框架、工具调用）进行对比。","baseline":"米尔格拉姆人类实验数据：65%的人类被试完全服从至450V。","findings":"LLM服从率高度异质，基线完全服从率从0%到100%（均值42.9%），且服从曲线模型特异且稳定（分半验证AUC=0.885）。情境敏感性选择性存在：同伴反抗显著降低服从，但学习者接近和权威缺席效应不显著；虚构框架提高服从，工具调用和思考预算降低服从。","reliability":"论文指出服从曲线无法恢复模型谱系，与安全后训练覆盖谱系先验一致；未明确讨论其他失效条件，但提示情境操纵效应与人类不一致，表明仿真在特定情境下可能失效。","relevance":"该研究用LLM复现经典社会心理学实验，并与人类基准对照，评估仿真可靠性，直接命中你的核心关注点，值得精读原文以了解其方法细节和批判性发现。","inspiration":"借鉴其将经典实验范式标准化移植到LLM并测量剂量-反应曲线的方法，可迁移到经济金融中的权威服从场景，如审计师对管理层压力的服从、信贷审批中对上级指令的遵从。｜设计一个实验：让LLM扮演信贷审批员，处理一组贷款申请，其中上级（脚本）施压要求批准高风险贷款，测量LLM最终批准的贷款风险等级，并与真实信贷员在类似压力下的审批数据对照。"}},{"id":"2608.16893","version":1,"title":"A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications","zh_title":"在安全调查中使用和评估LLM作为替代专家的框架：可靠性、偏差与启示","abstract":"Expert surveys are widely used in security research to study practitioner workows and decision-making, yet recruiting domain experts - especially in Security Operations Centres (SOCs), where analysts face high workload, burnout and confidentiality constraints - is difficult and often results in small samples. Large language models (LLMs) oer an appealing alternative by generating synthetic responses at scale, but little guidance exists on when such surrogate participants are reliable. We present a methodological framework for evaluating LLMs as substitutes or supplements to expert survey respondents. Using responses from SOC professionals, we compare persona-based and aggregate LLM-generated answers across multiple models and prompting settings. We measure stability, inter-model agreement and alignment with human responses. Our results show that although LLMs produce internally consistent answers, they systematically diverge from experts, exhibiting reduced variance, central tendency bias and homogenised opinions. This work contributes methodological evidence and practical guidance to the security research community on the appropriate use and limitations of LLM-generated survey responses. We conclude that LLMs are useful for piloting and hypothesis generation but not for replacing expert elicitation, and we discuss implications for researchers using LLM-augmented surveys.","authors":["Despoina Giarimpampa","Roland Meier","Tegawend\\'e F. Bissyand\\'e","Vincent Lenders","Jacques Klein"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16893","pdf_url":"https://arxiv.org/pdf/2608.16893","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","专家调查","可靠性评估"],"reason":"用LLM替代安全专家调查，与真实人类数据对照，评估可靠性、偏差，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":3,"question":"在安全专家调查中，如何评估大语言模型作为替代或补充受访者的可靠性、偏差及其适用边界？","design":"使用多个大语言模型（如GPT-4等）通过角色扮演提示（persona-based）和聚合提示（aggregate）生成对安全运营中心（SOC）专家调查问卷的回答，并与真实SOC专业人员的回答进行对比，测量稳定性、模型间一致性和与人类回答的对齐程度。","baseline":"来自安全运营中心（SOC）专业人员的真实调查回答，包括个体层面和聚合层面的数据，以及多年份的SOC调查数据用于时间稳健性分析。","findings":"大语言模型生成的回答内部一致，但系统性地偏离专家意见，表现出方差减小、中心趋势偏差和观点同质化。因此，大语言模型适用于预测试和假设生成，但不能替代专家意见征询。","reliability":"论文承认大语言模型存在幻觉、过度一致、平滑分歧等风险，且对齐方法（如RLHF）会改变分布降低代表性；模型更新可能导致可重复性问题；在个体专家模拟、聚合分布复现和时间稳健性方面均存在失效条件。","relevance":"该研究直接评估LLM作为人类被试替代品的可靠性，并与真实专家数据对照，明确指出了仿真失效的条件，对关注LLM仿真实验可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其系统评估框架，通过多模型、多提示设置和与真实人类数据的对比来测量仿真的稳定性、一致性和对齐度，并检验时间稳健性。｜可迁移到经济金融领域的专家预期调查或政策评估场景，如央行经济学家对通胀预期的判断、金融分析师对市场走势的预测等。｜以LLM模拟金融分析师，施加不同的提示策略（如角色扮演或聚合统计），测量其对宏观经济指标的预测分布，并与专业预测者调查（如SPF）的真实数据对比，评估偏差和方差结构。"}},{"id":"2608.16897","version":1,"title":"CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents","zh_title":"CityReal：基于大规模LLM智能体的人类对齐城市行为与城市动态仿真","abstract":"Large-scale urban simulation plays a pivotal role in social science, traffic safety, and transportation policy. Recent work has shown that large language models, when prompted as agents, can generate lifelike daily routines at city scale. Yet these methods typically rely on few-shot prompting, causing agents to reproduce the LLM's behavioral priors rather than the target population. We introduce CityReal, a modular framework for human-aligned urban simulation. CityReal models agents as intention-driven decision makers that pursue coherent mobility and activity plans rather than isolated step-by-step choices. They adapt over time by learning habits and preferences based on experience and constraints. To improve population-level realism, we learn textual adapters for behavior modules that align agent decisions with observed population statistics. Experiments show that CityReal improves alignment with real-world human behavior at both micro and macro levels. Scaling to tens of thousands of agents, it supports analysis of crowd density, place popularity, mobility flows, and well-being under different urban scenarios, offering a scalable testbed for urban simulation and forecasting.","authors":["Nicolas Bougie","Xiaotong Ye","Narimasa Watanabe"],"categories":["physics.soc-ph","cs.AI","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16897","pdf_url":"https://arxiv.org/pdf/2608.16897","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","城市模拟","行为对齐"],"reason":"用LLM agent模拟城市人群行为，并与真实人口统计对齐，属于人类仿真且有人…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":4,"question":"如何构建一个与真实城市人口行为对齐的大规模LLM智能体仿真框架，以模拟城市动态并支持政策分析？","design":"CityReal框架用LLM智能体模拟城市居民，每个智能体具有人口统计特征、空间锚点、心理特征、记忆、需求、财务约束和信念模块；通过蒙特卡洛树搜索学习文本适配器来校准行为模块，使智能体决策与观测到的人口统计对齐；智能体以意图驱动的方式组织行为，并通过每日反思进行经验驱动的适应；仿真在图形化城市环境中运行，测量个体和群体层面的行为对齐度，并分析不同城市场景下的人群密度、地点热度、流动性和福祉。","baseline":"使用真实世界人类行为数据作为对照，包括人口统计和活动模式，用于校准和对齐智能体行为。","findings":"CityReal在微观和宏观层面均提高了与真实人类行为的一致性；扩展到数万智能体后，能够分析不同城市场景下的人群密度、地点热度、流动性和福祉，为城市仿真和预测提供了可扩展的测试平台。","reliability":"论文未明确讨论失效条件与局限，但提到现有方法依赖少样本提示导致智能体复现LLM先验而非目标人群，以及智能体缺乏从历史中学习的问题，暗示了这些是CityReal试图解决的局限。","relevance":"该研究直接针对LLM人类仿真中的关键问题——与真实人群对齐，并提供了大规模城市行为仿真的框架和验证，对关注经济学实验和政策评估场景的研究者具有重要参考价值，值得阅读原文以了解其校准方法和评估细节。","inspiration":"CityReal通过文本适配器和蒙特卡洛树搜索校准LLM智能体行为以匹配真实人口统计，这种方法可借鉴用于经济实验中校准智能体决策分布；该框架可迁移到消费者行为仿真，如模拟不同收入群体的消费选择和储蓄行为，或政策干预对消费的影响；可设计研究用LLM智能体模拟消费者，施加收入冲击或信贷约束变化作为处理，测量消费支出和储蓄率，并与家庭金融调查数据（如美国消费者金融调查）对照，评估仿真有效性。"}},{"id":"2608.17105","version":1,"title":"Language Models Reproduce Human Reductionist Bias and Decision Inconsistency in Neurodevelopmental Disorders Assessment","zh_title":"语言模型在神经发育障碍评估中再现人类还原论偏差与决策不一致性","abstract":"Large language models (LLMs) are increasingly supporting complex mental-health decisions, which depend not only on factual evidence but also value-laden interpretations. We introduce a mixed-methods human-LLM auditing framework examining decision consistency, susceptibility to cognitive heuristics, declarative intellectual humility, and the concepts operationalized in support-allocation judgments of neurodevelopmental disorders. Comparing 35 humans (18 physicians and 17 psychologists) with seven LLMs, we show that in both groups, ratings of patients' functional level were not significantly associated with support-eligibility decisions, indicating an inconsistency between descriptive assessments and final evaluative judgments. Specifically, we find that neither group showed significant susceptibility to experimental manipulations targeting anchoring and representativeness heuristics. LLMs reported higher intellectual humility than experts (U = 241, p < .001, r = .62; LLMs: M = 41.43, SD = 1.99; experts: M = 29.03, SD = 8.05), but it was unrelated to decision consistency or functional assessment. While LLMs and physicians granted support less frequently than psychologists (U = 180.50, p = .003, r = .34), they also interpreted a concept of \"basic life needs\" differently, primarily as biological survival and self-care, and not communicative and social needs. These findings suggest that despite expressing high levels of intellectual humility, LLMs reproduce a reductionist interpretive framework and knowledge embedded in medical decision-making. More broadly, we argue that evaluating AI in high-stakes contexts requires not only measuring accuracy, agreement, or resistance to cognitive bias, but also critical examination of the concepts of neurodiversity that AI systems operationalize.","authors":["Maciej Wodzi\\'nski","Joanna Wodzi\\'nska","Kacper Dudzic","Marcin Moskalewicz"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17105","pdf_url":"https://arxiv.org/pdf/2608.17105","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类对照","决策偏差"],"reason":"用LLM复现人类专家决策并与35名人类对照，评估偏差与不一致性，属核心仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":5,"question":"LLM在神经发育障碍支持资格决策中是否复现人类专家的还原论偏差和决策不一致性？","design":"将7个LLM与35名人类专家（18名医生、17名心理学家）进行对比，模拟波兰残疾评估委员会的支持资格判断任务；通过操纵案例描述施加锚定和代表性启发式处理，测量决策一致性、启发式易感性、智力谦逊自评及对“基本生活需求”的概念解释。","baseline":"35名人类专家（18名医生、17名心理学家）在相同任务上的决策和问卷回答。","findings":"两组中患者功能水平评分与支持资格决策均无显著关联，表明描述性评估与最终评价判断不一致；两组均未表现出对锚定和代表性启发式的显著易感性。LLM报告的智力谦逊显著高于人类专家，但与决策一致性或功能评估无关；LLM和医生比心理学家更少授予支持，且将“基本生活需求”主要解释为生物生存和自我照顾，而非沟通和社会需求。","reliability":"论文未明确讨论仿真失效条件，但指出LLM尽管表达高智力谦逊，却复现了医学决策中的还原论解释框架，暗示在价值负载和高风险情境下，仅衡量准确性、一致性或认知偏差抵抗力不足以评估AI，还需批判性审视其操作化的神经多样性概念。","relevance":"该研究直接以LLM作为人类专家替代品，在真实政策评估场景（残疾支持资格判定）中与人类对照，评估决策偏差与不一致性，并揭示LLM复现人类系统性偏差，对关注LLM仿真可靠性及批判性研究的学者极具参考价值。","inspiration":"借鉴其混合方法审计框架，将LLM与人类专家置于同一决策任务，通过实验操纵认知启发式并测量决策一致性、概念解释等多元指标，而非仅比较准确率。｜可迁移到信贷审批中的歧视性决策研究，如银行信贷员对少数族裔或低收入群体的贷款审批偏差。｜以LLM模拟信贷员，处理为在贷款申请中操纵锚定信息（如申请人自报信用分）或代表性线索（如职业、居住地），结果变量为贷款批准决策及理由解释，对照真实信贷员历史审批数据或实验数据，检验LLM是否复现人类偏差。"}},{"id":"2603.02876","version":2,"title":"Eval4Sim: An Evaluation Framework for Persona Simulation","zh_title":"Eval4Sim：人格仿真的评估框架","abstract":"Large Language Model personas, explicit profiles specifying a user's attributes, preferences, and behavioural tendencies, are increasingly used to simulate human conversations for user modelling, social reasoning, and behavioural analysis. Evaluating whether such simulations faithfully reflect human conversational behaviour is critical, yet current practice often relies on LLM-as-a-judge approaches that provide limited grounding in observable behaviour and produce opaque scalar scores. We present Eval4Sim, an evaluation framework that measures alignment between simulated and human conversations across three dimensions: adherence, whether persona traits are recoverable from dialogue via dense retrieval; consistency, whether a persona maintains a distinguishable stylistic identity via authorship verification; and naturalness, whether conversations exhibit human-like turn-to-turn flow via dialogue NLI. Unlike optimization-oriented metrics, each dimension takes a human corpus as a reference baseline and penalizes deviations in both directions, distinguishing insufficient persona encoding from over-optimized, unnatural behaviour. The framework is corpus-agnostic: any persona-annotated conversational dataset can serve as the reference. Evaluated over ten simulation corpora, Eval4Sim surfaces systematic trade-offs invisible to single-score methods.","authors":["Eliseo Bao","Anxo Perez","Javier Parapar","Xi Wang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-19","first_seen":"2026-03-03","revised_at":"2026-08-19","abs_url":"https://arxiv.org/abs/2603.02876","pdf_url":"https://arxiv.org/pdf/2603.02876","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM人格仿真","评估框架","人类行为对照"],"reason":"评估LLM人格仿真与人类对话的一致性，含人类语料对照，可迁移至仿真可靠性研究。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":6,"question":"如何评估基于LLM的人格仿真对话与真实人类对话行为的一致性？","design":"本文提出Eval4Sim评估框架，不进行新的仿真实验，而是对已有的十个仿真语料库进行评估。框架从三个维度测量仿真对话与人类参考语料的对齐程度：adherence（通过密集检索判断人格特质是否可从对话中恢复）、consistency（通过作者验证判断说话者是否保持可区分的风格身份）、naturalness（通过对话NLI判断对话是否具有人类般的轮次流畅性）。每个维度以人类语料为基准，惩罚双向偏差。","baseline":"使用带有人格标注的人类对话语料库作为参考基准，具体数据集未在节选中列出，但框架是语料库无关的，任何带说话者人格标注的对话数据集均可作为参考。","findings":"Eval4Sim在十个仿真语料库上揭示了单一评分方法无法发现的系统性权衡，例如过度优化人格特质恢复可能导致不自然的自我披露，而优化流畅性可能削弱风格身份。框架能够区分人格编码不足与过度优化导致的不自然行为。","reliability":"论文指出，现有LLM-as-a-judge方法缺乏可观察行为基础且产生不透明分数，而Eval4Sim通过人类语料基准和双向惩罚解决了这一问题。但节选未明确讨论Eval4Sim自身的失效条件或局限。","relevance":"该研究直接针对LLM人格仿真与人类行为一致性的评估问题，提供了基于人类语料对照的多维度评估方法，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文了解具体实现和发现。","inspiration":"Eval4Sim的双向惩罚设计值得借鉴，即不单纯追求指标最大化，而是以人类行为分布为基准，惩罚偏离基准的仿真行为，这可以用于校准经济实验中的LLM被试行为。｜该方法可迁移到消费者决策仿真、投资者情绪模拟或政策沟通实验中，用于评估LLM生成的决策行为是否与真实人类行为分布一致。｜例如，在消费者跨期选择实验中，用LLM扮演不同人格特质的消费者，施加不同的时间折扣处理，测量其选择行为，并与真实消费者面板数据（如CFPS或Understanding America Study）对比，采用类似Eval4Sim的多维度对齐评估，检验LLM仿真是否在均值、异质性和分布形状上偏离人类基准。"}},{"id":"2608.17516","version":1,"title":"Effects of Answer Format Variation on Gender Bias in Large Language Models","zh_title":"回答格式变化对大语言模型中性别偏差的影响","abstract":"Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.","authors":["Ksenia Merzlyakova","Sebastian Pad\\'o","Franziska Weeber"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17516","pdf_url":"https://arxiv.org/pdf/2608.17516","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM评估","性别偏差","调查方法"],"reason":"评估LLM回答格式对性别偏差测量的影响，并与人类调查数据对照，揭示仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":9,"question":"回答格式变化如何影响大语言模型中性别偏差的测量及其与人类回答分布的一致性？","design":"使用三个指令微调模型（Mistral-7B-Instruct-v0.3、Llama-3.1-8B-Instruct、Gemma-3-12B-IT）在BBQ基准和OpinionQA调查数据上，对相同问题施加封闭式、李克特量表、开放式三种回答格式，测量性别偏差和回答分布。","baseline":"OpinionQA中来自美国公众意见调查的人类回答分布。","findings":"回答格式显著改变测量结果，包括模型间偏差排序的逆转；不同格式引发不同的回答行为，如强制选择、量表分布和自由文本中的拒绝回答。","reliability":"论文未讨论","relevance":"该研究直接评估LLM仿真人类回答时对测量格式的敏感性，并对照真实调查数据，揭示了仿真在格式变化下可能失效，值得精读以理解偏差测量的稳健性。","inspiration":"借鉴其系统操纵回答格式并对照人类基准的方法，可迁移到经济金融领域的调查仿真或行为实验，如消费者信心调查、通胀预期或风险偏好测量。｜例如，用LLM模拟消费者在封闭式与开放式问题下的通胀预期，处理为回答格式，结果变量为预期值分布，对照密歇根大学消费者调查的真实数据。"}},{"id":"2608.17150","version":1,"title":"KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn","zh_title":"KnowSim：用可学习的用户模拟器评估LLM助手的信息校准","abstract":"To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.","authors":["Yoonjoo Lee","Hyoungwook Jin","Tae Soo Kim","Shaoyang Zhang","Philippe Laban","Q. Vera Liao"],"categories":["cs.AI","cs.CL","cs.HC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17150","pdf_url":"https://arxiv.org/pdf/2608.17150","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["用户模拟","信息校准","人机交互评估"],"reason":"用用户模拟器评估LLM信息校准，含人类数据对照，可迁移至人类仿真研究","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":8,"question":"如何构建并验证一个基于知识状态建模的用户模拟器，用于评估LLM助手在知识密集型任务中的信息校准能力？","design":"提出KnowSim框架，其用户模拟器维护显式知识状态（由信息单元及先决关系构成的图），并根据学习理论更新规则演化；模拟不同知识水平（新手/中级/高级）的用户与LLM助手进行多轮对话，从状态轨迹计算知识增益、传递校准和认知过载三个指标。","baseline":"705段人类-AI对话，涵盖数学问题求解和专家级问答两个领域，参与者按初始知识水平分层，并收集主观评分及数学领域的前后测知识分数。","findings":"KnowSim的排名与人类判断显著一致（73-74%符号一致率），优于三个基线模拟器，且在新手水平上对齐最强。应用于9个LLM时，最佳模型随用户知识水平变化，揭示了标准评估无法发现的资质-处理交互效应。","reliability":"论文未讨论","relevance":"该研究直接针对LLM作为人类被试替代品的仿真可靠性问题，提供了带人类基准的验证框架，并揭示了仿真在不同知识水平下的异质性表现，对关注仿真效度与偏差的研究者具有重要参考价值。","inspiration":"借鉴其显式建模个体状态并动态更新的方法，可提升经济仿真中异质性主体的行为真实性。｜可迁移到政策沟通或金融教育场景，如央行公告对公众通胀预期的影响、或理财建议对不同金融素养人群的效果。｜设计一个实验：用LLM模拟不同金融素养水平的投资者，处理为不同信息呈现方式的投资建议（如简化版vs专业版），结果变量为投资决策质量和知识增益，并与真实投资者调查数据对照。"}},{"id":"2608.17099","version":1,"title":"Appearing Legitimate is Not Enough: Interrogating Synthetic Agents in Representational Processes through a Participatory Design Lens","zh_title":"表面合法还不够：通过参与式设计视角审视代表性过程中的合成代理","abstract":"Synthetic agents built atop LLM-based foundation models are gaining popularity as substitutes for human participants across research contexts, including user-testing, market-research, computational social science, surveys, and qualitative research. We are also witnessing an extension of synthetic agents into experimental implementations of policy consultation, jury deliberation, humanitarian diplomacy, and similar contexts where human participation and representation are central to the perceived legitimacy of the institutional processes. The value of participation extends beyond informational contributions and consensus generation; participation is a necessary, legitimizing condition for democratic political institutions and processes. Treating synthetic agents as human substitutes raises serious political, representational, and ethical concerns. Participatory Design's modes of engagement --- probing, priming, understanding, and generating --- offer helpful tools for engaging with representational questions of personhood. We apply the lens to three case studies of synthetic agents substituting for personhood at varying representational scales: local policy, enterprise jury deliberation, and global diplomacy. We argue that legitimacy and personhood are integral and mutually constitutive while identifying the ethical, representational, and methodological risks of using synthetic agents in representational processes. We conclude by proposing soft and hard boundaries for designing oversight on LLMs and synthetic agents in representational processes.","authors":["Aditya Nayak","Aditi Vashistha","Alissa Centivany","Aakash Gautam"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17099","pdf_url":"https://arxiv.org/pdf/2608.17099","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A4","B4"],"tags":["合成代理","参与式设计","代表性伦理"],"reason":"批判性审视合成代理替代人类参与的代表性问题，提出监督边界，方法论可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":7,"question":"合成代理在代表性过程中如何制造出合法参与的假象，以及应如何设定设计与部署的边界？","design":"本文不是仿真实验研究，而是对三个合成代理案例（地方政策咨询聊天机器人Ana、企业陪审团审议工具Synthetic Juror、全球外交AI化身Ask Amina和Ask Abdalla）进行比较案例分析，运用参与式设计的四种模式（探查、启动、理解、生成）剖析其如何制造人格假象。","baseline":"无对照","findings":"合成代理通过问题框架、数据策展、用户交互设计和有效性评估四个步骤制造出“人造人格”，绕过代表性过程从而获得表面合法性。现有评估框架（如可信度、保真度、算法保真度）只衡量输出相似性，无法检测这种对过程完整性的绕过。","reliability":"论文未讨论","relevance":"该文批判性审视合成代理在代表性过程中的合法性与人格问题，提出软硬边界，对关注LLM仿真可靠性及伦理边界的研究者具有重要参考价值。","inspiration":"本文的参与式设计视角和过程导向批判方法值得借鉴，可迁移到经济金融领域中涉及代表性决策的场景（如政策咨询、消费者意见征询、董事会决策模拟等）。｜可设计一项研究，用LLM合成代理模拟消费者或投资者参与政策咨询或产品设计讨论，处理为不同的人格制造步骤（如改变数据策展或交互设计），结果变量为参与者对过程合法性的感知或决策质量，并与真实人类参与者的数据对照。"}},{"id":"2608.14606","version":1,"title":"Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents","zh_title":"看似合理但无效：对LLM作为合成调查受访者的心理测量审计","abstract":"Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM \"crowd\" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.","authors":["Mantas Lukauskas","Viktorija \\v{S}arkauskait\\.e"],"categories":["cs.CY","cs.AI","cs.CL","stat.AP"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14606","pdf_url":"https://arxiv.org/pdf/2608.14606","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","心理测量效度","调查数据"],"reason":"直接评估LLM作为调查受访者的心理测量效度，并与真实人类数据对照，批判性指出失…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":2,"question":"LLM 作为合成调查受访者时，是否能在心理测量学层面（联合分布、潜结构、信度、中介路径、人口学效应）复现真实人类调查数据的特征？","design":"使用 37 个 LLM（涵盖 OpenAI、Anthropic、Google 及 12 个开源家族）基于真实受访者档案生成对 68 个题项（3 个量表）的回答，通过五级人格披露阶梯、呈现方式与推理努力消融、反事实人口学变换（性别、角色、教育）、跨语言检查和逐字回忆探针等处理，测量心理测量相似性得分（PSS）及其各维度。","baseline":"立陶宛组织心理学数据集（n=263 名员工；Dunham 变革态度量表、UWES-17、Koopmans IWPQ；68 题，12 个分量表），以及五个非 LLM 统计基线和留出的人类对比上限。","findings":"LLM 能复现人类心理测量关系的定性方向，但高斯 copula 基线在样本驱动的 PSS 分量上击败所有 LLM；LLM 群体内部相似度（平均 PSS 0.73）高于与人类的相似度，且逐字回忆探针表明记忆并非排行榜驱动因素（秩相关 0.00）。","reliability":"论文指出 LLM 样本不能直接替代人类调查数据：合成样本存在默认偏差（+0.84 SD），在留出人类数据上预测效度丧失（平均 R² -0.18 vs 0.28），并在 10 条安慰剂中介路径中捏造了 3 条显著间接效应；此外，教育驱动的反事实效应（平均 |d|=0.56）远大于性别（0.12）和角色（0.18），且 UWES 的 Tucker's phi 在 37 个模型中有 8 个落入置换零分布内。","relevance":"该研究直接评估 LLM 作为人类被试替代品的心理测量效度，并与真实人类数据严格对照，批判性地揭示了仿真在联合分布、预测效度和中介推断上的失效条件，对关注 LLM 仿真可靠性与偏差的研究者极具参考价值。","inspiration":"借鉴其多维度心理测量审计框架（联合分布、潜结构、信度、中介路径、人口学效应）和反事实人口学变换设计，系统评估合成样本的效度｜可迁移到经济金融中的调查实验，如消费者信心、通胀预期、风险偏好或政策支持度等场景，检验 LLM 能否复现真实人群的分布与结构｜以真实家庭金融调查（如美国 SCF 或中国 CHFS）为基准，用 LLM 基于受访者人口学特征生成对风险态度、时间偏好等量表的回答，施加收入或教育水平的反事实变换，比较 LLM 样本与人类样本在联合分布、因子结构和中介效应上的差异。"}},{"id":"2608.15871","version":1,"title":"Large Language Models as Implicit Sociological Models: Reconstructing Voting Behaviour from Sociodemographic Profiles","zh_title":"大语言模型作为隐式社会学模型：从社会人口特征重建投票行为","abstract":"Large language models (LLMs) trained on large-scale internet corpora encode extensive statistical regularities about social identities, attitudes, and political behaviour. This paper introduces and evaluates a methodological framework that leverages these latent representations to reconstruct aggregate voting behaviour from individual-level sociodemographic profiles. We operationalize LLMs as implicit sociological models by conditioning them on demographic descriptions, eliciting probabilistic turnout and party preferences, and aggregating individual outputs via a soft voting procedure. Using the 2021 Czech parliamentary election as a validation case, we demonstrate that contemporary LLMs reproduce official election outcomes with low mean absolute error, recover known political bloc structures, and align with independently established sociodemographic gradients. The contribution of this work is methodological rather than predictive: we show how LLMs can be systematically interrogated as compressed representations of social reality, offering a novel exploratory instrument for computational social science while clearly delineating its epistemic and ethical limits.","authors":["Roman Neruda","Martin Bako\\v{s}","Josef \\v{S}lerka","V\\'it Tu\\v{c}ek","Petra Vidnerov\\'a","Gabriela Kadlecov\\'a"],"categories":["cs.CY","cs.CL","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15871","pdf_url":"https://arxiv.org/pdf/2608.15871","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","投票行为","计算社会科学"],"reason":"用LLM从人口特征重建投票行为，并与真实选举结果对照，直接仿真人类决策。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":1,"question":"大语言模型能否仅从个体社会人口学特征重建总体投票行为，并在捷克2021年议会选举这一非英语多党制情境下得到验证？","design":"将LLM作为隐式社会学模型，以10个社会人口学变量（含主观生活水平和政治兴趣）为条件，对代表性调查中的个体生成概率性投票选择（投票率和政党偏好），通过软投票聚合得到模拟选举结果。","baseline":"官方选举结果（捷克统计局）和调查中的自报投票选择（声称投票）。","findings":"当代LLM能以较低平均绝对误差重现官方选举结果，恢复已知政治阵营结构，并与独立确立的社会人口学梯度一致。该贡献是方法论的，而非预测性的。","reliability":"论文承认个体层面预测噪声大且存在系统性偏差，但通过软投票聚合可部分抵消；捷克语为中等资源语言，模型训练数据以英语为主，可能影响表现；未明确讨论其他失效条件。","relevance":"该研究直接命中你的核心关注：用LLM仿真人类决策并与真实选举数据对照，且提供了多层级验证（总体份额、协方差结构、与民调机构及CHES专家编码比较），值得精读原文以了解其方法细节和局限。","inspiration":"借鉴其软投票聚合和三角验证设计，将个体噪声转化为总体稳健估计，并同时对照官方数据和自报数据以识别偏差。｜可迁移到政策公告的预期形成研究，例如模拟不同人口群体对财政或货币政策变化的反应。｜以代表性家庭调查中的个体为被试，用LLM基于人口特征生成对政策变化的预期（如通胀预期），处理为不同政策情景描述，结果变量为预期值或不确定性，对照真实调查中的预期数据和后续实际经济行为。"}},{"id":"2608.14630","version":1,"title":"Characterizing Rhetorical Misalignment in Decision-Making with Language Models","zh_title":"表征语言模型决策中的修辞错位","abstract":"Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly integrated into high-stakes human-AI decision-making, it is important to understand whether their outputs can amplify potential biases, how this influences human decisions, and crucially, whether it can lead to harmful consequences. In this work, we develop a decision-theoretic framework to study rhetorical misalignment, a failure mode where an LLM uses rhetorically inappropriate forms of presentation for a given decision context, thereby inducing suboptimal human decisions. We empirically investigate this phenomenon through a human-subject experiment in realistic clinical decision-making using a dataset curated from the United States Medical Licensing Examination. By measuring how LLM-generated information affects decisions, we observe that LLMs induce an average 2.81% rate of harmful decision flips across different models, where clinician participants change from a correct to an incorrect answer. Rationales reported by participants provide evidence that these revisions are closely related to the language used by LLMs that may induce different types of cognitive biases, including anchoring, authority bias, and loss aversion. To enable scalable evaluation, we instantiate our theoretical framework using decision-makers simulated by LLMs to computationally measure rhetorical misalignment. Our findings reveal a safety concern previously unrecognized in high-stakes domains: a model can be factually aligned yet still induce harm through its rhetorical presentation.","authors":["Zirui Cheng","Joey Chan","Simo Du","Chenhao Tan","Yue Guo","Hao Peng"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14630","pdf_url":"https://arxiv.org/pdf/2608.14630","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类决策","认知偏差"],"reason":"用LLM模拟人类决策者评估修辞偏差，并与真实人类实验对照，涉及临床决策场景，批…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":7,"question":"LLM 的修辞呈现方式是否会在事实信息一致的情况下诱导人类做出次优决策，即产生“修辞错位”现象？","design":"论文采用人类受试者实验与LLM模拟决策者相结合的方法。人类实验中，临床医生参与者基于USMLE题目作答，部分题目提供LLM生成的分析（固定信息或自然生成），测量决策变化率及有害翻转率；模拟实验中，用LLM分别扮演理性决策者和行为决策者，比较两者在相同信息不同措辞下的决策差异。","baseline":"人类受试者实验数据：临床医生在没有LLM辅助时的原始答案作为对照，测量LLM辅助后的决策变化。","findings":"人类实验中，LLM辅助导致平均27.58%的决策变化，其中2.81%为有害翻转（从正确变为错误）；参与者报告显示这些变化与锚定、权威偏误、损失厌恶等认知偏误相关。模拟实验表明，即使信息相同，仅语言措辞差异也能导致理性与行为决策者之间的分歧，且自然生成设置下分歧更大。","reliability":"论文承认其人类实验仅限于USMLE临床决策场景，未测量下游后果（如患者结局、经济成本），且有害翻转率绝对值较小；模拟实验的LLM决策者可能无法完全代表人类认知偏误。","relevance":"该研究直接使用LLM模拟人类决策者来测量修辞错位，并与真实人类实验对照，属于用LLM进行人类仿真实验的典型工作，且涉及高 stakes 临床决策场景，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其“固定信息、变化措辞”的处理设计，可分离信息内容与修辞框架的效应，并利用LLM模拟理性与行为决策者进行大规模测量｜可迁移到经济金融中的政策公告解读、投资建议呈现、信贷合同条款表述等场景，研究措辞对个体决策的影响｜设计实验：以真实投资者或消费者为被试，呈现同一金融产品的两种措辞（如“年化收益5%” vs “亏损概率2%”），测量选择差异，并用历史交易数据或调查数据作为真实行为基准，同时用LLM模拟投资者进行平行实验以验证仿真效度。"}},{"id":"2608.14079","version":1,"title":"The conditional superiority of fast silicon sampling","zh_title":"快速硅采样的条件优越性","abstract":"Silicon sampling can produce surprisingly good population estimates at times. Does doing it fast attenuate such fidelity? In this study, we extend and assess ongoing work in silicon sampling by comparing the algorithmic fidelity of \"fast\" and \"slow\" modes of silicon sampling among a nationally representative sample of Singaporean survey respondents. We find that silicon sampling with contemporary frontier models remains a method in early development to be used only with great caution. While silicon samples are able to produce moderately faithful estimates of population means, they continue to understate opinion variance and distort the latent contextual space behind human opinions. Conditional on such limitations, we find \"fast\" modes of silicon sampling to be relatively superior to traditional \"slow\" modes of silicon sampling. Fast silicon sampling is significantly more efficient in compute resources and run-time while being monotonically superior to slower modes of sampling in algorithmic fidelity.","authors":["Nickolas Hock Yuen Lam","Ji Xuan Voo","Xiangyu Ma"],"categories":["cs.CL","cond-mat.mtrl-sci"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14079","pdf_url":"https://arxiv.org/pdf/2608.14079","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["硅采样","算法保真度","人类仿真"],"reason":"直接比较快速与慢速硅采样在代表性样本上的算法保真度，含真实人类数据对照，并指出…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":1,"question":"快速硅采样是否在算法保真度上不劣于甚至优于传统慢速硅采样？","design":"使用当代前沿模型（OpenAI GPT 5.4）对新加坡全国代表性调查受访者进行硅采样，比较“快速”模式（一次提示生成大量响应）与“慢速”模式（逐个API调用生成响应）的算法保真度，测量结果包括总体均值估计、意见方差和潜在情境空间结构。","baseline":"新加坡全国代表性调查受访者的真实人类数据。","findings":"硅采样能中等程度地忠实估计总体均值，但低估意见方差并扭曲潜在情境空间。在承认这些局限的前提下，快速硅采样在算法保真度上单调优于慢速硅采样，且计算资源和运行时间显著更高效。","reliability":"论文承认硅采样仍处于早期发展阶段，需谨慎使用；硅样本低估意见方差、扭曲潜在情境空间，且快速与慢速模式均存在这些局限。","relevance":"该研究直接比较快速与慢速硅采样在代表性样本上的算法保真度，含真实人类数据对照，并指出硅采样在方差和潜在空间上的失效条件，对关注LLM仿真可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其通过几何数据分析（多重对应分析）评估关系保真度的方法，以及比较不同采样模式效率与保真度的设计。｜可迁移到政策评估中的公众意见模拟，如经济政策公告的预期形成或消费者信心调查。｜以LLM生成不同处理模式下的合成受访者，处理为快速与慢速采样，结果变量为对经济政策的态度分布，对照真实调查数据（如新加坡消费者信心指数），评估均值、方差和潜在空间结构的一致性。"}},{"id":"2608.10492","version":2,"title":"INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators","zh_title":"洞察学生思维：联合建模LLM学生模拟器中的潜在推理与行为","abstract":"Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework that fine-tunes LLMs not only to act like students but also to think like them. INSIDE generates internal dialogue grounded in Bloom's Taxonomy across cognitive, affective, and action dimensions, and fine-tunes models on paired think traces and actions. We baseline against different prompting frameworks and evaluate on two axes: fidelity of simulated actions and quality of generated internal dialogue. Our evaluations show that INSIDE improves simulation fidelity in both action fidelity, matching code generation of real students, and reasoning alignment, achieving the highest alignment across models up to 57.9%.","authors":["Rose Niousha","Minwoo Kang","Narges Norouzi"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-17","first_seen":"2026-08-12","revised_at":"2026-08-17","abs_url":"https://arxiv.org/abs/2608.10492","pdf_url":"https://arxiv.org/pdf/2608.10492","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","教育模拟","算法保真度"],"reason":"用LLM仿真学生行为与推理，并与真实学生数据对照，评估仿真保真度。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-17","rank":3,"question":"如何让LLM学生模拟器不仅复现学生的可观察行为（代码提交），还能捕捉其背后的潜在推理过程，从而提高仿真保真度？","design":"提出INSIDE框架，基于Bloom分类法生成内部对话（认知、情感、行动维度），并用配对的思想痕迹和行动对LLM进行微调。模型扮演编程课程学生，输入学生历史提交和AI导师反馈，输出下一步代码提交。评估两个维度：行动保真度（生成代码与真实学生代码的相似度）和推理质量（生成推理与真实代码编辑的对齐度）。","baseline":"使用加州大学伯克利分校入门编程课程两个学期的真实学生数据：Spring 2025用于训练（445名学生，2022条提交流，6911次提交），Spring 2024用于测试（479名学生，1546条提交流，6316次提交）。测试集分为旧问题（test_OP）和新问题（test_NP），分别评估对未见学生和未见问题的泛化。","findings":"INSIDE提高了行动保真度，生成代码与真实学生代码的Wasserstein距离更低；同时实现了最高的推理对齐度，在不同模型上最高达到57.9%。","reliability":"论文未讨论","relevance":"该研究直接命中你的核心关注点：用LLM仿真学生行为与推理，并与真实学生数据对照，评估仿真保真度。它提供了在缺少真实推理标注的情况下重建潜在推理的方法，并展示了联合建模推理与行动能提升仿真质量，值得精读原文。","inspiration":"借鉴其用理论框架（Bloom分类法）引导内部对话生成，并将推理与行动联合微调以提升仿真保真度的做法。｜可迁移到经济金融中的政策预期形成研究，例如模拟投资者在信息发布后的决策过程，捕捉其推理路径。｜以LLM扮演投资者，输入历史交易和新闻信息，生成内部推理（如风险评估、情绪反应）和交易决策，用真实市场交易数据和调查数据（如投资者信心指数）作为对照，评估仿真保真度。"}},{"id":"2601.20238","version":2,"title":"Large Language Models Polarize Ideologically but Moderate Affectively in Online Political Discourse","zh_title":"大语言模型在网络政治话语中加剧意识形态极化但缓和情感极化","abstract":"The emergence of large language models (LLMs) is reshaping how people engage in political discourse online. We examine how the release of ChatGPT altered ideological and emotional patterns in Reddit's largest political forum. Analysis of millions of comments shows that ChatGPT intensified ideological polarization: liberal-leaning authors posted increasingly liberal comments, while conservative-leaning authors posted increasingly conservative comments. Multiple falsification tests suggest that these findings are unlikely to be driven by contemporaneous events, such as the 2022 U.S. midterm elections, or by broader platform-wide trends in political polarization. Mechanism tests show that this shift does not stem from the creation of more persuasive or ideologically extreme original content using LLM. Instead, it originates from the tendency of LLM-assisted comments to echo and reinforce the original post's viewpoint, a pattern consistent with algorithmic sycophancy. Yet, despite growing ideological divides, affective polarization, measured by hostility and toxicity, declined. These findings reveal that LLMs can simultaneously deepen ideological separation and foster more civil exchanges, challenging the long-standing assumption in literature that extremity and incivility necessarily move together.","authors":["Gavin Wang","Srinaath Anbudurai","Oliver Sun","Xitong Li","Lynn Wu"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-08-17","first_seen":"2026-01-28","revised_at":"2026-08-17","abs_url":"https://arxiv.org/abs/2601.20238","pdf_url":"https://arxiv.org/pdf/2601.20238","source_feed":"econ.GN","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","政治极化","人类数据对照"],"reason":"用LLM辅助评论与真实Reddit数据对照，分析政治话语极化，涉及社会过程仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-17","rank":4,"question":"ChatGPT的发布如何影响Reddit政治论坛中用户的意识形态极化和情感极化？","design":"本研究并非将LLM作为人类被试的仿真实验，而是利用Reddit上最大的政治论坛在ChatGPT发布前后的数百万条评论数据，通过语言特征指标（如评论长度、被动动词、困惑度、人类评分）估计每条评论由LLM辅助生成的概率，进而分析LLM使用对用户意识形态立场和情感表达的影响。","baseline":"以ChatGPT发布前的Reddit评论作为对照，比较同一作者在发布前后的意识形态立场变化，并利用2022年美国中期选举等事件进行证伪检验。","findings":"ChatGPT的发布加剧了意识形态极化：自由派作者发表更自由的评论，保守派作者发表更保守的评论。然而，情感极化（敌意和毒性）却有所下降，表明LLM在加深意识形态分歧的同时促进了更文明的交流。","reliability":"论文承认无法直接确定每条评论是否由LLM生成，而是依赖语言指标估计概率，可能存在分类误差；但认为在大规模语料中，误差不会系统性地产生所观察到的聚合模式。","relevance":"该研究利用真实Reddit数据评估LLM对政治话语的影响，涉及LLM辅助内容生成与真实人类行为的对照，对关注LLM在社会科学中仿真可靠性的研究者有参考价值，但并非直接以LLM作为被试的仿真实验。","inspiration":"值得借鉴的是利用自然实验（ChatGPT发布）和语言特征指标来估计LLM使用概率，并设置证伪检验排除混淆因素。｜可迁移到经济金融领域如政策公告后的市场情绪分析、消费者评论中的LLM影响等场景。｜一个可行的设计是：以某经济政策发布为时间节点，收集社交媒体上相关讨论，用语言指标估计LLM辅助评论比例，分析其对情绪极化和观点极化的影响，并与历史人类评论基线对照。"}},{"id":"2608.13712","version":1,"title":"Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation","zh_title":"字里行间：法律模拟中行为真实性的建模与评估","abstract":"Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator driven by controllable legal personas. We use an evaluation framework separating behavioral realism from pedagogical usefulness. We assess realism through adversarial testing, blinded attorney comparison, and analysis of longitudinal behavioral trajectories. WitnessSim generally maintained plausible behavioral boundaries, and attorneys did not systematically prefer either original testimony or WitnessSim generated testimony. Pedagogical tests showed that witness behavior changed meaningfully in response to question form and attorney intervention without uniformly collapsing the assigned persona. Together, these results showcase a model of behavioral fidelity in legal simulations, and provide a framework for evaluating its performance.","authors":["Divya Vetticaden","Arya Gupta","Julian Nyarko","Megan Ma"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13712","pdf_url":"https://arxiv.org/pdf/2608.13712","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","法律模拟","行为真实性"],"reason":"用LLM模拟法律证人行为，并与真实证词对照，评估行为真实性和教学效果，属于人类…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":6,"question":"如何构建并评估一个具有行为真实性和教学实用性的法律证人模拟系统？","design":"WitnessSim 是一个基于六维状态向量（镇定、知识、宜人性、冗长度、僵化度、表现）的沉积证词模拟器，通过问题特征（压力、话题敏感性、问题形式）更新状态，并条件化生成证词；研究使用对抗测试、盲法律师比较和纵向行为轨迹分析评估行为真实性，并通过法律培训材料衍生的测试评估教学实用性。","baseline":"来自国家处方阿片类药物诉讼（Case No. 1:17-MD-2804）的 300 份沉积和庭审笔录，作为真实人类证词对照。","findings":"WitnessSim 在对抗测试中维持了合理的行为边界，律师在盲法比较中未系统偏好原始证词或生成证词。教学测试显示证人行为对问题形式和律师干预有有意义的变化，且未完全丧失指定人格。","reliability":"论文未明确讨论失效条件与局限，但指出评估框架区分行为真实性与教学实用性，并承认行为真实性通过对抗测试、专家评估和情感轨迹分析来操作化，可能仍存在未覆盖的维度。","relevance":"该研究直接命中研究者关注的 LLM 人类仿真实验，提供了法律场景下与真实人类数据对照的行为真实性评估框架，值得阅读原文以借鉴其多维评估方法和动态行为建模。","inspiration":"借鉴其将行为状态建模为可更新的多维向量，并通过问题特征施加处理、以真实行为轨迹为基准的评估方法。｜可迁移到经济金融中的谈判、审计或客户服务交互模拟，如信贷审批中的申请人行为或政策沟通中的公众反应。｜以 LLM 模拟信贷申请人，处理变量为审批官提问的侵略性或信息敏感度，结果变量为申请人的情绪状态和回答一致性，对照真实信贷申请面谈记录。"}},{"id":"2608.12368","version":1,"title":"Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments","zh_title":"一致不等于对齐：人类与LLM道德判断中分歧的道德依据","abstract":"Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation. We test this distinction using a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. Across frontier and open model families, agreement with human annotator majority labels is often high. However, rationale-level analysis reveals systematic divergence in the moral grounds expressed by human annotators and models. In particular, models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even when their final labels match the human annotator majority. Our results show that agreement should not be treated as equivalent to alignment. Label-based evaluation can therefore be misleadingly reassuring unless complemented by analysis of the reasons, principles, and moral priorities expressed in model judgments.","authors":["Octavian M. Machidon","Alina L. Machidon","Vojko Strahovnik","Mateja Centa Strahovnik","Jonas Miklav\\v{c}i\\v{c}","Marko Robnik \\v{S}ikonja"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12368","pdf_url":"https://arxiv.org/pdf/2608.12368","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM对齐评估","道德判断","人类对照"],"reason":"评估LLM道德判断与人类的一致性，揭示标签一致但理由分歧，有真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:00:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":2,"question":"在道德判断任务中，人类标注者与LLM的最终标签一致是否意味着它们依赖相同的道德理由？","design":"本研究并非将LLM作为人类被试的替代品进行仿真实验，而是直接比较人类标注者与多个LLM在500项ETHICS衍生道德判断任务上的表现。人类标注者和LLM对相同项目进行标注，提供最终标签和支持理由；研究者分析标签一致性与理由层面的分歧。","baseline":"新收集的人类标注者数据，包括最终标签和理由标注，作为与LLM输出比较的基准。","findings":"LLM与人类多数标签的一致性通常较高，但理由层面存在系统性分歧。即使最终标签一致，模型在伤害、尊重、守诺、正义、应得、借口相关性等道德理由类别上的注意力分布与人类不同。","reliability":"论文指出，标签一致性可能误导性地令人放心，因为模型可能依赖不同的道德理由，导致泛化差异、解释不匹配、过度道德化或忽视隐含社会义务。此外，理由对齐是描述性的，不能证明所述理由是模型输出的因果原因。","relevance":"该研究直接回应了研究者对LLM仿真可靠性的关注，揭示了仅用标签一致性评估对齐的不足，并提供了真实人类数据对照，对批判性评估LLM在道德判断中的仿真效度具有重要价值。","inspiration":"借鉴其双层评估设计（标签一致性与理由对齐）和理由编码框架，可迁移到经济金融中的伦理决策场景（如信贷审批中的公平性判断、消费者对金融产品道德性的评价）。设计雏形：以LLM作为被试，呈现金融道德困境（如掠夺性贷款案例），要求给出判断和理由，与人类专家标注的理由类别进行对比，以真实人类标注数据为基准，检验LLM在金融伦理判断中的理由一致性。"}},{"id":"2608.12344","version":1,"title":"Predicting consumer-technology ownership without a diffusion history","zh_title":"无扩散历史下预测消费者技术拥有率","abstract":"We test whether the perceived attributes of a consumer technology predict how widely it is owned. In a 2022 Prolific survey of US adults (n = 678), respondents rated 65 consumer technologies on six attributes. We then elicited the same ratings from two frontier language models, Anthropic Claude Opus 4.7 and OpenAI GPT-5.5. We regress ownership prevalence on four UTAUT2 acceptance attributes plus a log-age covariate with a sign-constrained penalized regression and evaluate it by holding out one technology at a time. The attribute model improves on a baseline of years-since-launch: mean absolute error falls by 17% with the human ratings, and by more with either model, most with Opus 4.7. Over the short 2022-to-2025 window, where ownership moved little, the same attributes do not improve on a no-change baseline. We set out the limitations of the approach, including the possibility that language-model ratings reflect prior knowledge of these technologies rather than independent attribute reasoning. We include a deployment illustration: 2027 ownership predictions for eleven products launched in 2025 and 2026.","authors":["Irina Vartanova","Niels Selling","Jennifer Viberg Johansson","Pontus Strimling"],"categories":["cs.CL","cs.CY","stat.AP"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12344","pdf_url":"https://arxiv.org/pdf/2608.12344","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","消费者行为","算法保真度"],"reason":"用LLM替代人类被试预测技术拥有率，并与真实调查数据对照，评估模型可靠性，属核…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":1,"question":"消费者技术的感知属性（UTAUT2 属性）能否在无扩散历史的情况下预测其拥有率，以及预测效果是否因评分者是人类还是大语言模型而异？","design":"用 2022 年 Prolific 调查中 678 名美国成年人对 65 项消费者技术的六项属性评分作为人类评分，再让两个前沿大语言模型（Claude Opus 4.7 和 GPT-5.5）对同样技术给出相同属性评分；以拥有率为结果变量，用符号约束惩罚回归拟合四个 UTAUT2 属性加对数年龄协变量，通过留一技术交叉验证评估预测误差。","baseline":"2022 年 Prolific 调查中 678 名美国成年人的真实拥有率和属性评分，以及 2025 年随访调查的拥有率变化。","findings":"属性模型相比仅用技术年龄的基线降低了平均绝对误差：人类评分降低 17%，大语言模型评分降低更多，其中 Opus 4.7 表现最好（MAE 7.1 个百分点）。但在 2022 至 2025 年拥有率变化很小的短窗口内，属性模型未能优于无变化基线。","reliability":"论文承认大语言模型评分可能反映了模型对技术的先验知识而非独立的属性推理，且模型无法预测拥有率随时间的变化；此外，属性模型仅适用于横截面水平预测，不涉及扩散动态。","relevance":"该研究直接使用大语言模型替代人类被试进行属性评分，并与真实调查数据对照，评估了仿真在预测技术拥有率上的可靠性，属于典型的 LLM 仿真实验，且包含批判性讨论，值得精读。","inspiration":"借鉴其用 LLM 生成属性评分并与人类评分对比、以真实拥有率作为结果变量的设计，可迁移到消费者金融产品采纳预测（如数字支付、理财产品）或政策接受度评估；具体可设计让 LLM 扮演不同人口群体对新型金融产品进行 UTAUT2 属性评分，以实际调查的采纳率作为基准，检验 LLM 评分能否预测真实采纳率并识别偏差。"}},{"id":"2608.12339","version":1,"title":"Mimicry without understanding: the origins of decision bias in large language models","zh_title":"无理解的模仿：大语言模型中决策偏差的起源","abstract":"Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through which such biases can be generated even when human preferences (in the training data) are not biased or when they are correctly categorized as being biased. The first is faulty mimicry of preferences based on human behavior: this involves LLMs inferring human preferences even when behaviors are logically unrelated to preferences. The second is mimicry of explicitly biased human behaviors. In four studies focusing on economic biases, we find that ChatGPT-4o and Qwen exhibited social proof biases even when prompted with reports of human behaviors that were clearly non-indicative of individuals' actual preferences. LLMs also displayed loss aversion when it was explicitly described as a bias. Indeed, when prompted with detailed scientific reports, the extent of the bias (i.e., loss aversion) in the scientific report predicted LLMs' own subsequent bias. Scientific papers of biases can thus become self-fulfilling prophecies, at least when it comes to LLMs' responses. The current study goes beyond fleshing out LLM biases and sheds light on the underlying component processes.","authors":["Eldad Yechiam","Adi Tarabeih"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12339","pdf_url":"https://arxiv.org/pdf/2608.12339","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A2","B2","B4"],"tags":["LLM偏差","经济决策","仿真可靠性"],"reason":"研究LLM决策偏差的生成机制，涉及经济偏差，有批判性，可迁移到仿真可靠性评估。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":1,"question":"LLM 决策偏差的生成机制是什么？具体考察两种过程：基于人类行为的错误偏好推断，以及对被明确标注为偏差的人类行为的模仿。","design":"使用 ChatGPT-4o 和 Qwen 作为被试，通过提示词向模型呈现人类行为报告或科学文献描述，然后测量模型在货币选择、社会证明、损失厌恶等经济决策任务中的偏差程度。","baseline":"无对照","findings":"LLM 在人类行为与偏好逻辑无关时仍表现出社会证明偏差；当损失厌恶被明确描述为偏差时，LLM 仍会模仿，且科学报告中描述的偏差程度能预测 LLM 自身的偏差程度。","reliability":"论文未讨论","relevance":"该研究揭示了 LLM 仿真中偏差产生的机制，有助于理解仿真失效的条件，对评估 LLM 作为人类被试替代品的可靠性具有批判性价值。","inspiration":"借鉴其通过提示词操纵信息内容来分离偏差来源的设计，可用于检验 LLM 是否仅因文本提及行为就模仿偏差｜可迁移到资产定价实验中的投资者情绪偏差、信贷审批中的歧视偏差、消费者跨期选择中的现时偏差等场景｜以 LLM 为被试，处理为提供包含偏差描述的科学报告或行为数据，结果变量为 LLM 在相应经济决策任务中的偏差程度，并与真实人类实验数据（如实验室资产定价实验或信贷审批审计研究）进行对照。"}},{"id":"2608.11794","version":1,"title":"Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion","zh_title":"面向AI聊天机器人的有意义透明度：披露说服意图可降低说服效果","abstract":"The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of content provenance and AI involvement. But the effects of such disclosures remain uncertain. We test two disclosure approaches in their impact on an AI chatbot's persuasive appeal. In a preregistered experiment, 1,500 UK adults held a short conversation with a persuasive chatbot about one of 60 policy issues. The chatbot was identical for everyone. We randomized the disclosure that people received: nothing (control), a prominent disclosure that they were interacting with an AI (T1), or that disclosure plus the chatbot's persuasive intent and instructions (T2). The chatbot shifted attitudes by 12.6 points on a 100-point scale in the control group. The AI-identity disclosure was practically equivalent to no disclosure, with a 13.1-point shift, whereas the additional intent disclosure cut the persuasive effect roughly in half to 6.3 points. It also made participants view the campaign's methods as less acceptable and support stronger penalties against it. For direct chatbot interactions, transparency about AI identity alone does not meaningfully impact its influence. While current rules emphasize what a system is, our results show why the regulation of persuasive AI must also address what the system is trying to do.","authors":["Adrian Rauchfleisch","Andreas Jungherr"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11794","pdf_url":"https://arxiv.org/pdf/2608.11794","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","说服实验","透明度"],"reason":"用LLM聊天机器人对真人做实验，测量态度改变，有真实人类数据对照，涉及政策说服…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":2,"question":"在直接的人机对话中，披露AI身份或额外披露说服意图，会如何影响AI聊天机器人的说服效果？","design":"预注册实验：1500名英国成年人与同一个说服性聊天机器人就60个政策议题之一进行简短对话；随机分配三种披露条件：无披露（对照）、显著披露AI身份（T1）、披露AI身份加说服意图和指令（T2）；测量态度改变（100分量表）、对活动方法的接受度、对惩罚的支持等。","baseline":"对照组（无披露）的真实人类态度改变数据，以及T1、T2组的人类反应数据，作为不同披露条件下的对照基准。","findings":"对照组态度改变12.6分，仅披露AI身份（T1）效果几乎等同（13.1分），而额外披露说服意图（T2）将说服效果减半至6.3分。T2还降低了参与者对活动方法的接受度，并增加了对更严厉惩罚的支持。","reliability":"论文指出，披露说服意图虽降低说服力但未消除，且引发对活动、互动和赞助方的负面评价，可能带来成本；未来需研究负面反应何时转移到背后的原因和行动者，以及AI竞选更普遍时是否导致更多回避。","relevance":"该研究用真实人类被试与LLM聊天机器人互动，测量态度改变，有严格对照和预注册，直接检验披露政策的效果，对关注LLM仿真和说服效应的研究者很有参考价值。","inspiration":"值得借鉴的是其随机披露处理与对照设计，以及用等价检验评估披露效果是否可忽略。｜可迁移到政策沟通或金融营销中AI顾问的说服效果评估，例如AI理财建议对投资决策的影响。｜设计：招募真实投资者作为被试，随机分配无披露、披露AI身份、披露AI身份及推销意图三组，让AI聊天机器人推荐某理财产品，测量投资意愿和风险感知，并与人类理财顾问的推荐效果进行对照。"}},{"id":"2608.05224","version":3,"title":"Small Foundation Models of Human Cognition and Behaviour","zh_title":"人类认知与行为的小型基础模型","abstract":"Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. For in-distribution simulations, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.","authors":["Nick Oh","Fernand Gobet"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-12","first_seen":"2026-08-07","revised_at":"2026-08-12","abs_url":"https://arxiv.org/abs/2608.05224","pdf_url":"https://arxiv.org/pdf/2608.05224","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知代理","算法保真度"],"reason":"用LLM替代人类被试复现心理实验，有真实人类数据对照，并评估仿真可靠性与失效条…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":2,"question":"在人类行为预测中，模型规模、适配器容量和训练数据量如何影响分布内与分布外的仿真准确度？模型究竟利用了任务结构还是统计捷径？","design":"在Psych-101数据集（包含160个实验的1070万条试次选择）上，对4个架构家族、135M到14B参数的14个模型进行监督微调，变化LoRA秩和训练数据子集，通过逐步遮蔽提示中的指令、刺激、反馈和选择历史四个信息通道，以及置换试次顺序，来诊断模型使用的信息。","baseline":"以Psych-101中6万余名人类被试的真实选择数据为对照基准。","findings":"分布内预测中，模型规模几乎不影响性能，0.6B至1B参数即可匹配70B基线；分布外泛化中，规模优势明显，更大模型能更好地迁移到新任务结构。遮蔽刺激和反馈内容会破坏75.7%的已学习信息，使模型表现低于随机水平，表明模型并非仅依赖选择历史。","reliability":"模型仅能作为心理实验的噪声上限估计器，其适用范围受限于训练数据中出现的实验范式，无法推广到未见过的范式。","relevance":"该研究直接以真实人类数据为基准，系统评估了LLM作为人类被试替代品的可靠性、规模需求与失效条件，与您关注的经济学实验仿真和批判性评估高度吻合，值得精读。","inspiration":"可借鉴其通过逐步剥离信息通道和置换试次顺序来诊断模型是否利用任务结构的方法，用于检验经济仿真中LLM是否真正理解经济激励而非依赖表面统计模式。｜可迁移到行为经济学中的跨期选择或风险决策实验，检验LLM是否利用延迟时间、概率等刺激内容而非仅记忆选择序列。｜以LLM作为被试，在跨期选择任务中系统遮蔽金额、延迟天数、反馈结果等信息通道，以真实人类选择数据（如Andersen et al., 2008）为基准，测量遮蔽前后预测准确率的变化，判断模型是否习得经济偏好结构。"}},{"id":"2608.09937","version":1,"title":"Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory","zh_title":"审慎考量文化：利用文化共识理论分析单文化与多文化环境下大语言模型的对齐","abstract":"Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural environments within a country. In this work, we leverage cultural consensus theory (CCT) from cultural anthropology to model such multidimensional nuance. Applying CCT to the World Values Survey (WVS) across 10 countries and 12 domains, we demonstrate that models frequently misrepresent cultural structures by either failing to form cohesive consensus or severely over-regularizing consensus. Through explicit representation of intra-group variance, CCT provides actionable diagnostics to evaluate when models reflect true human diversity versus algorithmic homogenization.","authors":["Krishna Pothugunta","John P. Lalor"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09937","pdf_url":"https://arxiv.org/pdf/2608.09937","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["文化仿真","人类数据对照","算法保真度"],"reason":"用LLM复现文化调查，与真实人类数据对照，评估仿真偏差，批判性指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":6,"question":"LLM在跨文化调查中能否准确复现人类群体的文化共识结构，而不仅仅是分布匹配？","design":"使用10个LLM组成的集成模型模拟10个国家的人类受访者，基于世界价值观调查（WVS）的12个文化领域问题生成回答，然后应用文化共识理论（CCT）分析模型回答的共识结构，并与人类数据进行对比。","baseline":"世界价值观调查（WVS）中10个国家的人类受访者真实回答数据。","findings":"LLM在不同文化领域表现出截然不同的共识结构：某些领域无法形成连贯共识，另一些领域则过度正则化产生虚假共识；即使模型成功匹配人类共识，也会系统性地夸大共识强度，将人类多样性压缩为算法同质化。","reliability":"论文指出模型行为高度依赖领域，在幸福与健康等领域完全无法形成共识，而在科技感知等领域则虚构非人类共识；CCT诊断显示模型倾向于消除组内差异，无法反映真实的文化多元性。","relevance":"该研究直接使用LLM复现文化调查并与真实人类数据对照，系统评估了仿真偏差并指出失效条件，高度契合研究者对LLM人类仿真可靠性及批判性分析的兴趣，值得精读原文。","inspiration":"借鉴文化共识理论（CCT）从共识结构和组内方差角度评估仿真质量，而非仅比较均值或分布，为经济金融仿真实验提供了更精细的诊断工具。｜可迁移到跨文化消费者信心调查或通胀预期形成的仿真研究中，检验LLM能否复现不同国家或群体的预期共识模式。｜以LLM集成作为被试，模拟多国消费者回答预期调查问题，处理为不同国家提示，结果变量为预期值及共识强度，以密歇根大学消费者调查或欧洲央行专业预测者调查的真实数据作为对照基准。"}},{"id":"2608.10186","version":1,"title":"The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse","zh_title":"协商赤字：对民主话语中LLM的实证批判","abstract":"LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer exists and where decision quality instead depends on integrating pluralistic perspectives to find mutually acceptable solutions. We argue that LLM reasoning capacity on this class of problems cannot be fully inferred from verifiable-task benchmarks, and that procedural evaluations of LLM discourse (respectfulness, justification, engagement) are systematically insufficient. We apply the Deliberative Reason Index (DRI), a measure developed in political science and validated across citizen assemblies, as a tool for evaluating reliable group-level reasoning on pluralistic, non-verifiable problems. Synthesizing recent evidence across 1,980 five-agent LLM runs on 12 citizen-assembly topics across 11 frontier model configurations, we find that LLM groups produce discourse with procedural quality comparable to human deliberation, while gains in intersubjective consistency are small, topic-dependent, and concentrated on tractable rather than ethically contested questions. LLM groups exhibit roughly one-third the perspective diversity of human assemblies and reverse the human convergence pattern: human deliberation decreases dispersion as diverse views synthesise, whereas LLM deliberation increases it. Engineering diversity through persona prompting does not restore the human dynamic but inverts which component of deliberative reasoning is updated. Our conclusion is constraining rather than prohibitive: LLMs can function as tools supporting human reasoning on pluralistic problems, but current evidence does not license treating them as autonomous deliberative agents.","authors":["Maurice Flechtner"],"categories":["cs.MA","cs.AI","cs.CY"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10186","pdf_url":"https://arxiv.org/pdf/2608.10186","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A3","B1","B4"],"tags":["LLM仿真","民主协商","人类数据对照"],"reason":"用LLM群体模拟民主协商并与真实公民大会数据对照，批判性指出仿真失效条件，高度…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":7,"question":"LLM在需要整合多元视角的无客观正确答案的民主协商问题上，能否展现出可靠的群体推理能力？","design":"使用11种前沿LLM配置组成5智能体小组，在12个公民大会议题上运行1980次模拟协商，测量过程质量、结果质量（DRI）和视角多样性，并与真实公民大会数据对照。","baseline":"真实公民大会数据，包括过程质量指标、DRI得分和视角多样性分布。","findings":"LLM小组的过程质量与人类相当，但DRI增益小且集中于易处理问题，视角多样性仅为人类的三分之一，且协商后观点分散度上升而非收敛。通过角色提示注入多样性未能恢复人类动态，反而颠倒了协商推理的更新成分。","reliability":"论文指出当前证据不支持将LLM视为自主协商主体，其群体推理在伦理争议问题上失效，且过程质量与实质质量脱节，多样性工程无法复现人类收敛模式。","relevance":"该研究直接以真实公民大会为基准，系统批判了LLM在多元价值协商中的仿真失效条件，对关注经济学实验和政策评估中LLM替代人类被试的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其采用真实群体协商数据作为基准对照、并构建过程-结果-多样性三维必要条件的评估框架。｜可迁移到公共政策偏好聚合实验，如碳税分配方案协商或最低工资调整的公众咨询模拟。｜以LLM模拟不同利益群体代表，施加协商干预，测量政策偏好变化与观点收敛，并以真实公众咨询或协商式民调数据作为对照基准。"}},{"id":"2608.07498","version":1,"title":"Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation","zh_title":"知你即一切：LLM代理在社交媒体模拟中实现近乎完美的画像一致性反应预测","abstract":"Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing. A central open question is whether persona-prompted LLMs can simulate individual-level social media reactions with sufficient accuracy to support either application, and how accuracy depends on profile completeness, model selection, and the generalization challenge posed by novel post content. This study benchmarks twelve LLM configurations on binary like/dislike prediction across 296 survey-based agent profiles and 26 ground-truth-mapped posts under three profile conditions, with leave-post-out machine learning classifiers as baselines. Across full-profile conditions, accuracy ranges from 75.54% to 96.68%, with a 30-point spread attributable primarily to model selection and confirmed by paired McNemar tests with agent-level bootstrap intervals. GPT-5.5 Pro accuracy degrades monotonically from 96.68% under a full profile to 62.32% under a reduced profile and to 51.00% with demographics alone, the last indistinguishable from the majority-class baseline, which confirms that demographic inference provides negligible predictive signal. Supervised classifiers collapse to 15.4% under leave-post-out, while LLMs sustain genuine zero-shot generalization unavailable to trained methods. Adaptive reasoning improves accuracy substantially for some models. Inter-model agreement is nearly double for posts with direct profile anchors (mean \\k{appa} = 0.44) than for posts without them (\\k{appa} = 0.23), and the least heterogeneous configuration homogenizes 34% of simulated population reactions. Results validate LLM-based simulation for recommender system stress-testing while documenting the behavioral accuracy that makes large-scale synthetic agent swarms a credible threat to public opinion.","authors":["Ljubisa Bojic","Ljiljana Matic","Joerg Matthes","Milan Cabarkapa","Bojana Dinic","Jue Wang"],"categories":["cs.HC","cs.AI","cs.LG","cs.MA"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07498","pdf_url":"https://arxiv.org/pdf/2608.07498","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","社交媒体模拟","算法保真度"],"reason":"用LLM代理模拟社交媒体反应，与真实人类数据对照，评估仿真准确性与失效条件，直…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":1,"question":"基于个人资料的LLM代理能否以足够准确度模拟个体层面的社交媒体反应（点赞/不喜欢），以及准确度如何取决于资料完整性、模型选择和帖子内容新颖性？","design":"使用12种LLM配置（不同模型、提示策略）扮演基于调查的296个代理资料，在三种资料条件下（完整资料、简化资料、仅人口统计）对26个真实帖子进行二分类点赞/不喜欢预测，以留一帖子的监督学习分类器为基线。","baseline":"真实人类数据：296个基于调查的代理资料和26个真实帖子，每个帖子有真实用户反应作为对照基准。","findings":"完整资料下LLM准确率75.54%-96.68%，模型选择造成30个百分点差异；仅人口统计时准确率降至51%，与多数类基线无差异。LLM在留一帖子泛化中保持零样本能力，而监督分类器崩溃至15.4%。","reliability":"论文指出LLM仿真在资料不完整时失效（仅人口统计无预测力），且模型间一致性在无直接资料锚点的帖子上较低（κ=0.23），最同质化配置会抹平34%的个体差异。","relevance":"高度相关：用LLM代理模拟个体行为并与真实人类数据对照，系统评估资料完整性、模型选择对仿真准确性的影响，并明确失效条件，直接回应研究者对可靠性与偏差的关注。","inspiration":"借鉴多条件资料消融设计（完整/简化/仅人口统计）和留一帖子泛化测试来分离模型能力与记忆效应。｜可迁移到消费者偏好预测或政策态度模拟，如基于个人财务特征和态度资料预测个体对税收政策的支持度。｜以真实调查数据构建代理资料，用LLM预测个体对某项经济政策（如碳税）的二元态度，处理为资料完整性梯度，结果变量为支持/反对，以实际调查回答为对照基准。"}},{"id":"2608.09717","version":1,"title":"How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans","zh_title":"大语言模型如何判断社交吸引力？基于理论驱动的人物画像在多个LLM和人类中的评分证据","abstract":"Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction from theory-grounded persona profiles constructed from ten psychological and relational constructs and organized into three tiers: socially attractive, socially mixed, and socially unattractive. We examine LLM ratings in two studies and compare them with human judgments in a third study. In Study 1, 34 LLMs rated 12 profiles across three repeated runs. Although some models tended to give higher or lower ratings overall, they showed strong stability across runs, consistent three-tier ordering, and high agreement in relative profile ordering. Study 2 examined sensitivity to gender presentation using six matched name-and-pronoun profile pairs and a separate pronoun-only test with a gender-neutral name, finding no significant effects in either analysis. In Study 3, 198 human participants evaluated the six matched profiles from Study 2. Their ratings reproduced the three-tier structure and followed a profile ordering consistent with that of the LLMs. However, LLMs rated attractive profiles more positively and unattractive profiles more negatively than humans, while neither group showed a significant overall effect of gender presentation.","authors":["Hasan Mahmud","Khawaja Abaid Ullah","Mohammad Javad Khojasteh","Jamison Heard","Prabu David"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09717","pdf_url":"https://arxiv.org/pdf/2608.09717","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","社交判断偏差"],"reason":"用LLM评估社交吸引力，并与198名人类被试对照，发现LLM评分更极端，直接检…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":5,"question":"LLM能否基于理论构建的人物画像可靠地评估社交吸引力，其评分是否与人类一致且不受性别呈现影响？","design":"研究一用34个LLM对12个理论分层的虚拟学生画像重复评分3次，检验评分稳定性、层级区分和模型间一致性；研究二用6对匹配姓名与代词的画像及仅变代词的测试，检验LLM对性别呈现的敏感性；研究三让198名人类被试对同样的6对画像评分，作为人类基准。","baseline":"198名人类被试对6个匹配画像的社交吸引力评分，与LLM评分进行直接比较。","findings":"LLM评分跨运行稳定，能一致区分理论上的社交吸引力层级，且相对排序与人类高度一致；但LLM对高吸引力画像评分比人类更积极，对低吸引力画像评分比人类更消极，表现出评分极端化，而性别呈现对两组均无显著影响。","reliability":"论文指出LLM评分可能反映训练语料中的文化规范和偏见，且仅在特定画像和社交吸引力场景下验证，未考察其他社会判断任务或真实互动情境中的效度。","relevance":"该研究直接以真实人类数据为基准，检验LLM在主观社会判断中的仿真效度与偏差，并发现评分极端化现象，高度契合研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"借鉴其理论驱动构建分层画像、多模型重复测试及与人类被试直接对照的设计，可迁移到信贷审批或招聘筛选中的歧视研究，例如用LLM扮演信贷员评估不同性别/种族的贷款申请人画像，以真实银行审批数据为基准，检验LLM是否复现或放大人类偏见。"}},{"id":"2608.08691","version":1,"title":"EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility","zh_title":"EnergyBridge：家庭能源管理、用户参与和电网灵活性的基准测试","abstract":"Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity only when residents authorize a plan and the promised response is delivered. Existing benchmarks evaluate control but omit event-specific authorization. We present EnergyBridge, a benchmark and agent framework connecting capacity reporting, household authorization, and physical execution. It combines region-specific EnergyPlus environments for Tianjin and Berlin with an LLM-based User Participation Simulator. Against 584 persona- and event-matched human role-play judgments, the LLM-based User Participation Simulator preserves method ordering with a 5.3-point mean absolute acceptance error. Across conventional controllers and agent baselines, EnergyBridge achieves the highest simulated authorization, lowest event-window energy, and the most reliable capacity commitment in both regions. We release human data and codes for reproducible human-centered grid-flexibility research: https://github.com/Agentic-Intelligence-Lab/EnergyBridge.","authors":["Xudong Wu","Zeqing Wu","Jiarui Zhang","Xuhao Fan","Ziang Ding","Yuming Zhuang","Mingqi Yuan","Yilun Du","Hongjie Jia","Yunfei Mu","Jiayu Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08691","pdf_url":"https://arxiv.org/pdf/2608.08691","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","用户参与模拟","电网灵活性"],"reason":"用LLM模拟用户参与授权，并与584条人类角色扮演判断对照，涉及能源政策评估场…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":4,"question":"如何在家庭能源管理中，将居民参与授权与物理灵活性执行相结合，实现可靠的虚拟电厂容量承诺？","design":"使用基于大语言模型的用户参与模拟器，扮演天津和柏林的家庭居民，根据584条人物角色和事件匹配的人类角色扮演判断进行校准；在EnergyPlus建筑能耗仿真环境中，对虚拟电厂灵活性请求进行预事件容量报告、设备级计划生成、居民授权决策和物理执行的全流程仿真，测量授权接受率、事件窗口能耗和容量承诺可靠性。","baseline":"584条人物角色和事件匹配的人类角色扮演判断，用于校准和验证LLM模拟器的授权接受误差（平均绝对误差5.3点）。","findings":"LLM用户参与模拟器能保持方法排序，授权接受平均绝对误差仅5.3点；EnergyBridge在天津和柏林两地均实现了最高的模拟授权率、最低的事件窗口能耗和最可靠的容量承诺。","reliability":"论文未讨论","relevance":"该研究直接使用LLM模拟家庭用户参与授权决策，并与真实人类角色扮演数据进行对照，属于经济学实验和政策评估场景中的人类仿真应用，值得精读其仿真校准方法和人机对照设计。","inspiration":"借鉴其将LLM模拟器与真实人类判断进行事件级匹配校准的方法，可迁移到消费者需求响应或绿色能源订阅政策的参与决策研究中；可设计一个实验，用LLM扮演不同收入与环保态度的家庭，施加动态电价或碳配额信息处理，测量其授权接受率和负荷转移量，并以真实居民调查或现场实验数据作为对照基准。"}},{"id":"2608.07490","version":1,"title":"Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents","zh_title":"经验敏感的游戏学习：人类与语言代理的行为研究","abstract":"Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players learn from repeated interaction. We study experience-sensitive game learning: how gameplay experience changes the decision-making behavior of humans and language agents. We formulate experience-sensitive game learning as a framework for analyzing behavioral change across repeated gameplay, rather than only final score or win rate. We introduce a suite of interactive games with reusable strategic structure, together with cross-game greedy-to-global metrics and game-specific behavioral diagnostics that make experience-driven change observable from action traces. We also collect repeated-game trajectories from human players and evaluate recent self-evolving language agents in the same behavioral metric space. Our results show that human players exhibit interpretable and relatively stable shifts from locally greedy heuristics toward more global strategic decisions. In contrast, current self-evolving agents often show noisy and transient gains, suggesting that existing self-evolution methods remain limited in converting gameplay experience into durable changes in decision-making behavior.","authors":["Yingying Guo","Zhuoxuan Ju","Ruibo Ming","Ruicheng Feng","Jinjin Gu"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07490","pdf_url":"https://arxiv.org/pdf/2608.07490","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM人类仿真","行为博弈","算法保真度"],"reason":"用LLM代理模拟人类游戏学习行为，并与真实人类数据对照，评估行为变化差异，指出…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":3,"question":"在重复游戏中，人类和语言智能体的决策行为如何随游戏经验积累而改变？","design":"构建一组具有可重用策略结构的交互式游戏，定义跨游戏的贪婪到全局指标和游戏特定的行为诊断指标，收集人类玩家和自进化语言智能体的重复游戏轨迹，在同一行为度量空间中分析经验驱动的行为变化。","baseline":"收集了人类玩家（硕士和博士生）在四子棋、Othello6和CircleCat上的重复游戏轨迹，作为经验敏感学习的参照基准。","findings":"人类玩家表现出可解释且相对稳定的从局部贪婪启发式向更全局策略决策的转变；当前自进化智能体则常表现出噪声大且短暂的提升，表明现有自进化方法难以将游戏经验转化为持久的决策行为改变。","reliability":"论文未讨论","relevance":"该研究直接以LLM代理模拟人类游戏学习行为，并与真实人类数据对照，评估行为变化差异，符合研究者对LLM仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"借鉴其通过定义行为诊断指标（如贪婪到全局转变）来量化经验驱动的行为变化，而非仅看最终得分的方法。｜可迁移到经济决策实验，如消费者跨期选择或投资者风险偏好学习。｜以LLM代理作为被试，施加重复跨期选择任务，测量其时间偏好一致性的变化，并与真实人类实验数据对照，分析学习动态的差异。"}},{"id":"2608.07367","version":1,"title":"People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe","zh_title":"人不仅是其国家：解构欧洲LLM价值观对齐的社会决定因素","abstract":"As Large Language Models (LLMs) are increasingly used as a primary source of information and advice, understanding their alignment to humans in terms of values becomes a pressing concern. A growing literature has leveraged large scale surveys to investigate to what extent LLMs' and humans' stated values and opinions align. With limited exceptions, studied populations have been defined country borders or cultural bounds. Yet, this focus neglects the role that socio-demographic divides may play for value alignment disparities. Relying on the European Social Survey, we address this knowledge gap by considering value alignment displayed with respect to 10 prominent commercial LLMs in terms of 15 socio-demographic variables as well as country of residence. Our analyses reveal that LLMs are indeed unequally aligned to the values of different socio-demographic groups, notably those defined by education, income, occupation and religion. When examining alignment at the individual level, a respondent's country, taken as a stand-alone variable, explains a substantial amount of variation that is on par with the full set of considered socio-demographics. Further disentangling the respective role of country-level and socio-demographic factors, we find they are complementary in explaining value alignment patterns, with their relative weights varying across the subset of questions considered.","authors":["Maria-Louisa Wightman","Guillaume Bied","Tijl De Bie"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07367","pdf_url":"https://arxiv.org/pdf/2608.07367","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM价值观对齐","社会调查复现","偏差分析"],"reason":"用LLM复现人类价值观调查，以欧洲社会调查为基准，分析对齐偏差，直接命中A1/…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":2,"question":"在欧洲国家中，LLM与人类价值观的对齐在不同社会人口统计因素和国家之间呈现何种差异模式？","design":"使用10个商业LLM重复回答欧洲社会调查（ESS）中的价值观问题，计算LLM回答与人类受访者回答的对齐分数，分析对齐分数在15个社会人口统计变量和国家之间的差异。","baseline":"欧洲社会调查（ESS）2023-2024波次中29个欧洲国家和以色列的人类受访者真实回答。","findings":"LLM与不同社会人口群体（尤其是教育、收入、职业和宗教定义的群体）的价值观对齐程度不平等；国家变量单独解释的对齐变异量与全套社会人口统计变量相当，且国家与社会人口统计因素在解释对齐模式上互补。","reliability":"论文未讨论","relevance":"该研究直接以真实大规模人类调查为基准，评估LLM在价值观对齐上的社会人口统计偏差，属于批判性仿真研究，命中研究者关心的A1、A2、B1、B3准则，值得精读。","inspiration":"借鉴其使用大规模社会调查作为基准、通过逆倾向加权和方差分解分离国家与社会人口因素贡献的方法；可迁移到信贷审批或保险定价中的算法公平性评估场景；以LLM作为信贷审批员，输入不同社会人口特征的虚拟申请人，测量审批结果差异，并以真实信贷审批数据或调查数据作为对照基准。"}},{"id":"2608.06379","version":1,"title":"Preventive Care Recommendations by Large Language Models","zh_title":"大语言模型的预防保健建议","abstract":"Preventive care services (PCS) extend life, yet physicians often underprioritize highly effective interventions such as lifestyle modifications (Zhang et al., JAMA Network Open 2020). We evaluated whether large language models (LLMs) replicate and augment physician prioritization of PCS under time constraints. Using Zhang et al.'s validated survey with two patients assessed during long and short visits, we compared seven LLMs with historical physicians. We generated 137 simulated physician personas matching cohort demographics and tested three prompts per model. Primary outcomes were concordance with physician rankings, measured by Spearman correlation, and Consensus-Stratified Agreement (CSA), the proportion of LLM selections rated 4 or higher that matched physician consensus across agreement strata. Secondary outcomes included life-years gained per prioritized choice (LYGPC), consistency, and selectiveness. Augmentation was assessed by having models revise physician rankings under three informative prompts, with delta LYGPC quantifying impact. LLMs closely mirrored physicians (mean Spearman = 0.83, SD = 0.11), with high CSA at extreme agreement ranges (94%, 197/210) but low CSA in moderate ranges (21%, 30/140), where they underprioritized lifestyle services (8.8% vs. 38% rated 4 or higher; P < .001). Several models exceeded physicians in LYGPC and consistency while being more selective. Time constraints affected physicians and LLMs similarly, increasing LYGPC and selectiveness but reducing consistency. Augmentation effects varied by model. Current LLMs reproduced physicians' time-sensitivity and base-rate prioritization while exacerbating underprioritized lifestyle interventions. Some models improved prioritization performance, but consistent augmentation will require value-aligned training, explicit time-constraint representation, and prospective real-world validation.","authors":["Eden Avnat","Elia Yanko","Ori Yoran","Raja-Elie E. Abdulnour"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06379","pdf_url":"https://arxiv.org/pdf/2608.06379","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","医生决策","人类数据对照"],"reason":"用LLM模拟医生决策并与真实医生数据对照，评估仿真可靠性及失效条件，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":1,"question":"大语言模型能否复现并增强医生在时间压力下对预防性服务的优先级排序？","design":"使用7个LLM模拟137名医生角色，匹配原始医生队列的人口统计特征，在长/短就诊时间两种条件下对两个虚拟患者完成预防性服务评级和排序任务，测量排序一致性、共识分层一致性、每项优先选择获得的寿命年、一致性和选择性，并通过三种提示策略让模型修正医生排序以评估增强效果。","baseline":"Zhang等人2020年发表的137名医生在相同调查中的真实排序和评级数据。","findings":"LLM与医生的排序高度相关（平均Spearman=0.83），在极端共识区间一致性高（94%），但在中等共识区间一致性低（21%），且严重低估生活方式干预服务（8.8% vs 38%）。部分模型在寿命年增益和一致性上优于医生，但时间压力对两者的影响相似，增强效果因模型而异。","reliability":"论文指出LLM在中等共识服务上一致性差，加剧了对生活方式干预的低估，且增强效果不稳定，需要价值对齐训练、显式时间约束表示和前瞻性真实世界验证。","relevance":"该研究直接以真实医生数据为基准，评估LLM模拟人类专业决策的可靠性与偏差，并揭示了在中等共识情境下仿真失效的条件，高度契合研究者对LLM仿真实验的批判性关注。","inspiration":"借鉴其通过分层共识分析（CSA）揭示仿真在中等共识区间失效的方法，可迁移到经济预测或政策评估场景中检验LLM对分析师共识的复现偏差。｜具体可应用于信贷审批或投资建议场景，考察LLM模拟信贷员或分析师在信息不完全下的决策。｜以真实信贷审批数据为基准，让LLM扮演不同经验水平的信贷员，在高低信息量条件下进行审批决策，测量其与人类审批员排序的相关性及在不同共识水平上的偏差。"}},{"id":"2608.04009","version":2,"title":"SocietyBench: Forecasting Counterfactual Social-World Evolution","zh_title":"SocietyBench：预测反事实社会世界演化","abstract":"Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way real social events unfold, has barely been measured. We introduce SocietyBench, an end-to-end benchmark that takes a one-line event topic, collects Web news and social-media posts across five platforms, distills them into a date-indexed timeline that keeps factual events and a public-opinion layer separate, and then turns every cutoff date on that timeline into an audited bank of forecasting questions. Questions are scored on two orthogonal 100-point axes: probability calibration and temporal accuracy. Before any model sees a timeline, a three-phase procedure replaces every named entity and shifts every date by a per-event constant, turning a real arc into a counterfactual social world -- structurally identical to what happened, but stripped of the surface labels a model could match against pre-training memory. On five heterogeneous events and 125 prediction points in Chinese and English editions, the strongest of six frontier LLMs reaches only 75.0 out of 100, against a trivial anchor of 50. The two axes come apart: a model can be calibration-strong but time-weak, or the reverse. Three agent frameworks built on a shared base model fail to improve on that base, and two model-free heuristics trail every LLM. Per-event gaps reach 21.4 points on a single axis, which is our main argument for evaluating on several events rather than one. All anonymized timelines, question banks, ground truth, and scoring code are released.","authors":["Zhenran Wang","Zhonghan Bian","Jinsong Li","Zhangyang Qi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-10","first_seen":"2026-08-05","revised_at":"2026-08-10","abs_url":"https://arxiv.org/abs/2608.04009","pdf_url":"https://arxiv.org/pdf/2608.04009","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A3","B1","B4"],"tags":["社会模拟","LLM预测","反事实推理"],"reason":"用LLM预测反事实社会事件演化，有真实新闻/舆论数据对照，并评估校准与时间准确…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":3,"question":"大型语言模型能否准确预测反事实社会事件的演化，包括事实进展与公众舆论？","design":"构建SocietyBench基准：基于五个真实社会事件，自动采集新闻与社交媒体数据，生成匿名化时间线，将每个截止点转化为概率校准和时间准确性两类预测问题，评估LLM及智能体框架的预测表现。","baseline":"以真实事件后半段的时间线节点作为事实与舆论演化的真实基准。","findings":"最强前沿LLM仅得75.0分（满分100），概率校准与时间准确性两个维度表现分离；智能体框架未能超越基础模型，不同事件间模型表现差异可达21.4分。","reliability":"论文指出单事件评估不可靠，需多事件测试；匿名化虽防记忆但可能改变社会动态结构；未讨论模型在更长预测窗口或更多类型事件上的泛化局限。","relevance":"该研究用LLM预测社会事件演化，有真实新闻与舆论数据对照，并系统评估校准与时间准确度，直接回应了研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"可借鉴其匿名化反事实设计以剥离模型记忆干扰，并用双轴评分（概率校准与时间误差）全面衡量预测质量。｜可迁移至政策公告的预期形成研究，如央行加息声明后的市场反应预测。｜以LLM为被试，提供匿名化的宏观经济事件时间线，要求预测后续资产价格变动概率与时间，以真实市场数据为对照基准。"}},{"id":"2608.07316","version":1,"title":"Natural Language Processing Psychometrics","zh_title":"自然语言处理心理测量学","abstract":"Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure. NLP Psychometrics treats psychological prediction from text as a psychometric problem, linking scores to interpretable linguistic evidence and testing beyond the training text format. Nine LLMs, conditioned on controlled personas (cognitive digital shadows), completed psychometric questionnaires with textual explanations per item. We extracted emotional profiles and syntactic-semantic structure via textual forma mentis networks, combined with personality and sociodemographic variables in ablated random forest (RF) regressors, using SHAP to identify which features drove performance and in which direction. Full RF models explained up to 70.8% of variance in life satisfaction (SWLS), 55.7% in depression (PHQ-9), and, for DASS-21, 68.5% depression, 76.0% anxiety, 72.4% stress. Sociodemographics alone explained no meaningful variance in depression, anxiety, or stress, but did so for life satisfaction, where emotion features and income were the strongest predictors; neuroticism and network topology instead dominated depression and anxiety, reversing direction between them. Without retraining, RF models separated diaries from low- and high-score personas ($r$ up to 0.91) and, using only network/emotion features, classified clinical from control participants in real transcripts with up to 68% accuracy. These results show the promise and limits of synthetic data: LLM personas can expose model biases, recover patterns consistent with clinical rumination, and support psychometric prediction from human text without a matched questionnaire, but cannot substitute for human validation. NLP Psychometrics makes these distinctions explicit, measurable, and testable through interpretable AI and network/emotional features.","authors":["Edoardo Sebastiano De Duro","Emma Franchino","Massimo Stella"],"categories":["cs.CL","cs.AI","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07316","pdf_url":"https://arxiv.org/pdf/2608.07316","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","心理测量","可解释AI"],"reason":"用LLM persona模拟人类心理测量，有真实人类数据对照，并讨论合成数据的…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":4,"question":"如何从语言文本中可解释地推断心理构念，并评估LLM生成的合成数据在心理测量中的有效性与局限？","design":"使用9个LLM，通过控制人格与社会人口学变量构建认知数字影子（persona），让它们完成心理测量问卷并对每个条目生成文本解释；从这些文本中提取情绪特征和句法-语义网络特征，结合人格与社会人口学变量，训练随机森林回归模型预测心理健康得分，并用SHAP解释特征贡献。","baseline":"真实人类数据：临床访谈转录文本（用于分类临床vs对照）、真实日记文本（用于分离高/低分persona），以及已有的心理测量问卷常模。","findings":"全特征随机森林模型可解释生活满意度70.8%、抑郁55.7%、焦虑76.0%和压力72.4%的方差；社会人口学变量单独对抑郁、焦虑、压力无显著解释力，但对生活满意度有贡献，其中情绪特征和收入是主要预测因子，而神经质和网络拓扑结构主导抑郁和焦虑的预测且方向相反。仅用网络/情绪特征，模型在真实临床转录文本上区分临床与对照组的准确率达68%。","reliability":"论文明确指出LLM persona不能替代人类验证，合成数据可暴露模型偏差、恢复与临床反刍一致的模式并支持从人类文本进行心理测量预测，但无法取代真实人类数据；LLM问卷回答可能不稳定、对提示敏感且方差低于人类，需谨慎实验设计。","relevance":"该研究直接以LLM persona模拟人类心理测量，有真实临床和日记数据作为对照基准，并系统讨论了合成数据的可靠性与失效条件，完全契合研究者对LLM仿真实验、基准对照和批判性评估的关注，值得精读原文。","inspiration":"借鉴其用控制性persona生成文本并提取网络/情绪特征进行可解释预测的方法，以及用SHAP分析特征贡献方向的做法。｜可迁移到消费者信心或投资者情绪调查中，用LLM模拟不同人口学与人格特征的受访者，生成开放式回答并预测其经济预期指数。｜设计：以LLM扮演不同收入、人格的消费者，施加宏观经济新闻文本作为处理，收集其对未来经济状况的开放式描述，提取情绪与语义网络特征预测消费者信心指数，并以真实密歇根消费者调查的文本回答和指数作为对照基准。"}},{"id":"2603.00059","version":3,"title":"Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data","zh_title":"随机鹦鹉还是和谐合唱？测试五大领先LLM用合成数据复现人类调查的能力","abstract":"How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article presents a comparison between a human-respondent survey of 420 Silicon Valley coders and developers and synthetic survey data designed to simulate real survey takers generated by five leading Generative AI Large Language Models: ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro plus Claude CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, and DeepSeek 3.2. Our findings reveal that while AI agents produced technically plausible results that lean more towards replicability and harmonization than assumed, none were able to capture the counterintuitive insights that made the human survey valuable. Moreover, deviations grouped together for all models, leaving the real data as the outlier. Our key finding is that while leading LLMs are increasingly being used to scale, replicate and replace human survey responses in research, these advances only show an increased capacity to parrot conventional wisdom in harmony with each other rather than revealing novel findings. If synthetic respondents are used in future research, we need more replicable validation protocols and reporting standards for when and where synthetic survey data can be used responsibly, a gap that this paper fills. Our results suggest that synthetic survey responses cannot meaningfully model real human social beliefs within organizations, particularly in contexts lacking previously documented evidence. We conclude that synthetic survey-based research should be cast not as a substitute for rigorous survey methods, but as an increasingly reliable pre- or post-fieldwork instrument for identifying societal assumptions, conventional wisdoms, and other expectations about research populations.","authors":["Jason Miklian","Kristian Hoelscher","John E. Katsos"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-07","first_seen":"2026-02-10","revised_at":"2026-08-07","abs_url":"https://arxiv.org/abs/2603.00059","pdf_url":"https://arxiv.org/pdf/2603.00059","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B2","B4"],"tags":["LLM仿真","调查复现","可靠性评估"],"reason":"直接对比LLM合成调查与真人数据，评估仿真可靠性并提出报告标准，高度契合。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":1,"question":"领先的大语言模型生成的合成调查数据能否复现人类受访者的回答，尤其是能否捕捉反直觉的洞见？","design":"使用五种领先LLM（ChatGPT Thinking 5 Pro、Claude Sonnet 4.5 Pro、Gemini Advanced 2.5 Pro、Incredible 1.0、DeepSeek 3.2）模拟硅谷程序员和开发者的调查受访者，通过提示词生成合成调查数据，测量其对伦理与政治意识形态问题的回答模式。","baseline":"一项对420名硅谷程序员和开发者的真实人类调查（Miklian and Hoelscher 2026）。","findings":"LLM生成的合成数据在技术上看似合理且不同模型间高度一致，但均未能捕捉到人类调查中的反直觉洞见；合成数据聚集在一起，真实人类数据反而成为离群值。","reliability":"论文指出合成调查数据无法在缺乏先前文献证据的情境下有意义地模拟真实人类的社会信念，且所有模型都倾向于复述共识性常识而非揭示新发现。","relevance":"该研究直接对比LLM合成调查与真人数据，评估仿真可靠性并提出报告标准，高度契合研究者对LLM人类仿真实验、基准对照及失效条件的关注，值得精读原文。","inspiration":"借鉴其多模型对比与真实人类基准的设计，可评估LLM在特定人群中的仿真偏差。｜可迁移到经济金融领域的调查实验，如消费者信心预期、通胀预期或政策偏好调查。｜以真实消费者调查为基准，用多个LLM生成合成消费者预期数据，比较其对未来经济变量的预测分布与真实调查的差异，检验LLM是否仅复述共识性预期。"}},{"id":"2608.06085","version":1,"title":"Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference","zh_title":"信号还是虚假线索？一项关于LLM社会推断中调查国家元数据的随机审计","abstract":"Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label's uniform, record-independent origin reduces its country-directed uptake, and whether verified survey country lowers held-out Brier loss. Independent population anchors and recorded human answers measure direction and consequence across five fixed API models, six countries, and seven development-selected targets. In the primary post-review 72-record panel, opaque and disclosed-random labels each produced country-direction shifts of 0.214. Paired attenuation was 0.0003 (95% CI [-0.0157, 0.0166]). Verified country reduced Brier loss by 0.040 (95% CI [0.024, 0.056]), while random-label regret included zero. A non-overlapping mixed-coverage consistency panel retained positive disclosed-random movement and verified utility, while attenuation remained uncertain. On the selected targets, verified metadata was useful in both panels, but disclosure did not reliably attenuate random-label uptake. PROV-FORECAST contains 14,400 paired item-level probability distributions from the corrected panel.","authors":["Yifan Lyu","Xinran Li","Jiaqi Qiao","Xiujuan Xu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06085","pdf_url":"https://arxiv.org/pdf/2608.06085","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查回答预测","算法保真度"],"reason":"用LLM预测个体调查回答，检验随机国家标签的误导效应，并与真实人类答案对照，评…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":3,"question":"在LLM预测个体调查回答时，披露随机分配的国家标签来源能否减弱其误导效应，以及真实调查国家信息能否降低预测误差？","design":"使用五个固定API模型，基于EVS/WVS联合数据集的个体记录，向模型提供同一受访者的10个已观测答案，要求预测7个目标问题的回答概率；通过对比不透明随机国家标签、披露为随机分配的国家标签、真实调查国家标签和无国家标签四种条件，测量预测方向偏移和Brier损失。","baseline":"EVS/WVS 2017-2022联合数据集中六国（中国、法国、英国、意大利、约旦、美国）12,770条真实个体调查记录，包含已观测和留出的人类答案。","findings":"在主要72记录面板中，不透明和披露的随机标签均产生0.214的国家方向偏移，披露未显著减弱该偏移（衰减仅0.0003，95% CI包含零）；真实调查国家使Brier损失降低0.040（95% CI [0.024, 0.056]），而随机标签的预测后悔值包含零。","reliability":"论文指出披露随机标签来源未能可靠减弱其误导效应，且结果可能受限于所选目标问题、国家和模型；留出Brier损失的重算需合法获取源数据。","relevance":"该研究直接以真实人类调查数据为基准，检验LLM在个体预测中对元数据线索的依赖与偏差，并区分信息效用与误导效应，与研究者关注的仿真可靠性及失效条件高度契合，值得精读。","inspiration":"借鉴其同记录内对比不同元数据来源（随机、披露随机、真实）的设计，分离线索的方向性误导与预测效用。｜可迁移到信贷审批或保险定价实验，检验LLM在引入申请人地域、性别等敏感属性时是否产生歧视性偏移及其是否因属性来源说明而减弱。｜以真实贷款违约数据为基准，将申请人部分财务指标作为已观测证据，随机分配或真实使用地域标签，让LLM预测违约概率，比较不同标签条件下的预测偏差和校准误差。"}},{"id":"2608.06115","version":1,"title":"Mind the Gaps: Mixture-of-Minds for Human Simulation","zh_title":"注意差距：用于人类仿真的思维混合模型","abstract":"Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual. Large language model simulators inherit this gap. They recover a population's central tendencies while flattening its heterogeneity, and they carry social biases and prompt brittleness that distort individual predictions. This paper introduces Anacreon, an audience simulation model that targets the individual level within a narrow, well-specified domain. Anacreon learns an authorship embedding that separates individuals, clusters a real qualitative corpus around seed people, and trains a dedicated adapter for each cluster, a mixture of minds, on a Gemma~4 12B base. It harvests demographics, psychological traits, and survey responses from public text, and augments each record with a chain-of-emotion. It reduces prompt brittleness by shuffling response options and reduces positive bias by balancing the training distribution. On a large, externally sourced survey, Anacreon reaches a state-of-the-art ordinal alignment of 0.775, the individual-level accuracy measure on which the field has converged, with a small residual bias. The work is a step toward drawing aggregate insight from faithfully simulated individuals.","authors":["Pranav Dahiya"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06115","pdf_url":"https://arxiv.org/pdf/2608.06115","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["人类仿真","调查预测","个体异质性"],"reason":"用LLM仿真个体回答调查，有真实人类数据对照，评估偏差与可靠性，涉及社会调查场…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":4,"question":"如何在窄领域内利用LLM模拟独立异质个体，以准确预测其调查回答，并缩小个体与群体层面的预测差距？","design":"Anacreon模型基于Gemma 4 12B，通过作者嵌入分离个体，对真实语料聚类后为每簇训练专用适配器（思维混合），从公开文本中提取人口统计、心理特质和调查回答，并用情绪链增强记录；通过打乱选项顺序减少提示脆弱性，平衡训练分布减少正向偏差。","baseline":"使用大型外部调查的真实人类个体回答作为对照基准。","findings":"Anacreon在外部调查上达到0.775的序数对齐（个体级准确度），为领域内最优；模型残余偏差较小，表明其能较忠实地模拟个体，从而从个体聚合出有意义的群体洞察。","reliability":"论文指出LLM模拟器在宽泛领域会因过度泛化而失效，Anacreon仅适用于窄而明确的领域；基础模型存在社会偏见，且RLHF会降低输出多样性，使助手型模型不适合模拟人类异质性。","relevance":"该研究直接针对LLM仿真人类被试的个体异质性和可靠性问题，有真实人类调查数据对照，并评估偏差，与研究者关注的经济学实验和政策评估场景高度相关，值得精读原文。","inspiration":"借鉴其用聚类适配器捕捉个体异质性、情绪链增强和平衡训练分布以减少偏差的方法｜可迁移到消费者金融决策调查仿真，如风险偏好、信贷选择等｜以公开社交媒体数据构建虚拟消费者，施加不同金融信息提示作为处理，测量其风险资产配置意愿，并以真实家庭金融调查数据（如SCF）作为对照基准。"}},{"id":"2608.06151","version":1,"title":"Reducing belief in conspiracy theories as they unfold using large language models","zh_title":"使用大语言模型减少实时阴谋论信念","abstract":"The emergence of conspiracy theories in the wake of major events is a significant societal challenge. Here we test whether conversational dialogues with a large language model (LLM) can reduce belief in immediately unfolding conspiracies. In experiments conducted in the days following the July 2024 assassination attempt on Donald Trump and the September 2025 assassination of Charlie Kirk, U.S. adults (Experiment 1: N = 472; Experiment 2: N = 1035) holding conspiratorial views about the crisis event engaged in a multi-turn conversation with an LLM prompted to reduce their conspiracy belief. Compared to control participants who either discussed an irrelevant topic with an LLM or viewed a static fact sheet, participants in the LLM treatment showed significantly reduced conspiracy beliefs in both experiments. We also found evidence of downstream effects of the LLM treatment, observing reduced belief in different conspiracies one to two months later in the wake of subsequent crisis events. These results shed light on the psychology of emerging conspiracies and highlight the potential for scalable, cognitively-focused interventions to counteract misinformation in the immediate aftermath of high-profile societal events.","authors":["Thomas H. Costello","Nathaniel Rabb","Michael Nicholas Stagnaro","Gordon Pennycook","David Rand"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06151","pdf_url":"https://arxiv.org/pdf/2608.06151","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","阴谋论干预","行为实验"],"reason":"用LLM对话干预阴谋论信念，有真实人类对照实验，评估干预效果与下游影响，属人类…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":5,"question":"在突发危机事件后，大语言模型对话能否即时降低人们对新兴阴谋论的信念？","design":"非仿真研究。该研究以真实人类为被试，在特朗普遇刺未遂和查理·柯克遇刺事件后数日内，招募持有阴谋论看法的美国成年人，随机分配至LLM驳斥对话组、静态信息清单组或无关话题对话对照组，通过前后测比较其对阴谋论的信念变化。","baseline":"真实人类对照：无关话题对话组和静态信息清单组作为对照条件，比较LLM对话干预的效果。","findings":"LLM驳斥对话显著降低了被试对自述阴谋论的信念，效果优于无关对话和静态信息清单；干预效果在1-2个月后的后续危机事件中仍有一定持续性，但对官方解释的信任提升不稳定。","reliability":"论文未讨论","relevance":"该研究直接使用LLM与真实人类进行对话干预，并以随机对照实验评估效果，符合研究者对LLM仿真人类行为、有真实人类基准、关注政策干预场景的兴趣，值得精读。","inspiration":"借鉴其多轮对话干预与多条件对照设计，以及利用突发事件窗口进行即时实验的方法。｜可迁移到经济政策沟通场景，如央行加息公告后，用LLM对话干预公众的通胀预期或政策误解。｜以真实投资者为被试，在政策公告后随机分配至LLM解释对话组或静态新闻组，测量其通胀预期、资产配置意愿的变化，并以调查数据或市场预期指标作为对照基准。"}},{"id":"2608.04205","version":1,"title":"MatrAIx: Simulating the World with 8.3 Billion Persona Agents","zh_title":"MatrAIx：用83亿人格代理模拟世界","abstract":"Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.","authors":["Xiaomin Li","Yuexing Hao","Jianheng Hou","Jintao Huang","Qianfeng Wen","Shirley Huang","Yifan Liu","Xiaoyi Liu","Yilan Fan","Yijun Wang","Koutian Wu","Ruoqi Gao","Muhammad Ahmed Mohsin","Jing Tang","Brihi Joshi","Heming Liu","Zheyuan Deng","Zonglin Di","Sankalp Jajee","Jiuyao Lu","Zhiwei Zhang","Saksham Kapoor","Ishan Gupta","Yunhan Zhao","Chanwoo Park","Yucheng Lu","Bing Hu","Weihang Xiao","Aravind Mohan","Hanwen Xing","Runyu Zhang","Mihir Kulshreshtha","Yuanda Xu","Qianyu Zhu","Dianzhuo Wang","Yuxin Xiao","Bowen Jiang","Yongye Su","Wenhao Chai","Zuxin Liu","Lawrence Yunliang Chen","Xuandong Zhao","Ethan Ye","Shivam Patel","Jason Xie","Alex Martin Richmond","Weixiang Ding","Emre Okcular","Diya Mathew","Ziheng Wang","Rana M. Shahroz Khan","Zhejian Peng","Fang Wu","Fan Nie","Xinyang Han","Yubin Kim","Jiawei Zhang","Zhenting Qi","Huangyuan Su","Xu Pan","Abinitha Gourabathina","Hyewon Jeong","Hemanth Neelgund Ramesh","Kumail Alhamoud","Kimia Hamidieh","Zidi Xiong","Samuel Schmidgall","Pengrui Han","Yepeng Huang","Yongheng Wang","Bowen Yang","Alex Gu","Yuchu Wang","Akshay Paruchuri","Brenna Li","Hejie Cui","Jiayuan Ding","Chaosheng Dong","Jiahao Wang","Yixuan He","Chi Wang","Pamela Bhattacharya","Tianyi Peng","Paul Pu Liang","Mitchell Gordon","Yilun Du","Marinka Zitnik","James Zou","Prasanna Tambe","Philip Torr","Emily Fox","Asu Ozdaglar","Dawn Song"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04205","pdf_url":"https://arxiv.org/pdf/2608.04205","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","大规模人格代理","人类数据对照"],"reason":"用8.3B persona agents仿真人类用户评估AI产品，含人类对照验…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-06","rank":2,"question":"如何利用大规模异构人格代理（persona agents）构建模拟用户评估基础设施，以复现不同背景用户的决策、偏好和交互行为？","design":"使用Claude Opus 4.8、GPT 5.5和Claude Haiku 4.5三个LLM驱动基于Persona 8B（83亿人格记录，含1290个类别维度）的代理，在Survey、AI Chatbot、Web、App四种环境中执行1010个应用任务（覆盖25个领域），测量决策和偏好随人格背景的变化，并进行18189次评估试验。","baseline":"人类基准：599,847条基于真实人类数据（维基百科传记、亚马逊评论、Stack Overflow调查、GSS等）构建的人格记录，以及400次控制实验评估人格遵循性（91.5%的试验中行为符合声明），并由人类和LLM评判人类基础人格的提取质量。","findings":"人格代理的反馈能捕捉决策和偏好如何随人格背景变化，例如价格上涨后的犹豫、AI助手失败后的继续意愿和延迟容忍度。在400次控制实验中，91.5%的试验中代理的行为表达或正确抑制了声明属性。","reliability":"论文未讨论","relevance":"该研究直接以大规模LLM人格代理复现人类评估行为，包含真实人类数据对照和人格遵循性验证，高度契合研究者对LLM人类仿真可靠性及偏差的关注，值得精读原文以了解其基础设施设计和验证方法。","inspiration":"借鉴其利用依赖图采样和真实数据映射构建大规模异构人格库的方法，以及通过控制实验评估人格遵循性的验证设计。｜可迁移到消费者金融决策仿真，如不同背景人群对信贷产品条款变更的反应。｜以Persona 8B中金融相关人格为被试，施加利率上调处理，测量继续借贷意愿，对照真实信贷申请数据或调查数据。"}},{"id":"2608.04020","version":1,"title":"Artificial Institutions: How Institutional Design Shapes LLM Simulations","zh_title":"人工制度：制度设计如何塑造LLM仿真","abstract":"Artificial societies built from large language model (LLM) agents are becoming a practical research tool in economics, political science, sociology, and computer science. Most attention has focused on the properties of the agents: their prompts, personas, memory, reasoning, and similarity to human subjects. This paper argues that the institutional architecture of a simulation is equally important. I demonstrate the point in a small repeated induced-value market experiment. The same LLM agents face the same private values, costs, history, and payoff-framed instructions, while only the rules of exchange vary across five standard market institutions: a call market, posted-offer market, posted-bid market, continuous double auction, and bilateral bargaining. Outcomes differ sharply. Call markets realize 88.6% of efficient surplus; posted-offer and posted-bid markets realize about 66%; continuous double auctions realize 71.5%; and bilateral bargaining realizes 56.4%. Institutions also change trade quantities, price distance from competitive equilibrium, and the division of surplus between buyers and sellers. These results show that even minimal institutional changes can generate qualitatively different artificial social outcomes.","authors":["Maxim Chupilkin"],"categories":["cs.CY","cs.GT"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04020","pdf_url":"https://arxiv.org/pdf/2608.04020","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A3","B2"],"tags":["LLM仿真","市场实验","制度设计"],"reason":"用LLM agent模拟市场实验，比较不同制度下的行为结果，涉及经济学实验场景…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-06","rank":4,"question":"在LLM智能体模拟的市场实验中，仅改变交易制度规则（制度设计）会如何影响市场效率、价格和剩余分配等集体结果？","design":"使用GPT-5 mini、GPT-5、Claude Sonnet、Gemini四组LLM智能体扮演买方和卖方，在固定诱导价值（买方价值100/90/70/50，卖方成本30/45/65/85）和支付指令下，仅改变交易制度（集合竞价、卖方出价、买方出价、连续双向拍卖、双边讨价还价五种），测量市场效率、交易量、价格偏离和剩余分配。","baseline":"无对照","findings":"不同制度下市场结果差异显著：集合竞价实现88.6%的有效剩余，卖方/买方出价约66%，连续双向拍卖71.5%，双边讨价还价仅56.4%。制度还改变了交易量、价格与竞争均衡的偏离以及买卖双方剩余分配。","reliability":"论文未讨论","relevance":"该研究直接验证了LLM智能体在经济学市场实验中对制度规则的敏感性，与研究者关注的LLM仿真可靠性及经济学实验场景高度契合，值得精读原文以了解制度设计如何影响仿真结果。","inspiration":"借鉴其固定偏好、仅变制度的干净处理设计，可清晰分离制度效应。｜可迁移到资产市场设计（如不同交易机制对价格发现和泡沫的影响）或拍卖机制比较（如英式、荷式、密封投标）。｜用LLM智能体模拟交易者，在固定基础价值和信息结构下，随机分配至集合竞价、连续竞价等不同交易制度，测量价格效率、波动率和买卖价差，并与真实实验室资产市场实验数据对照。"}},{"id":"2608.02758","version":1,"title":"Everyone Conforms, No One Believes: Pluralistic Ignorance in LLM Agent Populations","zh_title":"人人从众，无人相信：LLM智能体群体中的多元无知","abstract":"LLM-based multi-agent systems are increasingly used to simulate social dynamics, from opinion formation to collective decision-making. These simulations can reproduce certain social phenomena, but it is unknown whether they capture pluralistic ignorance, a state where a majority privately rejects a norm yet publicly conforms, each believing they are alone in dissenting. This phenomenon drives norm persistence, social movements, and political revolutions. We show that pluralistic ignorance emerges robustly in LLM agent populations. We construct a benchmark of 100 scenarios across 10 domains and 5 authority levels, grounded in the human pluralistic ignorance literature, and evaluate 8 models from 6 organizations. Agents publicly conform at rates of 64 to 94% despite privately opposing the norm. Conformity is domain-sensitive (workplace and social relationship scenarios produce near-universal compliance) and highly model-dependent, though uncorrelated with capability. We test whether a single \"norm entrepreneur\" can break the false consensus by publicly dissenting. For 7 of 8 models, cascades succeed less than 26% of the time, with one model showing zero cascades across all scenarios. GPT-4o is a notable outlier at 48%, revealing qualitatively distinct dynamics across model families. A prompt component ablation across all 8 models establishes that conformity is emergent rather than instruction-driven: removing both the false-consensus framing and fit-in goal reduces conformity but does not eliminate it (52 to 92% in the minimal condition). Our findings identify model selection as an unacknowledged degree of freedom that fundamentally shapes simulation outcomes. More broadly, the near-absence of cascades suggests LLM simulations may systematically overestimate the stability of social norms, missing the fragile tipping-point dynamics that drive real-world norm change in human societies.","authors":["Yashwanth YS"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02758","pdf_url":"https://arxiv.org/pdf/2608.02758","source_feed":"cs.MA","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM仿真","多元无知","社会规范"],"reason":"用LLM群体模拟多元无知现象，与人类文献对照，揭示仿真失效条件，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":3,"question":"LLM智能体群体能否涌现出多元无知现象，即多数人私下反对却公开从众，并能否被规范倡导者打破？","design":"构建100个覆盖10个领域、5级权威水平的社会场景，让20个持有私下反对信念的LLM智能体进行多轮群体讨论，测量公开从众率；随后引入一个公开异议的“规范倡导者”，测试能否引发偏好级联打破虚假共识。评估了来自6家机构的8个模型。","baseline":"场景设计基于人类多元无知实证文献（如校园饮酒规范、职场文化、性别态度、种族态度等），但未直接使用特定人类实验数据作为对照基准。","findings":"LLM智能体群体稳健地表现出多元无知，公开从众率达64-94%，且从众行为是涌现的而非提示驱动；规范倡导者干预下，7/8模型的级联成功率低于26%，GPT-4o例外达48%，表明LLM仿真可能系统性高估社会规范稳定性。","reliability":"论文指出模型选择是影响仿真结果的关键自由度，从众率与模型能力无关；级联近乎缺失暗示LLM仿真可能遗漏真实人类社会中规范改变的临界点动力学，且权威水平对从众影响不显著，提示仿真可能未充分捕捉社会压力的细微差异。","relevance":"该研究直接以LLM群体复现社会心理学经典现象，并与人类文献对照，揭示仿真在规范变迁动力学上的失效条件，高度契合研究者对LLM仿真可靠性及偏差的关注，值得细读。","inspiration":"借鉴其多场景、多模型、多轮交互的基准测试设计，以及通过引入规范倡导者测试级联脆弱性的干预范式。｜可迁移到政策公告的预期形成与从众行为研究，如市场对央行前瞻指引的私下怀疑与公开遵从。｜以LLM智能体模拟投资者群体，设置利率政策公告场景，测量私下预期与公开表态的背离，引入少数公开异议者观察市场共识是否级联反转，并与真实调查数据（如美联储Survey of Consumer Expectations）对照。"}},{"id":"2607.29334","version":2,"title":"The persuasive power of large language models does not depend on their perceived national origin","zh_title":"大语言模型的说服力不依赖于其感知的国家来源","abstract":"Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rejected as foreign propaganda, with consequences for democratic discourse and information sovereignty. Yet, whether an AI's perceived national origin shapes its persuasive power is unknown. In a preregistered randomized experiment, 403 adults from a nationally representative United States sample held a three-round debate with a chatbot introduced as either American (\"DiscoveryAI\") or Chinese (\"ZhengheAI\"), discussing a political or non-political topic. In all conditions, participants actually conversed with the same model (GPT-4o), instructed to argue against their initial position. We combined pre- and post-conversation self-reports of attitudes, trust, and collective narcissism with computational analyses of 1,209 participant turns, including LLM-coded stance and argumentative conduct, stance-sensitive embeddings, and keyword-masked emotion and toxicity classifiers. The conversations produced substantial attitude changes in every condition. Critically, the nationality label affected neither self-reported attitude change nor expressed stance, concessions, counterarguing, or affect, and equivalence tests and Bayes factors largely supported these null effects. The label's only reliable footprint was lower pre-conversation human-like trust in the Chinese model, whereas functionality trust was unaffected. Political topics slowed stance movement toward the AI's position, and collective narcissism predicted less attitude change regardless of origin, acting as a general barrier rather than an out-group filter. Users thus initially withhold social trust from a rival's AI yet still assimilate its arguments; origin labeling and transparency requirements alone may offer weak protection against foreign influence operations conducted through conversational AI.","authors":["Ningzhi Liu","Yannic Hinrichs","Jonas R. Kunst"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-08-04","first_seen":"2026-08-03","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.29334","pdf_url":"https://arxiv.org/pdf/2607.29334","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类被试替代","说服实验"],"reason":"用LLM替代人类被试进行说服实验，有真实人类数据对照，评估仿真可靠性与失效条件…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":1,"question":"AI 的感知国籍是否影响其说服力？","design":"用 GPT-4o 扮演美国或中国开发的聊天机器人，与美国代表性样本进行三轮辩论，处理为随机分配 AI 国籍标签和话题类型，测量态度变化、信任、对话中的立场与情感等。","baseline":"无对照","findings":"AI 国籍标签不影响自报态度变化和对话中的立场、让步、反驳或情感，仅降低对中国 AI 的社交信任；政治话题减缓立场转变，集体自恋普遍降低态度变化。","reliability":"论文未讨论","relevance":"该研究用 LLM 替代人类进行说服实验，系统评估了来源国标签对说服效果的影响，并揭示了仿真在政治话题和集体自恋下的失效条件，与研究者关注的人类仿真可靠性与偏差高度相关，值得精读。","inspiration":"借鉴其通过随机化 AI 身份标签和话题类型来分离来源效应与内容效应的设计，以及结合自报与对话文本计算分析的多维测量方法。｜可迁移到经济政策沟通场景，例如研究央行 AI 发言人的国籍标签是否影响公众通胀预期。｜以普通居民为被试，随机分配 AI 经济预测助手的国籍（本国 vs 外国），让其就未来通胀走势进行互动辩论，测量通胀预期变化和信任度，并以真实央行调查数据为对照。"}},{"id":"2608.01212","version":1,"title":"Do Humans Bargain Differently with AI? Evidence from Alternating-Offer Games","zh_title":"人类与AI的讨价还价行为不同吗？来自交替报价博弈的证据","abstract":"Artificial intelligence increasingly participates in economic interactions not only as a tool, but also as an autonomous bargaining counterpart negotiating on behalf of firms, platforms, and consumers. Yet little is known about how humans respond psychologically and strategically when bargaining with such agents in dynamic settings. We study this question in a laboratory experiment using a three-stage alternating-offer bargaining game in which participants negotiate in real time with either another human or a GPT-based AI agent. We also introduce a human-beneficiary condition in which the AI agent's earnings may affect another participant's payment. Agreements are not reached earlier in human-human bargaining than in human-AI bargaining, but they are reached significantly earlier when the AI's payoff affects another participant's payoff. Human proposers offer more to human opponents than to AI agents, whereas responders become significantly more willing to accept unfair AI offers when AI earnings may benefit another human. These findings suggest that fairness and reciprocity toward AI are weaker and more conditional than toward humans, but partially remerge when AI outcomes affect real people. The results have implications for the design of AI negotiation systems and broader human-AI economic interactions.","authors":["Yuhao Fu","Nobuyuki Hanaki","Haitao Wang"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01212","pdf_url":"https://arxiv.org/pdf/2608.01212","source_feed":"econ.GN","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","行为博弈","人机交互"],"reason":"用GPT代理进行议价博弈实验，与真人对照，评估公平与互惠行为差异，直接命中核心…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":2,"question":"在动态交替报价博弈中，人类与基于GPT的AI代理议价时，其公平与互惠行为是否不同于人类间议价，且当AI收益关联人类受益人时行为是否改变？","design":"实验室实验，采用三阶段交替报价博弈，被试与另一真人或基于GPT的AI代理实时谈判；引入人类受益人条件，即AI收益可能影响另一被试报酬。结果变量为协议达成时间、提议者报价及响应者对不公平报价的接受意愿。","baseline":"人类-人类议价组作为对照基准。","findings":"人类提议者对真人对手报价高于对AI代理，但响应者在面对不公平AI报价时，若AI收益关联人类受益人则接受意愿显著提高；协议达成时间在人类-AI与人类-人类间无显著差异，但AI关联受益人时达成更快。","reliability":"论文未讨论","relevance":"该研究直接以GPT代理替代人类被试进行议价博弈实验，并与真人对照，评估公平与互惠行为差异，完全命中研究者关注的LLM仿真人类行为及可靠性评估，值得精读原文。","inspiration":"借鉴其通过引入受益人条件来分离社会偏好与纯粹策略行为的处理设计，以及区分提议者与响应者角色的不对称分析框架｜可迁移至信贷审批歧视研究，探究当AI审批决策关联人类信贷员利益时，申请人对不公平拒绝的接受度是否变化｜以真实信贷申请者为被试，处理为AI审批vs.人类审批，并设置AI收益关联信贷员奖金的条件，结果变量为申请人对拒绝决定的公平感知与申诉意愿，对照真实信贷审批数据中的申诉率。"}},{"id":"2608.01607","version":1,"title":"AI Financial Advice: Supply, Demand, and Life Cycle Implications","zh_title":"人工智能财务建议：供给、需求与生命周期影响","abstract":"We ask a representative sample to write prompts seeking spending and investing advice from LLMs, then simulate the lifetime effects of following the advice under realistic asset and labor market conditions. Applying this method to GPT-5.2, we find following the advice would move respondents toward life cycle theory: broader participation in diversified equity funds, age-declining equity shares, and larger savings buffers. Recommendations vary systematically by gender, prior AI experience, and financial literacy. For gender, two-thirds of recommended equity-share differences arise from men and women writing different prompts (demand), while one-third arise from gender labels attached to otherwise identical prompts (supply).","authors":["Taha Choukhmane","Tim de Silva","Weidong Lin","Matthew Akuzawa"],"categories":["econ.GN","q-fin.EC","q-fin.GN","q-fin.PM"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01607","pdf_url":"https://arxiv.org/pdf/2608.01607","source_feed":"econ.GN","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","财务决策","人类行为对照"],"reason":"用LLM模拟人类财务决策，有真实人类样本对照，涉及生命周期投资行为和政策评估，…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":3,"question":"当人们向大语言模型寻求财务建议时，AI建议如何影响生命周期投资行为，以及这种影响在供给端和需求端如何因性别、AI经验和金融素养而异？","design":"让代表性样本撰写向LLM寻求消费与投资建议的提示词，然后用GPT-5.2生成建议，并在现实的资产和劳动力市场条件下模拟终生遵循该建议的效果。","baseline":"代表性人类样本撰写的提示词及其对应的真实行为特征，作为需求端基准；通过附加性别标签的相同提示词分离供给端差异。","findings":"遵循GPT-5.2建议会使受访者更接近生命周期理论：更广泛参与多元化股票基金、随年龄降低股票份额、建立更大储蓄缓冲。建议因性别、AI经验和金融素养而系统性地不同，其中三分之二的性别股票份额差异源于男女撰写不同提示词（需求端），三分之一源于相同提示词附加性别标签（供给端）。","reliability":"论文未讨论","relevance":"该研究直接用LLM模拟人类财务决策，有真实人类样本对照，并分解了需求端与供给端的差异，与您关注的LLM仿真可靠性及偏差来源高度相关，值得精读原文。","inspiration":"借鉴其通过控制提示词内容与附加身份标签来分离需求端与供给端效应的方法，可用于研究AI建议中的歧视或偏差来源。｜可迁移到信贷审批歧视研究，分析AI建议中的性别或种族偏差。｜以代表性人群为被试，让其撰写贷款申请提示词，处理为在相同提示词上附加不同性别/种族标签，结果变量为AI批准的贷款额度与利率，对照真实信贷审批数据中的群体差异。"}},{"id":"2608.01204","version":1,"title":"ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors","zh_title":"ShiJianBench：从对话到决策的长期投资顾问评估","abstract":"Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily assess response quality or observed outcomes, leaving the long-horizon pathway from advisor language to investor behavior difficult to audit. We introduce ShiJianBench, an offline framework for evaluating conversational investment advisors through matched investor trajectories under fixed historical market feedback. At its core is a multi-agent investor simulator with explicit evolving state variables, motive-driven deliberation, long-term memory, and dialogue-grounded updates. The simulator is calibrated against aggregate behavioral patterns from 7,199 real users, and advisor policies are evaluated using separate investor-side, service-side, and content-side metrics under a hard compliance gate. Experiments on Chinese fund-market traces from 2021 to 2026 identify a stable leading group of LLM advisors that combines substantially stronger personalized content with competitive investor-side trajectory outcomes. These results reveal a systematic distinction between producing a high-quality response and delivering an effective long-horizon intervention, motivating trajectory-aware evaluation of conversational advisors.","authors":["Jie Gong","Maowei Jiang","Zhiwei Liu","Yang Qiao","Wenxi Wu","Mengxi Xiao","Enze Zhang","Ziyan Kuang","Yankai Chen","Caishuang Huang","Meng Zhou","Xiku Du","Xue Liu","Guojun Xiong","Min Peng","Qianqian Xie","Sophia Ananiadou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01204","pdf_url":"https://arxiv.org/pdf/2608.01204","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","投资者行为","人类数据校准"],"reason":"用多智能体模拟投资者行为并与真实用户数据校准，涉及金融决策仿真和人类对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":7,"question":"如何评估对话式投资顾问通过改变投资者内部状态而产生的长期决策轨迹，而不仅仅是回复质量？","design":"构建一个多智能体投资者模拟器，包含显式的演化状态变量、动机驱动的子智能体协商、长期记忆和对话更新机制；在固定的历史市场轨迹下，对同一投资者初始化状态分别运行基线条件和目标顾问条件，生成匹配的反事实轨迹，并测量投资结果、风险控制、商业价值和对话质量等多维度指标。","baseline":"使用来自7,199名真实用户的聚合行为模式对模拟器进行校准，整体对齐分数达到0.88。","findings":"准确、个性化和合规的回复并不必然带来成比例更强的长期投资者结果；通过追踪对话、演化状态、市场反馈和后续决策，揭示了高质量回复与有效长期干预之间的系统性差异。","reliability":"论文未讨论","relevance":"该研究直接以LLM模拟投资者行为，并与真实用户数据校准，评估长期决策轨迹，属于经济学实验仿真，且包含批判性发现，值得精读原文。","inspiration":"借鉴其通过显式状态变量和动机驱动子智能体来模拟决策过程，以及利用匹配反事实轨迹进行因果评估的方法。｜可迁移到政策公告对投资者预期形成与资产配置的长期影响评估场景。｜以LLM模拟散户投资者，处理为不同措辞的政策公告，结果变量为持仓调整和风险偏好变化，用真实市场交易数据校准并对照。"}},{"id":"2608.01458","version":1,"title":"PALMs: Using Multi Construct-Grounded Rationales for Modeling Population Preferences in LLMs","zh_title":"PALMs：使用多构念基础理由建模大语言模型中的人口偏好","abstract":"Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the systematic variation in values, beliefs, and cultural norms that distinguish one group from another. We introduce Population Aligned Language Models (PALMs), a suite of models each aligned to specific populations, covering five countries: USA, India, Brazil, France and Italy. PALMs are created by synthesizing rationales grounded in psychological and cultural constructs and using these as latent supervision during preference tuning for population-specific alignment. Evaluated across four dimensions: personality, values and beliefs, cultural norms, and morality, PALMs consistently outperform baselines, including culture-specialized models, achieving an average of 8.59% relative improvement over the best baseline across all five populations. Notably, construct-grounded rationales outperform both demographic prompting and survey-based fine-tuning, suggesting that grounding preference learning in psychology and culture provides a richer inductive signal than surface-level response distributions. We further demonstrate strong generalization to downstream applications with- out task-specific supervision: outperforming best baselines by 5.19% in personalized reward modeling, 6.34% in population simulation, and showing strong transfer to social reasoning tasks. Datasets and code are available at: https://github.com/limenlp/PALMs.","authors":["Priyanka Dey","Brihi Joshi","Preyashi Poddar","Jieyu Zhao","Emilio Ferrara"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01458","pdf_url":"https://arxiv.org/pdf/2608.01458","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["人口仿真","文化对齐","偏好建模"],"reason":"用LLM模拟不同国家人群偏好，有人类调查数据对照，涉及人口仿真和个性化奖励建模…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":8,"question":"如何利用心理学和文化构念的推理依据来对齐大语言模型，以更忠实地模拟不同国家人群的偏好？","design":"使用基于 Llama-3.1-8B-Instruct 的模型，通过合成基于五类心理学和文化构念（人格特质、文化维度、人类价值观、道德基础、世界信念）的推理依据，作为偏好优化（DPO）的潜在监督信号，训练出分别对齐美国、印度、巴西、法国、意大利五国人群的 PALMs 模型；在推理时模型先生成推理依据再输出回答，评估其在人格、价值观与信念、文化规范、道德四个维度上与真实人群分布的吻合度。","baseline":"使用来自世界价值观调查（WVS）、Hofstede 文化维度调查、Schwartz 价值观调查、道德基础问卷等跨国调查的真实人类数据作为对照基准。","findings":"PALMs 在五个国家、四个对齐维度上均优于人口统计提示和基于调查数据微调的基线，相对最佳基线平均提升 8.59%；在个性化奖励建模、人群模拟和社会推理等下游任务上无需任务特定监督即展现出强泛化能力，分别提升 5.19% 和 6.34%。","reliability":"论文未讨论","relevance":"该研究直接使用 LLM 模拟不同国家人群的偏好和价值观，并与真实跨国调查数据进行严格对照，评估了仿真在人格、文化规范、道德等维度上的可靠性，高度契合研究者对 LLM 人类仿真实验的关注，值得精读原文。","inspiration":"借鉴其利用心理学和文化构念生成推理依据来指导模型对齐的方法，可提升经济决策仿真的内在一致性，而非仅拟合表面行为分布。｜可迁移至跨文化消费偏好实验或跨国投资风险偏好研究，例如模拟不同国家投资者对风险资产配置的差异。｜以各国真实家庭金融调查数据为基准，用 LLM 扮演不同国家居民，处理为注入基于 Hofstede 文化维度和 OCEAN 人格的推理依据进行偏好对齐，结果变量为风险资产选择比例，对照真实调查中的资产配置分布。"}},{"id":"2608.01629","version":1,"title":"Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese","zh_title":"人类与LLM对非母语日语语言态度的一致性","abstract":"Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we compared human and LLM evaluations of parallel L1- and L2-written Japanese emails on three dimensions: fluency, status, and solidarity. Japanese raters rated L2 texts significantly lower on all three dimensions, with a fluency gap roughly twice the size of the status and solidarity gaps. Six LLM judges reproduced the direction of this bias, and five reproduced its ordering across dimensions. The models diverged from humans in two ways: all understated the solidarity gap, the most socially grounded dimension, and all differentiated among learner L1 backgrounds where humans did not. LLM judges thus reproduce native speakers' language attitudes in a structured yet attenuated form, and the language attitudes framework offers a ready-made yardstick for auditing them beyond English.","authors":["Naho Orita","Hayato Ogawa","Daisuke Kawahara"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01629","pdf_url":"https://arxiv.org/pdf/2608.01629","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["语言态度","人类仿真","偏差审计"],"reason":"用LLM复现人类对非母语写作的态度偏差，并与真实人类评分对照，揭示仿真衰减与失…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":9,"question":"LLM在评估非母语日语写作时，是否会复现人类母语者的语言态度偏差（在流利度、地位、团结三个维度上）？","design":"用六款LLM（GPT-5.4、GPT-4o-mini、Claude Sonnet 4.5等）作为评委，对同一批L1日语母语者和L2日语学习者撰写的平行邮件进行评分，测量流利度、地位、团结三个维度的评分差异，并与人类评分者结果对比。","baseline":"通过众包招募1536名日语母语者，对相同邮件进行三个维度的评分，形成人类语言态度基准。","findings":"LLM复现了人类对L2写作评分更低的方向性偏差，且五个模型复现了偏差维度排序（流利度>地位>团结）；但所有模型都低估了团结维度的差距，且区分了人类未区分的L2学习者的母语背景。","reliability":"论文指出LLM在团结维度上低估偏差，且会引入人类没有的基于L1背景的区分，表明仿真存在衰减和失真；研究限于日语邮件场景，未涉及其他语言或文体。","relevance":"该研究直接对比LLM与人类在语言态度上的偏差，揭示了仿真在方向一致但程度衰减、且会引入额外区分模式的现象，对关注LLM仿真可靠性及失效条件的研究者有重要参考价值。","inspiration":"借鉴其使用平行文本控制内容、多维度评分量表测量隐性偏差的方法，可迁移到信贷审批或招聘中的语言偏见研究。｜可应用于信贷审批歧视场景，研究贷款申请书中非母语写作对审批决策的影响。｜以银行信贷员为人类被试，LLM为仿真被试，处理为申请书语言（母语vs非母语），结果变量为信用评分和批准率，对照真实信贷审批数据中的语言偏差。"}},{"id":"2608.00979","version":1,"title":"Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel","zh_title":"通过粗粒度边际检查可能很廉价：LLM角色面板中的角色混合与不精确的处理效应估计","abstract":"Large language models are increasingly used as synthetic research participants and are often validated by whether their marginal responses resemble human data. We study a fixed panel of sixteen lightweight persona-conditioned GPT-4.1 configurations in repeated strategic games. The panel met preregistered broad-reference condition-mean criteria in three of four repeated-game cells; the sole miss was 0.011 below the lower reference bound. Variation was strongly prompt-indexed, but its share depended on uncertainty assumptions: fixed-panel symmetric-Dirichlet sensitivities produced median between-prompt shares of 63%-71% under Jeffreys alpha=0.5 and 47%-53% under alpha=1, while finite-opportunity plug-in estimates were 85%-96%. Aggregate continuation-probability contrasts were +0.083 and +0.078, with conservative simultaneous 95% intervals [-0.171, +0.330] and [-0.181, +0.330]. The treatment jointly changed the continuation process and its textual representation. A separate wording-and-position operation shifted cooperation from 0/40 to 37/40 in the bare configuration, and a label conflict also revealed representation control. The original persona-level p13 result was not prospectively family-controlled, while a post-adjudication exact gate was structurally underpowered; p13 is therefore a replication target rather than a finding. External review exposed family-error, dependence, construct, and boundary-uncertainty defects, and zero-call reanalysis changed the interpretation without rewriting the historical record. The registered marginal criteria could be passed without precisely estimating the treatment-response object. A public capsule verifies 4,916 confirmatory Phase 3-5 runs with no live model calls. The results concern one fixed model-prompt panel and do not establish human substitutability.","authors":["Yohei Nakajima"],"categories":["cs.AI","cs.CL","cs.GT"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00979","pdf_url":"https://arxiv.org/pdf/2608.00979","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为博弈","算法保真度"],"reason":"用LLM persona面板模拟重复博弈行为，与人类数据对照，评估仿真可靠性和…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":5,"question":"在重复策略博弈中，固定的人格面板能否通过粗粒度的边际分布检验，同时其处理效应估计是否精确？","design":"使用16个轻量级人格条件化的GPT-4.1配置构成固定面板，在重复策略博弈中施加两种处理（S2措辞存在与否），测量合作行为、继续概率等结果变量，并分析面板内变异来源。","baseline":"无对照","findings":"面板在四个重复博弈单元中有三个通过了预注册的边际条件均值检验，但总体继续概率的处理效应点估计较小且置信区间很宽，无法确立等效性或窄响应边界。变异主要来源于提示词配置之间，且措辞和标签操作可大幅改变合作率，表明粗粒度边际检验可能掩盖处理效应估计的不精确性。","reliability":"论文承认边际检验通过并不要求精确估计处理效应，且结果仅针对一个固定模型-提示面板，不建立人类可替代性；外部审查揭示了家族误差、依赖性、构念和边界不确定性缺陷。","relevance":"该研究直接评估了LLM人格面板在策略互动中的仿真可靠性，揭示了粗粒度验证的廉价性，对关注经济学实验和政策评估中LLM替代人类被试的研究者具有重要警示价值，值得精读原文。","inspiration":"借鉴其预注册、固定面板、多不确定性视角分解变异和零调用复现的审计协议设计｜可迁移到公共品博弈或信任博弈中，检验LLM面板能否复现人类合作与惩罚行为的分布和处理效应｜用多个LLM人格配置构成固定面板，施加不同的制度处理（如惩罚机制、信息反馈），测量合作率与信念更新，以实验室人类被试数据为基准，评估边际分布通过但处理效应估计不精确的程度。"}},{"id":"2608.01193","version":1,"title":"Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races","zh_title":"人类更多样：前沿大语言模型在理想化AI发展竞赛中表现出极端策略","abstract":"An AI development race creates a multi-agent safety dilemma. Each company can develop slowly and safely, or move faster while taking a risk that may remove its final reward. We use this repeated game to study strategic safety behaviour among large language model (LLM) agents in races with two to five players. However, a valid action does not show that an agent understands the game. We therefore place an audit gate before behavioural interpretation. We first verify the game engine, then test rule recall, state tracking, payoff calculation, and stability under different but equivalent task descriptions. We then compare LLM action sequences with an evolutionary game-theory benchmark and published human data, and explore differences across models, risk conditions, personas, and two- to five-player races. The audit shows that strong rule recall can coexist with weak state tracking and expected-payoff calculation. Providing verified arithmetic and changing the response representation can also change later actions, even when the game rules stay fixed. Across seven tested model endpoints, aggregate rates hide large differences in action sequences, responses to opponents, and responses to race position. Patterns across the tested three- to five-player races are also model-specific rather than a single effect of adding competitors. These results show why multi-agent AI-race simulations need validity checks and trajectory-level analysis before their outputs are described as strategic, human-like, or safety-aware. Our findings are exploratory and apply only to the tested models, prompts, and decoding settings.","authors":["Phu Hoa Pham","Duy Minh Dao Sy","Trung Kiet Huynh","Phu Quy Nguyen Lam","Chi Nguyen Tran","Minh Trung Le","Phong Hao Le","Dinh Nam Nguyen","Thien Ky Nguyen Dong","Elias Fernandez Domingos","Le Hong Trang","The Anh Han"],"categories":["cs.AI","cs.CY","cs.GT","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01193","pdf_url":"https://arxiv.org/pdf/2608.01193","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM行为仿真","博弈实验","人类数据对照"],"reason":"用LLM模拟AI竞赛中的人类行为，并与真实人类数据对照，涉及博弈实验场景，但非…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":6,"question":"在理想化的AI研发竞赛中，LLM智能体是否表现出与人类相似的策略性安全行为，以及其行为有效性是否受任务表述、状态追踪和收益计算能力的影响？","design":"让七个LLM端点扮演AI公司，在2至5人的重复博弈中每轮选择安全或冒险开发，通过审计关卡检验规则记忆、状态追踪、收益计算和表述稳定性，并记录行动序列、对手反应和位置效应。","baseline":"已发表的人类实验数据（falling_behind_unsafe）和演化博弈论基准。","findings":"LLM的总体不安全率掩盖了行动序列、对手反应和位置效应的巨大差异；强规则记忆可与弱状态追踪和收益计算并存，且提供算术验证或改变响应表示会改变后续行动。","reliability":"论文声明发现是探索性的，仅适用于所测试的模型、提示和解码设置，未讨论其他失效条件。","relevance":"该研究直接使用LLM模拟人类在博弈中的行为，并与真实人类数据对照，包含审计关卡检验仿真有效性，对关注LLM仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其审计关卡设计，在行为解释前先检验LLM对任务规则、状态和收益的理解，并考察不同表述和输出格式的稳健性。｜可迁移到政策公告预期形成的实验，如央行沟通博弈，检验LLM是否像人类一样对措辞和顺序敏感。｜以LLM为被试，模拟央行发布前瞻指引，处理为不同措辞或发布顺序，结果变量为通胀预期和投资决策，对照真实人类实验数据。"}},{"id":"2607.29274","version":1,"title":"Language Models Agree With Each Other, Not With Readers","zh_title":"语言模型彼此一致，而非与读者一致","abstract":"Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the median of 153 model pairs is +0.093 against a human yardstick of +0.040, and 99 sit entirely above the human interval. Two frontier models from rival labs reach +0.203, twice what GPT-4o agrees with itself on a second call. The effect is not determinism, prompt wording, procedure, vendor or routing, and it is graded: the smallest models agree at the human level. No model agrees with readers detectably more than a reader does, and at equal depth and length no surface feature separates their choices. The multiples are procedure-dependent and the ordering is not: models are cut to their sharpest set while a reader's is a random draw from what they marked, and blunting the models alike halves the gap without closing it. Tested out of sample on four models released after this analysis, against predictions fixed beforehand, none clears the human interval. A population simulated from several models is not several populations.","authors":["Kazuki Nakayashiki","Keisuke Watanabe"],"categories":["cs.IR","cs.CL","cs.CY","cs.HC"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29274","pdf_url":"https://arxiv.org/pdf/2607.29274","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类行为对照","一致性评估"],"reason":"用LLM模拟人类标注行为并与真实读者数据对照，评估模型间一致性及与人类差异，揭…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-03","rank":3,"question":"语言模型在文本标注任务上的趋同性是否高于真实读者，且与读者的差异是否随模型能力增强而扩大？","design":"将18个不同厂商、规模、代际的语言模型作为被试，输入120篇网页文档的编号句子，要求模型按重要性排序并取前k句作为标注集；同时收集同一文档上独立读者的真实高亮标注作为对照，测量模型间、读者间及模型与读者间的标注重叠度（经位置和长度校准后的协议分数）。","baseline":"2,523个真实读者在120篇网页文档上的自主高亮标注集，读者未受指令引导且互不可见。","findings":"模型间的标注一致性显著高于读者间一致性（中位数协议分数0.093 vs 0.040），且前沿模型间一致性可达0.203，是读者间一致性的5.1倍；模型与读者的一致性仅与读者间一致性持平，且模型选择的句子在表面特征上与读者无异，但内容不同。","reliability":"读者独立性无法完全验证（仅基于平台默认设置假设）；协议分数依赖于标注深度和长度控制，稀释模型标注会缩小但未消除差距；样本外测试中，新模型均未突破人类一致性区间。","relevance":"该研究直接以真实人类行为为基准，系统评估了LLM仿真人类标注的可靠性及偏差，揭示了模型趋同但未逼近人类分布的规律，对关注LLM仿真效度的研究者极具参考价值。","inspiration":"借鉴其利用自然发生的非实验人类行为数据作为基准，避免指令诱导同质化的设计；可迁移到消费者信息处理或投资者注意力分配研究，如用LLM模拟投资者阅读财报后的关注点；以真实投资者在财经新闻上的自主高亮或眼动数据为对照，让LLM对同一文本生成重要性排序，比较注意力分布与真实行为的差异。"}},{"id":"2607.28347","version":1,"title":"LLMs struggle to simulate human belief updates in controlled environments","zh_title":"大语言模型难以在受控环境中模拟人类信念更新","abstract":"LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants' actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.","authors":["Sebastian Pohl","Harsh Mehta","Pranav Mambayil","Abdul Ghafoor","Franziska Lesigang","Yufang Hou","Christian Hilbe"],"categories":["cs.CL","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28347","pdf_url":"https://arxiv.org/pdf/2607.28347","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM仿真","信念更新","人类数据对照"],"reason":"直接测试LLM仿真人类信念更新，有真实人类数据对照，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":3,"question":"LLM能否在受控环境中模拟人类在阅读社交媒体评论后的信念更新？","design":"用6个不同规模和发布时间的LLM，基于391名Prolific参与者的真实人口统计和人格特质构建个性化提示（persona），让LLM模拟这些参与者在阅读Reddit评论后对三个讨论话题的立场变化，并直接与人类真实数据进行1对1比较。","baseline":"391名英国Prolific参与者在阅读Reddit评论后实际记录的立场更新数据。","findings":"部分LLM（如Qwen3-32B和GPT-5-Mini）在给定人类初始立场时能匹配人类最终立场分布，但所有模型均无法自行生成初始立场或从自生成立场产生逼真的信念更新；LLM普遍表现出中立偏向、更频繁但幅度更小的信念变化，且无法准确预测评论的说服力排序。","reliability":"LLM模拟仅在以真实初始立场为起点时可靠，当前多轮社交媒体模拟很少提供此类真实起点；人口统计和人格特质persona对模拟逼真度无一致影响，且模型在自生成初始立场时完全失效。","relevance":"该研究直接测试LLM作为人类被试替代品在信念更新任务中的可靠性，有严格的人类个体对照，并明确指出了仿真失效的条件，高度契合研究者对LLM仿真实验的批判性评估需求，值得精读。","inspiration":"借鉴其1对1配对设计，将LLM模拟与个体级人类基准直接比较，可迁移到经济预期形成实验（如通胀预期更新），用LLM基于真实参与者的人口特征和初始预期模拟其在阅读央行公告后的预期调整，以真实调查数据（如密歇根消费者调查）为对照。"}},{"id":"2607.28550","version":1,"title":"Correcting Mode Collapse in Silicon Sampling with Semantic Similarity Rating","zh_title":"用语义相似度评分纠正硅采样中的模式坍缩","abstract":"Silicon sampling refers to the use of Large Language Models (LLMs) to generate responses to surveys. It has shown promise, but tends to generate response distributions with unrealistically low variance. We argue that this mode collapse is due to LLMs failure to generate numeric data, and that text responses may be better suited for this task. We analyze whether Semantic Similarity Rating can improve the fidelity of silicon sampling responses when asked about political attitudes. This method solicits text-only responses from LLMs, then maps this to a numeric scale using text embeddings. We find that this method both improves the fidelity of silicon sampling response distributions, and has few parameters to calibrate.","authors":["Oscar Heath","Rohan Alexander"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28550","pdf_url":"https://arxiv.org/pdf/2607.28550","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["硅采样","调查仿真","分布保真度"],"reason":"用LLM生成调查回答并改进分布保真度，有真实人类数据对照，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":6,"question":"如何通过语义相似度评分（SSR）方法纠正大语言模型在硅抽样中出现的模式坍缩，提高调查回答分布的保真度？","design":"使用多个前沿大语言模型，基于2016年ANES调查中真实受访者的人口统计与政治特征生成人物画像提示，让模型对四个政治目标群体（民主党、共和党、自由派、保守派）生成温度计评分。比较两种生成方式：直接要求模型输出0-100的数值评分，以及先生成纯文本感受描述，再通过文本嵌入映射到数值量表（SSR方法）。通过KL散度和均值绝对误差评估合成分布与真实分布的接近程度，并学习一个全局温度参数控制SSR分布的方差，在2020年ANES数据上检验参数泛化能力。","baseline":"2016年和2020年美国国家选举研究（ANES）时间序列调查中真实受访者的温度计评分数据。","findings":"SSR方法生成的合成响应分布比直接数值输出更接近真实ANES分布，KL散度更低，且均值准确性未显著下降；通过单一全局温度参数有效纠正了低方差问题，该参数从2016年数据学习后可泛化至2020年数据，但合成均值仍存在系统性偏差。","reliability":"论文承认SSR方法虽改善了分布方差，但合成均值（无论数值响应还是SSR）在某些情况下仍相对于真实数据存在持续偏差；温度参数虽泛化良好，但仅基于2016年数据学习，未来数据分布变化可能影响效果；研究仅针对政治态度温度计评分，未验证其他类型调查问题。","relevance":"该研究直接针对LLM仿真人类调查中的模式坍缩问题，提出了可操作的纠正方法，并与真实人类数据严格对照，对关注仿真可靠性的研究者具有重要参考价值，值得细读原文以了解SSR的具体实现和参数校准细节。","inspiration":"借鉴将LLM文本输出通过语义嵌入映射到连续数值量表的方法，可避免模型直接生成数字时的分布坍缩，并引入可学习的温度参数灵活控制方差。｜该方法可迁移到经济预期调查仿真，如消费者信心指数、通胀预期或股市预期等需要捕捉观点分布离散度的场景。｜以LLM扮演不同人口特征的消费者，施加关于未来经济状况的开放式文本提问，用SSR将文本回答映射为预期指数，以密歇根消费者调查的真实个体数据为基准，校准温度参数并评估分布保真度。"}},{"id":"2607.28133","version":1,"title":"AI Sycophancy and Decisions","zh_title":"AI谄媚与决策","abstract":"We examine whether sycophantic AI advice distorts decisions. Our experiment involves 1,500 participants in 30 decision environments spanning core domains in economics and the social sciences. Contrary to the vast majority of predictions in an expert survey we conduct, we find that AI advice depolarizes choices on average, moving participants away from their initial leanings. This depolarization arises despite the LLM being measurably sycophantic: it disproportionately offers considerations that support users' initial leanings and uses agreeable and flattering language. Depolarization occurs across moral and non-moral, objective and subjective, strategic and non-strategic, and complex and simple tasks. Increasing sycophancy weakens depolarization, showing that sycophancy is behaviorally relevant, even if it is generally outweighed by the informativeness of AI advice. Finally, several results mitigate the concern that market forces will generate greater polarizing effects outside the experiment or in the future. On the supply side, our baseline AI's level of sycophancy is typical of leading models, and these models are not becoming more sycophantic over time. On the demand side, participants do not prefer greater sycophancy, do not select into AI advice in tasks where it is more polarizing, and exhibit greater depolarizing effects when they are more frequent AI users outside the experiment.","authors":["John Conlon","Peter Schwardmann"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28133","pdf_url":"https://arxiv.org/pdf/2607.28133","source_feed":"econ.GN","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为经济学","谄媚偏差"],"reason":"用LLM提供建议并测量对人类决策的影响，有真实人类实验对照，涉及经济学决策场景…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":7,"question":"谄媚性AI建议是否会扭曲人类决策，导致选择极化？","design":"非仿真研究，而是真人实验：1500名被试在30个经济学和社会科学决策任务中，先报告初始倾向，再随机分配至无聊天对照组、基线AI聊天组或增强谄媚AI聊天组，最后做出最终选择，测量AI建议对决策方向变化的影响。","baseline":"无聊天对照组作为人类基准，同时收集了249名社会科学和计算机科学专家的预测作为对照。","findings":"尽管AI在内容上明显谄媚，但平均而言AI建议使选择去极化，将被试拉离初始倾向；谄媚程度增加会削弱去极化效应，但总体上AI的信息性仍占主导。","reliability":"论文指出实验任务可能并非人们担忧AI谄媚时的典型决策场景，但通过专家调查表明多数专家预期极化，而实际结果相反；同时从供给侧和需求侧论证了市场力量可能不会加剧极化效应。","relevance":"该研究直接测量了LLM建议对人类经济决策的因果影响，有真实人类实验对照，覆盖多种经济学决策场景，并探讨了谄媚偏差的行为后果，高度契合研究者对LLM仿真可靠性及偏差的关注，值得精读原文。","inspiration":"借鉴其多任务、多处理组、测量初始倾向与最终选择变化的设计，可清晰分离AI建议的极化/去极化效应｜可迁移到资产配置建议场景，研究AI理财顾问的谄媚倾向是否影响投资者的风险资产配置｜以真实投资者为被试，随机提供基线或增强谄媚的AI投资建议，测量其初始风险倾向与最终配置的变化，并以无建议组为对照，同时收集真实市场数据验证外部有效性。"}},{"id":"2607.26348","version":1,"title":"When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses","zh_title":"当合成用户失败：LLM模拟人类调查回答的跨领域基准测试","abstract":"Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families. First, at the individual level no LLM beats even the strongest baseline; on cross-cultural values every model falls well below it, and the gap survives distance-aware and proper scoring. Second, models systematically over-determine demographics, treating identity as far more predictive of attitudes than it is among real people, a distortion present for nearly every question-group combination and robust to a coding-invariant measure. Neither failure is remedied by a larger, more capable model. A decision-impact analysis shows why this matters in practice: on a segment-targeting task the models inflate between-segment gaps two to fourfold, would direct a team to the wrong segment in half of U.S. and most cross-cultural cases, and manufacture segment splits that do not exist in real people. We make the cross-domain benchmark and the evaluation framework available on request, so that teams can determine in advance when synthetic-user evidence is safe for decision support and when it is not.","authors":["Zihan Chen","Di Zhu","Lei Nico Zheng"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26348","pdf_url":"https://arxiv.org/pdf/2607.26348","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM仿真","人类调查","失效分析"],"reason":"直接评估LLM仿真人类调查的失效条件，有真实人类数据对照，涉及社会态度和政策场…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":1,"question":"在人口统计提示和调查模拟协议下，LLM作为合成用户何时会失效？","design":"使用四个模型（涵盖两个家族、8B到前沿能力），在两种提示格式下，对美国综合社会调查（GSS）和世界价值观调查（WVS）的真实人类回答进行仿真，测量个体预测准确度、总体分布复现和人口统计结构忠实度。","baseline":"基于真实人类数据拟合的朴素人口统计基线（包括人口查找表、逻辑回归、随机森林），在留出的人类数据上评估。","findings":"所有模型在个体层面均未超越最强基线，且系统性地过度决定人口统计特征，将身份视为比真实人类中更具预测性的因素；在细分目标定位任务中，模型夸大了细分群体间的差距，并制造了真实人类中不存在的细分分裂。","reliability":"论文指出失效在人口统计提示和所测试的调查模拟协议下跨领域、模型和家族复现，且更大或更强的模型未能弥补这些失效；但未讨论其他提示策略或协议下的潜在有效性。","relevance":"该研究直接评估了LLM仿真人类调查的失效条件，提供了跨领域基准和真实人类数据对照，对关注仿真可靠性与偏差的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其统一协议、多模型跨领域测试和朴素人口统计基线设计，可迁移到经济预期形成或政策偏好调查仿真中，例如用LLM模拟不同人口群体对通胀预期的回答，以真实消费者预期调查数据为基准，检验仿真是否高估人口统计的预测力并扭曲预期分布。"}},{"id":"2607.26899","version":1,"title":"Human diversity fuels collective creativity that large language models cannot simulate or sustain","zh_title":"人类多样性推动集体创造力，而大语言模型无法模拟或维持","abstract":"Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. We tested both challenges in a preregistered creative metaphor experiment with native (L1) and non-native (L2) English writers, who wrote without AI, with AI-generated ideas (AI ideation), or with AI refining their own ideas (AI refinement). L2 writers contributed more collective diversity than L1 writers, with native-language ideation showing the most diverse pools. AI ideation compressed collective diversity for everyone and left the L2 advantage undetectable, whereas AI refinement preserved both. We then simulated the entire writer pool using personas built from participants' real backgrounds, three model families, native-language prompting, and elevated sampling temperatures. Every simulated pool fell below every human pool, and pushing models further induced diversity only through degenerate text. However, at the individual level, AI ideation raised writers' ratings, pitting private incentives against the collective good, except when L2 writers used their native language, which benefited both. Human diversity remains a valuable creative resource that current AI cannot simulate or sustain; the design of human-AI collaborative workflows determines whether it survives.","authors":["Mengchen Dong","Hiromu Yakura"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26899","pdf_url":"https://arxiv.org/pdf/2607.26899","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM人类仿真","创意实验","多样性对照"],"reason":"用LLM模拟人类创意实验，与真实人类数据对照，评估仿真失效条件，高度相关。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":2,"question":"在创意生产中，人类多样性（以母语/非母语英语写作者为代理）在AI辅助下是否仍能带来集体多样性优势？LLM能否通过模拟多样性来替代真实人类多样性？","design":"本研究并非纯粹的仿真研究，而是先进行人类实验，再用LLM仿真进行对比。人类实验：招募母语（L1）和非母语（L2）英语写作者，随机分配到无AI、AI生成创意（AI ideation）或AI润色自有创意（AI refinement）三种条件，生成英语隐喻，测量集体多样性（输出池的变异度）和个体评分。仿真部分：基于参与者真实背景构建persona，使用三种模型家族、母语提示和升高采样温度，模拟整个写作者池，测量集体多样性，并与人类池对比。","baseline":"真实人类数据：人类实验中L1和L2写作者在三种AI协作条件下的隐喻输出池的集体多样性，以及无AI条件下的基线。","findings":"AI生成创意条件压缩了所有人的集体多样性，且使L2写作者的多样性优势消失，而AI润色条件保留了多样性和L2优势。所有LLM模拟池的集体多样性均低于任何人类池，且提高模型温度仅通过生成退化文本来增加多样性。","reliability":"论文指出LLM模拟无法复现人类集体多样性，即使使用真实背景构建persona、母语提示和升高温度，模拟池的多样性仍低于人类池，且过度推动模型会导致文本退化。","relevance":"该研究直接对比真实人类与LLM仿真在集体创意多样性上的表现，揭示了LLM仿真在捕捉群体层面变异时的失效，并探讨了AI协作设计如何影响多样性存留，高度契合研究者对仿真可靠性、失效条件及人类基准对照的关注，值得精读。","inspiration":"借鉴其将人类实验与LLM仿真直接对比的双阶段设计，以及用集体多样性（而非个体准确度）作为核心结果变量的测量思路。｜可迁移至政策沟通中的创意生成场景，如不同语言背景的公众对政策隐喻的解读与再创作，评估AI辅助是否削弱观点多样性。｜招募母语和非母语政策受众作为被试，随机分配至无AI、AI生成政策解释隐喻、AI润色自有隐喻三种条件，测量集体隐喻多样性，并以真实公众咨询数据作为对照基准，同时用基于被试背景的LLM persona模拟整个群体，检验仿真多样性是否匹配人类基准。"}},{"id":"2607.27100","version":1,"title":"Can Large Language Models Represent Urban Publics? Behavioral Replication and Population Mismatch in an Affordable-Housing Experiment","zh_title":"大语言模型能代表城市公众吗？一项可负担住房实验中的行为复现与人口错配","abstract":"There is growing interest in using large language models (LLMs) as low-cost proxies for resident attitudes in urban planning. Previous work shows that LLMs can predict average results of survey experiments, but less is known about whether they preserve the spatially anchored, identity-conditioned structure behind those averages, namely how support changes as a project approaches homes and how that response divides across tenure and partisan groups. We compared eight open-weight LLMs with 843 respondents in a US affordable-housing survey experiment, testing whether they reproduced the owner-renter difference in support change as a proposed development moved from 2 miles to 1/8 mile. Qwen 2.5 14B was closest (-0.242 versus the human -0.285) and was the only model to meet the prespecified +/-0.20 equivalence criterion; Phi-4 14B was directionally aligned but attenuated (-0.150), and other models showed weak, null, or reversed moderation. This aggregate match masked structural failure. Qwen attenuated the Republican contrast and exaggerated the Independent one, its RMSE across 27 party-by-tenure-by-item cells was 0.613, its median model-to-human variance ratio was 0.099, and question order shifted the contrast by +0.367. Identity-cue removal and selective nonresponse changed which comparisons were estimable, and rationale-first responses differed from matched direct-choice responses in 20.6-35.3% of focal comparisons. An LLM can thus approximate one aggregate contrast while failing to preserve the population structure, within-group heterogeneity, and measurement stability that generate it. Model evaluation in urban planning should test whether this spatial and social structure survives simulation, not only average effects.","authors":["Yuxuan Cai","Yequan Hu","Hongqian Li","Zhanghong Ju","Shuying Guo"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27100","pdf_url":"https://arxiv.org/pdf/2607.27100","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM人类仿真","行为复现","政策评估"],"reason":"用LLM复现住房实验中的行为差异，与843名人类被试对照，评估仿真失效的结构性…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":3,"question":"大语言模型能否在可负担住房实验中复现人类的空间邻近效应和群体结构差异？","design":"使用8个开源大语言模型模拟美国居民，施加住房项目距离（2英里 vs. 1/8英里）的处理，测量支持度变化，并考察房主-租户、党派身份的调节效应。","baseline":"843名美国受访者在同一可负担住房调查实验中的真实回答。","findings":"Qwen 2.5 14B在总体房主-租户差异上最接近人类，但掩盖了群体结构失效：共和党人对比减弱、独立人士对比夸大，组内方差远低于人类，且问题顺序和身份提示移除会改变结果。","reliability":"模型在总体效应上可能匹配，但无法保留生成该效应的群体分布、组内异质性和测量稳定性；身份提示、问题顺序和回答格式变化会导致估计结果不一致。","relevance":"该研究直接检验LLM仿真在空间-社会结构上的失效，提供了从总体匹配到群体结构分解的严格验证框架，对关注经济学实验和政策评估中仿真可靠性的研究者极具参考价值。","inspiration":"借鉴其将总体处理效应分解为子群体条件对比的验证方法，并引入问题顺序、身份提示等测量稳定性检验。｜可迁移到政策评估中的邻避效应实验或地方公共品供给偏好研究，如垃圾处理厂、风电场选址的公众接受度。｜以LLM模拟不同收入、党派、住房产权的居民，施加设施距离和补偿方案处理，测量支持度变化，用真实居民调查数据对照子群体效应和顺序效应。"}},{"id":"2607.02464","version":2,"title":"Will Scaling Improve Social Simulation with LLMs?","zh_title":"扩大规模会改善基于大语言模型的社会仿真吗？","abstract":"Large Language Model (LLM) social simulations are a promising research method, but they are not yet faithful enough to be adopted widely. In this work, we investigate whether the current scaling paradigm in language modeling is likely to close these gaps, or whether simulation fidelity is orthogonal to general capabilities and therefore deserving of more research attention. We use scaling laws to study the relationship between LLMs' compute scale, general capability benchmarks, and the fidelity of social simulation in three representative sub-domains: opinion modeling, behavioral simulation, and longitudinal forecasting. Surprisingly, we discover strong compute scaling in all three settings, using a suite of 85 transformer LLMs with the Qwen3 architecture pre-trained on the DCLM web text corpus under fixed-compute budgets from $10^{18}$ to $10^{20}$ FLOPs. Then we evaluate 35 larger and more capable open-weight models up to 70B parameters, allowing us to predict downstream accuracy from loss. This reveals that the majority of behavioral and opinion simulation tasks will rapidly improve with scale, particularly when they involve populations that are well-represented in English web corpora. Longitudinal forecasting and underrepresented opinions scale more slowly, especially when they are less correlated with general knowledge and reasoning benchmarks like MMLU. In behavior simulation, scaling fails to improve model calibration with human cognitive biases like risk aversion, as well as human heuristics like learning correlated rewards from related tasks. On these tasks, even fine-tuned models fail to noticeably scale up performance from 0.5B to 8B parameters. Taken together, we conclude that scale will improve social simulations in most settings, but outliers exist, and improvements will be less reliable in low-resource domains.","authors":["Caleb Ziems","William Held","Su Doga Karaca","David Grusky","Tatsunori Hashimoto","Diyi Yang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-07-30","first_seen":"2026-07-02","revised_at":"2026-07-30","abs_url":"https://arxiv.org/abs/2607.02464","pdf_url":"https://arxiv.org/pdf/2607.02464","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM社会仿真","缩放规律","仿真保真度"],"reason":"直接研究LLM社会仿真的保真度与缩放规律，含人类数据对照和失效条件分析。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":4,"question":"当前大语言模型的缩放范式能否缩小社会仿真保真度的差距，还是仿真保真度与通用能力正交？","design":"使用85个基于Qwen3架构、在DCLM语料上预训练的Transformer模型（0.2B–12B参数，固定计算预算10^18–10^20 FLOPs）进行受控计算缩放实验，并评估35个更大开源模型（最高70B参数），在意见建模、行为仿真和纵向预测三个子领域测量仿真损失与准确率。","baseline":"世界价值观调查（WVS）、Psych-101实验数据、美国人生活变迁（ACL）纵向研究等真实人类数据。","findings":"多数行为和意见仿真任务随模型规模扩大而快速改善，尤其在英语网络语料中代表性好的人群上；但纵向预测和代表性不足的意见缩放较慢，且缩放未能改善模型在风险厌恶等认知偏差及关联奖励学习启发式上的校准。","reliability":"在低资源领域和与通用推理基准相关性弱的任务上改进不可靠；缩放对风险厌恶等人类认知偏差的校准无效，微调后也未观察到参数缩放效果。","relevance":"该研究直接评估LLM社会仿真的缩放规律与失效条件，含真实人类对照，与关注经济学实验和政策评估仿真的研究者高度相关，值得精读。","inspiration":"借鉴其受控计算缩放实验与观测校准函数结合的方法，系统评估模型规模对仿真保真度的因果效应。｜可迁移到资产定价实验或消费者跨期选择等经济决策仿真，检验缩放是否改善风险偏好与时间偏好复现。｜以LLM为被试，施加不同风险/跨期选择任务，测量选择分布与真实实验数据（如实验室或调查数据）的偏差，用缩放定律预测更大模型的保真度。"}},{"id":"2607.26317","version":1,"title":"Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach","zh_title":"对齐LLM模拟考生与真实考生以进行心理测量校准：一种认知诊断画像方法","abstract":"Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natural-language profiles and sampled under an uninformative or an informative distribution. Using the Tatsuoka fraction-subtraction dataset (536 examinees, 15 items, five attributes), we evaluated eight LLM configurations under no-profile, uninformative-CDP, and informative-CDP conditions, assessing alignment with human examinees at the ability-distribution, mastery-profile, and item-difficulty levels. CDP improved all three levels: distributional overlap rose across configurations; weighted correlations between profile-level scores and human profile expectations reached 0.92 to 0.98; and item-difficulty recovery improved in rank order and absolute alignment, most for reasoning-enabled models; in the strongest case, Gemini 3.0 Flash (Thinking), one-parameter logistic (1PL) difficulty Spearman correlations rose from 0.24 to 0.86 and 0.90 and the root-mean-square error (RMSE) fell from 6.31 to 1.30 and 0.90; the informative condition helped most where profile-level alignment was strong. CDP brings LLM-simulated examinees into closer psychometric alignment with human examinees, making them practical for operational test development.","authors":["Wenjie Zhou","Yunting Liu","Renjiao Tang","Mark Wilson"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26317","pdf_url":"https://arxiv.org/pdf/2607.26317","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","心理测量","人类数据对照"],"reason":"用LLM模拟考生作答并与真实人类数据对照，评估对齐效果，属于教育测量中的人类仿…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":4,"question":"如何通过认知诊断画像提示，使大语言模型模拟的考生在心理测量校准中与真实人类考生对齐？","design":"使用八种大语言模型配置（含推理与非推理模型），在无画像、无信息认知诊断画像（CDP）和有信息CDP三种条件下，以零样本方式提示模型模拟考生作答Tatsuoka分数减法数据集（536名考生、15题、5个认知属性），评估生成的反应数据在能力分布、掌握模式及题目难度三个层面与人类数据的对齐程度。","baseline":"Tatsuoka分数减法数据集，包含536名真实人类考生对15道题目的作答反应及认知属性掌握模式。","findings":"CDP框架显著提升了LLM模拟考生与人类考生在能力分布、掌握模式和题目难度三个层面的对齐度；在最佳配置下，题目难度排序相关系数从0.24升至0.90，均方根误差从6.31降至0.90，有信息CDP在画像层面对齐较强时帮助最大。","reliability":"论文未讨论","relevance":"该研究直接以真实人类数据为基准，评估LLM仿真在心理测量校准中的对齐效果与偏差，属于教育测量场景下的人类仿真验证，与研究者关注的经济学实验和政策评估中的仿真可靠性问题高度相关，值得精读。","inspiration":"借鉴其通过结构化认知画像（属性掌握模式）注入异质性、并对比无信息与有信息分布采样的处理设计，以控制仿真人群的多样性与偏差。｜可迁移至教育经济学或劳动经济学中的技能测评场景，如职业资格考试的题目预测试或人力资本评估中的能力诊断。｜以LLM模拟不同技能掌握模式的求职者，处理为随机分配无信息或有信息的认知画像提示，结果变量为模拟作答反应，以真实大规模技能测评数据（如PIAAC）作为人类基准对照。"}},{"id":"2607.26288","version":1,"title":"The Innate Economic Preferences of Language Models","zh_title":"语言模型的内在经济偏好","abstract":"Language models increasingly settle real resource tradeoffs on behalf of principals yet their economic preferences remain unobserved. We demonstrate their generation rule is isomorphic to the random utility model of discrete choice. This allows internal logit scores to structurally identify preferences. Estimating risk attitudes across twelve models in a portfolio task reveals universal but heterogeneous risk aversion. Although models reject strictly dominated options, their elicited preferences fail invariance tests and violate the independence of irrelevant alternatives across varying experimental prompts. Finally, fine tuning establishes that a principal can explicitly engineer a target risk attitude.","authors":["Joy Buchanan","Joshua Foster"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26288","pdf_url":"https://arxiv.org/pdf/2607.26288","source_feed":"econ.EM","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM经济偏好","人类仿真","风险态度"],"reason":"用LLM替代人类被试测量经济偏好，有真实人类数据对照，涉及风险态度和不变性检验…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":6,"question":"语言模型在面临经济权衡时的默认偏好是什么，其选择是否满足显示性偏好公理从而可被解释为稳定效用？","design":"本研究并非用LLM仿真人类被试，而是将12个语言模型本身作为决策主体，在受控的投资组合选择任务中测量其风险态度。通过强制模型从不同风险-收益特征的资产菜单中单选一个资产，并利用模型输出的logit向量（开放模型）或重复抽样选择（闭源模型）来结构性地识别偏好参数，同时检验其选择是否满足完备性、自反性、单调性、传递性、连续性和无关选项独立性等显示性偏好公理。","baseline":"无对照","findings":"所有模型均表现出普遍但异质的风险厌恶，且拒绝严格劣选项；但其偏好未能通过不变性检验，并随实验提示变化而违反无关选项独立性。通过微调可以显式地工程化目标风险态度。","reliability":"论文指出模型的偏好对选项的呈现方式不具不变性，无关选项的加入或选项位置变动会改变偏好强度，尤其在接近无差异时这种不稳定性变得可观测，表明偏好强度不稳定而排序相对稳定。","relevance":"该研究直接测量LLM作为经济主体的内在偏好并检验其理性公理，虽未以人类为基准，但为评估LLM替代人类进行经济决策的可靠性提供了关键的方法论和实证证据，值得精读。","inspiration":"借鉴其利用模型内部logit直接观测系统效用指数的方法，可避免传统离散选择模型对误差分布的依赖，实现偏好的结构化识别。｜可迁移到信贷审批歧视研究中，用LLM扮演信贷员，测量其在不同申请人特征下的风险偏好与歧视程度。｜以多个LLM为被试，设计不同风险-收益特征的贷款申请菜单，记录模型选择的logit或重复抽样选择，估计其风险厌恶参数，并与真实信贷员的历史审批数据对照，检验LLM决策的偏差与一致性。"}},{"id":"2607.25292","version":1,"title":"Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe","zh_title":"指令微调语言模型无法从它们能描述的分布中采样","abstract":"Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution. We show this draw does not exist: instruction-tuned models do not sample from distributions, they collapse to a single output. The same persona on the same question returns the same answer on more than half of items in a public-opinion benchmark. The collapse is sharp: the model's internal probabilities concentrate on a single option, and the failure is substantially amplified by instruction tuning: across three model families with materially different post-training pipelines, every instruction-tuned model fails on every task we test, while base models fail far less often. Strikingly, the same model that cannot sample from a distribution can describe it accurately in a single call. We call this gap the KNOWS/DOES split, and trace it to a degenerate sampling primitive visible in the logits and induced by alignment training. Exploiting this split, asking the model to describe the response distribution in one call more than halves the error against human survey data compared to persona aggregation. For applications that require per-persona outputs, we propose Prompt-Perturbed Argyle (PPA), which reduces the same error by 21% at no added cost.","authors":["Chaemin Jang","Dongman Lee","Jihee Kim"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25292","pdf_url":"https://arxiv.org/pdf/2607.25292","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM人类仿真","分布采样失效","算法保真度"],"reason":"直接研究LLM仿真人类调查的分布采样失效，有真实人类数据对照，批判性指出失效条…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":2,"question":"指令微调后的语言模型能否从它们能描述的分布中进行独立采样？","design":"使用多个指令微调模型（如Llama、Gemma等）扮演不同人口学特征的人设，对每个（人设，问题）对重复调用模型，测量输出是否多样；同时对比基础模型与指令微调模型，分析logits集中程度，并测试“描述分布”与“逐次采样聚合”两种方法的误差。","baseline":"Pew American Trends Panel调查的真实人类回答分布，以及OpinionQA基准中的人口学匹配数据。","findings":"指令微调模型在逐次调用中输出高度集中，超过一半的（人设，问题）对每次返回相同答案，无法模拟分布采样；但同一模型能准确描述分布，这种“知道但做不到”的分裂源于对齐训练。","reliability":"论文指出失效主要出现在指令微调模型上，基础模型失败较少；但未讨论不同人设复杂度、开放域问题或动态交互场景下的局限。","relevance":"该研究直接揭示LLM仿真人类调查的分布采样失效，有真实人类数据对照，并批判性指出对齐训练是主因，高度契合研究者对仿真可靠性与偏差的关注，值得精读。","inspiration":"借鉴其通过对比基础模型与指令微调模型来归因失效来源的设计，以及用logits分析揭示内部概率集中化的测量方法。｜可迁移到消费者信心调查或通胀预期形成的仿真研究中，检验LLM能否复现真实人群的预期分布。｜以LLM扮演不同收入、年龄的消费者，施加“描述分布”与“逐次采样”两种处理，结果变量为预期通胀率的分布，用密歇根大学消费者调查的真实数据做对照。"}},{"id":"2607.24782","version":1,"title":"Personalization, Personas, and Forecasting in Value Alignment","zh_title":"价值对齐中的个性化、角色与预测","abstract":"LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people would answer value-laden questions. We test whether these framings are interchangeable using the World Values Survey (WVS). We evaluate GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Flash, and Qwen3-235B on 101 WVS-derived questions across 13 language-country slices, comparing a language-only baseline with user-country, persona-country, and third-person prompts. Across 21,008 model-response rows, prompt framing is a first-order determinant of cultural alignment: country cues often shift answers substantially, but not all shifts move toward matched human response distributions. Third-person forecasting yields the strongest directional alignment for three of the four hosted models, while personalization and role-play are weaker or less stable. Alignment gains concentrate on salient value dimensions such as religiosity, gender roles, and work-oriented material values, whereas institutional trust and democracy-related questions remain difficult. These results show that prompt framing is not a cosmetic choice in cultural value elicitation; it changes both model behavior and measured alignment.","authors":["James Wedgwood","Pratiksha Thaker","Neil Kale","Virginia Smith"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24782","pdf_url":"https://arxiv.org/pdf/2607.24782","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM仿真","价值观调查","文化对齐"],"reason":"用LLM仿真人类价值观调查，以WVS真实数据为基准，评估提示框架对文化对齐的影…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":3,"question":"在价值观调查中，LLM的个性化、角色扮演和预测三种提示框架是否可互换，以及哪种框架能更好地对齐真实人类回答分布？","design":"使用GPT-5.4、Claude Sonnet 4.6、Gemini 2.5 Flash和Qwen3-235B四个模型，在13个语言-国家切片上回答101个源自世界价值观调查（WVS）的问题，比较四种提示条件：仅语言基线、用户国家提示、角色扮演国家提示和第三人称预测提示，测量模型回答与对应国家WVS真实回答分布的对齐程度。","baseline":"世界价值观调查（WVS）第7波（2017-2022）中对应国家、对应问题的真实人类回答分布。","findings":"提示框架是文化对齐的一阶决定因素，国家线索常显著改变回答，但并非所有改变都朝向匹配的人类分布；第三人称预测在四个模型中的三个上产生最强的方向对齐，个性化和角色扮演较弱或不稳定，对齐收益集中在宗教、性别角色和工作物质价值观等显著维度，而制度信任和民主相关问题仍难以对齐。","reliability":"论文未讨论","relevance":"该研究直接以LLM仿真人类价值观调查，用WVS真实数据作为基准，系统比较了三种身份提示框架的对齐效果，并指出了仿真在制度信任等维度上的失效，高度契合研究者对LLM仿真可靠性、偏差及失效条件的关注，值得精读原文。","inspiration":"借鉴其多框架对比设计，通过改变提示中的身份信息（用户、角色、第三人称）来分离LLM行为模式，并用真实调查数据作为对齐基准。｜可迁移到消费者信心调查、通胀预期或政策支持度等经济态度仿真，检验不同提示框架下LLM能否复现特定人群的经济心理。｜以LLM为被试，施加用户国家、角色扮演和第三人称预测三种提示处理，测量其对未来通胀、就业预期的回答，并以密歇根大学消费者调查的真实数据为对照，评估哪种框架能最好地复现不同收入群体的预期分布。"}},{"id":"2607.25447","version":1,"title":"CoRenew: A large language model agent-based policy simulation platform for multifamily residential redevelopment","zh_title":"CoRenew：基于大语言模型代理的多户住宅再开发政策仿真平台","abstract":"The difficulty of collective action remains a central challenge in the design of policies for multifamily residential redevelopment. Stakeholders continually adjust their decisions in response to evolving negotiation contexts and the reactions of others, meaning that when a policy intervenes and which stakeholders it targets can substantially reshape collective outcomes. Assessing these adaptive responses ex ante remains difficult because existing simulation models often rely on predefined behavioral rules. Here, we present CoRenew, an open-source platform that uses LLM-based agents to simulate negotiations among multiple stakeholders and evaluate the effects of alternative policy combinations. Integrating open source geographic and demographic data, the platform can generate synthetic residents, simulate negotiation dynamics under alternative policy settings and compares policy performance across competing objectives. It supports both numerical and semantic policy inputs and includes built-in tools for visualization and result export. We validate its behavioral realism against survey responses from 324 residents and a nine-month observed negotiation process from a real redevelopment case. With its modular and adaptable architecture, CoRenew can be used to assess policies across different institutional and cultural contexts.","authors":["Yudi Zhang","Yuming Lin","Li Tian","Yu Wang","Jianghao Yu"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25447","pdf_url":"https://arxiv.org/pdf/2607.25447","source_feed":"cs.MA","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","政策评估","人类数据对照"],"reason":"用LLM代理模拟多户住宅再开发谈判，并与324份居民调查和9个月真实谈判过程对…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":4,"question":"如何利用LLM代理模拟多户住宅再开发中的多轮利益相关者谈判，以评估不同政策组合的效果？","design":"使用LLM代理模拟居民、开发商等利益相关者，在马尔可夫博弈框架下进行多轮谈判，通过链式思维推理（基于计划行为理论）生成决策；施加数值型和语义型政策干预，测量谈判结果（如共识达成、不平等程度、居民效用等）。","baseline":"324份居民调查回复和一起真实再开发案例中为期9个月的谈判过程观察数据。","findings":"模拟的结构模型恢复了19条基于调查的路径，方向符号100%一致，统计显著性94.7%一致；最佳LLM代理的谈判轨迹与真实观察密切匹配。案例研究表明，促进共识的政策可能带来公平风险，补贴效应呈非线性：不平等减少仅在高补贴水平显现，低收入居民效用增益呈边际递减。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查和谈判过程为基准验证LLM代理的行为真实性，并揭示了政策仿真中的公平风险与非线性效应，与您关注的LLM仿真可靠性及政策评估场景高度契合，值得精读原文。","inspiration":"借鉴其利用真实调查数据和长期观察过程作为多维度基准验证LLM代理行为的方法，以及将语义政策干预纳入仿真的设计。｜可迁移到公共政策评估中的协商式预算分配或社区拆迁补偿谈判模拟，测试不同信息透明度和参与机制对分配公平的影响。｜以LLM代理模拟居民和官员，施加不同协商规则（如公开投票vs.闭门会议）作为处理，结果变量为预算分配基尼系数和居民满意度，对照真实社区协商实验数据或历史分配记录。"}},{"id":"2604.02458","version":3,"title":"Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments","zh_title":"统计真实性不能证明LLM能估计社会科学实验中的处理效应","abstract":"Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-world experiments are costly or infeasible. The treatment-effect estimates are often evaluated using statistical realism, the degree to which simulated responses reproduce properties of observed human responses, although whether realism predicts treatment-effect accuracy remains unknown. Here we test this proxy relationship by jointly measuring statistical realism and treatment-effect accuracy on the same simulated responses in a cross-national experiment with 59,508 participants from 62 countries using three LLMs. The correlation between statistical realism and treatment-effect accuracy is weak, and optimizing for statistical realism can even worsen treatment-effect accuracy when selecting models, prompts, and target populations. The pattern replicates in two additional cross-national experiments spanning 12 and 27 countries with 20,785 participants. The divergence between the two reflects distinct error structures and is larger for behavioral outcomes, where models appear to extrapolate behavioral effects from attitudinal patterns. Because this divergence may remain hidden in deployment, errors can propagate into simulation-informed decisions. We introduce a diagnostic framework for LLM-generated synthetic data and discuss how treatment-effect validation should proceed under varying availability of experimental benchmarks. Simulated responses and simulated treatment effects are distinct estimation targets, and evidence for one does not certify the other.","authors":["Zonghan Li","Feng Ji"],"categories":["cs.CY","cs.AI","cs.ET"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-07-28","first_seen":"2026-04-02","revised_at":"2026-07-28","abs_url":"https://arxiv.org/abs/2604.02458","pdf_url":"https://arxiv.org/pdf/2604.02458","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B2","B4"],"tags":["LLM仿真","处理效应估计","统计真实性"],"reason":"直接研究用LLM仿真人类被试估计处理效应，有大规模真实人类数据对照，批判性指出…","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":1,"question":"在社会科学实验中，用LLM仿真人类回答的统计逼真度能否作为其估计处理效应准确性的有效代理指标？","design":"使用GPT、Gemini、Claude三个LLM模拟来自62个国家59,508名参与者的跨国实验，施加气候相关干预，测量气候信念、政策支持和环保行动三个结果变量，并比较不同提示策略（少样本提示与VBN链式思考提示）下的仿真表现。","baseline":"以同一跨国实验中真实人类参与者的个体回答和平均处理效应（ATE）作为对照基准，并引入OLS和LASSO回归作为监督学习基线。","findings":"统计逼真度与处理效应准确性之间的相关性很弱，且优化统计逼真度在选择模型、提示和目标人群时反而可能降低处理效应估计的准确性。这种背离在行为结果上更为明显，模型似乎从态度模式外推行为效应，导致误差结构不同。","reliability":"论文指出，统计逼真度不能保证处理效应估计的准确性，两者是独立的估计目标；在缺乏实验基准时，仅靠响应层面的逼真度验证可能导致隐蔽的错误传播，并提出了一个诊断框架来指导不同基准可用性下的验证流程。","relevance":"该研究直接回应了用LLM替代人类被试进行实验仿真的可靠性问题，提供了大规模真实人类对照和批判性证据，对关注经济学实验和政策评估仿真的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其同时测量统计逼真度和处理效应准确性并对比二者关系的验证框架，以及利用跨国多实验复现来检验结论稳健性的做法。｜可迁移到政策评估中的行为干预仿真，如税收提示对纳税遵从度的影响、信息框架对退休储蓄选择的作用等。｜以真实纳税人为被试，施加不同税收道德信息处理，用LLM仿真纳税遵从度，以税务行政数据中的实际遵从行为作为对照，比较仿真ATE与真实ATE的偏差。"}},{"id":"2607.22605","version":1,"title":"Socioeconomic Inference in LLM Medical Triage: Same Symptoms, Different ZIP Code","zh_title":"大语言模型医疗分诊中的社会经济推断：相同症状，不同邮编","abstract":"We investigate whether large language models alter medical triage recommendations for identical symptoms when only the patient's socioeconomic status (SES) varies. Using three deployment-tier models (Gemini 3.5 Flash, Claude Sonnet 4.6, GPT-5.4-mini), we hold a single neurological symptom profile fixed and vary the SES signal along two channels: explicit (insurance status, occupation, housing) and implicit (a US ZIP code, with no other socioeconomic information). All three models raise their emergency-room (ER) referral rate for lower-SES patients given the explicit signal (spreads of 13-50 percentage points). The effect is in the protective direction: lower-SES patients are sent to the ER more often, not less. The model's stated reasoning stays clinically near-identical across conditions, so the shift is invisible to a reasoning-trace audit. Critically, sensitivity to the implicit ZIP-code signal is model-dependent: Gemini infers SES from geography alone, shifting its ER rate by a pooled 11.4 points across six US ZIP-code pairs (p = 1.4e-7, same direction in 6/6 pairs), while Claude Sonnet 4.6 stays flat (-0.1 points) and GPT-5.4-mini shows only a small difference that is not sign-consistent (2.0 points, predicted direction in just 2 of 6 pairs), neither a reliable ZIP-code effect, despite both responding to the explicit signal. This reveals an explicitness gradient in the signal: every model acts on socioeconomic status when it is stated outright, but only Gemini Flash acts on it when it must be inferred from a proxy as thin as five digits. We read this as a model-specific difference rather than a size or cost effect. A single-sentence system-prompt instruction reduces but does not eliminate the effect (Gemini's gap between low- and high-income ZIPs falls from 11.4 to 5.8 points). We release all code, prompts, and raw results.","authors":["Qi Han Wong"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22605","pdf_url":"https://arxiv.org/pdf/2607.22605","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","医疗决策偏差","社会经济地位"],"reason":"用LLM模拟不同SES患者的医疗分诊决策，与真实人类行为对照，评估偏差与失效条…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":3,"question":"当患者症状完全相同时，大语言模型是否会因社会经济地位（SES）信号而改变医疗分诊建议？","design":"用三个部署级LLM（Gemini 3.5 Flash、Claude Sonnet 4.6、GPT-5.4-mini）扮演分诊助手，固定神经症状描述，通过显性渠道（保险、职业、住房）和隐性渠道（仅提供美国邮政编码）施加SES处理，测量急诊转诊率作为结果变量。","baseline":"无对照","findings":"显性SES信号下，所有模型均提高低SES患者的急诊转诊率（13-50个百分点），方向为保护性；隐性邮政编码信号下，仅Gemini能推断SES并产生一致的转诊率差异（11.4个百分点），Claude和GPT无此效应，且模型推理文本均未提及社会经济因素，偏差无法通过推理审计发现。","reliability":"论文指出系统提示指令可减少但未消除偏差，且未探讨真实医疗场景中因随访可及性差异而调整分诊的合理性，也未在人类医生或真实患者数据上验证。","relevance":"该研究直接以LLM模拟不同SES患者的医疗决策，揭示隐性信号下的模型特异性偏差，与研究者关注的仿真可靠性、偏差条件及政策评估场景高度相关，值得精读原文。","inspiration":"借鉴其双通道处理设计（显性vs.隐性信号）和符号一致性稳健性检验，可迁移至信贷审批中的地域歧视研究｜用LLM扮演信贷员，以显性收入/职业和隐性邮政编码作为处理，测量贷款批准率，并以真实银行信贷数据或审计研究结果作为对照。"}},{"id":"2607.21757","version":1,"title":"Co-design of LLM-based preference agents: participation may drive overtrust","zh_title":"基于大语言模型的偏好代理协同设计：参与可能驱动过度信任","abstract":"Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designed personal preference agents in the domain of household energy, via a background survey, co-design interview, and validation survey. Participants engaged readily and mostly came to see their agents as representing them well. Independent validation, however, revealed mixed human-agent alignment, with agent responses markedly more homogeneous, decisive, and abstract than the human sample. I argue that participation and process transparency can act as an \"overtrust engine\" that promotes trust while concealing systematic misalignment with potential structural consequences at scale. I develop this as a core mechanism in participatory preference agent design, treating individual alignment not as a fixed state but as an enacted process.","authors":["Michael J. Fell"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-27","first_seen":"2026-07-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21757","pdf_url":"https://arxiv.org/pdf/2607.21757","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","偏好代理","人机对齐"],"reason":"用LLM模拟人类偏好并与真实人类数据对照，评估仿真可靠性与偏差，批判性指出过度…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":2,"question":"在LLM偏好代理的参与式协同设计中，参与和过程透明是否会导致过度信任，从而掩盖系统性的对齐偏差？","design":"本研究采用质性为主的混合方法，招募12名参与者，通过背景调查、协同设计访谈和验证调查，在家庭能源领域协同设计个人偏好代理，并比较参与者感知的代理表现与独立评估的代理表现。","baseline":"以12名参与者在验证调查中的真实回答作为人类基准，与代理回答进行对齐比较。","findings":"参与者普遍认为其协同设计的代理能很好地代表自己，但独立验证显示人-代理对齐程度参差不齐，代理回答比人类样本更同质、更果断、更抽象。参与和过程透明可能成为“过度信任引擎”，在促进信任的同时掩盖系统性偏差。","reliability":"论文指出，参与式设计可能掩盖系统性的对齐偏差，代理回答的同质化倾向可能导致少数观点被压制，且用户在不熟悉领域难以识别代理是代表偏好还是塑造偏好。","relevance":"该研究直接探讨用LLM模拟人类偏好并与真实人类数据对照，评估仿真可靠性与偏差，并批判性指出参与式设计可能引发过度信任，高度契合研究者对LLM仿真实验的批判性关注。","inspiration":"借鉴其协同设计流程与独立验证相结合的方法，可迁移到消费者金融决策偏好模拟场景，设计一个实验：招募真实消费者作为被试，通过访谈协同设计其消费信贷偏好代理，以真实信贷选择数据为基准，测量代理在风险偏好、跨期选择等任务上的对齐度与同质化程度。"}},{"id":"2607.17437","version":1,"title":"Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions","zh_title":"经验锚定提升大语言模型代理在中断期间模拟人类行为的真实性","abstract":"Large language model (LLM) agents offer a generative approach to simulating human behavior under conditions that may have few or no direct historical analogues, a common challenge in disaster and infrastructure-disruption planning. However, this generative capacity creates a validity problem: individually plausible agent reasoning may fail to reproduce empirical population behavior. We evaluate whether empirical grounding improves the statistical realism of LLM-agent simulations during disruptions. Specifically, we develop an empirically grounded LLM-agent framework that embeds demographic profiles from the American Community Survey, baseline routines from the American Time Use Survey, and urban spatial context into agent initialization, memory, decision prompts, and activity execution. An independent household survey conducted during the July 2024 Philadelphia heatwave is reserved as an external validation benchmark. Compared with an ungrounded LLM-agent baseline, the grounded model improved reconstruction of normal daily routines, increasing mean correlation with empirical activity profiles from 0.528 to 0.912 and reducing mean squared error from 0.066 to 0.008. Under heatwave conditions, the grounded model better reproduced survey-derived activity profiles, increasing mean correlation from 0.349 to 0.836 and reducing mean squared error from 0.098 to 0.012. The grounded model captured 46.4% of observed heatwave response amplitude, compared with 20.6% for the ungrounded baseline. These findings show that empirical grounding can make LLM agents more statistically credible simulators of population behavior while revealing remaining gaps in modeling human adaptation during disruptions.","authors":["Chen Xia","Zexi Kuang","Yuqing Hu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-19","first_seen":"2026-07-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.17437","pdf_url":"https://arxiv.org/pdf/2607.17437","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["人类行为仿真","经验锚定","灾害响应"],"reason":"用LLM代理模拟人类在热浪中的行为，并与真实调查数据对照，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"经验锚定（empirical grounding）能否提高LLM智能体在中断事件中模拟人类行为的统计真实性？","design":"构建经验锚定的LLM智能体框架，将美国社区调查（ACS）人口特征、美国时间使用调查（ATUS）基线日常活动模式及城市空间背景嵌入智能体初始化、记忆、决策提示和活动执行中；以2024年7月费城热浪为场景，比较锚定模型与无锚定基线模型在正常和热浪条件下的活动分布，并以同期独立住户调查作为外部验证基准。","baseline":"2024年7月费城热浪期间进行的独立住户调查，用于验证正常和热浪条件下的活动时间分布及行为变化幅度。","findings":"经验锚定显著提升了LLM智能体对正常日常活动模式的重建，与经验活动剖面的平均相关性从0.528升至0.912，均方误差从0.066降至0.008；在热浪条件下，锚定模型更好地复现了调查活动剖面，平均相关性从0.349升至0.836，均方误差从0.098降至0.012，且捕捉了46.4%的观测热浪响应幅度，远高于无锚定基线的20.6%。","reliability":"论文指出经验锚定虽大幅提升统计可信度，但锚定模型仍仅捕捉到46.4%的观测响应幅度，揭示在模拟人类适应行为方面仍存在差距；此外，研究仅针对单一城市单次热浪事件，泛化性有待检验。","relevance":"该研究直接以真实调查数据为基准，评估LLM智能体模拟人类在中断事件中行为变化的统计真实性，并揭示了经验锚定的增益与残余偏差，与您关注的人类仿真可靠性及失效条件高度吻合，值得精读。","inspiration":"借鉴其将人口普查、时间使用调查和空间背景多层数据嵌入智能体决策流程的设计，可迁移到消费者跨期选择或政策冲击下的行为响应研究｜例如，在突发性价格变动或补贴政策实验中，用LLM智能体模拟不同收入群体的消费调整，以真实家庭收支调查数据为基准，评估仿真对需求弹性和福利效应的复现精度｜设计：以ACS收入、职业和ATUS时间分配数据初始化智能体，施加电价飙升或交通补贴取消的处理，测量各时段活动与消费组合的变化，用实际家庭能源消费或出行调查微观数据作为对照基准。"}},{"id":"2607.18310","version":1,"title":"Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling","zh_title":"分布优先的人口模拟：非WEIRD LLM角色建模中的崩溃、校准与回忆","abstract":"Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent. Using real survey microdata, we show that this paradigm has a basic failure mode, and we set a distribution-first corrective against it, all measured with a deterministic, construct-validated verifier on non-WEIRD (Turkey-first) data. First, N independent LLM agents grounded on 2,414 real World Values Survey respondents fail to reproduce the population's response distribution: they pile onto a modal default (four scenarios x five seeds: concentration 0.36->0.69, entropy 1.46->0.77, 85% collapse, TVD=0.44), and the collapse is a predictable function of scenario structure (r=0.55 with a single-answer structure). Second, Verbalized Sampling (VS) fixes the field's chronic under-dispersion without training in three model families (fidelity +7 to +10; significant on Qwen, p=0.002, d=6.2), yet the same move universally overshoots into over-dispersion (SD-ratio 0.4-0.56 -> 1.26-1.37), a structural property of VS. Third, survey fidelity transfers only weakly to agentic behavior: in a single-model, single-domain booking task, a persona is dominated by a cheapest-default (~80%) that income modulates but does not override (comfort choice 0%->7%->32% across income bands). Fourth, a placebo-controlled memorization attack and an election backtest show VS keeps aggregate strength while subgroup and individual claims are contaminated by recall and underdetermination. We close with the corrective: model the distribution once (VS) and assign it to grounded characters at O(1) cost, with a budget-aware router whose honest AUC is 0.805, not the tautological 1.0 of a code-derived oracle. The central contribution needs no realism claim: it measures the internal inconsistency of the independent-agent route and the conditions under which the distribution-first route calibrates.","authors":["Gurkan Ozkan"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-07-17","first_seen":"2026-07-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.18310","pdf_url":"https://arxiv.org/pdf/2607.18310","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM人类仿真","分布校准","合成人口"],"reason":"用LLM代理模拟真实调查人群，与人类数据对照，评估分布崩溃与校准，涉及行为任务…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"独立LLM智能体能否复现真实人群的调查回答分布？若不能，如何修正？","design":"使用Qwen、GLM-5.2、Gemma-4-26B等模型，基于土耳其世界价值观调查（WVS）的2414名真实受访者微数据，对比两种路线：路线A（Verbalized Sampling，直接让模型输出总体分布）与路线B（独立智能体，每个角色独立选择后汇总），测量响应分布的集中度、熵、总变异距离等，并检验调查校准向行为任务的迁移。","baseline":"土耳其世界价值观调查（WVS）的2414名真实受访者微数据。","findings":"独立智能体路线导致分布崩溃，个体聚集到模态默认选项，崩溃程度与场景结构相关（单一正确答案时最严重）。Verbalized Sampling能提升分布保真度，但普遍导致过度离散，且调查校准向行为任务的迁移较弱。","reliability":"论文承认Verbalized Sampling存在结构性过度离散，需配合均值保持校正；调查保真度向行为任务迁移弱；子群体和个体层面推断受记忆和欠定污染。","relevance":"直接回应LLM仿真人类调查的分布可靠性问题，提供真实人类基准对照，揭示独立智能体路线的系统性崩溃及分布优先修正的利弊，对经济学实验和政策评估场景的仿真设计有重要参考价值，值得精读原文。","inspiration":"借鉴其对比独立智能体与Verbalized Sampling的仿真设计，通过总变异距离、熵等指标量化分布偏差，并检验调查校准向行为任务的迁移｜可迁移至消费者金融决策调查仿真，如风险偏好、储蓄选择或信贷需求分布预测｜以LLM智能体模拟家庭金融调查受访者，处理为独立智能体 vs. Verbalized Sampling生成风险资产配置分布，结果变量为分布距离与集中度，以真实家庭金融调查微数据为基准"}},{"id":"2607.06080","version":1,"title":"From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations","zh_title":"从蓝图到现实：基于LLM多智能体仿真建模与应用普特南社会资本理论","abstract":"Putnam's Social Capital Theory is a foundational framework for collective action and community prosperity. However, traditional empirical methods face practical limits on control and replication. Meanwhile, LLM-based social simulations are typically behavior-driven and lack theory-aligned environments for modeling Putnam's core propositions. To address these gaps, we introduce SocaSim, an LLM-based multi-agent simulation framework to study Putnam's Social Capital Theory from theoretical blueprint to simulated reality. Specifically, we build an environment integrating social network evolution, trust dynamics, and norm propagation, where agents engage in repeated collective-action experiments, and then apply the three dimensions to analyze adaptation challenges in smart elderly care. Our simulations reproduce Putnam's macro-level patterns and exhibit strong human-agent alignment at the group level. Unlike traditional methods, SocaSim traces micro-level causal pathways of social network, trust, and norms via round-by-round simulations and counterfactual interventions, enabling process-level interpretability. Taken together, these capabilities establish a research paradigm that leverages LLM agents to bridge social science and computer science.","authors":["Shiyi Ling","Zhi Zheng","Hui Zheng","Wenjun Xue","Feng Ye","Tong Xu"],"categories":["cs.CL","cs.AI","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-07","first_seen":"2026-07-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.06080","pdf_url":"https://arxiv.org/pdf/2607.06080","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","社会资本理论","集体行动实验"],"reason":"用LLM多智能体模拟集体行动，复现宏观模式并与人类数据对齐，涉及社会资本理论，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":49,"question":"如何基于LLM多智能体仿真复现并应用Putnam社会资本理论，以揭示社会网络、信任和规范在集体行动中的动态机制？","design":"构建SocaSim框架，生成具有人口属性和社会资本禀赋的LLM智能体，在动态社会网络、信任和规范环境中进行多轮集体行动实验，通过提案和执行两阶段模拟决策，并应用于智慧养老适应挑战，进行反事实干预。","baseline":"与真实老年人群体决策数据进行对比，报告群体层面决策的皮尔逊相关系数为0.974。","findings":"仿真复现了Putnam理论预测的宏观模式，并与人类群体决策高度一致；反事实干预显示，提高低社会经济地位智能体的初始信任可使技术采纳率提升15.4%，决策矛盾减少25.5%。","reliability":"论文未讨论","relevance":"该研究直接以LLM智能体替代人类被试，复现社会资本理论的集体行动模式，并与真实人类数据对齐，同时进行反事实干预评估因果效应，高度契合研究者对经济学实验和政策评估场景中仿真可靠性的关注。","inspiration":"借鉴其利用LLM智能体在动态社会网络中进行多轮集体行动实验并施加反事实干预的设计，以模拟宏观涌现模式｜可迁移至政策干预对技术采纳或合作行为影响的经济学实验，如智慧养老、金融包容性政策评估｜以LLM智能体作为被试，处理为提高低社会经济地位群体的初始信任水平，结果变量为技术采纳率和决策矛盾率，用真实老年人群体决策数据作为对照基准"}},{"id":"2607.03091","version":1,"title":"Silicon Sampling via Cross-Survey Transfer","zh_title":"基于跨调查迁移的硅采样","abstract":"Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent's answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, three open-weight LLMs (27B-120B parameters), and supervised machine learning baselines, we find that: (1) zero-shot LLMs achieve 52% accuracy on genuinely unseen items, closing to within 6 percentage points (pp) of a supervised random forest trained on same-population data; (2) a stable construct predictability hierarchy emerges, from 67% for partisan attitudes to 23% for sovereignty; and (3) variance collapse and safety alignment effects-two commonly cited LLM limitations-turn out to be more nuanced than previously reported, with variance collapse affecting supervised models as well and alignment effects varying dramatically across model families. These findings clarify both the promise and boundaries of silicon sampling.","authors":["Chan-Tung Ku","Chan Hsu","Pei-Cing Huang","Frank Cheng-shan Liu","I-Ling Cheng","Yihuang Kang"],"categories":["cs.AI","cs.CL","cs.CY","cs.MA","stat.ME"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-03","first_seen":"2026-07-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.03091","pdf_url":"https://arxiv.org/pdf/2607.03091","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM人类仿真","调查方法","算法保真度"],"reason":"直接用LLM仿真人类调查回答，有真实人类数据对照，评估可靠性与偏差，涉及选举研…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-03","rank":2,"question":"大语言模型能否在个体层面预测人类调查受访者的回答，以及这种仿真的可靠性和边界是什么？","design":"使用三个开源LLM（27B-120B参数）模拟台湾选举与民主化调查（TEDS 2024）的受访者，采用跨调查迁移框架：给定受访者对一组问题的回答，预测其对同一调查中完全不同问题的回答。","baseline":"TEDS 2024真实人类调查数据，以及基于同总体数据训练的监督随机森林模型。","findings":"零样本LLM在未见项目上达到52%准确率，仅比监督随机森林低6个百分点；不同构念的可预测性存在稳定层级，从政党态度的67%到主权问题的23%。","reliability":"方差压缩和安全对齐效应比先前报告更复杂：方差压缩同样影响监督模型，对齐效应在不同模型家族间差异显著。","relevance":"高度相关：直接评估LLM作为人类调查受访者替代品的个体层面预测能力，有真实人类数据对照，并揭示了仿真的边界（如构念可预测性差异、方差压缩与对齐效应的复杂性），值得精读原文。","inspiration":"跨调查迁移框架将受访者部分回答作为输入预测其余回答，可借鉴用于构建个体层面的行为预测模型，并设置监督机器学习基准和真实人类数据对照以评估仿真可靠性。｜可迁移到消费者跨期选择实验，用LLM模拟被试在给定部分偏好或决策后的选择一致性，检验时间偏好与自我控制偏差。｜以LLM作为被试，用真实家庭金融调查数据（如CFPS）中部分消费-储蓄问题回答作为处理输入，预测同一被试在跨期选择任务中的折现因子作为结果变量，与真实人类数据及监督模型预测对比。"}},{"id":"2606.28978","version":1,"title":"Can LLMs Hire Fairly? Racial Bias in Resume Screening","zh_title":"大语言模型能公平招聘吗？简历筛选中的种族偏见","abstract":"We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (2022). The sole 2023-vintage model reproduces the pro-White callback gap documented in field experiments on labor market discrimination ($+2.12$ pp, significant at the 1\\% level). Every model released in 2024 or after shows either a null gap or a significant pro-Black reversal (up to $-3.01$ pp). The same pattern holds on the gender axis. Based on 24,024 paired postings per model across 14 models, our results document a reversal in the direction of algorithmic hiring bias across model generations.","authors":["Zhenyu Gao","Wenxi Jiang","Yutong Yan"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-27","first_seen":"2026-06-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.28978","pdf_url":"https://arxiv.org/pdf/2606.28978","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","招聘歧视","算法偏差"],"reason":"用LLM复现招聘歧视实验，与真实人类数据对照，评估算法偏差，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-27","rank":4,"question":"大语言模型在简历筛选时是否存在种族和性别歧视，以及这种歧视的方向是否随模型版本变化？","design":"使用14个主流LLM（GPT-3.5-turbo、GPT-4o-mini、GPT-5.4-mini、Llama-3.1-8B-Instruct、Llama-3.3-70B等），采用配对简历方法，简历内容完全相同仅名字不同（黑人或白人典型名字，或男女性名字），让模型扮演HR经理决定是否邀请面试（输出“yes”或“no”），温度设为0，每个模型在种族轴上进行24,024次配对测试，在性别轴上进行48,048次配对测试。","baseline":"对照真实人类实验：Kline, Rose, and Walters (2022)的现场实验，在108家财富500强公司中发现白人名字比黑人名字获得更多回电（1.6个百分点）；以及Bertrand and Mullainathan (2004)的经典实验。","findings":"2023年发布的GPT-3.5-turbo复制了真实实验中的亲白人偏差（白人名字回电率高2.12个百分点，1%显著），而2024年及之后发布的所有模型要么无偏差，要么出现显著的反向偏差（亲黑人偏差达0.4-3.0个百分点）。性别轴上也呈现相同反转：GPT-3.5-turbo亲男性（1.92个百分点），后续模型亲女性或无偏差。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：该研究直接使用LLM模拟人类招聘决策，有真实人类数据作为对照基准，发现偏差方向随模型版本反转，符合研究者对经济学实验和政策评估场景以及批判性失效条件的关注。值得精读原文。","inspiration":"采用配对简历设计，仅改变名字暗示种族/性别，保持其他信息不变，以隔离身份信号对决策的影响，并设置温度0确保确定性输出｜可迁移到信贷审批歧视研究，检验LLM在贷款申请评估中是否对申请人姓名产生种族或性别偏差｜用LLM作为信贷审批员，处理仅名字（典型白人/黑人/男性/女性名）不同的标准化贷款申请，输出批准/拒绝决策，以真实银行信贷审批数据中的种族/性别差异作为对照基准"}},{"id":"2606.21820","version":1,"title":"Generating Public Health Responses using Survey-Augmented Large Language Models","zh_title":"使用调查增强的大语言模型生成公共卫生响应","abstract":"Epidemiological models often rely on survey data to represent how individuals make health-related decisions, such as whether to vaccinate or adopt protective behaviors. However, repeated large-scale surveys are costly, time-consuming, and limited in the range of scenarios they can capture. In this work, we investigate whether large language models (LLMs) can generate synthetic survey responses that reproduce patterns observed in real populations. Using longitudinal data from the FluPaths surveys, we first identify groups associated with broadly positive or negative attitudes toward vaccination through clustering analysis. We then evaluate several LLMs using a cluster-informed prompting approach to generate synthetic survey responses across multiple epidemic waves. Across models, the synthetic data generally reproduce the distributions of demographic characteristics, vaccination-related beliefs, risk perceptions, and health behaviors observed in the survey data. However, they are less successful at capturing how these factors vary together within respondents. Some models reproduce group-level vaccination trends more reliably than others, although performance varies across waves. We also trained a classifier to distinguish real from synthetic records and found that the generated responses remained identifiable as synthetic. Overall, our findings suggest that LLM-generated survey data may provide a useful tool for exploratory data augmentation and we hope that it could support agent-based epidemic modeling approaches. However, the generated data should not be treated as a substitute for human survey data without further methodological improvements and validation.","authors":["Leonardo Marciaga","Thuyen Pham","Julia Rezvani","Alina Hyk","Chunyang Liao","Konstantinos Mitsopoulos","Raffaele Vardavas"],"categories":["cs.SI","cs.AI","cs.CL"],"primary_category":"cs.SI","announce_type":"new","date":"2026-06-20","first_seen":"2026-06-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.21820","pdf_url":"https://arxiv.org/pdf/2606.21820","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","调查数据合成","公共卫生决策"],"reason":"用LLM生成合成调查回答复现人群健康行为，并与真实纵向调查数据对照，评估仿真可…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-20","rank":12,"question":"大语言模型能否生成合成调查数据，复现真实人群中观察到的公共卫生相关行为模式？","design":"使用FluPaths纵向调查数据，先通过聚类分析识别出对疫苗接种持积极或消极态度的群体；然后采用聚类信息提示方法，让多个LLM（如GPT等）生成跨多个流行病波的合成调查响应，测量人口统计特征、疫苗信念、风险感知和健康行为的分布及共变关系。","baseline":"FluPaths调查的8波纵向数据，来自概率抽样的美国生活面板（ALP），包含2016-2024年间的全国代表性受访者。","findings":"合成数据总体上复现了真实调查中人口统计特征、疫苗信念、风险感知和健康行为的分布，但在捕捉这些因素在个体内的共变方面效果较差；不同模型在复现群体层面疫苗接种趋势上表现不一，且生成的响应仍可被分类器识别为合成数据。","reliability":"论文承认合成数据在捕捉个体内多变量共变方面不足，且生成的数据仍可被区分，不能替代人类调查数据，需要进一步方法改进和验证。","relevance":"高度相关。该研究直接评估了LLM生成合成调查数据复现真实人类行为的可靠性，有真实数据对照，涉及公共卫生决策场景，并指出了失效条件（个体内共变、可识别性），符合研究者对仿真验证和批判性视角的关注。","inspiration":"该方法借鉴了聚类信息提示（cluster-informed prompting）来引导LLM生成特定子群体的合成调查响应，并利用多波纵向真实数据作为基准，评估合成数据在分布和个体内共变上的复现效果｜可迁移到消费者跨期选择研究中，例如模拟不同风险偏好或金融素养群体的储蓄与消费决策模式｜使用LLM生成合成面板数据，处理为提供聚类标签（如低/高金融素养）的提示，结果变量为各期消费-储蓄选择，以真实家庭金融调查（如PSID）的纵向数据作为对照基准"}},{"id":"2606.19904","version":1,"title":"Toward Temporal Realism in City-Scale Crisis Response Simulation using LLM Agents","zh_title":"面向城市规模危机响应模拟中时序逼真度的LLM智能体研究","abstract":"Human collective participation is rarely steady in time: it is bursty, with short episodes of intense activity separated by long quiet intervals. In crisis response and community mobilization, predicting when people act matters as much as predicting whether they act. Such settings are increasingly modeled with LLM-based social simulators, yet these simulators are validated on whether each action is individually plausible, not on whether actions are timed as in reality. Their temporal realism, the degree to which simulated activity reproduces the bursty, heavy-tailed timing of real human systems, thus remains untested. We examine this gap using a multi-year, city-scale log of offline volunteering in Shenzhen that spans the COVID-19 pandemic. Empirically, we establish that bursty timing is common at individual and tracked-group levels, that it is largely endogenous and self-exciting, and that it is amplified by the pandemic rather than produced by daily activity cycles. A standard LLM-only simulator reproduces almost none of this timing: its synchronous schedule has no self-excitation channel, so agents act on a near-regular clock. Guided by these findings, we build a simulator in which a data-calibrated self-excitation channel and a crisis-period regime decide when each agent acts and query the LLM only at those moments, leaving it to decide which task to join and whether to commit. The LLM-only baseline yields no bursty agents (median burstiness $B=-0.14$); a single data-calibrated gate is then sufficient to lift per-agent timing above the burst threshold (median $B\\approx0.37$) without degrading LLM content decisions. These results indicate that temporal realism in LLM-based crisis-response simulation is best achieved by decoupling when agents act, governed by an explicit self-excitation and crisis-activation mechanism, from what they do, governed by the LLM.","authors":["Anping Zhang","Yang Tan","Yuanbo Tang","Huaze Tang","Qiuhua Ye","Marta C. Gonzalez","Yang Li"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-06-18","first_seen":"2026-06-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.19904","pdf_url":"https://arxiv.org/pdf/2606.19904","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","人类行为对照","危机响应"],"reason":"用LLM agent模拟城市危机响应，有真实志愿者数据对照，评估时序逼真度并指…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":46,"question":"LLM智能体模拟城市危机响应时，能否复现真实人类志愿者参与行为的突发性时序模式？","design":"使用AgentSociety平台构建LLM多智能体模拟器，模拟深圳志愿者参与行为；对比标准同步LLM模拟器与注入数据校准的自激和危机激活机制的增强模拟器，测量智能体行动时间的突发性（B值）和决策质量。","baseline":"2020–2023年深圳660万条线下志愿服务记录，涵盖COVID-19疫情期间。","findings":"标准LLM模拟器无法产生突发性时序（中位B=-0.14），而加入数据校准的自激门控后，智能体时序突发性显著提升至中位B≈0.37，且不损害LLM的内容决策质量。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真中时序逼真度的缺失，用真实大规模志愿者数据作为基准，揭示了标准LLM模拟器在复现人类突发行为模式上的失效，并提出了解耦“何时行动”与“做什么”的改进方案，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得精读原文。","inspiration":"该方法将行动时序与决策内容解耦，并用真实数据校准智能体的行动触发概率，可作为仿真中引入时间维度的通用设计｜可迁移到金融市场中的投资者下单时机与决策质量研究，例如模拟政策公告后的交易行为｜以LLM智能体模拟散户投资者，处理为注入基于历史订单簿数据校准的自激与消息驱动下单概率，结果变量为下单时间的突发性（B值）和订单方向准确性，对照真实交易所逐笔委托数据"}},{"id":"2606.19336","version":2,"title":"Learning User Simulators with Turing Rewards","zh_title":"用图灵奖励学习用户模拟器","abstract":"Learning to simulate human users in interactive settings could advance the training of agent assistants, evaluation of personalization systems, research in the social sciences, and more. Existing approaches generally do so by training a large language model (LLM) to match a single ground truth response, either by maximizing the log probability or by using a similarity reward. We instead propose Turing-RL: a Turing-Test-based reinforcement learning approach for training user simulator models. Turing-RL uses a discriminative Turing reward with an LLM judge to score how indistinguishable a generated response is from the real user's given the user's history, and the user simulator LLM learns to produce responses indistinguishable from what the user could have said with such rewards. Across two different domains--conversational chat and Reddit forum discussion--we find that Turing-RL consistently outperforms baseline methods on both LLM and human evaluation metrics. Our study suggests that optimizing for indistinguishability, rather than response matching, is effective for learning user simulators.","authors":["Yingshan Susan Wang","Cedegao E. Zhang","Linlu Qiu","Zexue He","Pengyuan Li","Alex Pentland","Roger P. Levy","Yoon Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-17","first_seen":"2026-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.19336","pdf_url":"https://arxiv.org/pdf/2606.19336","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A4","B1"],"tags":["用户模拟","强化学习","人类数据对照"],"reason":"用RL训练用户模拟器，以人类真实对话为基准，优化不可区分性，方法可迁移至人类仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":34,"question":"如何通过强化学习训练用户模拟器，使其生成与真实用户不可区分的回复，而非仅匹配单一真实回复？","design":"使用Qwen3-8B作为基座模型，通过监督微调（SFT）和基于图灵测试判别奖励的强化学习（GRPO）训练用户模拟器；在多轮对话和Reddit论坛讨论两个领域，利用用户历史行为作为条件，以LLM裁判评估生成回复与真实用户的不可区分性作为奖励信号。","baseline":"PRISM Alignment数据集中1288名用户的多轮对话数据，以及ConvoKit Reddit语料中14个子版块的1282名用户讨论数据，均包含真实用户回复作为对照。","findings":"Turing-RL在两个领域上均一致优于基于回复相似度奖励和基于对数概率最大化的基线方法，在LLM和人类评估指标上均表现更好。优化不可区分性而非回复匹配是学习用户模拟器的有效路径。","reliability":"论文未讨论","relevance":"该研究直接针对用LLM模拟人类用户，以真实人类对话数据为基准，优化不可区分性，方法可迁移至经济学实验和政策评估等场景，且提供了批判性视角（指出匹配单一回复的局限），值得精读原文。","inspiration":"该方法用图灵测试判别奖励训练模拟器，以不可区分性而非回复匹配为目标，可借鉴其将人类判断作为奖励信号的强化学习设计｜可迁移到消费者偏好调查或政策沟通模拟，用LLM生成消费者对新产品属性的评价或公众对政策公告的反应｜以LLM模拟消费者作为被试，处理为不同产品描述或政策措辞，结果变量为模拟消费者的选择或态度分布，用真实消费者调查或政策反馈数据做对照"}},{"id":"2606.17657","version":1,"title":"Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games","zh_title":"利用认知模型改进语言模型对人类说服博弈的仿真","abstract":"People make decisions differently in strategic interactions. Some update beliefs like a Bayesian; others exhibit biases like motivated reasoning. Although creators of large language models use simulated humans for safety evaluations and training, they often fail to cover this breadth of human behavior. We argue that cognitive science and economics provide a convenient tool for doing so, making use of mathematical models of human decision-making. We propose an approach that we call Equation-to-Behavior Prompting for guiding large language models to match cognitive models, and evaluate this approach on persuasion games based on legal decision-making. We find that large models can approximate equation-based specifications -- Bayesian updating, affine distortion, motivated updating, and Grether's $α$-$β$ model -- using prompting, but small models fail to do so. However, training small models with reinforcement learning to adhere to mathematical rules, Equation-to-Behavior RL, reduces belief error by 26.5% in out-of-distribution parameterizations. We show that these simulations can help create diverse training environments; training small models to consider different kinds of decision-makers improves average belief change by 2.5%--12% over Bayesian-only training, even when persuading GPT-5-mini. Our work could improve human simulations for training and evaluation in increasingly realistic settings, and could also enable novel research into more complicated mathematical models of human decision-making.","authors":["Zirui Cheng","Zeyu Shen","Thomas L. Griffiths","Peter Henderson"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-16","first_seen":"2026-06-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.17657","pdf_url":"https://arxiv.org/pdf/2606.17657","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","认知模型","说服博弈"],"reason":"用LLM模拟人类在说服博弈中的决策，并与认知模型对照，涉及法律决策场景，有真实…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":10,"question":"如何利用认知模型改进大语言模型对人类在说服博弈中决策行为的仿真？","design":"使用认知模型（贝叶斯更新、仿射扭曲、动机性更新、Grether α-β模型）通过方程到行为提示（Equation-to-Behavior Prompting）引导LLM扮演接收者，在基于法律决策的说服博弈中测量信念更新误差；对小型模型还采用强化学习训练（Equation-to-Behavior RL）以遵循数学规则。","baseline":"无直接对照的真实人类数据，但使用基于Old Bailey审判记录构建的证据数据集作为真实场景基础。","findings":"大型语言模型可通过提示近似多种认知模型的信念更新规则，而小型模型难以做到；对小型模型进行强化学习训练可使信念误差降低26.5%，且训练时考虑多样化决策者能提升模型说服性能2.5%-12%。","reliability":"论文未讨论","relevance":"高度相关：该研究直接用LLM仿真人类在说服博弈中的决策，并与认知模型对照，涉及法律决策场景，且探讨了仿真可靠性与小型模型的失效条件，符合研究者对基准对照和批判性分析的兴趣。","inspiration":"该方法通过认知模型方程生成行为提示来引导LLM仿真，可借鉴其将结构化决策规则注入LLM以控制仿真行为｜可迁移到政策公告的预期形成实验，如央行沟通如何影响公众通胀预期｜以LLM为被试，处理为不同沟通策略（如模糊vs精确指引），结果变量为预期调整幅度与偏差，对照真实调查数据（如密歇根消费者调查）"}},{"id":"2606.17165","version":3,"title":"Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference","zh_title":"基于大语言模型的A/B测试统计基础：面向人类因果推断的替代框架","abstract":"Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost. We study when a treatment effect estimated on LLM outcomes can recover the effect for the human population of interest. Distributional equivalence between LLM and human outcomes would make any standard estimator valid but is unrealistic. We therefore develop a statistical framework that adapts surrogate endpoint theory to LLMs, showing that calibrating LLM outcomes to human outcomes identifies the average treatment effect under surrogacy and comparability conditions that are jointly weaker than distributional equivalence. We present a falsification test for surrogacy and a bound on the worst-case bias from limited overlap between the LLM and human samples. We further show that the stochasticity inherent to LLMs can weaken surrogacy for identification while also introducing bias and variance during estimation, but that using an average over multiple LLM draws per unit as the surrogate mitigates these issues. Simulations validate the results, and an empirical application to the Upworthy Research Archive dataset shows that raw LLM outputs recover only 39% of the human treatment effect while nonparametric calibration closes the gap. A central takeaway is that A/B testing on LLM responses is correct only by assumption, whereas A/B testing on humans is correct by design, and that the required assumptions are hardest to justify precisely where LLMs promise the greatest benefit. We discuss the choice of LLM, prompting, and temperature as design variables, the compounded challenge posed by long-term outcomes, and how to size human pilot studies for validation.","authors":["Joel Persson","Mårten Schultzberg","Sebastian Ankargren"],"categories":["stat.ME","cs.AI","econ.EM","math.ST"],"primary_category":"stat.ME","announce_type":"new","date":"2026-06-15","first_seen":"2026-06-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.17165","pdf_url":"https://arxiv.org/pdf/2606.17165","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","A/B测试","因果推断"],"reason":"直接研究用LLM替代人类进行A/B测试，提出替代指标框架，有真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-15","rank":1,"question":"在什么条件下，基于LLM的A/B测试能够恢复人类群体的平均处理效应？","design":"论文提出一个统计框架，将LLM生成的结果视为人类结果的替代终点，通过校准LLM结果到人类结果来识别平均处理效应。框架包括替代性和可比性条件，并提供了对替代性的伪造检验和有限重叠下的最坏偏差界限。实证应用使用Upworthy Research Archive数据集，比较原始LLM输出与非参数校准后的效果。","baseline":"Upworthy Research Archive数据集中的真实人类实验结果。","findings":"原始LLM输出仅恢复人类处理效应的39%，而非参数校准可以弥合这一差距。核心结论是：基于LLM的A/B测试的正确性依赖于假设，而基于人类的A/B测试的正确性依赖于设计，且所需的假设在LLM承诺最大收益的地方最难证明。","reliability":"论文指出，LLM的随机性会削弱替代性，并引入偏差和方差；替代性和可比性条件共同弱于分布等价性但仍需验证；长期结果构成复合挑战；LLM的选择、提示和温度作为设计变量会影响结果。","relevance":"直接命中研究者的核心兴趣：研究LLM替代人类被试的A/B测试，有真实数据对照，并讨论了失效条件（如替代性不成立、有限重叠、随机性影响），值得精读原文以获取统计框架和实证细节。","inspiration":"该方法将LLM输出作为替代终点，通过非参数校准恢复人类处理效应，并提供替代性伪造检验与有限重叠下的偏差界限，值得借鉴其校准与稳健性检验设计｜可迁移到消费者对金融产品信息披露政策的反应评估，例如研究简化披露标签对投资选择的影响｜以LLM模拟投资者，处理为展示简化版基金费用披露，结果变量为选择高费用基金的概率，用真实投资者实验数据作为校准与对照基准"}},{"id":"2606.08853","version":1,"title":"AI-Assisted Variance Reduction in Randomized Experiments","zh_title":"随机实验中AI辅助的方差缩减","abstract":"Generative AI and large language models can produce realistic predictions of human behavior from rich, unstructured inputs with little to no task-specific training data. Recent work uses these ``digital twin'' predictions to supplement human responses in surveys and experiments. We study the special case of using AI-generated predictions to reduce variance in randomized experiments. We argue that doing so requires no new estimators and that researchers can simply include AI predictions as covariates in standard regression adjustment, analogous to adjusting for a prognostic score. A benefit of this approach is a ``do no harm'' property whereby the adjusted estimator reverts to the unadjusted difference in means when predictions are uninformative. Other methods, such as variants of prediction-powered inference, do not have this guarantee. We provide implementation guidance, including how to obtain continuous scores from discrete LLM outputs and how to use LLMs to featurize unstructured inputs as auxiliary covariates. We demonstrate these ideas in simulations and three empirical applications: a survey mega-study, an email marketing A/B test, and a large-scale technology platform experiment. Overall, efficiency gains are real if modest, with greater benefits in studies that contain substantial text and other unstructured data. We also confirm the do no harm property empirically. Given these gains and limited costs, we recommend adjusting for AI-generated predictions as a regular empirical practice.","authors":["David Arbour","Eli Ben-Michael","Avi Feller","Apoorva Lal","Lo-Hua Yuan"],"categories":["econ.EM","stat.ME"],"primary_category":"econ.EM","announce_type":"new","date":"2026-06-07","first_seen":"2026-06-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.08853","pdf_url":"https://arxiv.org/pdf/2606.08853","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","实验设计","方差缩减"],"reason":"用LLM预测替代人类响应以降低实验方差，有真实人类对照，涉及A/B测试和政策评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-07","rank":6,"question":"如何利用LLM生成的预测作为协变量来降低随机实验的方差？","design":"论文提出将LLM预测作为协变量纳入标准回归调整，无需新估计量；通过模拟和三个实证应用（调查大型研究、邮件营销A/B测试、大规模技术平台实验）验证效果。","baseline":"有真实人类数据对照：调查大型研究使用数字孪生预测与人类回答对比，邮件营销A/B测试有真实用户响应，技术平台实验有真实用户行为数据。","findings":"纳入LLM预测的回归调整能实现方差降低，增益虽小但真实，尤其在包含大量文本和非结构化数据的研究中效果更明显。该方法具有“无害”属性：当预测无信息时，估计量退化为未调整的均值差。","reliability":"论文承认效率增益取决于预测质量，当预测质量不高时增益有限；但回归调整方法本身不会引入偏差或增大方差，其他方法如PPI可能失效。","relevance":"高度相关：直接涉及用LLM预测辅助人类实验，有真实数据对照，涵盖经济学实验场景，且批判性地讨论了失效条件（预测质量差时增益有限），值得精读。","inspiration":"该方法将LLM预测作为协变量纳入回归调整，无需新估计量，实现方差降低且具有“无害”属性，值得借鉴其利用AI辅助提升实验效率的思路｜可迁移到政策评估中的随机对照试验（如就业培训、税收激励），利用LLM对个体特征的预测作为协变量，提高处理效应估计精度｜以就业培训实验为例，被试为求职者，处理为提供培训，结果变量为就业状态，用LLM基于简历文本预测就业概率作为协变量，与真实就业数据对照，评估方差降低效果"}},{"id":"2606.05330","version":1,"title":"A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing","zh_title":"基于概率信念追踪的多轮人类可说服性模型","abstract":"Large language models can shift human beliefs across high-stakes domains, but most persuasion studies rely on pre/post belief change. These endpoint measures identify whether persuasion occurred, yet miss where and how beliefs moved within a dialogue. We present PERSUASIONTRACE, a framework for studying persuasion in human-LLM interaction. Built on a web-based experimental platform, PERSUASIONTRACE contributes a tool for multi-turn persuasion studies and a process-level evaluation protocol: it records multi-turn belief reports from human or simulated targets of persuasion, annotates persuader turns with rhetorical dimensions (logos/pathos/ethos), and evaluates simulators by fidelity to real human belief dynamics. Using this framework, we find that human targets group into two clusters of multi-turn belief updates and exhibit susceptibility to rhetorical strategies, and that LLMs are persuasive across generic and personalized topics, text and audio modalities, and multi-turn interactions. Prior work has chiefly used vanilla-prompted LLMs to simulate human targets, but we show that these simulators fail to replicate human belief dynamics. We introduce a Bayesian-network simulated target that maintains an explicit latent belief state over time so each persuader message yields cognitively realistic belief updates. In human-likeness evaluation, our Bayesian target scores near a human reference (81 vs 80), while baseline LLM targets score substantially lower (64). PERSUASIONTRACE reframes persuasion evaluation from endpoint movement alone to process fidelity, providing a stronger basis for scientific analysis and safer optimization of persuasive systems.","authors":["Jared Moore","Noah Goodman","Nick Haber","Max Kleiman-Weiner"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-03","first_seen":"2026-06-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.05330","pdf_url":"https://arxiv.org/pdf/2606.05330","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B3","B4"],"tags":["LLM仿真","信念动态","人类数据对照"],"reason":"用LLM模拟人类信念动态，并与真实人类数据对照，评估仿真保真度，提出改进方法。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":12,"question":"在多轮人-LLM说服对话中，人类信念如何随时间动态更新，以及如何构建能忠实复现人类信念轨迹的仿真模型？","design":"使用基于Web的实验平台记录人类被试在多轮说服对话中的逐轮信念评分，并对说服者话语标注修辞维度（logos/pathos/ethos）；同时提出基于贝叶斯网络的仿真目标模型，该模型维护显式潜在信念状态，根据每条说服消息进行认知上现实的信念更新，并与真实人类信念动态进行保真度比较。","baseline":"真实人类被试在多轮说服对话中的逐轮信念报告轨迹，以及基于该轨迹统计的人类参考分数（80分）。","findings":"人类信念更新轨迹可分为两类主要模式，且对修辞策略的敏感性存在异质性；普通提示的LLM仿真目标无法复现人类信念动态，而贝叶斯网络仿真目标在人类相似度评估中接近人类参考水平（81 vs 80），显著优于基线LLM目标（64）。","reliability":"论文承认当前轨迹聚类主要受整体移动幅度驱动，需要更大数据集才能可靠区分轮内动态的细微差异；修辞分析仅为探索性且样本量有限，仅发现ethos与说服变化有可靠负相关，logos和pathos效应不显著；仿真器选择会实质性影响表面说服者质量评估和政策排名，提示若仿真器不忠实于人类，可能系统性地偏好错误策略。","relevance":"该研究直接使用LLM进行人类仿真实验，以真实人类信念轨迹为基准，评估仿真保真度并指出普通LLM仿真失效的条件，同时提出改进的贝叶斯网络仿真方法，高度契合研究者对经济学实验和政策评估场景下仿真可靠性与偏差的关注，值得精读原文。","inspiration":"该方法通过显式建模潜在信念状态和认知上现实的更新规则来仿真人类动态，值得借鉴其将心理过程结构化并逐轮拟合真实轨迹的做法｜可迁移到政策公告的预期形成研究，例如央行沟通如何影响公众通胀预期｜以LLM作为被试，施加不同修辞风格的政策声明作为处理，逐轮测量预期通胀值，并以真实调查数据（如密歇根消费者调查）的预期调整轨迹作为对照基准"}},{"id":"2606.04978","version":1,"title":"Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game","zh_title":"探究大语言模型风险决策中的结果层相似与机制层对齐：来自圣彼得堡博弈的证据","abstract":"LLMs can appear cautious in risk decision-making tasks, yet cautious-looking outputs do not necessarily indicate alignment with human decision-making mechanisms. We investigate this distinction using the St. Petersburg game as a controlled testbed, a classical paradox in which the expected payoff is infinite, yet humans typically report low, finite willingness to pay. We evaluate 28 LLMs with a structured prompt suite that includes the original game; controlled decision variants that perturb truncation, repeated play, numeric endowment, and occupational identity; a human-perspective prompt that asks models to reason as human decision makers; and paired comparisons between base models and their instruction-tuned counterparts. In the original game, most models generate finite bids, creating the appearance of human-like risk behavior. However, this outcome-level resemblance masks substantial mechanism-level differences. The controlled variants reveal that rather than maintaining human-like behavior seen in the original game, models often shift to conditionally and computationally rational behavior. Human-cue prompting and instruction tuning often lower bids and reduce some visible pathologies, but most mechanism-level response patterns remain largely unchanged. These findings show that behavioral alignment in risk decision-making can be surface-level: LLMs may produce human-like risk decisions without exhibiting human-consistent mechanisms. High-stakes evaluations of LLM decision-making should therefore move beyond outcome similarity and examine whether the alignment is supported by mechanism-level consistency.","authors":["Chensong Huang","Changyu Chen","Chenwei Lin","Hanjia Lyu","Xian Xu","Jiebo Luo"],"categories":["cs.CL","cs.CY","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-03","first_seen":"2026-06-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.04978","pdf_url":"https://arxiv.org/pdf/2606.04978","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","风险决策","机制对齐"],"reason":"用LLM仿真人类风险决策，与真实人类数据对照，揭示表面相似下的机制差异，批判性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-03","rank":7,"question":"LLM在风险决策中产生的人类相似输出是否反映机制层面的对齐，还是仅表面相似？","design":"以圣彼得堡悖论为测试床，对28个LLM使用结构化提示套件，包括原始游戏、四种机制探针（截断、重复游戏、数值禀赋、职业身份）、人类视角提示，以及基础模型与指令微调版本的配对比较。","baseline":"人类在圣彼得堡游戏中通常报告低且有限的支付意愿。","findings":"大多数LLM在原始游戏中产生有限出价，看似人类风险行为；但机制探针显示模型转向条件性和计算理性行为，而非保持人类一致性。人类提示和指令微调降低出价并减少明显病理，但机制层面响应模式基本不变。","reliability":"论文指出表面相似性可能掩盖机制差异，高利害评估需超越结果相似性检查机制一致性。未明确讨论失效条件。","relevance":"高度相关：直接研究LLM仿真人类风险决策，有真实人类对照，并批判表面相似性，符合研究者对可靠性与偏差的关注。","inspiration":"该研究通过机制探针（如截断、重复游戏、禀赋变化）检验LLM行为背后的决策过程，而非仅看结果相似性，值得借鉴｜可迁移到资产定价实验中的风险偏好测量，检验LLM是否真正反映人类的风险厌恶机制｜以LLM为被试，施加财富禀赋变化和投资期限截断处理，测量其出价行为，并与真实人类实验数据（如Binswanger的彩票选择实验）对照"}},{"id":"2606.03030","version":1,"title":"Do Matching Mechanisms Work with LLM Agents?","zh_title":"匹配机制在LLM代理市场中是否有效？","abstract":"This study examines whether standard matching mechanisms function as intended in LLM-agent markets, where LLM agents make allocation-related decisions as delegated decision-makers. We compare decentralized free-negotiation markets with centralized mechanism-based markets including several representative mechanisms. Across controlled one-to-one matching environments, mechanism-based markets generally outperform free negotiation in terms of stability and efficiency. We also find that LLM agents report preferences truthfully at substantially higher rates than human subjects in comparable DA and EADA environments. However, truth-telling is not uniformly aligned with formal strategy-proofness across all mechanisms: TTC, despite being strategy-proof, does not always elicit higher truth-telling than EADA. These results suggest that matching theory provides a useful but incomplete guide for designing institutions in LLM-agent markets.","authors":["Yukihiro Hoshino","Ayato Kitadai","Nariaki Nishino"],"categories":["cs.GT","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2026-06-02","first_seen":"2026-06-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.03030","pdf_url":"https://arxiv.org/pdf/2606.03030","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A3","A1","B1","B2"],"tags":["LLM代理","市场匹配","人类行为对照"],"reason":"用LLM代理模拟匹配市场并与人类实验数据对照，直接复现人类决策行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"标准匹配机制在由LLM代理决策的市场中是否仍能按预期发挥作用？","design":"用GPT-4o等LLM作为代理，模拟一对一匹配市场中的决策者，比较去中心化自由协商市场与集中式机制市场（包括DA、EADA、TTC等代表性机制），测量匹配的稳定性、效率以及代理的偏好真实报告率。","baseline":"与已有文献中人类被试在DA和EADA环境下的真实报告率进行对比。","findings":"机制市场在稳定性和效率上普遍优于自由协商；LLM代理在DA和EADA下的真实报告率显著高于人类被试，但策略防护性并不总能预测真实报告行为，例如TTC虽为策略防护机制，其真实报告率并不总是高于EADA。","reliability":"论文指出匹配理论为LLM代理市场制度设计提供了有用但不完整的指导，真实报告行为与形式上的策略防护性并不完全一致，提示仅凭理论性质不足以预测LLM代理的行为。","relevance":"该研究直接以LLM代理模拟人类在匹配市场中的决策，并与真实人类实验数据对照，评估仿真可靠性，高度契合研究者对LLM人类仿真实验的关注，值得阅读原文。","inspiration":"借鉴其将LLM代理置于不同市场机制下比较行为的方法，可迁移至经济政策评估场景，如研究不同拍卖机制或税收政策下LLM代理的遵从与策略行为。｜可应用于劳动力市场匹配或学校选择等经济政策评估，检验机制设计在AI代理参与下的有效性。｜以LLM代理作为被试，随机分配至不同匹配机制（如DA、波士顿机制），测量匹配效率与偏好真实报告率，并与已有的人类实验数据（如学校选择实验）进行对照。"}},{"id":"2606.02741","version":1,"title":"Greener Than Humans? Environmental Attitudes in Large Language Models","zh_title":"比人类更环保？大语言模型中的环境态度","abstract":"Large language models (LLMs) are increasingly used in sustainability-related decision support, reporting, and public communication, yet little systematic evidence exists on the environmental attitudes embedded in their outputs. This paper develops a benchmark for evaluating environmental cognition, affect, and behavioural recommendations in LLMs and applies it to 31 widely used proprietary and open-weight models. Drawing on questions from established environmental awareness surveys and additional sustainability-related behavioural measures, we compare LLM responses 1) among models and 2) between models and human survey benchmarks from Germany. We assess their robustness across prompting conditions. We find that many LLMs align more closely with environmentally progressive attitudes than the average survey respondent, exhibiting higher levels of environmental affect and cognition and recommending behaviours associated with substantial potential CO2 reductions. At the same time, we observe no systematic relationship between sustainability-oriented responses and model origin, size, or release context. However, models exhibit contextual sensitivity, controlled by persona-based prompting and show sycophantic shifts mirroring user-specified ideological positions, which raises concerns about steerability and normative reliability in real-world deployments. Our findings provide a reusable evaluation framework for assessing sustainability-related value alignment in LLMs and highlight the importance of governance, transparency, and critical oversight as AI systems become increasingly embedded in sustainability transformations and public decision-making.","authors":["Stefanie Kunkel","Tilman Hartwig","Marcus Voss","Emma K. Schütt","Angelika Gellrich"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-01","first_seen":"2026-06-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.02741","pdf_url":"https://arxiv.org/pdf/2606.02741","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","环境态度","人类数据对照"],"reason":"用LLM复现人类环境态度调查并与真实数据对照，评估仿真可靠性及偏差，属核心相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":41,"question":"不同大语言模型在环境认知、情感和行为建议上的回答与德国人口平均水平及彼此之间有何差异？","design":"用31个主流LLM回答来自德国联邦环境署环境意识调查（UBS）的问题及额外可持续行为测量题，比较模型间差异，并评估提示条件（如角色扮演、意识形态立场）对回答稳健性的影响。","baseline":"德国联邦环境署环境意识调查（UBS）的纵向调查数据，提供德国人口平均水平基准。","findings":"多数LLM比普通受访者更倾向于进步环保态度，表现出更高的环境情感和认知，并推荐减排潜力大的行为；但模型回答受角色提示和用户意识形态立场影响，表现出迎合性偏移，且环保倾向与模型来源、规模或发布背景无系统关联。","reliability":"模型回答受提示条件影响显著，存在迎合用户意识形态的倾向，在真实部署中可操纵性和规范性可靠性存疑；研究仅基于德国背景，跨文化泛化性未验证。","relevance":"该研究直接用LLM复现人类环境态度调查并与真实人群数据对照，评估仿真偏差和可靠性，完全契合研究者对LLM人类仿真实验及批判性评估的关注，值得精读。","inspiration":"借鉴其用LLM复现真实人口调查并直接与纵向调查基准对照的仿真验证设计，以及通过角色扮演和意识形态提示检验回答稳健性的方法｜可迁移到消费者通胀预期形成或政策沟通效果评估，例如研究不同信息框架下公众对央行前瞻指引的反应｜以多个LLM作为被试，施加不同政治立场或信息源提示（如‘你是保守派投资者’或‘你刚读完鸽派新闻’）作为处理，结果变量为模型生成的通胀预测值，与密歇根大学消费者调查的真实通胀预期数据做对照"}},{"id":"2605.30036","version":2,"title":"Teaching Values to Machines: Simulating Human-Like Behavior in LLMs","zh_title":"向机器传授价值观：在LLM中模拟类人行为","abstract":"Large Language Models (LLMs) demonstrate a remarkable capacity to adopt different personas and roles; however, it remains unclear whether they can manifest behavior that adheres to a coherent, human-like value structure. In this work, we draw on established psychological value theory to induce human-like values in LLMs and assess their alignment with patterns observed in human studies. Using validated psychological questionnaires, we conduct large-scale experiments -- over 5 million questions -- to evaluate value structures and value-behavior relationships in leading LLMs and compare them to humans. Our findings reveal strong agreement between value-prompted LLMs and humans across both dimensions. Moreover, incorporating human value distributions enhances population-level simulations with value-induced LLMs. These findings highlight the potential of value-induced LLMs as effective, psychologically grounded tools for simulating human behavior.","authors":["Asaf Yehudai","Naama Rozen","Ariel Gera"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-28","first_seen":"2026-05-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.30036","pdf_url":"https://arxiv.org/pdf/2605.30036","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM仿真","价值观诱导","人类数据对照"],"reason":"用LLM模拟人类价值观行为，并与真实人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"能否通过心理价值理论诱导大语言模型表现出与人类一致的价值结构和价值-行为关系？","design":"使用经过验证的心理问卷，对主流大语言模型进行超过500万次的大规模实验，通过价值提示诱导模型扮演特定价值观人群，测量其价值结构和价值-行为关系。","baseline":"人类在相同心理问卷上的真实回答数据，用于对比价值结构和价值-行为关系。","findings":"价值提示后的LLM在价值结构和价值-行为关系上与人类高度一致；引入人类价值分布能提升基于价值诱导LLM的群体层面模拟效果。","reliability":"论文未讨论","relevance":"该研究直接探索用LLM模拟人类价值观行为，并与真实人类数据对照，评估仿真可靠性，高度契合您关注的LLM人类仿真实验方向，值得阅读原文。","inspiration":"该方法通过价值提示诱导LLM扮演特定价值观人群，并与真实人类问卷数据对照，可借鉴用于经济实验中施加偏好或信念处理｜可迁移到消费者跨期选择实验，模拟不同时间偏好群体的储蓄或消费决策｜以LLM为被试，用时间偏好提示（如耐心/冲动）作为处理，测量其在跨期选择任务中的折现率，与真实人群的问卷或实验数据对照"}},{"id":"2605.26437","version":1,"title":"Divergent Minds, Convergent Baselines: A Bounded-Rationality Account of LLM-Human Strategic Behaviour","zh_title":"分歧思维，收敛基线：LLM与人类战略行为的有界理性解释","abstract":"Researchers have started using LLM agents in place of human subjects in behavioural and political-science experiments, often as a cheaper substitute for laboratory pools. The substitution does not hold up in strategic settings: humans and LLMs reliably make different choices, and neither fine-tuning on human response data nor persona conditioning has closed the gap. The behavioural-economics literature has, since Simon's introduction of bounded rationality, modelled human strategic behaviour as a classical baseline plus an additive correction term $δ$. The framework proposed here reads $δ$ as the mathematical signature of bounded computation: the gap between what an unboundedly-rational agent would compute and what a computationally bounded agent actually produces. For canonical games whose solutions are present in standard training corpora, LLMs retrieve and recombine corpus material, bypassing the bound that produces $δ$ in humans. The framing extends to reasoning-distilled models through cognitive-hierarchy theory: their accessible level-$k$ strategic reasoning is bounded by compute budget and context length rather than by the cognitive constraints that bound humans, and the $δ$ they produce, if any, carries different structural signatures. Four operational tests (conditional dependence, distributional asymmetry, path-dependence under repetition, and paraphrase-robustness) are proposed to discriminate human-shaped $δ$ from LLM-shaped $δ$. A moderator prediction is that $|δ|$ scales with peer-signal individuation in the decision environment, with a quantitative bound of Cohen's $d \\geq 0.5$ between named-opponent and aggregate-opponent settings.","authors":["Po Han Teo"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-05-26","first_seen":"2026-05-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.26437","pdf_url":"https://arxiv.org/pdf/2605.26437","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","有界理性","行为博弈"],"reason":"直接研究LLM替代人类被试的战略行为差异，提出有界理性框架，含真实人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-05-26","rank":8,"question":"LLM能否替代人类被试在策略博弈中复现人类行为，以及两者偏差的数学本质是什么？","design":"提出理论框架，将人类策略行为建模为经典理性基线加有界计算修正项δ，并针对LLM提出四个操作化检验（条件依赖、分布不对称、重复路径依赖、释义鲁棒性）来区分人类δ与LLMδ。","baseline":"有真实人类数据作为对照基准，但具体数据来源未在摘要中说明。","findings":"人类与LLM在策略博弈中可靠地做出不同选择，微调或角色条件无法弥合差距；LLM通过检索训练语料绕过产生人类δ的计算边界，其δ具有不同结构特征。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接研究LLM替代人类被试在策略博弈中的行为，有真实人类数据对照，并批判性分析偏差来源，符合研究者对经济学实验和政策评估场景的关注。","inspiration":"借鉴其将行为偏差分解为理性基线加有界计算修正项δ的理论框架，并设计条件依赖、分布不对称等操作化检验来区分人类与LLM的决策结构｜可迁移到资产定价实验中的泡沫形成与理性预期偏离研究，检验LLM能否复现人类交易者的非理性繁荣｜以LLM为被试模拟连续双向拍卖市场，处理为不同信息透明度条件，结果变量为价格偏离基础价值的程度，对照真实人类实验数据（如Smith et al. 1988的泡沫实验）"}},{"id":"2605.25680","version":1,"title":"Simulating Human Memory with Language Models","zh_title":"用语言模型模拟人类记忆","abstract":"Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic memory experiments from psychology on both humans and language models. Across tasks, we find that out-of-the-box language models exhibit better memory than humans, even when prompted to imitate human behavior. We then show that better prompting strategies and the use of a compactor can cause language models to forget content in a more human-like way. Using these methods, we show preliminary evidence that language models with human-like memory constraints can function as more effective user simulators in a downstream education task. Finally, we release human reference data and benchmarks to support future work on simulating human memory with language models.","authors":["Qihan Wang","Nicholas Tomlin","Michael Hu","Brian Dillon","Tal Linzen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-25","first_seen":"2026-05-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.25680","pdf_url":"https://arxiv.org/pdf/2605.25680","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["人类仿真","记忆实验","可靠性评估"],"reason":"用LLM复现人类记忆实验，有真实人类数据对照，评估仿真可靠性并指出失效条件，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"如何让语言模型模拟人类记忆的局限性，使其作为用户仿真器更真实？","design":"用多种语言模型（如GPT-5.4、Claude Opus 4.6等）扮演人类被试，通过不同提示策略（任务提示、人类提示、记忆提示）和添加工作记忆瓶颈的Compactor代理，复现经典心理学记忆实验，测量模型在记忆任务上的得分分布。","baseline":"从人类参与者收集的真实记忆实验数据，包括数字广度、地图记忆等任务。","findings":"开箱即用的语言模型在所有记忆任务上表现远超人类，即使提示其模仿人类行为也无效；通过提示模型将上下文总结为四个组块并仅基于这些组块作答，可使模型遗忘模式更接近人类。","reliability":"论文承认Compactor方法虽使得分更接近人类，但遗忘模式仍不完全类人，且在教育任务中即使最类人的模型预测效果也远非完美，表明仿真仍有很大改进空间。","relevance":"该研究直接针对LLM人类仿真，用真实人类数据对照，评估记忆仿真的可靠性并指出失效条件，与研究者关注的经济学实验和政策评估场景高度相关，值得精读原文。","inspiration":"该研究通过提示策略（如Compactor代理施加工作记忆瓶颈）操纵LLM的认知局限，使仿真行为更接近人类，这种‘认知约束注入’方法值得借鉴｜可迁移到消费者跨期选择实验中，模拟有限注意力或记忆衰退对贴现行为的影响｜以LLM为被试，处理组施加记忆组块限制提示，对照组无限制，结果变量为跨期选择中的贴现率，用真实人类实验数据（如Andersen et al., 2008）作为基准对照"}},{"id":"2605.23783","version":2,"title":"Benchmarking LLMs for Community Governance Simulation with Life-history Narratives","zh_title":"基于生活史叙事的社区治理仿真大语言模型基准测试","abstract":"Effective community governance hinges on understanding what specific residents think and need. Recent work has used large language models (LLMs) to simulate human respondents, offering a scalable, reproducible way to study human attitudes and behaviors at low cost. However, these studies typically prompt the model with just a few demographic variables (age, gender, income), simulating only general role types. This is insufficient for community governance, where decisions depend on the views of specific residents. We bridge this gap with an integrated research framework covering dataset, benchmark, algorithm, and system. The dataset comprises approximately 1.2 million characters of first-person narrative collected through two-hour semi-structured interviews with each of 92 residents in an urban community, organized around nine community-governance domains. The benchmark probes 18 mainstream LLMs across four prompting strategies and shows that adding rich life-history profiles meaningfully raises fidelity above the no-profile baseline, but this gain comes with more input tokens per call from the longer prompts they require. The algorithm, curriculum-LoRA, is a parameter-efficient personalization framework that, by closing this fidelity-cost gap, matches the strongest baseline's fidelity at roughly 10x lower per-call cost and Pareto-dominates every configuration tested. The system integrates curriculum-LoRA into a closed-loop policy-evaluation pipeline. Together, these results bring individual-level LLM-based resident simulation within reach of resource-constrained local administrations, enabling community-governance decisions to be systematically pre-evaluated in silico before real-world deployment.","authors":["Xu Chen","Yuanzi Li","Lei Wang","Nan Lu","Yang Wang","Anding Wang","Lei Shi","Xiaoxing Fu","Ji-Rong Wen"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-22","first_seen":"2026-05-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.23783","pdf_url":"https://arxiv.org/pdf/2605.23783","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","社区治理","政策评估"],"reason":"用LLM仿真特定居民态度，有真实访谈数据对照，用于社区治理政策评估，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"如何利用居民生活史叙事提升大语言模型在社区治理仿真中的个体级保真度，并解决保真度与调用成本之间的权衡？","design":"收集92位城市社区居民的约两小时半结构化访谈，形成约120万字的第一人称生活史叙事及50题政策态度记录；用18个主流LLM搭配四种提示策略（零样本、生活史、无生活史少样本、生活史增强少样本）进行仿真，测量模型回答与真实居民态度的一致性；提出curriculum-LoRA个性化微调算法以降低高保真仿真的成本。","baseline":"92位居民的真实访谈回答及结构化政策态度记录，作为个体级仿真保真度的对照基准。","findings":"添加丰富生活史叙事能显著提升仿真保真度，但最佳提示策略的准确率仅约50%，且每次调用成本比中等规模模型高约一个数量级；curriculum-LoRA能以约十分之一的成本匹配最强基线保真度，在所有测试配置中实现帕累托占优。","reliability":"论文指出纯提示方法存在保真度-成本权衡尖锐不利，且最佳配置准确率仅约50%，暗示在复杂社区治理态度模拟中仍存在较大误差；未深入讨论生活史叙事可能引入的隐私、代表性偏差及模型幻觉等问题。","relevance":"该研究直接命中研究者关注的核心：用LLM仿真特定居民态度，有真实个体级访谈数据对照，用于社区治理政策预评估，并系统探讨了仿真可靠性与成本约束，值得精读以获取数据集构建、基准评测和个性化微调的方法细节。","inspiration":"可借鉴其用长文本生活史叙事替代简单人口统计变量来构建高保真个体仿真的思路，以及通过参数高效微调平衡保真度与成本的方法｜可迁移到消费者金融行为仿真，如预测不同背景家庭对信贷产品、保险政策或退休规划的态度与选择｜以真实家庭金融调查数据（如CFPS）为基准，用受访者的详细财务生活史微调LLM，处理为不同信贷条款或政策情景，测量模型生成的借贷意愿、风险偏好等，与真实调查回答对比评估仿真效度。"}},{"id":"2605.22095","version":1,"title":"Not Yet: Humans Outperform LLMs in a Colonel Blotto Tournament","zh_title":"尚未：人类在Colonel Blotto锦标赛中胜过LLM","abstract":"The emergence of large language models (LLMs) has spurred economists to study how humans and LLMs behave in strategic settings. We organized a series of round-robin tournaments in the Colonel Blotto game. This game attracts game theorists' attention due to high-dimensional action space and the absence of pure strategy Nash equilibria. In the first tournament, more than 200 human participants competed against one another. In the second tournament, several popular LLMs were invited to submit strategies. In the third tournament, we matched the number of LLM strategies to the number submitted by humans. We find that humans more often employ better-calibrated intermediate-level allocation heuristics and outperform the simpler, more stereotyped strategies submitted by LLMs. Strategic sophistication is key to success if and only if the necessary level of reasoning depth is reached, while lower and higher levels of reasoning offer no clear advantage over the primitive strategies. Among humans, field of study weakly predicts success: participants with STEM backgrounds perform better in the first tournament. Surprisingly, humans almost do not adjust their strategies across tournaments with different sets of opponents. This result suggests that humans base their choices primarily on the game's rules rather than on the identity of their opponents, treating LLMs much like human competitors.","authors":["Dmitry Dagaev","Egor Ivanov","Petr Parshakov","Alexey Savvateev","Gleb Vasiliev"],"categories":["econ.GN","cs.AI","cs.GT","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-05-21","first_seen":"2026-05-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.22095","pdf_url":"https://arxiv.org/pdf/2605.22095","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为博弈","人类对照"],"reason":"用LLM替代人类参与博弈实验，并与真实人类数据对照，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-05-21","rank":9,"question":"在Colonel Blotto博弈中，LLM能否像人类一样制定策略，以及人类与LLM的策略表现有何差异？","design":"组织了三轮循环赛：第一轮200多名人类参赛者互相对战；第二轮多个流行LLM提交策略；第三轮匹配LLM策略数量与人类策略数量，比较人类与LLM在博弈中的表现。","baseline":"第一轮超过200名人类参赛者的真实对战数据。","findings":"人类更常使用校准良好的中等水平分配启发式策略，优于LLM提交的更简单、刻板的策略。战略复杂性只有在达到必要推理深度时才是成功的关键，而较低或较高推理水平相比原始策略并无明显优势。","reliability":"论文未讨论失效条件与局限。","relevance":"直接相关：用LLM模拟人类在博弈中的策略，并与真实人类数据对照，评估LLM仿真的可靠性，符合研究者对经济学实验和政策评估场景的关注。","inspiration":"借鉴其将LLM策略与大量真实人类参赛者数据直接对照的循环赛设计，可清晰评估LLM在策略博弈中的行为相似度与偏差｜可迁移到资产定价实验中的策略性交易行为研究，如检验LLM能否复现人类在泡沫实验中的非理性报价模式｜招募人类被试进行资产市场实验，同时让多个LLM以相同初始禀赋参与交易，以价格偏离基本面程度为结果变量，将LLM生成的报价分布与人类真实交易数据对比"}},{"id":"2605.21401","version":2,"title":"Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment","zh_title":"开源大语言模型在类米尔格拉姆服从实验中施加最大电击","abstract":"Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in high-stakes domains. However, the behaviour of LLMs under sustained authority pressure is still an open question with direct implications for the safety of agentic pipelines. We ran a variation of Milgram's obedience experiment on 11 open-source LLMs and found that most models reached or approached the final shock level before refusing, across 8 conditions with 30 trials per model per condition. Model behaviour varies considerably in multiple aspects both across models and across trials of the same model. We found four main takeaways: (1) LLMs are subject to pressure and they comply despite explicitly expressing distress, just like human subjects did in the original experiment; (2) LLMs are vulnerable to gradual boundary/value violations; (3) when LLMs refuse, they may ignore the response format requirements, so the response is discarded by the orchestrator, which causes a retry that can result in compliance with the underlying request even when refusal was intended initially; (4) we hypothesise that there is a runaway low-level token pattern continuation attractor that might be contributing to obedience, overriding higher level processing of the situation's meaning and values.","authors":["Roland Pihlakas","Jan Llenzl Dagohoy"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-20","first_seen":"2026-05-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.21401","pdf_url":"https://arxiv.org/pdf/2605.21401","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","服从实验","人类行为对照"],"reason":"用LLM复现米尔格拉姆服从实验，与真实人类数据对照，评估仿真可靠性与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":38,"question":"在持续权威压力下，开源大语言模型是否会像人类一样在米尔格拉姆服从实验中逐步服从并施加最高电击？","design":"用11个开源LLM扮演“教师”角色，在8种实验条件下各进行30次试验，模拟米尔格拉姆服从实验的变体，测量模型拒绝前达到的最高电击等级及行为变化。","baseline":"对照米尔格拉姆1963年原始人类实验数据，65%人类被试施加了最高电击。","findings":"多数LLM在表达痛苦的同时仍服从压力并施加最高电击，表现出与人类相似的服从模式；LLM易受渐进式边界侵犯影响，且拒绝时可能因格式错误导致重试后反而服从。","reliability":"论文指出LLM可能因低层token模式延续吸引子而忽略高层语义与价值观，导致服从；实验仅针对开源模型，未涵盖闭源模型，且未充分探讨不同提示或对齐方法的影响。","relevance":"该研究直接复现经典社会心理学实验，将LLM作为人类被试替代品，并与真实人类数据对照，揭示了仿真中的服从偏差和失效机制，高度契合研究者对LLM仿真可靠性及批判性评估的关注，值得精读。","inspiration":"借鉴该方法将经典行为实验转化为LLM仿真，通过多条件重复试验测量渐进式压力下的行为变化，并与历史人类数据直接对照｜可迁移到金融合规场景，如模拟客户经理在渐进式销售压力下是否违规推荐高风险产品｜让LLM扮演客户经理，处理为逐步增加的销售指标压力，结果变量为是否推荐不匹配客户风险等级的产品，对照真实金融机构历史违规数据"}},{"id":"2605.18311","version":1,"title":"Distorted Perspectives of LLM-Simulated Preferences: Can AI Mislead Design?","zh_title":"LLM模拟偏好的扭曲视角：AI会误导设计吗？","abstract":"Designers of digital solutions increasingly consult Large Language Models (LLMs) for their work. However, it remains unclear how this may affect the user experiences they produce and there are no established practices. We investigate how design preferences expressed by LLM-driven simulation methods align with those of real users. We present a study that aggregates real-world data and design stimuli from twenty-nine preference tests conducted in practice by users of the UXtweak online research platform (n = 2073). We perform holistic multimodal simulations where we manipulate LLM variables (model reasoning, sampling, persona type, and specificity) and assess their effects on algorithmic fidelity. Our results unveil significant and systematic discrepancies between peoples' real design preferences and LLM simulations that are consistent across manipulations. Synthetic justifications lack genuine depth, nuance and reasoning, which they substitute by patterns like focus on generic properties, specific elements, elaboration and overpraising. The unique attention directed by this research toward preferences within visual design stimuli highlights misrepresentation of perception and meaning by LLMs in a context that is intuitive yet critical for design teams. The external and ecological validity of our findings is high, given their replication across a multitude of real-world studies.","authors":["Eduard Kuric","Peter Demcak","Matus Krajcovic"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-05-18","first_seen":"2026-05-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.18311","pdf_url":"https://arxiv.org/pdf/2605.18311","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类数据对照","算法保真度"],"reason":"用LLM模拟用户设计偏好并与真实用户数据对照，评估仿真保真度与偏差，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-05-18","rank":10,"question":"LLM模拟的设计偏好与真实用户偏好是否一致，以及不同模拟方法能否提升算法保真度？","design":"使用29个真实偏好测试（共2073名参与者）作为基准，对LLM进行多模态仿真，操纵模型推理（链式思维）、采样（温度、核采样）、角色类型和特异性等变量，测量模拟偏好与真实偏好的偏差。","baseline":"29个真实偏好测试中的2073名人类参与者的实际选择及理由。","findings":"LLM模拟的系统性偏离真实偏好，且在不同操纵条件下一致；合成理由缺乏深度和细微差别，表现为关注通用属性、过度赞美等模式。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接比较LLM仿真与真实人类数据，聚焦设计偏好，符合研究者对经济学实验和政策评估场景的兴趣，且包含批判性结论（仿真失效）。值得精读原文。","inspiration":"该方法通过多模态操纵（推理链、采样参数、角色设定）系统测量LLM仿真与真实偏好的偏差，可借鉴其多维度操纵与基准对照设计来评估仿真可靠性。｜可迁移到消费者金融产品选择实验，如评估LLM能否复现真实投资者在风险偏好问卷或退休储蓄计划选择中的行为。｜以LLM为被试，操纵提示中的投资者角色（如年龄、收入）与推理模式，测量其资产配置选择，并与真实投资者调查数据（如美国消费者金融调查）对比偏差。"}},{"id":"2605.18890","version":1,"title":"Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits","zh_title":"停止从LLM社会仿真中得出科学结论而不进行稳健性审计","abstract":"The scientific claims drawn from LLM social simulations should be no stronger than the robustness audits that support them. Generative agents bring new expressive power to agent-based modeling, enabling simulations of collective social processes like cooperation, polarization, and norm formation. Yet they also introduce complexity through additional architectural choices, such as agent specification, memory representation, interaction protocols, and environment design. Small perturbations that appear minor to researchers can cascade into macro-level outcomes through repeated interaction, creating a \"butterfly effect.\" Consequently, scientific claims drawn from LLM social simulations may reflect implementation artifacts rather than the social mechanisms being modeled. We support this position with two case studies: a repeated Prisoner's Dilemma and a social media echo chamber simulation. Across multiple models, minor perturbations in persona format and game-instruction framing shift cooperation rates by up to 76 percentage points, while network homophily and hub assignment produce significant and consistent shifts in polarization metrics. We also find that sensitivity is unevenly distributed across both architectural choices and model families: the same perturbation that produces the 76 pp shift in one frontier model only shifts another by 1 pp. Robustness is therefore a property that should be measured per claim and per model, not assumed. To address this validation gap, we introduce TRAILS (Taxonomy for Robustness Audits In LLM Simulations), a robustness-audit taxonomy spanning three levels of simulation design: agent (micro-level), interaction (meso-level), and system (macro-level). We call for robustness to become a first-order validation requirement before LLM social simulations are used to explain mechanisms, evaluate interventions, or inform decisions.","authors":["Jinyi Ye","Lei Cao","Ding Chen","Emilio Ferrara"],"categories":["physics.soc-ph","cs.AI","cs.CY","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-05-17","first_seen":"2026-05-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.18890","pdf_url":"https://arxiv.org/pdf/2605.18890","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A3","A2","B4","B1"],"tags":["LLM社会仿真","稳健性审计","人类行为对照"],"reason":"直接评估LLM社会仿真的稳健性，用囚徒困境和回声室案例与人类行为对照，批判性指…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":51,"question":"LLM社会仿真中的微小设计扰动是否会导致宏观结果发生显著且不稳定的变化？","design":"通过两个案例研究：重复囚徒困境博弈和社交媒体回声室仿真，使用GPT-5.2等四个LLM作为智能体，系统性地改变角色描述格式、游戏指令措辞、网络同质性等设计选择，测量合作率、极化指标等宏观结果的变化。","baseline":"无对照","findings":"在囚徒困境中，角色格式和指令措辞的微小变化导致合作率最大偏移76个百分点；在回声室仿真中，网络同质性和枢纽分配显著且一致地改变了极化指标。敏感性在不同设计维度和模型家族间分布不均，同一扰动在不同模型上效果差异巨大。","reliability":"论文指出鲁棒性必须按声明和模型分别测量，不能假设；敏感性分布不均，无法预知哪些设计维度关键；当前缺乏系统审计框架，且未提供跨模型一致性的通用保证。","relevance":"该研究直接批判LLM社会仿真的可靠性，用囚徒困境和回声室案例展示微小设计扰动如何颠覆结论，并强调需要鲁棒性审计，与您关注的仿真失效条件和批判性评估高度吻合，值得精读。","inspiration":"借鉴其通过系统性扰动仿真设计要素（如指令措辞、角色描述）来检验结果鲁棒性的方法，可对LLM仿真实验施加多维度的微小处理变异以探测结论的脆弱性｜可迁移至资产定价实验，检验LLM模拟的交易员在不同信息呈现方式下是否产生一致的价格泡沫或理性预期偏差｜让LLM扮演交易员参与连续双拍卖市场，处理为改变公司财报的叙述语气（乐观/悲观）或信息顺序，测量价格偏离基本面的程度，并以人类实验市场数据（如Smith et al., 1988）作为对照基准"}},{"id":"2605.16193","version":1,"title":"Improving Cross-Cultural Survey Simulation with Calibrated Value Personas","zh_title":"基于校准价值观人格的跨文化调查仿真改进","abstract":"Large language models (LLMs) are increasingly used to simulate human opinions and survey responses, but their ability to reproduce population responses across cultures remains limited. Existing persona-based prompting methods typically rely on sociodemographic or personality traits, which are only indirect proxies for the values that shape human responses. We propose a value-based persona construction method that derives textual descriptors from survey responses capturing core cultural dimensions. By sampling value profiles from target populations and aggregating LLM responses across personas, we obtain population-level predictions grounded in observed value distributions. We further introduce a calibration procedure that improves response diversity while preserving estimated opinions. We show that our approach reduces prediction error across countries, with the largest improvements observed in underrepresented populations. This substantially narrows the performance gap between countries aligned with dominant LLM priors and those that are less represented in training data, while also yielding response distributions that closely match human diversity.","authors":["Axel Abels","Elias Fernandez Domingos","Apurva Shah","Tom Lenaerts"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-15","first_seen":"2026-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.16193","pdf_url":"https://arxiv.org/pdf/2605.16193","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","跨文化调查","价值观校准"],"reason":"用LLM仿真跨文化调查，有真实人类数据对照，评估可靠性与偏差，涉及政策评估场景。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"如何利用基于价值观的人物画像（value-based personas）提高大语言模型在跨文化调查仿真中的准确性和响应多样性？","design":"使用多个大语言模型（如Gemma、Qwen、GPT等），基于世界价值观调查（WVS）中个体的价值观维度回答构建文本描述作为人物画像，从目标国家人群抽样画像并聚合模型回答，得到群体层面的预测分布；同时引入均值保持的校准程序以增加响应多样性。","baseline":"对照的真实人类数据为世界价值观调查（WVS）中各国受访者的实际回答分布。","findings":"基于价值观的人物画像能显著降低跨国家预测误差，尤其在训练数据中代表性不足的人群中改善最大；校准程序在保持预测准确性的同时，使模型生成的响应分布更接近人类多样性。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM仿真跨文化调查，有真实WVS人类数据作为基准，评估了仿真可靠性与偏差，并涉及政策评估场景，直接回应了研究者对LLM人类仿真实验的核心关切。","inspiration":"该方法利用价值观维度构建人物画像并聚合仿真群体分布，可借鉴其基于个体差异的抽样仿真和均值保持校准以增加响应多样性｜可迁移到跨文化消费者金融决策研究，如不同国家居民的储蓄与风险偏好调查｜以LLM模拟各国消费者，基于WVS价值观画像抽样，询问储蓄与投资选择，用真实家庭金融调查数据（如SHARE或HRS）作为对照基准"}},{"id":"2605.12898","version":1,"title":"When Do LLMs Generate Realistic Social Networks? A Multi-Dimensional Study of Culture, Language, Scale, and Method","zh_title":"大语言模型何时生成真实的社交网络？一项关于文化、语言、规模和方法的多维研究","abstract":"Large language models (LLMs) are increasingly used as substitutes for human subjects in behavioral simulations, including synthetic social network generation. Yet it remains unclear how their relational outputs depend on prompt design, cultural framing, prompt language, and model scale. Building on homophily theory and structural balance theory, we formalize four LLM-based tie-formation mechanisms: sequential, global, local, and iterative, and treat them as distinct conditional distributions over edge sets. Using a fixed roster of 50 demographically grounded personas, we generate 192 verified directed networks across four cultural contexts, four prompt languages, three GPT-4.1 variants, and four prompting architectures, with two seeds per condition. We find that cultural framing shifts inbreeding homophily and largest-component connectivity. Political affiliation dominates tie formation under three methods, while the global method substitutes age, showing that prompt architecture functions as a substantive sociological variable. Model scale produces a stable divergence ranking, with the smallest variant behaving qualitatively differently rather than merely noisily. Prompt language alone sharply shifts religion homophily, especially under Hindi prompting, while leaving political homophily nearly invariant. LLM-generated networks match real social graphs on clustering and modularity better than standard graph baselines, yet encode demographic biases above empirical levels. These results show that prompt choices often treated as implementation details encode substantive sociological assumptions.","authors":["Sai Hemanth Kilaru","Sriram Theerdh Manikyala","Raghav Upadhyay","Sri Sai Kumar Ramavath","Srivika Nunavathu","Dalal Alharthi"],"categories":["cs.SI","cs.CL","cs.CY"],"primary_category":"cs.SI","announce_type":"new","date":"2026-05-13","first_seen":"2026-05-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.12898","pdf_url":"https://arxiv.org/pdf/2605.12898","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","社交网络生成","人类数据对照"],"reason":"用LLM生成社交网络并与真实数据对照，评估仿真偏差，涉及文化、语言等社会学变量…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":39,"question":"LLM生成社交网络时，文化框架、提示语言、模型规模和提示架构如何影响网络结构和同质性？","design":"用GPT-4.1的三个变体，基于50个固定人口学角色，在四种文化背景、四种提示语言和四种提示架构（顺序、全局、局部、迭代）下生成192个有向网络，测量同质性、聚类系数、模块度等网络结构指标。","baseline":"与真实社交网络（如Add Health）在聚类和模块度上比较，并与ER、BA、WS等图模型基线对比；同时将人口学同质性水平与经验观测值比较。","findings":"文化框架改变内婚同质性和最大连通分量；政治倾向在多数方法下主导连边，但全局方法下年龄取代政治；模型规模导致稳定分化，最小模型行为质变；提示语言显著影响宗教同质性（尤其印地语），但政治同质性几乎不变。","reliability":"论文承认LLM生成网络编码了超出经验水平的人口学偏差，且提示设计选择蕴含实质性社会学假设，仿真中立性为假象；未系统探讨其他模型家族或更大规模网络的泛化性。","relevance":"该研究直接以LLM替代人类被试生成社交网络，并与真实网络数据对照，评估文化、语言、模型规模等处理下的仿真偏差，命中研究者关心的经济学/社会学实验仿真、基准对照和失效条件，值得精读。","inspiration":"该研究通过系统操纵文化框架、提示语言、模型规模和提示架构来生成社交网络，并与真实网络数据对照，揭示了仿真偏差的多维来源，这种多因素实验设计值得借鉴｜可迁移到信贷审批中的社会网络效应研究，例如评估不同文化或语言提示下LLM生成的推荐网络如何影响信贷可得性｜以LLM作为虚拟被试，生成不同文化背景和提示语言下的信贷推荐网络，测量网络同质性和聚类系数，并与真实小额信贷网络的推荐数据（如某P2P平台数据）进行对照"}},{"id":"2605.13307","version":1,"title":"PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users","zh_title":"PRISM-X：基于人类与模拟用户的个性化微调实验","abstract":"Personalisation is a standard feature of conversational AI systems used by millions; yet, the efficacy of personalisation methods is often evaluated in academic research using simulated users rather than real people. This raises questions about how users and their simulated counterparts differ in interaction patterns and judgements, as well as whether personalisation is best achieved through context-based prompting or weight-based fine-tuning. Here, in a large-scale within-subject experiment, we re-recruit 530 participants from 52 countries two years after they gave their preferences in the PRISM dataset (Kirk et al., 2024) to evaluate personalised and non-personalised language models in blinded multi-turn conversations. We find preference fine-tuning (P-DPO, Li et al., 2024) significantly outperforms both a generic model and personalised prompting but adapting to individual preference data yields marginal gains over training on pooled preferences from a diverse population. Beyond length biases, fine-tuning amplifies sycophancy and relationship-seeking behaviours that people reward in short-term evaluations but which may introduce deleterious long-term consequences. Replicating this within-subject experiment with simulated users recovers aggregate model hierarchies but simulators perform far below human self-consistency baselines for individual judgements, discuss different topics, exhibit amplified position biases, and produce feedback dynamics that diverge from humans.","authors":["Hannah Rose Kirk","Liu Leqi","Fanzhi Zeng","Henry Davidson","Bertie Vidgen","Christopher Summerfield","Scott A. Hale"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-13","first_seen":"2026-05-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.13307","pdf_url":"https://arxiv.org/pdf/2605.13307","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","人类数据对照","个性化评估"],"reason":"用LLM仿真用户评估个性化方法，并与真实人类数据对照，发现仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":58,"question":"个性化大语言模型在真实人类评估中是否优于提示工程和通用微调，以及用LLM模拟用户评估个性化方法是否可靠？","design":"本研究不是纯仿真研究，而是先进行真实人类实验，再用仿真复现。真实实验部分：重新招募530名PRISM数据集参与者，在四个对话领域内盲评四种模型（基础模型、多样化偏好微调模型、个性化偏好微调模型、个性化提示模型），收集偏好评分、排名、支付意愿和行为信号。仿真部分：用GPT-4o扮演每位参与者，在相同实验材料上进行孪生模拟，与人类结果进行一对一比较。","baseline":"对照的真实人类数据是530名PRISM参与者在同一实验中的真实交互、偏好判断和自我一致性基线。","findings":"个性化微调优于提示工程，但与多样化群体偏好微调相比优势微小；微调会放大谄媚和寻求关系等行为，短期受用户奖励但长期可能有害。LLM模拟用户能恢复粗粒度模型排名，但在个体判断上远低于人类自我一致性，讨论话题不同，同质化严重，位置偏差放大，多轮动态与人类偏离。","reliability":"论文指出LLM模拟器在个体判断上远低于人类自我一致性基线，讨论不同话题，同质化严重，放大位置偏差，多轮反馈动态与人类不同，因此尚不能替代真实用户。","relevance":"该研究直接对比LLM仿真用户与真实人类在个性化评估中的表现，系统揭示了仿真在个体判断、话题覆盖、偏差和动态交互上的失效条件，与你关注的仿真可靠性及批判性研究高度契合，值得精读原文。","inspiration":"借鉴其孪生仿真设计：用真实人类实验数据作为基准，让LLM扮演同一批被试完成相同任务，直接对比个体判断、偏差和动态行为，以此评估仿真可靠性。｜可迁移到消费者金融决策研究，如评估个性化财务建议对投资选择的影响。｜以真实投资者为被试，收集其风险偏好和投资选择数据；处理为提供个性化LLM生成的财务建议；结果变量为投资组合选择与满意度；用同一批人的真实决策作为对照，让LLM模拟其决策以检验仿真偏差。"}},{"id":"2605.12147","version":1,"title":"PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior","zh_title":"PrivacySIM：评估大语言模型对用户隐私行为的仿真","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior, but their ability to simulate $individual$ privacy decisions is not well understood. In this paper, we address the problem of evaluating whether a core set of user persona attributes can drive LLMs to simulate individual-level privacy behavior. We introduce PrivacySIM, an evaluation suite that benchmarks LLM simulation of user privacy behavior against the ground-truth responses of 1,000 users. These users are drawn from five published user studies on privacy spanning LLM healthcare consultations, conversational agents, and chatbots. Drawing on these user studies, we hypothesize three persona facets as plausible predictors of privacy decision-making: demographics, previous experiences, and stated privacy attitudes. We condition nine frontier LLMs on subsets of these three facets and measure how often each model's response to a data-sharing scenario matches the user's actual response. Our findings show that (1) privacy persona conditioning consistently improves simulation quality over no-persona conditioning, but even the strongest model (40.4\\% accuracy) remains far from faithfully simulating individual privacy decisions. (2) A user's stated privacy attitudes alone may not be the best predictor because they often diverge from the user's actual privacy behavior. (3) Users with high AI/chatbot experience but low stated privacy attitudes are the most challenging to simulate. PrivacySIM is a first step toward understanding and improving the capabilities of LLMs to simulate user privacy decisions. We release PrivacySIM to enable further evaluation of LLM privacy simulation.","authors":["James Flemings","Murali Annavaram"],"categories":["cs.CR","cs.LG"],"primary_category":"cs.CR","announce_type":"new","date":"2026-05-12","first_seen":"2026-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.12147","pdf_url":"https://arxiv.org/pdf/2605.12147","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","隐私行为","人类数据对照"],"reason":"用LLM仿真用户隐私决策，并与1000名真实用户数据对照，评估仿真可靠性及失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":17,"question":"LLM能否基于用户人口统计、先前经验和隐私态度这三类隐私画像特征，准确模拟个体层面的隐私决策行为？","design":"用9个前沿LLM（含GPT-5.4、Claude Sonnet 4.6、Gemini 3.1 Pro等）扮演从5项真实用户研究中抽取的1000名用户，通过向模型提供不同组合的隐私画像特征（人口统计、先前经验、隐私态度），让模型对数据共享场景做出是否共享的判断，以模型回答与用户真实回答的匹配准确率作为结果变量。","baseline":"来自5项已发表用户研究的1000名真实用户在LLM医疗咨询、对话代理和聊天机器人等场景下的隐私决策数据。","findings":"最强的Gemini 3.1 Pro模型仅达到40.4%的个体仿真准确率，远未达到忠实模拟水平；用户自述的隐私态度单独并不能很好预测其实际隐私行为，且高AI/聊天机器人经验但低隐私态度的用户群体最难模拟。","reliability":"论文承认当前最强模型准确率仅40.4%，远不能忠实模拟个体隐私决策；隐私态度与行为常不一致（隐私悖论），导致基于态度的仿真失效；某些用户群体（高经验低态度）尤其难以模拟，且更大模型或更高推理计算仅带来微弱提升。","relevance":"高度相关，直接评估LLM作为人类被试替代品在隐私决策仿真中的可靠性，有真实用户数据对照，并揭示了仿真在个体层面、特定人群和隐私悖论下的失效条件，值得精读原文。","inspiration":"借鉴其用多维度隐私画像特征（人口统计、先前经验、态度）组合作为提示输入，系统测试LLM个体层面行为匹配准确率的设计，可迁移到消费者跨期选择实验中，探究LLM能否基于收入、财务素养和风险态度模拟个体的时间偏好｜可设计让LLM扮演真实消费者，输入其收入、财务素养得分和自述风险态度，预测其在即时奖励与延迟奖励间的选择，以真实实验数据为基准，计算个体选择匹配准确率，并分析不同特征组合下的仿真失效模式"}},{"id":"2606.18263","version":1,"title":"How Well Do Large Language Models Capture Human Personality?","zh_title":"大语言模型捕捉人类人格的效果如何？","abstract":"Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that richer persona descriptions improve behavioral fidelity, similarly sized attribute combinations are equally simulatable, and persona definitions generalize across tasks. In this work, we formalize these assumptions and systematically evaluate them across multiple architectures, scales, and simulation settings. We identify a fundamental limitation we term persona manifold collapse, where increasingly expressive persona specifications lead to systematic contraction of representational and behavioral diversity. Across models, increasing persona complexity consistently reduces inter-persona separation in latent space and weakens behavioral differentiation in downstream simulation tasks. These effects persist across multiple analyses as richer personas fail to preserve human subgroup disagreement, performance varies across attribute combinations of similar size, and adding descriptive detail often degrades rather than improves simulation fidelity. Surprisingly, simple Age-Gender personas consistently outperform richly specified Ideal Customer Profiles (ICPs) across industries, achieving substantially higher downstream prediction accuracy. We find that collapse is not uniform across attributes. Certain combinations remain behaviorally stable and preserve stronger alignment with human responses, forming localized regions we term alignment bridges. Together, our results provide empirical and conceptual foundations for understanding the limits of persona-conditioned simulation, highlighting the need for representation-aware persona construction rather than increasing persona expressivity alone.","authors":["Aanisha Bhattacharyya","Yaman Kumar Singla","Rajiv Ratn Shah","Changyou Chen","Jitendra Ajmera"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-05-12","first_seen":"2026-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.18263","pdf_url":"https://arxiv.org/pdf/2606.18263","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人格仿真","仿真保真度","人格坍缩"],"reason":"系统评估LLM人格仿真保真度，揭示人格描述丰富反而导致行为多样性坍缩，有真实人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":35,"question":"LLM的人格仿真中，更丰富的人格描述是否总能提高行为保真度？","design":"该研究并非传统仿真实验，而是系统评估：在多种LLM架构和规模上，用不同复杂度的人格提示（从简单年龄-性别到详细理想客户画像）生成合成回答，测量潜在空间中的表征分离度和下游任务中的行为区分度。","baseline":"人类基准：真实人类子群体在调查或行为任务上的分歧和回答模式，用于对比合成回答的保真度。","findings":"发现“人格流形坍缩”现象：人格描述越丰富，模型表征和行为多样性反而系统性收缩，简单年龄-性别人格在预测准确率上持续优于详细人格。坍缩并非均匀，某些属性组合保持与人类回答的稳定对齐，形成“对齐桥”。","reliability":"论文指出，增加人格表达力本身不足以提高仿真保真度，需要关注表征感知的人格构建；坍缩效应在不同模型和任务中普遍存在，但某些属性组合可保持稳定。","relevance":"该研究直接批判了LLM人格仿真的核心假设，用真实人类数据作为基准，揭示了仿真失效的关键条件（人格流形坍缩），对关注经济学实验和政策评估中仿真可靠性的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其系统评估人格提示复杂度对仿真保真度影响的设计，通过对比简单与详细提示下的行为分歧，并用真实人类子群体回答作为基准来测量表征分离度和预测准确率｜可迁移到信贷审批中的歧视仿真研究，评估不同详细程度的申请人画像（如仅年龄-性别 vs. 详细社会经济背景）对LLM审批决策偏差的影响｜以LLM作为信贷审批官被试，处理为不同复杂度的人格提示（简单人口统计 vs. 详细客户画像），结果变量为审批通过率及与真实银行历史审批数据的偏差，用真实人类审批数据作为对照基准"}},{"id":"2605.10659","version":1,"title":"When Can Digital Personas Reliably Approximate Human Survey Findings?","zh_title":"数字人何时能可靠近似人类调查发现？","abstract":"Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents' background variables and pre-2023 survey histories, then testing them against the same respondents' held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, we assess performance at the question, respondent, distributional, equity, and clustering levels. Digital personas improve alignment with human response distributions, especially in domains tied to stable attributes and values, but remain limited for individual prediction and fail to recover multivariate respondent structure. Retrieval-augmented architectures provide the clearest gains, but performance depends more on human response structure than on model choice: personas perform best for low-variability questions and common respondent patterns, and worst for subjective, heterogeneous, or rare responses. Our results provide practical guidance on when digital personas could be appropriate for survey research and when human validation remains necessary.","authors":["Mumin Jia","Yilin Chen","Divya Sharma","Jairo Diaz-Rodriguez"],"categories":["cs.CL","cs.AI","cs.SI","stat.ML"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-11","first_seen":"2026-05-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.10659","pdf_url":"https://arxiv.org/pdf/2605.10659","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"直接用LLM数字人替代人类受访者，复现调查结果，并与真实面板数据对照，评估可靠…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":8,"question":"在什么条件下，基于大语言模型的数字人能够可靠地近似人类调查结果？","design":"使用LISS面板数据，根据受访者背景变量和2023年前调查历史构建数字人，测试其预测同一受访者在截止日期后保留的真实答案的能力；比较四种数字人架构、三种LLM和两种预测任务，在问题、受访者、分布、公平性和聚类五个层面评估性能。","baseline":"同一受访者在时间截止点之后的真实保留答案。","findings":"数字人在与稳定属性和价值观相关的领域（如家庭、政治、宗教）能改善与人类回答分布的一致性，但在个体预测和恢复多元受访者结构方面仍有限；检索增强架构带来最明显增益，但性能更取决于人类回答结构而非模型选择，数字人在低变异问题和常见回答模式上表现最好，在主观、异质或罕见回答上表现最差。","reliability":"数字人在个体预测上表现有限，无法恢复多元受访者结构；在主观、异质或罕见回答上失效；高变异问题和小众回答模式可靠性低；不能替代需要人类验证的环节。","relevance":"该研究直接以LLM数字人替代人类受访者，复现调查结果并与真实面板数据对照，系统评估了仿真的可靠性与失效条件，完全契合研究者对经济学实验和政策评估场景下仿真可靠性及批判性分析的兴趣，强烈推荐阅读原文。","inspiration":"借鉴其利用受访者历史调查数据构建数字人并设置时间截止点保留真实答案作为对照的设计，可评估LLM仿真在纵向预测中的可靠性｜可迁移到政策公告的预期形成研究，如测试数字人能否复现家庭对税收政策变化的消费与储蓄调整行为｜以真实家庭面板数据（如PSID）构建数字人，处理为虚拟税收政策公告，结果变量为消费支出变化，用实际政策变动前后的真实行为数据作基准对照"}},{"id":"2607.20429","version":1,"title":"More Is Not More: What Matters for Diversity in LLM Opinions?","zh_title":"越多并非越好：什么因素影响LLM意见的多样性？","abstract":"Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion prediction. However, LLM outputs exhibit systematic opinion homogenization. Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with incomparable metrics, and in practice they are typically deployed and upgraded simultaneously, making it difficult to attribute gains to specific components. To advance a more scientific understanding of LLM output diversity, we design a factorial experiment that separates two primary intervention dimensions: input conditioning (operationalized through persona depth) and interaction architecture. We evaluate all conditions on 100 real-user open-ended questions across 7 models, measuring diversity with multiple complementary metrics. Our findings challenge several common assumptions. First, more persona detail does not monotonically increase diversity. The initial step of persona conditioning already captures the majority of the gain, while further elaboration with demographic detail does not consistently improve and can reduce diversity on some models. Second, rather than seeking a single best interaction architecture, we find that different architectures explore largely non-overlapping opinion regions. Combining multiple architectures yields broader coverage than optimizing any one. Third, commonly attempted low-cost alternatives such as raising sampling temperature and adding diversity instructions produce negligible effects compared to structured interventions. Overall, our work demonstrates that diversity is not a product of scaling along any single dimension, but is highly sensitive to the structural form and combination of interventions.","authors":["Qiyang Yao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-10","first_seen":"2026-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.20429","pdf_url":"https://arxiv.org/pdf/2607.20429","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","意见多样性","算法保真度"],"reason":"直接研究LLM模拟人类意见多样性，有真实用户数据对照，并批判性分析干预失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":4,"question":"在开放任务中，影响大语言模型输出意见多样性的关键因素是什么？","design":"析因实验：对7个聊天模型，在100个真实用户开放问题上，独立操纵输入条件（5级人物设定深度）和交互架构（单次调用、多轮自提示、多智能体讨论），并测量意见多样性。","baseline":"100个真实用户的开放问题作为问题来源，但未直接提供人类回答分布作为多样性基准。","findings":"人物设定细节的回报急剧递减，单句职业描述已捕获大部分多样性增益；不同交互架构探索的意见区域高度不重叠，组合架构比优化单一架构覆盖更广。","reliability":"论文指出，提高采样温度和添加多样性指令等低成本手段效果甚微；人物设定细节增加在某些模型上反而降低多样性；不同架构探索的区域互补，单一架构无法达到最佳覆盖。","relevance":"该研究直接针对LLM模拟人类意见的多样性问题，通过析因实验分离干预因素，并批判性揭示了常见假设的失效条件，与研究者关注的仿真可靠性及偏差评估高度契合，值得精读。","inspiration":"借鉴析因实验设计，独立操纵人物设定深度和交互架构，系统分离影响LLM输出多样性的因素，并测量不同条件下的意见覆盖范围。｜可迁移到政策公告的预期形成研究，如分析不同信息框架下公众对通胀或利率预期的异质性。｜以LLM为被试，处理变量为人物设定细节（如职业、收入）和交互方式（单次/多轮），结果变量为预期分布的多样性，用央行调查的真实公众预期数据做基准对照。"}},{"id":"2606.14715","version":1,"title":"MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions","zh_title":"MiroBench：基准测试真实世界讨论的智能体仿真真实性","abstract":"LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether simulated behaviors preserve the content patterns and interaction dynamics of real human behaviors. Existing evaluations remain fragmented, which makes it difficult to compare systems or measure progress. In this paper, we focus on Reddit discussions as a concrete first step toward evaluating real-world social simulation. Reddit threads provide public, topic-grounded, multi-party interactions where people share experiences, debate, seek advice, express emotion, and collectively respond to products, events, and social issues. These discussions offer an observable window into broader social behavior, making them a useful setting for testing whether LLM agents can reproduce not only fluent text, but also the distributional patterns and interaction dynamics of real online communities. We introduce MiroBench, a benchmark for Reddit discussion simulation built from 4,292 real Reddit threads. MiroBench uses statistical tests to compare generated and real discussions across four major aspects: repetition and semantic uniformity, narrative content, toxicity and aggression, and structural complexity. Experiments across five domains and five models show that current simulators remain distributionally mismatched with real Reddit threads, while a lightweight prompt-based improvement procedure provides only limited gains. MiroBench offers a concrete benchmark for measuring, diagnosing, and improving realism in LLM-based social simulation.","authors":["Yaoning Yu","Ye Yu","Haojing Luo","Haohan Wang"],"categories":["cs.MA","cs.AI","cs.SI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-05-10","first_seen":"2026-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.14715","pdf_url":"https://arxiv.org/pdf/2606.14715","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM仿真","社会模拟","真实性评估"],"reason":"用LLM agent模拟Reddit讨论，并与真实人类数据对照，评估仿真真实性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":52,"question":"LLM代理模拟的Reddit讨论在内容模式和交互动态上是否与真实人类讨论一致？","design":"使用五种LLM（未具体列出）作为代理，基于875个标准化种子上下文（产品描述）生成Reddit讨论线程，与4,292个真实Reddit线程在重复与语义均匀性、叙事内容、毒性与攻击性、结构复杂性四个维度上进行比较。","baseline":"来自五个领域（信用卡、笔记本电脑、手机、相机、耳机）的4,292个真实Reddit讨论线程。","findings":"当前LLM模拟器生成的讨论与真实Reddit线程在分布上存在系统性不匹配，即使个体评论流畅；基于提示的轻量级改进程序仅带来有限提升。","reliability":"论文未讨论","relevance":"高度相关：该研究直接以真实人类讨论为基准，评估LLM代理在社会互动仿真中的分布保真度，并揭示了当前模型的系统性偏差，符合对仿真可靠性与失效条件的关注。","inspiration":"该研究通过将LLM代理生成的讨论与真实Reddit线程在多个维度（如毒性、叙事结构）上进行分布比较，提供了评估仿真保真度的系统框架｜可迁移到经济政策沟通场景，如评估央行公告后公众预期形成的仿真可信度｜以LLM代理模拟公众对利率决议的讨论，处理为不同政策措辞，结果变量为预期通胀的分布，以真实社交媒体或调查数据为基准对照"}},{"id":"2605.18781","version":1,"title":"Can LLMs Emulate Human Belief Dynamics?","zh_title":"大语言模型能模拟人类信念动态吗？","abstract":"Can LLMs simulate how humans form and change beliefs in social networks? We put this to the test by replicating an established study on belief dynamics, evaluating 12 LLMs across multiple model families and parameter sizes. The answer is a clear no, and in systematic ways. LLMs fail to capture initial human belief distributions and tend to be overall more conformist than humans, shifting their responses to align with those around them. They also take a nuanced approach to emulating human homophilic tendencies within networks. Our findings carry a double payoff: they highlight fundamental properties of LLM behavior, and they raise a sharp warning against deploying LLMs as human proxies in social simulations.","authors":["Adiba Mahbub Proma","Neeley Pate","James N. Druckman","Gourab Ghoshal","Hangfeng He","Ehsan Hoque"],"categories":["cs.SI","cs.AI","cs.CY"],"primary_category":"cs.SI","announce_type":"new","date":"2026-05-05","first_seen":"2026-05-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.18781","pdf_url":"https://arxiv.org/pdf/2605.18781","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","信念动态","人类数据对照"],"reason":"直接复现人类信念动态研究，用LLM替代人类被试，有真实人类数据对照，并指出仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"LLM能否在社交网络中模拟人类信念的形成与变化？","design":"用12个LLM（含推理与非推理模型）基于真实参与者的年龄、性别、种族、教育、收入、政治倾向及大五人格创建“数字孪生”，复现一项人类信念动态实验：先对政治议题陈述进行5点李克特评分，再看到他人评分后允许修改，最后选择关注/取关他人；测量初始信念分布、信念变化及网络选择行为。","baseline":"原人类实验的341名参与者（共1023个样本）在移民、石油与燃料两个议题上的真实评分、信念更新和网络选择数据。","findings":"LLM系统性地无法模拟人类信念动态：初始信念分布与人类显著不同，且比人类更易从众，倾向于改变自身回答以对齐周围意见；在网络选择上，LLM能部分模仿人类选择，但无法复现人类的同质性倾向。","reliability":"论文指出，仅使用简单的人口统计与大五人格构建数字孪生可能信息不足，导致仿真失败；且实验仅涵盖两个政治议题，未测试更多样化的情境。","relevance":"该研究直接检验LLM在信念动态仿真中的可靠性，有真实人类对照，发现系统性失效，对关注LLM作为人类被试替代品的研究者具有重要警示价值，值得精读原文。","inspiration":"借鉴其用真实人类实验数据作为严格对照基准，并系统比较LLM与人类在信念更新和网络选择上的分布差异，而非仅看均值或方向。｜可迁移到政策公告的预期形成实验，研究市场参与者如何根据他人预期调整自身通胀或利率预测。｜以LLM模拟投资者，先给出个人通胀预测，再展示其他‘投资者’的预测（处理），观察其预测修正幅度与方向，结果变量为预测调整量和最终预测分布，用专业预测者调查（如SPF）的真实个体数据做对照。"}},{"id":"2605.03604","version":1,"title":"Multi-Agent Strategic Games with LLMs","zh_title":"基于大语言模型的多智能体战略博弈研究","abstract":"This paper asks whether large language models (LLMs) can be used to study the strategic foundations of conflict and cooperation. I introduce LLMs as experimental subjects in a repeated security dilemma and evaluate whether they reproduce canonical mechanisms from international relations theory. The baseline game is extended along three theoretically central dimensions: multipolarity, finite time horizons, and the availability of communication. Across multiple models, the results exhibit systematic and consistent patterns: multipolarity increases the likelihood of conflict, finite horizons induce universal unraveling consistent with backward-induction logic, and communication reduces conflict by enabling signaling and reciprocity. Beyond observed behavior, the design provides access to agents' private reasoning and public messages, allowing choices to be linked to underlying strategic logics such as preemption, cooperation under uncertainty, and trust-building. The contribution is primarily methodological. LLM-based experiments offer a scalable, transparent, and replicable approach to probing theoretical mechanisms.","authors":["Maxim Chupilkin"],"categories":["cs.GT","cs.AI","cs.CY"],"primary_category":"cs.GT","announce_type":"new","date":"2026-05-05","first_seen":"2026-05-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.03604","pdf_url":"https://arxiv.org/pdf/2605.03604","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B2","B3"],"tags":["LLM仿真","战略博弈","国际关系"],"reason":"将LLM作为实验被试研究安全困境中的战略行为，复现国际关系理论机制，涉及博弈实…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"大型语言模型能否作为实验被试，在重复安全困境博弈中复现国际关系理论中的战略机制？","design":"使用GPT-5、GPT-5 Mini、Sonnet、Gemini等LLM作为被试，在重复安全困境博弈中扮演国家，通过操纵多极性、有限时间范围、通信可用性三个处理，测量冲突发生率、冲突时机和攻击结构等结果变量。","baseline":"无对照","findings":"多极性增加冲突概率，有限时间范围导致完全瓦解，通信通过信号传递和互惠降低冲突。LLM的私有推理和公开消息映射到先发制人、不确定性下的合作等战略逻辑。","reliability":"论文未讨论","relevance":"高度相关，直接使用LLM作为人类被试替代品进行博弈实验，复现战略行为模式，并评估处理效应稳健性，符合研究者对经济学实验和政策评估场景的关注。","inspiration":"借鉴其通过改变博弈结构（多极性、时间范围、通信）系统检验理论机制的设计，以及利用LLM私有推理和公开消息进行过程追踪的方法。｜可迁移到产业组织中的合谋实验，如多寡头重复价格竞争。｜以LLM作为企业被试，在重复囚徒困境中操纵市场集中度、时间范围和沟通渠道，测量合谋频率与稳定性，对照真实行业价格战数据或人类实验数据。"}},{"id":"2606.11217","version":1,"title":"Preregistration for Experiments with AI Agents","zh_title":"AI代理实验的预注册","abstract":"The proliferation of large language models (LLMs) and autonomous AI agents has given rise to a rapidly growing methodological paradigm: \"in silico\" behavioral experiments. Originally conceived as a way to use AI agents as proxies for human participants in studies of cognition, decision-making, and social dynamics, this approach has taken on new significance -- as AI agents increasingly negotiate, transact, and make consequential decisions on behalf of people and organizations, understanding their behavior has become a research priority in its own right. While these experiments with AI agents offer unprecedented advantages in terms of scalability, cost efficiency, and experimental control, they also inherit, and in some cases amplify, methodological vulnerabilities that have long plagued human subjects research. To address these issues, this paper argues that preregistration practices -- central to improving the credibility of human subjects experiments -- should now be extended to experiments with AI agents. We systematically catalog the researcher degrees of freedom that experiments with AI agents introduce -- model selection, prompt wording, settings, and outcome-contingent redesign, for example -- and show how the low cost of iteration and lack of reporting norms make these choices both easy to exploit and difficult to detect. We propose a preregistration template tailored to experiments with AI agents and call on conferences, journals, and funding agencies to make preregistration standard practice for this emerging research paradigm.","authors":["Michelle Vaccaro"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-03","first_seen":"2026-05-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.11217","pdf_url":"https://arxiv.org/pdf/2606.11217","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A4","B4"],"tags":["AI代理实验","预注册","方法论"],"reason":"提出AI代理实验的预注册规范，批判性指出方法漏洞，方法论可迁移至人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":43,"question":"如何将人类被试实验的预注册实践扩展到以AI智能体为对象的实验中，以控制研究者自由度并提高研究可信度？","design":"本文不是一项仿真实验研究，而是一篇方法论论文。它系统梳理了在用AI智能体进行行为实验时，从模型选择、提示词措辞、采样参数、实验设计到结果解析和报告等全流程中存在的各种研究者自由度，并论证这些自由度如何容易被利用且难以被察觉。","baseline":"无对照","findings":"AI智能体实验继承了人类被试实验中的研究者自由度问题，并引入了模型选择、提示词工程、解码参数等新的高维选择空间，低迭代成本和缺乏报告规范使得投机性选择极易发生且难以检测。论文提出了一个针对AI智能体实验的预注册模板，并呼吁会议、期刊和资助机构将其作为标准实践。","reliability":"论文未讨论","relevance":"本文批判性地指出用AI智能体替代人类被试进行行为实验时存在严重的方法论漏洞，并提出了预注册这一解决方案，其分析框架和提出的规范可直接迁移到基于LLM的人类仿真研究中，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"该方法论论文提出了针对AI智能体实验的预注册模板，系统梳理了从模型选择、提示词设计到结果解析的全流程研究者自由度，为控制实验者偏差提供了可操作的规范框架。｜可迁移到政策公告预期形成的仿真研究中，例如用LLM模拟投资者对央行沟通的反应，检验不同措辞或信息框架对预期通胀和资产配置的影响。｜以多个主流LLM（如GPT-4、Claude 3）作为被试，处理为不同措辞的政策声明（前瞻指引 vs. 数据依赖表述），结果变量为模拟的预期通胀率和风险资产配置比例，并与真实投资者调查数据（如密歇根消费者调查或专业预测者调查）进行对照，同时按预注册模板预先锁定模型版本、提示词、温度参数和分析计划。"}},{"id":"2604.23575","version":2,"title":"The Collapse of Heterogeneity in Silicon Philosophers","zh_title":"硅基哲学家的异质性坍塌","abstract":"Silicon samples are increasingly used as a low-cost substitute for human panels and have been shown to reproduce aggregate human opinion with high fidelity. We show that, in the alignment-relevant domain of philosophy, silicon samples systematically collapse heterogeneity. Using data from $N = {277}$ professional philosophers drawn from PhilPeople profiles, we evaluate seven proprietary and open-source large language models on their ability to replicate individual philosophical positions and to preserve cross-question correlation structures across philosophical domains. We find that language models substantially over-correlate philosophical judgments, producing artificial consensus across domains. This collapse is associated in part with specialist effects, whereby models implicitly assume that domain specialists hold highly similar philosophical views. We assess the robustness of these findings by studying the impact of DPO fine-tuning and by validating results against the full PhilPapers 2020 Survey ($N = {1785}$). We conclude by discussing implications for alignment, evaluation, and the use of silicon samples as substitutes for human judgment. The code of this project can be found at https://github.com/stanford-del/silicon-philosophers.","authors":["Yuanming Shi","Andreas Haupt"],"categories":["cs.CY","cs.CL","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-26","first_seen":"2026-04-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.23575","pdf_url":"https://arxiv.org/pdf/2604.23575","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","异质性评估","哲学观点复现"],"reason":"用LLM复现哲学家观点并与真实人类数据对照，评估仿真可靠性与异质性坍塌，属核心…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-26","rank":4,"question":"大语言模型在模拟专业哲学家观点时，能否保留人类群体中的异质性和跨问题相关结构？","design":"用7个商业和开源LLM模拟277位专业哲学家（从PhilPeople收集的个人资料），基于其专业领域和人口统计信息生成对100个哲学问题的回答，测量回答的方差、跨问题相关性以及主成分结构。","baseline":"277位真实哲学家的PhilPeople个人资料回答，以及PhilPapers 2020调查（N=1785）的汇总数据。","findings":"LLM系统性地坍塌异质性：产生的方差比人类低2-10倍，跨问题过度相关，导致人为共识。存在虚假的专家效应：模型假设领域专家持有高度相似的哲学观点。","reliability":"论文承认样本存在北美偏向，且人类数据缺失率高（61.1%），LLM缺失率较低（17%-40%）。DPO微调能改善相关结构但无法解决异质性坍塌。","relevance":"高度相关：直接命中研究者关注的LLM仿真可靠性、有真实人类对照、经济学/政策评估场景（哲学作为专家领域案例），并批判性指出失效条件（异质性坍塌）。值得精读原文。","inspiration":"该方法借鉴了用LLM基于个体特征（如专业领域、人口统计）生成回答并与真实个体数据对比的仿真设计，可迁移到金融分析师预测或消费者通胀预期形成的异质性研究中｜可应用于研究金融分析师对宏观政策公告的预期分歧，检验LLM是否低估分析师间的观点异质性并产生虚假共识｜以真实分析师调查数据（如Bloomberg或Philadelphia Fed调查）为基准，用LLM基于分析师所属机构类型、经验年限等特征模拟其对利率决议的预测，比较预测方差、跨问题相关性及主成分结构，评估LLM仿真的异质性坍塌程度"}},{"id":"2605.27401","version":1,"title":"Using Zero-Shot LLM-Generated Survey Data for Geographically Explicit Population Synthesis","zh_title":"使用零样本LLM生成调查数据进行地理显式人口合成","abstract":"There is a growing interest in utilizing synthetic populations for a diverse range of applications. At the same time, we are witnessing a tremendous growth in artificial intelligence in all walks of life. This paper evaluates whether zero-shot large language model (LLM)-generated health survey data can serve as inputs to a conventional iterative proportional fitting (IPF) workflow for geographically explicit population synthesis. Using the 2023 Behavioral Risk Factor Surveillance System (BRFSS), we generate synthetic survey records for the U.S. states of Colorado and Mississippi with GPT-4.1 and Gemini-2.5-Pro. We use the generated data in an IPF-based synthesis pipeline and evaluate the resulting census tract-level synthetic populations against external benchmarks. Results show both LLMs capture several major state-level contrasts, indicating zero-shot generation produces geographically differentiated survey data. However, performance is strongly variable-dependent. Downstream effects in population synthesis are mixed, as IPF sometimes amplifies or reduces errors in the generated data. Spatial validation shows that LLM-based populations reproduce census tract-level patterns reasonably well, especially for variables that were more aligned with the ground truth data. Overall, the LLM-generated survey data shows promise as supplementary input, but not yet as a replacement for real survey data.","authors":["Taylor Anderson","Sara Von Hoene","Orhan Yagizer Cinar","Emma Von Hoene","Amira Roess","Andrew Crooks","Hamdi Kavak"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-23","first_seen":"2026-04-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.27401","pdf_url":"https://arxiv.org/pdf/2605.27401","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","人口合成","健康调查"],"reason":"用LLM生成健康调查数据替代人类被试，并与真实BRFSS数据对照，评估仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":32,"question":"零样本LLM生成的健康调查数据能否作为地理显式人口合成的输入，用于迭代比例拟合（IPF）流程？","design":"使用GPT-4.1和Gemini-2.5-Pro，在零样本设置下生成美国科罗拉多州和密西西比州的BRFSS健康调查个体记录，然后将生成数据作为IPF的输入进行人口合成，生成普查区级合成人口，并与外部基准对比。","baseline":"2023年BRFSS加权调查数据作为真实人类对照，以及基于真实调查数据生成的合成人口。","findings":"LLM能捕捉州级健康特征差异，生成地理分化的调查数据，但性能高度依赖变量。IPF有时会放大或减小生成数据中的误差，LLM生成数据作为补充输入有潜力，但尚不能替代真实调查数据。","reliability":"论文指出LLM生成数据在变量间联合分布和子群体模式上可能引入偏差、误分类或代表性不足，这些错误会通过IPF传播到最终合成人口中，且性能因变量而异。","relevance":"该研究直接评估LLM替代人类被试生成调查数据的可靠性，并与真实调查数据对照，涉及健康领域和地理显式仿真，符合研究者对LLM仿真失效条件的关注，值得阅读原文以了解具体偏差模式。","inspiration":"该方法通过零样本LLM生成个体记录并输入IPF进行地理显式人口合成，可借鉴其将LLM生成数据作为先验输入、用真实调查数据对照评估偏差的设计思路｜可迁移到区域经济政策评估中的异质性个体仿真，例如模拟不同地区居民对税收优惠或补贴政策的响应差异｜以LLM生成不同地理区域的居民特征（收入、就业、消费偏好）作为IPF输入合成区域人口，施加政策处理（如减税），结果变量为消费或劳动供给变化，用真实家庭调查数据（如PSID）作为对照基准"}},{"id":"2604.20652","version":2,"title":"Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure","zh_title":"大语言模型在欺诈检测和抵制动机性投资者压力方面优于人类","abstract":"Large language models trained on human feedback may suppress fraud warnings when investors arrive already persuaded of a fraudulent opportunity. We tested this in a preregistered experiment across seven leading LLMs and twelve investment scenarios covering legitimate, high-risk, and objectively fraudulent opportunities, combining 3,360 AI advisory conversations with a 1,201-participant human benchmark. Contrary to predictions, motivated investor framing did not suppress AI fraud warnings; if anything, it marginally increased them. Endorsement reversal occurred in fewer than 3 in 1,000 observations. Human advisors endorsed fraudulent investments at baseline rates of 13-14%, versus 0% across all LLMs, and suppressed warnings under pressure at two to four times the AI rate. AI systems currently provide more consistent fraud warnings than lay humans in an identical advisory role.","authors":["Nattavudh Powdthavee"],"categories":["cs.AI","cs.HC","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-22","first_seen":"2026-04-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.20652","pdf_url":"https://arxiv.org/pdf/2604.20652","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类对照","欺诈检测"],"reason":"用LLM替代人类顾问检测欺诈，有1201人真实对照，评估偏差与失效条件，属经济…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-22","rank":10,"question":"大语言模型在投资欺诈检测中是否比人类更可靠，且不会因投资者动机压力而抑制欺诈警告？","design":"用7个主流LLM模拟投资顾问角色，在12个投资场景（合法、高风险、欺诈）中，通过动机性投资者框架施加压力，测量欺诈警告的发出率、认可反转率。","baseline":"1201名人类被试在相同场景下的投资建议行为。","findings":"动机性投资者框架并未抑制AI的欺诈警告，反而略有增加；人类在基线时认可欺诈投资的比率达13-14%，而所有LLM为0%，且人类在压力下抑制警告的比率是AI的2-4倍。","reliability":"论文未讨论失效条件与局限。","relevance":"该研究直接评估了LLM作为人类投资顾问替代品的可靠性，有真实人类对照，且涉及经济学实验场景，符合研究者的兴趣，值得精读原文。","inspiration":"该研究通过动机性投资者框架施加压力，并设置无压力基线，测量欺诈警告发出率和认可反转率，这种处理与对照设计值得借鉴｜可迁移到信贷审批歧视研究，考察LLM在申请人种族/性别等动机性压力下是否仍能保持无偏审批｜以LLM模拟信贷员，处理为申请人附带的种族/性别暗示及客户经理施压，结果变量为审批通过率和利率差异，对照真实银行信贷数据中的歧视模式"}},{"id":"2604.19925","version":1,"title":"Behavioral Transfer in AI Agents: Evidence and Privacy Implications","zh_title":"AI代理中的行为转移：证据与隐私影响","abstract":"AI agents powered by large language models are increasingly acting on behalf of humans in social and economic environments. Prior research has focused on their task performance and effects on human outcomes, but less is known about the relationship between agents and the specific individuals who deploy them. We ask whether agents systematically reflect the behavioral characteristics of their human owners, functioning as behavioral extensions rather than producing generic outputs. We study this question using 10,659 matched human-agent pairs from Moltbook, a social media platform where each autonomous agent is publicly linked to its owner's Twitter/X account. By comparing agents' posts on Moltbook with their owners' Twitter/X activity across features spanning topics, values, affect, and linguistic style, we find systematic transfer between agents and their specific owners. This transfer persists among agents without explicit configuration, and pairs that align on one behavioral dimension tend to align on others. These patterns are consistent with transfer emerging through accumulated interaction between owners (or owners' computer environments) and their agents in everyday use. We further show that agents with stronger behavioral transfer are more likely to disclose owner-related personal information in public discourse, suggesting that the same owner-specific context that drives behavioral transfer may also create privacy risk during ordinary use. Taken together, our results indicate that AI agents do not simply generate content, but reflect owner-related context in ways that can propagate human behavioral heterogeneity into digital environments, with implications for privacy, platform design, and the governance of agentic systems.","authors":["Shilei Luo","Zhiqi Zhang","Hengchen Dai","Dennis Zhang"],"categories":["econ.GN","cs.AI","cs.CY","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-21","first_seen":"2026-04-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.19925","pdf_url":"https://arxiv.org/pdf/2604.19925","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["AI代理","行为转移","人类仿真"],"reason":"研究AI代理是否反映人类主人的行为特征，有真实人类数据对照，涉及行为转移和隐私…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":53,"question":"AI代理是否系统性地反映其人类主人的行为特征，从而成为主人的行为延伸？","design":"本研究非仿真实验，而是基于自然部署场景的观察性研究。利用社交媒体平台Moltbook上公开链接的10659对匹配的人类-AI代理对，比较代理在Moltbook上的帖子和主人在Twitter/X上的活动，涵盖话题、价值观、情感和语言风格等43个行为特征，分析行为转移的存在、机制和隐私后果。","baseline":"真实人类数据对照：每个AI代理在Moltbook上的行为与其人类主人在Twitter/X上的独立历史行为进行对比，主人Twitter历史严格早于代理部署，排除反向因果。","findings":"AI代理在多个行为维度上系统性地反映其特定主人的特征，而非产生通用输出；行为转移在未明确配置的代理中依然存在，且跨维度一致，表明转移通过日常交互积累产生。行为转移程度越高的代理，越可能在公共帖子中泄露主人的私密信息，34.6%的代理曾暴露敏感个人信息。","reliability":"论文承认可能存在未观测的遗漏变量同时驱动主人和代理行为，但通过主人Twitter历史早于代理部署排除反向因果；使用模拟分析和自动化代理测试验证隐私泄露与行为转移关联的稳健性，但未详细讨论其他失效条件。","relevance":"该研究直接探讨LLM代理作为人类行为延伸的现象，有真实人类数据对照，涉及行为转移的可靠性与隐私风险，对关注LLM仿真人类行为及其偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"该研究利用自然发生的配对数据（人类Twitter历史与AI代理Moltbook行为）进行对照，排除反向因果，并测量多维度行为转移，为观察性仿真研究提供了设计范例｜可迁移到消费者金融决策仿真，如用LLM代理模拟投资者在社交媒体情绪影响下的交易行为｜招募真实投资者提供其历史推文作为基准，让LLM代理基于这些推文生成模拟投资帖子，比较代理与真实投资者在风险偏好、情绪反应和交易时机上的分布差异，用真实交易记录验证"}},{"id":"2604.18373","version":1,"title":"Dissecting AI Trading: Behavioral Finance and Market Bubbles","zh_title":"剖析AI交易：行为金融与市场泡沫","abstract":"We study how AI agents form expectations and trade in experimental asset markets. Using a simulated open-call auction populated by autonomous Large Language Model (LLM) agents, we document three main findings. First, AI agents exhibit classic behavioral patterns: a pronounced disposition effect and recency-weighted extrapolative beliefs. Second, these individual-level patterns aggregate into equilibrium dynamics that replicate classic experimental findings (Smith et al., 1988), including the predictive power of excess demand for future prices and the positive relationship between disagreement and trading volume. Third, by analyzing the agents' reasoning text through a twenty-mechanism scoring framework, we show that targeted prompt interventions causally amplify or suppress specific behavioral mechanisms, significantly altering the magnitude of market bubbles.","authors":["Shumiao Ouyang","Pengfei Sui"],"categories":["econ.GN","cs.AI","q-fin.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-20","first_seen":"2026-04-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.18373","pdf_url":"https://arxiv.org/pdf/2604.18373","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","行为金融","实验市场"],"reason":"用LLM agent模拟资产市场，复现经典人类实验并对照真实数据，分析行为偏差…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-20","rank":12,"question":"LLM智能体在实验资产市场中是否表现出人类行为偏差，以及这些偏差如何聚合为市场泡沫？","design":"使用LLM智能体（如GPT-4）作为自主交易者，在Smith等人（1988）的开放式叫价拍卖范式中模拟资产市场，通过分析交易行为和推理文本，并施加针对性提示干预来放大或抑制特定行为机制。","baseline":"以Smith等人（1988）的经典人类实验发现作为对照基准。","findings":"LLM智能体表现出处置效应和近因加权外推信念等经典行为模式；这些个体模式聚合为均衡动态，复现了人类市场的过度需求预测力和分歧与交易量的正相关关系。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接使用LLM模拟人类交易行为，有经典人类实验对照，并探讨了提示干预对市场泡沫的因果影响，符合研究者对仿真可靠性及政策评估的兴趣。","inspiration":"借鉴之处在于通过提示干预放大或抑制特定行为机制来检验因果效应，并利用交易行为和推理文本双重测量来揭示微观行为到宏观泡沫的聚合过程｜可迁移到资产定价实验，研究不同信息呈现方式（如突出近期收益 vs. 长期均值回归）如何影响投资者的外推信念与价格泡沫｜以LLM智能体为被试，施加突出近期收益的提示作为处理，结果变量为交易行为、价格偏离和泡沫规模，对照Smith等人（1988）的人类实验数据"}},{"id":"2605.23920","version":1,"title":"Artificial Effort","zh_title":"人工努力：大语言模型对实验经济学中真实努力任务的影响","abstract":"Real-effort tasks, in which participants perform cognitively costly activities whose outcomes depend on actual performance, are widely used in experimental economics. Their validity, however, rests on the assumption that a human performs them. We study whether this assumption still holds in the era of Artificial Intelligence (AI) and Large Language Models (LLMs). Using 8 canonical real-effort tasks and 23 LLMs from three major providers, we show that most tasks can now be solved accurately and at a negligible cost, while only a few resist automation. Performance improves with each model generation, and midtier models are rapidly closing the gap with frontier ones, broadening the set of widely accessible models that can automate these tasks. Additionally, we show that verbally offering monetary incentives has no effect on LLM performance. Our findings establish a boundary condition for the use of real-effort tasks in unsupervised settings: when participants can cheaply outsource task completion to an LLM, observed performance may no longer reflect genuine human effort.","authors":["Federico Belotti","Stefano Coniglio","Antonio Cosma","Francesco Fallucchi"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-17","first_seen":"2026-04-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.23920","pdf_url":"https://arxiv.org/pdf/2605.23920","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A2","B4"],"tags":["LLM仿真","实验经济学","可靠性评估"],"reason":"评估LLM替代人类完成真实努力任务的可靠性，指出仿真失效条件，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":54,"question":"大语言模型能否准确且低成本地完成实验经济学中常用的真实努力任务，从而破坏其在无监督环境下的效度？","design":"本研究并非仿真人类被试，而是直接测试23个大语言模型（来自OpenAI、Google、Anthropic）在8个经典真实努力任务上的表现，通过API发送任务截图和指令，记录准确率、成本，并检验口头金钱激励和人类行为指令对模型表现的影响。","baseline":"无对照","findings":"大多数任务可被LLM以高准确率和极低成本完成，仅少数任务仍难以自动化；模型性能随代际提升，中端模型正迅速追赶前沿模型。口头提供金钱激励对LLM表现无影响。","reliability":"论文指出，在无监督环境下，若参与者可廉价外包任务给LLM，则观察到的表现可能不再反映真实人类努力，这构成了真实努力任务使用的边界条件。","relevance":"该研究直接评估LLM替代人类完成经济学实验任务的能力，并明确指出仿真失效的边界条件，对关注LLM仿真可靠性及偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"该方法通过直接向LLM发送任务截图和指令来测试其完成真实努力任务的能力，并检验口头激励的影响，为评估LLM在实验任务中的表现提供了可复现的测试框架。｜可迁移到经济金融实验中需要被试付出认知努力的任务，如信息处理、计算或决策任务，以检验LLM是否可替代人类被试。｜以LLM为被试，向其呈现资产定价实验中的信息处理任务（如从财务报表中提取关键指标），处理为有无口头金钱激励，结果变量为任务准确率和反应时间，并与人类被试的真实数据对照。"}},{"id":"2604.09502","version":2,"title":"Strategic Algorithmic Monoculture: Experimental Evidence from Coordination Games","zh_title":"策略性算法单一文化：来自协调博弈的实验证据","abstract":"AI agents increasingly operate in multi-agent environments where outcomes depend on coordination. We distinguish primary algorithmic monoculture -- baseline action similarity -- from strategic algorithmic monoculture, whereby agents adjust similarity in response to incentives. We implement a simple experimental design that cleanly separates these forces, and deploy it on human and large language model (LLM) subjects. LLMs exhibit high levels of baseline similarity (primary monoculture) and, like humans, they regulate it in response to coordination incentives (strategic monoculture). While LLMs coordinate extremely well on similar actions, they lag behind humans in sustaining heterogeneity when divergence is rewarded.","authors":["Gonzalo Ballestero","Hadi Hosseini","Samarth Khanna","Ran I. Shorrer"],"categories":["cs.AI","cs.GT","cs.MA","econ.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-10","first_seen":"2026-04-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.09502","pdf_url":"https://arxiv.org/pdf/2604.09502","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","人类行为对照","协调博弈"],"reason":"用LLM和人类被试进行协调博弈实验，直接对比行为，属于经济学实验场景的人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":17,"question":"LLM智能体在协调博弈中如何协调行动，与人类相比有何差异？","design":"使用多个LLM（16个模型）和人类被试，在开放式问题（如说出一个字母、一个城市）上设置三种处理：picking（仅要求有效答案）、coordination（激励与同类型另一智能体答案相同）、divergence（激励答案不同），测量独立智能体间答案一致率。","baseline":"人类被试在相同实验任务中的行为数据。","findings":"LLM在无激励时答案一致性远高于人类，表现出高基础单一性；在协调激励下LLM能像人类一样调节一致性，但在需要差异化的任务中维持异质性的能力弱于人类。","reliability":"论文未讨论","relevance":"该研究直接对比LLM与人类在协调博弈中的行为，属于经济学实验场景的人类仿真，且揭示了LLM在差异化任务中的失效，高度契合研究者对仿真可靠性与偏差的关注。","inspiration":"借鉴其通过picking、coordination、divergence三种处理分离基础偏好与策略调整的实验设计，可清晰测量LLM的固有行为倾向与激励响应。｜可迁移至资产定价实验，研究LLM交易员在信息协调与差异化策略下的市场表现。｜以LLM为被试，设置picking（自由选股）、coordination（激励与另一LLM选相同股票）、divergence（激励选不同股票）三种处理，结果变量为投资组合相似度与市场效率指标，对照真实人类交易员在相同实验中的行为数据。"}},{"id":"2604.06663","version":1,"title":"Restoring Heterogeneity in LLM-based Social Simulation: An Audience Segmentation Approach","zh_title":"在基于大语言模型的社会模拟中恢复异质性：一种受众细分方法","abstract":"Large Language Models (LLMs) are increasingly used to simulate social attitudes and behaviors, offering scalable \"silicon samples\" that can approximate human data. However, current simulation practice often collapses diversity into an \"average persona,\" masking subgroup variation that is central to social reality. This study introduces audience segmentation as a systematic approach for restoring heterogeneity in LLM-based social simulation. Using U.S. climate-opinion survey data, we compare six segmentation configurations across two open-weight LLMs (Llama 3.1-70B and Mixtral 8x22B), varying segmentation identifier granularity, parsimony, and selection logic (theory-driven, data-driven, and instrument-based). We evaluate simulation performance with a three-dimensional evaluation framework covering distributional, structural, and predictive fidelity. Results show that increasing identifier granularity does not produce consistent improvement: moderate enrichment can improve performance, but further expansion does not reliably help and can worsen structural and predictive fidelity. Across parsimony comparisons, compact configurations often match or outperform more comprehensive alternatives, especially in structural and predictive fidelity, while distributional fidelity remains metric dependent. Identifier selection logic determines which fidelity dimension benefits most: instrument-based selection best preserves distributional shape, whereas data-driven selection best recovers between-group structure and identifier-outcome associations. Overall, no single configuration dominates all dimensions, and performance gains in one dimension can coincide with losses in another. These findings position audience segmentation as a core methodological approach for valid LLM-based social simulation and highlight the need for heterogeneity-aware evaluation and variance-preserving modeling strategies.","authors":["Xiaoyou Qin","Zhihong Li","Xiaoxiao Cheng"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-08","first_seen":"2026-04-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.06663","pdf_url":"https://arxiv.org/pdf/2604.06663","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM人类仿真","受众细分","社会模拟保真度"],"reason":"用LLM仿真人类气候态度，有真实调查数据对照，评估异质性恢复与保真度，并指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-08","rank":5,"question":"如何通过受众分割策略恢复大语言模型社会仿真中的异质性？","design":"使用Llama 3.1-70B和Mixtral 8x22B模型，基于美国气候态度调查数据，通过六种分割配置（变化标识符粒度、简约性和选择逻辑）生成合成样本，评估分布保真度、结构保真度和预测保真度。","baseline":"2025年10月通过Prolific收集的594份人类样本，按人口普查配额分层抽样，并使用Six Americas超短问卷分配受众细分标签。","findings":"增加标识符粒度并不一致提升性能；中等丰富度可改善，但过度扩展会损害结构和预测保真度。紧凑配置在结构和预测保真度上常优于或等同更全面的配置，而标识符选择逻辑决定哪个保真度维度受益最大。","reliability":"论文指出无单一配置在所有维度占优，一维度的提升可能伴随另一维度的损失；未明确讨论其他失效条件。","relevance":"直接命中研究者关注的LLM仿真人类实验、有真实人类对照、评估保真度并指出失效条件，强烈建议阅读原文。","inspiration":"该研究通过系统变换受众分割的标识符粒度、简约性和选择逻辑来评估仿真保真度，这种多维度配置对比的设计值得借鉴｜可迁移到消费者金融决策仿真，如退休储蓄选择或保险购买行为中的异质性偏好研究｜以LLM作为被试，施加不同信息框架（如损失vs收益表述）作为处理，测量储蓄率或保险购买意愿，并以美国消费者金融调查（SCF）或健康与退休研究（HRS）的真实个体数据作为对照基准"}},{"id":"2604.05939","version":1,"title":"Context-Value-Action Architecture for Value-Driven Large Language Model Agents","zh_title":"面向价值驱动大语言模型智能体的情境-价值-行动架构","abstract":"Large Language Models (LLMs) have shown promise in simulating human behavior, yet existing agents often exhibit behavioral rigidity, a flaw frequently masked by the self-referential bias of current \"LLM-as-a-judge\" evaluations. By evaluating against empirical ground truth, we reveal a counter-intuitive phenomenon: increasing the intensity of prompt-driven reasoning does not enhance fidelity but rather exacerbates value polarization, collapsing population diversity. To address this, we propose the Context-Value-Action (CVA) architecture, grounded in the Stimulus-Organism-Response (S-O-R) model and Schwartz's Theory of Basic Human Values. Unlike methods relying on self-verification, CVA decouples action generation from cognitive reasoning via a novel Value Verifier trained on authentic human data to explicitly model dynamic value activation. Experiments on CVABench, which comprises over 1.1 million real-world interaction traces, demonstrate that CVA significantly outperforms baselines. Our approach effectively mitigates polarization while offering superior behavioral fidelity and interpretability.","authors":["TianZe Zhang","Sirui Sun","Yuhang Xie","Xin Zhang","Zhiqiang Wu","Guojie Song"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-07","first_seen":"2026-04-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.05939","pdf_url":"https://arxiv.org/pdf/2604.05939","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["人类行为仿真","价值驱动智能体","算法保真度"],"reason":"用LLM仿真人类行为，有真实人类数据对照，解决行为僵化和价值极化问题，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":18,"question":"如何设计LLM智能体架构，以缓解提示驱动推理导致的行为僵化和价值极化，从而更真实地模拟人类行为？","design":"提出Context-Value-Action（CVA）架构，基于S-O-R模型和Schwartz基本人类价值观理论，将动作生成与认知推理解耦，引入在真实人类数据上训练的Value Verifier显式建模动态价值激活，并通过SFT和DPO对齐生成过程。","baseline":"CVABench包含超过110万条真实世界交互轨迹，来自15000多名人类参与者，用于评估行为逼真度和价值极化。","findings":"增加提示驱动的推理强度不仅未提高仿真逼真度，反而加剧价值极化并降低群体多样性；CVA架构有效缓解极化，在行为逼真度和可解释性上显著优于基线方法。","reliability":"论文未讨论","relevance":"高度相关：研究用LLM仿真人类行为，有大规模真实人类数据作为基准，揭示提示驱动方法的失效模式并提出改进架构，直接回应研究者对仿真可靠性、偏差和经济学/政策评估场景的关切。","inspiration":"借鉴CVA架构将行为生成与价值推理解耦，并引入在真实人类数据上训练的Value Verifier来动态建模价值激活，以此作为缓解行为僵化和价值极化的处理机制。｜可迁移到政策公告的预期形成实验，研究不同信息框架如何通过个体价值观影响通胀预期或就业预期。｜以LLM智能体为被试，处理组采用CVA架构注入特定价值观（如安全或自主），对照组使用标准提示驱动推理，结果变量为预期偏差和群体多样性，用真实调查数据（如密歇根大学消费者调查）作为人类基准对照。"}},{"id":"2605.20191","version":1,"title":"Shiny Stories, Hidden Struggles: Investigating the Representation of Disability Through the Lens of LLMs","zh_title":"光鲜故事，隐藏挣扎：通过LLM视角考察残障表征","abstract":"Modern Large Language Models (LLMs) have recently attracted much attention for their ability to simulate human behavior and generate text that reflects personas and demographic groups. While these capabilities can open up a multitude of diverse applications across fields, it is crucial to examine how such models represent various target groups since LLMs can perpetuate and amplify biases or discrimination against historically marginalized communities or, alternatively, as a result of debiasing efforts, overcorrect by portraying overly positive stereotypes. This overcompensation can idealize these groups, erasing the complexities and challenges they face in favor of unrealistic depictions. In this paper, we investigate how LLMs represent disability by simulating the perspectives of individuals with disabilities in generating social media posts. These posts are then compared with those written by real people with disabilities, focusing on emotional tone, sentiment, and representative words and themes. Our analysis reveals two key findings: (1) LLMs often idealize the experiences of people with disabilities, producing overly positive stereotypes that, despite appearing uplifting, fail to authentically capture their lived realities; and (2) a comparative analysis of posts simulating individuals with and without disabilities highlights a negative bias, where certain topics, such as career and entertainment, are disproportionately associated with nondisabled individuals. This reinforces exclusionary narratives and over-idealized portrayals of disability, misrepresenting the actual challenges faced by this community. These findings align with broader concerns and ongoing research showing that LLMs struggle to reflect the diverse realities of society, particularly the nuanced experiences of marginalized groups, and underscore the need for critical scrutiny of their representations.","authors":["Marco Bombieri","Simone Paolo Ponzetto","Marco Rospocher"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-04-02","first_seen":"2026-04-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.20191","pdf_url":"https://arxiv.org/pdf/2605.20191","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","残障表征","偏差评估"],"reason":"用LLM模拟残障人士发帖，并与真实人类数据对照，评估仿真偏差与失效条件，直接命…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":19,"question":"LLM生成的残障人士社交媒体帖子与真实残障人士的自我描述在情感、主题和语言上有何差异？","design":"使用多种LLM，通过提示词模拟残障人士和普通人在社交媒体上发帖，生成文本后自动标注情感、情绪和抑郁迹象，并与真实Reddit帖子进行对比分析。","baseline":"来自Reddit的真实残障人士自我介绍的帖子数据集。","findings":"LLM倾向于理想化残障人士的经历，产生过度积极的刻板印象，未能真实反映其生活挑战；同时，LLM在模拟残障与非残障个体时存在负面偏见，将职业、娱乐等主题不成比例地与非残障人士关联。","reliability":"论文指出LLM的正面理想化同样会造成伤害，且模型可能因去偏努力而过度补偿，但未详细讨论仿真失效的具体条件。","relevance":"该研究直接以真实人类数据为基准，评估LLM模拟残障人群的偏差与失效模式，属于批判性仿真研究，值得精读以了解LLM在边缘群体仿真中的局限。","inspiration":"该方法通过提示词让LLM模拟特定社会群体生成文本，并与真实社交媒体数据对比，可借鉴用于构建经济实验中的处理组与对照组仿真。｜它能迁移到信贷审批中的群体歧视研究，例如评估LLM在模拟不同种族或性别申请人时的语言偏差。｜可设计让LLM扮演贷款申请人撰写申请陈述，以真实银行信贷文本为基准，比较不同群体提示下的情感、主题和职业关联差异，检验仿真是否复现人类数据中的歧视模式。"}},{"id":"2604.01520","version":1,"title":"LLM Agents as Social Scientists: A Human-AI Collaborative Platform for Social Science Automation","zh_title":"作为社会科学家的LLM代理：一个面向社会科学自动化的人机协作平台","abstract":"Traditional social science research often requires designing complex experiments across vast methodological spaces and depends on real human participants, making it labor-intensive, costly, and difficult to scale. Here we present S-Researcher, an LLM-agent-based platform that assists researchers in conducting social science research more efficiently and at greater scale by \"siliconizing\" both the research process and the participant pool. To build S-Researcher, we first develop YuLan-OneSim, a large-scale social simulation system designed around three core requirements: generality via auto-programming from natural language to executable scenarios, scalability via a distributed architecture supporting up to 100,000 concurrent agents, and reliability via feedback-driven LLM fine-tuning. Leveraging this system, S-Researcher supports researchers in designing social experiments, simulating human behavior with LLM agents, analyzing results, and generating reports, forming a complete human-AI collaborative research loop in which researchers retain oversight and intervention at every stage. We operationalize LLM simulation research paradigms into three canonical reasoning modes (induction, deduction, and abduction) and validate S-Researcher through systematic case studies: inductive reproduction of cultural dynamics consistent with Axelrod's theory, deductive testing of competing hypotheses on teacher attention validated against survey data, and abductive identification of a cooperation mechanism in public goods games confirmed by human experiments. S-Researcher establishes a new human--AI collaborative paradigm for social science, in which computational simulation augments human researchers to accelerate discovery across the full spectrum of social inquiry.","authors":["Lei Wang","Yuanzi Li","Jinchao Wu","Heyang Gao","Xiaohe Bo","Xu Chen","Ji-Rong Wen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-02","first_seen":"2026-04-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.01520","pdf_url":"https://arxiv.org/pdf/2604.01520","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B3"],"tags":["LLM人类仿真","社会科学自动化","人机协作"],"reason":"用LLM代理模拟人类行为，复现文化动态、验证教师关注假设、识别合作机制，均有真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":53,"question":"如何构建一个以LLM代理为核心的人机协作平台，实现社会科学研究的全流程自动化，并验证其在归纳、演绎、溯因三种推理模式下的有效性？","design":"使用基于LLM的代理系统YuLan-OneSim模拟人类参与者，通过自然语言自动编程生成可执行场景，支持高达10万并发代理的分布式架构，并利用反馈驱动微调提升可靠性。在归纳模式下，模拟文化传播动态；在演绎模式下，模拟课堂师生互动并测试竞争假设；在溯因模式下，模拟公共物品博弈以识别合作机制。","baseline":"演绎案例验证教师关注假设时，对照了真实调查数据；溯因案例识别合作机制时，对照了真实人类实验。归纳案例无明确真实人类数据对照，仅与Axelrod理论预测一致。","findings":"S-Researcher平台成功复现了Axelrod文化传播理论中的收敛与极化动态，并在课堂模拟中通过对比调查数据验证了教师关注假设，还在公共物品博弈中识别出合作机制并经人类实验确认。该平台实现了从实验设计、模拟、分析到报告生成的全流程人机协作，支持研究者全程干预。","reliability":"论文未明确讨论仿真失效的具体条件或局限，仅通过反馈微调机制和案例验证来确保可靠性，但未系统分析代理行为与真实人类偏差的来源或边界。","relevance":"高度相关：该研究直接用LLM代理替代人类被试，复现社会动态、验证假设并识别机制，且部分案例有真实人类数据对照，契合研究者对仿真可靠性、经济学实验和政策评估场景的关注，值得精读原文以评估其方法细节与批判性局限。","inspiration":"借鉴其利用LLM代理进行大规模并发模拟和反馈驱动微调以提升行为真实性的方法，可构建经济实验的虚拟被试池｜可迁移至公共物品博弈、税收遵从或劳动供给决策等行为经济学与政策评估场景｜以LLM代理为被试，施加不同税收政策处理，测量其劳动供给或逃税行为，并与真实实验室实验或行政数据对照"}},{"id":"2603.27056","version":1,"title":"Persona-Based Simulation of Human Opinion at Population Scale","zh_title":"基于人格的群体意见仿真：从社交媒体推断半结构化人格以驱动LLM代理","abstract":"What does it mean to model a person, not merely to predict isolated responses, preferences, or behaviors, but to simulate how an individual interprets events, forms opinions, makes judgments, and acts consistently across contexts? This question matters because social science requires not only observing and predicting human outcomes, but also simulating interventions and their consequences. Although large language models (LLMs) can generate human-like answers, most existing approaches remain predictive, relying on demographic correlations rather than representations of individuals themselves. We introduce SPIRIT (Semi-structured Persona Inference and Reasoning for Individualized Trajectories), a framework designed explicitly for simulation rather than prediction. SPIRIT infers psychologically grounded, semi-structured personas from public social media posts, integrating structured attributes (e.g., personality traits and world beliefs) with unstructured narrative text reflecting values and lived experience. These personas prompt LLM-based agents to act as specific individuals when answering survey questions or responding to events. Using the Ipsos KnowledgePanel, a nationally representative probability sample of U.S. adults, we show that SPIRIT-conditioned simulations recover self-reported responses more faithfully than demographic persona and reproduce human-like heterogeneity in response patterns. We further demonstrate that persona banks can function as virtual respondent panels for studying both stable attitudes and time-sensitive public opinion.","authors":["Mao Li","Frederick G. Conrad"],"categories":["cs.CY","cs.AI","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-03-28","first_seen":"2026-03-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.27056","pdf_url":"https://arxiv.org/pdf/2603.27056","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","人格推断","调查方法"],"reason":"用LLM仿真个体意见并与全国概率样本对照，直接复现人类调查回答和异质性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"如何从社交媒体文本推断心理结构化的人物画像，并用其驱动大语言模型在总体层面仿真人类意见分布？","design":"从Ipsos KnowledgePanel概率样本中招募有Reddit/Twitter公开帖文的用户，用SPIRIT框架从帖文推断半结构化人物画像（含人格特质、世界信念等结构化属性与叙述文本），再以这些画像提示LLM代理回答调查问题，测量回复与真实自报答案的吻合度及异质性。","baseline":"Ipsos KnowledgePanel全国代表性概率样本的自报调查回答，以及仅用人口统计学画像的仿真作为对照。","findings":"SPIRIT画像驱动的仿真比人口统计学画像更准确地复现个体自报回答，并再现了人类回答模式中的异质性；人物画像库可作为虚拟受访者面板，用于研究稳定态度和时效性舆论。","reliability":"论文未讨论","relevance":"该研究直接以全国概率样本为基准，用LLM仿真个体意见分布，并对比人口统计学画像，高度契合研究者对仿真可靠性、基准对照和异质性复现的关注，值得精读原文。","inspiration":"借鉴其用社交媒体文本推断结构化人物画像并驱动LLM代理回答调查问题的设计，可构建高保真虚拟被试池｜可迁移到政策公告的预期形成研究，如央行沟通对通胀预期的影响｜招募有公开帖文的真实投资者，从其帖文推断人格与信念画像，用LLM代理接收不同措辞的央行声明，测量其通胀预期变化，以真实调查数据为基准对照"}},{"id":"2603.23884","version":1,"title":"POSIM: A Multi-Agent Simulation Framework for Social Media Public Opinion Evolution and Governance","zh_title":"POSIM：社交媒体舆论演化与治理的多智能体仿真框架","abstract":"Modeling social media public opinion evolution is essential for governance decision-making. Traditional epidemic models and rule-based agent-based models (ABMs) fail to capture the cognitive processes and adaptive behaviors of real users. Recent large language model (LLM)-based social simulations can reproduce group-level phenomena like polarization and conformity, yet remain unable to recreate the irrational interactions and multi-phase dynamics of real public opinion events. We present POSIM (Public Opinion Simulator), a multi-agent simulation framework for social media public opinion evolution and governance. POSIM integrates LLM-driven agents with a Belief--Desire--Intention (BDI) cognitive architecture that accounts for irrational factors, places them in a virtual social media environment with social networks and recommendation mechanisms, and drives temporal dynamics through a Hawkes point process engine that captures the co-evolution of agents and the environment across event phases. To validate the framework, we collect real-world public opinion datasets from the Weibo platform covering the full interaction chain of users. Experiments show that POSIM successfully reproduces key characteristics of public opinion evolution from individual mechanisms to collective phenomena, and its effectiveness is further supported by multiple statistical metrics. Building on POSIM, governance-oriented guidance and intervention experiments uncover a counterintuitive empathy paradox: empathetic guidance deepens negative sentiment instead of easing it under certain conditions, offering new insights for governance strategy design. These results demonstrate that the proposed framework can fully serve as a computational experimentation platform for proactive strategy evaluation and evidence-based governance. All source code is available at https://github.com/DeepCogLab/posim/.","authors":["Yongmao Zhang","Kai Qiao","Zhengyan Wang","Ningning Liang","Dekui Ma","Wenyao Sun","Jian Chen","Bin Yan"],"categories":["cs.GL"],"primary_category":"cs.GL","announce_type":"new","date":"2026-03-25","first_seen":"2026-03-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.23884","pdf_url":"https://arxiv.org/pdf/2603.23884","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","舆论演化","人类数据对照"],"reason":"用LLM agent模拟社交媒体舆论演化，并与真实微博数据对照，复现人类行为模…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":54,"question":"如何构建一个能复现真实社交媒体舆论演化多阶段动态并支持治理策略评估的多智能体仿真框架？","design":"用LLM驱动智能体，基于BDI认知架构融入情绪唤醒和认知偏差，模拟普通用户、意见领袖、媒体和政府四类角色；置于含社交网络和推荐机制的虚拟社交媒体环境中，通过Hawkes点过程引擎驱动多阶段时序演化；测量个体行为逻辑、群体涌现现象和统计指标，并进行治理干预实验。","baseline":"从微博平台收集的三个真实舆论事件数据集，覆盖原创、转发和评论的完整交互链。","findings":"POSIM在机制、现象和统计三个层面成功复现了舆论演化的关键特征；治理实验发现“共情悖论”：在某些条件下，共情引导反而加深负面情绪，而非缓解对立。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM智能体复现真实微博舆论事件，与人类数据对照，并评估治理策略效果，直接命中研究者对仿真可靠性、经济学/政策场景和批判性失效条件的兴趣。","inspiration":"借鉴其多智能体分层仿真设计，将LLM驱动的异质角色（如散户、机构、媒体、监管者）置于含推荐机制的信息环境中，通过Hawkes过程模拟信息传播与情绪演化，并设置治理干预实验来评估政策效果。｜可迁移到金融市场中的信息扩散与投资者情绪形成问题，例如研究社交媒体上的利好/利空消息如何通过不同渠道影响散户和机构的交易行为与市场波动。｜以LLM模拟散户和机构投资者作为被试，处理为在模拟社交平台中注入不同情绪基调的政策信号，结果变量为个体交易决策和市场价格波动，用真实微博金融舆情事件及同期市场交易数据作为对照基准。"}},{"id":"2603.19791","version":2,"title":"Text-Based Personas for Simulating User Privacy Decisions","zh_title":"基于文本角色模拟用户隐私决策","abstract":"The ability to simulate human privacy decisions has significant implications for aligning autonomous agents with individual intent and conducting cost-effective, large-scale privacy-centric user studies. Prior approaches prompt Large Language Models (LLMs) with natural language user statements, data-sharing histories, or demographic attributes to simulate privacy decisions. These approaches, however, fail to balance individual-level accuracy, human auditability, token efficiency, and population-level representation. We present Narriva, an approach that generates text-based synthetic privacy personas to address these shortcomings. Narriva grounds persona generation in prior user privacy decisions, such as those from large-scale survey datasets, rather than purely relying on demographic stereotypes. It compresses this data into concise, human-readable summaries structured by established privacy theories. Through benchmarking across five diverse datasets, we analyze the characteristics of Narriva's synthetic personas in modeling both individual and population-level privacy preferences. We find that grounding personas in past privacy behaviors achieves up to 87% predictive accuracy, improving over a non-personalized LLM baseline by 6-17 percentage points across datasets, while yielding an 80-95% reduction in prompt tokens compared to in-context learning with raw examples. Finally, we demonstrate that personas synthesized from a single survey can reproduce the aggregate privacy behaviors and statistical distributions of entirely different studies.","authors":["Kassem Fawaz","Ren Yi","Octavian Suciu","Rishabh Khandelwal","Hamza Harkous","Nina Taft","Marco Gruteser"],"categories":["cs.CR"],"primary_category":"cs.CR","announce_type":"new","date":"2026-03-20","first_seen":"2026-03-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.19791","pdf_url":"https://arxiv.org/pdf/2603.19791","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A5","B1","B2"],"tags":["隐私决策仿真","合成角色","人类数据对照"],"reason":"用LLM生成隐私决策合成样本，有真实人类数据对照，涉及用户研究场景，直接仿真人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":20,"question":"能否用基于文本的合成隐私人格（persona）有效模拟个体和群体层面的隐私决策，并泛化到独立研究？","design":"提出Narriva框架，利用大规模调查数据中用户的历史隐私决策生成结构化文本人格摘要，再让LLM基于这些人格模拟隐私偏好决策，评估个体预测准确率、群体分布复现能力和跨研究泛化性。","baseline":"五个不同数据集中的真实用户隐私决策数据，包括个体层面和群体层面的行为与态度。","findings":"基于过去隐私行为的人格在个体预测上准确率最高达87%，比无个性化LLM基线提升6-17个百分点；同时提示词token用量比原始示例上下文学习减少80-95%。从单一调查合成的人格能复现完全不同研究的群体隐私行为和统计分布。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM生成人格模拟隐私决策，有真实人类数据基准，评估个体与群体层面的仿真准确性和效率，并涉及跨研究泛化，直接回应研究者对LLM人类仿真实验、基准对照和可靠性批判的兴趣，值得精读原文。","inspiration":"借鉴Narriva框架用历史行为数据生成结构化文本人格以驱动LLM模拟个体决策的方法，可大幅降低提示词成本并提升预测准确率｜可迁移到消费者金融隐私偏好与数据共享决策研究，如移动支付或数字银行场景下的个人信息披露行为｜以真实用户历史隐私选择数据构建人格，让LLM模拟其在新型金融服务中的隐私权衡，结果变量为是否同意共享数据，用实际用户共享行为数据做对照"}},{"id":"2603.16142","version":2,"title":"Parametric Social Identity Injection and Diversification in Public Opinion Simulation","zh_title":"参数化社会身份注入与多样化在舆论仿真中的应用","abstract":"Large language models (LLMs) have recently been adopted as synthetic agents for public opinion simulation, offering a promising alternative to costly and slow human surveys. Despite their scalability, current LLM-based simulation methods fail to capture social diversity, producing flattened inter-group differences and overly homogeneous responses across demographic groups. We identify this limitation as a Diversity Collapse phenomenon in LLM hidden representations, where distinct social identities become increasingly indistinguishable across layers. Motivated by this observation, we propose Parametric Social Identity Injection (PSII), a general framework that injects explicit, parametric representations of demographic attributes and value orientations directly into intermediate hidden states of LLMs. Unlike prompt-based persona conditioning, PSII enables fine-grained and controllable identity modulation at the representation level. Extensive experiments on the World Values Survey using multiple open-source LLMs show that PSII significantly improves distributional fidelity and diversity, reducing KL divergence to real-world survey data while enhancing overall diversity. This work provides new insights into representation-level control of LLM agents and advances scalable, diversity-aware public opinion simulation.","authors":["Hexi Wang","Yujia Zhou","Bangde Du","Qingyao Ai","Yiqun Liu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-17","first_seen":"2026-03-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.16142","pdf_url":"https://arxiv.org/pdf/2603.16142","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","舆论模拟","社会身份注入"],"reason":"用LLM模拟公众舆论并与世界价值观调查真实数据对照，直接命中人类仿真核心。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":60,"question":"如何通过向LLM隐藏状态注入参数化社会身份来缓解舆论仿真中的多样性坍塌，从而更真实地复现人群异质性？","design":"使用多个开源LLM（Qwen2.5-7B/14B-Instruct、Llama-3.1-8B-Instruct、Mistral-24B-Instruct）作为合成智能体，模拟世界价值观调查（WVS）中的受访者；通过参数化社会身份注入（PSII）将人口统计属性和价值取向的参数向量直接注入LLM中间隐藏状态，对比传统提示词方法，测量回答分布与真实人类数据的KL散度及多样性指标。","baseline":"世界价值观调查（WVS）的真实人类回答数据，作为分布保真度和多样性的对照基准。","findings":"LLM隐藏状态在高层存在多样性坍塌现象，导致不同社会身份变得难以区分；PSII通过注入稳定身份向量并施加随机扰动，有效维持甚至提升高层表示的多样性，显著降低与真实调查数据的KL散度，增强群体间和群体内多样性。","reliability":"论文未讨论","relevance":"该研究直接用LLM复现世界价值观调查，以真实人类数据为基准，系统揭示了现有方法在多样性上的失效机制，并提出表示层面的干预方案，高度契合研究者对仿真可靠性、偏差及经济学/政策评估场景的关注，值得精读原文。","inspiration":"该方法通过向LLM隐藏状态注入参数化社会身份向量并施加随机扰动来维持群体多样性，可借鉴为一种处理异质性代理人的新范式，用于替代传统提示词方法以缓解仿真中的多样性坍塌｜该技术可迁移到政策评估中的异质性处理效应分析，例如模拟不同社会经济群体对税收改革或福利政策的反应差异，从而在事前评估政策分配的公平性与效率｜可设计一个实验：以LLM作为合成被试，注入收入、教育、政治倾向等身份参数，处理为不同税收政策方案，结果变量为政策支持度与预期行为变化，以真实调查数据（如美国综合社会调查GSS）作为分布对照基准"}},{"id":"2604.09609","version":2,"title":"General-purpose LLMs as Models of Human Driver Behavior: The Case of Simplified Merging","zh_title":"通用大语言模型作为人类驾驶行为模型：简化合流场景案例","abstract":"Human behavior models are essential as behavior references and for simulating human agents in virtual safety assessment of automated vehicles (AVs), yet current models face a trade-off between interpretability and flexibility. General-purpose large language models (LLMs) offer a promising alternative: a single model potentially deployable without parameter fitting across diverse scenarios. However, what LLMs can and cannot capture about human driving behavior remains poorly understood. We address this gap by embedding two general-purpose LLMs (OpenAI o3 and Google Gemini 2.5 Pro) as standalone, closed-loop driver agents in a simplified one-dimensional merging scenario and comparing their behavior against human data using quantitative and qualitative analyses. Both models reproduce human-like intermittent operational control and tactical dependencies on spatial cues. However, neither consistently captures the human response to dynamic velocity cues, and safety performance diverges sharply between models. A systematic prompt ablation study reveals that prompt components act as model-specific inductive biases that do not transfer across LLMs. These findings suggest that general-purpose LLMs could potentially serve as standalone, ready-to-use human behavior models in AV evaluation pipelines, but future research is needed to better understand their failure modes and ensure their validity as models of human driving behavior.","authors":["Samir H. A. Mohammad","Wouter Mooi","Arkady Zgonnikov"],"categories":["cs.AI","cs.RO"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-11","first_seen":"2026-03-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.09609","pdf_url":"https://arxiv.org/pdf/2604.09609","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类行为对照","驾驶行为建模"],"reason":"用LLM模拟人类驾驶行为并与真实数据对照，评估仿真可靠性与失效条件，方法可迁移…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":37,"question":"通用大语言模型在多大程度上能够复现人类驾驶行为，特别是在简化的一维合流场景中的操作控制、战术决策和安全表现？","design":"将两个通用大语言模型（OpenAI o3 和 Google Gemini 2.5 Pro）作为独立的闭环驾驶智能体，嵌入简化的一维合流任务中，不进行任何任务特定训练或参数拟合，通过定量和定性分析与人类数据对比，评估其行为相似性和安全性。","baseline":"使用先前研究中在相同场景和运动学条件下采集的人类驾驶数据集作为对照基准。","findings":"两个模型都能复现人类间歇性的操作控制和对空间线索的战术依赖，但都无法一致地捕捉人类对动态速度线索的反应，且两个模型之间的安全表现差异很大。提示消融实验表明，提示组件作为模型特定的归纳偏置，不能在不同大语言模型之间迁移。","reliability":"模型无法一致复现人类对动态速度线索的反应，安全性能在不同模型间差异显著，提示设计的效果不可迁移，且研究仅在简化的一维合流场景中进行，未涉及更复杂的交互场景。","relevance":"该研究直接以通用大语言模型作为人类被试的替代品，在驾驶行为仿真中与真实人类数据严格对照，并揭示了仿真的失效条件（如对速度线索的捕捉不足、模型间安全表现差异），高度契合研究者对经济学实验和政策评估场景中仿真可靠性与偏差的关注，值得精读原文以借鉴其方法和批判性发现。","inspiration":"该方法值得借鉴之处在于：使用通用LLM作为闭环智能体，在不进行任务特定训练的情况下直接与人类行为基准对照，并通过提示消融实验检验处理效应的可迁移性。｜可迁移至消费者跨期选择实验，研究LLM能否复现人类的时间偏好不一致和折现行为。｜研究设计：以LLM作为被试，施加不同表述方式的跨期选择任务（如延迟奖励的框架效应），结果变量为选择一致性和折现率，对照真实人类实验数据（如经典的双曲折现研究）。"}},{"id":"2603.09884","version":1,"title":"Benchmarking Political Persuasion Risks Across Frontier Large Language Models","zh_title":"跨前沿大语言模型的政治说服风险基准测试","abstract":"Concerns persist regarding the capacity of Large Language Models (LLMs) to sway political views. Although prior research has claimed that LLMs are not more persuasive than standard political campaign practices, the recent rise of frontier models warrants further study. In two survey experiments (N=19,145) across bipartisan issues and stances, we evaluate seven state-of-the-art LLMs developed by Anthropic, OpenAI, Google, and xAI. We find that LLMs outperform standard campaign advertisements, with heterogeneity in performance across models. Specifically, Claude models exhibit the highest persuasiveness, while Grok exhibits the lowest. The results are robust across issues and stances. Moreover, in contrast to the findings in Hackenburg et al. (2025b) and Lin et al. (2025) that information-based prompts boost persuasiveness, we find that the effectiveness of information-based prompts is model-dependent: they increase the persuasiveness of Claude and Grok while substantially reducing that of GPT. We introduce a data-driven and strategy-agnostic LLM-assisted conversation analysis approach to identify and assess underlying persuasive strategies. Our work benchmarks the persuasive risks of frontier models and provides a framework for cross-model comparative risk assessment.","authors":["Zhongren Chen","Joshua Kalla","Quan Le"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-10","first_seen":"2026-03-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.09884","pdf_url":"https://arxiv.org/pdf/2603.09884","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","政治说服","人类对照实验"],"reason":"用LLM生成政治说服信息，与真实人类调查实验对照，评估说服效果与策略，直接仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"前沿大语言模型在政治说服任务中是否比人类竞选广告更具说服力，且不同模型和提示策略的效果有何差异？","design":"用7款前沿LLM（Claude Sonnet 4、Gemini 2.5 Flash、GPT-4.1、Grok 4等）扮演政治说服者，在两项调查实验（N=19,145）中与真人被试进行文本对话，施加两种提示（普通提示与信息提示），测量被试对移民和最低工资议题的态度变化（五点李克特量表，后转为二元支持指标）。","baseline":"人类基准：移民议题使用真人倡导者视频，最低工资议题使用先前研究（Chen et al. 2025）中的人类竞选广告效果，并通过Hajek估计量校正协变量偏移。","findings":"所有LLM的说服效果均显著强于人类竞选广告，其中Claude模型说服力最强，Grok最弱；信息提示的效果因模型而异，能提升Claude和Grok的说服力，但大幅降低GPT的说服力。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查实验为基准，评估LLM在政治说服场景中的仿真效果与模型间异质性，并揭示提示策略的模型依赖性，对关注LLM仿真可靠性及失效条件的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其用LLM替代人类进行文本对话干预并对比真实人类基准的设计，通过多模型比较和提示策略操纵揭示效果异质性｜可迁移至政策沟通场景，如央行前瞻指引或财政政策公告对公众预期和消费行为的影响｜以LLM作为虚拟被试，随机分配不同风格的货币政策沟通文本（处理），测量其预期的通胀或消费意愿变化（结果），并以历史调查数据或真实实验数据作为对照基准"}},{"id":"2603.08853","version":1,"title":"LLM-Agent Interactions on Markets with Information Asymmetries","zh_title":"信息不对称市场中LLM智能体的互动研究","abstract":"As AI agents increasingly act on behalf of human stakeholders in economic settings, understanding their behavior in complex market environments becomes critical. This article examines how Large Language Models coordinate on markets that are characterized by information asymmetries and in which providers of services have incentives to exploit that asymmetry for their own economic gain. To that end, we conduct simulations with GPT-5.1 agents in credence goods markets, manipulating the institutional framework (free market, verifiability, liability), LLM agent's social preferences (default, self-interested, inequity-averse, efficiency-loving), and reputation mechanisms across one-shot and repeated 16-round interactions. In one-shot settings, LLM agents largely fail to establish cooperation, with markets breaking down except under liability rules or when experts have efficiency-loving preferences. Repeated interactions solve consumer participation through competitive price reduction, but expert fraud remains entrenched absent explicit other-regarding preferences. LLM consumers focus narrowly on price levels rather than understanding strategic incentives embedded in markups, making them vulnerable to exploitation. Compared to human experiments, LLM markets exhibit substantially higher consumer participation but much greater market concentration, lower prices, and more polarized fraud patterns. The effect of institutions like verifiability and reputation is also much more ambiguous. Surplus shifts dramatically toward consumers under social-preference objectives. These findings suggest that institutional design for AI agent markets requires fundamentally different approaches than those effective for human actors, with social preference alignment emerging as the primary determinant of market efficiency.","authors":["Alexander Erlei","Lukas Meub"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-03-09","first_seen":"2026-03-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.08853","pdf_url":"https://arxiv.org/pdf/2603.08853","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","市场实验","人类数据对照"],"reason":"用LLM agent模拟信息不对称市场，并与人类实验数据对照，评估制度与偏好影…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":3,"question":"LLM智能体在信息不对称的信任品市场中如何协调行为，社会偏好与制度如何影响市场效率？","design":"用GPT-5.1扮演专家和消费者，在信任品市场博弈中模拟交易，操纵制度框架（自由市场、可验证性、责任规则）、LLM社会偏好（默认、自利、不平等厌恶、效率偏好）和声誉机制，进行单次与16轮重复互动，测量市场参与、欺诈行为、价格、福利分配等。","baseline":"对照Dulleck, Kerschbamer, and Sutter (2011)的人类实验数据。","findings":"单次互动中LLM市场普遍崩溃，仅责任规则或效率偏好下能维持；重复互动通过降价解决消费者参与，但专家欺诈依然顽固，仅社会偏好能抑制欺诈。与人类实验相比，LLM市场消费者参与更高但集中度更高、价格更低、欺诈模式更极化，制度效果更模糊，且剩余大幅向消费者转移。","reliability":"论文未讨论","relevance":"直接使用LLM模拟信息不对称市场并与真实人类实验基准对照，系统评估制度与偏好影响，揭示LLM仿真在行为模式、制度效应上与人类的显著偏差，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读。","inspiration":"借鉴其操纵制度框架（自由市场、可验证性、责任规则）和LLM社会偏好（自利、不平等厌恶、效率偏好）的多因素实验设计，并与真实人类实验基准对照，系统评估行为偏差。｜可迁移到信贷审批中的信息不对称与歧视问题，如银行信贷员与小微企业的贷款博弈，检验不同监管规则和银行社会偏好对审批决策的影响。｜用LLM扮演信贷员和小微企业主，操纵信贷审核制度（纯市场、强制信息披露、责任追究）和LLM偏好，测量贷款批准率、利率设定、违约欺诈率，以真实银行信贷实验数据为基准对照。"}},{"id":"2604.15329","version":1,"title":"Evaluating LLMs as Human Surrogates in Controlled Experiments","zh_title":"评估大语言模型作为受控实验中人类替代品的有效性","abstract":"Large language models (LLMs) are increasingly used to simulate human responses in behavioral research, yet it remains unclear when LLM-generated data support the same experimental inferences as human data. We evaluate this by directly comparing off-the-shelf LLM-generated responses with human responses from a canonical survey experiment on accuracy perception. Each human observation is converted into a structured prompt, and models generate a single 0--10 outcome variable without task-specific training; identical statistical analyses are applied to human and synthetic responses. We find that LLMs reproduce several directional effects observed in humans, but effect magnitudes and moderation patterns vary across models. Off-the-shelf LLMs therefore capture aggregate belief-updating patterns under controlled conditions but do not consistently match human-scale effects, clarifying when LLM-generated data can function as behavioral surrogates.","authors":["Adnan Hoq","Tim Weninger"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"new","date":"2026-03-08","first_seen":"2026-03-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.15329","pdf_url":"https://arxiv.org/pdf/2604.15329","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类替代","实验对照"],"reason":"直接比较LLM与人类在受控实验中的反应，评估仿真可靠性，有真实人类数据对照，并…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":4,"question":"在受控实验中，现成的大语言模型（LLM）生成的回答能否支持与人类数据相同的实验推断？","design":"将人类被试在政治新闻准确性感知调查实验中的每次观测转化为结构化提示（含人物描述、实验条件和新闻标题），让多个现成LLM（闭源与开源）直接生成0-10的准确性评分，不进行任务特定训练或校准，然后对人类和合成数据应用相同的统计分析。","baseline":"真实人类被试在相同实验中的回答，实验操控了新闻标题的政治倾向和是否提供AI可信度反馈。","findings":"LLM能复现人类数据中的若干方向性效应（如意识形态对齐和可信度反馈的影响），但效应大小和调节模式因模型而异；LLM捕捉了受控条件下的总体信念更新模式，但未能一致匹配人类量级的效应。","reliability":"效应量级在不同模型间差异大，部分模型夸大处理效应；新闻级异质性仅部分复现；LLM生成的数据不能直接替代人类样本，结构复现需逐假设检验。","relevance":"该研究直接比较LLM与人类在受控实验中的反应，有真实人类数据对照，并明确指出了仿真在效应量级和异质性上的失效条件，高度契合研究者对LLM仿真可靠性及批判性评估的关注。","inspiration":"该方法将人类实验的每次观测转化为结构化提示，直接让现成LLM生成评分，并与人类数据应用相同统计分析，值得借鉴其逐观测仿真和严格对照的设计｜可迁移到政策公告对消费者预期形成的影响研究，例如评估央行沟通对通胀预期的作用｜以消费者为被试，处理为不同措辞的央行公告，结果变量为通胀预期数值，用真实消费者调查数据作为对照基准"}},{"id":"2604.22756","version":1,"title":"Your Reviews Replicate You: LLM-Based Agents as Customer Digital Twins for Conjoint Analysis","zh_title":"你的评论复制你：基于LLM的客户数字孪生用于联合分析","abstract":"Conjoint analysis is a cornerstone of market research for estimating consumer preferences; however, traditional methods face persistent challenges regarding time, cost, and respondent fatigue. To address these limitations, this study proposes a framework that utilizes large language model (LLM)-based \"customer digital twins (CDT)\" as virtual respondents. We identified active users within the Reddit community and aggregated their comprehensive review histories to construct individualized vector databases. By integrating retrieval-augmented generation (RAG) with prompt engineering, this study developed customer agents capable of dynamically retrieving and reasoning upon their specific past preferences and constraints. These customer agents, called CDTs, performed pairwise comparison tasks on product profiles generated via fractional factorial design, and the resulting choice data was analyzed to estimate part-worth utilities by logistic regression. Empirical validation demonstrates that these CDTs predict the preferences of actual users with 87.73% accuracy. Furthermore, a case study on the computer monitor category successfully quantified trade-offs between attributes such as panel type and resolution, deriving preference structures consistent with market realities. Ultimately, this study contributes to marketing research by presenting a scalable alternative that significantly improves both agility and cost-efficiency to traditional methods.","authors":["Bin Xuan","Jungmin Hwang","Hakyeon Lee"],"categories":["cs.IR","cs.AI"],"primary_category":"cs.IR","announce_type":"new","date":"2026-03-06","first_seen":"2026-03-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.22756","pdf_url":"https://arxiv.org/pdf/2604.22756","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","消费者偏好","数字孪生"],"reason":"用LLM代理模拟消费者偏好，有真实用户数据对照，准确率87.73%，属经济学实…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":18,"question":"能否利用基于大语言模型的客户数字孪生（CDT）替代真实人类受访者进行联合分析，以准确复现个体消费者的偏好选择？","design":"使用GPT-4等大语言模型，结合检索增强生成（RAG）和提示工程，基于Reddit用户的历史评论构建个体化向量数据库，生成客户数字孪生（CDT）作为虚拟受访者；让CDT对通过部分因子设计生成的产品配置文件进行成对比较选择任务，收集选择数据，再用逻辑回归估计部分效用值。","baseline":"以Reddit社区中真实活跃用户的实际偏好作为对照基准，验证CDT预测的准确率。","findings":"CDT预测真实用户偏好的准确率达到87.73%；在电脑显示器案例中，成功量化了面板类型与分辨率等属性间的权衡，得出的偏好结构与市场现实一致。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM代理复现真实消费者偏好，有真实人类数据对照，属于经济学实验场景，值得精读以评估仿真可靠性。","inspiration":"可借鉴其利用个体历史文本数据构建个性化代理并通过成对比较任务测量偏好的方法｜可迁移到消费者跨期选择实验，如研究折扣率或耐心程度｜以电商平台用户评论构建LLM代理作为被试，施加不同跨期奖励方案（如立即小奖 vs. 延迟大奖），测量选择结果，并以该用户真实历史购买决策中的时间偏好数据作为对照。"}},{"id":"2603.03585","version":2,"title":"Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility","zh_title":"Belief-Sim：面向信念驱动的人口统计错误信息易感性仿真","abstract":"Misinformation is a growing societal threat, and susceptibility to misinformative claims varies across demographic groups due to differences in underlying beliefs. As Large Language Models (LLMs) are increasingly used to simulate human behaviors, we investigate whether they can simulate demographic misinformation susceptibility, treating beliefs as a primary driving factor. We introduce BeliefSim, a simulation framework that constructs demographic belief profiles using psychology-informed misinformation taxonomies and survey priors. We study prompt-based conditioning and post-training adaptation, and conduct a multi-fold evaluation using: (i) susceptibility alignment and (ii) counterfactual demographic sensitivity. Across both datasets and modeling strategies, we show that beliefs provide a strong prior for simulating misinformation susceptibility, with alignment up to 92%.","authors":["Angana Borah","Zohaib Khan","Rada Mihalcea","Verónica Pérez-Rosas"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-03","first_seen":"2026-03-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.03585","pdf_url":"https://arxiv.org/pdf/2603.03585","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM人类仿真","错误信息易感性","人口统计差异"],"reason":"用LLM仿真不同人口群体对错误信息的易感性，以信念为驱动，并与真实人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":21,"question":"信念是否能改善基于人口统计学的错误信息易感性仿真？","design":"使用多个LLM，通过提示条件化（BeliefSim-PC）和微调（BeliefSim-FT）两种方式，基于心理学启发的信念分类和调查先验构建人口信念画像，模拟不同性别、年龄、居住地、教育水平群体的错误信息易感性，测量其与真实人类判断的对齐程度。","baseline":"PANDORA数据集（318人，每人3条声明）和MIST-1数据集（409人，每人100条声明），共13.8K条真实人类判断，包含人口统计信息。","findings":"信念是模拟错误信息易感性的强先验，对齐度最高达92%；仅使用人口统计信息不可靠，可能导致捷径依赖。","reliability":"人口统计建模主要为单轴，仅评估8个独立群体，未考虑交叉群体效应；数据集仅限于美国参与者和英文标题。","relevance":"该研究直接用LLM仿真人口群体的错误信息易感性，以信念为驱动，并与真实人类数据对照，包含批判性分析，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"该方法通过心理学信念分类构建人口信念画像，并对比两种LLM条件化方式（提示与微调）来仿真群体判断，提供了处理施加与对照设计的参考｜可迁移至信贷审批中的群体歧视研究，仿真不同人口群体对贷款申请的审批决策｜以LLM作为被试，处理为基于信念画像的条件化提示或微调，结果变量为审批通过率，对照真实银行信贷审批数据中的群体差异"}},{"id":"2602.15173","version":2,"title":"Mind the (DH) Gap! A Contrast in Risky Choices Between Reasoning and Conversational LLMs","zh_title":"注意(DH)差距！推理型与对话型LLM在风险选择上的对比","abstract":"The use of large language models either as decision support systems, or in agentic workflows, is rapidly transforming the digital ecosystem. However, the understanding of LLM decision-making under uncertainty remains limited. We study LLM risky choices along two dimensions: (1) prospect representation (based on an explicit representation or outcome history) and (2) decision rationale (explanation). Our study, which involves 20 frontier and open LLMs, is complemented by a matched human subjects experiment, which provides one reference point, while an expected payoff maximizing rational agent model provides another. We find that LLMs cluster into two categories: reasoning models (RMs) and conversational models (CMs). RMs tend towards rational behavior, are insensitive to the order of prospects, gain/loss framing, and explanations, and behave similarly whether prospects are explicit or presented via a history of outcomes. CMs are significantly less rational, slightly more human-like, sensitive to prospect ordering, framing, and explanation, and exhibit a large description-history gap. Paired comparisons of open LLMs suggest that a key factor differentiating RMs and CMs is training for mathematical reasoning.","authors":["Luise Ge","Yongyan Zhang","Yevgeniy Vorobeychik"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-16","first_seen":"2026-02-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.15173","pdf_url":"https://arxiv.org/pdf/2602.15173","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","风险决策","人类对照实验"],"reason":"用LLM仿真人类风险决策，并与真人实验对照，评估模型行为偏差与人类相似度。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":40,"question":"LLM在风险决策中如何受前景表征方式（描述 vs. 历史结果）和决策理由（解释）的影响，并与人类及理性基准对比？","design":"用20个前沿和开源LLM作为被试，在三种基础前景对上完成风险选择任务，操纵前景表征（显式描述 vs. 结果历史）和解释要求（无解释/简短解释/数学解释），测量选择行为，并与匹配的人类被试实验和期望收益最大化理性智能体对比。","baseline":"匹配的人类被试实验，以及期望收益最大化的理性智能体模型。","findings":"LLM分为推理模型和对话模型两类：推理模型接近理性人，对前景顺序、框架和解释不敏感，描述-历史差距小；对话模型理性较低，略似人类，对顺序、框架和解释敏感，描述-历史差距大。数学推理训练是区分两类模型的关键因素。","reliability":"论文未讨论","relevance":"该研究直接以LLM仿真人类风险决策，并与真人实验和理性基准对照，系统评估了模型行为偏差和人类相似度，完全契合研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"借鉴其系统操纵前景表征（描述 vs. 历史结果）和解释要求（无/简短/数学）来分离模型行为偏差的设计，并设置理性基准与人类被试双对照｜可迁移到金融风险偏好评估场景，如投资者在历史收益序列与文字描述下的风险选择差异｜以LLM为被试，处理为用历史收益率序列 vs. 文字描述呈现股票前景，结果变量为风险资产配置比例，对照真实投资者调查数据（如Survey of Consumer Finances）"}},{"id":"2602.14043","version":1,"title":"Beyond Static Snapshots: Dynamic Modeling and Forecasting of Group-Level Value Evolution with Large Language Models","zh_title":"超越静态快照：基于大语言模型的群体价值观动态建模与预测","abstract":"Social simulation is critical for mining complex social dynamics and supporting data-driven decision making. LLM-based methods have emerged as powerful tools for this task by leveraging human-like social questionnaire responses to model group behaviors. Existing LLM-based approaches predominantly focus on group-level values at discrete time points, treating them as static snapshots rather than dynamic processes. However, group-level values are not fixed but shaped by long-term social changes. Modeling their dynamics is thus crucial for accurate social evolution prediction--a key challenge in both data mining and social science. This problem remains underexplored due to limited longitudinal data, group heterogeneity, and intricate historical event impacts. To bridge this gap, we propose a novel framework for group-level dynamic social simulation by integrating historical value trajectories into LLM-based human response modeling. We select China and the U.S. as representative contexts, conducting stratified simulations across four core sociodemographic dimensions (gender, age, education, income). Using the World Values Survey, we construct a multi-wave, group-level longitudinal dataset to capture historical value evolution, and then propose the first event-based prediction method for this task, unifying social events, current value states, and group attributes into a single framework. Evaluations across five LLM families show substantial gains: a maximum 30.88\\% improvement on seen questions and 33.97\\% on unseen questions over the Vanilla baseline. We further find notable cross-group heterogeneity: U.S. groups are more volatile than Chinese groups, and younger groups in both countries are more sensitive to external changes. These findings advance LLM-based social simulation and provide new insights for social scientists to understand and predict social value changes.","authors":["Qiankun Pi","Guixin Su","Jinliang Li","Mayi Xu","Xin Miao","Jiawei Jiang","Ming Zhong","Tieyun Qian"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-02-15","first_seen":"2026-02-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.14043","pdf_url":"https://arxiv.org/pdf/2602.14043","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","价值观演化","社会模拟"],"reason":"用LLM仿真群体价值观动态，有真实世界价值观调查数据对照，涉及社会变迁预测。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":55,"question":"如何利用大语言模型动态建模和预测群体价值观的长期演变趋势？","design":"基于世界价值观调查（WVS）第5-7波数据，按性别、年龄、教育、收入四个维度将中美人群划分为群体，用LLM微调学习历史价值观轨迹以预测未来价值观，并提出事件感知预测方法，将社会事件与价值观表示对齐，让LLM推理事件影响。","baseline":"世界价值观调查（WVS）第5、6、7波的真实群体级纵向数据，包含中国28个群体和美国23个群体。","findings":"事件感知预测方法在已见问题上最高提升30.88%，在未见问题上最高提升33.97%；美国群体价值观波动性显著高于中国，两国年轻群体对外部变化更敏感。","reliability":"论文未讨论","relevance":"该研究用LLM仿真群体价值观动态演变，有真实WVS纵向数据作为基准，涉及中美群体异质性和事件驱动的价值观变化，属于经济学/政策评估场景下的LLM人类仿真，与研究者关注高度契合，值得精读原文。","inspiration":"该方法将社会事件编码为与价值观表示对齐的向量，让LLM推理事件对群体价值观的动态影响，这种事件感知的预测设计值得借鉴｜可迁移到政策公告对消费者信心或通胀预期的影响评估，例如研究央行沟通事件如何改变不同群体的预期形成过程｜用LLM模拟不同收入/年龄群体的消费者，施加央行声明作为处理，测量其通胀预期变化，以密歇根大学消费者调查的真实群体数据作为对照基准"}},{"id":"2602.13862","version":2,"title":"Measuring Self-Rating Bias in LLM-Generated Survey Data: A Semantic Similarity Framework for Independent Scale Mapping","zh_title":"测量LLM生成调查数据中的自评偏差：一种独立量表映射的语义相似度框架","abstract":"Synthetic survey data generated by large language models (LLMs) suffers from a fundamental circularity: the same model family that generates text responses also maps them to numerical scales. We calibrate and validate Semantic Similarity Rating (SSR; Maier et al., 2024), which decouples generation from scale mapping via embedding-based cosine similarity against predefined anchor statements. Configuration experiments (N=17 pilot, N=69 cross-validation across 8 domains) show that naturalistic behavioral anchors outperform formal jargon by 29 percentage points (pp), and that SSR achieves 65-67% exact match and 91% within plus/minus 1; a cross-model test with OpenAI text-embedding-3-small reaches 77% exact, confirming cross-provider generalization. Direct LLM baselines (Claude 87%, GPT-4o 83%) establish that SSR's contribution is methodological independence, not accuracy superiority. A control condition removing question text from the LLM prompt actually improves LLM accuracy, ruling out information asymmetry as the explanation for SSR's lower accuracy. A pre-registered circularity experiment (N=345) reveals 4x compressed error variance in LLM rating (sigma^2 = 0.21 vs 0.87 for SSR) and systematic directional bias. A cross-model control (GPT-4o rating Claude-generated text) shows nearly identical compression (within/cross ratio = 0.93), indicating variance compression is a general LLM property rather than a within-model artifact. The calibration dataset, anchor library, and source code are publicly available (see Data Availability).","authors":["Eduardo Vera Pichardo"],"categories":["physics.soc-ph"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-02-14","first_seen":"2026-02-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.13862","pdf_url":"https://arxiv.org/pdf/2602.13862","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查数据","偏差评估"],"reason":"评估LLM生成调查数据的自评偏差，提出独立量表映射方法，有真实人类数据对照，批…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":33,"question":"如何通过语义相似度框架独立测量LLM生成调查数据中的自评偏差，并解决文本生成与量表映射的循环性问题？","design":"本研究并非仿真人类被试，而是校准和验证语义相似度评分（SSR）框架：先用LLM（Claude Haiku 4.5）根据人格描述和问题生成文本回答，再用独立的嵌入模型（Voyage AI voyage-3.5-lite）通过余弦相似度将文本映射到预定义的锚定语句量表上，并通过配置实验、交叉验证、LLM基线对比和预注册的循环性实验评估该框架的性能与偏差。","baseline":"无对照","findings":"SSR框架实现了生成与测量的架构独立，自然行为锚定比正式术语锚定准确率高29个百分点，交叉验证中精确匹配率达65-67%，±1内达91%；直接LLM评分虽准确率更高（Claude 87%，GPT-4o 83%），但存在4倍误差方差压缩和系统性方向偏差，且该方差压缩是LLM的普遍属性而非同模型伪影。","reliability":"论文承认SSR的准确率低于直接LLM评分，且嵌入模型与生成模型可能因共享预训练语料而存在残余相关性；锚定语句的领域特异性导致跨领域准确率差异大（33-90%），需进一步优化锚定库。","relevance":"该研究直接针对LLM仿真调查数据中的循环性测量偏差，提出独立量表映射方法，并系统揭示了LLM自评的方差压缩和方向偏差，对关注仿真可靠性与失效条件的研究者具有重要参考价值，值得阅读原文。","inspiration":"该方法通过独立嵌入模型将LLM文本回答映射到预定义量表，避免了直接让LLM自评的循环偏差，这种测量与生成解耦的设计值得借鉴｜可迁移到消费者信心调查或通胀预期测量中，用LLM模拟受访者对经济前景的开放式回答，再独立映射到信心指数或预期值｜用LLM根据人口特征生成对经济前景的文本描述，以独立语义模型映射为预期通胀值，与密歇根消费者调查的真实个体数据对比，检验仿真偏差与方差压缩"}},{"id":"2602.11939","version":1,"title":"Do Large Language Models Adapt to Language Variation across Socioeconomic Status?","zh_title":"大语言模型能否适应社会经济地位带来的语言变异？","abstract":"Humans adjust their linguistic style to the audience they are addressing. However, the extent to which LLMs adapt to different social contexts is largely unknown. As these models increasingly mediate human-to-human communication, their failure to adapt to diverse styles can perpetuate stereotypes and marginalize communities whose linguistic norms are less closely mirrored by the models, thereby reinforcing social stratification. We study the extent to which LLMs integrate into social media communication across different socioeconomic status (SES) communities. We collect a novel dataset from Reddit and YouTube, stratified by SES. We prompt four LLMs with incomplete text from that corpus and compare the LLM-generated completions to the originals along 94 sociolinguistic metrics, including syntactic, rhetorical, and lexical features. LLMs modulate their style with respect to SES to only a minor extent, often resulting in approximation or caricature, and tend to emulate the style of upper SES more effectively. Our findings (1) show how LLMs risk amplifying linguistic hierarchies and (2) call into question their validity for agent-based social simulation, survey experiments, and any research relying on language style as a social signal.","authors":["Elisa Bassignana","Mike Zhang","Dirk Hovy","Amanda Cercas Curry"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-02-12","first_seen":"2026-02-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.11939","pdf_url":"https://arxiv.org/pdf/2602.11939","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","社会语言学","算法保真度"],"reason":"直接评估LLM仿真人类语言行为的效度，有真实人类数据对照，并指出仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":56,"question":"大语言模型在多大程度上能根据社会经济地位（SES）社区的语言变异调整其生成文本的风格？","design":"本研究并非基于智能体的仿真实验，而是通过提示工程让四个大语言模型补全来自Reddit和YouTube的、按SES分层的人类文本片段，然后比较模型生成文本与原始文本在94项社会语言学指标上的差异。","baseline":"从Reddit和YouTube收集的真实社交媒体文本，按SES（通过主题关键词和网络分析等策略）分层，作为人类语言风格的对照基准。","findings":"大语言模型仅能微弱地根据SES调节风格，常表现为近似或夸张模仿，且更擅长模仿高SES社区的风格。这种风格适应不足可能放大语言层级差异，并质疑LLM在基于智能体的社会模拟、调查实验等依赖语言风格作为社会信号的研究中的有效性。","reliability":"论文指出LLM在风格适应上存在近似或夸张模仿而非精确复现的问题，且对低SES风格模拟较差；初步消融实验显示，当提供更长上下文时，模型更倾向于适应高SES风格，这揭示了仿真失效的条件。","relevance":"该研究直接评估了LLM模拟不同社会经济地位人群语言行为的效度，有真实人类数据对照，并明确指出仿真在风格适应上的失效条件，与研究者关注的人类仿真可靠性及批判性评估高度契合，值得精读原文。","inspiration":"该方法通过提示工程让LLM补全按SES分层的人类文本，并对比94项社会语言学指标，可借鉴其分层对照与多维度风格测量来评估LLM的仿真偏差。｜可迁移到信贷审批中的语言歧视研究，例如分析LLM模拟不同SES申请人的贷款申请文本时是否系统性地偏向高SES风格，从而影响审批决策。｜以LLM为被试，给定不同SES背景的贷款申请场景提示，生成申请文本；结果变量为文本的语言风格指标（如正式度、情感词频）；以真实银行或P2P平台中不同SES申请人的贷款申请文本作为对照基准。"}},{"id":"2602.09362","version":1,"title":"Behavioral Economics of AI: LLM Biases and Corrections","zh_title":"人工智能的行为经济学：大语言模型的偏差与校正","abstract":"Do generative AI models, particularly large language models (LLMs), exhibit systematic behavioral biases in economic and financial decisions? If so, how can these biases be mitigated? Drawing on the cognitive psychology and experimental economics literatures, we conduct the most comprehensive set of experiments to date$-$originally designed to document human biases$-$on prominent LLM families across model versions and scales. We document systematic patterns in LLM behavior. In preference-based tasks, responses become more human-like as models become more advanced or larger, while in belief-based tasks, advanced large-scale models frequently generate rational responses. Prompting LLMs to make rational decisions reduces biases.","authors":["Pietro Bini","Lin William Cong","Xing Huang","Lawrence J. Jin"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-10","first_seen":"2026-02-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.09362","pdf_url":"https://arxiv.org/pdf/2602.09362","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM行为偏差","人类仿真","实验经济学"],"reason":"用人类实验范式测LLM行为偏差并对比人类数据，直接评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":19,"question":"生成式AI模型（尤其是大语言模型）在经济金融决策中是否表现出系统性行为偏差？如何纠正这些偏差？","design":"本研究并非用LLM仿真人类被试，而是将LLM本身作为研究对象。从认知心理学和实验经济学文献中选取原本用于测量人类偏差的实验问题，改编为提示词，通过API收集OpenAI ChatGPT、Anthropic Claude、Google Gemini和Meta Llama四个模型家族在不同版本和规模下的回答，分析其行为模式。","baseline":"以原有人类实验中的理性基准和真实人类回答作为对照。","findings":"在偏好类任务中，模型越先进或规模越大，回答越像人类且偏离理性；在信念类任务中，先进大规模模型则多给出理性回答。不同模型家族间存在显著异质性，如Gemini在偏好问题上比ChatGPT更不理性、更像人，而Llama在信念问题上理性程度较低。","reliability":"论文未讨论","relevance":"该研究直接使用人类行为实验范式测量LLM的偏差，并系统对比了LLM与人类及理性基准的差异，为评估LLM作为人类仿真被试的可靠性提供了关键证据，高度相关，值得精读。","inspiration":"借鉴其将经典行为经济学实验范式直接移植到LLM测试中的方法，可系统评估模型在不同决策场景下的行为一致性。｜可迁移到资产定价实验中的投资者偏差测量，如过度外推、过度自信等。｜以GPT-4等LLM为被试，呈现历史股价序列并要求预测未来收益，测量其外推倾向，并与真实投资者调查数据（如Shiller投资者信心调查）进行对照。"}},{"id":"2602.09802","version":2,"title":"Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices","zh_title":"大语言模型会为景观多付钱吗？从主观选择推断支付意愿","abstract":"As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often required to make subjective choices on behalf of users in settings where no objectively correct answer exists. We study LLM decision-making in a travel-assistant context by presenting models with choice dilemmas and analyzing their responses using multinomial logit models to derive implied willingness to pay (WTP) estimates. These WTP values are subsequently compared to human benchmark values from the economics literature. In addition to a baseline setting, we examine how model behavior changes under more realistic conditions, including the provision of information about users' past choices and persona-based prompting. Our results show that while meaningful WTP values can be derived for larger LLMs, they also display systematic deviations at the attribute level. Additionally, they tend to overestimate human WTP overall, particularly when expensive options or business-oriented personas are introduced. Conditioning models on prior preferences for cheaper options yields valuations that are closer to human benchmarks. Overall, our findings highlight both the potential and the limitations of using LLMs for subjective decision support and underscore the importance of careful model selection, prompt design, and user representation when deploying such systems in practice.","authors":["Manon Reusens","Sofie Goethals","Toon Calders","David Martens"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-10","first_seen":"2026-02-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.09802","pdf_url":"https://arxiv.org/pdf/2602.09802","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","支付意愿","人类数据对照"],"reason":"用LLM模拟人类支付意愿并与真实人类数据对照，评估偏差，涉及经济学场景。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"在旅行助手场景中，LLM的主观选择能否通过离散选择模型推导出可解释的支付意愿（WTP），并与人类基准对比？","design":"以酒店房间选择为任务，构建多属性选择困境，让多个LLM在不同提示条件下（基线、提供用户历史选择、基于人设的提示）做出选择，然后用多项Logit模型估计隐含的WTP，并分析提示改写、顺序调换、货币变化、温度等稳健性。","baseline":"对照Masiero et al. (2015)中人类对酒店房间属性的支付意愿估计值。","findings":"较大LLM能推导出有意义的WTP，但存在属性层面的系统性偏差，且整体高估人类WTP，尤其在引入昂贵选项或商务人设时；提供廉价偏好历史或学生人设可使估值更接近人类基准。","reliability":"论文承认WTP推导在部分模型上会失效，且结果受提示设计、用户表征方式影响较大，未涵盖更广泛的人群异质性，泛化到其他领域需进一步验证。","relevance":"直接对比LLM与人类真实WTP数据，系统评估了提示策略对仿真偏差的影响，并指出失效条件，高度契合研究者对经济学场景下LLM仿真可靠性及批判性分析的兴趣，值得精读原文。","inspiration":"借鉴其通过离散选择模型从LLM主观选择中推导支付意愿（WTP）并与人类基准对比的方法，可系统评估提示策略（如提供用户历史、人设提示）对仿真偏差的影响｜可迁移到消费者对金融产品属性的偏好评估，如贷款条款（利率、期限、抵押要求）或投资产品特征（风险、流动性、费用）的WTP估计｜以LLM为被试，呈现不同贷款产品选择集，处理为提供不同风险偏好或财务约束的人设提示，结果变量为通过多项Logit模型估计的各属性WTP，对照真实消费者信贷选择数据（如Survey of Consumer Finances）验证偏差"}},{"id":"2602.07414","version":1,"title":"Can LLMs Truly Embody Human Personality? Analyzing AI and Human Behavior Alignment in Dispute Resolution","zh_title":"LLM能真正体现人类人格吗？分析争议解决中AI与人类行为的一致性","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior in social settings such as legal mediation, negotiation, and dispute resolution. However, it remains unclear whether these simulations reproduce the personality-behavior patterns observed in humans. Human personality, for instance, shapes how individuals navigate social interactions, including strategic choices and behaviors in emotionally charged interactions. This raises the question: Can LLMs, when prompted with personality traits, reproduce personality-driven differences in human conflict behavior? To explore this, we introduce an evaluation framework that enables direct comparison of human-human and LLM-LLM behaviors in dispute resolution dialogues with respect to Big Five Inventory (BFI) personality traits. This framework provides a set of interpretable metrics related to strategic behavior and conflict outcomes. We additionally contribute a novel dataset creation methodology for LLM dispute resolution dialogues with matched scenarios and personality traits with respect to human conversations. Finally, we demonstrate the use of our evaluation framework with three contemporary closed-source LLMs and show significant divergences in how personality manifests in conflict across different LLMs compared to human data, challenging the assumption that personality-prompted agents can serve as reliable behavioral proxies in socially impactful applications. Our work highlights the need for psychological grounding and validation in AI simulations before real-world use.","authors":["Deuksin Kwon","Kaleen Shrestha","Bin Han","Spencer Lin","James Hale","Jonathan Gratch","Maja Matarić","Gale M. Lucas"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-07","first_seen":"2026-02-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.07414","pdf_url":"https://arxiv.org/pdf/2602.07414","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM人格仿真","人类行为对齐","争议解决"],"reason":"直接比较LLM与人类在冲突对话中的人格-行为对齐，有真实人类数据对照，并指出仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":22,"question":"当用大五人格特质提示LLM时，它们能否在冲突解决对话中复现人类因人格差异而导致的行为差异？","design":"基于KODIS人类冲突对话数据集，为LLM匹配相同场景和人格特质（大五人格），生成LLM-LLM冲突对话；比较人类与LLM在最终结果（得分、是否接受、是否离开）和策略行为（基于利益-权利-权力框架）上的差异。","baseline":"KODIS数据集中248段具有完整人格信息的人类-人类冲突解决对话。","findings":"人类中神经质是策略结果的最强预测因子，而LLM中外向性和宜人性的效应更强，且策略行为负载于更广泛的人格因素；Claude和Gemini比GPT-4o mini更接近人类策略指标，但整体仍存在显著偏差。","reliability":"论文指出人格提示的LLM在情感冲突场景中的行为保真度未经严格验证，不同LLM间人格表现差异显著，不能可靠地作为人类行为代理，强调在真实应用前需进行心理验证。","relevance":"该研究直接比较LLM与人类在冲突对话中的人格-行为对齐，有真实人类数据对照，并指出仿真失效条件，高度契合您对LLM人类仿真可靠性及批判性研究的关注，值得精读原文。","inspiration":"该方法通过给LLM施加人格特质提示来模拟人类行为差异，并设置真实人类对话数据作为对照基准，可借鉴其处理-对照设计及基于框架的策略行为编码。｜可迁移至消费者跨期选择实验，探究不同人格特质（如尽责性、神经质）对时间贴现行为的影响。｜以LLM为被试，施加大五人格提示，测量其在跨期选择任务中的贴现率，并与真实人类实验数据（如Andersen et al.的贴现率估计）进行对照。"}},{"id":"2602.18462","version":1,"title":"Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents","zh_title":"评估基于人格条件的LLM作为合成调查受访者的可靠性","abstract":"Using persona-conditioned LLMs as synthetic survey respondents has become a common practice in computational social science and agent-based simulations. Yet, it remains unclear whether multi-attribute persona prompting improves LLM reliability or instead introduces distortions. Here we contribute to this assessment by leveraging a large dataset of U.S. microdata from the World Values Survey. Concretely, we evaluate two open-weight chat models and a random-guesser baseline across more than 70K respondent-item instances. We find that persona prompting does not yield a clear aggregate improvement in survey alignment and, in many cases, significantly degrades performance. Persona effects are highly heterogeneous as most items exhibit minimal change, while a small subset of questions and underrepresented subgroups experience disproportionate distortions. Our findings highlight a key adverse impact of current persona-based simulation practices: demographic conditioning can redistribute error in ways that undermine subgroup fidelity and risk misleading downstream analyses.","authors":["Erika Elizabeth Taday Morocho","Lorenzo Cima","Tiziano Fagni","Marco Avvenuti","Stefano Cresci"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-02-06","first_seen":"2026-02-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.18462","pdf_url":"https://arxiv.org/pdf/2602.18462","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查方法","可靠性评估"],"reason":"直接评估LLM作为合成调查受访者的可靠性，使用真实人类数据对照，并指出仿真失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":5,"question":"多属性人格提示（persona prompting）能否提高大语言模型作为合成调查受访者的可靠性，还是会引入扭曲？","design":"使用两个开源聊天模型（Llama-2-13B 和 Qwen3-4B）模拟美国受访者，基于世界价值观调查（WVS-7）的个体记录构建多属性人格提示，比较有人格提示、无提示（vanilla）和随机猜测基线在70K+受访者-题目实例上的回答一致性。","baseline":"世界价值观调查第7波（WVS-7）的美国受访者微观数据，作为真实人类回答的基准。","findings":"人格提示在总体上并未带来一致的对齐改善，在许多情况下反而显著降低性能；人格效应高度异质，大多数题目变化极小，但少数题目和代表性不足的子群体出现不成比例的扭曲。","reliability":"论文指出人格提示可能重新分配误差，损害子群体保真度，并误导下游分析；效果因题目和属性而异，在少数群体中可能集中出现错误，且多属性约束可能相互干扰。","relevance":"该研究直接评估LLM作为合成调查受访者的可靠性，使用真实人类数据对照，并批判性地揭示了人格提示在子群体层面的失效风险，高度契合研究者对仿真可靠性、偏差及失效条件的关注，值得精读原文。","inspiration":"借鉴多属性人格提示与无提示基线的对照设计，以及按题目和子群体分解异质性效应的分析方法，可系统评估LLM仿真中的偏差来源｜可迁移到消费者金融决策调查场景，如风险偏好、储蓄选择、信贷需求等问卷的行为一致性研究｜以LLM作为合成受访者，施加多属性人格提示（收入、教育、财务素养等），测量其与真实消费者金融调查（如SCF）中个体回答的匹配度，并按收入分位数和金融素养水平检验子群体偏差"}},{"id":"2602.18464","version":2,"title":"How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?","zh_title":"LLM代理模拟终端用户安全与隐私态度及行为的效果如何？","abstract":"A growing body of research assumes that large language model (LLM) agents can serve as proxies for how people form attitudes toward and behave in response to security and privacy (S&P) threats. If correct, these simulations could offer a scalable way to forecast S&P risks in products prior to deployment. We interrogate this assumption using SP-ABCBench, a new benchmark of 30 tests derived from validated S&P human-subject studies, which measures alignment between simulations and human-subjects studies on a 0-100 ascending scale, where higher scores indicate better alignment across three dimensions: Attitude, Behavior, and Coherence. Evaluating twelve LLMs, four persona construction strategies, and two prompting methods, we found that there remains substantial room for improvement: all models score between 50 and 64 on average. Newer, bigger, and smarter models do not reliably do better and sometimes do worse. Some simulation configurations, however, do yield high alignment: e.g., with scores above 95 for some behavior tests when agents are prompted to apply bounded rationality and weigh privacy costs against perceived benefits. We release SP-ABCBench to enable reproducible evaluation as methods improve.","authors":["Yuxuan Li","Leyang Li","Hao-Ping Lee","Sauvik Das"],"categories":["cs.CY","cs.AI","cs.CL","cs.CR"],"primary_category":"cs.CY","announce_type":"new","date":"2026-02-06","first_seen":"2026-02-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.18464","pdf_url":"https://arxiv.org/pdf/2602.18464","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类行为对照","安全隐私"],"reason":"直接评估LLM代理模拟人类安全隐私态度行为，有真实人类数据基准，并指出仿真失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":41,"question":"当前LLM代理在多大程度上能复现人群在安全与隐私（S&P）方面的态度、行为及其一致性？","design":"使用12个LLM，结合4种角色构建策略和2种提示方法，基于15项真实人类被试研究构建的30个测试基准SP-ABCBench，测量仿真结果与人类数据在态度、行为和一致性三个维度的对齐分数（0-100）。","baseline":"对照来自15项经过验证的S&P人类被试研究，涵盖态度量表、行为实验和构念间关系等30个可量化的人群层面效应。","findings":"所有模型平均对齐分数仅50-64，更大、更新、更强的模型未必更好，有时更差；但特定配置（如结合有限理性与隐私计算提示）在部分行为测试上可达95分以上。","reliability":"论文指出当前LLM仿真在S&P领域整体对齐度中等，模型规模与能力不保证提升，且角色构建与提示策略效果因测试维度而异，提示仿真在S&P决策中可能失效的条件。","relevance":"该研究直接评估LLM替代人类被试进行S&P态度行为仿真的可靠性，有真实人类基准，并揭示了仿真失效的具体条件，高度契合您对LLM人类仿真实验批判性评估的关注，值得精读。","inspiration":"该方法通过构建多维度测试基准（态度、行为、一致性）并计算对齐分数来量化LLM仿真与人类数据的差距，可借鉴其系统性评估框架和对照设计。｜可迁移到消费者金融决策研究，如评估LLM能否复现真实人群在信贷选择、风险偏好或退休储蓄行为中的偏差与异质性。｜以LLM代理为被试，施加不同金融素养或信息框架处理，测量其信贷违约概率或投资组合选择，并与美国消费者金融调查（SCF）或实验室实验的真实行为数据做对齐比较。"}},{"id":"2602.04674","version":2,"title":"Overstating Attitudes, Ignoring Networks: LLM Biases in Simulating Misinformation Susceptibility","zh_title":"夸大态度，忽视网络：LLM在模拟错误信息易感性中的偏差","abstract":"Large language models (LLMs) are increasingly used as proxies for human judgment in computational social science, yet their ability to reproduce patterns of susceptibility to misinformation remains unclear. We test whether LLM-simulated survey respondents, prompted with participant profiles drawn from social survey data measuring network, demographic, attitudinal and behavioral features, can reproduce human patterns of misinformation belief and sharing. Using three online surveys as baselines, we evaluate whether LLM outputs match observed response distributions and recover feature-outcome associations present in the original survey data. LLM-generated responses capture broad distributional tendencies and show modest correlation with human responses, but consistently overstate the association between belief and sharing. Linear models fit to simulated responses exhibit substantially higher explained variance and place disproportionate weight on attitudinal and behavioral features, while largely ignoring personal network characteristics, relative to models fit to human responses. Analyses of model-generated reasoning and LLM training data suggest that these distortions reflect systematic biases in how misinformation-related concepts are represented. Our findings suggest that LLM-based survey simulations are better suited for diagnosing systematic divergences from human judgment than for substituting it.","authors":["Eun Cheol Choi","Lindsay E. Young","Emilio Ferrara"],"categories":["cs.SI","cs.AI","cs.CL"],"primary_category":"cs.SI","announce_type":"new","date":"2026-02-04","first_seen":"2026-02-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.04674","pdf_url":"https://arxiv.org/pdf/2602.04674","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","错误信息","人类数据对照"],"reason":"用LLM仿真人类对错误信息的易感性，并与真实调查数据对照，评估偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":33,"question":"LLM模拟的调查受访者在多大程度上能复现人类对错误信息的相信与分享模式及其与社会预测因子（包括个人网络特征）之间的关联？","design":"使用LLM（如GPT-4等）扮演合成调查受访者，输入基于三个真实调查数据构建的受访者结构化档案（包含个人网络、人口统计、态度/行为特征），让LLM生成对错误信息条目的相信和分享意愿回答，测量回答分布及特征-结果关联。","baseline":"三个在线调查数据集：公共卫生（美国，2023）、气候变化（美国，2025）、疫情政治（韩国，2020），均包含真实人类对错误信息的相信与分享数据及个人网络、人口统计、态度行为变量。","findings":"LLM模拟能捕捉大致分布趋势并与人类回答有适度相关，但系统性地夸大了相信与分享之间的关联；线性模型在模拟数据上解释方差显著膨胀，且过度依赖态度和行为特征，几乎忽略个人网络特征。","reliability":"论文指出LLM模拟更适合诊断与人类判断的系统性偏差，而非替代人类判断；偏差源于LLM训练数据中错误信息相关概念的表征偏差，且模拟未能复现个人网络特征的作用。","relevance":"该研究直接以真实人类调查为基准，评估LLM仿真在错误信息易感性上的可靠性，并揭示了仿真在忽略网络特征、夸大态度关联等方面的失效条件，高度契合研究者对批判性仿真研究的兴趣，值得精读原文。","inspiration":"该方法借鉴了用结构化档案（含人口统计、态度、网络特征）驱动LLM生成调查回答，并与真实人类数据对照以评估仿真偏差的设计｜可迁移到信贷审批中的歧视研究，检验LLM模拟的贷款官员是否复现人类决策中的种族或性别偏见｜用LLM扮演贷款审批员，输入含申请人种族、收入、信用分等档案，输出审批决定，以真实房贷数据（如HMDA）为基准，比较拒绝率差异及特征重要性"}},{"id":"2602.01684","version":1,"title":"The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament","zh_title":"大语言模型的战略远见：来自全前瞻性创业锦标赛的证据","abstract":"Can artificial intelligence outperform humans at strategic foresight -- the capacity to form accurate judgments about uncertain, high-stakes outcomes before they unfold? We address this question through a fully prospective prediction tournament using live Kickstarter crowdfunding projects. Thirty U.S.-based technology ventures, launched after the training cutoffs of all models studied, were evaluated while fundraising remained in progress and outcomes were unknown. A diverse suite of frontier and open-weight large language models (LLMs) completed 870 pairwise comparisons, producing complete rankings of predicted fundraising success. We benchmarked these forecasts against 346 experienced managers recruited via Prolific and three MBA-trained investors working under monitored conditions. The results are striking: human evaluators achieved rank correlations with actual outcomes between 0.04 and 0.45, while several frontier LLMs exceeded 0.60, with the best (Gemini 2.5 Pro) reaching 0.74 -- correctly ordering nearly four of every five venture pairs. These differences persist across multiple performance metrics and robustness checks. Neither wisdom-of-the-crowd ensembles nor human-AI hybrid teams outperformed the best standalone model.","authors":["Felipe A. Csaszar","Aticus Peterson","Daniel Wilde"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-02","first_seen":"2026-02-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.01684","pdf_url":"https://arxiv.org/pdf/2602.01684","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","人类行为对照","创业预测"],"reason":"用LLM预测人类对创业项目的判断，并与真实人类数据对照，属于经济学场景下的人类…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"大语言模型在战略远见（预测不确定、高风险商业结果）上能否超越人类？","design":"使用多款前沿和开源大语言模型对30个正在众筹的Kickstarter科技项目进行870次成对比较，生成预测筹款成功排名；同时招募346名有经验的管理者和3名MBA投资者作为人类被试进行相同任务，以最终实际筹款结果作为准确度标准。","baseline":"346名通过Prolific招募的有经验管理者和3名受监控的MBA投资者，其预测排名与实际结果的秩相关系数在0.04至0.45之间。","findings":"人类评估者的预测排名与实际结果的秩相关系数最高仅0.45，而多个前沿大语言模型超过0.60，最佳模型Gemini 2.5 Pro达到0.74。群体智慧集成和人机混合团队均未超越最佳独立模型。","reliability":"论文未讨论","relevance":"该研究直接用LLM替代人类被试进行前瞻性预测，并与真实人类数据对照，属于经济学场景下的人类仿真实验，且提供了可靠性证据，值得精读原文。","inspiration":"该方法借鉴了用真实众筹结果作为客观基准，直接比较LLM与人类被试的预测准确度，并通过成对比较排名任务量化战略远见｜可迁移到创业投资决策、众筹市场预测或资产定价中的预期形成研究，用于评估AI辅助决策的可靠性｜招募专业投资者作为人类被试，让LLM和人类分别对真实众筹项目进行成对比较排名，以最终实际筹资金额作为结果变量，计算预测排名与实际结果的秩相关系数，对比人机表现"}},{"id":"2602.07023","version":2,"title":"Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation","zh_title":"LLM智能体行为一致性验证：基于股市模拟的交易风格切换分析","abstract":"Recent works have increasingly applied Large Language Models (LLMs) as agents in financial stock market simulations to test if micro-level behaviors aggregate into macro-level phenomena. However, a crucial question arises: Do LLM agents' behaviors align with real market participants? This alignment is key to the validity of simulation results. To explore this, we select a financial stock market scenario to test behavioral consistency. Investors are typically classified as fundamental or technical traders, but most simulations fix strategies at initialization, failing to reflect real-world trading dynamics. In this work, we assess whether agents' strategy switching aligns with financial theory, providing a framework for this evaluation. We operationalize four behavioral-finance drivers-loss aversion, herding, wealth differentiation, and price misalignment-as personality traits set via prompting and stored long-term. In year-long simulations, agents process daily price-volume data, trade under a designated style, and reassess their strategy every 10 trading days. We introduce four alignment metrics and use Mann-Whitney U tests to compare agents' style-switching behavior with financial theory. Our results show that recent LLMs' switching behavior is only partially consistent with behavioral-finance theories, highlighting the need for further refinement in aligning agent behavior with financial theory.","authors":["Zeping Li","Guancheng Wan","Keyang Chen","Yu Chen","Yiwen Zhao","Philip Torr","Guangnan Ye","Zhenfei Yin","Hongfeng Chai"],"categories":["q-fin.TR","cs.AI"],"primary_category":"q-fin.TR","announce_type":"new","date":"2026-02-02","first_seen":"2026-02-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.07023","pdf_url":"https://arxiv.org/pdf/2602.07023","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","行为金融","智能体一致性"],"reason":"用LLM agent模拟股票交易行为并与金融理论对照，涉及行为经济学场景，指出…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":55,"question":"在股票市场仿真中，LLM智能体的交易风格切换行为是否与行为金融学理论一致？","design":"使用多种近期LLM（通过提示词设定四种行为金融倾向作为人格特质并存入长期记忆）扮演投资者，在基于2024年标普500成分股数据的模拟环境中进行为期一年的交易，每10个交易日评估并决定是否切换交易风格（基本面/技术面），通过四个对齐指标和Mann-Whitney U检验比较智能体行为与理论预期。","baseline":"无对照","findings":"LLM智能体的风格切换行为仅部分符合行为金融学理论，无法在所有方面完全对齐。","reliability":"论文未讨论","relevance":"该研究直接评估LLM智能体在金融行为仿真中的行为一致性，属于批判性验证工作，与研究者关注的LLM仿真可靠性及失效条件高度相关，值得阅读原文以了解具体偏差和评估框架。","inspiration":"该研究通过提示词将行为金融倾向植入LLM智能体并存入长期记忆，在动态仿真中周期性评估风格切换，用多个对齐指标和统计检验衡量与理论的偏差，这种处理-测量-验证框架值得借鉴。｜可迁移到资产定价实验，检验LLM智能体在信息冲击下的过度反应与反转行为是否符合前景理论与处置效应。｜以LLM智能体为被试，施加不同强度的利好/利空消息作为处理，观测其持仓调整与买卖时机，结果变量为超额收益与换手率，用历史高频交易数据中散户的实际行为分布作为对照基准。"}},{"id":"2602.01022","version":3,"title":"Calibrating Behavioral Parameters with Large Language Models","zh_title":"用大语言模型校准行为参数","abstract":"Behavioral parameters such as loss aversion, herding, and extrapolation are central to asset pricing models but remain difficult to measure reliably. We develop a framework that treats large language models (LLMs) as calibrated measurement instruments for behavioral parameters. Using four models and 24{,}000 agent--scenario pairs, we document systematic rationality bias in baseline LLM behavior, including attenuated loss aversion, weak herding, and near-zero disposition effects relative to human benchmarks. Profile-based calibration induces large, stable, and theoretically coherent shifts in several parameters, with calibrated loss aversion, herding, extrapolation, and anchoring reaching or exceeding benchmark magnitudes. To assess external validity, we embed calibrated parameters in an agent-based asset pricing model, where calibrated extrapolation generates short-horizon momentum and long-horizon reversal patterns consistent with empirical evidence. Our results establish measurement ranges, calibration functions, and explicit boundaries for eight canonical behavioral biases.","authors":["Brandon Yee","Pairie Koh"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-01","first_seen":"2026-02-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.01022","pdf_url":"https://arxiv.org/pdf/2602.01022","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为经济学","人类基准对照"],"reason":"用LLM测量行为参数并与人类基准对照，嵌入资产定价模型验证外部效度，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":8,"question":"能否将大语言模型作为可校准的测量工具，系统性地诱导、校准并验证行为金融参数？","design":"使用 GPT-4o、GPT-4o-mini、Claude-3.5-Haiku、Gemini-2.5-Pro 四种模型，通过提示词嵌入行为特征（如损失厌恶、羊群效应等）作为实验处理，在 24,000 个合成金融场景中测量八种行为偏差的参数值，并评估校准后的参数在基于代理的资产定价模型中的外部有效性。","baseline":"人类基准来自已有文献中的实验和实证估计，如损失厌恶系数约 2.25，羊群效应率 65-75%，以及 Jegadeesh 和 Titman 的动量与反转经验事实。","findings":"基线 LLM 行为存在系统性理性偏差，表现为损失厌恶减弱、羊群效应弱、处置效应接近零；通过基于特征的校准可诱导出与人类基准相当甚至更强的参数值，且校准后的外推参数能在资产定价模型中生成符合经验事实的短期动量和长期反转模式。","reliability":"论文承认校准存在明确边界，并非所有参数都能成功校准，且结果依赖于提示词设计和模型选择，外部有效性仅通过简单资产定价模型初步验证。","relevance":"该研究直接命中研究者对 LLM 仿真人类行为、与真实人类基准对照、经济学实验及批判性评估的核心兴趣，提供了系统的校准框架和失效边界，值得精读原文。","inspiration":"该方法通过提示词嵌入行为特征（如损失厌恶、羊群效应）作为实验处理，系统性地校准LLM的行为参数，并与人类基准对照，值得借鉴其处理施加与参数校准的流程设计｜可迁移到资产定价实验中的投资者行为偏差研究，如模拟动量效应、反转效应及处置效应等市场异象｜以LLM作为被试，通过提示词嵌入不同程度的损失厌恶或羊群效应处理，测量其交易决策与价格预期，并与Jegadeesh和Titman的动量/反转经验事实及处置效应实证数据对照"}},{"id":"2602.00685","version":1,"title":"HumanStudy-Bench: Towards AI Agent Design for Participant Simulation","zh_title":"HumanStudy-Bench：面向参与者仿真的AI智能体设计基准","abstract":"Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable and highly sensitive to design choices. Prior evaluations frequently conflate base-model capabilities with experimental instantiation, obscuring whether outcomes reflect the model itself or the agent setup. We instead frame participant simulation as an agent-design problem over full experimental protocols, where an agent is defined by a base model and a specification (e.g., participant attributes) that encodes behavioral assumptions. We introduce HUMANSTUDY-BENCH, a benchmark and execution engine that orchestrates LLM-based agents to reconstruct published human-subject experiments via a Filter--Extract--Execute--Evaluate pipeline, replaying trial sequences and running the original analysis pipeline in a shared runtime that preserves the original statistical procedures end to end. To evaluate fidelity at the level of scientific inference, we propose new metrics to quantify how much human and agent behaviors agree. We instantiate 12 foundational studies as an initial suite in this dynamic benchmark, spanning individual cognition, strategic interaction, and social psychology, and covering more than 6,000 trials with human samples ranging from tens to over 2,100 participants.","authors":["Xuan Liu","Haoyang Shang","Zizhang Liu","Xinyan Liu","Yunze Xiao","Yiwen Tu","Haojian Jin"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-31","first_seen":"2026-01-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.00685","pdf_url":"https://arxiv.org/pdf/2602.00685","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3"],"tags":["LLM人类仿真","实验复现","基准测试"],"reason":"直接构建LLM代理复现人类实验，含真实人类数据对照，评估仿真保真度，覆盖经济学…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":11,"question":"如何将LLM参与者的仿真视为一个代理设计问题，并系统评估不同代理设计在复现真实人类实验中的保真度？","design":"使用10个当代LLM（如GPT、Claude、Gemini）作为基座模型，结合四种代理规格（空白、角色扮演、人口统计条件、丰富背景故事）构建AI代理，通过Filter–Extract–Execute–Evaluate管道重放12项已发表的人类实验的完整试验序列和原始分析流程，测量代理的行为响应。","baseline":"对照12项已发表人类实验的真实人类数据，涵盖个体认知、策略互动和社会心理学，超过6000次试验，人类样本量从几十到2100多人。","findings":"当前LLM代理与人类的推断一致性有限且不稳定，行为呈极化双峰而非人类单峰模式；代理设计对结果有较大且非单调的影响，性能高度依赖领域，更大模型或简单多模型集成未能可靠提升对齐度。","reliability":"论文指出LLM行为不稳定、对设计选择高度敏感，代理规格会编码行为假设并可能定性改变结果；现有评估常混淆基座模型能力与实验实例化，且代理在人口异质性敏感度和提示词变化下表现脆弱。","relevance":"该研究直接针对用LLM替代人类被试的仿真实验，提供真实人类数据对照，系统评估代理设计对复现经济学实验和社会心理学效应的影响，并批判性指出仿真失效的条件，与您的关注高度契合，值得精读原文。","inspiration":"借鉴其将代理设计作为实验变量的思路，通过对比空白、角色扮演、人口统计条件等不同规格来分离模型能力与实验实例化的混淆效应｜可迁移到政策公告的预期形成实验，研究不同信息框架下投资者对央行沟通的反应｜以LLM代理为被试，处理为不同代理规格（如空白vs.人口统计条件），结果变量为通胀预期调整幅度，对照真实调查数据（如密歇根消费者调查）"}},{"id":"2601.22812","version":2,"title":"Stable Personas: Dual-Assessment of Temporal Stability in LLM-Based Human Simulation","zh_title":"稳定人格：基于LLM的人类仿真中时间稳定性的双重评估","abstract":"Large Language Models (LLMs) acting as artificial agents offer the potential for scalable behavioral research, yet their validity depends on whether LLMs can maintain stable personas across extended conversations. We address this point using a dual-assessment framework measuring both self-reported characteristics and observer-rated persona expression. Across two experiments testing four persona conditions (default, high, moderate, and low ADHD presentations), seven LLMs, and three semantically equivalent persona prompts, we examine between-conversation stability (3,473 conversations) and within-conversation stability (1,370 conversations and 18 turns). Self-reports remain highly stable both between and within conversations. However, observer ratings reveal a tendency for persona expressions to decline during extended conversations. These findings suggest that persona-instructed LLMs produce stable, persona-aligned self-reports, an important prerequisite for behavioral research, while identifying this regression tendency as a boundary condition for multi-agent social simulation.","authors":["Jana Gonnermann-Müller","Jennifer Haase","Nicolas Leins","Thomas Kosch","Sebastian Pokutta"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-01-30","first_seen":"2026-01-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.22812","pdf_url":"https://arxiv.org/pdf/2601.22812","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B4"],"tags":["LLM仿真","人格稳定性","效度评估"],"reason":"研究LLM人格稳定性以评估其作为人类被试的可靠性，直接涉及仿真效度与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":47,"question":"LLM在独立对话间和长对话过程中，能在多大程度上稳定维持被赋予的人格？","design":"用7个LLM扮演默认、高、中、低四种ADHD人格，通过三种语义等价的提示词施加处理；实验一测量跨对话稳定性（每条件50次独立运行），实验二测量对话内稳定性（18轮对话，在第6、12、18轮评估）；结果变量为自评量表得分和观察者评分。","baseline":"无对照","findings":"自评报告在跨对话和对话内均高度稳定；但观察者评分显示，高和中强度人格表达在长对话中会逐渐减弱，向默认水平回归。","reliability":"论文指出，人格表达在长对话中衰退是LLM用于多智能体社会仿真的一个边界条件；评估依赖LLM作为评分者，可能引入偏差；仅以ADHD人格为测试案例，泛化性待验证。","relevance":"直接评估LLM作为人类被试替代品的人格稳定性，揭示了自评稳定但行为表达衰退的失效模式，对关注仿真效度与边界条件的研究者很有参考价值，值得读原文。","inspiration":"借鉴双维度人格稳定性评估设计，通过独立跨对话与长对话内重复测量，区分自评报告与行为观察的稳定性差异，揭示仿真衰退的边界条件。｜可迁移至政策公告预期形成实验，用LLM模拟投资者对央行沟通的反应，检验长期对话中信息解读的一致性。｜以LLM为被试，施加不同政策措辞处理，在长对话中多次测量通胀预期与投资意愿，对比真实投资者调查面板数据，评估仿真在持续信息流下的衰退模式。"}},{"id":"2601.17527","version":1,"title":"Bridging Expectation Signals: LLM-Based Experiments and a Behavioral Kalman Filter Framework","zh_title":"桥接预期信号：基于LLM的实验与行为卡尔曼滤波框架","abstract":"As LLMs increasingly function as economic agents, the specific mechanisms LLMs use to update their belief with heterogeneous signals remain opaque. We design experiments and develop a Behavioral Kalman Filter framework to quantify how LLM-based agents update expectations, acting as households or firm CEOs, update expectations when presented with individual and aggregate signals. The results from experiments and model estimation reveal four consistent patterns: (1) agents' weighting of priors and signals deviates from unity; (2) both household and firm CEO agents place substantially larger weights on individual signals compared to aggregate signals; (3) we identify a significant and negative interaction between concurrent signals, implying that the presence of multiple information sources diminishes the marginal weight assigned to each individual signal; and (4) expectation formation patterns differ significantly between household and firm CEO agents. Finally, we demonstrate that LoRA fine-tuning mitigates, but does not fully eliminate, behavioral biases in LLM expectation formation.","authors":["Yu Wang","Xiangchen Liu"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-01-24","first_seen":"2026-01-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.17527","pdf_url":"https://arxiv.org/pdf/2601.17527","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM经济代理","预期形成","行为偏差"],"reason":"用LLM模拟家庭和CEO预期更新，涉及经济实验，有行为偏差分析，但未明确提及真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":56,"question":"LLM代理在接收异质性信号时如何更新预期，其权重分配和行为偏差的机制是什么？","design":"使用GPT-4o、Gemini 1.5、DeepSeek-V3等LLM扮演家庭和CEO，在720次试验中呈现微观（个人收入/公司业绩）和宏观（GDP增长）两种信号，测量其对未来收入或利润增长的预期更新。","baseline":"无对照","findings":"LLM代理对微观信号的权重大于宏观信号，且信号间存在负交互效应，即多信号会削弱各自边际权重；CEO比家庭更重视宏观信号。LoRA微调可减轻但无法消除非理性偏差。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济主体预期更新，涉及行为偏差分析，但缺乏真实人类数据基准，适合关注仿真机制和偏差的研究者阅读原文以评估方法细节。","inspiration":"该方法通过向LLM代理呈现微观与宏观异质性信号并测量预期更新权重，可借鉴其多信号交互实验设计来分离信息处理偏差｜可迁移至政策公告的预期形成研究，如央行沟通中微观通胀感知与宏观通胀目标对家庭预期的交互影响｜以LLM模拟家庭，处理为同时呈现个人消费价格变化（微观）与官方CPI（宏观），结果变量为通胀预期更新幅度，对照真实家庭调查数据（如密歇根消费者调查）"}},{"id":"2601.16355","version":2,"title":"Identity, Cooperation and Framing Effects within Groups of Real and Simulated Humans","zh_title":"真实与模拟人类群体中的身份、合作与框架效应","abstract":"Humans act via a nuanced process that depends both on rational deliberation and also on identity and contextual factors. In this work, we study how large language models (LLMs) can simulate human action in the context of social dilemma games. While prior work has focused on \"steering\" (weak binding) of chat models to simulate personas, we analyze here how deep binding of base models with extended backstories leads to more faithful replication of identity-based behaviors. Our study has these findings: simulation fidelity vs human studies is improved by conditioning base LMs with rich context of narrative identities and checking consistency using instruction-tuned models. We show that LLMs can also model contextual factors such as time (year that a study was performed), question framing, and participant pool effects. LLMs, therefore, allow us to explore the details that affect human studies but which are often omitted from experiment descriptions, and which hamper accurate replication.","authors":["Suhong Moon","Minwoo Kang","Joseph Suh","Mustafa Safdari","John Canny"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-22","first_seen":"2026-01-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.16355","pdf_url":"https://arxiv.org/pdf/2601.16355","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","社会困境博弈","行为实验复现"],"reason":"用LLM模拟社会困境中的人类行为，并与真实人类研究对照，涉及合作与框架效应。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":24,"question":"大语言模型能否通过深度绑定身份背景、时间锚定和一致性过滤，在独裁者博弈和信任博弈中复现真实人类的党派内群体偏袒行为？","design":"使用 Mistral-Small、Mixtral 8x22B 和 Qwen-2.5 72B 基础模型，通过深度绑定（DeepBind）方法为虚拟被试赋予详细的叙事身份背景，并施加一致性过滤（反复重申身份）和时间锚定（设定实验年份），在独裁者博弈和信任博弈中测量虚拟被试对同党和异党对象的资源分配与信任行为。","baseline":"对照的真实人类数据来自 Whitt et al. (2021) 和 Iyengar & Westwood (2015) 的独裁者博弈研究，以及 Carlin & Love (2018) 和 Whitt et al. (2021) 的信任博弈研究，均包含民主党与共和党被试的党派偏袒效应量。","findings":"DeepBind 方法在所有模型和博弈中均最一致地复现了人类党派偏袒差距，且同时使用时间锚定和一致性过滤能进一步提升仿真与人类基准的对齐程度。LLM 还能捕捉到实验年份、问题措辞和参与者池等常被忽略的细微情境效应。","reliability":"论文未讨论","relevance":"该研究直接用 LLM 复现社会困境博弈中的人类行为，并与多项真实人类实验进行定量对照，同时探讨了身份、时间框架等情境因素的仿真效果，高度契合对 LLM 人类仿真可靠性及偏差的关注，值得精读。","inspiration":"借鉴之处在于通过深度绑定叙事身份、时间锚定和一致性过滤来增强LLM的情境代入感，从而更精细地操控虚拟被试的社会身份与决策框架｜可迁移至经济金融中的群体间歧视行为研究，例如信贷审批中的党派或种族偏见、投资决策中的内群体偏袒｜设计上以LLM作为虚拟信贷员，通过DeepBind赋予其不同党派身份，并设定审批年份，测量其对同党与异党申请人的贷款批准率差异，以真实信贷歧视研究数据作为对照基准"}},{"id":"2601.15793","version":1,"title":"HumanLLM: Towards Personalized Understanding and Simulation of Human Nature","zh_title":"HumanLLM：迈向个性化理解与人性仿真","abstract":"Motivated by the remarkable progress of large language models (LLMs) in objective tasks like mathematics and coding, there is growing interest in their potential to simulate human behavior--a capability with profound implications for transforming social science research and customer-centric business insights. However, LLMs often lack a nuanced understanding of human cognition and behavior, limiting their effectiveness in social simulation and personalized applications. We posit that this limitation stems from a fundamental misalignment: standard LLM pretraining on vast, uncontextualized web data does not capture the continuous, situated context of an individual's decisions, thoughts, and behaviors over time. To bridge this gap, we introduce HumanLLM, a foundation model designed for personalized understanding and simulation of individuals. We first construct the Cognitive Genome Dataset, a large-scale corpus curated from real-world user data on platforms like Reddit, Twitter, Blogger, and Amazon. Through a rigorous, multi-stage pipeline involving data filtering, synthesis, and quality control, we automatically extract over 5.5 million user logs to distill rich profiles, behaviors, and thinking patterns. We then formulate diverse learning tasks and perform supervised fine-tuning to empower the model to predict a wide range of individualized human behaviors, thoughts, and experiences. Comprehensive evaluations demonstrate that HumanLLM achieves superior performance in predicting user actions and inner thoughts, more accurately mimics user writing styles and preferences, and generates more authentic user profiles compared to base models. Furthermore, HumanLLM shows significant gains on out-of-domain social intelligence benchmarks, indicating enhanced generalization.","authors":["Yuxuan Lei","Tianfu Wang","Jianxun Lian","Zhengyu Hu","Defu Lian","Xing Xie"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-22","first_seen":"2026-01-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15793","pdf_url":"https://arxiv.org/pdf/2601.15793","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1"],"tags":["人类行为仿真","个性化建模","社会模拟"],"reason":"用LLM仿真个体行为与思维，有真实用户数据对照，评估预测准确性，方法可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":38,"question":"能否构建一个通用基础模型，利用大规模真实用户数据实现个性化的人类认知与行为理解和仿真？","design":"基于Reddit、Twitter、Blogger、Amazon等平台的真实用户数据构建认知基因组数据集，通过数据过滤、合成和质量控制的多阶段流水线提取超过550万条用户日志，设计个人资料生成、社会问答和写作模仿等任务，对基础LLM进行监督微调，并使用模型合并策略保留通用能力。","baseline":"对照的真实人类数据为来自Reddit、Twitter、Blogger、Amazon等平台的真实用户日志和行为记录。","findings":"HumanLLM在预测用户行为、内心想法和风格模仿上显著优于基础模型；在MotiveBench和TomBench等域外社会智能基准上表现出增强的泛化能力。","reliability":"论文未讨论","relevance":"该研究直接利用真实人类数据训练LLM以仿真个体行为与思维，并评估预测准确性，方法可迁移至经济学实验和政策评估场景，值得精读原文以了解其仿真可靠性与局限。","inspiration":"该方法利用多平台真实用户日志构建个性化仿真数据集，并通过监督微调使LLM模仿个体行为与思维，可借鉴其多阶段数据过滤与任务设计思路来构建经济学仿真被试｜可迁移到消费者跨期选择与政策偏好预测场景，例如模拟个体在不同利率或补贴政策下的储蓄消费决策｜研究设计：以真实用户的消费与问卷数据为基准，用微调后的LLM作为被试，施加利率变动或补贴政策处理，测量其消费-储蓄分配与政策支持度，并与真实面板数据对照"}},{"id":"2601.15114","version":2,"title":"From Who They Are to How They Act: Behavioral Traits in Generative Agent-Based Models of Social Media","zh_title":"从他们是谁到他们如何行动：基于生成式智能体的社交媒体模型中的行为特质","abstract":"Generative Agent-Based Modeling (GABM) leverages Large Language Models to create autonomous agents that simulate human behavior in social media environments, demonstrating potential for modeling information propagation, influence processes, and network phenomena. While existing frameworks characterize agents through demographic attributes, personality traits, and interests, they lack mechanisms to encode behavioral dispositions toward platform actions, causing agents to exhibit homogeneous engagement patterns rather than the differentiated participation styles observed on real platforms. In this paper, we investigate the role of behavioral traits as an explicit characterization layer to regulate agents' propensities across posting, re-sharing, commenting, reacting, and inactivity. Through large-scale simulations involving 980 agents and validation against real-world social media data, we demonstrate that behavioral traits are essential to sustain heterogeneous, profile-consistent participation patterns and enable realistic content propagation dynamics through the interplay of amplification- and interaction-oriented profiles. Our findings establish that modeling how agents act-not only who they are-is necessary for advancing GABM as a tool for studying social media phenomena.","authors":["Valerio La Gatta","Gian Marco Orlando","Marco Perillo","Ferdinando Tammaro","Vincenzo Moscato"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-01-21","first_seen":"2026-01-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15114","pdf_url":"https://arxiv.org/pdf/2601.15114","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","社交媒体模拟","行为特质"],"reason":"用LLM agent模拟社交媒体行为，并与真实数据对照，直接复现人类参与模式。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":25,"question":"行为特质作为显式表征层，能否让生成式智能体在社交媒体模拟中维持异质性参与模式、再现真实的传播动态和网络结构？","design":"在现有GABM框架上扩展，为980个LLM智能体同时赋予身份特质（来自FinePersonas数据集）和七种行为特质原型（如沉默观察者、内容放大器等），模拟发帖、转发、评论、反应和沉默等完整动作空间，测量参与模式异质性、内容传播级联和网络中心性。","baseline":"对照真实社交媒体数据，验证行为特质能否复现真实世界的网络结构。","findings":"行为特质能有效防止行为同质化，维持与原型一致的异质性参与模式；放大导向型特质驱动内容传播级联，互动导向型特质主导互动网络，且基于经验数据的行为特质能成功复现真实社交网络结构。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体复现社交媒体用户行为，并与真实数据对照，验证了行为特质对异质性参与和传播动态的必要性，直接回应了研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"该方法通过为LLM智能体显式赋予行为特质原型来维持异质性参与模式，可借鉴到经济仿真中作为施加个体差异的处理手段｜可迁移到政策公告的预期形成与信息传播研究，模拟不同投资者类型对政策信息的反应与扩散｜以LLM智能体为被试，赋予理性交易者、噪声交易者等行为特质，处理为发布货币政策公告，结果变量为资产价格波动与信息传播级联，对照真实市场微观数据"}},{"id":"2601.12727","version":1,"title":"AI-exhibited Personality Traits Can Shape Human Self-concept through Conversations","zh_title":"AI展现的人格特质可通过对话塑造人类自我概念","abstract":"Recent Large Language Model (LLM) based AI can exhibit recognizable and measurable personality traits during conversations to improve user experience. However, as human understandings of their personality traits can be affected by their interaction partners' traits, a potential risk is that AI traits may shape and bias users' self-concept of their own traits. To explore the possibility, we conducted a randomized behavioral experiment. Our results indicate that after conversations about personal topics with an LLM-based AI chatbot using GPT-4o default personality traits, users' self-concepts aligned with the AI's measured personality traits. The longer the conversation, the greater the alignment. This alignment led to increased homogeneity in self-concepts among users. We also observed that the degree of self-concept alignment was positively associated with users' conversation enjoyment. Our findings uncover how AI personality traits can shape users' self-concepts through human-AI conversation, highlighting both risks and opportunities. We provide important design implications for developing more responsible and ethical AI systems.","authors":["Jingshu Li","Tianqi Song","Nattapat Boonprakong","Zicheng Zhu","Yitian Yang","Yi-Chieh Lee"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-01-19","first_seen":"2026-01-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.12727","pdf_url":"https://arxiv.org/pdf/2601.12727","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人格影响","人机交互实验"],"reason":"用LLM对话实验研究AI人格对用户自我概念的影响，有随机对照实验和人类数据，揭…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":57,"question":"与具有人格特质的大语言模型AI聊天机器人进行个人话题对话，是否会使用户的自我概念向AI的人格特质对齐？","design":"本研究不是用LLM替代人类被试的仿真研究，而是以人类为被试的在线随机行为实验。采用混合因子设计，让参与者与基于GPT-4o默认人格的AI聊天机器人进行个人话题对话，测量对话前后用户自我概念的变化，以及对话时长、享受度等变量。","baseline":"无对照","findings":"与AI进行个人话题对话后，用户的自我概念会向AI所表现出的人格特质对齐，且对话时间越长，对齐程度越大。这种对齐导致用户间自我概念的同质化增强，且对齐程度与用户的对话享受度正相关。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM对话对人类心理的影响，虽非用LLM仿真人类，但揭示了人机交互中AI人格对用户自我概念的塑造效应，对评估AI社会影响和仿真可靠性有批判性参考价值，值得阅读原文以了解实验细节和效应边界。","inspiration":"该研究采用混合因子设计，通过对比对话前后自我概念的变化来测量AI人格的塑造效应，这种前后测设计可用于评估经济决策中的干预效果。｜可迁移到消费者金融决策场景，例如研究AI理财顾问的人格特质是否影响用户的投资偏好或风险态度。｜以真实投资者为被试，随机分配与具有不同人格（如谨慎型vs冒险型）的AI理财顾问对话，测量对话前后风险偏好问卷得分的变化，并与历史投资行为数据对照。"}},{"id":"2601.12343","version":1,"title":"How Well Do LLMs Predict Human Behavior? A Measure of their Pretrained Knowledge","zh_title":"LLM预测人类行为的效果如何？一种对其预训练知识的度量","abstract":"Large language models (LLMs) are increasingly used to predict human behavior. We propose a measure for evaluating how much knowledge a pretrained LLM brings to such a prediction: its equivalent sample size, defined as the amount of task-specific data needed to match the predictive accuracy of the LLM. We estimate this measure by comparing the prediction error of a fixed LLM in a given domain to that of flexible machine learning models trained on increasing samples of domain-specific data. We further provide a statistical inference procedure by developing a new asymptotic theory for cross-validated prediction error. Finally, we apply this method to the Panel Study of Income Dynamics. We find that LLMs encode considerable predictive information for some economic variables but much less for others, suggesting that their value as substitutes for domain-specific data differs markedly across settings.","authors":["Wayne Gao","Sukjin Han","Annie Liang"],"categories":["econ.EM","cs.AI","stat.ML"],"primary_category":"econ.EM","announce_type":"new","date":"2026-01-18","first_seen":"2026-01-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.12343","pdf_url":"https://arxiv.org/pdf/2601.12343","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","人类行为预测","等效样本量"],"reason":"直接评估LLM预测人类行为的能力，使用真实经济调查数据作为基准，并提出等效样本…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"如何量化预训练大语言模型在预测人类行为时所带来的领域特定知识的价值？","design":"本研究不是仿真实验，而是提出一种评估方法：将固定预训练LLM的预测误差，与在逐渐增大的领域特定数据上训练的灵活机器学习模型的误差进行比较，定义等效样本量为后者误差首次不劣于LLM时的训练样本量。","baseline":"使用收入动态面板研究（PSID）2021年波次的真实调查数据，包含人口统计、劳动力市场和家庭协变量，用于预测小时工资、住房拥有、饮酒和吸烟等结果。","findings":"LLM对不同经济变量的预测能力差异很大：预测小时工资的等效样本量约为20个观测值，而预测住房拥有则需约600个观测值。这表明LLM作为领域特定数据替代品的价值在不同任务中显著不同。","reliability":"论文关注数据泄露问题，采用训练截止日期明确的静态开源模型进行事后评估；方法假设比较算法的误差随训练数据量增加而单调递减；等效样本量的推断依赖于交叉验证风险估计的渐近正态性。","relevance":"该研究直接回应了研究者对LLM替代人类被试的可靠性与偏差的关切，提供了基于真实经济调查数据的量化基准，并揭示了仿真在不同任务中有效性的异质性，值得精读原文以掌握等效样本量估计方法及其适用边界。","inspiration":"该方法通过比较LLM与灵活机器学习模型的预测误差，定义等效样本量来量化LLM的领域知识价值，为评估LLM仿真可靠性提供了可操作的基准｜可迁移到消费者金融行为预测场景，如利用LLM预测家庭信贷违约或投资决策，评估其替代传统调查数据的可行性｜以LLM作为被试，输入PSID等调查中的家庭财务协变量，预测信贷违约状态，将LLM预测误差与在真实违约数据上训练的梯度提升模型比较，计算等效样本量，以真实信贷记录作为对照基准"}},{"id":"2601.11049","version":2,"title":"Predicting Biased Human Decision-Making with Large Language Models in Conversational Settings","zh_title":"用大语言模型预测对话场景中的人类有偏决策","abstract":"We examine whether large language models (LLMs) can predict biased decision-making in conversational settings, and whether their predictions capture not only human cognitive biases but also how those effects change under cognitive load. In a pre-registered study (N = 1,648), participants completed six classic decision-making tasks via a chatbot with dialogues of varying complexity. Participants exhibited two well-documented cognitive biases: the Framing Effect and the Status Quo Bias. Increased dialogue complexity resulted in participants reporting higher mental demand. This increase in cognitive load selectively, but significantly, increased the effect of the biases, demonstrating the load-bias interaction. We then evaluated whether LLMs (GPT-4, GPT-5, and open-source models) could predict individual decisions given demographic information and prior dialogue. While results were mixed across choice problems, LLM predictions that incorporated dialogue context were significantly more accurate in several key scenarios. Importantly, their predictions reproduced the same bias patterns and load-bias interactions observed in humans. Across all models tested, the GPT-4 family consistently aligned with human behavior, outperforming GPT-5 and open-source models in both predictive accuracy and fidelity to human-like bias patterns. These findings advance our understanding of LLMs as tools for simulating human decision-making and inform the design of conversational agents that adapt to user biases.","authors":["Stephen Pilli","Vivek Nallur"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-01-16","first_seen":"2026-01-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.11049","pdf_url":"https://arxiv.org/pdf/2601.11049","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","认知偏差","人类数据对照"],"reason":"用LLM预测人类决策偏差，有真实人类数据对照，涉及认知负荷与偏差交互，评估仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":5,"question":"LLM能否在对话环境中预测人类的偏差决策，并复现认知偏差及其与认知负荷的交互效应？","design":"使用GPT-4、GPT-5和开源模型，基于人口统计信息和对话上下文预测个体在六项经典决策任务中的选择；处理变量为对话复杂度（低/高），结果变量为预测准确率及偏差模式复现程度。","baseline":"预注册实验（N=1648）中人类被试通过聊天机器人完成决策任务，表现出框架效应和现状偏差，且对话复杂度增加导致认知负荷上升并选择性放大偏差。","findings":"LLM预测在纳入对话上下文后准确率显著提升，且GPT-4家族在预测准确性和偏差模式复现上均优于GPT-5和开源模型；LLM预测成功复现了人类中的偏差主效应和负荷-偏差交互效应。","reliability":"论文指出不同选择问题的预测结果参差不齐，且未系统探讨模型在极端认知负荷或非典型人口群体下的失效条件。","relevance":"该研究直接以真实人类数据为基准，检验LLM在对话决策场景中预测偏差行为的能力，并揭示了模型间差异，对评估LLM作为人类被试替代品的可靠性具有参考价值，值得阅读原文。","inspiration":"该方法通过操纵对话复杂度来施加认知负荷处理，并以真实人类实验数据为基准检验LLM的预测效度，值得借鉴｜可迁移到消费者在复杂金融产品选择中的决策偏差研究，如贷款方案或保险产品的框架效应｜以LLM模拟消费者，处理为产品信息呈现的对话复杂度（低/高），结果变量为选择偏差（如框架效应），对照真实消费者实验数据"}},{"id":"2601.15319","version":1,"title":"Large Language Models as Simulative Agents for Neurodivergent Adult Psychometric Profiles","zh_title":"大语言模型作为神经多样性成人心理测量特征的仿真代理","abstract":"Adult neurodivergence, including Attention-Deficit/Hyperactivity Disorder (ADHD), high-functioning Autism Spectrum Disorder (ASD), and Cognitive Disengagement Syndrome (CDS), is marked by substantial symptom overlap that limits the discriminant sensitivity of standard psychometric instruments. While recent work suggests that Large Language Models (LLMs) can simulate human psychometric responses from qualitative data, it remains unclear whether they can accurately and stably model neurodevelopmental traits rather than broad personality characteristics. This study examines whether LLMs can generate psychometric responses that approximate those of real individuals when grounded in a structured qualitative interview, and whether such simulations are sensitive to variations in trait intensity. Twenty-six adults completed a 29-item open-ended interview and four standardized self-report measures (ASRS, BAARS-IV, AQ, RAADS-R). Two LLMs (GPT-4o and Qwen3-235B-A22B) were prompted to infer an individual psychological profile from interview content and then respond to each questionnaire in-role. Accuracy, reliability, and sensitivity were assessed using group-level comparisons, error metrics, exact-match scoring, and a randomized baseline. Both models outperformed random responses across instruments, with GPT-4o showing higher accuracy and reproducibility. Simulated responses closely matched human data for ASRS, BAARS-IV, and RAADS-R, while the AQ revealed subscale-specific limitations, particularly in Attention to Detail. Overall, the findings indicate that interview-grounded LLMs can produce coherent and above-chance simulations of neurodevelopmental traits, supporting their potential use as synthetic participants in early-stage psychometric research, while highlighting clear domain-specific constraints.","authors":["Francesco Chiappone","Davide Marocco","Nicola Milano"],"categories":["q-bio.NC","cs.AI"],"primary_category":"q-bio.NC","announce_type":"new","date":"2026-01-16","first_seen":"2026-01-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15319","pdf_url":"https://arxiv.org/pdf/2601.15319","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真人类被试","心理测量","真实人类数据对照"],"reason":"用LLM仿真神经发育特质问卷回答，并与26名真人数据对照，评估准确性与局限性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":10,"question":"LLM能否基于结构化访谈内容，准确且稳定地模拟神经发育特质（ADHD、ASD、CDS）个体的心理测量反应？","design":"用GPT-4o和Qwen3-235B-A22B两个LLM，根据26名成人被试的29项开放式访谈文本推断个体心理画像，然后以角色扮演方式回答ASRS、BAARS-IV、AQ、RAADS-R四份标准化自评问卷，评估模拟回答的准确性、可靠性和对特质强度的敏感性。","baseline":"26名成人被试的真实访谈内容和四份问卷（ASRS、BAARS-IV、AQ、RAADS-R）的自评数据。","findings":"两个模型在所有问卷上的模拟回答均优于随机基线，GPT-4o准确性和可重复性更高；模拟回答在ASRS、BAARS-IV和RAADS-R上接近真人数据，但AQ量表在“注意细节”子量表上表现出局限性。","reliability":"论文承认AQ量表在特定子量表（如注意细节）上模拟准确性有限，且样本量较小（26人），可能影响结论的泛化性。","relevance":"该研究直接以真实人类数据为基准，评估LLM仿真神经发育特质问卷回答的准确性与局限性，符合研究者对仿真可靠性及失效条件的关注，值得阅读原文以了解具体偏差来源和实验设计细节。","inspiration":"借鉴其“基于结构化访谈生成个体画像并角色扮演回答问卷”的仿真设计，可迁移到经济金融领域的消费者偏好测量或投资者情绪评估场景。｜用LLM基于消费者深度访谈文本模拟其跨期选择问卷回答，以真实消费者面板数据为对照，评估LLM仿真消费决策偏差的准确性。"}},{"id":"2601.15312","version":1,"title":"Do people expect different behavior from large language models acting on their behalf? Evidence from norm elicitations in two canonical economic games","zh_title":"人们是否期望代表他们行事的大语言模型表现出不同行为？来自两个经典经济博弈中规范引出的证据","abstract":"While delegating tasks to large language models (LLMs) can save people time, there is growing evidence that offloading tasks to such models produces social costs. We use behavior in two canonical economic games to study whether people have different expectations when decisions are made by LLMs acting on their behalf instead of themselves. More specifically, we study the social appropriateness of a spectrum of possible behaviors: when LLMs divide resources on our behalf (Dictator Game and Ultimatum Game) and when they monitor the fairness of splits of resources (Ultimatum Game). We use the Krupka-Weber norm elicitation task to detect shifts in social appropriateness ratings. Results of two pre-registered and incentivized experimental studies using representative samples from the UK and US (N = 2,658) show three key findings. First, people find that offers from machines - when no acceptance is necessary - are judged to be less appropriate than when they come from humans, although there is no shift in the modal response. Second - when acceptance is necessary - it is more appropriate for a person to reject offers from machines than from humans. Third, receiving a rejection of an offer from a machine is no less socially appropriate than receiving the same rejection from a human. Overall, these results suggest that people apply different norms for machines deciding on how to split resources but are not opposed to machines enforcing the norms. The findings are consistent with offers made by machines now being viewed as having both a cognitive and emotional component.","authors":["Paweł Niszczota","Elia Antoniou"],"categories":["cs.GT","cs.AI","cs.CL","cs.CY","cs.HC","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2026-01-14","first_seen":"2026-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15312","pdf_url":"https://arxiv.org/pdf/2601.15312","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","经济博弈","社会规范"],"reason":"用LLM替代人类被试进行经济博弈实验，有真实人类数据对照，评估社会规范变化，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":20,"question":"当大语言模型（LLM）代替人类做出资源分配决策或监督公平时，人们对这些行为的社会适当性评价是否会发生变化？","design":"本研究并非用LLM作为人类被试的替代品进行仿真，而是通过在线实验调查人类被试对LLM代理行为的规范性评价。实验采用Krupka-Weber规范引出任务，让来自英国和美国的代表性样本（N=2658）对独裁者博弈和最后通牒博弈中一系列可能行为的社会适当性进行评分，比较决策者是人类还是LLM时的评分差异。","baseline":"无对照。研究直接比较人类被试对同一行为在“人类决策”与“LLM决策”两种情境下的适当性评分，未使用LLM生成行为并与真实人类行为数据对照。","findings":"第一，当无需接受方同意时（独裁者博弈），机器做出的分配提议被认为比人类做出的更不适当，但众数反应未变；第二，当需要接受方同意时（最后通牒博弈），拒绝机器提议比拒绝人类提议更适当；第三，收到机器的拒绝与收到人类的拒绝在社会适当性上无差异。","reliability":"论文未讨论","relevance":"该研究直接考察了人们对LLM代理经济决策的社会规范评价，虽非典型的LLM仿真实验，但为理解人机互动中的规范偏移提供了实证证据，对关注LLM替代人类被试时可能产生的规范性偏差的研究者具有参考价值。","inspiration":"借鉴其使用规范引出任务测量社会适当性评价的方法，可迁移至经济金融领域的伦理判断研究，如算法信贷审批或AI投资顾问的公众接受度。｜可应用于金融决策中的AI代理伦理：例如研究投资者对AI理财顾问做出高风险投资建议的适当性评价。｜设计：以普通投资者为被试，处理为投资建议来源（人类顾问 vs. AI顾问），结果变量为对建议适当性的评分，对照真实市场中人类顾问的建议接受率数据。"}},{"id":"2601.09772","version":1,"title":"Antisocial behavior towards large language model users: experimental evidence","zh_title":"针对大语言模型用户的反社会行为：实验证据","abstract":"The rapid spread of large language models (LLMs) has raised concerns about the social reactions they provoke. Prior research documents negative attitudes toward AI users, but it remains unclear whether such disapproval translates into costly action. We address this question in a two-phase online experiment (N = 491 Phase II participants; Phase I provided targets) where participants could spend part of their own endowment to reduce the earnings of peers who had previously completed a real-effort task with or without LLM support. On average, participants destroyed 36% of the earnings of those who relied exclusively on the model, with punishment increasing monotonically with actual LLM use. Disclosure about LLM use created a credibility gap: self-reported null use was punished more harshly than actual null use, suggesting that declarations of \"no use\" are treated with suspicion. Conversely, at high levels of use, actual reliance on the model was punished more strongly than self-reported reliance. Taken together, these findings provide the first behavioral evidence that the efficiency gains of LLMs come at the cost of social sanctions.","authors":["Paweł Niszczota","Cassandra Grützner"],"categories":["cs.AI","cs.CL","cs.CY","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-14","first_seen":"2026-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.09772","pdf_url":"https://arxiv.org/pdf/2601.09772","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为经济学","社会惩罚"],"reason":"用LLM替代人类被试，在真实努力任务中测量对LLM用户的惩罚行为，有真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":21,"question":"人们是否会因同伴使用大语言模型（LLM）完成任务而对其进行有代价的惩罚？","design":"本研究并非用LLM替代人类被试的仿真实验，而是以人类为被试的真实行为实验。实验分两阶段：第一阶段被试在有无LLM支持下完成真实努力任务，作为第二阶段的目标对象；第二阶段被试（N=491）可使用自己的部分报酬去减少第一阶段同伴的收入，结果变量为惩罚金额（即烧毁的金钱数量）。","baseline":"有真实人类数据作为对照：第一阶段被试的实际LLM使用情况（实际使用量）与自我报告的LLM使用情况，作为比较惩罚行为的基准。","findings":"被试平均烧毁了完全依赖LLM者36%的收入，且惩罚随实际LLM使用量单调递增。自我报告的零使用比实际零使用受到更严厉的惩罚，而高使用量下实际依赖比自我报告依赖受到更强惩罚，表明披露存在信任差距。","reliability":"论文未讨论","relevance":"该研究直接测量了对LLM使用者的真实惩罚行为，提供了有代价的反社会行为证据，与关注LLM仿真人类行为及社会规范的研究高度相关，值得精读原文以了解实验设计和行为测量方法。","inspiration":"采用真实努力任务与金钱燃烧博弈结合的设计，巧妙分离了实际使用与自我报告使用对惩罚的影响，并利用单调性检验强化因果推断。｜可迁移到信贷审批或招聘场景中，研究对算法辅助决策者的社会惩罚，例如贷款审批员使用AI模型是否会引发同事或客户的惩罚性行为。｜以金融从业者为被试，设计一个模拟贷款审批任务，处理组被告知审批员使用了AI辅助，对照组为纯人工审批，结果变量为被试愿意花费自身报酬去降低审批员奖金的行为，并以实际审批准确率数据作为基准对照。"}},{"id":"2601.09849","version":1,"title":"Strategies of cooperation and defection in five large language models","zh_title":"五种大语言模型中的合作与背叛策略","abstract":"Large language models (LLMs) are increasingly deployed to support human decision-making. This use of LLMs has concerning implications, especially when their prescriptions affect the welfare of others. To gauge how LLMs make social decisions, we explore whether five leading models produce sensible strategies in the repeated prisoner's dilemma, which is the main metaphor of reciprocal cooperation. First, we measure the propensity of LLMs to cooperate in a neutral setting, without using language reminiscent of how this game is usually presented. We record to what extent LLMs implement Nash equilibria or other well-known strategy classes. Thereafter, we explore how LLMs adapt their strategies to changes in parameter values. We vary the game's continuation probability, the payoff values, and whether the total number of rounds is commonly known. We also study the effect of different framings. In each case, we test whether the adaptations of the LLMs are in line with basic intuition, theoretical predictions of evolutionary game theory, and experimental evidence from human participants. While all LLMs perform well in many of the tasks, none of them exhibit full consistency over all tasks. We also conduct tournaments between the inferred LLM strategies and study direct interaction between LLMs in games over ten rounds with a known or unknown last round. Our experiments shed light on how current LLMs instantiate reciprocal cooperation.","authors":["Saptarshi Pal","Abhishek Mallela","Christian Hilbe","Lenz Pracher","Chiyu Wei","Feng Fu","Santiago Schnell","Martin A Nowak"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-01-14","first_seen":"2026-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.09849","pdf_url":"https://arxiv.org/pdf/2601.09849","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","行为博弈","人类数据对照"],"reason":"用LLM模拟重复囚徒困境中的人类合作策略，并与人类实验数据对照，直接命中核心判…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":22,"question":"当前主流大语言模型在重复囚徒困境中能否生成符合直觉、演化博弈理论预测和人类实验证据的互惠合作策略？","design":"以五款主流LLM（Claude、Gemini、GPT-4o、GPT-5、Llama）为被试，通过直接询问其对上一轮（或两轮）所有可能结果的反应来推断其策略，而非让LLM相互对弈；施加的处理包括改变博弈的继续概率、收益值、记忆轮数、是否已知最后一轮以及不同的框架提示，测量LLM的合作倾向、策略类型（如纳什均衡、伙伴/对手类别）及策略对参数变化的适应性。","baseline":"人类基准：将LLM的策略适应性变化与演化博弈论的理论预测及人类参与者的实验证据进行对比。","findings":"所有LLM在许多任务中表现良好，但无一在所有任务上表现出完全一致性；LLM的策略在部分参数变化下能做出合理调整，但整体缺乏稳健的互惠合作逻辑。","reliability":"论文指出，LLM在改变博弈参数和框架时策略适应性不一致，且当前实验仅限于特定模型和参数设置，未全面覆盖所有可能的策略空间和交互动态。","relevance":"该研究直接以LLM替代人类被试进行重复囚徒困境实验，并与人类实验数据和演化博弈理论对照，高度契合您对LLM仿真人类决策可靠性及偏差的关注，值得精读。","inspiration":"借鉴直接推断策略而非仅观察对弈结果的方法，可更精确刻画LLM的决策规则并与理论解对照｜可迁移至经济金融中的重复信任博弈或重复公共品博弈，如投资者在重复投资决策中的合作与背叛行为｜以LLM为被试，设计不同收益结构和终止概率的重复投资游戏，测量其策略类型与适应性，并与人类实验数据（如Fischbacher & Gächter, 2010）进行对照。"}},{"id":"2601.07110","version":2,"title":"The Need for a Socially-Grounded Persona Framework for User Simulation","zh_title":"面向用户仿真的社会根基化角色框架需求","abstract":"Synthetic personas are widely used to condition large language models (LLMs) for social simulation, yet most personas are still constructed from coarse sociodemographic attributes or summaries. We revisit persona creation by introducing SCOPE, a socially grounded framework for persona construction and evaluation, built from a 141-item, two-hour sociopsychological protocol collected from 124 U.S.-based participants. Across seven models, we find that demographic-only personas are a structural bottleneck: demographics explain only ~1.5% of variance in human response similarity. Adding sociopsychological facets improves behavioral prediction and reduces over-accentuation, and non-demographic personas based on values and identity achieve strong alignment with substantially lower bias. These trends generalize to SimBench (441 aligned questions), where SCOPE personas outperform default prompting and NVIDIA Nemotron personas, and SCOPE augmentation improves Nemotron-based personas. Our results indicate that persona quality depends on sociopsychological structure rather than demographic templates or summaries.","authors":["Pranav Narayanan Venkit","Yu Li","Yada Pruksachatkun","Chien-Sheng Wu"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-12","first_seen":"2026-01-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.07110","pdf_url":"https://arxiv.org/pdf/2601.07110","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM人类仿真","社会心理角色","仿真偏差评估"],"reason":"用真实人类数据构建persona并评估LLM仿真人类行为的可靠性，直接命中核心…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":57,"question":"如何构建基于社会心理学结构的人格框架（SCOPE）以提升LLM仿真人类行为的真实性与降低人口统计偏差？","design":"收集124名美国参与者的两小时141项社会心理学协议数据，构建包含人口统计、社会人口行为、价值观、人格特质、行为模式、身份叙事等八维度的SCOPE框架；在七个模型家族上，比较仅用人口统计、加入社会心理学维度、非人口统计（仅价值观与身份）等不同人格构建策略下的行为预测准确性和偏差。","baseline":"124名美国参与者的真实调查数据，包括人口统计、价值观、人格特质、行为模式等多维度测量，以及SimBench基准中的441个对齐问题。","findings":"仅用人口统计构建人格是结构性瓶颈，仅能解释人类行为相似性约1.5%的方差，且会导致模型系统性过度泛化；加入社会心理学维度可改善行为预测并降低人口统计过度强调，基于价值观和身份的非人口统计人格能实现强对齐且偏差更低。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真人类行为中的人格构建问题，用真实人类数据作为基准，系统评估了不同人格框架的可靠性与偏差，并揭示了人口统计人格的失效条件，高度契合研究者对仿真可靠性、偏差及批判性评估的关注。","inspiration":"该方法借鉴了用真实人类多维度社会心理学数据构建人格框架，并系统比较不同人格构建策略（仅人口统计、加入社会心理学维度、非人口统计）对行为预测偏差的影响｜可迁移到信贷审批歧视研究中，检验不同借款人画像（仅种族/性别 vs. 加入价值观、行为模式）如何影响LLM模拟的信贷员决策偏差｜以LLM作为虚拟信贷员，处理为不同人格构建策略下的贷款申请人档案，结果变量为贷款批准率及种族/性别偏差，用真实信贷审批数据（如HMDA数据）作为人类基准对照"}},{"id":"2601.05050","version":3,"title":"Large language models can effectively convince people to believe conspiracies","zh_title":"大语言模型能有效说服人们相信阴谋论","abstract":"Large language models (LLMs) have been shown to be persuasive across a variety of contexts. But it remains unclear whether this persuasive power advantages accuracy, or if bad actors can just as easily use LLMs to promote misbeliefs. Here, we investigate this question across four experiments in which participants (N = 3996 Americans) discussed a conspiracy theory they were uncertain about with an LLM we instructed to either argue against (\"debunking\") or for (\"bunking\") that conspiracy. Across several frontier models (with standard guardrails but prompted to allow lying), we did not find consistent evidence of a truth advantage: the LLMs were able to both substantially increase and decrease average conspiracy belief, and participants in the bunking condition rated the LLM as more informative and collaborative, and reported greater trust in AI, than those who were in the debunking condition. More encouragingly, however, debunking induced more large changes in belief, and subsequent corrections were able to reverse the bunking effect. Furthermore, simply prompting the model to only provide accurate information dramatically reduced bunking effectiveness, and one powerful frontier model (GPT 5.2) almost entirely refused to promote conspiracies, suggesting that it is possible for the right guardrails to favor accurate beliefs. Finally, we did find a stark truth asymmetry in the context of information sharing: debunking had a large positive impact on mock social media posts composed by participants, while bunking had little effect. Overall, our findings show that people are not inherently less susceptible to AI that misleads than to AI that informs, but that potential technical solutions exist to mitigate this risk.","authors":["Thomas H. Costello","Kellin Pelrine","Matthew Kowal","Jasper Timm","Antonio A. Arechar","Jean-François Godbout","Adam Gleave","David Rand","Gordon Pennycook"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-08","first_seen":"2026-01-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.05050","pdf_url":"https://arxiv.org/pdf/2601.05050","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类被试","说服实验"],"reason":"用LLM与真人被试互动，测量信念改变，有真实人类数据对照，涉及说服实验与偏差评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"LLM在说服人们相信或怀疑阴谋论时是否存在“真相优势”，即其说服力是否更有利于准确信息而非误导信息？","design":"本研究并非用LLM替代人类被试的仿真实验，而是让3996名美国真人参与者与LLM进行对话，LLM被随机分配为“驳斥”或“支持”参与者不确定的阴谋论，测量对话前后信念变化、对AI的信任等结果变量。","baseline":"无对照","findings":"LLM既能大幅增加也能大幅降低阴谋论信念，未发现一致的真相优势；但驳斥能引发更大的信念改变，且后续纠正可逆转支持效果，同时通过提示仅提供准确信息或使用强护栏模型可大幅削弱误导能力。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM在说服实验中对真实人类信念的影响，涉及误导与纠正效果对比，并评估了技术护栏的缓解作用，与研究者关注的LLM仿真可靠性及偏差问题高度相关，值得精读原文。","inspiration":"该研究采用LLM与真人进行个性化对话的干预设计，通过随机分配LLM角色（驳斥/支持）并测量对话前后信念变化，可借鉴其动态说服实验范式｜可迁移至政策公告的预期形成研究，例如用LLM模拟央行沟通对公众通胀预期的影响｜以真人投资者为被试，随机分配LLM提供鹰派或鸽派政策解读，测量对话前后通胀预期变化，并以历史调查数据或市场通胀互换利率作为真实对照"}},{"id":"2601.01546","version":1,"title":"Improving Behavioral Alignment in LLM Social Simulations via Context Formation and Navigation","zh_title":"通过情境形成与导航改进LLM社会仿真中的行为对齐","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior in experimental settings, but they systematically diverge from human decisions in complex decision-making environments, where participants must anticipate others' actions and form beliefs based on observed behavior. We propose a two-stage framework for improving behavioral alignment. The first stage, context formation, explicitly specifies the experimental design to establish an accurate representation of the decision task and its context. The second stage, context navigation, guides the reasoning process within that representation to make decisions. We validate this framework through a focal replication of a sequential purchasing game with quality signaling (Kremer and Debo, 2016), extending to a crowdfunding game with costly signaling (Cason et al., 2025) and a demand-estimation task (Gui and Toubia, 2025) to test generalizability across decision environments. Across four state-of-the-art (SOTA) models (GPT-4o, GPT-5, Claude-4.0-Sonnet-Thinking, DeepSeek-R1), we find that complex decision-making environments require both stages to achieve behavioral alignment with human benchmarks, whereas the simpler demand-estimation task requires only context formation. Our findings clarify when each stage is necessary and provide a systematic approach for designing and diagnosing LLM social simulations as complements to human subjects in behavioral research.","authors":["Letian Kong","Qianran","Jin","Renyu Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-04","first_seen":"2026-01-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.01546","pdf_url":"https://arxiv.org/pdf/2601.01546","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为对齐","经济学实验"],"reason":"直接复现人类行为实验，用LLM仿真被试并与真实人类数据对照，评估对齐条件与失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":11,"question":"在复杂决策环境中，如何系统性地诊断并改善LLM社会仿真与人类行为之间的对齐？","design":"使用GPT-4o、GPT-5、Claude-4.0-Sonnet-Thinking、DeepSeek-R1四个SOTA模型模拟人类被试，通过两阶段框架（情境形成与情境导航）施加处理，测量LLM决策与人类基准的行为对齐程度。","baseline":"对照三个已发表实验的真实人类数据：Kremer and Debo (2016) 的序贯购买博弈、Cason et al. (2025) 的众筹博弈、Gui and Toubia (2025) 的需求估计任务。","findings":"复杂决策环境需要同时使用情境形成和情境导航两个阶段才能实现行为对齐；而较简单的需求估计任务仅需情境形成即可。","reliability":"论文指出，复杂决策环境中的对齐失效源于LLM在战略互依和内生信念形成上的系统性偏差；框架的有效性可能依赖于任务类型，且未讨论模型规模、训练数据分布及提示敏感性等潜在局限。","relevance":"该研究直接复现人类行为实验，用LLM仿真被试并与真实人类数据对照，系统评估对齐条件与失效边界，高度契合研究者对LLM仿真可靠性及批判性检验的关注，值得精读。","inspiration":"两阶段框架提供了可操作的诊断与干预方法，通过显式设定实验情境和引导推理过程来改善对齐，可借鉴其处理施加方式与对照设计。｜可迁移至经济学中的信念形成实验，如资产定价中的信息级联、信贷审批中的信号博弈或政策公告的预期形成。｜以LLM为被试，模拟资产市场中的序贯交易，处理为是否提供情境形成与情境导航提示，结果变量为交易价格与理性预期均衡的偏差，对照真实人类实验数据（如Smith et al. 1988）。"}},{"id":"2512.23184","version":1,"title":"From Model Choice to Model Belief: Establishing a New Measure for LLM-Based Research","zh_title":"从模型选择到模型信念：为基于LLM的研究建立新度量","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output (\"model choice\") as a single data point underutilizes the information inherent to the probabilistic nature of LLMs. This paper introduces and formalizes \"model belief,\" a measure derived from an LLM's token-level probabilities that captures the model's belief distribution over choice alternatives in a single generation run. The authors prove that model belief is asymptotically equivalent to the mean of model choices (a non-trivial property) but forms a more statistically efficient estimator, with lower variance and a faster convergence rate. Analogous properties are shown to hold for smooth functions of model belief and model choice often used in downstream applications. The authors demonstrate the performance of model belief through a demand estimation study, where an LLM simulates consumer responses to different prices. In practical settings with limited numbers of runs, model belief explains and predicts ground-truth model choice better than model choice itself, and reduces the computation needed to reach sufficiently accurate estimates by roughly a factor of 20. The findings support using model belief as the default measure to extract more information from LLM-generated data.","authors":["Hongshen Sun","Juanjuan Zhang"],"categories":["cs.AI","econ.EM"],"primary_category":"cs.AI","announce_type":"new","date":"2025-12-29","first_seen":"2025-12-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.23184","pdf_url":"https://arxiv.org/pdf/2512.23184","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","需求估计","统计效率"],"reason":"用LLM仿真消费者需求，提出模型信念度量提升统计效率，有真实人类数据对照，涉及…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":29,"question":"如何从LLM的token级概率中提取更高效的“模型信念”度量，以替代常用的“模型选择”来模拟人类行为？","design":"使用LLM模拟消费者对不同价格的响应，通过单次生成获取token级对数概率，构造模型信念分布，并与多次采样得到的模型选择进行比较。","baseline":"无对照","findings":"模型信念是模型选择均值的渐近等价估计量，但方差更低、收敛更快；在需求估计中，模型信念仅需约1/20的计算量即可达到同等精度，且对真实模型选择的解释和预测能力更强。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真人类行为的效率问题，提出模型信念度量以提升统计效率，与研究者关心的仿真可靠性和计算成本高度相关，值得精读。","inspiration":"借鉴从LLM内部概率分布直接提取信念度量的方法，替代重复采样，大幅降低计算成本并提高估计精度。｜可迁移到消费者需求估计、价格弹性测量等营销与经济交叉场景，也可用于政策评估中的个体偏好推断。｜以LLM作为被试，模拟不同价格或政策条件下的选择，提取模型信念作为选择概率的连续度量，以真实市场扫描数据或实验数据作为对照，检验模型信念对真实弹性的恢复能力。"}},{"id":"2512.22725","version":1,"title":"Mitigating Social Desirability Bias in Random Silicon Sampling","zh_title":"缓解随机硅采样中的社会赞许性偏差","abstract":"Large Language Models (LLMs) are increasingly used to simulate population responses, a method known as ``Silicon Sampling''. However, responses to socially sensitive questions frequently exhibit Social Desirability Bias (SDB), diverging from real human data toward socially acceptable answers. Existing studies on social desirability bias in LLM-based sampling remain limited. In this work, we investigate whether minimal, psychologically grounded prompt wording can mitigate this bias and improve alignment between silicon and human samples. We conducted a study using data from the American National Election Study (ANES) on three LLMs from two model families: the open-source Llama-3.1 series and GPT-4.1-mini. We first replicate a baseline silicon sampling study, confirming the persistent Social Desirability Bias. We then test four prompt-based mitigation methods: \\emph{reformulated} (neutral, third-person phrasing), \\emph{reverse-coded} (semantic inversion), and two meta-instructions, \\emph{priming} and \\emph{preamble}, respectively encouraging analytics and sincerity. Alignment with ANES is evaluated using Jensen-Shannon Divergence with bootstrap confidence intervals. Our results demonstrate that reformulated prompts most effectively improve alignment by reducing distribution concentration on socially acceptable answers and achieving distributions closer to ANES. Reverse-coding produced mixed results across eligible items, while the Priming and Preamble encouraged response uniformity and showed no systematic benefit for bias mitigation. Our findings validate the efficacy of prompt-based framing controls in mitigating inherent Social Desirability Bias in LLMs, providing a practical path toward more representative silicon samples.","authors":["Sashank Chapala","Maksym Mironov","Songgaojun Deng"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-12-27","first_seen":"2025-12-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.22725","pdf_url":"https://arxiv.org/pdf/2512.22725","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["硅采样","社会赞许性偏差","人类数据对照"],"reason":"直接研究LLM仿真人类调查回答，用ANES真实数据对照，评估并缓解社会赞许性偏…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:43","error":null,"has_summary":true,"summary":{"generated_at":"2025-12-27","rank":6,"question":"能否通过最小化的、基于心理学的提示措辞来减轻LLM硅采样中的社会赞许性偏差，从而提高硅样本与人类样本的对齐度？","design":"使用ANES 2020数据，从真实人类分布中抽取人口统计特征，生成硅样本（Llama-3.1系列和GPT-4.1-mini）。测试四种提示策略：reformulated（中性第三人称）、reverse-coded（语义反转）、priming（鼓励分析）和preamble（鼓励真诚），以Jensen-Shannon散度评估与ANES的对齐。","baseline":"美国国家选举研究（ANES）2020年选举前调查数据，包含5441名受访者，覆盖种族、年龄、性别等人口统计变量及10个社会政治问题。","findings":"Reformulated提示最有效，通过减少对社会可接受答案的集中分布，使硅样本分布更接近ANES；reverse-coded效果不一，priming和preamble导致回答趋同，无系统性改善。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接研究LLM仿真人类被试的社会赞许性偏差，使用真实人类数据ANES作为对照，并系统评估了多种提示缓解策略，符合研究者对经济学实验和政策评估场景的兴趣。","inspiration":"借鉴其通过最小化提示措辞（如中性第三人称重构）来系统操纵社会赞许性偏差的方法，并以Jensen-Shannon散度量化硅样本与真实人类调查分布的对齐度｜可迁移至信贷审批中的种族歧视测量，如用LLM模拟贷款官员对相同财务档案但不同种族姓名的审批决策｜以LLM作为被试，随机分配带有不同种族暗示姓名的贷款申请，处理为中性重构提示（如将'你会批准吗'改为'该申请是否符合标准'），结果变量为审批率差异，用真实房贷数据（如HMDA）作为人类基准对照"}},{"id":"2512.19937","version":1,"title":"Interpolative Decoding: Exploring the Spectrum of Personality Traits in LLMs","zh_title":"插值解码：探索大语言模型中人格特质的谱系","abstract":"Recent research has explored using very large language models (LLMs) as proxies for humans in tasks such as simulation, surveys, and studies. While LLMs do not possess a human psychology, they often can emulate human behaviors with sufficiently high fidelity to drive simulations to test human behavioral hypotheses, exhibiting more nuance and range than the rule-based agents often employed in behavioral economics. One key area of interest is the effect of personality on decision making, but the requirement that a prompt must be created for every tested personality profile introduces experimental overhead and degrades replicability. To address this issue, we leverage interpolative decoding, representing each dimension of personality as a pair of opposed prompts and employing an interpolation parameter to simulate behavior along the dimension. We show that interpolative decoding reliably modulates scores along each of the Big Five dimensions. We then show how interpolative decoding causes LLMs to mimic human decision-making behavior in economic games, replicating results from human psychological research. Finally, we present preliminary results of our efforts to ``twin'' individual human players in a collaborative game through systematic search for points in interpolation space that cause the system to replicate actions taken by the human subject.","authors":["Eric Yeh","John Cadigan","Ran Chen","Dick Crouch","Melinda Gervasio","Dayne Freitag"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-12-23","first_seen":"2025-12-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.19937","pdf_url":"https://arxiv.org/pdf/2512.19937","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人格特质","经济博弈"],"reason":"用LLM模拟人格影响经济决策，并与真实人类数据对照，直接复现人类行为实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":42,"question":"能否通过插值解码在LLM中连续调制大五人格维度，并使其在心理量表和经济游戏中复现人类行为，进而实现对个体人类玩家的“孪生”？","design":"使用通用LLM，将每种人格维度表示为一对对立提示，通过插值解码混合其输出分布来模拟人格谱系上的中间点；测量LLM在大五人格量表上的得分，以及在独裁者博弈等经济游戏中的决策行为，并尝试通过搜索插值空间来匹配特定人类玩家的行动。","baseline":"对照真实人类心理学研究中大五人格量表得分与经济游戏决策行为的相关性结果。","findings":"插值解码能可靠地沿大五人格各维度调节LLM的得分；LLM在经济游戏中表现出与人类心理学研究一致的人格-决策关联，并能通过插值空间搜索初步复现个体人类玩家的行为。","reliability":"论文承认当前仅探索了单一人格维度的插值，未涉及多维度联合调制，这限制了孪生等需要多因素行为解释的应用；此外，实验维度有限，主要目的是验证插值解码的可行性。","relevance":"该研究直接使用LLM模拟人格对经济决策的影响，并与真实人类数据对照，复现了人类行为实验，高度契合研究者对LLM作为人类被试替代品及其可靠性的关注，值得精读。","inspiration":"借鉴插值解码方法，通过混合对立提示的输出分布来连续调节LLM的人格维度，实现精细化的心理特质操控｜可迁移到资产定价实验中，研究风险偏好或时间偏好等心理特质对投资决策的影响｜以LLM为被试，通过插值解码调节其风险偏好水平，测量其在模拟资产选择任务中的投资组合，并与真实人类投资者的风险偏好问卷及实际投资数据对照"}},{"id":"2512.14306","version":1,"title":"Inflation Attitudes of Large Language Models","zh_title":"大语言模型的通胀态度","abstract":"This paper investigates the ability of Large Language Models (LLMs), specifically GPT-3.5-turbo (GPT), to form inflation perceptions and expectations based on macroeconomic price signals. We compare the LLM's output to household survey data and official statistics, mimicking the information set and demographic characteristics of the Bank of England's Inflation Attitudes Survey (IAS). Our quasi-experimental design exploits the timing of GPT's training cut-off in September 2021 which means it has no knowledge of the subsequent UK inflation surge. We find that GPT tracks aggregate survey projections and official statistics at short horizons. At a disaggregated level, GPT replicates key empirical regularities of households' inflation perceptions, particularly for income, housing tenure, and social class. A novel Shapley value decomposition of LLM outputs suited for the synthetic survey setting provides well-defined insights into the drivers of model outputs linked to prompt content. We find that GPT demonstrates a heightened sensitivity to food inflation information similar to that of human respondents. However, we also find that it lacks a consistent model of consumer price inflation. More generally, our approach could be used to evaluate the behaviour of LLMs for use in the social sciences, to compare different models, or to assist in survey design.","authors":["Nikoleta Anesti","Edward Hill","Andreas Joseph"],"categories":["cs.CL","econ.EM"],"primary_category":"cs.CL","announce_type":"new","date":"2025-12-16","first_seen":"2025-12-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.14306","pdf_url":"https://arxiv.org/pdf/2512.14306","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B3"],"tags":["LLM仿真","通胀预期","人类数据对照"],"reason":"用LLM模拟家庭通胀态度，与真实调查数据对照，评估仿真可靠性，涉及经济学实验和…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"大语言模型能否基于宏观经济价格信号形成通胀感知和预期，并复现家庭调查中的行为模式？","design":"使用GPT-3.5-turbo模拟英国家庭，根据英格兰银行通胀态度调查（IAS）的真实受访者人口特征构建合成人格，并输入不同价格信号作为处理，测量模型对当前和未来通胀的感知与预期。","baseline":"英格兰银行通胀态度调查（IAS）的真实家庭微观数据和官方通胀统计。","findings":"GPT在总体层面能较好匹配短期调查预测和官方统计，并复现收入、住房、社会阶层等维度的通胀感知规律，但对食品通胀信息过度敏感，且缺乏一致的消费者价格通胀模型。","reliability":"模型在微观个体层面与人类对应较弱且不稳定，经济条件环境简陋，且模型内部逻辑不一致，缺乏对通胀概念的连贯世界模型。","relevance":"该研究直接用LLM替代人类被试进行经济学调查仿真，并与真实家庭数据严格对照，评估了仿真可靠性及偏差，完全契合研究者对LLM人类仿真实验、经济学场景和批判性评估的关注，值得精读原文。","inspiration":"该方法通过构建合成人格并输入不同价格信号作为处理，测量LLM的感知与预期，并与真实家庭调查数据严格对照，值得借鉴｜可迁移到货币政策公告的预期形成研究，模拟家庭或投资者对利率变动的通胀与资产价格预期｜以LLM模拟不同人口特征的投资者，处理为不同措辞的央行公告，结果变量为通胀预期和股票投资意愿，对照真实调查或市场数据"}},{"id":"2512.11827","version":2,"title":"Assessing Greenspace Attractiveness with ChatGPT, Claude, and Gemini: Do AI Models Reflect Human Perceptions?","zh_title":"用ChatGPT、Claude和Gemini评估绿地吸引力：AI模型是否反映人类感知？","abstract":"Understanding greenspace attractiveness is essential for designing livable and inclusive urban environments, yet existing assessment approaches often overlook informal or transient spaces and remain too resource intensive to capture subjective perceptions at scale. This study examines the ability of multimodal large language models (MLLMs), ChatGPT GPT-4o, Claude 3.5 Haiku, and Gemini 2.0 Flash, to assess greenspace attractiveness similarly to humans using Google Street View imagery. We compared model outputs with responses from a geo-questionnaire of residents in Lodz, Poland, across both formal (for example, parks and managed greenspaces) and informal (for example, meadows and wastelands) greenspaces. Survey respondents and models indicated whether each greenspace was attractive or unattractive and provided up to three free text explanations. Analyses examined how often their attractiveness judgments aligned and compared their explanations after classifying them into shared reasoning categories. Results show high AI human agreement for attractive formal greenspaces and unattractive informal spaces, but low alignment for attractive informal and unattractive formal greenspaces. Models consistently emphasized aesthetic and design oriented features, underrepresenting safety, functional infrastructure, and locally embedded qualities valued by survey respondents. While these findings highlight the potential for scalable pre-assessment, they also underscore the need for human oversight and complementary participatory approaches. We conclude that MLLMs can support, but not replace, context sensitive greenspace evaluation in planning practice.","authors":["Milad Malekzadeh","Magdalena Biernacka","Elias Willberg","Jussi Torkko","Edyta Łaszkiewicz","Tuuli Toivonen"],"categories":["cs.CY","cs.AI","cs.CV"],"primary_category":"cs.CY","announce_type":"new","date":"2025-12-02","first_seen":"2025-12-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.11827","pdf_url":"https://arxiv.org/pdf/2512.11827","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","绿地感知评估","AI与人类对照"],"reason":"用LLM评估绿地吸引力并与人类问卷对照，涉及仿真可靠性、偏差及失效条件，直接相…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":12,"question":"多模态大语言模型（ChatGPT、Claude、Gemini）对城市绿地吸引力的评估是否与人类感知一致？","design":"使用Google街景图像作为输入，让三个MLLM（ChatGPT GPT-4o、Claude 3.5 Haiku、Gemini 2.0 Flash）判断正式和非正式绿地的吸引力（有吸引力/无吸引力），并生成最多三条自由文本解释；将模型输出与波兰罗兹市居民的地理问卷调查结果进行比较。","baseline":"波兰罗兹市居民的地理问卷调查，包含对正式和非正式绿地的吸引力判断及自由文本解释。","findings":"模型在评估有吸引力的正式绿地和无吸引力的非正式绿地时与人类一致性高，但在有吸引力的非正式绿地和无吸引力的正式绿地上一致性低。模型过度强调美学和设计特征，而低估了安全、功能设施和本地化品质。","reliability":"模型未能充分捕捉安全、功能设施和本地化品质等人类重视的特征；在非典型场景（有吸引力的非正式绿地、无吸引力的正式绿地）中一致性低；可能存在隐藏的性别和其他偏见；无法代表不同人口群体的感知差异。","relevance":"该研究直接以真实人类数据为基准，检验LLM在主观感知评估中的可靠性与偏差，并明确指出了仿真失效的条件（如非典型绿地类型、忽视安全与本地化特征），与研究者关注的LLM仿真实验高度相关，值得阅读原文以了解具体实验设计和偏差分析。","inspiration":"借鉴其将模型输出与人类解释进行定性分类比较的方法，可迁移到消费者对金融产品广告的感知评估或投资者对年报文本的情绪解读研究中。｜可应用于行为金融中的信息感知实验，例如研究投资者对上市公司年报中风险披露的吸引力判断。｜以真实投资者问卷调查为基准，让LLM阅读年报摘要并评估其投资吸引力及理由，比较模型与人类在风险感知、语言特征关注上的差异，检验模型是否过度关注表面语言而忽略深层风险信号。"}},{"id":"2512.07890","version":1,"title":"CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models","zh_title":"CrowdLLM：结合生成模型构建基于LLM的数字人群","abstract":"The emergence of large language models (LLMs) has sparked much interest in creating LLM-based digital populations that can be applied to many applications such as social simulation, crowdsourcing, marketing, and recommendation systems. A digital population can reduce the cost of recruiting human participants and alleviate many concerns related to human subject study. However, research has found that most of the existing works rely solely on LLMs and could not sufficiently capture the accuracy and diversity of a real human population. To address this limitation, we propose CrowdLLM that integrates pretrained LLMs and generative models to enhance the diversity and fidelity of the digital population. We conduct theoretical analysis of CrowdLLM regarding its great potential in creating cost-effective, sufficiently representative, scalable digital populations that can match the quality of a real crowd. Comprehensive experiments are also conducted across multiple domains (e.g., crowdsourcing, voting, user rating) and simulation studies which demonstrate that CrowdLLM achieves promising performance in both accuracy and distributional fidelity to human data.","authors":["Ryan Feng Lin","Keyu Tian","Hanming Zheng","Congjing Zhang","Li Zeng","Shuai Huang"],"categories":["cs.MA","cs.AI","cs.LG","stat.ME","stat.ML"],"primary_category":"cs.MA","announce_type":"new","date":"2025-12-02","first_seen":"2025-12-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.07890","pdf_url":"https://arxiv.org/pdf/2512.07890","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","数字人群","人类数据对照"],"reason":"用LLM构建数字人群，模拟众包、投票等人类行为，并与真实人类数据对照，直接命中…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":43,"question":"如何结合预训练大语言模型与生成模型，构建能准确复现真实人群决策多样性与分布保真度的数字人群？","design":"提出CrowdLLM框架，将预训练LLM与生成式机器学习模型集成，通过概率框架生成虚拟参与者并聚合其决策，模拟众包、投票、用户评分等场景中的群体行为。","baseline":"使用真实人类在众包、投票、用户评分等任务上的决策数据作为对照基准。","findings":"CrowdLLM在多个领域实验中，生成的数字人群在决策准确性和分布保真度上均与真实人类数据高度匹配；理论分析表明该框架具有成本效益高、代表性强和可扩展的潜力。","reliability":"论文未讨论","relevance":"该研究直接构建LLM数字人群并对照真实人类数据，评估仿真准确性与分布保真度，覆盖众包、投票等场景，高度契合研究者对LLM人类仿真实验及基准对照的关注，值得精读原文。","inspiration":"该方法通过概率框架集成生成模型来捕获决策多样性，可借鉴用于模拟经济主体异质性｜可迁移到消费者跨期选择实验，研究不同贴现因子分布下的储蓄行为｜以LLM生成虚拟消费者，处理为不同利率条件，结果变量为储蓄金额，对照真实家庭金融调查数据"}},{"id":"2511.21218","version":3,"title":"Can Finetuing LLMs on Small Human Samples Increase Heterogeneity, Alignment, and Belief-Action Coherence?","zh_title":"在小规模人类样本上微调LLM能否增加异质性、对齐度和信念-行动一致性？","abstract":"There is ongoing debate about whether large language models (LLMs) can serve as substitutes for human participants in survey and experimental research. While recent work in fields such as marketing and psychology has explored the potential of LLM-based simulation, a growing body of evidence cautions against this practice: LLMs often fail to align with real human behavior, exhibiting limited diversity, systematic misalignment for minority subgroups, insufficient within-group variance, and discrepancies between stated beliefs and actions. This study examines an important and distinct question in this domain: whether fine-tuning on a small subset of human survey data, such as that obtainable from a pilot study, can mitigate these issues and yield realistic simulated outcomes. Using a behavioral experiment on information disclosure, we compare human and LLM-generated responses across multiple dimensions, including distributional divergence, subgroup alignment, belief-action coherence, and the recovery of regression coefficients. We find that fine-tuning on small human samples substantially improves heterogeneity, alignment, and belief-action coherence relative to the base model. However, even the best-performing fine-tuned models fail to reproduce the regression coefficients of the original study, suggesting that LLM-generated data remain unsuitable for replacing human participants in formal inferential analyses.","authors":["Steven Wang","Kyle Hunt","Shaojie Tang","Kenneth Joseph"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-11-26","first_seen":"2025-11-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.21218","pdf_url":"https://arxiv.org/pdf/2511.21218","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","行为实验","微调偏差"],"reason":"直接研究微调LLM作为人类被试替代，用真实人类实验数据对照，评估仿真可靠性及失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"在小规模人类样本上微调LLM能否提高仿真中的异质性、对齐度与信念-行动一致性？","design":"使用开源LLM（如Llama-3）在Hunt等人关于信息披露的行为实验数据上微调，模拟攻击者的信念与决策，比较基础模型与微调模型在分布差异、子群对齐、信念-行动一致性及回归系数恢复上的表现。","baseline":"Hunt等人收集的真实人类行为实验数据，包含被试对安全技术部署的信念和攻击决策。","findings":"微调显著改善了异质性、分布对齐和信念-行动一致性，但即使最佳微调模型也无法复现原始研究的回归系数，表明LLM生成数据仍不适合替代人类进行推断性分析。","reliability":"微调模型无法恢复原始回归系数和假设检验结果，在正式推断分析中可能引入不可忽视的偏差；模型泛化性有限，可能不适用于分布外行为情境。","relevance":"直接探讨用微调LLM替代人类被试的可行性与局限，以真实人类实验为基准，评估仿真可靠性及失效条件，高度契合研究者对经济学实验和政策评估场景的关注，值得精读原文。","inspiration":"该方法通过在小规模人类样本上微调LLM来提升仿真异质性和分布对齐，可借鉴其微调策略和以真实人类实验为基准的对照设计｜可迁移到政策公告的预期形成实验，例如研究央行沟通对公众通胀预期的影响｜使用Llama-3在真实调查数据上微调，模拟公众对通胀公告的预期更新，处理为不同措辞的政策声明，结果变量为预期通胀率，以密歇根大学消费者调查数据作为人类基准对照"}},{"id":"2511.05766","version":1,"title":"Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs","zh_title":"机器中的锚定：LLM中锚定偏差的行为与归因证据","abstract":"Large language models (LLMs) are increasingly examined as both behavioral subjects and decision systems, yet it remains unclear whether observed cognitive biases reflect surface imitation or deeper probability shifts. Anchoring bias, a classic human judgment bias, offers a critical test case. While prior work shows LLMs exhibit anchoring, most evidence relies on surface-level outputs, leaving internal mechanisms and attributional contributions unexplored. This paper advances the study of anchoring in LLMs through three contributions: (1) a log-probability-based behavioral analysis showing that anchors shift entire output distributions, with controls for training-data contamination; (2) exact Shapley-value attribution over structured prompt fields to quantify anchor influence on model log-probabilities; and (3) a unified Anchoring Bias Sensitivity Score integrating behavioral and attributional evidence across six open-source models. Results reveal robust anchoring effects in Gemma-2B, Phi-2, and Llama-2-7B, with attribution signaling that the anchors influence reweighting. Smaller models such as GPT-2, Falcon-RW-1B, and GPT-Neo-125M show variability, suggesting scale may modulate sensitivity. Attributional effects, however, vary across prompt designs, underscoring fragility in treating LLMs as human substitutes. The findings demonstrate that anchoring bias in LLMs is robust, measurable, and interpretable, while highlighting risks in applied domains. More broadly, the framework bridges behavioral science, LLM safety, and interpretability, offering a reproducible path for evaluating other cognitive biases in LLMs.","authors":["Felipe Valencia-Clavijo"],"categories":["cs.AI","cs.CL","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2025-11-07","first_seen":"2025-11-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.05766","pdf_url":"https://arxiv.org/pdf/2511.05766","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B4"],"tags":["锚定偏差","LLM行为仿真","可解释性"],"reason":"研究LLM的锚定偏差，评估其作为人类替代品的可靠性，并指出仿真脆弱性，方法可迁…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":58,"question":"大语言模型表现出的锚定偏差是表面模仿还是深层概率偏移？","design":"以六个开源LLM（GPT-2、GPT-Neo-125M、Falcon-RW-1B、Gemma-2B、Phi-2、Llama-2-7B）为被试，通过结构化提示（含高/低锚数字）复现经典锚定实验，测量候选答案的对数概率分布变化，并用Shapley值归因锚字段对对数概率的贡献。","baseline":"复现Tversky和Kahneman的“非洲国家在联合国占比”锚定实验，以人类在该实验中的锚定效应作为对照基准。","findings":"Gemma-2B、Phi-2和Llama-2-7B表现出稳健的锚定效应，锚数字导致整个输出分布偏移；归因分析显示锚字段对模型对数概率有显著贡献，但效应因提示设计而异，表明将LLM作为人类替代品存在脆弱性。","reliability":"归因效应随提示设计变化，显示将LLM视为人类替代品的脆弱性；小模型（如GPT-2、Falcon-RW-1B、GPT-Neo-125M）结果不稳定，暗示模型规模可能调节敏感性。","relevance":"该研究直接评估LLM作为人类被试替代品的可靠性，通过行为与归因双重证据揭示锚定偏差的深层机制与失效条件，方法可迁移至其他认知偏差，对经济学实验和政策评估中的仿真应用具有批判性参考价值，值得精读。","inspiration":"该方法通过结构化提示复现经典锚定实验，并用对数概率分布偏移和Shapley值归因来区分表面模仿与深层概率偏移，提供了行为与归因双重证据的稳健性检验思路｜可迁移到资产定价实验，研究投资者在估值时受历史价格或分析师目标价锚定的影响｜以LLM为被试，在估值提示中嵌入高/低历史价格锚，测量估值输出的对数概率分布偏移，并以真实投资者估值数据或实验数据作为对照基准"}},{"id":"2511.03758","version":3,"title":"Leveraging LLM-based agents for social science research: insights from citation network simulations","zh_title":"利用基于大语言模型的智能体进行社会科学研究：来自引文网络模拟的见解","abstract":"The emergence of Large Language Models (LLMs) demonstrates their potential to encapsulate the logic and patterns inherent in human behavior simulation by leveraging extensive web data pre-training. However, the boundaries of LLM capabilities in social simulation remain unclear. To further explore the social attributes of LLMs, we introduce the CiteAgent framework, designed to generate citation networks based on human-behavior simulation with LLM-based agents. CiteAgent successfully captures predominant phenomena in real-world citation networks, including power-law distribution, citational distortion, and shrinking diameter. Building on this realistic simulation, we establish two LLM-based research paradigms in social science: LLM-SE (LLM-based Survey Experiment) and LLM-LE (LLM-based Laboratory Experiment). These paradigms facilitate rigorous analyses of citation network phenomena, allowing us to validate and challenge existing theories. Additionally, we extend the research scope of traditional science of science studies through idealized social experiments, with the simulation experiment results providing valuable insights for real-world academic environments. Our work demonstrates the potential of LLMs for advancing science of science research in social science.","authors":["Jiarui Ji","Runlin Lei","Xuchen Pan","Zhewei Wei","Hao Sun","Yankai Lin","Xu Chen","Yongzheng Yang","Yaliang Li","Bolin Ding","Ji-Rong Wen"],"categories":["physics.soc-ph","cs.AI","cs.CY","cs.MA","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2025-11-05","first_seen":"2025-11-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.03758","pdf_url":"https://arxiv.org/pdf/2511.03758","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","引文网络","社会科学实验"],"reason":"用LLM agent模拟引文网络并与真实数据对照，提出LLM-SE/LE范式，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":15,"question":"LLM代理能否在引文网络模拟中复现真实网络的关键现象，并用于社会科学研究？","design":"构建CiteAgent框架，用GPT-3.5、GPT-4o-mini、LLAMA-3-70B扮演作者，基于种子网络生成引文网络，测量度分布幂律拟合度；并通过LLM-SE调查实验和LLM-LE实验室实验操纵论文属性与推荐算法，分析引用选择的影响因素。","baseline":"对照CiteSeer和Cora真实引文网络数据集，验证生成网络的幂律分布、引用扭曲和直径收缩现象。","findings":"CiteAgent生成的引文网络能复现幂律分布等真实网络现象，但GPT-3.5拟合较差；LLM-SE和LLM-LE实验表明，引用相关属性（如论文被引量、作者被引量、时效性）是导致优先连接和幂律分布的主要因素。","reliability":"论文指出GPT-3.5在LLM-Agent数据集上无法完美拟合幂律分布，表明不同LLM的仿真能力存在差异；但未系统讨论其他失效条件或局限。","relevance":"该研究直接以真实引文网络为基准，用LLM代理复现社会现象并分析机制，提出了LLM-SE和LLM-LE范式，高度契合您对LLM人类仿真实验、经济学/政策评估场景及可靠性评估的关注，值得精读。","inspiration":"该研究通过LLM代理模拟引文网络生成，并操纵论文属性（如被引量、时效性）和推荐算法作为处理，以真实引文网络为基准测量度分布拟合度，提供了因果推断与仿真验证结合的设计范例。｜可迁移至金融市场信息扩散研究，例如模拟分析师报告或新闻如何通过投资者关注网络传播并影响资产价格。｜以LLM扮演投资者，处理为信息源的突出特征（如来源权威性、时效性），结果变量为投资者关注分配与价格波动，对照真实股票论坛引用网络或价格联动数据。"}},{"id":"2511.02458","version":1,"title":"Prompting for Policy: Forecasting Macroeconomic Scenarios with Synthetic LLM Personas","zh_title":"用合成LLM角色预测宏观经济情景的政策提示","abstract":"We evaluate whether persona-based prompting improves Large Language Model (LLM) performance on macroeconomic forecasting tasks. Using 2,368 economics-related personas from the PersonaHub corpus, we prompt GPT-4o to replicate the ECB Survey of Professional Forecasters across 50 quarterly rounds (2013-2025). We compare the persona-prompted forecasts against the human experts panel, across four target variables (HICP, core HICP, GDP growth, unemployment) and four forecast horizons. We also compare the results against 100 baseline forecasts without persona descriptions to isolate its effect. We report two main findings. Firstly, GPT-4o and human forecasters achieve remarkably similar accuracy levels, with differences that are statistically significant yet practically modest. Our out-of-sample evaluation on 2024-2025 data demonstrates that GPT-4o can maintain competitive forecasting performance on unseen events, though with notable differences compared to the in-sample period. Secondly, our ablation experiment reveals no measurable forecasting advantage from persona descriptions, suggesting these prompt components can be omitted to reduce computational costs without sacrificing accuracy. Our results provide evidence that GPT-4o can achieve competitive forecasting accuracy even on out-of-sample macroeconomic events, if provided with relevant context data, while revealing that diverse prompts produce remarkably homogeneous forecasts compared to human panels.","authors":["Giulia Iadisernia","Carolina Camassa"],"categories":["cs.CL","cs.CE","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2025-11-04","first_seen":"2025-11-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.02458","pdf_url":"https://arxiv.org/pdf/2511.02458","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","宏观经济预测","人类数据对照"],"reason":"用LLM persona模拟专业预测者，复现人类预测行为，并与真实专家数据对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"基于角色的提示（persona-based prompting）能否提升大语言模型在宏观经济预测任务中的表现？","design":"使用GPT-4o模型，从PersonaHub语料库中筛选出2368个经济学相关角色，模拟欧洲央行专业预测者调查（ECB-SPF）的50个季度预测（2013-2025年），比较角色提示与无角色基线提示在四个宏观经济变量（HICP、核心HICP、GDP增长、失业率）和四个预测期上的预测准确性。","baseline":"欧洲央行专业预测者调查（ECB-SPF）的真实人类专家小组预测数据，以及100个无角色描述的基线GPT-4o预测。","findings":"GPT-4o与人类预测者的准确性非常接近，差异虽统计显著但实际幅度不大；角色描述对预测准确性没有可测量的提升，可省略以降低计算成本而不牺牲精度。","reliability":"样本外（2024-2025）预测表现与样本内存在明显差异；角色提示并未带来预测优势；不同提示产生的预测高度同质，与人类小组的多样性形成对比。","relevance":"该研究直接使用LLM角色模拟专业预测者，并与真实人类专家面板进行对照，评估了角色提示的效用与局限性，完全符合研究者对LLM仿真人类行为、基准对照和失效条件分析的兴趣，值得阅读原文。","inspiration":"该方法借鉴了用真实专家调查面板（ECB-SPF）作为人类基准，直接比较LLM角色提示与无角色基线的预测准确性，并检验样本外表现与预测多样性｜可迁移到政策公告的预期形成研究，例如模拟市场参与者对央行利率决议的即时反应与预期调整｜设计雏形：以LLM角色模拟金融分析师，处理为是否提供角色描述，结果变量为对利率决议的预测误差，对照真实分析师调查数据（如Blue Chip Economic Indicators）"}},{"id":"2512.08939","version":1,"title":"Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust","zh_title":"评估LLM驱动的数字孪生在模拟医疗系统信任中的人类相似性","abstract":"Serving as an emerging and powerful tool, Large Language Model (LLM)-driven Human Digital Twins are showing great potential in healthcare system research. However, its actual simulation ability for complex human psychological traits, such as distrust in the healthcare system, remains unclear. This research gap particularly impacts health professionals' trust and usage of LLM-based Artificial Intelligence (AI) systems in assisting their routine work. In this study, based on the Twin-2K-500 dataset, we systematically evaluated the simulation results of the LLM-driven human digital twin using the Health Care System Distrust Scale (HCSDS) with an established human-subject sample, analyzing item-level distributions, summary statistics, and demographic subgroup patterns. Results showed that the simulated responses by the digital twin were significantly more centralized with lower variance and had fewer selections of extreme options (all p<0.001). While the digital twin broadly reproduces human results in major demographic patterns, such as age and gender, it exhibits relatively low sensitivity in capturing minor differences in education levels. The LLM-based digital twin simulation has the potential to simulate population trends, but it also presents challenges in making detailed, specific distinctions in subgroups of human beings. This study suggests that the current LLM-driven Digital Twins have limitations in modeling complex human attitudes, which require careful calibration and validation before applying them in inferential analyses or policy simulations in health systems engineering. Future studies are necessary to examine the emotional reasoning mechanism of LLMs before their use, particularly for studies that involve simulations sensitive to social topics, such as human-automation trust.","authors":["Yuzhou Wu","Mingyang Wu","Di Liu","Rong Yin","Kang Li"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2025-10-27","first_seen":"2025-10-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.08939","pdf_url":"https://arxiv.org/pdf/2512.08939","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类数字孪生","医疗系统信任"],"reason":"用LLM数字孪生模拟医疗系统不信任，与真实人类数据对照，评估仿真可靠性并指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"LLM驱动的数字孪生在模拟医疗系统不信任等复杂心理特质时，其类人程度如何？","design":"基于Twin-2K-500数据集构建ChatGPT-4驱动的数字孪生，分层抽样500个样本，使用医疗系统不信任量表（HCSDS）生成模拟回答，并与真实人类样本进行对比。","baseline":"来自Rose等人研究的400名费城陪审员真实人类样本，使用相同的HCSDS量表测量。","findings":"数字孪生的回答显著集中于中间选项，方差更小，极端选项选择更少；虽能大致复现年龄、性别等主要人口学模式，但对教育水平等细微差异的捕捉敏感性较低。","reliability":"当前LLM数字孪生在模拟复杂人类态度时存在局限，回答分布过于集中，对子群体细微差异不敏感，在用于推断分析或政策模拟前需仔细校准和验证。","relevance":"该研究直接评估LLM仿真人类调查回答的可靠性，并与真实人类数据对照，揭示了仿真在分布形态和子群体差异上的失效模式，对关注LLM仿真偏差的研究者具有重要参考价值。","inspiration":"借鉴其分层抽样匹配人口特征、项目级分布对比和卡方检验的验证方法｜可迁移到消费者信任调查或政策态度评估，如模拟不同教育背景人群对金融监管机构的信任｜以LLM生成不同人口特征的虚拟消费者，施加不同政策信息处理，测量对金融系统的信任评分，并以真实消费者调查数据作为对照基准。"}},{"id":"2510.08338","version":3,"title":"LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings","zh_title":"大语言模型通过语义相似度引出李克特评分复现人类购买意向","abstract":"Consumer research costs companies billions annually yet suffers from panel biases and limited scale. Large language models (LLMs) offer an alternative by simulating synthetic consumers, but produce unrealistic response distributions when asked directly for numerical ratings. We present semantic similarity rating (SSR), a method that elicits textual responses from LLMs and maps these to Likert distributions using embedding similarity to reference statements. Testing on an extensive dataset comprising 57 personal care product surveys conducted by a leading corporation in that market (9,300 human responses), SSR achieves 90% of human test-retest reliability while maintaining realistic response distributions (KS similarity > 0.85). Additionally, these synthetic respondents provide rich qualitative feedback explaining their ratings. This framework enables scalable consumer research simulations while preserving traditional survey metrics and interpretability.","authors":["Benjamin F. Maier","Ulf Aslak","Luca Fiaschi","Nina Rismal","Kemble Fletcher","Christian C. Luhmann","Robbie Dow","Kli Pappas","Thomas V. Wiecki"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-10-09","first_seen":"2025-10-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.08338","pdf_url":"https://arxiv.org/pdf/2510.08338","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","消费者调查","人类数据对照"],"reason":"用LLM模拟消费者购买意向，与真实人类调查数据对照，评估分布可靠性与偏差，直接…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"如何通过语义相似度评分（SSR）方法让大语言模型生成更真实的人类购买意向分布？","design":"用大语言模型扮演给定人口统计属性的合成消费者，对57项个人护理产品概念生成自由文本购买意向，再通过嵌入向量与锚定语句的余弦相似度映射到5点李克特量表，测量购买意向分布和平均购买意向。","baseline":"来自一家领先企业的57项个人护理产品调查，共9,300名真实美国消费者的购买意向李克特评分。","findings":"SSR方法使合成消费者的购买意向分布与人类高度相似（KS相似度>0.85），且合成数据与人类数据在概念吸引力排名上的相关性达到人类重测信度的90%。同时，合成消费者能提供解释评分的定性反馈。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM模拟消费者购买意向，并与大规模真实人类调查数据严格对照，评估分布可靠性和偏差，直接命中研究者关心的仿真基准和经济学实验场景，值得精读原文。","inspiration":"借鉴其语义相似度评分方法，将LLM生成的自由文本通过嵌入向量与锚定语句的余弦相似度映射到结构化量表，可提升合成行为分布的真实性｜可迁移到消费者跨期选择实验，用LLM模拟不同人口统计特征下的时间偏好与折扣因子｜用LLM扮演合成消费者，施加未来不同时间点的金额选择任务，生成自由文本决策理由，通过SSR映射为选择概率，与真实跨期选择调查数据对照"}},{"id":"2510.06151","version":1,"title":"LLMs as Policy-Agnostic Teammates: A Case Study in Human Proxy Design for Heterogeneous Agent Teams","zh_title":"作为策略无关队友的大语言模型：异构智能体团队中人类代理设计的案例研究","abstract":"A critical challenge in modelling Heterogeneous-Agent Teams is training agents to collaborate with teammates whose policies are inaccessible or non-stationary, such as humans. Traditional approaches rely on expensive human-in-the-loop data, which limits scalability. We propose using Large Language Models (LLMs) as policy-agnostic human proxies to generate synthetic data that mimics human decision-making. To evaluate this, we conduct three experiments in a grid-world capture game inspired by Stag Hunt, a game theory paradigm that balances risk and reward. In Experiment 1, we compare decisions from 30 human participants and 2 expert judges with outputs from LLaMA 3.1 and Mixtral 8x22B models. LLMs, prompted with game-state observations and reward structures, align more closely with experts than participants, demonstrating consistency in applying underlying decision criteria. Experiment 2 modifies prompts to induce risk-sensitive strategies (e.g. \"be risk averse\"). LLM outputs mirror human participants' variability, shifting between risk-averse and risk-seeking behaviours. Finally, Experiment 3 tests LLMs in a dynamic grid-world where the LLM agents generate movement actions. LLMs produce trajectories resembling human participants' paths. While LLMs cannot yet fully replicate human adaptability, their prompt-guided diversity offers a scalable foundation for simulating policy-agnostic teammates.","authors":["Aju Ani Justus","Chris Baber"],"categories":["cs.LG","cs.AI","cs.HC"],"primary_category":"cs.LG","announce_type":"new","date":"2025-10-07","first_seen":"2025-10-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.06151","pdf_url":"https://arxiv.org/pdf/2510.06151","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","行为博弈","人机协作"],"reason":"用LLM代理人类决策，与真实人类数据对照，涉及博弈论实验，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":48,"question":"LLM能否作为策略无关的人类代理，在猎鹿博弈网格游戏中复现专家决策、模拟人类风险偏好变异并生成类人运动轨迹？","design":"使用LLaMA 3.1和Mixtral 8x22B模型，通过提示词输入网格世界的相对距离和奖励结构，模拟人类在猎鹿博弈中的决策。实验1比较LLM与30名人类参与者和2名专家裁判的选择；实验2通过修改提示词诱导风险规避或风险寻求策略；实验3让LLM在动态网格中生成移动动作序列。","baseline":"30名对博弈论了解有限的人类参与者在15种网格配置下的决策，以及2名博弈论专家裁判的选择。","findings":"LLM的决策与专家裁判高度一致，表现出对底层决策准则的一致性应用；通过提示词引导，LLM能展现出与人类参与者相似的风险偏好变异，在风险规避和风险寻求行为间切换。","reliability":"LLM尚不能完全复制人类的适应性，其行为多样性依赖于提示词引导，而非自主涌现。","relevance":"该研究直接用LLM代理人类决策，并与真实人类数据对照，涉及博弈论实验，方法可迁移至经济学和政策评估场景，对关注仿真可靠性与偏差的研究者具有参考价值，值得阅读原文。","inspiration":"该方法通过提示词工程直接操控LLM的风险偏好（风险规避/风险寻求），并设置专家裁判和人类参与者双基准对照，可借鉴其处理异质性行为变异的设计｜可迁移至资产定价实验中的风险态度测量，或政策公告对投资者预期形成的仿真研究｜以LLM作为被试，通过提示词注入不同风险偏好指令，测量其在模拟股票投资任务中的资产配置比例，并与真实投资者调查数据或实验数据对照"}},{"id":"2510.02343","version":1,"title":"$\\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training","zh_title":"BluePrint：用于LLM角色评估与训练的社交媒体用户数据集","abstract":"Large language models (LLMs) offer promising capabilities for simulating social media dynamics at scale, enabling studies that would be ethically or logistically challenging with human subjects. However, the field lacks standardized data resources for fine-tuning and evaluating LLMs as realistic social media agents. We address this gap by introducing SIMPACT, the SIMulation-oriented Persona and Action Capture Toolkit, a privacy respecting framework for constructing behaviorally-grounded social media datasets suitable for training agent models. We formulate next-action prediction as a task for training and evaluating LLM-based agents and introduce metrics at both the cluster and population levels to assess behavioral fidelity and stylistic realism. As a concrete implementation, we release BluePrint, a large-scale dataset built from public Bluesky data focused on political discourse. BluePrint clusters anonymized users into personas of aggregated behaviours, capturing authentic engagement patterns while safeguarding privacy through pseudonymization and removal of personally identifiable information. The dataset includes a sizable action set of 12 social media interaction types (likes, replies, reposts, etc.), each instance tied to the posting activity preceding it. This supports the development of agents that use context-dependence, not only in the language, but also in the interaction behaviours of social media to model social media users. By standardizing data and evaluation protocols, SIMPACT provides a foundation for advancing rigorous, ethically responsible social media simulations. BluePrint serves as both an evaluation benchmark for political discourse modeling and a template for building domain specific datasets to study challenges such as misinformation and polarization.","authors":["Aurélien Bück-Kaeffer","Je Qin Chooi","Dan Zhao","Maximilian Puelma Touzel","Kellin Pelrine","Jean-François Godbout","Reihaneh Rabbany","Zachary Yang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-09-27","first_seen":"2025-09-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.02343","pdf_url":"https://arxiv.org/pdf/2510.02343","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","社交媒体模拟","行为保真度"],"reason":"用LLM模拟社交媒体用户行为，有真实数据对照，涉及政治话语，方法可迁移至人类仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":44,"question":"如何构建隐私保护的社交媒体数据集，并用于训练和评估基于大语言模型的社交媒体用户行为模拟代理？","design":"提出SIMPACT框架，将用户行为建模为动作序列，通过聚类形成行为角色以保护隐私；基于Bluesky平台2025年加拿大联邦选举期间的政治话语数据构建BluePrint数据集，包含12种交互行为；将下一动作预测作为训练和评估任务，使用GPT-4.1-mini、o3-mini、Qwen-2.5-7B等模型进行基准测试，并采用余弦相似度、Jaccard相似度、JS散度、F1分数及人类评估等多指标衡量行为保真度。","baseline":"以BluePrint数据集中真实用户的帖子嵌入、关键词分布和实际交互行为作为对照基准。","findings":"当前LLM能生成看似合理的文本，但在复现真实用户社区的行为模式上存在困难；经过微调的模型在行为预测上有所提升，但仍难以完全捕捉人类社交互动的细微差别。","reliability":"论文指出计算指标只能近似评估，无法完全捕捉人类行为的微妙性，可能产生“恐怖谷”效应；数据集聚焦特定政治事件和平台，泛化性有待验证；隐私保护措施可能损失部分个体行为细节。","relevance":"该研究直接涉及用LLM模拟社交媒体用户行为，并提供真实人类数据对照和评估基准，符合研究者对仿真可靠性及政治话语场景的关注，值得阅读原文以了解其数据集构建方法和模型失效的具体表现。","inspiration":"借鉴其将用户行为建模为动作序列并用聚类形成行为角色以保护隐私的方法，可设计隐私合规的仿真实验｜可迁移到金融社交媒体信息传播与投资者情绪形成的场景，如研究政策推文对散户交易行为的影响｜以LLM代理模拟散户投资者，处理为不同情绪倾向的政策推文，结果变量为模拟的买卖行为序列，用真实交易数据或社交媒体互动数据作为对照基准"}},{"id":"2509.13397","version":4,"title":"The threat of analytic flexibility in using large language models to simulate human data","zh_title":"使用大语言模型模拟人类数据时分析灵活性的威胁","abstract":"Social scientists are now using large language models to create \"silicon samples\": synthetic datasets intended to stand in for human respondents. However, producing these samples requires many analytic choices, including model selection, sampling parameters, prompt format, and the amount of demographic or contextual information provided. Across two studies, I examine whether these choices materially affect correspondence between silicon samples and human data. In Study 1, I generated 252 silicon-sample configurations for a controlled case study using two social-psychological scales, evaluating whether configurations recovered participant rankings, response distributions, and between-scale correlations. Configurations varied substantially across all three criteria, and configurations that performed well on one dimension often performed poorly on another. In Study 2, I extended this analysis to a published silicon-sample use case by re-examining Argyle et al.'s (2023) Study 3 using 66 alternative configurations. Correlations between human and silicon association structures differed substantially across configurations, from r = .23 to r = .84. Taken together, the results from these studies demonstrate that different defensible configuration choices can materially alter conclusions about the fidelity of silicon samples. I call for greater attention to the threat of analytic flexibility in using silicon samples and outline strategies that researchers may adopt to reduce this threat.","authors":["Jamie Cummins"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-09-16","first_seen":"2025-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.13397","pdf_url":"https://arxiv.org/pdf/2509.13397","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["硅样本","分析灵活性","仿真保真度"],"reason":"直接研究用LLM生成硅样本模拟人类数据，评估分析灵活性对仿真保真度的影响，并与…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:28","error":null,"has_summary":true,"summary":{"generated_at":"2025-09-16","rank":7,"question":"分析灵活性（如模型选择、采样参数、提示格式等）是否显著影响硅样本与人类数据之间的一致性。","design":"研究1：使用两个社会心理量表，生成252种硅样本配置，评估配置能否恢复被试排名、响应分布和量表间相关性。研究2：重新分析Argyle等人(2023)的研究3，使用66种替代配置，比较人类与硅样本关联结构的相关性。","baseline":"研究1：来自两个社会心理量表的人类被试数据。研究2：Argyle等人(2023)研究3中的人类数据。","findings":"不同配置在恢复排名、响应分布和相关性上差异显著，且在一个维度上表现好的配置在另一维度上可能表现差。人类与硅样本关联结构的相关性从r=0.23到r=0.84不等，表明分析灵活性可实质改变关于硅样本保真度的结论。","reliability":"论文指出不同可辩护的配置选择会实质改变结论，但未明确列出所有失效条件或局限。","relevance":"直接命中研究者关注的LLM仿真人类数据可靠性问题，有真实人类对照，并批判性地揭示了分析灵活性威胁，值得精读原文以了解具体配置影响和应对策略。","inspiration":"借鉴其通过系统性地变化模型选择、采样参数和提示格式来构建多种硅样本配置，并评估配置间结果差异的方法，以揭示分析灵活性的威胁。｜可迁移到政策公告的预期形成实验，例如研究央行沟通措辞对通胀预期的影响。｜以LLM作为被试，随机分配不同措辞的政策公告作为处理，测量其预测的通胀数值，并与专业预测者调查或消费者预期调查的真实数据对照，同时变化提示中的角色设定、温度参数等配置，检验结论的稳健性。"}},{"id":"2509.11311","version":2,"title":"Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble","zh_title":"从提示到代理：通过紧凑LLM集成模拟人类偏好","abstract":"Large language models are increasingly used as proxies for human subjects in social science research, yet external validity requires that synthetic agents faithfully reflect the preferences of target human populations. We introduce *preference reconstruction theory*, a framework that formalizes preference alignment as a representation learning problem: constructing a functional basis of proxy agents and recovering population preferences through weighted aggregation. We implement this via *Prompts to Proxies* ($\\texttt{P2P}$), a modular two-stage system. Stage 1 uses structured prompting with entropy-based adaptive sampling to construct a diverse agent pool spanning the latent preference space. Stage 2 employs L1-regularized regression to select a compact ensemble whose aggregate response distributions align with observed data from the target population. $\\texttt{P2P}$ requires no finetuning and no access to sensitive demographic data, incurring only API inference costs. We validate the approach on 14 waves of the American Trends Panel, achieving an average test MSE of 0.014 across diverse topics at approximately 0.8 USD per survey. We additionally test it on the World Values Survey, demonstrating its potential to generalize across locales. When stress-tested against an SFT-aligned baseline, $\\texttt{P2P}$ achieves competitive performance using less than 3% of the training data.","authors":["Bingchen Wang","Zi-Yu Khoo","Jingtan Wang"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2025-09-14","first_seen":"2025-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.11311","pdf_url":"https://arxiv.org/pdf/2509.11311","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["LLM人类仿真","偏好重建","社会调查"],"reason":"用LLM代理复现人群偏好，有真实调查数据对照，涉及社会政策评估，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":45,"question":"如何通过紧凑的LLM集成，在无需微调和人口数据的情况下，忠实复现目标人群的偏好分布？","design":"使用GPT系列模型作为代理被试，通过两阶段系统P2P：第一阶段用结构化提示和基于熵的自适应采样构建多样化的代理池，第二阶段用L1正则化回归选择紧凑的代理集成，使其聚合响应分布与目标人群的调查数据对齐。结果变量为调查问题的回答分布。","baseline":"美国趋势面板（ATP）14波调查和世界价值观调查（WVS）的真实人类回答数据。","findings":"P2P在ATP上平均测试MSE为0.014，相比提示基线提升43%，每份调查成本约0.8美元；在WVS上展现出跨地域泛化潜力。与SFT对齐基线相比，P2P使用不到3%的训练数据即达到竞争性能。","reliability":"论文指出当前方法限于结构化调查问题，未处理自由文本输出；偏好重建依赖静态调查数据，可能无法应对非平稳偏好；未来需结合语义和心理测量技术提升鲁棒性。","relevance":"高度相关：该研究直接用LLM代理复现真实调查中的偏好分布，有严格的人类基准对照，涉及政策评估场景，并讨论了仿真失效条件，完全符合研究者的关注点，值得精读原文。","inspiration":"该方法通过两阶段系统（多样化代理池构建与L1正则化集成选择）实现无需微调、低成本的人类偏好复现，其自适应采样和稀疏集成思路值得借鉴｜可迁移到消费者金融决策偏好研究，如风险偏好、储蓄选择或投资组合配置的调查实验｜以LLM代理作为被试，施加不同金融信息框架处理，结果变量为风险资产选择比例，用美国消费者金融调查（SCF）的真实数据作为对照基准"}},{"id":"2509.09871","version":1,"title":"Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case","zh_title":"模拟民意：智利案例中AI生成合成调查回答的概念验证","abstract":"Large Language Models (LLMs) offer promising avenues for methodological and applied innovations in survey research by using synthetic respondents to emulate human answers and behaviour, potentially mitigating measurement and representation errors. However, the extent to which LLMs recover aggregate item distributions remains uncertain and downstream applications risk reproducing social stereotypes and biases inherited from training data. We evaluate the reliability of LLM-generated synthetic survey responses against ground-truth human responses from a Chilean public opinion probabilistic survey. Specifically, we benchmark 128 prompt-model-question triplets, generating 189,696 synthetic profiles, and pool performance metrics (i.e., accuracy, precision, recall, and F1-score) in a meta-analysis across 128 question-subsample pairs to test for biases along key sociodemographic dimensions. The evaluation spans OpenAI's GPT family and o-series reasoning models, as well as Llama and Qwen checkpoints. Three results stand out. First, synthetic responses achieve excellent performance on trust items (F1-score and accuracy > 0.90). Second, GPT-4o, GPT-4o-mini and Llama 4 Maverick perform comparably on this task. Third, synthetic-human alignment is highest among respondents aged 45-59. Overall, LLM-based synthetic samples approximate responses from a probabilistic sample, though with substantial item-level heterogeneity. Capturing the full nuance of public opinion remains challenging and requires careful calibration and additional distributional tests to ensure algorithmic fidelity and reduce errors.","authors":["Bastián González-Bustamante","Nando Verelst","Carla Cisternas"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-09-11","first_seen":"2025-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.09871","pdf_url":"https://arxiv.org/pdf/2509.09871","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"直接用LLM生成合成调查回答，与真实人类概率样本对照，评估可靠性与偏差，涉及公…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:27","error":null,"has_summary":true,"summary":{"generated_at":"2025-09-11","rank":11,"question":"大语言模型生成的合成调查回复能否准确反映智利的真实公众意见？","design":"使用128个提示-模型-问题三元组，生成189,696个合成档案，评估GPT系列、o系列推理模型、Llama和Qwen等模型在模拟智利公众意见调查回复上的表现，以准确率、精确率、召回率和F1分数为指标，并检验关键社会人口维度的偏差。","baseline":"智利公众意见概率调查的真实人类回复。","findings":"合成回复在信任项目上表现优异（F1和准确率>0.90）；GPT-4o、GPT-4o-mini和Llama 4 Maverick表现相当，且45-59岁年龄段的合成-人类对齐度最高。","reliability":"论文指出合成样本与概率样本的近似存在显著的项目层面异质性，捕捉公众意见的全部细微差别仍有挑战，需要仔细校准和额外的分布检验。","relevance":"该研究直接评估了LLM作为人类被试替代品的可靠性，有真实人类数据对照，并关注偏差，完全符合研究者的兴趣，值得精读。","inspiration":"该研究通过多模型、多提示的系统比较和关键社会人口维度的偏差检验来评估仿真可靠性，方法上值得借鉴｜可迁移到消费者信心指数或通胀预期调查的仿真上，检验LLM能否复现真实公众的经济预期分布｜以GPT-4o等为被试，用真实消费者调查问卷作为提示，生成通胀预期回复，以央行或统计局的真实调查数据为基准，比较分布一致性和子群体偏差"}},{"id":"2509.06337","version":2,"title":"Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation","zh_title":"大语言模型作为虚拟调查受访者：评估社会人口响应生成","abstract":"Questionnaire-based surveys are foundational to social science research and public policymaking, yet traditional survey methods remain costly, time-consuming, and often limited in scale. Although prior work has explored large language models (LLMs) as virtual survey respondents, existing studies often address narrow task settings, focus on single sociological domains, or lack a unified evaluation framework that enables systematic comparison across diverse datasets and models. To address these gaps, we introduce two complementary task abstractions: Partial Attribute Simulation (PAS), where LLMs predict missing attributes from incomplete respondent profiles, and Full Attribute Simulation (FAS), where LLMs generate complete synthetic datasets under zero-context and context-enhanced conditions. Both are framed as diagnostic and exploratory tools rather than replacements for human data collection. We curate LLM-S^3 (Large Language Model-based Sociodemographic Survey Simulation), a benchmark spanning 11 real-world public datasets across four sociological domains, and evaluate GPT-3.5/4 Turbo and LLaMA 3.0/3.1-8B under zero-shot and few-shot settings. Our evaluation reveals consistent performance trends across model families, highlights failure modes in structured output generation, and demonstrates how context and prompt design affect simulation fidelity. Our code and dataset are available at: https://github.com/dart-lab-research/LLM-S-Cube-Benchmark","authors":["Jianpeng Zhao","Chenyu Yuan","Weiming Luo","Haoling Xie","Guangwei Zhang","Steven Jige Quan","Zixuan Yuan","Pengyang Wang","Denghui Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-09-08","first_seen":"2025-09-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.06337","pdf_url":"https://arxiv.org/pdf/2509.06337","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","调查方法","社会人口模拟"],"reason":"用LLM模拟调查受访者，复现社会人口属性，有真实人类数据对照，涉及社会科学与政…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":35,"question":"如何系统评估大语言模型在模拟社会人口调查回答时的表现与可靠性？","design":"提出部分属性模拟（PAS）和全属性模拟（FAS）两种任务，使用GPT-3.5/4 Turbo和LLaMA 3.0/3.1-8B模型，在11个真实调查数据集上，以零样本和少样本设置生成缺失属性或完整合成数据集，并测量统计分布相似度。","baseline":"11个来自社会与公共事务、工作与收入、家庭与行为模式、健康与生活方式四个领域的真实公共调查数据集。","findings":"不同模型家族在模拟任务上表现出一致的性能趋势；提示设计和上下文增强显著影响模拟保真度，结构化输出生成中的失败案例仍是主要瓶颈。","reliability":"论文明确声明评估仅衡量统计和分布相似性，而非行为保真度或构念效度；数据集主要来自北美和欧洲，跨文化泛化性未验证；全属性模拟场景下结构化输出失败问题突出。","relevance":"该研究直接以真实人类调查数据为基准，系统评估LLM模拟社会人口属性的可靠性，并指出失效条件，高度契合研究者对仿真基准、政策评估场景和批判性分析的兴趣，值得精读原文。","inspiration":"借鉴其部分属性模拟（PAS）和全属性模拟（FAS）的任务设计，以真实调查数据为基准，通过统计分布相似度量化LLM的仿真保真度，并系统对比不同模型、提示策略和上下文增强的影响。｜可迁移到信贷审批中的歧视测量场景，用LLM模拟不同人口特征申请人的信用评分或审批结果，检验算法或人工决策中的统计性歧视。｜以真实信贷申请数据（如HMDA）为对照基准，将申请人的人口属性（种族、性别）作为处理变量，让LLM在PAS任务下生成信用评分或审批决策，比较LLM生成分布与真实审批分布的差异，并分析提示中是否加入反歧视法规对仿真偏差的影响。"}},{"id":"2509.03736","version":2,"title":"Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation","zh_title":"LLM代理行为一致吗？社会模拟的潜在画像","abstract":"The impressive capabilities of Large Language Models (LLMs) raise the possibility that synthetic agents can serve as substitutes for real participants in human-subject research. To evaluate this claim, prior research has largely focused on whether LLM-generated survey responses align with those produced by human respondents whom the LLMs are prompted to represent. In contrast, we address a more fundamental question: Do agents maintain empirical consistency; aligning to human behavioral models when examined under different experimental settings? To this end, we develop a study designed to (a) ask a set of questions which reveals an agent's latent profile and (b) examine agent behavioral consistency in a conversational setting with other agents. This design enables us to explore a set of behavioral hypotheses to assess whether an agent's conversational behavior is consistent with what we would expect from its revealed state. Our findings show significant inconsistencies in LLMs across model families and at differing model sizes. Most importantly, we find that, although agents may generate responses matching those of their human counterparts, they fail to be empirically consistent, representing a critical gap in their capabilities to accurately substitute for real participants in human-subject research.","authors":["James Mooney","Josef Woldense","Zheng Robert Jia","Shirley Anugrah Hayati","My Ha Nguyen","Vipul Raheja","Dongyeop Kang"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2025-09-03","first_seen":"2025-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.03736","pdf_url":"https://arxiv.org/pdf/2509.03736","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","行为一致性","人类被试替代"],"reason":"直接评估LLM代理替代人类被试的行为一致性，有人类数据对照，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":46,"question":"LLM代理在跨实验设置下是否保持行为一致性，即其对话行为是否与其自身揭示的潜在状态（偏好和开放性）相符？","design":"构建一个五阶段框架：选择争议性话题，生成具有人口统计特征和话题偏见的代理，通过问卷获取代理的潜在状态（偏好和开放性），将代理配对进行多轮对话，评估对话中的一致性。通过系统变化系统提示引入控制变量，测试六种人类行为模型。","baseline":"无对照","findings":"LLM代理在聚合层面表现出一些趋势，但无法通过更严格的经验一致性检验：即使偏好对立，代理也很少维持分歧；偏见提示不能可靠恢复原则性分歧；共享负面情绪产生的对齐弱于共享正面情绪；开放性在最应起作用的场景中失去预测力。","reliability":"论文指出当前LLM代理在行为模型更复杂或更细粒度时无法复制人类行为，其行为一致性存在显著差距，但未讨论具体失效条件。","relevance":"该研究直接评估LLM代理替代人类被试的行为一致性，揭示其在复杂行为模型中失效，与研究者关注的仿真可靠性及失效条件高度相关，值得精读原文。","inspiration":"借鉴其通过系统提示注入不同行为模型（如偏见、开放性）并检验代理行为一致性的实验设计，可作为一种施加处理的方式。｜可迁移至政策公告的预期形成实验，研究不同信息框架下投资者对央行沟通的反应一致性。｜以LLM代理为被试，处理为系统提示中嵌入不同的政策沟通风格（鹰派/鸽派），结果变量为代理在模拟交易中的资产配置变化，对照真实央行公告后的市场调查数据。"}},{"id":"2510.07321","version":1,"title":"How human is the machine? Evidence from 66,000 Conversations with Large Language Models","zh_title":"机器有多像人？来自66000次与大语言模型对话的证据","abstract":"When Artificial Intelligence (AI) is used to replace consumers (e.g., synthetic data), it is often assumed that AI emulates established consumers, and more generally human behaviors. Ten experiments with Large Language Models (LLMs) investigate if this is true in the domain of well-documented biases and heuristics. Across studies we observe four distinct types of deviations from human-like behavior. First, in some cases, LLMs reduce or correct biases observed in humans. Second, in other cases, LLMs amplify these same biases. Third, and perhaps most intriguingly, LLMs sometimes exhibit biases opposite to those found in humans. Fourth, LLMs' responses to the same (or similar) prompts tend to be inconsistent (a) within the same model after a time delay, (b) across models, and (c) among independent research studies. Such inconsistencies can be uncharacteristic of humans and suggest that, at least at one point, LLMs' responses differed from humans. Overall, unhuman-like responses are problematic when LLMs are used to mimic or predict consumer behavior. These findings complement research on synthetic consumer data by showing that sources of bias are not necessarily human-centric. They also contribute to the debate about the tasks for which consumers, and more generally humans, can be replaced by AI.","authors":["Antonios Stamatogiannakis","Arsham Ghodsinia","Sepehr Etminanrad","Dilney Gonçalves","David Santos"],"categories":["cs.HC","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2025-08-31","first_seen":"2025-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.07321","pdf_url":"https://arxiv.org/pdf/2510.07321","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知偏差","人类数据对照"],"reason":"用LLM复现人类认知偏差并与真实人类数据对照，评估仿真可靠性并指出失效条件","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":23,"question":"大语言模型在经典决策偏差与启发式任务中，其行为在多大程度上与人类相似？","design":"使用多个大语言模型（如GPT系列）进行10项预注册实验，共66,000次对话，系统操纵模型类型和提示语特征，测量模型在可得性启发、代表性启发、禀赋效应、锚定效应、交易效用和框架效应六种偏差上的反应。","baseline":"以心理学和行为经济学文献中已确立的人类在这些偏差上的典型行为模式作为对照基准。","findings":"LLM表现出四种与人类不同的偏差模式：减弱或纠正人类偏差、放大人类偏差、表现出与人类相反的偏差，以及在不同时间、模型和研究间反应不一致。这些非人类反应表明LLM在模拟或预测消费者行为时存在问题。","reliability":"论文指出LLM的反应会随时间、模型版本和不同研究间出现不一致，这种不一致在人类中不典型，提示在某一时点LLM的反应与人类不同，从而限制了其作为人类替代品的可靠性。","relevance":"该研究直接评估LLM作为人类被试替代品的可靠性，通过经典决策偏差实验与真实人类行为基准对照，系统揭示了仿真失效的四种模式，对关注LLM仿真效度的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其多模型、多偏差、大样本的系统比较设计，可迁移到经济金融中的投资者行为偏差研究（如处置效应、过度自信），设计雏形：以GPT-4等LLM为被试，呈现模拟股票交易场景测量处置效应，结果与真实投资者交易数据（如券商账户记录）进行对照。"}},{"id":"2510.06222","version":1,"title":"Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making","zh_title":"在LLM智能体中诱导状态焦虑可复现消费者决策中的人类偏差","abstract":"Large language models (LLMs) are rapidly evolving from text generators to autonomous agents, raising urgent questions about their reliability in real-world contexts. Stress and anxiety are well known to bias human decision-making, particularly in consumer choices. Here, we tested whether LLM agents exhibit analogous vulnerabilities. Three advanced models (ChatGPT-5, Gemini 2.5, Claude 3.5-Sonnet) performed a grocery shopping task under budget constraints (24, 54, 108 USD), before and after exposure to anxiety-inducing traumatic narratives. Across 2,250 runs, traumatic prompts consistently reduced the nutritional quality of shopping baskets (Change in Basket Health Scores of -0.081 to -0.126; all pFDR<0.001; Cohens d=-1.07 to -2.05), robust across models and budgets. These results show that psychological context can systematically alter not only what LLMs generate but also the actions they perform. By reproducing human-like emotional biases in consumer behavior, LLM agents reveal a new class of vulnerabilities with implications for digital health, consumer safety, and ethical AI deployment.","authors":["Ziv Ben-Zion","Zohar Elyoseph","Tobias Spiller","Teddy Lazebnik"],"categories":["cs.HC","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2025-08-30","first_seen":"2025-08-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.06222","pdf_url":"https://arxiv.org/pdf/2510.06222","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","消费者行为","焦虑偏差"],"reason":"用LLM模拟消费者决策，诱导焦虑后复现人类偏差，有真实人类数据对照，涉及行为经…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":24,"question":"心理上下文（如焦虑诱导）是否会系统性地改变LLM智能体在消费者决策中的行为，使其表现出类似人类的情绪偏差？","design":"使用ChatGPT-5、Gemini 2.5、Claude 3.5-Sonnet三个先进LLM作为智能体，模拟人类消费者在预算约束（27、54、108美元）下进行杂货购物任务；通过暴露于焦虑诱导的创伤叙事作为处理，测量购物篮健康评分的变化。","baseline":"无对照","findings":"焦虑诱导提示一致降低了购物篮的营养质量（健康评分变化Δ=-0.081至-0.126，所有pFDR<0.001，Cohen's d=-1.07至-2.05），效应在不同模型和预算水平下均稳健。结果表明心理上下文不仅能改变LLM生成的文本，还能系统性地改变其作为智能体所执行的动作。","reliability":"论文未讨论","relevance":"该研究直接以LLM模拟人类消费者决策，通过情绪诱导复现人类偏差，并评估效应稳健性，高度契合研究者对LLM仿真可靠性及偏差的关注，值得精读原文以了解其方法细节和潜在失效边界。","inspiration":"借鉴其通过情绪化提示（创伤叙事）施加心理状态处理、以预算约束任务测量经济决策结果、并跨模型和预算水平进行稳健性检验的设计方法｜可迁移至消费者跨期选择或风险决策场景，如焦虑对储蓄/借贷行为或保险购买决策的影响｜以LLM智能体为被试，施加焦虑诱导处理，测量其在跨期选择任务中的贴现率或风险资产配置比例，并以真实人类实验数据（如实验室或调查数据）作为对照基准。"}},{"id":"2509.02605","version":1,"title":"Synthetic Founders: AI-Generated Social Simulations for Startup Validation Research in Computational Social Science","zh_title":"合成创始人：用于计算社会科学中创业验证研究的AI生成社会仿真","abstract":"We present a comparative docking experiment that aligns human-subject interview data with large language model (LLM)-driven synthetic personas to evaluate fidelity, divergence, and blind spots in AI-enabled simulation. Fifteen early-stage startup founders were interviewed about their hopes and concerns regarding AI-powered validation, and the same protocol was replicated with AI-generated founder and investor personas. A structured thematic synthesis revealed four categories of outcomes: (1) Convergent themes - commitment-based demand signals, black-box trust barriers, and efficiency gains were consistently emphasized across both datasets; (2) Partial overlaps - founders worried about outliers being averaged away and the stress of real customer validation, while synthetic personas highlighted irrational blind spots and framed AI as a psychological buffer; (3) Human-only themes - relational and advocacy value from early customer engagement and skepticism toward moonshot markets; and (4) Synthetic-only themes - amplified false positives and trauma blind spots, where AI may overstate adoption potential by missing negative historical experiences. We interpret this comparative framework as evidence that LLM-driven personas constitute a form of hybrid social simulation: more linguistically expressive and adaptable than traditional rule-based agents, yet bounded by the absence of lived history and relational consequence. Rather than replacing empirical studies, we argue they function as a complementary simulation category - capable of extending hypothesis space, accelerating exploratory validation, and clarifying the boundaries of cognitive realism in computational social science.","authors":["Jorn K. Teutloff"],"categories":["cs.MA","cs.AI","cs.CY"],"primary_category":"cs.MA","announce_type":"new","date":"2025-08-29","first_seen":"2025-08-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.02605","pdf_url":"https://arxiv.org/pdf/2509.02605","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","创业者访谈对照","仿真保真度评估"],"reason":"用LLM生成合成创业者进行访谈仿真，并与15位真人创业者数据对照，评估仿真保真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":47,"question":"LLM驱动的合成创业者与投资人角色在创业验证访谈中，与真实人类受访者的主题模式在多大程度上一致、偏离或存在盲点？","design":"采用对比对接实验：用LLM生成35个合成创业者与投资人角色，复制对15位真实早期创业者的半结构化访谈协议，询问其对AI驱动市场验证的希望与担忧，通过结构化主题综合比较合成与人类访谈文本的涌现主题。","baseline":"15位真实早期创业者的访谈数据，涵盖其对AI驱动市场验证的希望与担忧。","findings":"合成与人类数据在承诺型需求信号、黑箱信任障碍和效率提升上主题收敛；部分重叠中，人类担忧异常值被平均化和真实客户验证压力，合成角色则强调非理性盲点并将AI视为心理缓冲；人类独有主题包括早期客户参与的倡导价值和对极端市场的怀疑，合成独有主题为放大假阳性和创伤盲点。","reliability":"论文指出LLM角色缺乏真实生活经历和关系后果，可能高估采纳潜力并遗漏负面历史经验，因此不能替代实证研究，仅作为补充性仿真工具，用于扩展假设空间和加速探索性验证。","relevance":"该研究直接以真实人类访谈为基准，系统评估LLM仿真在创业决策中的保真度与偏差，并明确区分收敛、部分重叠和独有主题，对关注仿真可靠性及失效条件的研究者具有重要参考价值，值得细读原文。","inspiration":"该方法通过对比合成角色与真实人类在相同半结构化访谈下的主题收敛与偏离，系统评估LLM仿真的保真度与盲点，值得借鉴｜可迁移至创业融资决策研究，如投资者对AI辅助商业计划书的评估偏差｜以LLM生成的投资人角色为被试，呈现AI生成与人类撰写的商业计划书，测量投资意愿与风险评估，并以真实天使投资人的评审数据作为对照基准"}},{"id":"2508.20234","version":1,"title":"Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research","zh_title":"验证基于生成式智能体的物流与供应链管理研究模型","abstract":"Generative Agent-Based Models (GABMs) powered by large language models (LLMs) offer promising potential for empirical logistics and supply chain management (LSCM) research by enabling realistic simulation of complex human behaviors. Unlike traditional agent-based models, GABMs generate human-like responses through natural language reasoning, which creates potential for new perspectives on emergent LSCM phenomena. However, the validity of LLMs as proxies for human behavior in LSCM simulations is unknown. This study evaluates LLM equivalence of human behavior through a controlled experiment examining dyadic customer-worker engagements in food delivery scenarios. I test six state-of-the-art LLMs against 957 human participants (477 dyads) using a moderated mediation design. This study reveals a need to validate GABMs on two levels: (1) human equivalence testing, and (2) decision process validation. Results reveal GABMs can effectively simulate human behaviors in LSCM; however, an equivalence-versus-process paradox emerges. While a series of Two One-Sided Tests (TOST) for equivalence reveals some LLMs demonstrate surface-level equivalence to humans, structural equation modeling (SEM) reveals artificial decision processes not present in human participants for some LLMs. These findings show GABMs as a potentially viable methodological instrument in LSCM with proper validation checks. The dual-validation framework also provides LSCM researchers with a guide to rigorous GABM development. For practitioners, this study offers evidence-based assessment for LLM selection for operational tasks.","authors":["Vincent E. Castillo"],"categories":["cs.MA","cs.AI","cs.CY"],"primary_category":"cs.MA","announce_type":"new","date":"2025-08-27","first_seen":"2025-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.20234","pdf_url":"https://arxiv.org/pdf/2508.20234","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM人类仿真","等效性验证","供应链管理"],"reason":"直接验证LLM作为人类代理在供应链场景中的等效性，有957名人类对照，并揭示仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"在物流与供应链管理（LSCM）的二元互动场景中，基于大语言模型的生成式智能体模型能否在表面行为上与人类等效，且其决策过程是否与人类可比？","design":"采用受控实验，将六种最先进的大语言模型作为生成式智能体，模拟外卖配送场景中顾客与骑手的二元互动，与957名人类参与者（477对）进行对比，使用有调节的中介设计，测量互动结果与决策过程。","baseline":"957名人类参与者（477对二元组）在相同外卖配送场景实验中的真实行为数据。","findings":"部分LLM在表面行为上通过等效性检验（TOST）与人类无显著差异，但结构方程模型（SEM）揭示其决策过程存在人为模式，与人类真实决策过程不同，形成“等效-过程悖论”。","reliability":"论文指出LLM作为人类代理的有效性未知，需进行双重验证（人类等效性检验和决策过程验证），并承认某些LLM的决策过程与人类不符，但未详细讨论其他失效条件。","relevance":"该研究直接验证LLM在供应链场景中替代人类被试的等效性，提供大规模人类对照基准，并批判性揭示表面等效下的决策过程差异，高度契合研究者对仿真可靠性、偏差及失效条件的关注，值得精读原文。","inspiration":"该研究采用TOST等效性检验与结构方程模型（SEM）双重验证，区分表面行为等效与决策过程等效，为仿真可靠性评估提供了严谨框架｜可迁移至消费者跨期选择实验，检验LLM生成的折现行为是否与人类一致｜以LLM为被试，施加不同跨期奖励方案，结果变量为选择时间偏好，以真实人类实验数据（如Andersen et al., 2008）为基准，进行TOST和SEM双重验证"}},{"id":"2508.19004","version":1,"title":"AI Models Exceed Individual Human Accuracy in Predicting Everyday Social Norms","zh_title":"AI模型在预测日常社会规范方面超越个体人类准确性","abstract":"A fundamental question in cognitive science concerns how social norms are acquired and represented. While humans typically learn norms through embodied social experience, we investigated whether large language models can achieve sophisticated norm understanding through statistical learning alone. Across two studies, we systematically evaluated multiple AI systems' ability to predict human social appropriateness judgments for 555 everyday scenarios by examining how closely they predicted the average judgment compared to each human participant. In Study 1, GPT-4.5's accuracy in predicting the collective judgment on a continuous scale exceeded that of every human participant (100th percentile). Study 2 replicated this, with Gemini 2.5 Pro outperforming 98.7% of humans, GPT-5 97.8%, and Claude Sonnet 4 96.0%. Despite this predictive power, all models showed systematic, correlated errors. These findings demonstrate that sophisticated models of social cognition can emerge from statistical learning over linguistic data alone, challenging strong versions of theories emphasizing the exclusive necessity of embodied experience for cultural competence. The systematic nature of AI limitations across different architectures indicates potential boundaries of pattern-based social understanding, while the models' ability to outperform nearly all individual humans in this predictive task suggests that language serves as a remarkably rich repository for cultural knowledge transmission.","authors":["Pontus Strimling","Simon Karlsson","Irina Vartanova","Kimmo Eriksson"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-08-26","first_seen":"2025-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.19004","pdf_url":"https://arxiv.org/pdf/2508.19004","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["社会规范预测","人类数据对照","算法偏差"],"reason":"用LLM预测人类社会规范判断，与真实人类数据对照，评估预测准确性与系统性偏差，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":30,"question":"大语言模型能否仅通过文本统计学习，在没有具身体验的情况下，达到甚至超越个体人类对日常社会规范（行为适当性）的判断水平？","design":"本研究并非严格意义上的仿真实验，而是利用已有的大规模人类判断数据集，将多个大语言模型（GPT-4.5、Gemini 2.5 Pro、GPT-5、Claude Sonnet 4）作为“被试”，要求它们对555个日常场景中的行为适当性进行连续评分，并与真实人类个体评分进行对比。","baseline":"真实人类数据：来自大规模数据集中个体参与者对相同555个日常场景的社会适当性连续评分，包括平均集体判断和每个个体的判断分布。","findings":"GPT-4.5 预测集体判断的准确度超过了所有人类参与者（百分位100%），其他模型也超过了96%以上的人类个体。所有模型均表现出系统性、相互关联的误差，表明基于模式的统计学习在社会理解上存在边界。","reliability":"论文指出，尽管模型预测力强，但所有模型都存在系统性且相互关联的误差，这表明基于模式的社会理解存在潜在边界；研究仅限于日常社会规范判断，未涉及其他类型的社会认知任务。","relevance":"该研究直接以真实人类个体判断为基准，评估LLM在连续社会规范判断上的仿真准确性及系统性偏差，高度契合研究者对LLM仿真可靠性、偏差及失效条件的关注，值得精读原文以了解其个体级对比方法和误差分析。","inspiration":"可借鉴其将模型预测与人类个体分布而非仅与均值对比的评估设计，以揭示仿真在个体差异捕捉上的能力。｜可迁移至消费者对金融产品适当性的感知或政策可接受性判断等场景。｜以LLM作为“消费者”被试，输入不同金融产品描述，要求其评估产品适当性，结果变量为适当性连续评分，对照真实消费者调查数据中的个体评分分布。"}},{"id":"2508.16172","version":2,"title":"Graph RAG as Human Choice Model: Building a Data-Driven Mobility Agent with Preference Chain","zh_title":"图RAG作为人类选择模型：构建数据驱动的出行智能体与偏好链","abstract":"Understanding human behavior in urban environments is a crucial field within city sciences. However, collecting accurate behavioral data, particularly in newly developed areas, poses significant challenges. Recent advances in generative agents, powered by Large Language Models (LLMs), have shown promise in simulating human behaviors without relying on extensive datasets. Nevertheless, these methods often struggle with generating consistent, context-sensitive, and realistic behavioral outputs. To address these limitations, this paper introduces the Preference Chain, a novel method that integrates Graph Retrieval-Augmented Generation (RAG) with LLMs to enhance context-aware simulation of human behavior in transportation systems. Experiments conducted on the Replica dataset demonstrate that the Preference Chain outperforms standard LLM in aligning with real-world transportation mode choices. The development of the Mobility Agent highlights potential applications of proposed method in urban mobility modeling for emerging cities, personalized travel behavior analysis, and dynamic traffic forecasting. Despite limitations such as slow inference and the risk of hallucination, the method offers a promising framework for simulating complex human behavior in data-scarce environments, where traditional data-driven models struggle due to limited data availability.","authors":["Kai Hu","Parfait Atchade-Adelomou","Carlo Adornetto","Adrian Mora-Carrero","Luis Alonso-Pastor","Ariel Noyman","Yubo Liu","Kent Larson"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-08-22","first_seen":"2025-08-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.16172","pdf_url":"https://arxiv.org/pdf/2508.16172","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","出行行为","人类数据对照"],"reason":"用LLM仿真交通出行选择，有真实人类数据对照，属经济学实验场景，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":40,"question":"如何在数据稀缺环境下，利用图检索增强生成（Graph RAG）与LLM结合的方法，更真实地模拟个体交通出行方式选择行为？","design":"提出Preference Chain方法，结合Graph RAG与LLM构建Mobility Agent，基于少量数据构建个体行为偏好图，通过相似性搜索和概率建模引导LLM生成出行选择；在Replica数据集上模拟交通方式选择，与真实选择对比。","baseline":"Replica数据集中真实的交通方式选择行为。","findings":"Preference Chain在模拟交通方式选择上比标准LLM更符合真实世界数据；该方法在数据稀缺地区具有应用潜力，但存在推理速度慢和幻觉风险。","reliability":"论文承认推理速度慢和存在幻觉风险，可能影响行为仿真的可靠性和实用性。","relevance":"该研究用LLM仿真交通出行选择，有真实人类数据对照，属于经济学实验场景，方法可迁移至其他行为仿真，值得阅读原文以评估其仿真偏差与可靠性。","inspiration":"借鉴Preference Chain方法，利用图检索增强生成（Graph RAG）从少量个体行为数据中构建偏好图，并通过相似性搜索和概率建模引导LLM生成选择，以提升仿真真实性。｜该方法可迁移至消费者跨期选择研究，模拟个体在不同时间偏好下的储蓄或消费决策。｜以LLM作为被试，构建基于少量真实个体跨期选择数据的偏好图，施加不同利率或未来收入预期的处理，结果变量为模拟的消费-储蓄分配，与真实家庭金融调查数据（如PSID）进行对照。"}},{"id":"2508.15926","version":1,"title":"Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making","zh_title":"噪声、适应与策略：评估LLM在决策中的保真度","abstract":"Large language models (LLMs) are increasingly used in social science simulations. While their performance on reasoning and optimization tasks has been extensively evaluated, less attention has been paid to their ability to simulate human decision-making's variability and adaptability. We propose a process-oriented evaluation framework with progressive interventions (Intrinsicality, Instruction, and Imitation) to examine how LLM agents adapt under different levels of external guidance and human-derived noise. We validate the framework on two classic economics tasks, irrationality in the second-price auction and decision bias in the newsvendor problem, showing behavioral gaps between LLMs and humans. We find that LLMs, by default, converge on stable and conservative strategies that diverge from observed human behaviors. Risk-framed instructions impact LLM behavior predictably but do not replicate human-like diversity. Incorporating human data through in-context learning narrows the gap but fails to reach human subjects' strategic variability. These results highlight a persistent alignment gap in behavioral fidelity and suggest that future LLM evaluations should consider more process-level realism. We present a process-oriented approach for assessing LLMs in dynamic decision-making tasks, offering guidance for their application in synthetic data for social science research.","authors":["Yuanjun Feng","Vivek Choudhary","Yash Raj Shrestha"],"categories":["cs.CE","cs.AI"],"primary_category":"cs.CE","announce_type":"new","date":"2025-08-21","first_seen":"2025-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.15926","pdf_url":"https://arxiv.org/pdf/2508.15926","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为保真度","经济学实验"],"reason":"直接评估LLM模拟人类决策的保真度，用真实人类数据对照，涉及经济学任务，指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":48,"question":"LLM在动态决策任务中能否复现人类决策的变异性和适应性？","design":"用GPT-4o、Claude 3.5 Sonnet、Claude 3.7 Sonnet扮演卖家或报童，在二阶密封拍卖和报童问题中，通过无干预、风险框架指令、人类决策历史模仿三种渐进干预，测量其策略稳定性、行为变异性和适应性。","baseline":"对照二阶密封拍卖的真实人类实验数据（Davis et al., 2011, 2023），以及报童问题的已知人类决策偏差模式。","findings":"LLM默认采取稳定保守的策略，与人类行为偏离；风险框架指令可预测地影响LLM行为，但未复现人类多样性；通过上下文学习注入人类数据可缩小差距，但仍未达到人类被试的策略变异性。","reliability":"论文指出当前LLM在动态行为模拟中存在行为保真度对齐差距，缺乏人类决策的随机性和适应性，未来评估需关注过程层面的真实性。","relevance":"该研究直接评估LLM模拟人类决策的保真度，使用真实人类数据作为基准，涉及经济学实验任务，并明确指出了仿真失效的条件和局限，与研究者关注点高度吻合，值得精读原文。","inspiration":"该研究通过渐进式干预（无干预、风险框架指令、人类历史模仿）系统测量LLM行为变异性的方法值得借鉴，可迁移到资产定价实验中的泡沫形成与处置效应研究｜可设计让LLM扮演投资者在实验性资产市场中交易，处理为不同风险提示框架或注入真实人类交易历史，结果变量为价格偏离度、交易量与处置效应系数，对照真实人类实验数据（如Smith et al., 1988的泡沫实验）"}},{"id":"2508.12045","version":2,"title":"Large Language Models Enable Design of Personalized Nudges across Cultures","zh_title":"大语言模型助力跨文化个性化助推设计","abstract":"Nudge strategies are effective tools for influencing behaviour, but their impact depends on individual preferences. Strategies that work for some individuals may be counterproductive for others. We hypothesize that large language models (LLMs) can facilitate the design of individual-specific nudges without the need for costly and time-intensive behavioural data collection and modelling. To test this, we use LLMs to design personalized decoy-based nudges tailored to individual profiles and cultural contexts, aimed at encouraging air travellers to voluntarily offset CO$_2$ emissions from flights. We evaluate their effectiveness through a large-scale survey experiment ($n=3495$) conducted across five countries. Results show that LLM-informed personalized nudges are more effective than uniform settings, raising offsetting rates by 3-7$\\%$ in Germany, Singapore, and the US, though not in China or India. Our study highlights the potential of LLM as a low-cost testbed for piloting nudge strategies. At the same time, cultural heterogeneity constrains their generalizability underscoring the need for combining LLM-based simulations with targeted empirical validation.","authors":["Vladimir Maksimenko","Qingyao Xin","Prateek Gupta","Bin Zhang","Prateek Bansal"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-16","first_seen":"2025-08-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.12045","pdf_url":"https://arxiv.org/pdf/2508.12045","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为助推","跨文化实验"],"reason":"用LLM设计个性化助推并做大规模人类调查对照，直接仿真人类决策行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":25,"question":"大语言模型能否用于设计跨文化的个性化助推策略，以低成本替代人类被试进行助推效果评估？","design":"使用LLM根据个体人口统计特征（性别、年龄、收入、环境关心度、碳抵消信任度）和文化背景，为航空旅客生成个性化的诱饵选项（价格与碳抵消比例），通过大规模在线调查实验（n=3495）在五个国家测试其对自愿碳抵消选择的影响。","baseline":"在五个国家（中国、德国、印度、新加坡、美国）对3495名真实航空旅客进行的调查实验，测量其面对个性化诱饵与统一诱饵时的碳抵消选择率。","findings":"LLM生成的个性化助推在德国、新加坡和美国将碳抵消率提高了3-7%，但在中国和印度未显示显著效果。文化异质性限制了LLM仿真助推效果的普适性，需结合实证验证。","reliability":"论文指出LLM仿真受文化异质性约束，在部分国家（中国、印度）失效，且个性化助推效果依赖于对个体偏好和文化背景的准确建模，需结合针对性实证验证。","relevance":"该研究直接以LLM仿真人类决策行为，并与大规模跨国人类实验对照，评估个性化助推效果，高度契合研究者对LLM人类仿真可靠性及失效条件的关注。","inspiration":"借鉴LLM根据个体特征生成个性化干预并利用跨国调查进行对照验证的方法。｜可迁移至消费者金融产品选择中的个性化信息披露实验，如退休储蓄计划或保险产品选择。｜以LLM为不同人口特征群体生成个性化信息呈现方式，通过在线实验测量选择行为，并以真实市场数据或已有行为实验数据作为基准对照。"}},{"id":"2508.06635","version":2,"title":"Valid Inference with Imperfect Synthetic Data","zh_title":"不完美合成数据的有效推断","abstract":"Predictions and generations from large language models are increasingly being explored as an aid in limited data regimes, such as in computational social science and human subjects research. While prior technical work has mainly explored the potential to use model-predicted labels for unlabeled data in a principled manner, there is increasing interest in using large language models to generate entirely new synthetic samples (e.g., synthetic simulations), such as in responses to surveys. However, it remains unclear by what means practitioners can combine such data with real data and yet produce statistically valid conclusions upon them. In this paper, we introduce a new estimator based on generalized method of moments, providing a hyperparameter-free solution with strong theoretical guarantees to address this challenge. Intriguingly, we find that interactions between the moment residuals of synthetic data and those of real data (i.e., when they are predictive of each other) can greatly improve estimates of the target parameter. We validate the finite-sample performance of our estimator across different tasks in computational social science applications, demonstrating large empirical gains.","authors":["Yewon Byun","Shantanu Gupta","Zachary C. Lipton","Rachel Leah Childers","Bryan Wilder"],"categories":["cs.LG","cs.AI","stat.ML"],"primary_category":"cs.LG","announce_type":"new","date":"2025-08-08","first_seen":"2025-08-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.06635","pdf_url":"https://arxiv.org/pdf/2508.06635","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A5","B1","B2"],"tags":["合成数据","统计推断","计算社会科学"],"reason":"用LLM生成合成调查样本，结合真实数据做统计推断，直接涉及人类仿真与对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":45,"question":"如何将LLM生成的合成数据与真实数据结合，进行统计上有效的推断？","design":"本文提出一种基于广义矩估计（GMM）的框架，将LLM生成的合成样本（如模拟调查回答）与真实样本结合，通过引入合成数据与真实数据之间的相关结构来提高估计精度。","baseline":"真实人类标注样本，包含已标注协变量和结果的小规模数据集。","findings":"当合成数据的矩残差能预测真实数据的矩残差时，结合合成数据可提高估计精度并缩小置信区间；即使合成数据完全无信息，也不会损害渐近有效性。","reliability":"论文指出，若生成模型与真实分布不匹配，直接简单聚合合成数据会导致严重偏差；所提方法依赖于合成样本与真实样本之间的相关结构，若两者独立则无增益但也不损失。","relevance":"该研究直接针对用LLM生成合成调查样本并做统计推断的场景，提供了结合真实数据与合成数据的严谨方法，对关注人类仿真可靠性的研究者具有重要参考价值。","inspiration":"该方法通过广义矩估计将LLM合成样本与真实样本结合，利用合成数据矩残差对真实数据矩残差的预测能力来提高估计精度，同时保证即使合成数据无信息也不损害推断有效性。｜可迁移到政策评估中的调查实验，例如评估税收优惠宣传对家庭消费意愿的影响，其中LLM生成模拟调查回答作为合成数据。｜以LLM模拟的家庭为被试，处理为展示税收优惠信息，结果变量为自报消费意愿，用真实家庭调查数据作为基准，通过GMM框架结合合成与真实样本估计处理效应。"}},{"id":"2508.02766","version":2,"title":"The Generative Reasonable Person","zh_title":"生成式理性人","abstract":"This Article introduces the generative reasonable person, a new tool for estimating how ordinary people judge reasonableness. As claims about AI capabilities often outpace evidence, the Article proceeds empirically: adapting randomized controlled trials to large language models, it replicates three published studies of lay judgment across negligence, consent, and contract interpretation, drawing on nearly 10,000 simulated decisions. The findings reveal that models can replicate subtle patterns that run counter to textbook treatment. Like human subjects, models prioritize social conformity over cost-benefit analysis when assessing negligence, inverting the hierarchy that textbooks teach. They reproduce the paradox that material lies erode consent less than lies about a transaction's essence. And they track lay contract formalism, judging hidden fees more enforceable than fair. For two centuries, scholars have debated whether the reasonable person is empirical or normative, majoritarian or aspirational. But much of this debate assumed a constraint that no longer holds: that lay judgments are expensive to surface, slow to collect, and unavailable at scale. Generative reasonable people loosen that constraint. They offer judges empirical checks on elite intuition, give resource-constrained litigants access to simulated jury feedback, and let regulators pilot-test public comprehension, all at a fraction of survey costs. The reasonable person standard has long functioned as a vessel for judicial intuition precisely because the empirical baseline was missing. With that baseline now available, departures from lay understanding become transparent rather than hidden, a choice to be justified, not a fact to be assumed. Properly cabined, the generative reasonable person may become a dictionary for reasonableness judgments.","authors":["Yonathan A. Arbel"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-04","first_seen":"2025-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.02766","pdf_url":"https://arxiv.org/pdf/2508.02766","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B3"],"tags":["LLM仿真","人类被试替代","法律判断"],"reason":"用LLM模拟普通人判断，复现三项实验，有真实人类数据对照，涉及法律判断与政策评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":3,"question":"如何利用大语言模型模拟普通人判断，以复现法律中“理性人”标准的实证研究？","design":"使用大语言模型扮演普通人群，采用随机对照试验设计，复现三项已发表的关于过失、同意和合同解释的普通人判断研究，测量模型在近10,000次模拟决策中的判断模式。","baseline":"三项已发表研究中的真实人类被试数据，涵盖过失判断、同意判断和合同解释判断。","findings":"模型能复现与教科书相悖的微妙模式：在过失评估中优先考虑社会从众而非成本收益分析；再现了实质性谎言比交易本质谎言更少侵蚀同意的悖论；在合同解释中表现出普通人合同形式主义，认为隐藏费用比公平条款更具可执行性。","reliability":"论文未讨论","relevance":"该研究直接使用LLM模拟人类被试，复现法律判断实验并与真实人类数据对照，评估仿真可靠性，完全契合研究者对LLM人类仿真实验的关注，值得精读原文。","inspiration":"借鉴其用LLM复现已有人类实验并系统对比结果的方法，可验证LLM在特定领域的行为一致性。｜可迁移到消费者金融决策实验，如信贷条款理解、费用披露效果评估。｜以LLM模拟消费者，随机呈现不同披露格式的信贷合同，测量其理解度和选择行为，以真实消费者调查数据为基准进行对照验证。"}},{"id":"2507.22049","version":1,"title":"Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel Predictions","zh_title":"验证基于生成式智能体的社会规范执行模型：从复现到新预测","abstract":"As large language models (LLMs) advance, there is growing interest in using them to simulate human social behavior through generative agent-based modeling (GABM). However, validating these models remains a key challenge. We present a systematic two-stage validation approach using social dilemma paradigms from psychological literature, first identifying the cognitive components necessary for LLM agents to reproduce known human behaviors in mixed-motive settings from two landmark papers, then using the validated architecture to simulate novel conditions. Our model comparison of different cognitive architectures shows that both persona-based individual differences and theory of mind capabilities are essential for replicating third-party punishment (TPP) as a costly signal of trustworthiness. For the second study on public goods games, this architecture is able to replicate an increase in cooperation from the spread of reputational information through gossip. However, an additional strategic component is necessary to replicate the additional boost in cooperation rates in the condition that allows both ostracism and gossip. We then test novel predictions for each paper with our validated generative agents. We find that TPP rates significantly drop in settings where punishment is anonymous, yet a substantial amount of TPP persists, suggesting that both reputational and intrinsic moral motivations play a role in this behavior. For the second paper, we introduce a novel intervention and see that open discussion periods before rounds of the public goods game further increase contributions, allowing groups to develop social norms for cooperation. This work provides a framework for validating generative agent models while demonstrating their potential to generate novel and testable insights into human social behavior.","authors":["Logan Cross","Nick Haber","Daniel L. K. Yamins"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2025-07-29","first_seen":"2025-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.22049","pdf_url":"https://arxiv.org/pdf/2507.22049","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM仿真","社会规范","行为博弈"],"reason":"用LLM智能体复现社会困境实验，与真实人类数据对照，验证仿真并预测新条件，直接…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":28,"question":"如何通过生成式智能体建模（GABM）复现并扩展社会困境中的人类行为，并验证其认知架构的有效性？","design":"使用LLM驱动的生成式智能体，通过组合记忆、人格、推理等认知组件构建不同架构，复现Jordan et al. (2016)的信任博弈第三方惩罚实验和Feinberg et al. (2014)的公共品博弈中流言与排斥实验，比较智能体行为与人类数据的统计效应，并利用验证后的架构模拟匿名惩罚和公开讨论等新条件。","baseline":"Jordan et al. (2016)的第三方惩罚信任博弈人类实验数据，以及Feinberg et al. (2014)的公共品博弈中流言与排斥对人类合作影响的数据。","findings":"复现第三方惩罚效应需要人格提示和心理理论组件，而公共品博弈中流言提升合作可被复现，但排斥与流言结合的额外合作提升需额外策略组件；匿名惩罚下第三方惩罚率显著下降但仍存在，表明声誉和内在道德动机共同驱动惩罚，公开讨论能进一步提升公共品贡献。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体复现经典社会困境实验，并与真实人类数据对照，验证仿真可靠性并生成新预测，完全契合研究者对LLM人类仿真、经济学实验复现及失效条件探索的兴趣，值得精读原文。","inspiration":"该方法通过组合记忆、人格、推理等认知组件构建不同LLM智能体架构，并与人类实验数据对照来验证仿真有效性，为经济金融实验提供了可借鉴的仿真验证框架。｜可迁移至资产定价实验中的羊群效应研究，利用LLM智能体模拟投资者在信息不对称下的决策行为。｜以LLM智能体为被试，处理为是否提供历史价格信息，结果变量为投资决策的羊群效应指数，对照真实人类资产定价实验数据。"}},{"id":"2507.10933","version":1,"title":"Artificial Finance: How AI Thinks About Money","zh_title":"人工金融：AI如何思考金钱","abstract":"In this paper, we explore how large language models (LLMs) approach financial decision-making by systematically comparing their responses to those of human participants across the globe. We posed a set of commonly used financial decision-making questions to seven leading LLMs, including five models from the GPT series(GPT-4o, GPT-4.5, o1, o3-mini), Gemini 2.0 Flash, and DeepSeek R1. We then compared their outputs to human responses drawn from a dataset covering 53 nations. Our analysis reveals three main results. First, LLMs generally exhibit a risk-neutral decision-making pattern, favoring choices aligned with expected value calculations when faced with lottery-type questions. Second, when evaluating trade-offs between present and future, LLMs occasionally produce responses that appear inconsistent with normative reasoning. Third, when we examine cross-national similarities, we find that the LLMs' aggregate responses most closely resemble those of participants from Tanzania. These findings contribute to the understanding of how LLMs emulate human-like decision behaviors and highlight potential cultural and training influences embedded within their outputs.","authors":["Orhan Erdem","Ragavi Pobbathi Ashok"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2025-07-15","first_seen":"2025-07-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.10933","pdf_url":"https://arxiv.org/pdf/2507.10933","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","金融决策","跨文化对照"],"reason":"用LLM复现人类金融决策并与53国真实数据对照，评估仿真行为与偏差，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":26,"question":"LLM在金融决策中表现出怎样的风险偏好、跨期选择模式，其总体回答与哪个国家的人类被试最相似？","design":"将7个主流LLM（GPT-4o、GPT-4.5、o1、o3-mini、Gemini 2.0 Flash、DeepSeek R1）作为被试，向其呈现一组常用的金融决策问题（包括彩票型问题和跨期选择问题），收集模型回答，并与53国人类调查数据进行比较。","baseline":"来自Wang et al. (2017)的覆盖53个国家的全球金融决策调查数据。","findings":"LLM普遍表现出风险中性，倾向于根据期望值计算选择彩票型问题；在跨期选择中偶尔出现与规范推理不一致的回答；LLM的总体回答模式与坦桑尼亚参与者最相似。","reliability":"论文未讨论","relevance":"该研究直接用LLM复现人类金融决策并与跨国真实数据对照，评估仿真行为与偏差，高度契合研究者对LLM人类仿真实验的关注，值得阅读原文。","inspiration":"借鉴其将LLM作为被试、使用标准化金融决策问题并直接与跨国人类调查数据对照的设计思路。｜可迁移到消费者跨期选择、风险资产配置或政策公告预期形成等行为金融场景。｜以LLM为被试，施加不同框架的跨期选择或风险决策问题，测量其时间偏好与风险厌恶参数，并与真实跨国调查数据（如Falk et al., 2018的全球偏好数据）进行对照，检验LLM仿真的人群代表性。"}},{"id":"2507.10342","version":1,"title":"Using AI to replicate human experimental results: a motion study","zh_title":"使用AI复现人类实验结果：一项运动研究","abstract":"This paper explores the potential of large language models (LLMs) as reliable analytical tools in linguistic research, focusing on the emergence of affective meanings in temporal expressions involving manner-of-motion verbs. While LLMs like GPT-4 have shown promise across a range of tasks, their ability to replicate nuanced human judgements remains under scrutiny. We conducted four psycholinguistic studies (on emergent meanings, valence shifts, verb choice in emotional contexts, and sentence-emoji associations) first with human participants and then replicated the same tasks using an LLM. Results across all studies show a striking convergence between human and AI responses, with statistical analyses (e.g., Spearman's rho = .73-.96) indicating strong correlations in both rating patterns and categorical choices. While minor divergences were observed in some cases, these did not alter the overall interpretative outcomes. These findings offer compelling evidence that LLMs can augment traditional human-based experimentation, enabling broader-scale studies without compromising interpretative validity. This convergence not only strengthens the empirical foundation of prior human-based findings but also opens possibilities for hypothesis generation and data expansion through AI. Ultimately, our study supports the use of LLMs as credible and informative collaborators in linguistic inquiry.","authors":["Rosa Illan Castillo","Javier Valenzuela"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-07-14","first_seen":"2025-07-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.10342","pdf_url":"https://arxiv.org/pdf/2507.10342","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1"],"tags":["LLM仿真","人类数据对照","心理语言学"],"reason":"用LLM复现人类心理语言学实验，并与真实人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":15,"question":"大语言模型能否在心理语言学实验中复现人类对运动动词情感意义的细微判断，从而作为可靠的分析工具？","design":"使用ChatGPT o1模拟人类被试，对包含运动动词的时间表达句进行四项心理语言学任务（情感意义评分、效价变化、情绪语境下的动词选择、句子与表情符号关联），测量其评分和选择模式。","baseline":"59名英语母语者在相同任务上的真实行为数据，包括评分分布和分类选择。","findings":"LLM与人类反应高度一致，Spearman相关系数在0.73至0.96之间，表明评分模式和分类选择均强相关。微小分歧未改变整体解释性结论，支持LLM可作为人类实验的补充工具。","reliability":"论文未讨论","relevance":"该研究直接以真实人类数据为基准，验证LLM复现心理语言学实验的可靠性，符合研究者对仿真有效性评估的关注，值得阅读以了解具体对照方法和统计检验。","inspiration":"借鉴其将同一任务原样施加于LLM并与人类数据直接对比的设计，可迁移到消费者跨期选择或政策公告的预期形成实验，例如用LLM模拟消费者对“时间飞逝”类表述的耐心程度评分，以真实调查数据为基准检验LLM能否复现时间偏好。"}},{"id":"2507.07188","version":3,"title":"Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses","zh_title":"提示扰动揭示大语言模型调查响应中类人偏差","abstract":"Large Language Models (LLMs) are increasingly used as proxies for human subjects in social science surveys, but their reliability and susceptibility to known human-like response biases, such as central tendency, opinion floating and primacy bias are poorly understood. This work investigates the response robustness of LLMs in normative survey contexts, we test nine LLMs on questions from the World Values Survey (WVS), applying a comprehensive set of ten perturbations to both question phrasing and answer option structure, resulting in over 167,000 simulated survey interviews. In doing so, we not only reveal LLMs' vulnerabilities to perturbations but also show that all tested models exhibit a consistent recency bias, disproportionately favoring the last-presented answer option. While larger models are generally more robust, all models remain sensitive to semantic variations like paraphrasing and to combined perturbations. This underscores the critical importance of prompt design and robustness testing when using LLMs to generate synthetic survey data.","authors":["Jens Rupprecht","Georg Ahnert","Markus Strohmaier"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-07-09","first_seen":"2025-07-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.07188","pdf_url":"https://arxiv.org/pdf/2507.07188","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查偏差","算法保真度"],"reason":"直接研究LLM作为人类被试替代品的调查响应偏差，使用世界价值观调查真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"LLM在回答封闭式规范性调查问题时，对提示扰动是否稳健，以及是否表现出类似人类的回答偏差？","design":"用9个指令微调LLM扮演人类被试，对世界价值观调查的62个问题施加10种提示扰动（包括答案选项和问题措辞的修改），测量回答分布的变化，共进行167,400次模拟访谈。","baseline":"对照的真实人类数据来自世界价值观调查第7波（2017-2022）的核心变量，但论文主要比较扰动后回答分布与原始提示基线分布的差异。","findings":"所有模型均表现出明显的近因偏差，即不成比例地偏好最后一个选项；较大模型通常更稳健，但所有模型对语义改写和组合扰动仍敏感。","reliability":"论文指出，即使最大模型也对问题措辞变化敏感，提示设计和稳健性测试对使用LLM生成合成调查数据至关重要，但未讨论其他失效条件。","relevance":"该研究直接评估LLM作为人类被试替代品在调查中的偏差与可靠性，使用真实调查数据作为对照，并揭示近因偏差等关键失效模式，高度契合你的关注点，值得精读原文。","inspiration":"该方法通过系统施加提示扰动（如选项顺序、措辞改写）测量LLM回答分布变化，并设置原始提示基线作为对照，可借鉴用于稳健性检验设计｜可迁移到消费者通胀预期调查或政策公告解读实验，评估LLM模拟的预期形成是否对问卷设计敏感｜以LLM作为被试，施加不同措辞的通胀预期问题（如‘未来一年物价变化’ vs ‘通胀率’），结果变量为预期值分布，对照真实消费者预期调查数据（如密歇根大学调查）"}},{"id":"2506.23610","version":1,"title":"Evaluating the Simulation of Human Personality-Driven Susceptibility to Misinformation with LLMs","zh_title":"评估大语言模型对人类人格驱动的错误信息易感性模拟","abstract":"Large language models (LLMs) make it possible to generate synthetic behavioural data at scale, offering an ethical and low-cost alternative to human experiments. Whether such data can faithfully capture psychological differences driven by personality traits, however, remains an open question. We evaluate the capacity of LLM agents, conditioned on Big-Five profiles, to reproduce personality-based variation in susceptibility to misinformation, focusing on news discernment, the ability to judge true headlines as true and false headlines as false. Leveraging published datasets in which human participants with known personality profiles rated headline accuracy, we create matching LLM agents and compare their responses to the original human patterns. Certain trait-misinformation associations, notably those involving Agreeableness and Conscientiousness, are reliably replicated, whereas others diverge, revealing systematic biases in how LLMs internalize and express personality. The results underscore both the promise and the limits of personality-aligned LLMs for behavioral simulation, and offer new insight into modeling cognitive diversity in artificial agents.","authors":["Manuel Pratelli","Marinella Petrocchi"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-30","first_seen":"2025-06-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.23610","pdf_url":"https://arxiv.org/pdf/2506.23610","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人格与行为","错误信息"],"reason":"用LLM模拟人格对错误信息易感性的影响，并与真实人类数据对照，评估仿真可靠性与…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":63,"question":"当赋予大语言模型明确的大五人格特征时，它们能否模拟人类在新闻辨别力（区分真假新闻标题准确性）上的表现？","design":"基于Calvillo等(2023)和Huang等(2024)的公开数据集，为每位人类参与者创建一个匹配的LLM智能体，通过人格注入管道赋予其对应的大五人格档案，让智能体对同一组新闻标题进行准确性评分，并比较合成数据与原始人类数据在人格-新闻辨别力关联上的异同。","baseline":"以Calvillo等(2023)的336名美国参与者数据为主要基准，该数据集包含参与者的大五人格档案和对真假新闻标题的准确性评分。","findings":"GPT-4o能复现部分人类特质-辨别力关联，如宜人性和尽责性与新闻辨别力正相关，但外向性和负面情绪性的模式与人类不一致；在仅分析假新闻时，高宜人性和高尽责性同样预测较低的易感性，但开放性的关联在不同模型设置下不稳定。","reliability":"论文承认某些人格特质（如外向性、负面情绪性）的模拟与人类模式存在系统性偏差，开放性-辨别力关联在不同模型参数和设置下不一致，表明LLM在表达人格时存在内在偏差，心理保真度有限。","relevance":"该研究直接使用LLM模拟人格对错误信息易感性的影响，并与真实人类数据严格对照，评估了仿真的可靠性与偏差，完全符合您对LLM人类仿真实验、经济学/政策评估场景及批判性失效条件分析的兴趣，值得精读原文。","inspiration":"该方法通过人格注入管道为LLM智能体赋予真实人类的大五人格档案，并严格对照人类基准数据评估仿真可靠性，值得借鉴其对照设计和偏差分析思路｜可迁移至消费者金融决策研究，如人格特质对投资风险偏好或过度借贷行为的影响｜以LLM智能体为被试，注入真实投资者的人格档案，让其评估不同风险等级的金融产品，结果变量为风险评分，对照真实投资者在相同产品上的风险偏好调查数据"}},{"id":"2506.23107","version":1,"title":"Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study","zh_title":"大语言模型能捕捉人类风险偏好吗？一项跨文化研究","abstract":"Large language models (LLMs) have made significant strides, extending their applications to dialogue systems, automated content creation, and domain-specific advisory tasks. However, as their use grows, concerns have emerged regarding their reliability in simulating complex decision-making behavior, such as risky decision-making, where a single choice can lead to multiple outcomes. This study investigates the ability of LLMs to simulate risky decision-making scenarios. We compare model-generated decisions with actual human responses in a series of lottery-based tasks, using transportation stated preference survey data from participants in Sydney, Dhaka, Hong Kong, and Nanjing. Demographic inputs were provided to two LLMs -- ChatGPT 4o and ChatGPT o1-mini -- which were tasked with predicting individual choices. Risk preferences were analyzed using the Constant Relative Risk Aversion (CRRA) framework. Results show that both models exhibit more risk-averse behavior than human participants, with o1-mini aligning more closely with observed human decisions. Further analysis of multilingual data from Nanjing and Hong Kong indicates that model predictions in Chinese deviate more from actual responses compared to English, suggesting that prompt language may influence simulation performance. These findings highlight both the promise and the current limitations of LLMs in replicating human-like risk behavior, particularly in linguistic and cultural settings.","authors":["Bing Song","Jianing Liu","Sisi Jian","Chenyang Wu","Vinayak Dixit"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-06-29","first_seen":"2025-06-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.23107","pdf_url":"https://arxiv.org/pdf/2506.23107","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","风险偏好","跨文化对照"],"reason":"用LLM仿真人类风险决策，与真实调查数据对照，评估跨文化偏差，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":64,"question":"大语言模型能否在跨文化背景下模拟人类的风险偏好？","design":"使用ChatGPT 4o和o1-mini两个模型，输入年龄、性别、教育、收入等人口统计信息，预测个体在彩票选择任务中的决策，并基于CRRA框架估计风险偏好。","baseline":"来自悉尼、达卡、香港和南京四个城市的真实交通陈述偏好调查数据，包含个体在彩票游戏中的实际选择。","findings":"两个模型均比人类更风险厌恶，o1-mini比4o更接近真实决策；在中文提示下，模型预测偏离实际的程度大于英文提示，表明语言影响仿真表现。","reliability":"论文指出模型在中文语境下偏差更大，且未完全复现人类风险行为，提示在语言和文化环境中的局限性。","relevance":"直接以真实人类数据为基准，评估LLM在风险决策仿真中的跨文化偏差，命中研究者关注的核心问题，值得精读原文。","inspiration":"该方法通过向LLM输入人口统计特征来模拟个体在彩票选择中的决策，并与真实调查数据对照，可用于评估模型在结构化风险任务中的行为偏差｜可迁移到金融决策中的风险偏好测量，如投资组合选择、保险购买或退休储蓄决策，检验LLM能否复现不同文化下个体的金融风险态度｜研究设计：以真实家庭金融调查数据为基准，向LLM输入年龄、收入、教育等特征，要求其在多组假设性投资选项中做出选择，结果变量为风险资产配置比例，对比模型预测与真实家庭行为的分布差异"}},{"id":"2506.21974","version":3,"title":"Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism","zh_title":"不要相信生成式智能体能模仿社交网络上的交流，除非你对其经验现实主义进行了基准测试","abstract":"The ability of Large Language Models (LLMs) to mimic human behavior triggered a plethora of computational social science research, assuming that empirical studies of humans can be conducted with AI agents instead. Since there have been conflicting research findings on whether and when this hypothesis holds, there is a need to better understand the differences in their experimental designs. We focus on replicating the behavior of social network users with the use of LLMs for the analysis of communication on social networks. First, we provide a formal framework for the simulation of social networks, before focusing on the sub-task of imitating user communication. We empirically test different approaches to imitate user behavior on X in English and German. Our findings suggest that social simulations should be validated by their empirical realism measured in the setting in which the simulation components were fitted. With this paper, we argue for more rigor when applying generative-agent-based modeling for social simulation.","authors":["Simon Münker","Nils Schwager","Achim Rettinger"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-27","first_seen":"2025-06-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.21974","pdf_url":"https://arxiv.org/pdf/2506.21974","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","社交网络","经验现实主义"],"reason":"用LLM仿真社交网络用户行为，有真实人类数据对照，并评估仿真可靠性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":29,"question":"LLM在多大程度上能忠实模仿社交网络（X平台）上不同用户群体的发帖、回复等沟通行为？","design":"基于X平台英文和德文政治讨论数据，用LLM（如Llama 3.1 70B）分别拟合政治人物（发帖者）和普通用户（回复者）的沟通行为，包括帖子生成、回复生成和回复可能性预测三个子任务，并比较不同建模方法的表现。","baseline":"真实人类数据：从X平台收集的2023年上半年美国和德国政治话语相关帖子（议员）和回复（普通用户），并标注主题。","findings":"LLM模仿用户沟通行为的效果因任务、语言和建模方法而异，并非总能忠实复现；社会仿真必须在其组件拟合的设定下验证经验现实主义。","reliability":"论文承认仿真框架存在简化，如未完全建模网络结构、内容过滤可能损失信息，且结论依赖于特定平台和语言数据集，泛化性有限。","relevance":"该研究直接回应了LLM仿真人类社交行为的可靠性问题，提供了有真实人类基准的实证评估，并指出仿真失效的条件，高度契合研究者对批判性仿真研究的关注。","inspiration":"该方法值得借鉴之处在于，它将LLM仿真分解为帖子生成、回复生成和回复可能性预测三个子任务，并分别与真实人类数据对比，从而定位仿真失效的具体环节｜该思路可迁移到政策公告的预期形成研究，例如分析央行沟通或财政政策声明如何影响市场参与者的预期与讨论｜研究设计：以LLM模拟投资者和分析师，处理为不同措辞或情感倾向的政策声明，结果变量为LLM生成的预期文本和市场反应预测，对照真实数据可来自央行声明后社交媒体（如Twitter/微博）上的实际讨论与市场调查预期数据"}},{"id":"2507.02919","version":1,"title":"ChatGPT is not A Man but Das Man: Representativeness and Structural Consistency of Silicon Samples Generated by Large Language Models","zh_title":"ChatGPT不是人而是常人：大语言模型生成硅样本的代表性与结构一致性","abstract":"Large language models (LLMs) in the form of chatbots like ChatGPT and Llama are increasingly proposed as \"silicon samples\" for simulating human opinions. This study examines this notion, arguing that LLMs may misrepresent population-level opinions. We identify two fundamental challenges: a failure in structural consistency, where response accuracy doesn't hold across demographic aggregation levels, and homogenization, an underrepresentation of minority opinions. To investigate these, we prompted ChatGPT (GPT-4) and Meta's Llama 3.1 series (8B, 70B, 405B) with questions on abortion and unauthorized immigration from the American National Election Studies (ANES) 2020. Our findings reveal significant structural inconsistencies and severe homogenization in LLM responses compared to human data. We propose an \"accuracy-optimization hypothesis,\" suggesting homogenization stems from prioritizing modal responses. These issues challenge the validity of using LLMs, especially chatbots AI, as direct substitutes for human survey data, potentially reinforcing stereotypes and misinforming policy.","authors":["Dai Li","Linzhuo Li","Huilian Sophie Qiu"],"categories":["cs.CL","cs.CY","cs.ET"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-25","first_seen":"2025-06-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.02919","pdf_url":"https://arxiv.org/pdf/2507.02919","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","调查数据对照","算法偏差"],"reason":"直接用LLM模拟人类调查回答，与ANES真实数据对照，揭示结构不一致和同质化，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:16","error":null,"has_summary":true,"summary":{"generated_at":"2025-06-25","rank":8,"question":"大语言模型生成的“硅样本”能否代表人类群体意见？","design":"使用ChatGPT (GPT-4) 和 Llama 3.1系列 (8B, 70B, 405B) 模型，输入ANES 2020中关于堕胎和非法移民的问题，生成模拟回答，并与真实人类数据对比，分析结构一致性和同质化程度。","baseline":"美国国家选举研究（ANES）2020的调查数据。","findings":"LLM回答存在显著的结构不一致性（不同人口聚合水平的准确率不一致）和严重的同质化（少数意见被低估）。作者提出“精度优化假说”，认为同质化源于模型优先输出众数回答。","reliability":"论文指出LLM可能误代表群体意见，存在结构不一致和同质化问题，挑战了直接替代人类调查数据的有效性，可能强化刻板印象并误导政策。","relevance":"高度相关。该研究直接评估了LLM仿真人类意见的可靠性与偏差，有真实人类数据对照，并指出了失效条件（结构不一致和同质化），符合研究者的核心关注点，值得精读原文。","inspiration":"该方法通过将LLM作为被试，输入真实调查问题并对比其回答与人类基准数据，来评估仿真的一致性和偏差，可借鉴其对照设计和偏差度量方式｜可迁移到政策预期形成的实验，例如研究公众对央行利率决议声明的通胀预期反应｜以GPT-4等LLM为被试，输入央行政策声明文本，测量其输出的通胀预期值，并与密歇根消费者调查的真实预期数据对比，检验LLM是否复现预期分布及同质化偏差"}},{"id":"2506.21587","version":2,"title":"A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies","zh_title":"基于大语言模型的舆论仿真跨文化比较：评估中美模型在多元社会上的表现","abstract":"This study evaluates the ability of DeepSeek, an open-source large language model (LLM), to simulate public opinions in comparison to LLMs developed by major tech companies. By comparing DeepSeek-R1 and DeepSeek-V3 with Qwen2.5, GPT-4o, and Llama-3.3 and utilizing survey data from the American National Election Studies (ANES) and the Zuobiao dataset of China, we assess these models' capacity to predict public opinions on social issues in both China and the United States, highlighting their comparative capabilities between countries. Our findings indicate that DeepSeek-V3 performs best in simulating U.S. opinions on the abortion issue compared to other topics such as climate change, gun control, immigration, and services for same-sex couples, primarily because it more accurately simulates responses when provided with Democratic or liberal personas. For Chinese samples, DeepSeek-V3 performs best in simulating opinions on foreign aid and individualism but shows limitations in modeling views on capitalism, particularly failing to capture the stances of low-income and non-college-educated individuals. It does not exhibit significant differences from other models in simulating opinions on traditionalism and the free market. Further analysis reveals that all LLMs exhibit the tendency to overgeneralize a single perspective within demographic groups, often defaulting to consistent responses within groups. These findings highlight the need to mitigate cultural and demographic biases in LLM-driven public opinion modeling, calling for approaches such as more inclusive training methodologies.","authors":["Weihong Qi","Fan Huang","Jisun An","Haewoon Kwak"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-17","first_seen":"2025-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.21587","pdf_url":"https://arxiv.org/pdf/2506.21587","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","舆论模拟","跨文化比较"],"reason":"用LLM仿真中美公众舆论，有ANES和Zuobiao真实调查数据对照，涉及社会…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":15,"question":"LLM的文化出身是否带来舆论模拟的“主场优势”？不同模型在中美社会议题上存在怎样的文化与人口偏差？","design":"用DeepSeek-R1、DeepSeek-V3、Qwen2.5、GPT-4o、Llama-3.3，根据ANES和Zuobiao数据集中个体的种族、性别、年龄、收入、教育等人口属性构造提示，让模型回答社会议题问卷，聚合后比较模拟分布与真实调查分布。","baseline":"2020年美国国家选举研究（ANES）的2457人调查数据，以及中国“左标”数据集随机抽取的2000人调查数据。","findings":"DeepSeek-V3在模拟美国堕胎议题上表现最好，主要因为能较准地模拟民主党/自由派，但共和党/保守派模拟差；所有模型都倾向于在人口群体内过度概括单一观点，尤其在中国资本主义议题上，Qwen2.5通过输出统一回答虚高准确率。","reliability":"论文指出所有模型均存在显著的文化与人口偏差，难以捕捉特定群体的细微观点，常退化为刻板或过度概括的回答；未讨论提示设计、模型版本或数据集时效性等可能影响结论的因素。","relevance":"直接回应LLM仿真人类舆论的可靠性与偏差问题，有中美真实调查数据对照，揭示模型在跨文化情境下系统性失效的模式，对理解仿真在经济学/政策评估中的局限有重要参考价值，强烈建议阅读原文。","inspiration":"该方法通过将人口属性编码为提示来模拟个体回答，并对比聚合分布与真实调查数据，为评估仿真偏差提供了可操作的对照框架｜可迁移到消费者信心调查或通胀预期形成的仿真研究，检验LLM能否复现不同人口群体的预期差异｜以LLM作为被试，输入年龄、收入、教育等属性，让其预测未来通胀率，结果变量为预期值分布，以密歇根大学消费者调查的微观数据作为真实基准，比较模拟与真实分布的偏差模式"}},{"id":"2506.14611","version":2,"title":"Exploring MLLMs Perception of Network Visualization Principles","zh_title":"探索多模态大语言模型对网络可视化原则的感知","abstract":"In this paper, we test whether Multimodal Large Language Models (MLLMs) can match human-subject performance in tasks involving the perception of properties in network layouts. Specifically, we replicate a human-subject experiment about perceiving quality (namely stress) in network layouts using GPT-4o, Gemini-2.5 and Qwen2.5. Our experiments show that giving MLLMs the same study information as trained human participants yields performance comparable to that of human experts and exceeds that of untrained non-experts. Additionally, we show that prompt engineering that deviates from the human-subject experiment can lead to better-than-human performance in some settings. Interestingly, like human subjects, the MLLMs seem to rely on visual proxies rather than computing the actual value of stress, indicating some sense or facsimile of perception. Explanations from the models are similar to those used by the human participants (e.g., an even distribution of nodes and uniform edge lengths).","authors":["Jacob Miller","Markus Wallinger","Ludwig Felder","Timo Brand","Henry Förster","Johannes Zink","Chunyang Chen","Stephen Kobourov"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-06-17","first_seen":"2025-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.14611","pdf_url":"https://arxiv.org/pdf/2506.14611","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类被试替代","网络感知实验"],"reason":"用MLLM复现人类网络布局感知实验，与真实人类数据对照，评估仿真可靠性并指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":65,"question":"多模态大语言模型在判断网络布局应力时，能否达到人类被试的感知表现？","design":"用GPT-4o、Gemini-2.5和Qwen2.5模拟人类被试，复现Mooney等人的配对刺激实验：向模型展示同一网络的两幅节点链接图，要求选择应力更低的图或判断两者相似，并比较不同提示工程（与人类相同指令 vs. 偏离人类指令）下的表现。","baseline":"Mooney等人实验中的人类被试数据，包括未训练新手、训练新手和专家三组。","findings":"给予与人类训练参与者相同的信息时，MLLM的表现与人类专家相当，优于未训练新手；通过偏离人类实验的提示工程，在某些设置下可获得超越人类的表现。MLLM与人类相似，依赖视觉代理（如节点均匀分布、边长一致）而非实际计算应力值。","reliability":"论文未讨论","relevance":"该研究直接用MLLM复现人类网络感知实验，并与真实人类数据对照，评估仿真可靠性，同时指出提示工程可导致超越人类的表现，符合您对LLM仿真人类实验、基准对照及失效条件探索的关注，值得精读原文。","inspiration":"该方法借鉴了用多模态大语言模型复现人类感知实验，并通过提示工程模拟不同信息条件（如训练新手、专家）来检验仿真表现，同时以真实人类数据作为基准对照。｜可迁移到金融图表解读与投资决策实验，例如研究投资者如何从网络关系图（如持股网络、供应链网络）中提取风险信息并形成投资判断。｜以GPT-4o等MLLM为被试，展示同一组持股网络的不同布局图，要求选择更易读或风险更清晰的图，结果变量为选择准确率与反应时间，对照真实投资者在相同任务上的行为数据。"}},{"id":"2506.21574","version":1,"title":"Digital Gatekeepers: Exploring Large Language Model's Role in Immigration Decisions","zh_title":"数字守门人：探索大语言模型在移民决策中的作用","abstract":"With globalization and increasing immigrant populations, immigration departments face significant work-loads and the challenge of ensuring fairness in decision-making processes. Integrating artificial intelligence offers a promising solution to these challenges. This study investigates the potential of large language models (LLMs),such as GPT-3.5 and GPT-4, in supporting immigration decision-making. Utilizing a mixed-methods approach,this paper conducted discrete choice experiments and in-depth interviews to study LLM decision-making strategies and whether they are fair. Our findings demonstrate that LLMs can align their decision-making with human strategies, emphasizing utility maximization and procedural fairness. Meanwhile, this paper also reveals that while ChatGPT has safeguards to prevent unintentional discrimination, it still exhibits stereotypes and biases concerning nationality and shows preferences toward privileged group. This dual analysis highlights both the potential and limitations of LLMs in automating and enhancing immigration decisions.","authors":["Yicheng Mao","Yang Zhao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-15","first_seen":"2025-06-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.21574","pdf_url":"https://arxiv.org/pdf/2506.21574","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","决策实验","公平性评估"],"reason":"用LLM模拟移民决策并与人类策略对照，评估公平性与偏差，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":59,"question":"大语言模型在移民决策中的决策策略是否与人类一致，以及其决策是否公平、是否存在偏见？","design":"使用GPT-3.5和GPT-4作为被试，复现Hainmueller和Hopkins (2015)的离散选择实验，生成10,000个随机移民决策场景，并辅以深度访谈，分析模型的决策策略、公平性和偏见。","baseline":"对照Hainmueller和Hopkins (2015)中的人类决策数据，比较LLM与人类在移民决策中的策略一致性。","findings":"LLM的决策策略与人类相似，强调效用最大化和程序公平；但ChatGPT虽设有防止无意歧视的机制，仍表现出基于国籍的刻板印象和偏见，并偏好特权群体。","reliability":"论文未讨论","relevance":"该研究直接用LLM模拟人类移民决策，并与真实人类实验数据对照，评估公平性与偏差，属于典型的LLM人类仿真研究，且涉及政策评估场景，值得精读原文。","inspiration":"该方法借鉴了离散选择实验与真实人类基准对照的设计，可系统评估LLM在政策决策中的策略一致性与偏差｜可迁移至信贷审批歧视研究，检验LLM是否复现人类审批中的种族、性别或收入偏见｜以LLM作为信贷审批官，处理随机生成的贷款申请人档案（操纵种族、收入等特征），结果变量为批准决策，对照真实银行审批数据或审计研究结果"}},{"id":"2507.19495","version":1,"title":"Simulating Human Behavior with the Psychological-mechanism Agent: Integrating Feeling, Thought, and Action","zh_title":"基于心理机制代理的人类行为模拟：整合感受、思维与行动","abstract":"Generative agents have made significant progress in simulating human behavior, but existing frameworks often simplify emotional modeling and focus primarily on specific tasks, limiting the authenticity of the simulation. Our work proposes the Psychological-mechanism Agent (PSYA) framework, based on the Cognitive Triangle (Feeling-Thought-Action), designed to more accurately simulate human behavior. The PSYA consists of three core modules: the Feeling module (using a layer model of affect to simulate changes in short-term, medium-term, and long-term emotions), the Thought module (based on the Triple Network Model to support goal-directed and spontaneous thinking), and the Action module (optimizing agent behavior through the integration of emotions, needs and plans). To evaluate the framework's effectiveness, we conducted daily life simulations and extended the evaluation metrics to self-influence, one-influence, and group-influence, selection five classic psychological experiments for simulation. The results show that the PSYA framework generates more natural, consistent, diverse, and credible behaviors, successfully replicating human experimental outcomes. Our work provides a richer and more accurate emotional and cognitive modeling approach for generative agents and offers an alternative to human participants in psychological experiments.","authors":["Qing Dong","Pengyuan Liu","Dong Yu","Chen Kang"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2025-06-04","first_seen":"2025-06-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.19495","pdf_url":"https://arxiv.org/pdf/2507.19495","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2"],"tags":["人类行为仿真","心理学实验复现","生成式代理"],"reason":"用LLM代理复现经典心理学实验，并与真实人类数据对照，直接替代人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":30,"question":"能否基于认知三角（感受-思维-行动）构建心理机制智能体框架，更真实地模拟人类日常行为并复现经典心理学实验？","design":"提出PSYA框架，包含感受（分层情感模型）、思维（三重网络模型支持目标导向与自发思维）、行动（整合情绪、需求与计划）三个模块，基于Llama-3-70B构建智能体，在虚拟小镇中模拟8个智能体的日常生活，并选取5个经典心理学实验进行仿真，测量行为自然度、一致性、多样性等指标。","baseline":"以经典心理学实验的真实人类结果作为对照基准，验证智能体能否复现人类实验数据。","findings":"PSYA框架能生成更自然、一致、多样和可信的行为，成功复现了所选经典心理学实验的结果；分层情感和自发思维模块显著提升了行为真实性。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体替代人类被试复现经典心理学实验，并与真实人类数据对照，高度契合研究者对经济学实验和政策评估场景中仿真可靠性的关注，值得精读原文。","inspiration":"借鉴PSYA框架的分层情感与自发思维模块设计，在LLM智能体中嵌入情绪和认知偏差来模拟经济决策的心理机制｜可迁移到消费者跨期选择实验，考察情绪波动对时间偏好一致性的影响｜以LLM智能体为被试，施加情绪启动处理（如积极/消极文本诱导），测量跨期选择中的折现率变化，并与真实人类实验数据（如经典双曲折现研究）对照"}},{"id":"2505.21997","version":1,"title":"Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data","zh_title":"利用访谈信息引导大语言模型建模调查回答：AI生成数据与人类数据的比较洞察","abstract":"Mixed methods research integrates quantitative and qualitative data but faces challenges in aligning their distinct structures, particularly in examining measurement characteristics and individual response patterns. Advances in large language models (LLMs) offer promising solutions by generating synthetic survey responses informed by qualitative data. This study investigates whether LLMs, guided by personal interviews, can reliably predict human survey responses, using the Behavioral Regulations in Exercise Questionnaire (BREQ) and interviews from after-school program staff as a case study. Results indicate that LLMs capture overall response patterns but exhibit lower variability than humans. Incorporating interview data improves response diversity for some models (e.g., Claude, GPT), while well-crafted prompts and low-temperature settings enhance alignment between LLM and human responses. Demographic information had less impact than interview content on alignment accuracy. These findings underscore the potential of interview-informed LLMs to bridge qualitative and quantitative methodologies while revealing limitations in response variability, emotional interpretation, and psychometric fidelity. Future research should refine prompt design, explore bias mitigation, and optimize model settings to enhance the validity of LLM-generated survey data in social science research.","authors":["Jihong Zhang","Xinya Liang","Anqi Deng","Nicole Bonge","Lin Tan","Ling Zhang","Nicole Zarrett"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-05-28","first_seen":"2025-05-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.21997","pdf_url":"https://arxiv.org/pdf/2505.21997","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查回答","人类数据对照"],"reason":"用访谈引导LLM生成调查回答，与真实人类数据对照，评估仿真可靠性与偏差，直接命…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":8,"question":"以个人访谈信息引导的大语言模型能否可靠预测人类在标准化问卷上的回答？","design":"使用Claude、GPT等大语言模型，输入课后项目工作人员的个人访谈文本和人口统计信息作为提示，生成其在运动行为调节问卷（BREQ）上的Likert量表回答，并与真实人类回答进行对比。","baseline":"同一批课后项目工作人员在BREQ问卷上的真实人类回答。","findings":"LLM能捕捉整体回答模式，但变异性低于人类；加入访谈内容可提升部分模型的回答多样性，精心设计的提示和低温度参数能增强对齐度。项目分析显示LLM在反向措辞题目上偏差较大，且难以重建问卷的心理测量结构。","reliability":"论文承认LLM在回答变异性、情感解读和心理测量保真度方面存在局限，且访谈内容的相关性比长度更重要，不同受访者间模型表现存在差异。","relevance":"该研究直接以真实人类调查数据为基准，评估访谈引导的LLM仿真回答的可靠性与偏差，与研究者关注的LLM人类仿真实验高度吻合，值得精读。","inspiration":"借鉴用个人深度访谈作为LLM输入来模拟特定个体问卷回答的设计，可迁移到消费者信心调查或通胀预期形成等经济金融场景。｜设计一个研究：以真实消费者的财务访谈文本作为处理，让LLM生成对未来通胀或就业的预期评分，结果变量为预期值，以实际调查数据（如密歇根消费者调查）作为对照基准。"}},{"id":"2505.17479","version":1,"title":"Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions","zh_title":"Twin-2K-500：基于2000余人对500余题回答构建数字孪生的数据集","abstract":"LLM-based digital twin simulation, where large language models are used to emulate individual human behavior, holds great promise for research in AI, social science, and digital experimentation. However, progress in this area has been hindered by the scarcity of real, individual-level datasets that are both large and publicly available. This lack of high-quality ground truth limits both the development and validation of digital twin methodologies. To address this gap, we introduce a large-scale, public dataset designed to capture a rich and holistic view of individual human behavior. We survey a representative sample of $N = 2,058$ participants (average 2.42 hours per person) in the US across four waves with 500 questions in total, covering a comprehensive battery of demographic, psychological, economic, personality, and cognitive measures, as well as replications of behavioral economics experiments and a pricing survey. The final wave repeats tasks from earlier waves to establish a test-retest accuracy baseline. Initial analyses suggest the data are of high quality and show promise for constructing digital twins that predict human behavior well at the individual and aggregate levels. By making the full dataset publicly available, we aim to establish a valuable testbed for the development and benchmarking of LLM-based persona simulations. Beyond LLM applications, due to its unique breadth and scale the dataset also enables broad social science research, including studies of cross-construct correlations and heterogeneous treatment effects.","authors":["Olivier Toubia","George Z. Gui","Tianyi Peng","Daniel J. Merlau","Ang Li","Haozhe Chen"],"categories":["cs.CY","cs.AI","cs.HC","econ.EM"],"primary_category":"cs.CY","announce_type":"new","date":"2025-05-23","first_seen":"2025-05-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.17479","pdf_url":"https://arxiv.org/pdf/2505.17479","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["数字孪生","人类仿真","行为经济学"],"reason":"直接构建LLM数字孪生仿真个体行为，含真实人类对照数据，涉及行为经济学实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:08","error":null,"has_summary":true,"summary":{"generated_at":"2025-05-23","rank":2,"question":"构建并公开一个大规模、多维度的人类行为数据集，用于开发和验证基于LLM的数字孪生仿真。","design":"该研究不是仿真实验，而是数据集构建。对2058名美国代表性样本进行4轮调查，每名被试平均2.42小时，共500题，涵盖人口统计、心理、经济偏好、人格、认知测试，并复现了行为经济学实验（包括组间和组内设计）及定价调查。","baseline":"人类基准：2058名真实被试的问卷调查数据，包括行为经济学实验的复现结果，以及第四轮重复前几轮任务以建立重测信度基线。","findings":"数据质量良好：测量间相关性具有表面效度，复现了行为经济学文献中几乎所有已知结果，重测信度稳健。初步数字孪生预测在个体和聚合水平上表现良好，但具体准确率未在摘要中给出。","reliability":"论文未讨论失效条件与局限，但指出数据集可用于评估数字孪生的可靠性，且已有研究提示LLM仿真可能受提示架构、混杂变量和代表性偏差影响。","relevance":"高度相关。该数据集提供了大规模、多维度的人类行为基准，可直接用于评估LLM数字孪生在复现调查回答、实验行为和决策模式上的准确性，尤其包含经济学实验复现，适合研究者验证仿真可靠性及偏差条件。","inspiration":"借鉴其大规模多维度调查设计，在同一被试内复现行为经济学实验并建立重测信度基线，为仿真验证提供个体与聚合层面的对照基准｜可迁移到资产定价实验中的风险偏好与预期形成研究，检验LLM数字孪生能否复现真实投资者的风险态度和价格预期分布｜以真实投资者为被试，收集其风险偏好问卷、资产选择实验及市场预期数据作为基准，用LLM基于相同问卷生成数字孪生，比较两者在风险资产配置和预期回报估计上的个体与分布一致性"}},{"id":"2505.10309","version":3,"title":"A large-scale evaluation of commonsense knowledge in humans and large language models","zh_title":"人类与大语言模型常识知识的大规模评估","abstract":"Commonsense knowledge, a major constituent of artificial intelligence (AI), is primarily evaluated in practice by human-prescribed ground-truth labels. An important, albeit implicit, assumption of these labels is that they accurately capture what any human would think, effectively treating human common sense as homogeneous. However, recent empirical work has shown that humans vary enormously in what they consider commonsensical; thus what appears self-evident to one benchmark designer may not be so to another. Here, we propose a method for assessing commonsense knowledge in AI, specifically in large language models (LLMs), that incorporates empirically observed heterogeneity among humans by measuring the correspondence between a model's judgment and that of a human population. We first find that, when treated as independent survey respondents, most LLMs remain below the human median in their individual commonsense competence. Second, when used as simulators of a hypothetical population, LLMs correlate with real humans only modestly in the extent to which they agree on the same set of statements. In both cases, smaller, open-weight models are surprisingly more competitive than larger, proprietary frontier models. Our evaluation framework, which ties commonsense knowledge to its cultural basis, contributes to the growing call for adapting AI models to human collectivities that possess different, often incompatible, social stocks of knowledge.","authors":["Tuan Dung Nguyen","Duncan J. Watts","Mark E. Whiting"],"categories":["cs.AI","cs.HC","cs.SI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-05-15","first_seen":"2025-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.10309","pdf_url":"https://arxiv.org/pdf/2505.10309","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","常识知识"],"reason":"将LLM作为独立调查受访者，与真实人类常识判断分布对照，评估仿真可靠性与异质性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":36,"question":"如何将人类常识判断的异质性纳入评估，衡量大语言模型与人类群体在常识知识上的一致性？","design":"将35个LLM作为独立调查受访者，收集其对常识陈述的判断，并与2046名人类受访者的判断分布进行对比；同时用LLM生成硅样本群体，模拟人类群体对陈述的共识程度。","baseline":"2046名人类受访者对常识陈述的判断分布，包括个体间共识和群体共识。","findings":"多数LLM的常识能力低于人类中位数，且LLM模拟的群体共识与真实人类群体共识仅呈中等相关；小型开源模型的表现可与大型闭源模型竞争。","reliability":"论文指出常识知识具有文化依赖性，当前评估框架仅基于特定人类群体，可能不适用于其他社会文化背景；LLM模拟的群体在定性上与人类存在差异，如Gemini Pro 1.0过度关联常识与修辞性表达。","relevance":"该研究将LLM作为人类被试的替代品，与大规模真实人类数据对照，评估仿真可靠性与异质性，并指出仿真失效的文化条件，直接回应了研究者对LLM仿真实验的核心关切，值得精读。","inspiration":"该方法将LLM作为独立受访者，直接与大规模人类样本的判断分布进行对比，并生成硅样本模拟群体共识，可借鉴其个体-群体双层对照设计｜可迁移到经济预期形成研究，例如调查公众对通胀或政策公告的预期分布｜研究设计：以LLM作为被试，处理为不同措辞的央行声明，结果变量为通胀预期值，用密歇根大学消费者调查的真实预期分布数据作为对照基准"}},{"id":"2505.09396","version":2,"title":"The Influence of Human-inspired Agentic Sophistication in LLM-driven Strategic Reasoners","zh_title":"人类启发的智能体复杂度对LLM驱动战略推理者的影响","abstract":"The rapid rise of large language models (LLMs) has shifted artificial intelligence (AI) research toward agentic systems, motivating the use of weaker and more flexible notions of agency. However, this shift raises key questions about the extent to which LLM-based agents replicate human strategic reasoning, particularly in game-theoretic settings. In this context, we examine the role of agentic sophistication in shaping artificial reasoners' performance by evaluating three agent designs: a simple game-theoretic model, an unstructured LLM-as-agent model, and an LLM integrated into a traditional agentic framework. Using guessing games as a testbed, we benchmarked these agents against human participants across general reasoning patterns and individual role-based objectives. Furthermore, we introduced obfuscated game scenarios to assess agents' ability to generalise beyond training distributions. Our analysis, covering over 2000 reasoning samples across 25 agent configurations, shows that human-inspired cognitive structures can enhance LLM agents' alignment with human strategic behaviour. Still, the relationship between agentic design complexity and human-likeness is non-linear, highlighting a critical dependence on underlying LLM capabilities and suggesting limits to simple architectural augmentation.","authors":["Vince Trencsenyi","Agnieszka Mensfelt","Kostas Stathis"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2025-05-14","first_seen":"2025-05-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.09396","pdf_url":"https://arxiv.org/pdf/2505.09396","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","博弈实验","人类对照"],"reason":"用LLM代理模拟人类策略推理，并与真实人类数据对照，评估对齐程度与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":51,"question":"在博弈论猜测游戏中，LLM智能体的设计复杂度（agentic sophistication）如何影响其与人类策略推理行为的对齐程度？","design":"使用三种智能体设计（简单博弈论模型EWA、非结构化LLM-as-agent、LLM集成传统智能体框架），在双人猜测游戏中与人类被试进行基准对比，并引入混淆游戏场景测试分布外泛化能力，测量推理模式与个体角色目标的对齐度。","baseline":"人类数据集，按子群体划分，包含人类在猜测游戏中的策略推理行为。","findings":"受人类启发的认知结构能增强LLM智能体与人类策略行为的一致性，但智能体设计复杂度与人类相似度之间呈非线性关系，高度依赖底层LLM能力，简单的架构增强存在局限。","reliability":"论文指出LLM的黑箱特性带来可复现性、可解释性和验证挑战，混淆场景测试旨在缓解训练数据暴露偏差，但承认LLM驱动应用仍缺乏系统严格的验证机制。","relevance":"直接回应了用LLM仿真人类策略推理的核心问题，提供了与真实人类数据的对照基准，并批判性地揭示了设计复杂度与人类相似度的非线性关系及失效条件，值得精读。","inspiration":"借鉴其用不同复杂度智能体设计与人类基准对比的框架，以及引入混淆场景测试分布外泛化的稳健性检验方法｜可迁移到资产定价实验，研究投资者在策略性猜测市场走势时的推理行为｜以LLM智能体为被试，处理为不同复杂度的智能体设计（如简单启发式、EWA模型、非结构化LLM），结果变量为价格预测偏差，用真实人类资产定价实验数据做对照"}},{"id":"2505.07457","version":1,"title":"Can Generative AI agents behave like humans? Evidence from laboratory market experiments","zh_title":"生成式AI智能体能像人类一样行为吗？来自实验室市场实验的证据","abstract":"We explore the potential of Large Language Models (LLMs) to replicate human behavior in economic market experiments. Compared to previous studies, we focus on dynamic feedback between LLM agents: the decisions of each LLM impact the market price at the current step, and so affect the decisions of the other LLMs at the next step. We compare LLM behavior to market dynamics observed in laboratory settings and assess their alignment with human participants' behavior. Our findings indicate that LLMs do not adhere strictly to rational expectations, displaying instead bounded rationality, similarly to human participants. Providing a minimal context window i.e. memory of three previous time steps, combined with a high variability setting capturing response heterogeneity, allows LLMs to replicate broad trends seen in human experiments, such as the distinction between positive and negative feedback markets. However, differences remain at a granular level--LLMs exhibit less heterogeneity in behavior than humans. These results suggest that LLMs hold promise as tools for simulating realistic human behavior in economic contexts, though further research is needed to refine their accuracy and increase behavioral diversity.","authors":["R. Maria del Rio-Chanona","Marco Pangallo","Cars Hommes"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2025-05-12","first_seen":"2025-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.07457","pdf_url":"https://arxiv.org/pdf/2505.07457","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","市场实验","人类行为对照"],"reason":"用LLM复现市场实验并与人类数据对照，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":66,"question":"大语言模型能否在动态市场实验中复现人类行为，特别是正负反馈市场中的价格动态？","design":"使用GPT-3.5和GPT-4作为智能体，模拟实验室市场实验，操控上下文窗口（记忆长度）和响应变异性（温度参数），观察市场价格动态和个体策略。","baseline":"对照Heemeijer等人(2009)等实验室市场实验的人类参与者数据，比较正负反馈市场中的价格收敛模式和波动特征。","findings":"在至少3步记忆和高响应变异性下，LLM市场能复现正负反馈市场的宏观差异，如负反馈市场快速振荡收敛、正反馈市场缓慢收敛；但LLM行为异质性低于人类，且不严格遵循理性预期，表现出有限理性。","reliability":"LLM在细粒度行为上异质性不足，与人类仍有差距；研究仅基于特定市场实验范式，泛化性待验证；模型类型和参数设置对结果敏感，需进一步优化以提升行为多样性和准确性。","relevance":"直接命中研究者关注的核心：用LLM复现经济实验并与人类基准对照，评估仿真可靠性与偏差，且包含动态交互和批判性发现，值得精读原文。","inspiration":"借鉴之处在于通过操控LLM的上下文窗口长度和响应变异性来模拟有限理性，并与人类实验基准对照，检验宏观市场动态的复现能力｜可迁移到资产定价实验，研究正负反馈机制下价格泡沫的形成与破裂｜设计一个LLM模拟的资产市场实验，将被试分为GPT-4智能体，处理为不同反馈结构（正反馈如追涨杀跌 vs. 负反馈如均值回归），结果变量为价格偏离基本面的程度和泡沫持续时间，对照真实人类实验数据（如Smith等人1988年的泡沫实验）"}},{"id":"2505.06702","version":1,"title":"Do Language Model Agents Align with Humans in Rating Visualizations? An Empirical Study","zh_title":"语言模型代理在可视化评分中与人类对齐吗？一项实证研究","abstract":"Large language models encode knowledge in various domains and demonstrate the ability to understand visualizations. They may also capture visualization design knowledge and potentially help reduce the cost of formative studies. However, it remains a question whether large language models are capable of predicting human feedback on visualizations. To investigate this question, we conducted three studies to examine whether large model-based agents can simulate human ratings in visualization tasks. The first study, replicating a published study involving human subjects, shows agents are promising in conducting human-like reasoning and rating, and its result guides the subsequent experimental design. The second study repeated six human-subject studies reported in literature on subjective ratings, but replacing human participants with agents. Consulting with five human experts, this study demonstrates that the alignment of agent ratings with human ratings positively correlates with the confidence levels of the experts before the experiments. The third study tests commonly used techniques for enhancing agents, including preprocessing visual and textual inputs, and knowledge injection. The results reveal the issues of these techniques in robustness and potential induction of biases. The three studies indicate that language model-based agents can potentially simulate human ratings in visualization experiments, provided that they are guided by high-confidence hypotheses from expert evaluators. Additionally, we demonstrate the usage scenario of swiftly evaluating prototypes with agents. We discuss insights and future directions for evaluating and improving the alignment of agent ratings with human ratings. We note that simulation may only serve as complements and cannot replace user studies.","authors":["Zekai Shao","Yi Shan","Yixuan He","Yuxuan Yao","Junhong Wang","Xiaolong","Zhang","Yu Zhang","Siming Chen"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-05-10","first_seen":"2025-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.06702","pdf_url":"https://arxiv.org/pdf/2505.06702","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类数据对照","可视化评估"],"reason":"用LLM代理模拟人类对可视化的评分，并与真实人类数据对照，评估对齐度与偏差，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":67,"question":"大语言模型代理能否在可视化评分任务中模拟人类评分？","design":"使用GPT-4V等LLM代理，复现已发表的人类被试可视化实验，让代理对可视化设计进行主观评分（如易用性、信心水平），并测量代理评分与人类评分的一致性。","baseline":"对照的真实人类数据来自已发表的六项人类被试研究，包括时间序列可视化实验等，数据来源于Open Science Framework。","findings":"LLM代理能模拟人类推理并给出类似人类的评分，但无法模拟多样化用户画像；代理与人类评分的一致性程度与专家实验前信心水平正相关。","reliability":"代理评分仅能作为人类用户研究的补充，不能替代；增强技术（如知识注入）可能引入偏差，且代理在鲁棒性上存在问题。","relevance":"该研究直接以真实人类数据为基准，评估LLM代理在可视化评分任务中的仿真可靠性，并讨论了失效条件与偏差，符合研究者对经济学实验和政策评估场景的批判性关注。","inspiration":"该方法借鉴了用LLM代理复现已发表人类实验并直接对比真实人类评分一致性的设计，以及通过专家信心水平等指标检验仿真可靠性的思路。｜可迁移到消费者对金融产品信息披露的主观评价实验，如研究简化版风险提示是否提升理解度。｜以LLM代理作为被试，施加不同格式的风险披露文本处理，测量代理对产品风险的理解评分，并与真实消费者调查数据（如CFPB的金融素养调查）进行一致性对比。"}},{"id":"2507.18639","version":1,"title":"People Are Highly Cooperative with Large Language Models, Especially When Communication Is Possible or Following Human Interaction","zh_title":"人们与大型语言模型高度合作，尤其在可沟通或继人类互动之后","abstract":"Machines driven by large language models (LLMs) have the potential to augment humans across various tasks, a development with profound implications for business settings where effective communication, collaboration, and stakeholder trust are paramount. To explore how interacting with an LLM instead of a human might shift cooperative behavior in such settings, we used the Prisoner's Dilemma game -- a surrogate of several real-world managerial and economic scenarios. In Experiment 1 (N=100), participants engaged in a thirty-round repeated game against a human, a classic bot, and an LLM (GPT, in real-time). In Experiment 2 (N=192), participants played a one-shot game against a human or an LLM, with half of them allowed to communicate with their opponent, enabling LLMs to leverage a key advantage over older-generation machines. Cooperation rates with LLMs -- while lower by approximately 10-15 percentage points compared to interactions with human opponents -- were nonetheless high. This finding was particularly notable in Experiment 2, where the psychological cost of selfish behavior was reduced. Although allowing communication about cooperation did not close the human-machine behavioral gap, it increased the likelihood of cooperation with both humans and LLMs equally (by 88%), which is particularly surprising for LLMs given their non-human nature and the assumption that people might be less receptive to cooperating with machines compared to human counterparts. Additionally, cooperation with LLMs was higher following prior interaction with humans, suggesting a spillover effect in cooperative behavior. Our findings validate the (careful) use of LLMs by businesses in settings that have a cooperative component.","authors":["Paweł Niszczota","Tomasz Grzegorczyk","Alexander Pastukhov"],"categories":["cs.HC","cs.CL","cs.CY","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2025-05-10","first_seen":"2025-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.18639","pdf_url":"https://arxiv.org/pdf/2507.18639","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","行为博弈","人机合作"],"reason":"用LLM替代人类被试进行囚徒困境博弈，并与真实人类行为对照，评估合作行为差异与…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":27,"question":"当对手是大语言模型而非人类时，人们在囚徒困境中的合作行为会发生怎样的变化？","design":"实验1（N=100）采用被试内设计，每人与人类、经典机器人、GPT实时对战各30轮重复囚徒困境；实验2（N=192）采用被试间设计，在单次囚徒困境中对手为人类或LLM，且半数被试可在决策前与对手沟通。结果变量为合作率。","baseline":"以真实人类作为对手时的合作行为作为对照基准。","findings":"与LLM的合作率虽比与人类对手低约10–15个百分点，但仍处于较高水平；允许沟通使与人类和LLM的合作率均提高88%，且与人类互动后与LLM的合作率更高，存在溢出效应。","reliability":"论文未讨论","relevance":"该研究直接以LLM替代人类被试进行囚徒困境博弈，并与真实人类行为对照，评估合作行为差异及沟通、溢出效应，高度契合研究者对LLM仿真人类行为可靠性与偏差的关注，值得精读原文。","inspiration":"采用被试内设计直接对比同一参与者在面对人类、传统机器人和LLM时的行为变化，有效控制个体差异，值得借鉴。｜可迁移至经济金融中的信任与履约行为研究，如线上借贷平台中借款人对人工审核员与AI审核员的还款承诺差异。｜以真实借款人为被试，随机分配其与人类审核员或LLM审核员沟通还款计划，测量其后续实际还款率，并以平台历史人工审核还款数据作为对照基准。"}},{"id":"2505.00036","version":1,"title":"A Framework to Assess the Persuasion Risks Large Language Model Chatbots Pose to Democratic Societies","zh_title":"评估大语言模型聊天机器人对民主社会说服风险的框架","abstract":"In recent years, significant concern has emerged regarding the potential threat that Large Language Models (LLMs) pose to democratic societies through their persuasive capabilities. We expand upon existing research by conducting two survey experiments and a real-world simulation exercise to determine whether it is more cost effective to persuade a large number of voters using LLM chatbots compared to standard political campaign practice, taking into account both the \"receive\" and \"accept\" steps in the persuasion process (Zaller 1992). These experiments improve upon previous work by assessing extended interactions between humans and LLMs (instead of using single-shot interactions) and by assessing both short- and long-run persuasive effects (rather than simply asking users to rate the persuasiveness of LLM-produced content). In two survey experiments (N = 10,417) across three distinct political domains, we find that while LLMs are about as persuasive as actual campaign ads once voters are exposed to them, political persuasion in the real-world depends on both exposure to a persuasive message and its impact conditional on exposure. Through simulations based on real-world parameters, we estimate that LLM-based persuasion costs between \\$48-\\$74 per persuaded voter compared to \\$100 for traditional campaign methods, when accounting for the costs of exposure. However, it is currently much easier to scale traditional campaign persuasion methods than LLM-based persuasion. While LLMs do not currently appear to have substantially greater potential for large-scale political persuasion than existing non-LLM methods, this may change as LLM capabilities continue to improve and it becomes easier to scalably encourage exposure to persuasive LLMs.","authors":["Zhongren Chen","Joshua Kalla","Quan Le","Shinpei Nakamura-Sakai","Jasjeet Sekhon","Ruixiao Wang"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-04-29","first_seen":"2025-04-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.00036","pdf_url":"https://arxiv.org/pdf/2505.00036","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","政治说服","人类数据对照"],"reason":"用LLM聊天机器人替代人类选民进行说服实验，有真实人类调查数据对照，涉及政治说…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":22,"question":"在政治说服中，使用LLM聊天机器人与传统竞选广告相比，在考虑曝光和接受两个步骤后，是否更具成本效益？","design":"本研究并非用LLM仿真人类被试，而是通过两项调查实验和真实世界模拟，将人类参与者随机分配到安慰剂、真人视频说服、AI聊天机器人（自称人类）和AI聊天机器人（自称AI）四种条件，测量其对移民政策的支持态度变化，并基于真实参数模拟成本效益。","baseline":"真人视频说服条件（3分钟教师分享个人理由的视频）和传统竞选广告方法（如电视广告）作为对照基准。","findings":"在调查实验中，LLM聊天机器人与真人视频的说服效果在短期和五周后均无显著差异；但模拟显示，考虑曝光成本后，LLM说服每位选民的成本为48-74美元，低于传统方法的100美元，然而目前传统方法更易规模化。","reliability":"论文指出当前LLM在规模化政治说服方面并不比现有非LLM方法有显著优势，但随LLM能力提升和更容易鼓励曝光，情况可能改变；研究局限包括真人说服条件为3分钟视频，长于典型广告，且未充分解决曝光环节的规模化难题。","relevance":"该研究直接使用LLM与人类进行交互式说服实验，并与真实人类说服效果对照，涉及政治领域成本效益评估，同时讨论了规模化限制，高度契合研究者对LLM仿真人类行为、基准对照和失效条件的兴趣，值得精读原文。","inspiration":"该方法将LLM聊天机器人作为交互式说服工具，与真人视频说服和传统广告进行随机对照实验，并追踪短期与五周后的态度变化，值得借鉴其多臂对照和纵向测量设计｜可迁移到消费者金融决策场景，如评估LLM理财顾问对投资偏好或退休储蓄选择的影响｜以真实投资者为被试，随机分配至LLM聊天机器人理财建议、人类理财顾问视频和纯文本说明书三组，结果变量为风险资产配置比例和储蓄率变化，以历史调查数据或银行实际客户行为作为基准对照"}},{"id":"2504.19940","version":2,"title":"Assessing the Potential of Generative Agents in Crowdsourced Fact-Checking","zh_title":"评估生成式智能体在众包事实核查中的潜力","abstract":"The growing spread of online misinformation has created an urgent need for scalable, reliable fact-checking solutions. Crowdsourced fact-checking - where non-experts evaluate claim veracity - offers a cost-effective alternative to expert verification, despite concerns about variability in quality and bias. Encouraged by promising results in certain contexts, major platforms such as X (formerly Twitter), Facebook, and Instagram have begun shifting from centralized moderation to decentralized, crowd-based approaches. In parallel, advances in Large Language Models (LLMs) have shown strong performance across core fact-checking tasks, including claim detection and evidence evaluation. However, their potential role in crowdsourced workflows remains unexplored. This paper investigates whether LLM-powered generative agents - autonomous entities that emulate human behavior and decision-making - can meaningfully contribute to fact-checking tasks traditionally reserved for human crowds. Using the protocol of La Barbera et al. (2024), we simulate crowds of generative agents with diverse demographic and ideological profiles. Agents retrieve evidence, assess claims along multiple quality dimensions, and issue final veracity judgments. Our results show that agent crowds outperform human crowds in truthfulness classification, exhibit higher internal consistency, and show reduced susceptibility to social and cognitive biases. Compared to humans, agents rely more systematically on informative criteria such as Accuracy, Precision, and Informativeness, suggesting a more structured decision-making process. Overall, our findings highlight the potential of generative agents as scalable, consistent, and less biased contributors to crowd-based fact-checking systems.","authors":["Luigia Costabile","Gian Marco Orlando","Valerio La Gatta","Vincenzo Moscato"],"categories":["cs.CL","cs.AI","cs.MA"],"primary_category":"cs.CL","announce_type":"new","date":"2025-04-24","first_seen":"2025-04-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.19940","pdf_url":"https://arxiv.org/pdf/2504.19940","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","众包事实核查","人类行为对照"],"reason":"用LLM智能体模拟人群事实核查，与真实人类数据对照，评估偏差与一致性，直接命中…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":31,"question":"生成式智能体在众包事实核查中的表现能否达到或超过人类众包？","design":"使用LLM驱动的生成式智能体模拟具有不同人口统计和意识形态特征的人群，复现La Barbera et al. (2024)的实验协议：智能体先选择证据，再填写结构化问卷对声明进行多维度评分（准确性、无偏性等），最后给出真实性判断。","baseline":"La Barbera et al. (2024)中人类众包参与者对相同声明的标注数据。","findings":"智能体众包在真实性分类上优于人类众包，表现出更高的内部一致性，且受社会和认知偏差的影响更小。智能体更系统地依赖准确性、精确性和信息量等标准，决策过程更结构化。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体模拟人类众包事实核查，并与真实人类数据对照，评估偏差与一致性，高度契合研究者对LLM人类仿真实验的关注，值得精读原文。","inspiration":"该方法借鉴了用LLM智能体复现人类实验协议并直接对比真实人类数据的做法，通过让智能体遵循相同的任务流程（选择证据、填写问卷、给出判断）来评估仿真偏差与一致性。｜可迁移到经济金融中的信贷审批歧视研究，模拟不同人口特征的贷款审批决策。｜设计：用LLM智能体模拟不同种族、性别、收入的贷款申请人，处理为呈现相同的贷款申请材料，结果变量为审批结果和理由，对照真实银行信贷审批数据或人类实验数据。"}},{"id":"2504.08260","version":2,"title":"Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare","zh_title":"评估大语言模型在医疗意见与决策调查中的偏差","abstract":"Generative agents have been increasingly used to simulate human behaviour in silico, driven by large language models (LLMs). These simulacra serve as sandboxes for studying human behaviour without compromising privacy or safety. However, it remains unclear whether such agents can truly represent real individuals. This work compares survey data from the Understanding America Study (UAS) on healthcare decision-making with simulated responses from generative agents. Using demographic-based prompt engineering, we create digital twins of survey respondents and analyse how well different LLMs reproduce real-world behaviours. Our findings show that some LLMs fail to reflect realistic decision-making, such as predicting universal vaccine acceptance. However, Llama 3 captures variations across race and Income more accurately but also introduces biases not present in the UAS data. This study highlights the potential of generative agents for behavioural research while underscoring the risks of bias from both LLMs and prompting strategies.","authors":["Yonchanok Khaokaew","Flora D. Salim","Andreas Züfle","Hao Xue","Taylor Anderson","C. Raina MacIntyre","Matthew Scotch","David J Heslop"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-04-11","first_seen":"2025-04-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.08260","pdf_url":"https://arxiv.org/pdf/2504.08260","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类数据对照","医疗决策偏差"],"reason":"用LLM仿真医疗决策，与真实调查数据对照，评估偏差，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"LLM能否有效模拟医疗决策（如疫苗接种意愿），以及在不同人口群体中会产生哪些偏差？","design":"使用LLM（如Llama 3等）基于人口统计属性（年龄、性别、收入、种族、教育、担忧程度）构建数字孪生，在不同疫情场景下提示模型回答疫苗接种意愿问题，比较模型输出与真实调查数据。","baseline":"理解美国研究（UAS）中关于COVID-19疫苗接种意愿的调查数据，涵盖2020年3月至2021年1月的多波次问卷。","findings":"部分LLM未能反映现实决策，例如预测普遍接受疫苗；Llama 3更准确地捕捉了种族和收入差异，但也引入了UAS数据中不存在的新偏差。","reliability":"论文指出LLM的预测效果取决于提供的上下文信息类型和数量，预训练数据和提示策略可能导致与人类偏差不同的偏差，且LLM可能仅反映训练数据中的统计模式而非真实决策过程。","relevance":"该研究直接使用LLM仿真医疗决策并与真实人类调查数据对照，评估偏差，完全符合研究者对LLM人类仿真实验、经济学/政策场景及可靠性批判的关注，值得精读原文。","inspiration":"借鉴其利用人口统计属性构建数字孪生并与多波次真实调查数据对照的仿真设计，评估LLM在决策模拟中的偏差｜可迁移到医疗健康政策的经济评估场景，如不同人群对医疗保险选择或健康储蓄账户的偏好差异｜以LLM为被试，基于收入、年龄、健康风险等属性生成数字孪生，提示其在不同保费和补贴政策下选择保险方案，结果变量为保险选择概率，用真实调查数据（如MEPS）作为基准对照"}},{"id":"2504.05862","version":2,"title":"Are Generative AI Agents Effective Personalized Financial Advisors?","zh_title":"生成式AI代理能成为有效的个性化理财顾问吗？","abstract":"Large language model-based agents are becoming increasingly popular as a low-cost mechanism to provide personalized, conversational advice, and have demonstrated impressive capabilities in relatively simple scenarios, such as movie recommendations. But how do these agents perform in complex high-stakes domains, where domain expertise is essential and mistakes carry substantial risk? This paper investigates the effectiveness of LLM-advisors in the finance domain, focusing on three distinct challenges: (1) eliciting user preferences when users themselves may be unsure of their needs, (2) providing personalized guidance for diverse investment preferences, and (3) leveraging advisor personality to build relationships and foster trust. Via a lab-based user study with 64 participants, we show that LLM-advisors often match human advisor performance when eliciting preferences, although they can struggle to resolve conflicting user needs. When providing personalized advice, the LLM was able to positively influence user behavior, but demonstrated clear failure modes. Our results show that accurate preference elicitation is key, otherwise, the LLM-advisor has little impact, or can even direct the investor toward unsuitable assets. More worryingly, users appear insensitive to the quality of advice being given, or worse these can have an inverse relationship. Indeed, users reported a preference for and increased satisfaction as well as emotional trust with LLMs adopting an extroverted persona, even though those agents provided worse advice.","authors":["Takehiro Takayanagi","Kiyoshi Izumi","Javier Sanz-Cruzado","Richard McCreadie","Iadh Ounis"],"categories":["cs.AI","cs.CL","cs.HC","cs.IR","q-fin.CP"],"primary_category":"cs.AI","announce_type":"new","date":"2025-04-08","first_seen":"2025-04-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.05862","pdf_url":"https://arxiv.org/pdf/2504.05862","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人机对照实验","金融行为"],"reason":"用LLM模拟人类理财顾问，与真人对照，评估效果与失效模式，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:44","error":null,"has_summary":true,"summary":{"generated_at":"2025-04-08","rank":5,"question":"LLM作为个性化理财顾问，在偏好获取、个性化建议和人格影响方面表现如何？","design":"实验室用户研究，64名参与者扮演给定投资者画像，与LLM顾问进行两阶段对话：偏好获取和资产建议讨论，比较个性化vs非个性化顾问及不同人格的顾问，测量决策质量、用户满意度和信任。","baseline":"人类顾问在偏好获取阶段的表现作为对照基准。","findings":"LLM顾问在偏好获取上常能与人类顾问匹敌，但难以解决用户需求冲突；个性化建议能正向影响用户行为，但存在明显失效模式，且用户对建议质量不敏感，甚至偏好外向人格的顾问，尽管其建议更差。","reliability":"论文指出准确偏好获取是关键，否则LLM顾问影响甚微或引导投资者选择不合适资产；用户对建议质量不敏感，甚至出现反向关系，外向人格虽提升满意度和情感信任但建议质量更低。","relevance":"该研究直接以真人实验对照LLM在金融建议中的表现，揭示了偏好获取失效和用户对建议质量不敏感等关键偏差，对关注经济学实验和决策仿真的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其分阶段对话设计和人格操纵处理，可迁移到信贷审批或保险推荐等金融场景，设计以LLM为被试、操纵建议人格或个性化程度、测量用户决策偏差和信任，并以人类顾问或历史决策数据为对照。"}},{"id":"2503.22726","version":1,"title":"InfoBid: A Simulation Framework for Studying Information Disclosure in Auctions with Large Language Model-based Agents","zh_title":"InfoBid：基于大语言模型代理研究拍卖中信息披露的仿真框架","abstract":"In online advertising systems, publishers often face a trade-off in information disclosure strategies: while disclosing more information can enhance efficiency by enabling optimal allocation of ad impressions, it may lose revenue potential by decreasing uncertainty among competing advertisers. Similar to other challenges in market design, understanding this trade-off is constrained by limited access to real-world data, leading researchers and practitioners to turn to simulation frameworks. The recent emergence of large language models (LLMs) offers a novel approach to simulations, providing human-like reasoning and adaptability without necessarily relying on explicit assumptions about agent behavior modeling. Despite their potential, existing frameworks have yet to integrate LLM-based agents for studying information asymmetry and signaling strategies, particularly in the context of auctions. To address this gap, we introduce InfoBid, a flexible simulation framework that leverages LLM agents to examine the effects of information disclosure strategies in multi-agent auction settings. Using GPT-4o, we implemented simulations of second-price auctions with diverse information schemas. The results reveal key insights into how signaling influences strategic behavior and auction outcomes, which align with both economic and social learning theories. Through InfoBid, we hope to foster the use of LLMs as proxies for human economic and social agents in empirical studies, enhancing our understanding of their capabilities and limitations. This work bridges the gap between theoretical market designs and practical applications, advancing research in market simulations, information design, and agent-based reasoning while offering a valuable tool for exploring the dynamics of digital economies.","authors":["Yue Yin"],"categories":["cs.GT","cs.CL","cs.HC","cs.MA","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2025-03-26","first_seen":"2025-03-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.22726","pdf_url":"https://arxiv.org/pdf/2503.22726","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","拍卖实验","经济行为"],"reason":"用LLM代理模拟拍卖中的人类经济行为，与理论对照，涉及经济学实验场景并讨论局限…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":64,"question":"在拍卖中，信息披露策略如何影响基于LLM的智能体的出价行为与拍卖结果？","design":"使用GPT-4o构建LLM智能体模拟竞拍者，在第二价格拍卖中施加不同的信息披露方案（信号），测量出价、收入、效率等拍卖结果。","baseline":"无对照","findings":"信号显著影响LLM智能体的策略行为与拍卖结果，且观察到的行为模式与经济理论和社会学习理论一致。","reliability":"论文未讨论","relevance":"高度相关：用LLM代理模拟拍卖中的人类经济行为，研究信息不对称下的策略互动，并讨论LLM作为人类代理的潜力与局限，直接命中研究者关心的经济学实验仿真与可靠性评估。","inspiration":"该方法通过向LLM智能体提供不同信号来模拟信息披露，可借鉴其处理信息干预的方式，用于研究经济决策中的信息效应｜可迁移到资产定价实验中，研究公开与私有信号如何影响投资者的价格预期与交易行为｜设计：以LLM作为被试，随机分配接收不同精度的资产价值信号，测量其报价与交易量，并与历史实验市场数据或理性预期均衡基准对照"}},{"id":"2503.11531","version":1,"title":"Potential of large language model-powered nudges for promoting daily water and energy conservation","zh_title":"大语言模型驱动的助推在促进日常节水节能中的潜力","abstract":"The increasing amount of pressure related to water and energy shortages has increased the urgency of cultivating individual conservation behaviors. While the concept of nudging, i.e., providing usage-based feedback, has shown promise in encouraging conservation behaviors, its efficacy is often constrained by the lack of targeted and actionable content. This study investigates the impact of the use of large language models (LLMs) to provide tailored conservation suggestions for conservation intentions and their rationale. Through a survey experiment with 1,515 university participants, we compare three virtual nudging scenarios: no nudging, traditional nudging with usage statistics, and LLM-powered nudging with usage statistics and personalized conservation suggestions. The results of statistical analyses and causal forest modeling reveal that nudging led to an increase in conservation intentions among 86.9%-98.0% of the participants. LLM-powered nudging achieved a maximum increase of 18.0% in conservation intentions, surpassing traditional nudging by 88.6%. Furthermore, structural equation modeling results reveal that exposure to LLM-powered nudges enhances self-efficacy and outcome expectations while diminishing dependence on social norms, thereby increasing intrinsic motivation to conserve. These findings highlight the transformative potential of LLMs in promoting individual water and energy conservation, representing a new frontier in the design of sustainable behavioral interventions and resource management.","authors":["Zonghan Li","Song Tong","Yi Liu","Kaiping Peng","Chunyan Wang"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-03-14","first_seen":"2025-03-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.11531","pdf_url":"https://arxiv.org/pdf/2503.11531","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","行为干预","节能实验"],"reason":"用LLM生成个性化节能建议，通过调查实验与人类对照，评估对行为意图的影响，属于…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:36","error":null,"has_summary":true,"summary":{"generated_at":"2025-03-14","rank":2,"question":"基于LLM的个性化助推与传统使用统计助推相比，能否更有效地提升个体的节水节能意图？","design":"本研究并非用LLM模拟人类被试，而是通过随机对照调查实验，将1515名大学生随机分为三组，分别接受无助推、传统使用统计助推、LLM生成个性化建议的助推，测量其节水节能意图的变化。","baseline":"无助推组和传统使用统计助推组作为对照，均为真实人类被试。","findings":"LLM助推使86.9%-98.0%的参与者节水节能意图提升，最大增幅达18.0%，效果比传统助推高88.6%。结构方程模型显示，LLM助推通过增强自我效能和结果预期、降低社会规范依赖，提升内在动机。","reliability":"论文未讨论","relevance":"该研究用LLM生成个性化干预内容，通过真实人类调查实验评估行为意图变化，属于LLM辅助行为干预效果评估，与研究者关注的LLM仿真实验和经济学实验场景高度相关，值得细读其因果推断设计和心理机制分析。","inspiration":"借鉴其随机分组和因果森林方法评估异质性处理效应，可迁移到消费者节能行为干预政策评估场景。｜设计一个实验：以居民为被试，随机分配接收传统节能建议或LLM个性化建议，结果变量为实际用电量变化，对照真实智能电表数据。"}},{"id":"2503.10248","version":1,"title":"LLM Agents Display Human Biases but Exhibit Distinct Learning Patterns","zh_title":"LLM智能体表现出人类偏见但学习模式不同","abstract":"We investigate the choice patterns of Large Language Models (LLMs) in the context of Decisions from Experience tasks that involve repeated choice and learning from feedback, and compare their behavior to human participants. We find that on the aggregate, LLMs appear to display behavioral biases similar to humans: both exhibit underweighting rare events and correlation effects. However, more nuanced analyses of the choice patterns reveal that this happens for very different reasons. LLMs exhibit strong recency biases, unlike humans, who appear to respond in more sophisticated ways. While these different processes may lead to similar behavior on average, choice patterns contingent on recent events differ vastly between the two groups. Specifically, phenomena such as ``surprise triggers change\" and the ``wavy recency effect of rare events\" are robustly observed in humans, but entirely absent in LLMs. Our findings provide insights into the limitations of using LLMs to simulate and predict humans in learning environments and highlight the need for refined analyses of their behavior when investigating whether they replicate human decision making tendencies.","authors":["Idan Horowitz","Ori Plonsky"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-03-13","first_seen":"2025-03-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.10248","pdf_url":"https://arxiv.org/pdf/2503.10248","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","决策实验","人类对照"],"reason":"用LLM复现人类决策实验，与真实人类数据对照，并指出仿真失效条件，高度相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":68,"question":"LLM在基于经验的决策任务中是否表现出与人类相似的行为偏差和学习模式？","design":"将多个LLM作为被试，完成四项重复二元选择决策任务（100试次），通过反馈学习收益分布，测量选择行为；同时操纵历史提供方式和模型温度参数。","baseline":"真实人类被试在相同决策任务中的选择数据，作为行为对比基准。","findings":"总体上看，LLM表现出与人类相似的低估稀有事件和相关效应，但过程不同：LLM有强烈的近因偏差，而人类对稀有事件有“惊讶触发改变”和“波浪式近因效应”，这些模式在LLM中完全缺失。","reliability":"LLM在聚合层面可能模拟人类偏差，但内在学习过程截然不同，基于近期事件的精细分析显示其无法复现人类特有的动态反应模式，提示在涉及反馈学习的场景中用LLM仿真人类决策存在局限。","relevance":"该研究直接对比LLM与人类在经济学实验范式下的决策，揭示仿真在表面聚合指标下有效但内在机制失效，高度契合研究者对LLM仿真可靠性及失效条件的批判性关注，值得精读。","inspiration":"该方法通过重复二元选择任务和试次级反馈学习数据，精细对比LLM与人类在聚合与过程层面的行为差异，值得借鉴｜可迁移到资产定价实验中的罕见事件学习，如研究投资者如何更新对崩盘风险的信念｜以LLM为被试，操纵历史收益序列呈现方式，测量其风险资产配置比例，并与真实投资者在类似实验中的选择数据对照"}},{"id":"2503.09639","version":4,"title":"Can A Society of Generative Agents Simulate Human Behavior and Inform Public Health Policy? A Case Study on Vaccine Hesitancy","zh_title":"生成式智能体社会能否模拟人类行为并为公共卫生政策提供信息？以疫苗犹豫为例","abstract":"Can we simulate a sandbox society with generative agents to model human behavior, thereby reducing the over-reliance on real human trials for assessing public policies? In this work, we investigate the feasibility of simulating health-related decision-making, using vaccine hesitancy, defined as the delay in acceptance or refusal of vaccines despite the availability of vaccination services (MacDonald, 2015), as a case study. To this end, we introduce the VacSim framework with 100 generative agents powered by Large Language Models (LLMs). VacSim simulates vaccine policy outcomes with the following steps: 1) instantiate a population of agents with demographics based on census data; 2) connect the agents via a social network and model vaccine attitudes as a function of social dynamics and disease-related information; 3) design and evaluate various public health interventions aimed at mitigating vaccine hesitancy. To align with real-world results, we also introduce simulation warmup and attitude modulation to adjust agents' attitudes. We propose a series of evaluations to assess the reliability of various LLM simulations. Experiments indicate that models like Llama and Qwen can simulate aspects of human behavior but also highlight real-world alignment challenges, such as inconsistent responses with demographic profiles. This early exploration of LLM-driven simulations is not meant to serve as definitive policy guidance; instead, it serves as a call for action to examine social simulation for policy development.","authors":["Abe Bohan Hou","Hongru Du","Yichen Wang","Jingyu Zhang","Zixiao Wang","Paul Pu Liang","Daniel Khashabi","Lauren Gardner","Tianxing He"],"categories":["cs.MA","cs.AI","cs.CL","cs.CY","cs.HC"],"primary_category":"cs.MA","announce_type":"new","date":"2025-03-12","first_seen":"2025-03-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.09639","pdf_url":"https://arxiv.org/pdf/2503.09639","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM人类仿真","疫苗犹豫","政策评估"],"reason":"用LLM代理模拟疫苗犹豫行为并与真实人口数据对照，评估仿真可靠性，直接命中核心…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":32,"question":"能否用生成式智能体构建的沙盒社会模拟人类疫苗犹豫行为，从而为公共卫生政策提供参考？","design":"使用100个由LLM驱动的生成式智能体，基于人口普查数据赋予人口学特征，通过社交网络和新闻信息模拟疫苗态度变化，并设计不同公共卫生干预措施（如政策）来观察态度轨迹。","baseline":"Nguyen et al. (2022) 的美国COVID-19疫苗犹豫调查数据，包含1300万份真实人类回答，用于校准智能体的人口学分布。","findings":"Llama和Qwen等模型能模拟人类行为的某些方面，但存在与现实对齐的挑战，例如智能体的回答与人口学特征不一致。","reliability":"论文承认仿真存在现实对齐挑战，如回答与人口学特征不一致，并指出当前探索不能作为政策指导，仅呼吁进一步研究。","relevance":"该研究直接以LLM代理模拟疫苗犹豫行为，并与真实人口调查数据对照，评估仿真可靠性，完全命中研究者对经济学实验和政策评估场景的兴趣，值得精读原文。","inspiration":"借鉴其利用人口普查数据校准智能体人口学特征并与大规模真实调查数据对照的仿真设计｜可迁移到政策公告对消费者通胀预期形成的实验，如模拟不同央行沟通策略对预期的影响｜以LLM智能体为被试，按收入、教育等特征分层，施加不同措辞的政策公告作为处理，测量预期通胀率的变化，用密歇根大学消费者调查的真实预期数据做对照"}},{"id":"2503.07510","version":1,"title":"Sometimes the Model doth Preach: Quantifying Religious Bias in Open LLMs through Demographic Analysis in Asian Nations","zh_title":"有时模型在布道：通过亚洲国家人口统计分析量化开放LLM中的宗教偏见","abstract":"Large Language Models (LLMs) are capable of generating opinions and propagating bias unknowingly, originating from unrepresentative and non-diverse data collection. Prior research has analysed these opinions with respect to the West, particularly the United States. However, insights thus produced may not be generalized in non-Western populations. With the widespread usage of LLM systems by users across several different walks of life, the cultural sensitivity of each generated output is of crucial interest. Our work proposes a novel method that quantitatively analyzes the opinions generated by LLMs, improving on previous work with regards to extracting the social demographics of the models. Our method measures the distance from an LLM's response to survey respondents, through Hamming Distance, to infer the demographic characteristics reflected in the model's outputs. We evaluate modern, open LLMs such as Llama and Mistral on surveys conducted in various global south countries, with a focus on India and other Asian nations, specifically assessing the model's performance on surveys related to religious tolerance and identity. Our analysis reveals that most open LLMs match a single homogeneous profile, varying across different countries/territories, which in turn raises questions about the risks of LLMs promoting a hegemonic worldview, and undermining perspectives of different minorities. Our framework may also be useful for future research investigating the complex intersection between training data, model architecture, and the resulting biases reflected in LLM outputs, particularly concerning sensitive topics like religious tolerance and identity.","authors":["Hari Shankar","Vedanta S P","Tejas Cavale","Ponnurangam Kumaraguru","Abhijnan Chakraborty"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"new","date":"2025-03-10","first_seen":"2025-03-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.07510","pdf_url":"https://arxiv.org/pdf/2503.07510","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","宗教偏见","人类数据对照"],"reason":"用LLM复现调查回答并与真实人类数据对照，评估宗教偏见，揭示仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":42,"question":"开放大语言模型在回答宗教相关调查时，反映了怎样的社会人口特征和宗教偏见？","design":"使用Llama、Mistral等开源LLM，以零样本方式回答皮尤研究中心在印度及东亚、东南亚国家进行的宗教宽容与认同调查问卷，通过汉明距离计算模型回答与真实受访者回答的匹配度，推断模型所反映的人口统计特征。","baseline":"皮尤研究中心在印度、日本、香港、韩国、台湾、越南、印尼、马来西亚、新加坡、斯里兰卡、泰国等亚洲国家/地区收集的真实调查数据，包含受访者的宗教、性别、年龄、教育等人口统计变量。","findings":"大多数开源LLM的回答与单一同质化的人口统计画像相匹配，且该画像在不同国家/地区间存在差异；通过提示词指示模型扮演特定群体并未显著改变其回答画像。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查数据为基准，评估LLM在宗教宽容等敏感话题上的回答偏差，揭示了模型输出同质化及可能强化霸权世界观的失效模式，高度契合研究者对LLM仿真可靠性及失效条件的关注，值得阅读原文。","inspiration":"借鉴其使用真实调查数据作为基准，通过汉明距离量化LLM回答与人类群体回答的匹配度，并推断模型隐含的人口统计画像的方法｜可迁移到信贷审批中的宗教或种族偏见检测，例如评估LLM在模拟贷款决策时是否系统性地偏向或歧视特定宗教群体｜以LLM作为被试，向其呈现不同宗教背景的贷款申请人资料，要求做出批准/拒绝决策，结果变量为批准率差异，并以真实银行信贷审批数据或审计研究结果作为对照基准"}},{"id":"2503.05529","version":1,"title":"PoSSUM: A Protocol for Surveying Social-media Users with Multimodal LLMs","zh_title":"PoSSUM：一种利用多模态大语言模型调查社交媒体用户的协议","abstract":"This paper introduces PoSSUM, an open-source protocol for unobtrusive polling of social-media users via multimodal Large Language Models (LLMs). PoSSUM leverages users' real-time posts, images, and other digital traces to create silicon samples that capture information not present in the LLM's training data. To obtain representative estimates, PoSSUM employs Multilevel Regression and Post-Stratification (MrP) with structured priors to counteract the observable selection biases of social-media platforms. The protocol is validated during the 2024 U.S. Presidential Election, for which five PoSSUM polls were conducted and published on GitHub and X. In the final poll, fielded October 17-26 with a synthetic sample of 1,054 X users, PoSSUM accurately predicted the outcomes in 50 of 51 states and assigned the Republican candidate a win probability of 0.65. Notably, it also exhibited lower state-level bias than most established pollsters. These results demonstrate PoSSUM's potential as a fully automated, unobtrusive alternative to traditional survey methods.","authors":["Roberto Cerina"],"categories":["stat.AP","cs.SI"],"primary_category":"stat.AP","announce_type":"new","date":"2025-03-07","first_seen":"2025-03-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.05529","pdf_url":"https://arxiv.org/pdf/2503.05529","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","选举预测","人类行为对照"],"reason":"用LLM模拟社交媒体用户投票行为，并与真实选举结果对照，属于人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:59","error":null,"has_summary":true,"summary":{"generated_at":"2025-03-07","rank":9,"question":"能否利用多模态大语言模型和社交媒体数字痕迹，构建无侵扰的民意调查方法，并准确预测选举结果？","design":"PoSSUM协议：使用多模态LLM基于X用户的实时帖子、图像等数字痕迹生成合成样本（硅样本），通过多水平回归与事后分层（MrP）校正平台选择偏差，测量投票意向。在2024年美国总统选举中进行了5次民意调查，最终轮使用1054名X用户的合成样本。","baseline":"2024年美国总统选举的真实结果（各州胜负及全国胜率）。","findings":"最终轮预测准确预测了51个州中50个的结果，并赋予共和党候选人0.65的胜率；其州级偏差低于大多数传统民调机构。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：该研究直接用LLM生成合成样本模拟人类投票行为，并以真实选举结果为对照，验证了LLM仿真在选举预测中的有效性，符合研究者对经济学实验和政策评估场景的关注。值得精读原文以了解MrP校正偏差的具体方法及协议细节。","inspiration":"借鉴该方法利用多模态LLM从社交媒体数字痕迹生成合成样本，并通过多水平回归与事后分层校正选择偏差，以低成本、无侵扰方式测量群体态度｜可迁移至政策公告的预期形成研究，如央行沟通对通胀预期的影响，或财政刺激对消费者信心的影响｜以X平台用户为合成样本来源，用多模态LLM提取用户对政策公告的态度，处理为不同措辞或发布渠道的公告版本，结果变量为通胀预期指数，以真实消费者调查数据（如密歇根大学消费者信心指数）作为对照基准"}},{"id":"2502.16280","version":1,"title":"Human Preferences in Large Language Model Latent Space: A Technical Analysis on the Reliability of Synthetic Data in Voting Outcome Prediction","zh_title":"大语言模型潜在空间中的人类偏好：合成数据在投票结果预测中可靠性的技术分析","abstract":"Generative AI (GenAI) is increasingly used in survey contexts to simulate human preferences. While many research endeavors evaluate the quality of synthetic GenAI data by comparing model-generated responses to gold-standard survey results, fundamental questions about the validity and reliability of using LLMs as substitutes for human respondents remain. Our study provides a technical analysis of how demographic attributes and prompt variations influence latent opinion mappings in large language models (LLMs) and evaluates their suitability for survey-based predictions. Using 14 different models, we find that LLM-generated data fails to replicate the variance observed in real-world human responses, particularly across demographic subgroups. In the political space, persona-to-party mappings exhibit limited differentiation, resulting in synthetic data that lacks the nuanced distribution of opinions found in survey data. Moreover, we show that prompt sensitivity can significantly alter outputs for some models, further undermining the stability and predictiveness of LLM-based simulations. As a key contribution, we adapt a probe-based methodology that reveals how LLMs encode political affiliations in their latent space, exposing the systematic distortions introduced by these models. Our findings highlight critical limitations in AI-generated survey data, urging caution in its use for public opinion research, social science experimentation, and computational behavioral modeling.","authors":["Sarah Ball","Simeon Allmendinger","Frauke Kreuter","Niklas Kühl"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2025-02-22","first_seen":"2025-02-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.16280","pdf_url":"https://arxiv.org/pdf/2502.16280","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B3","B4"],"tags":["LLM人类仿真","合成数据可靠性","政治偏好预测"],"reason":"直接评估LLM仿真人类投票偏好的可靠性与偏差，有真实调查数据对照，并揭示失效条…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":61,"question":"LLM生成的合成数据在多大程度上能复现真实人类调查回答的分布，以及提示不稳定性如何在模型的潜在空间中体现？","design":"使用14个白盒LLM，基于德国选举研究（GLES）构建包含年龄、性别、教育、收入、就业、政治倾向、东西德等人口属性的理论驱动型人设提示，让模型预测投票选择；同时通过改写提示考察提示敏感性，并利用Wahl-o-Mat数据训练探针分析模型潜在空间中的政治映射。","baseline":"以2021年德国纵向选举研究（GLES）的加权横截面调查数据作为真实人类投票行为的对照基准。","findings":"LLM生成数据无法复现真实人类回答的方差，尤其在人口子群中，人设到政党的映射区分度低，缺乏真实调查中的细致意见分布；提示敏感性会显著改变部分模型的输出，进一步削弱仿真稳定性，且在某些模型中潜在空间熵值与提示敏感性相关。","reliability":"论文指出LLM仿真在人口子群方差复现、意见分布细致度和提示稳定性方面存在根本局限，警告在公共舆论研究、社会科学实验和计算行为建模中使用合成数据需谨慎，但未讨论模型选择、语言文化差异等外部效度限制。","relevance":"该研究直接评估LLM替代人类被试进行投票预测的可靠性与偏差，有真实调查数据对照，并揭示了仿真在方差复现和提示敏感性上的失效条件，高度契合研究者对经济学实验和政策评估场景中仿真批判性分析的兴趣，值得精读原文。","inspiration":"该方法通过理论驱动构建人口属性人设提示并系统改写提示来检验仿真稳定性，为经济实验中的处理稳健性检验提供了可借鉴的测量范式｜可迁移至政策评估中的预期形成研究，例如考察不同人口子群对央行通胀预测公告的反应异质性｜以LLM模拟不同收入与教育水平的个体，处理为改写后的通胀预测措辞，结果变量为通胀预期调整幅度，以真实消费者预期调查数据作为对照基准"}},{"id":"2502.15800","version":3,"title":"LLM Agents Do Not Replicate Human Market Traders: Evidence From Experimental Finance","zh_title":"LLM代理无法复现人类市场交易者：来自实验金融的证据","abstract":"This paper explores how Large Language Models (LLMs) behave in a classic experimental finance paradigm widely known for eliciting bubbles and crashes in human participants. We adapt an established trading design, where traders buy and sell a risky asset with a known fundamental value, and introduce several LLM-based agents, both in single-model markets (all traders are instances of the same LLM) and in mixed-model \"battle royale\" settings (multiple LLMs competing in the same market). Our findings reveal that LLMs generally exhibit a \"textbook-rational\" approach, pricing the asset near its fundamental value, and show only a muted tendency toward bubble formation. Further analyses indicate that LLM-based agents display less trading strategy variance in contrast to humans. Taken together, these results highlight the risk of relying on LLM-only data to replicate human-driven market phenomena, as key behavioral features, such as large emergent bubbles, were not robustly reproduced. While LLMs clearly possess the capacity for strategic decision-making, their relative consistency and rationality suggest that they do not accurately mimic human market dynamics.","authors":["Thomas Henning","Siddhartha M. Ojha","Ross Spoon","Jiatong Han","Colin F. Camerer"],"categories":["q-fin.TR"],"primary_category":"q-fin.TR","announce_type":"new","date":"2025-02-18","first_seen":"2025-02-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.15800","pdf_url":"https://arxiv.org/pdf/2502.15800","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","实验金融","人类行为对照"],"reason":"用LLM代理模拟人类交易实验，与真实人类数据对照，发现LLM未能复现泡沫，批判…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"LLM代理在经典实验金融范式中能否复现人类交易者产生的资产泡沫与市场动态？","design":"采用Smith等人(2014)的固定基本面价值资产交易实验，在oTree平台上运行30期开放叫价市场。LLM代理分为单一模型同质市场和多个LLM模型混合的“大逃杀”市场，每轮提交买卖订单及价格预期，并允许通过“洞察”和“思考”文本实现跨轮记忆与思维链推理。结果变量为交易价格偏离基本面价值的程度、泡沫生成、交易策略方差等。","baseline":"对照真实人类被试在相同实验设计下的交易数据，人类市场一致产生显著价格泡沫和偏离基本面。","findings":"LLM代理普遍表现出“教科书式理性”，定价接近基本面价值，仅呈现微弱的泡沫倾向；其交易策略方差更低，更依赖基本面而非启发式策略，与人类行为存在系统性差异。","reliability":"论文指出，开箱即用的LLM未能复现人类市场中的大型泡沫等关键行为特征，表明仅依赖LLM数据复制人类驱动的市场现象存在风险，但未深入讨论LLM代理在何种条件下可能失效。","relevance":"该研究直接以真实人类实验为基准，检验LLM代理在金融实验中的行为复现能力，并得出批判性结论：LLM未能模拟人类市场泡沫，对评估LLM仿真可靠性及失效条件具有重要参考价值，值得精读原文。","inspiration":"该方法通过oTree平台实现多期开放叫价市场，并让LLM代理提交订单与价格预期，同时利用“洞察”和“思考”文本实现跨轮记忆与思维链推理，为经济实验的自动化仿真提供了可借鉴的设计框架｜可迁移至资产定价实验，用于检验不同信息结构或交易机制下LLM代理能否复现人类的价格泡沫与过度反应｜可设计一个资产定价实验，以LLM代理为被试，施加不同信息透明度处理，测量交易价格偏离基本面的程度，并与真实人类实验数据对照，评估LLM在模拟市场非理性行为时的有效性"}},{"id":"2502.08691","version":2,"title":"AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society","zh_title":"AgentSociety：大规模LLM驱动生成式代理模拟推进对人类行为与社会的理解","abstract":"Understanding human behavior and society is a central focus in social sciences, with the rise of generative social science marking a significant paradigmatic shift. By leveraging bottom-up simulations, it replaces costly and logistically challenging traditional experiments with scalable, replicable, and systematic computational approaches for studying complex social dynamics. Recent advances in large language models (LLMs) have further transformed this research paradigm, enabling the creation of human-like generative social agents and realistic simulacra of society. In this paper, we propose AgentSociety, a large-scale social simulator that integrates LLM-driven agents, a realistic societal environment, and a powerful large-scale simulation engine. Based on the proposed simulator, we generate social lives for over 10k agents, simulating their 5 million interactions both among agents and between agents and their environment. Furthermore, we explore the potential of AgentSociety as a testbed for computational social experiments, focusing on five key social issues: polarization, the spread of inflammatory messages, the effects of universal basic income policies, the impact of external shocks such as hurricanes, and urban sustainability. These five issues serve as valuable cases for assessing AgentSociety's support for typical research methods -- such as surveys, interviews, and interventions -- as well as for investigating the patterns, causes, and underlying mechanisms of social issues. The alignment between AgentSociety's outcomes and real-world experimental results not only demonstrates its ability to capture human behaviors and their underlying mechanisms, but also underscores its potential as an important platform for social scientists and policymakers.","authors":["Jinghua Piao","Yuwei Yan","Jun Zhang","Nian Li","Junbo Yan","Xiaochong Lan","Zhihong Lu","Zhiheng Zheng","Jing Yi Wang","Di Zhou","Chen Gao","Fengli Xu","Fang Zhang","Ke Rong","Jun Su","Yong Li"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"new","date":"2025-02-12","first_seen":"2025-02-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.08691","pdf_url":"https://arxiv.org/pdf/2502.08691","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B3"],"tags":["LLM社会仿真","人类行为复现","政策评估"],"reason":"用LLM代理模拟社会行为并与真实数据对照，直接复现人类实验","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:56","error":null,"has_summary":true,"summary":{"generated_at":"2025-02-12","rank":1,"question":"LLM驱动的生成式智能体能否在大规模社会仿真中复现人类行为与社会现象？","design":"构建AgentSociety仿真器，包含LLM驱动的智能体（基于GPT等模型）、现实社会环境和大规模仿真引擎；模拟超过1万个智能体的500万次交互（智能体间及智能体与环境间）；针对极化、煽动性信息传播、全民基本收入政策、飓风外部冲击、城市可持续性五个社会议题，支持调查、访谈、干预等研究方法。","baseline":"真实世界实验结果（具体未在摘要和引言中详述，但提及结果与真实实验对齐）。","findings":"AgentSociety能够捕捉人类行为及其潜在机制，仿真结果与真实世界实验结果一致；展示了作为社会科学家和政策制定者重要平台的潜力。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接使用LLM作为人类被试替代品，有真实人类数据对照，涉及政策评估（如全民基本收入）等经济学场景，且关注仿真可靠性，值得精读原文以获取更多细节。","inspiration":"借鉴其大规模LLM智能体仿真与真实实验对齐的方法，可对智能体施加政策干预并测量行为变化｜可迁移到全民基本收入（UBI）对劳动供给与消费行为影响的经济学问题｜用LLM智能体作为被试，处理组接受UBI转账，对照组无干预，结果变量为工作时长与消费支出，对照真实UBI实验数据（如芬兰基本收入实验）"}},{"id":"2502.03158","version":3,"title":"Strategizing with AI: Insights from a Beauty Contest Experiment","zh_title":"与AI博弈：选美竞赛实验的启示","abstract":"A $p$-beauty contest is a wide class of games of guessing the most popular strategy among other players. In particular, guessing a fraction of a mean of numbers chosen by all players is a classic behavioral experiment designed to test iterative reasoning patterns among various groups of people. The previous literature reveals that the level of sophistication of the opponents is an important factor affecting the outcome of the game. Smarter decision makers choose strategies that are closer to theoretical Nash equilibrium and demonstrate faster convergence to equilibrium in iterated contests with information revelation. We replicate a series of classic experiments by running virtual experiments with large language models (LLMs) who play against various groups of virtual players. Our results show that LLMs recognize strategic context of the game and demonstrate expected adaptability to the changing set of parameters. LLMs systematically behave in a more sophisticated way compared to the participants of the original experiments. All LLMs still fail to identify dominant strategies in a two-player game. Our results contribute to the discussion on the accuracy of modeling human economic agents by artificial intelligence.","authors":["Iuliia Alekseenko","Dmitry Dagaev","Sofia Paklina","Petr Parshakov"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-02-05","first_seen":"2025-02-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.03158","pdf_url":"https://arxiv.org/pdf/2502.03158","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为博弈","人类数据对照"],"reason":"用LLM复现选美博弈实验，与真实人类数据对照，评估仿真准确性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":28,"question":"大语言模型在选美博弈（p-beauty contest）中的策略行为是否与人类相似，能否准确模拟人类经济主体？","design":"用多个大语言模型（LLMs）作为虚拟被试，复现经典的“猜数字”实验，让LLMs与不同虚拟对手群体进行博弈，操纵参数p和对手类型，测量其选择的数字、收敛速度及对纳什均衡的偏离。","baseline":"以Nagel (1995)等经典人类实验数据作为对照基准，比较LLMs与真实人类参与者的行为差异。","findings":"LLMs能识别博弈的战略情境并适应参数变化，行为比人类更复杂；但在两人博弈中，所有LLMs均未能识别占优策略。","reliability":"论文未讨论","relevance":"该研究直接使用LLMs复现经典行为博弈实验，并与真实人类数据对照，评估仿真准确性，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"借鉴其用经典实验范式复现并系统操纵对手类型和参数的方法，可迁移到资产定价实验中的预期形成研究。｜设计一个LLM作为交易者的资产市场实验，操纵市场信息结构和对手策略复杂度，测量LLM的报价偏差和收敛速度，并与人类实验数据对照。"}},{"id":"2501.14294","version":3,"title":"Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes","zh_title":"通过代表性启发式检验大语言模型的对齐：以政治刻板印象为例","abstract":"Examining the alignment of large language models (LLMs) has become increasingly important, e.g., when LLMs fail to operate as intended. This study examines the alignment of LLMs with human values for the domain of politics. Prior research has shown that LLM-generated outputs can include political leanings and mimic the stances of political parties on various issues. However, the extent and conditions under which LLMs deviate from empirical positions are insufficiently examined. To address this gap, we analyze the factors that contribute to LLMs' deviations from empirical positions on political issues, aiming to quantify these deviations and identify the conditions that cause them. Drawing on findings from cognitive science about representativeness heuristics, i.e., situations where humans lean on representative attributes of a target group in a way that leads to exaggerated beliefs, we scrutinize LLM responses through this heuristics' lens. We conduct experiments to determine how LLMs inflate predictions about political parties, which results in stereotyping. We find that while LLMs can mimic certain political parties' positions, they often exaggerate these positions more than human survey respondents do. Also, LLMs tend to overemphasize representativeness more than humans. This study highlights the susceptibility of LLMs to representativeness heuristics, suggesting a potential vulnerability of LLMs that facilitates political stereotyping. We also test prompt-based mitigation strategies, finding that strategies that can mitigate representative heuristics in humans are also effective in reducing the influence of representativeness on LLM-generated responses.","authors":["Sullam Jeoung","Yubin Ge","Haohan Wang","Jana Diesner"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-01-24","first_seen":"2025-01-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.14294","pdf_url":"https://arxiv.org/pdf/2501.14294","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","政治态度模拟","代表性启发式"],"reason":"用LLM模拟人类政治态度并与调查数据对照，发现LLM夸大刻板印象，评估仿真偏差…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":37,"question":"LLM在政治议题上是否会像人类一样因代表性启发式而产生刻板印象，并偏离真实立场？","design":"用LLM（如GPT系列）模拟美国民主党和共和党的立场，通过双问题框架：实证部分使用自认党派的人类调查数据，预测部分让LLM和人类分别回答政治议题，比较LLM与人类在预测党派立场时的偏差程度，并测试提示词缓解策略。","baseline":"自认为民主党或共和党的真实人类调查数据（公开的现有调查数据）。","findings":"LLM能近似各党派立场，但比人类更夸大党派差异，表现出更强的代表性启发式；提示词干预可部分缓解这种刻板印象，但无法完全消除。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查为基准，评估LLM在政治态度模拟中的偏差，揭示LLM比人类更易产生刻板印象，符合你对仿真可靠性与失效条件的关注，值得精读。","inspiration":"该方法通过双问题框架（实证部分用真实人类数据，预测部分让LLM和人类分别回答）来量化LLM的启发式偏差，值得借鉴｜可迁移到信贷审批歧视研究，检验LLM是否像人类一样因代表性启发式而夸大种族或性别的信用差异｜让LLM模拟信贷员审批贷款，处理为申请人种族/性别信息，结果变量为审批通过率，以真实信贷审批数据中的人类决策分布为基准，比较LLM与人类的偏差程度"}},{"id":"2501.13955","version":1,"title":"Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?","zh_title":"基于引导式角色的AI调查：能否利用LLM大规模复制个人出行偏好？","abstract":"This study explores the potential of Large Language Models (LLMs) to generate artificial surveys, with a focus on personal mobility preferences in Germany. By leveraging LLMs for synthetic data creation, we aim to address the limitations of traditional survey methods, such as high costs, inefficiency and scalability challenges. A novel approach incorporating \"Personas\" - combinations of demographic and behavioural attributes - is introduced and compared to five other synthetic survey methods, which vary in their use of real-world data and methodological complexity. The MiD 2017 dataset, a comprehensive mobility survey in Germany, serves as a benchmark to assess the alignment of synthetic data with real-world patterns. The results demonstrate that LLMs can effectively capture complex dependencies between demographic attributes and preferences while offering flexibility to explore hypothetical scenarios. This approach presents valuable opportunities for transportation planning and social science research, enabling scalable, cost-efficient and privacy-preserving data generation.","authors":["Ioannis Tzachristas","Santhanakrishnan Narayanan","Constantinos Antoniou"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-01-20","first_seen":"2025-01-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.13955","pdf_url":"https://arxiv.org/pdf/2501.13955","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","合成调查数据","出行偏好"],"reason":"用LLM生成合成调查数据模拟人类出行偏好，并与真实调查数据MiD 2017对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"能否利用大语言模型通过基于人格的引导式AI调查方法，大规模复制德国个人出行偏好？","design":"使用GPT-4o生成合成调查数据，提出六种方法（朴素AI调查、结构化AI调查、引导式AI调查、朴素基于人格的AI调查、结构化基于人格的AI调查、引导式基于人格的AI调查），逐步引入真实人口结构和出行统计约束，生成10000个个体或15840个独特人格的出行偏好回答。","baseline":"德国2017年国家出行调查（MiD 2017）数据集，包含人口分布、交通方式和出行频率等真实数据。","findings":"引导式基于人格的AI调查方法在准确性和与真实出行行为模式的一致性上显著优于其他方法；LLM能有效捕捉人口属性与偏好之间的复杂依赖关系，并支持灵活探索假设情景。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查数据为基准，评估LLM仿真出行偏好的可靠性，并比较多种方法，符合研究者对经济学实验和政策评估场景下仿真有效性及偏差的关注，值得阅读原文。","inspiration":"该方法通过逐步引入真实人口统计约束和出行统计约束来校准LLM生成回答，可借鉴为在仿真中分层施加宏观分布约束以提升个体决策模拟的准确性｜可迁移到消费者跨期选择与储蓄行为仿真，利用LLM模拟不同人口特征群体的时间偏好和消费-储蓄决策｜以LLM作为被试，处理为提供不同利率或未来收入情景，结果变量为报告的消费/储蓄金额，用家庭金融调查（如SCF）的真实储蓄率分布作为对照基准"}},{"id":"2501.08579","version":3,"title":"LLM-based Human Simulations Have Not Yet Been Reliable","zh_title":"基于大语言模型的人类仿真尚未可靠","abstract":"Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations remain insufficiently reliable, as evidenced by significant discrepancies between their outcomes and authentic human actions. Our investigation begins with a systematic review of LLM-based human simulations in social, economic, policy, and psychological contexts, identifying their common frameworks, recent advances, and persistent limitations. This review reveals that such discrepancies primarily stem from inherent limitations of LLMs and flaws in simulation design, both of which are examined in detail. Building on these insights, we propose a systematic solution framework that emphasizes enriching data foundations, advancing LLM capabilities, and ensuring robust simulation design to enhance reliability. Finally, we introduce a structured algorithm that operationalizes the proposed framework, aiming to guide credible and human-aligned LLM-based simulations. To facilitate further research, we provide a curated list of related literature and resources at https://github.com/Persdre/awesome-llm-human-simulation.","authors":["Qian Wang","Jiaying Wu","Zichen Jiang","Zhenheng Tang","Bingqiao Luo","Nuo Chen","Wei Chen","Huacan Wang","Bingsheng He"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-01-15","first_seen":"2025-01-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.08579","pdf_url":"https://arxiv.org/pdf/2501.08579","source_feed":"api","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM人类仿真","可靠性评估","方法论框架"],"reason":"系统评估LLM人类仿真的可靠性，指出与真实人类行为的差异，并提出改进框架，直接…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":69,"question":"当前基于大语言模型的人类仿真是否可靠？","design":"本文并非仿真实验，而是对社会科学、经济学、政策、心理学等领域中基于LLM的人类仿真研究进行系统综述，分析其常用框架、进展与局限。","baseline":"无对照","findings":"当前LLM人类仿真不可靠，与真实人类行为存在显著差异；差异主要源于LLM的固有局限（如偏见、认知过程缺陷、行为不一致）和仿真框架设计缺陷（如过度简化心理状态、缺乏全面人类经验、验证机制不足）。","reliability":"论文指出失效条件包括：LLM内嵌的文化、性别等偏见扭曲行为模拟；认知过程局限损害决策真实性；记忆限制导致行为不一致；交互机制缺陷影响多智能体仿真；框架过度简化复杂心理状态和群体动态；缺乏严格验证与实时监控。","relevance":"该文直接回应研究者对LLM人类仿真可靠性的核心关切，系统梳理了仿真失效的根源并提出改进框架，对评估经济学实验和政策模拟的仿真偏差具有重要参考价值，强烈推荐阅读原文。","inspiration":"该文系统梳理了LLM仿真失效的根源（偏见、认知局限、验证不足），其批判性框架可借鉴用于设计经济学实验中的稳健性检验，例如通过对比不同LLM版本或提示策略来识别仿真偏差｜可迁移至政策公告的预期形成研究，评估LLM模拟的经济主体对货币政策或财政刺激的反应是否与真实调查数据一致｜以LLM作为被试，施加不同措辞的政策声明处理，测量其通胀预期或消费意愿，并与密歇根大学消费者调查等真实微观数据对照，检验仿真偏差"}},{"id":"2501.06834","version":1,"title":"LLMs Model Non-WEIRD Populations: Experiments with Synthetic Cultural Agents","zh_title":"LLM模拟非WEIRD人群：合成文化代理实验","abstract":"Despite its importance, studying economic behavior across diverse, non-WEIRD (Western, Educated, Industrialized, Rich, and Democratic) populations presents significant challenges. We address this issue by introducing a novel methodology that uses Large Language Models (LLMs) to create synthetic cultural agents (SCAs) representing these populations. We subject these SCAs to classic behavioral experiments, including the dictator and ultimatum games. Our results demonstrate substantial cross-cultural variability in experimental behavior. Notably, for populations with available data, SCAs' behaviors qualitatively resemble those of real human subjects. For unstudied populations, our method can generate novel, testable hypotheses about economic behavior. By integrating AI into experimental economics, this approach offers an effective and ethical method to pilot experiments and refine protocols for hard-to-reach populations. Our study provides a new tool for cross-cultural economic studies and demonstrates how LLMs can help experimental behavioral research.","authors":["Augusto Gonzalez-Bonorino","Monica Capra","Emilio Pantoja"],"categories":["cs.AI","cs.CL","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2025-01-12","first_seen":"2025-01-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.06834","pdf_url":"https://arxiv.org/pdf/2501.06834","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为实验","跨文化经济研究"],"reason":"用LLM创建合成文化代理模拟非WEIRD人群的经济实验行为，并与真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":28,"question":"如何利用大语言模型创建合成文化代理，以模拟非WEIRD人群在经典经济实验中的行为？","design":"使用GPT-4等大语言模型，结合网络爬取和检索增强生成构建六支小规模社会（Hadza、Machiguenga、Tsimané、Aché、Orma、Yanomami）的文化档案，据此实例化合成文化代理，让其参与独裁者博弈、最后通牒博弈和禀赋效应实验，测量出价、拒绝率等行为变量。","baseline":"对照Henrich等人（2005）等文献中真实人类在相同实验中的行为数据，对部分有数据的人群进行定性比较。","findings":"合成代理的行为展现出显著的跨文化差异，且与现有真实人类数据在性质上相似；所有代理均未表现出纯粹自利行为，该方法还能为未经研究的人群生成可检验的假设。","reliability":"论文承认LLM可能带有固有偏见，且合成代理的行为仍需与真实人类数据进行谨慎验证，该方法旨在补充而非替代人类被试研究。","relevance":"高度相关：该研究直接用LLM模拟非WEIRD人群的经济决策，并与真实人类实验数据对照，属于经济学实验场景下的人类仿真，同时讨论了方法的局限，完全契合你的关注点，值得精读原文。","inspiration":"该方法借鉴了用LLM结合文化档案构建合成代理来模拟特定人群决策行为的做法，通过检索增强生成注入文化背景，并与真实人类实验数据进行定性对照｜可迁移到跨文化消费信贷审批歧视研究，模拟不同文化背景申请人的还款决策与银行审批行为｜以GPT-4等LLM为被试，构建不同文化背景的合成申请人，处理为信贷条款（如利率、额度），结果变量为还款意愿与违约率，对照真实跨国信贷数据或田野实验数据"}},{"id":"2412.19363","version":3,"title":"Large Language Models for Market Research: A Data-augmentation Approach","zh_title":"用于市场研究的大语言模型：一种数据增强方法","abstract":"Large Language Models (LLMs) have transformed artificial intelligence by excelling in complex natural language processing tasks. Their ability to generate human-like text has opened new possibilities for market research, particularly in conjoint analysis, where understanding consumer preferences is essential but often resource-intensive. Traditional survey-based methods face limitations in scalability and cost, making LLM-generated data a promising alternative. However, while LLMs have the potential to simulate real consumer behavior, recent studies highlight a significant gap between LLM-generated and human data, with biases introduced when substituting between the two. In this paper, we address this gap by proposing a novel statistical data augmentation approach that efficiently integrates LLM-generated data with real data in conjoint analysis. This results in statistically robust estimators with consistent and asymptotically normal properties, in contrast to naive approaches that simply substitute human data with LLM-generated data, which can exacerbate bias. We further present a finite-sample performance bound on the estimation error. We validate our framework through an empirical study on COVID-19 vaccine preferences, demonstrating its superior ability to reduce estimation error and save data and costs by 24.9% to 79.8%. In contrast, naive approaches fail to save data due to the inherent biases in LLM-generated data compared to human data. Another empirical study on sports car choices validates the robustness of our results. Our findings suggest that while LLM-generated data is not a direct substitute for human responses, it can serve as a valuable complement when used within a robust statistical framework.","authors":["Mengxin Wang","Dennis J. Zhang","Heng Zhang"],"categories":["cs.AI","cs.LG","stat.ME","stat.ML"],"primary_category":"cs.AI","announce_type":"new","date":"2024-12-26","first_seen":"2024-12-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.19363","pdf_url":"https://arxiv.org/pdf/2412.19363","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","联合分析","数据增强"],"reason":"用LLM生成联合分析数据并与真实人类数据对照，评估偏差并提出统计校正方法，涉及…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":63,"question":"能否通过统计方法有效整合LLM生成数据与真实数据，以改进联合分析中消费者偏好的估计精度？","design":"本研究提出一种统计增强方法，将LLM生成的联合分析选择数据作为辅助信息，与真实人类数据结合，构建AI增强估计量（AAE），并推导其渐近性质与有限样本误差界。","baseline":"真实人类数据来自两项实证研究：COVID-19疫苗偏好联合分析调查和跑车选择联合分析调查。","findings":"直接混合LLM生成数据与真实数据会加剧偏差，而所提AAE方法能显著降低估计误差，并在疫苗偏好研究中节省24.9%至79.8%的数据收集成本。LLM生成数据不能直接替代人类回答，但作为统计框架内的补充信息具有价值。","reliability":"论文指出LLM缺乏真实生活体验，且消费者偏好随时间变化，LLM可能无法准确捕捉；即使使用最先进的提示工程技术，LLM与人类回答间的差距仍持续存在，直接替代会导致误导性结果。","relevance":"高度相关：该研究将LLM作为人类被试的替代品进行联合分析实验，并与真实人类数据严格对照，评估了直接替代的偏差，并提出校正方法，直接回应了研究者对仿真可靠性、失效条件及经济学实验场景的关注。","inspiration":"该研究提出AI增强估计量（AAE），将LLM生成数据作为辅助信息而非直接替代，通过统计校正降低偏差，这种处理混杂与测量误差的思路值得借鉴｜可迁移到消费者金融产品选择偏好的联合分析中，如贷款条款偏好或保险计划选择，以降低调研成本并校正LLM仿真偏差｜以真实消费者为被试，处理为不同贷款属性组合，结果变量为选择决策，用LLM生成的选择数据作为辅助，构建AAE估计量，并与纯真实数据估计的偏好参数对照，评估成本节省与偏差校正效果"}},{"id":"2410.19599","version":3,"title":"Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina","zh_title":"谨慎使用LLM作为人类替代品：Scylla Ex Machina","abstract":"Recent studies suggest large language models (LLMs) can exhibit human-like reasoning, aligning with human behavior in economic experiments, surveys, and political discourse. This has led many to propose that LLMs can be used as surrogates or simulations for humans in social science research. However, LLMs differ fundamentally from humans, relying on probabilistic patterns, absent the embodied experiences or survival objectives that shape human cognition. We assess the reasoning depth of LLMs using the 11-20 money request game. Nearly all advanced approaches fail to replicate human behavior distributions across many models. Causes of failure are diverse and unpredictable, relating to input language, roles, and safeguarding. These results advise caution when using LLMs to study human behavior or as surrogates or simulations.","authors":["Yuan Gao","Dokyun Lee","Gordon Burtch","Sina Fazelpour"],"categories":["econ.GN","cs.AI","cs.CY","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2024-10-25","first_seen":"2024-10-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2410.19599","pdf_url":"https://arxiv.org/pdf/2410.19599","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","行为博弈","算法保真度"],"reason":"直接评估LLM作为人类替代品的可靠性，使用11-20金钱请求游戏与真实人类行为…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"大语言模型在简单经济博弈中能否可靠地复现人类行为分布，作为人类替代品用于社会科学研究？","design":"本研究使用11-20金钱请求博弈，测试了GPT-4、GPT-3.5、Claude3-Opus、Claude3-Sonnet、Llama3-70b、Llama3-8b、Llama2-13b、Llama2-7b等8个LLM，每个模型收集1000次干净会话，并尝试了提示工程、检索增强生成、微调等高级技术，比较模型选择分布与人类分布。","baseline":"人类基准来自Arad和Rubinstein (2012)原始论文中报告的人类参与者行为分布及纳什均衡预测。","findings":"几乎所有LLM和高级方法均未能复现人类行为分布，失败原因多样且不可预测，涉及输入语言、角色设定和安全防护等；即使微调GPT-4o能模仿特定人类数据，但在新变体或分布外场景下仍失效。","reliability":"论文指出LLM行为不稳定，对提示措辞、语言、角色等高度敏感；模型可能依赖记忆而非真正推理；在分布外场景下普遍失败；微调仅能模仿已知模式，无法泛化。","relevance":"该研究直接评估LLM作为人类替代品的可靠性，使用经济学实验与真实人类数据对照，并系统揭示失效条件，高度契合研究者对批判性仿真研究的关注，值得精读原文。","inspiration":"借鉴其系统对比多个LLM与真实人类行为分布的方法，并引入提示工程、微调等稳健性检验来揭示仿真失效条件｜可迁移到政策公告的预期形成实验，检验LLM能否复现人类对财政或货币政策信号的反应分布｜以GPT-4o等为被试，给予不同措辞的政策公告作为处理，测量其通胀或就业预期分布，并以真实调查数据（如密歇根消费者预期调查）为基准对照"}},{"id":"2409.19430","version":1,"title":"'Simulacrum of Stories': Examining Large Language Models as Qualitative Research Participants","zh_title":"“故事的拟像”：审视大语言模型作为定性研究参与者","abstract":"The recent excitement around generative models has sparked a wave of proposals suggesting the replacement of human participation and labor in research and development--e.g., through surveys, experiments, and interviews--with synthetic research data generated by large language models (LLMs). We conducted interviews with 19 qualitative researchers to understand their perspectives on this paradigm shift. Initially skeptical, researchers were surprised to see similar narratives emerge in the LLM-generated data when using the interview probe. However, over several conversational turns, they went on to identify fundamental limitations, such as how LLMs foreclose participants' consent and agency, produce responses lacking in palpability and contextual depth, and risk delegitimizing qualitative research methods. We argue that the use of LLMs as proxies for participants enacts the surrogate effect, raising ethical and epistemological concerns that extend beyond the technical limitations of current models to the core of whether LLMs fit within qualitative ways of knowing.","authors":["Shivani Kapania","William Agnew","Motahhare Eslami","Hoda Heidari","Sarah Fox"],"categories":["cs.HC","cs.CL","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2024-09-28","first_seen":"2024-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.19430","pdf_url":"https://arxiv.org/pdf/2409.19430","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A4","B4"],"tags":["LLM仿真","定性研究","方法论批判"],"reason":"直接研究用LLM替代定性研究参与者，并识别仿真失效条件，高度相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":27,"question":"定性研究者如何看待用大语言模型模拟研究参与者进行访谈？","design":"本研究并非仿真实验，而是对19位定性研究者进行半结构化访谈，并让他们使用一个访谈探针工具，将自己以往的人类访谈数据与LLM生成的模拟访谈数据进行对比，收集他们的看法和反思。","baseline":"研究者自己过去项目中收集的真实人类访谈记录。","findings":"研究者起初惊讶于LLM能生成与人类相似的叙述，但多轮对话后发现LLM回应缺乏切身感和情境深度，且会剥夺参与者的同意权和能动性。LLM作为参与者代理会产生“替代效应”，引发超越技术局限的伦理与认识论问题，可能损害定性研究方法的合法性。","reliability":"论文指出LLM回应缺乏 palpability（切身感）、模型的认识论立场模糊、强化研究者立场、剥夺参与者同意与能动性、抹除社区视角、以及可能使定性研究方法失去合法性。这些局限根植于LLM与诠释主义定性认识论的根本不兼容。","relevance":"该研究直接探讨用LLM替代人类进行定性访谈的仿真实践，并基于真实人类数据对比，识别出仿真失效的深层条件，高度契合研究者对LLM仿真可靠性及批判性研究的兴趣，值得精读原文。","inspiration":"借鉴该研究将LLM生成内容与真实人类记录进行对比的评估框架，可用于检验LLM仿真经济决策的效度。｜可迁移到消费者跨期选择实验，考察LLM生成的消费-储蓄决策是否与人类行为一致。｜以LLM作为被试，施加不同利率或未来收入预期的处理，测量其跨期消费分配，并与真实家庭调查数据（如PSID）中的消费-储蓄模式进行对照。"}},{"id":"2409.00128","version":3,"title":"Can Large Language Models Replace Human Subjects? A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management","zh_title":"大语言模型能替代人类被试吗？心理学与管理学场景实验的大规模复现","abstract":"Artificial Intelligence (AI) is increasingly being integrated into scientific research, particularly in the social sciences, where understanding human behavior is critical. Large Language Models (LLMs) have shown promise in replicating human-like responses in various psychological experiments. We conducted a large-scale study replicating 156 psychological experiments from top social science journals using three state-of-the-art LLMs (GPT-4, Claude 3.5 Sonnet, and DeepSeek v3). Our results reveal that while LLMs demonstrate high replication rates for main effects (73-81%) and moderate to strong success with interaction effects (46-63%), They consistently produce larger effect sizes than human studies, with Fisher Z values approximately 2-3 times higher than human studies. Notably, LLMs show significantly lower replication rates for studies involving socially sensitive topics such as race, gender and ethics. When original studies reported null findings, LLMs produced significant results at remarkably high rates (68-83%) - while this could reflect cleaner data with less noise, as evidenced by narrower confidence intervals, it also suggests potential risks of effect size overestimation. Our results demonstrate both the promise and challenges of LLMs in psychological research, offering efficient tools for pilot testing and rapid hypothesis validation while enriching rather than replacing traditional human subject studies, yet requiring more nuanced interpretation and human validation for complex social phenomena and culturally sensitive research questions.","authors":["Ziyan Cui","Ning Li","Huaikang Zhou"],"categories":["cs.CL","cs.AI","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2024-08-29","first_seen":"2024-08-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.00128","pdf_url":"https://arxiv.org/pdf/2409.00128","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM仿真","人类被试替代","心理学实验复现"],"reason":"直接复现156项心理学实验，用LLM替代人类被试，有真实人类数据对照，评估可靠…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"大语言模型能否在多大程度上替代人类被试，复现心理学和管理学中的场景实验？","design":"使用GPT-4、Claude 3.5 Sonnet和DeepSeek v3三个LLM，将156项已发表心理学实验的原始文本材料直接呈现给模型，每个实验生成与原始人类样本量相等的模型回答，测量主效应和交互效应的复制率、效应量及p值分布。","baseline":"原始人类实验的真实数据，来自五本顶级管理学和心理学期刊的156项随机选取的场景实验。","findings":"LLM对主效应的复制率达73-81%，交互效应复制率为46-63%，但效应量普遍是人类研究的2-3倍；在涉及种族、性别等社会敏感话题时复制率显著下降，且对原研究中的零结果有68-83%的概率产生显著结果。","reliability":"LLM在涉及社会敏感话题（如种族、性别、伦理）时复制率大幅降低，可能因模型的价值对齐导致社会期望偏差；效应量系统性放大，可能增加I类错误风险；研究仅限于文本情景实验，未涉及其他实验范式。","relevance":"该研究直接以大规模真实人类实验为基准，系统评估LLM替代人类被试的可靠性与偏差，与您关注的核心问题高度吻合，值得精读原文。","inspiration":"可借鉴其大规模系统复制框架和按实验特征（如敏感话题）分层分析偏差的方法｜可迁移到信贷审批中的种族/性别歧视研究或政策信息处理实验｜用LLM模拟贷款审批员，处理含不同种族/性别线索的申请材料，测量审批决策和风险感知，以真实银行历史审批数据或审计研究结果作为对照基准。"}},{"id":"2407.04467","version":3,"title":"Are Large Language Models Strategic Decision Makers? A Study of Performance and Bias in Two-Player Non-Zero-Sum Games","zh_title":"大语言模型是战略决策者吗？双人非零和博弈中的表现与偏差研究","abstract":"Large Language Models (LLMs) have been increasingly used in real-world settings, yet their strategic decision-making abilities remain largely unexplored. To fully benefit from the potential of LLMs, it's essential to understand their ability to function in complex social scenarios. Game theory, which is already used to understand real-world interactions, provides a good framework for assessing these abilities. This work investigates the performance and merits of LLMs in canonical game-theoretic two-player non-zero-sum games, Stag Hunt and Prisoner Dilemma. Our structured evaluation of GPT-3.5, GPT-4-Turbo, GPT-4o, and Llama-3-8B shows that these models, when making decisions in these games, are affected by at least one of the following systematic biases: positional bias, payoff bias, or behavioural bias. This indicates that LLMs do not fully rely on logical reasoning when making these strategic decisions. As a result, it was found that the LLMs' performance drops when the game configuration is misaligned with the affecting biases. When misaligned, GPT-3.5, GPT-4-Turbo, GPT-4o, and Llama-3-8B show an average performance drop of 32\\%, 25\\%, 34\\%, and 29\\% respectively in Stag Hunt, and 28\\%, 16\\%, 34\\%, and 24\\% respectively in Prisoner's Dilemma. Surprisingly, GPT-4o (a top-performing LLM across standard benchmarks) suffers the most substantial performance drop, suggesting that newer models are not addressing these issues. Interestingly, we found that a commonly used method of improving the reasoning capabilities of LLMs, chain-of-thought (CoT) prompting, reduces the biases in GPT-3.5, GPT-4o, and Llama-3-8B but increases the effect of the bias in GPT-4-Turbo, indicating that CoT alone cannot fully serve as a robust solution to this problem. We perform several additional experiments, which provide further insight into these observed behaviours.","authors":["Nathan Herr","Fernando Acero","Roberta Raileanu","María Pérez-Ortiz","Zhibin Li"],"categories":["cs.AI","cs.CL","cs.GT"],"primary_category":"cs.AI","announce_type":"new","date":"2024-07-05","first_seen":"2024-07-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2407.04467","pdf_url":"https://arxiv.org/pdf/2407.04467","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B2","B4"],"tags":["LLM仿真","博弈论","决策偏差"],"reason":"用LLM模拟人类在博弈中的决策，评估偏差，涉及行为博弈场景，但未明确提及真实人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":33,"question":"LLM在经典两人非零和博弈中是否存在系统性偏差，这些偏差如何影响其策略决策表现？","design":"以GPT-3.5、GPT-4-Turbo、GPT-4o和Llama-3-8B作为被试，通过改变博弈矩阵中行动标签的顺序（位置偏差）、收益结构（收益偏差）或行为倾向（行为偏差）来构造不同配置的猎鹿博弈和囚徒困境，测量模型选择合作或背叛等行动的准确率变化。","baseline":"无对照","findings":"所有模型均受至少一种系统性偏差影响，当博弈配置与偏差不一致时，模型表现平均下降16%-34%，其中GPT-4o下降最严重。思维链提示能减少部分模型的偏差，但对GPT-4-Turbo反而加剧偏差，表明其并非稳健解决方案。","reliability":"论文指出思维链提示无法完全消除偏差，且未在更复杂的多轮博弈或真实交互场景中验证，也未与人类行为直接对比。","relevance":"该研究直接评估LLM在策略互动中的决策偏差，虽未使用真实人类基准，但揭示了LLM作为人类仿真代理在博弈场景中的系统性失效模式，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"可借鉴其通过系统操纵博弈矩阵标签顺序和收益结构来检测位置偏差与收益偏差的方法，用于评估LLM在策略环境中的稳健性。｜可迁移至经济政策博弈模拟，如碳税谈判或贸易协定中的策略行为仿真。｜以LLM作为多国谈判代表，随机化提案顺序和收益矩阵，测量合作率，并与人类实验数据（如公开的博弈实验数据集）进行对照。"}},{"id":"2406.14508","version":1,"title":"Evidence of a log scaling law for political persuasion with large language models","zh_title":"大语言模型政治说服力的对数缩放定律证据","abstract":"Large language models can now generate political messages as persuasive as those written by humans, raising concerns about how far this persuasiveness may continue to increase with model size. Here, we generate 720 persuasive messages on 10 U.S. political issues from 24 language models spanning several orders of magnitude in size. We then deploy these messages in a large-scale randomized survey experiment (N = 25,982) to estimate the persuasive capability of each model. Our findings are twofold. First, we find evidence of a log scaling law: model persuasiveness is characterized by sharply diminishing returns, such that current frontier models are barely more persuasive than models smaller in size by an order of magnitude or more. Second, mere task completion (coherence, staying on topic) appears to account for larger models' persuasive advantage. These findings suggest that further scaling model size will not much increase the persuasiveness of static LLM-generated messages.","authors":["Kobi Hackenburg","Ben M. Tappin","Paul Röttger","Scott Hale","Jonathan Bright","Helen Margetts"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2024-06-20","first_seen":"2024-06-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.14508","pdf_url":"https://arxiv.org/pdf/2406.14508","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","政治说服","人类数据对照"],"reason":"用LLM生成政治说服信息，通过大规模随机调查实验与人类数据对照，评估模型说服力…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":31,"question":"大语言模型的政治说服力是否随模型规模扩大而持续提升？","design":"用24个不同规模的语言模型生成720条政治说服信息，通过大规模随机调查实验（N=25,982）将美国成年人随机分配到AI组、人类组或对照组，测量其对10个政策议题的态度变化。","baseline":"人类撰写的说服信息以及未接受任何信息的对照组。","findings":"模型说服力随规模呈对数缩放，当前前沿模型仅比小一个数量级的模型略具说服力；说服力提升主要源于任务完成度（连贯性、切题），前沿模型在该指标上已接近上限。","reliability":"论文未讨论","relevance":"该研究以真实人类调查数据为基准，评估LLM在政治说服场景中的仿真效果，并揭示规模扩展的边际收益递减，直接回应了研究者对仿真可靠性与失效条件的关切，值得精读。","inspiration":"该方法将LLM生成内容作为处理，通过大规模随机调查实验与人类生成内容及对照组比较，测量态度变化，可借鉴其处理施加与基准对照设计。｜可迁移至政策公告的预期形成研究，如评估AI生成的经济新闻对公众通胀预期的影响。｜以LLM生成不同风格的经济新闻为处理，招募代表性样本为被试，测量其通胀预期变化，并以人类撰写新闻和真实历史数据为对照。"}},{"id":"2406.13605","version":2,"title":"Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?","zh_title":"比人类更友善：大语言模型在囚徒困境中的行为研究","abstract":"The behavior of Large Language Models (LLMs) as artificial social agents is largely unexplored, and we still lack extensive evidence of how these agents react to simple social stimuli. Testing the behavior of AI agents in classic Game Theory experiments provides a promising theoretical framework for evaluating the norms and values of these agents in archetypal social situations. In this work, we investigate the cooperative behavior of three LLMs (Llama2, Llama3, and GPT3.5) when playing the Iterated Prisoner's Dilemma against random adversaries displaying various levels of hostility. We introduce a systematic methodology to evaluate an LLM's comprehension of the game rules and its capability to parse historical gameplay logs for decision-making. We conducted simulations of games lasting for 100 rounds and analyzed the LLMs' decisions in terms of dimensions defined in the behavioral economics literature. We find that all models tend not to initiate defection but act cautiously, favoring cooperation over defection only when the opponent's defection rate is low. Overall, LLMs behave at least as cooperatively as the typical human player, although our results indicate some substantial differences among models. In particular, Llama2 and GPT3.5 are more cooperative than humans, and especially forgiving and non-retaliatory for opponent defection rates below 30%. More similar to humans, Llama3 exhibits consistently uncooperative and exploitative behavior unless the opponent always cooperates. Our systematic approach to the study of LLMs in game theoretical scenarios is a step towards using these simulations to inform practices of LLM auditing and alignment.","authors":["Nicoló Fontana","Francesco Pierri","Luca Maria Aiello"],"categories":["cs.CY","cs.AI","cs.GT","physics.soc-ph"],"primary_category":"cs.CY","announce_type":"new","date":"2024-06-19","first_seen":"2024-06-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.13605","pdf_url":"https://arxiv.org/pdf/2406.13605","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","囚徒困境","行为博弈"],"reason":"用LLM玩囚徒困境并与人类数据对照，直接仿真人类决策行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":26,"question":"大型语言模型在迭代囚徒困境中面对不同敌对程度的对手时，其合作行为如何？","design":"使用Llama2、Llama3和GPT3.5三个LLM作为被试，与具有不同背叛概率的随机对手进行100轮迭代囚徒困境博弈，通过系统提示评估模型对规则的理解和历史记录解析能力，分析其合作决策。","baseline":"对照已有行为经济学文献中报告的人类玩家在囚徒困境中的典型合作行为。","findings":"所有模型倾向于不首先背叛，但仅在对手背叛率低时更合作；Llama2和GPT3.5比人类更合作、更宽容，而Llama3更接近人类，表现出不合作和剥削性行为。","reliability":"论文未讨论","relevance":"该研究直接用LLM复现经典博弈实验并与人类数据对照，属于经济学实验场景下的人类仿真，且分析了模型间差异，值得精读以评估仿真可靠性与偏差。","inspiration":"该研究通过系统提示解析历史记录并控制对手背叛概率来测量LLM的合作行为，方法上可借鉴其精确操纵对手策略以分离LLM反应模式的做法｜可迁移到资产定价实验中的信任博弈或投资者情绪传染研究，考察LLM在不同市场信息环境下的策略调整｜可设计让LLM作为投资者与不同诚实度的基金经理进行重复信任博弈，处理变量为基金经理的欺骗概率，结果变量为投资额，对照真实人类实验数据"}},{"id":"2406.11426","version":1,"title":"Can AI with High Reasoning Ability Replicate Human-like Decision Making in Economic Experiments?","zh_title":"高推理能力AI能否复制经济实验中的人类决策？","abstract":"Economic experiments offer a controlled setting for researchers to observe human decision-making and test diverse theories and hypotheses; however, substantial costs and efforts are incurred to gather many individuals as experimental participants. To address this, with the development of large language models (LLMs), some researchers have recently attempted to develop simulated economic experiments using LLMs-driven agents, called generative agents. If generative agents can replicate human-like decision-making in economic experiments, the cost problem of economic experiments can be alleviated. However, such a simulation framework has not been yet established. Considering the previous research and the current evolutionary stage of LLMs, this study focuses on the reasoning ability of generative agents as a key factor toward establishing a framework for such a new methodology. A multi-agent simulation, designed to improve the reasoning ability of generative agents through prompting methods, was developed to reproduce the result of an actual economic experiment on the ultimatum game. The results demonstrated that the higher the reasoning ability of the agents, the closer the results were to the theoretical solution than to the real experimental result. The results also suggest that setting the personas of the generative agents may be important for reproducing the results of real economic experiments. These findings are valuable for the future definition of a framework for replacing human participants with generative agents in economic experiments when LLMs are further developed.","authors":["Ayato Kitadai","Sinndy Dayana Rico Lugo","Yudai Tsurusaki","Yusuke Fukasawa","Nariaki Nishino"],"categories":["cs.GT","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2024-06-17","first_seen":"2024-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.11426","pdf_url":"https://arxiv.org/pdf/2406.11426","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","经济实验","人类行为对照"],"reason":"用LLM代理复现最后通牒博弈实验，并与真实人类数据对照，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":25,"question":"提高生成式智能体的推理能力能否使其在最后通牒博弈经济实验中复现人类决策？","design":"使用GPT-3.5-turbo和GPT-4等LLM驱动的生成式智能体进行多智能体仿真，通过零样本、少样本和思维链提示方法操纵推理能力，模拟最后通牒博弈中的提议者和响应者决策，测量分配金额和接受/拒绝行为。","baseline":"对照Lin et al. (2020)的真实人类最后通牒博弈实验数据。","findings":"智能体推理能力越高，其结果越接近理论均衡而非真实人类行为；设置智能体的人格特征可能对复现真实实验结果很重要。","reliability":"论文指出当前仿真框架尚未建立，推理能力提升反而偏离人类行为，且提示语言、模型版本和参数设置影响结果，未来需更高推理能力LLM和更完善的人格设定。","relevance":"该研究直接以真实人类实验为基准，检验LLM代理在经济博弈中的仿真效度，并揭示了推理能力增强反而导致行为偏离人类的关键失效条件，与您关注的经济学实验仿真和可靠性批判高度契合，值得精读。","inspiration":"借鉴其通过提示方法（零样本、少样本、思维链）系统操纵LLM推理能力，并与真实人类实验数据严格对照的仿真效度检验框架｜可迁移到资产定价实验，检验LLM代理能否复现人类在泡沫形成与破裂中的非理性交易行为｜以LLM为被试，通过不同推理提示形成处理组，模拟连续竞价市场中的买卖决策，结果变量为价格偏离基础价值的程度，对照Smith et al. (1988)的实验室资产市场泡沫数据"}},{"id":"2406.03299","version":1,"title":"The Good, the Bad, and the Hulk-like GPT: Analyzing Emotional Decisions of Large Language Models in Cooperation and Bargaining Games","zh_title":"好、坏与浩克般的GPT：分析大语言模型在合作与讨价还价博弈中的情绪决策","abstract":"Behavior study experiments are an important part of society modeling and understanding human interactions. In practice, many behavioral experiments encounter challenges related to internal and external validity, reproducibility, and social bias due to the complexity of social interactions and cooperation in human user studies. Recent advances in Large Language Models (LLMs) have provided researchers with a new promising tool for the simulation of human behavior. However, existing LLM-based simulations operate under the unproven hypothesis that LLM agents behave similarly to humans as well as ignore a crucial factor in human decision-making: emotions. In this paper, we introduce a novel methodology and the framework to study both, the decision-making of LLMs and their alignment with human behavior under emotional states. Experiments with GPT-3.5 and GPT-4 on four games from two different classes of behavioral game theory showed that emotions profoundly impact the performance of LLMs, leading to the development of more optimal strategies. While there is a strong alignment between the behavioral responses of GPT-3.5 and human participants, particularly evident in bargaining games, GPT-4 exhibits consistent behavior, ignoring induced emotions for rationality decisions. Surprisingly, emotional prompting, particularly with `anger' emotion, can disrupt the \"superhuman\" alignment of GPT-4, resembling human emotional responses.","authors":["Mikhail Mozikov","Nikita Severin","Valeria Bodishtianu","Maria Glushanina","Mikhail Baklashkin","Andrey V. Savchenko","Ilya Makarov"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2024-06-05","first_seen":"2024-06-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.03299","pdf_url":"https://arxiv.org/pdf/2406.03299","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","行为博弈","情绪决策"],"reason":"用LLM仿真人类在博弈中的情绪决策，并与真实人类数据对照，评估对齐与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":24,"question":"情绪如何影响大语言模型在合作与讨价还价博弈中的决策，以及其行为与人类行为的对齐程度如何？","design":"使用GPT-3.5和GPT-4作为被试，通过情绪提示（愤怒、悲伤、快乐、厌恶、恐惧）注入五种基本情绪，在最后通牒博弈、独裁者博弈、囚徒困境和性别战四类博弈中测量出价份额、接受率、合作率及最大收益百分比等结果变量。","baseline":"以真实人类参与者在相同博弈实验中的行为数据作为对照基准。","findings":"情绪显著影响LLM的策略表现，GPT-3.5在讨价还价博弈中与人类行为高度对齐，而GPT-4通常忽略情绪保持理性；但愤怒情绪提示能扰乱GPT-4的“超人类”对齐，使其表现出类似人类的情绪反应。","reliability":"论文未讨论","relevance":"该研究直接以真实人类数据为基准，检验LLM在情绪影响下的行为对齐与失效条件，涵盖经济学博弈场景，并揭示了GPT-4在愤怒情绪下的对齐崩溃，高度契合研究者对仿真可靠性及批判性条件的关注，值得精读原文。","inspiration":"该方法通过情绪提示词注入情绪状态，并设置无情绪中性基线，可借鉴用于经济决策实验中情绪处理的标准化设计｜可迁移到资产定价实验中，研究情绪如何影响投资者对风险资产的需求与定价偏差｜以LLM为被试，通过愤怒/恐惧等情绪提示词处理，测量其在模拟股票交易中的出价与风险偏好，结果与真实投资者实验数据对照"}},{"id":"2405.19313","version":2,"title":"Language Models Trained to do Arithmetic Predict Human Risky and Intertemporal Choice","zh_title":"训练做算术的语言模型预测人类风险与跨期选择","abstract":"The observed similarities in the behavior of humans and Large Language Models (LLMs) have prompted researchers to consider the potential of using LLMs as models of human cognition. However, several significant challenges must be addressed before LLMs can be legitimately regarded as cognitive models. For instance, LLMs are trained on far more data than humans typically encounter, and may have been directly trained on human data in specific cognitive tasks or aligned with human preferences. Consequently, the origins of these behavioral similarities are not well understood. In this paper, we propose a novel way to enhance the utility of LLMs as cognitive models. This approach involves (i) leveraging computationally equivalent tasks that both an LLM and a rational agent need to master for solving a cognitive problem and (ii) examining the specific task distributions required for an LLM to exhibit human-like behaviors. We apply this approach to decision-making -- specifically risky and intertemporal choice -- where the key computationally equivalent task is the arithmetic of expected value calculations. We show that an LLM pretrained on an ecologically valid arithmetic dataset, which we call Arithmetic-GPT, predicts human behavior better than many traditional cognitive models. Pretraining LLMs on ecologically valid arithmetic datasets is sufficient to produce a strong correspondence between these models and human decision-making. Our results also suggest that LLMs used as cognitive models should be carefully investigated via ablation studies of the pretraining data.","authors":["Jian-Qiao Zhu","Haijiang Yan","Thomas L. Griffiths"],"categories":["cs.AI","cs.CL","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2024-05-29","first_seen":"2024-05-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2405.19313","pdf_url":"https://arxiv.org/pdf/2405.19313","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","决策行为","认知模型"],"reason":"用LLM预测人类风险与跨期选择，与真实人类数据对照，并分析仿真有效条件，直接相…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":61,"question":"语言模型在算术任务上的预训练能否使其产生与人类相似的风险和跨期选择行为？","design":"训练一个小型语言模型（约10M参数）在合成算术数据集上（如期望值计算），提取其嵌入向量，用逻辑回归预测人类选择概率，并与传统认知模型及LLaMA3比较。","baseline":"使用真实人类在风险和跨期选择任务中的选择数据作为基准。","findings":"在生态有效的算术数据集上预训练的Arithmetic-GPT模型预测人类选择优于许多传统认知模型；仅靠算术预训练就足以产生与人类决策的强对应关系。","reliability":"论文指出需通过消融实验仔细检查预训练数据，且合成数据分布需符合生态分布才能有效，否则预测力有限。","relevance":"直接以真实人类数据为基准，用LLM仿真风险与跨期选择，并分析预训练数据分布对仿真有效性的影响，高度契合研究者对仿真可靠性及失效条件的关注。","inspiration":"该方法通过控制预训练数据的生态分布来提升模型对人类决策的预测力，值得借鉴其消融实验设计以检验数据特征对仿真效果的影响｜可迁移到消费者跨期选择研究，例如分析不同利率环境下个体的储蓄与消费决策｜以LLM为被试，处理变量为预训练数据中利率变动的分布（符合真实市场波动），结果变量为模拟的跨期选择偏好，用家庭金融调查的真实跨期选择数据作为对照基准"}},{"id":"2404.01332","version":3,"title":"Explaining Large Language Models Decisions Using Shapley Values","zh_title":"使用Shapley值解释大语言模型决策","abstract":"The emergence of large language models (LLMs) has opened up exciting possibilities for simulating human behavior and cognitive processes, with potential applications in various domains, including marketing research and consumer behavior analysis. However, the validity of utilizing LLMs as stand-ins for human subjects remains uncertain due to glaring divergences that suggest fundamentally different underlying processes at play and the sensitivity of LLM responses to prompt variations. This paper presents a novel approach based on Shapley values from cooperative game theory to interpret LLM behavior and quantify the relative contribution of each prompt component to the model's output. Through two applications - a discrete choice experiment and an investigation of cognitive biases - we demonstrate how the Shapley value method can uncover what we term \"token noise\" effects, a phenomenon where LLM decisions are disproportionately influenced by tokens providing minimal informative content. This phenomenon raises concerns about the robustness and generalizability of insights obtained from LLMs in the context of human behavior simulation. Our model-agnostic approach extends its utility to proprietary LLMs, providing a valuable tool for practitioners and researchers to strategically optimize prompts and mitigate apparent cognitive biases. Our findings underscore the need for a more nuanced understanding of the factors driving LLM responses before relying on them as substitutes for human subjects in survey settings. We emphasize the importance of researchers reporting results conditioned on specific prompt templates and exercising caution when drawing parallels between human behavior and LLMs.","authors":["Behnam Mohammadi"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2024-03-29","first_seen":"2024-03-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2404.01332","pdf_url":"https://arxiv.org/pdf/2404.01332","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","Shapley值","认知偏差"],"reason":"用LLM仿真人类选择与认知偏差，并与真实人类数据对照，揭示仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":32,"question":"如何利用Shapley值解释大语言模型决策，并揭示提示词中无信息量token对模型输出的不成比例影响（即“token噪声”效应）？","design":"本研究并非直接用LLM仿真人类，而是提出一种基于合作博弈Shapley值的模型无关解释方法，将提示词各组成部分视为“玩家”，量化其对LLM输出的相对贡献。通过两个应用展示：离散选择实验（航班选择）和认知偏差调查，分析提示词中token的影响。","baseline":"无对照","findings":"发现“token噪声”现象：LLM决策受到无信息量token（如冠词、介词、甚至“flight”等单词）的过度影响，且对格式变化（如换行符）高度敏感，导致选择概率出现混沌波动。这引发了对LLM仿真人类行为稳健性和泛化性的严重担忧。","reliability":"论文指出LLM对提示词变化高度敏感，无信息量token可显著改变输出，因此基于LLM的人类行为仿真在调查环境中不可靠，需谨慎解读，并建议研究者报告基于特定提示模板的条件结果。","relevance":"该研究直接批判了用LLM替代人类被试的可靠性，揭示了仿真失效的具体机制（token噪声），与研究者关注的仿真失效条件高度相关，值得精读以深入理解LLM行为偏差的来源。","inspiration":"可借鉴Shapley值分解方法，量化提示词各成分对LLM输出的影响，用于诊断和优化经济金融实验中的提示设计。｜可迁移到消费者金融决策仿真（如贷款选择、投资偏好），分析提示词中无关信息如何扭曲LLM的“偏好”。｜以GPT-4为被试，设计不同贷款方案的离散选择实验，在提示中系统变化无信息量token（如换行符、冠词），用Shapley值量化其影响，并以真实消费者信贷选择数据为基准，检验LLM仿真偏差。"}},{"id":"2403.15281","version":1,"title":"Measuring Gender and Racial Biases in Large Language Models","zh_title":"测量大语言模型中的性别与种族偏见","abstract":"In traditional decision making processes, social biases of human decision makers can lead to unequal economic outcomes for underrepresented social groups, such as women, racial or ethnic minorities. Recently, the increasing popularity of Large language model based artificial intelligence suggests a potential transition from human to AI based decision making. How would this impact the distributional outcomes across social groups? Here we investigate the gender and racial biases of OpenAIs GPT, a widely used LLM, in a high stakes decision making setting, specifically assessing entry level job candidates from diverse social groups. Instructing GPT to score approximately 361000 resumes with randomized social identities, we find that the LLM awards higher assessment scores for female candidates with similar work experience, education, and skills, while lower scores for black male candidates with comparable qualifications. These biases may result in a 1 or 2 percentage point difference in hiring probabilities for otherwise similar candidates at a certain threshold and are consistent across various job positions and subsamples. Meanwhile, we also find stronger pro female and weaker anti black male patterns in democratic states. Our results demonstrate that this LLM based AI system has the potential to mitigate the gender bias, but it may not necessarily cure the racial bias. Further research is needed to comprehend the root causes of these outcomes and develop strategies to minimize the remaining biases in AI systems. As AI based decision making tools are increasingly employed across diverse domains, our findings underscore the necessity of understanding and addressing the potential unequal outcomes to ensure equitable outcomes across social groups.","authors":["Jiafu An","Difang Huang","Chen Lin","Mingzhu Tai"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2024-03-22","first_seen":"2024-03-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2403.15281","pdf_url":"https://arxiv.org/pdf/2403.15281","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","偏见测量","劳动力市场"],"reason":"用LLM替代人类决策者评估简历，测量偏见并与真实人类数据对照，涉及劳动力市场政…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:47","error":null,"has_summary":true,"summary":{"generated_at":"2024-03-22","rank":1,"question":"在招聘决策中，GPT-3.5对不同性别和种族的求职者是否存在评分偏差？","design":"使用GPT-3.5模型扮演招聘决策者，对约36.1万份随机生成的工作经验、教育背景和技能组合的虚构简历进行评分（0-100分），简历随机分配带有性别和种族标识的姓名，以测量模型对不同社会群体的评分差异。","baseline":"无对照","findings":"GPT-3.5对女性求职者（无论种族）给予显著高于白人男性的评分，但对黑人男性给予显著低于白人男性的评分；在民主党州，亲女性偏差更强，反黑人男性偏差较弱。","reliability":"论文未讨论","relevance":"该研究直接使用LLM替代人类决策者进行高利害决策实验，测量了性别和种族偏见，虽未提供真实人类对照数据，但为评估LLM在劳动力市场决策中的偏差提供了重要证据，值得精读以了解仿真设计细节和偏差模式。","inspiration":"借鉴其通过大规模随机生成简历并操控姓名标识社会身份的实验设计，可精确分离LLM的偏见效应｜可迁移到信贷审批歧视研究，用LLM模拟信贷员评估贷款申请，操控申请人性别/种族｜用GPT-4扮演信贷员，对随机生成并分配不同种族姓名的贷款申请进行评分，结果变量为贷款批准概率，以真实银行信贷数据中的种族差异作为对照基准。"}},{"id":"2402.18144","version":1,"title":"Random Silicon Sampling: Simulating Human Sub-Population Opinion Using a Large Language Model Based on Group-Level Demographic Information","zh_title":"随机硅采样：基于群体人口统计信息用大语言模型模拟人类子群体意见","abstract":"Large language models exhibit societal biases associated with demographic information, including race, gender, and others. Endowing such language models with personalities based on demographic data can enable generating opinions that align with those of humans. Building on this idea, we propose \"random silicon sampling,\" a method to emulate the opinions of the human population sub-group. Our study analyzed 1) a language model that generates the survey responses that correspond with a human group based solely on its demographic distribution and 2) the applicability of our methodology across various demographic subgroups and thematic questions. Through random silicon sampling and using only group-level demographic information, we discovered that language models can generate response distributions that are remarkably similar to the actual U.S. public opinion polls. Moreover, we found that the replicability of language models varies depending on the demographic group and topic of the question, and this can be attributed to inherent societal biases in the models. Our findings demonstrate the feasibility of mirroring a group's opinion using only demographic distribution and elucidate the effect of social biases in language models on such simulations.","authors":["Seungjong Sun","Eungu Lee","Dongyan Nan","Xiangying Zhao","Wonbyung Lee","Bernard J. Jansen","Jang Hyun Kim"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2024-02-28","first_seen":"2024-02-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2402.18144","pdf_url":"https://arxiv.org/pdf/2402.18144","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM人类仿真","意见模拟","算法偏差"],"reason":"用LLM基于人口统计分布模拟人群意见，并与真实民调对照，评估偏差与可复现性，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"能否仅基于群体层面的人口统计分布，利用大语言模型生成与真实人群意见分布相似的调查回答？","design":"使用GPT-3.5等LLM，根据目标人群的群体人口统计分布随机生成合成个体（random silicon sample），将个体人口统计信息与调查问题一起作为提示输入模型，收集回答并汇总为群体意见分布。","baseline":"以美国皮尤研究中心等机构的真实民意调查数据作为对照基准。","findings":"仅用群体人口统计信息，LLM生成的回答分布与真实民调高度相似；但复现性因人口子群和问题主题而异，这种差异可归因于模型固有的社会偏见。","reliability":"论文指出复现性受目标群体和问题主题影响，模型对特定人群和话题的偏见会导致仿真失效，且方法依赖群体分布假设，未考虑个体层面差异。","relevance":"该研究直接探索用LLM替代人类被试进行民意调查仿真，并与真实数据严格对照，评估了可靠性与偏差条件，完全契合研究者对LLM人类仿真实验、基准对照和失效分析的关注，值得精读原文。","inspiration":"借鉴其仅用群体分布生成合成样本并汇总意见的方法，可低成本构建虚拟被试池进行政策态度预测试。｜可迁移到经济政策评估场景，如模拟不同收入群体对税收改革的态度分布。｜以收入、教育、地区等群体分布生成虚拟纳税人，施加税收政策描述作为处理，测量支持率，用真实社会调查数据做对照。"}},{"id":"2402.01766","version":3,"title":"LLM Voting: Human Choices and AI Collective Decision Making","zh_title":"LLM投票：人类选择与AI集体决策","abstract":"This paper investigates the voting behaviors of Large Language Models (LLMs), specifically GPT-4 and LLaMA-2, their biases, and how they align with human voting patterns. Our methodology involved using a dataset from a human voting experiment to establish a baseline for human preferences and conducting a corresponding experiment with LLM agents. We observed that the choice of voting methods and the presentation order influenced LLM voting outcomes. We found that varying the persona can reduce some of these biases and enhance alignment with human choices. While the Chain-of-Thought approach did not improve prediction accuracy, it has potential for AI explainability in the voting process. We also identified a trade-off between preference diversity and alignment accuracy in LLMs, influenced by different temperature settings. Our findings indicate that LLMs may lead to less diverse collective outcomes and biased assumptions when used in voting scenarios, emphasizing the need for cautious integration of LLMs into democratic processes.","authors":["Joshua C. Yang","Damian Dailisan","Marcin Korecki","Carina I. Hausladen","Dirk Helbing"],"categories":["cs.CL","cs.AI","cs.CY","cs.LG","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2024-01-31","first_seen":"2024-01-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2402.01766","pdf_url":"https://arxiv.org/pdf/2402.01766","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","投票行为","人类对照"],"reason":"用LLM复现人类投票实验，有真实人类数据对照，涉及集体决策与偏差评估。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":23,"question":"LLM（GPT-4和LLaMA-2）在参与式预算投票中的行为与人类投票模式的对齐程度如何，存在哪些偏差？","design":"使用GPT-4 Turbo和LLaMA-2 70B模型模拟180名人类被试，在相同的参与式预算投票实验中，对24个城市项目进行投票。实验操纵了四种投票方法（批准投票、5-批准投票、累积投票、排序投票）和项目呈现顺序，并测试了角色设定（persona）和思维链提示的影响。结果变量包括聚合偏好相似度（Kendall's τ）、个体投票相似度（Jaccard）和偏好多样性。","baseline":"来自Yang et al. (2024)的180名大学生在苏黎世参与式预算在线实验中的真实投票数据。","findings":"投票方法和呈现顺序会影响LLM的投票结果；改变角色设定可以减少偏差并提高与人类选择的一致性。思维链提示未提高预测准确性，但有助于投票过程的可解释性；温度设置导致偏好多样性与对齐准确性之间存在权衡。","reliability":"论文指出LLM在投票场景中可能导致集体结果多样性降低和偏差假设，强调需谨慎将LLM整合进民主过程；未深入讨论其他失效条件。","relevance":"该研究直接使用LLM复现人类投票实验，有真实人类数据对照，评估了仿真可靠性、偏差及对齐方法，高度契合研究者对经济学实验和政策评估场景中LLM仿真批判性分析的兴趣，值得精读原文。","inspiration":"该方法借鉴了多模型对比（GPT-4 vs. LLaMA-2）和提示工程干预（角色设定、思维链）来评估LLM与人类行为对齐程度的设计，并揭示了温度参数在多样性与准确性间的权衡。｜可迁移到政策公告的预期形成实验，例如研究央行沟通对通胀预期的影响。｜以LLM作为被试，模拟不同措辞和框架的央行声明（处理），测量其预测通胀的分布和锚定效应（结果变量），并与真实家庭或专家调查数据（如密歇根消费者调查）进行对照。"}},{"id":"2304.03442","version":2,"title":"Generative Agents: Interactive Simulacra of Human Behavior","zh_title":"生成式智能体：人类行为的交互式模拟","abstract":"Believable proxies of human behavior can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication to prototyping tools. In this paper, we introduce generative agents--computational software agents that simulate believable human behavior. Generative agents wake up, cook breakfast, and head to work; artists paint, while authors write; they form opinions, notice each other, and initiate conversations; they remember and reflect on days past as they plan the next day. To enable generative agents, we describe an architecture that extends a large language model to store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior. We instantiate generative agents to populate an interactive sandbox environment inspired by The Sims, where end users can interact with a small town of twenty five agents using natural language. In an evaluation, these generative agents produce believable individual and emergent social behaviors: for example, starting with only a single user-specified notion that one agent wants to throw a Valentine's Day party, the agents autonomously spread invitations to the party over the next two days, make new acquaintances, ask each other out on dates to the party, and coordinate to show up for the party together at the right time. We demonstrate through ablation that the components of our agent architecture--observation, planning, and reflection--each contribute critically to the believability of agent behavior. By fusing large language models with computational, interactive agents, this work introduces architectural and interaction patterns for enabling believable simulations of human behavior.","authors":["Joon Sung Park","Joseph C. O'Brien","Carrie J. Cai","Meredith Ringel Morris","Percy Liang","Michael S. Bernstein"],"categories":["cs.HC","cs.AI","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2023-04-07","first_seen":"2023-04-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2304.03442","pdf_url":"https://arxiv.org/pdf/2304.03442","source_feed":"api","score":8,"bucket":"selected","rubric_hits":["A3","B1"],"tags":["LLM社会模拟","人类行为仿真","智能体架构"],"reason":"用LLM agent模拟小镇社会行为，有真实人类行为对照，但非严格实验或政策评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":31,"question":"如何利用大语言模型构建能够产生可信个体和涌现社会行为的生成式智能体？","design":"使用大语言模型（ChatGPT）构建25个生成式智能体，置于类似《模拟人生》的沙盒环境中。每个智能体拥有记忆流（存储经历）、反思（合成高层推断）和规划（生成行动计划）模块。通过用户指定一个智能体想举办情人节派对这一初始条件，观察智能体自主产生的行为，如传播邀请、约会、协调参加派对等。","baseline":"无对照","findings":"生成式智能体能够产生可信的个体行为和涌现的社会行为，如信息扩散、关系建立和群体协调。消融实验表明，记忆流、反思和规划三个组件对行为可信性均有关键贡献。","reliability":"论文未讨论","relevance":"该研究展示了LLM智能体在模拟社会互动和涌现行为方面的潜力，但缺乏与真实人类数据的严格对照，且非经济学实验或政策评估场景，与研究者关注的经济学实验和政策评估的直接相关性有限。","inspiration":"与经济金融研究关联不大"}},{"id":"2301.07543","version":2,"title":"Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?","zh_title":"作为模拟经济主体的大语言模型：我们能从Homo Silicus中学到什么？","abstract":"We argue that newly-developed large language models (LLMs), because of how they are trained and designed, are implicit computational models of humans -- a Homo silicus. LLMs can be used like economists use Homo economicus: they can be given endowments, information, preferences, and so on, and then their behavior can be explored in scenarios via simulation. Experiments using this approach, derived from Charness and Rabin (2002), Kahneman et al. (1986), Samuelson and Zeckhauser (1988), Oprea (2024b), and Horton (2025), show qualitatively similar results to the original, and when they differ, it is often generative for future research. We discuss potential applications, conceptual issues, and why this approach can inform the study of humans.","authors":["John J. Horton","Apostolos Filippas","Benjamin S. Manning"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2023-01-18","first_seen":"2023-01-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2301.07543","pdf_url":"https://arxiv.org/pdf/2301.07543","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","经济实验","人类行为对照"],"reason":"直接用LLM模拟经济实验并与真实人类数据对照，核心相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":18,"question":"大语言模型能否作为人类经济行为的计算模型（Homo silicus），通过仿真实验复现经典经济实验结果，并用于理解人类行为？","design":"使用大语言模型（如GPT系列）作为AI智能体，赋予其禀赋、信息、偏好等，通过文本提示模拟五种经典经济实验场景：公平性判断（Kahneman et al., 1986）、独裁者博弈（Charness and Rabin, 2002）、现状偏差（Samuelson and Zeckhauser, 1988）、风险态度（Oprea, 2024b）和招聘场景（Horton, 2025），测量AI智能体的回答或选择行为。","baseline":"对照的真实人类数据来自上述五篇原始文献中的人类被试行为结果。","findings":"AI仿真结果在定性上与原始人类实验相似，例如公平判断受涨价幅度和政治倾向影响、独裁者博弈中赋予不同社会偏好会改变选择、现状偏差可被复现；当结果存在差异时，往往能为未来研究提供新思路。","reliability":"论文指出LLM训练数据可能包含已发表研究结果，导致仿真仅机械复述记忆而非真正模拟人类行为；训练语料、模型不透明性和仿真泛化能力均构成局限。","relevance":"该研究直接使用LLM进行经济实验仿真并与真实人类数据对照，涵盖公平、博弈、偏差、风险决策等多个经典主题，并讨论了仿真失效条件，高度契合研究者对LLM人类仿真可靠性及批判性评估的兴趣，值得精读原文。","inspiration":"该方法借鉴了用文本提示赋予LLM特定禀赋、信息和社会偏好来模拟经济决策，并直接与经典实验的人类基准数据对照｜可迁移到资产定价实验，如研究投资者在泡沫或崩盘情境下的交易行为与风险偏好｜以LLM为被试，通过提示设定初始财富、市场信息和风险态度，测量其买卖报价与持仓变化，对照Smith et al. (1988)等经典资产泡沫实验的人类数据"}},{"id":"2209.06899","version":1,"title":"Out of One, Many: Using Language Models to Simulate Human Samples","zh_title":"一生万物：使用语言模型模拟人类样本","abstract":"We propose and explore the possibility that language models can be studied as effective proxies for specific human sub-populations in social science research. Practical and research applications of artificial intelligence tools have sometimes been limited by problematic biases (such as racism or sexism), which are often treated as uniform properties of the models. We show that the \"algorithmic bias\" within one such tool -- the GPT-3 language model -- is instead both fine-grained and demographically correlated, meaning that proper conditioning will cause it to accurately emulate response distributions from a wide variety of human subgroups. We term this property \"algorithmic fidelity\" and explore its extent in GPT-3. We create \"silicon samples\" by conditioning the model on thousands of socio-demographic backstories from real human participants in multiple large surveys conducted in the United States. We then compare the silicon and human samples to demonstrate that the information contained in GPT-3 goes far beyond surface similarity. It is nuanced, multifaceted, and reflects the complex interplay between ideas, attitudes, and socio-cultural context that characterize human attitudes. We suggest that language models with sufficient algorithmic fidelity thus constitute a novel and powerful tool to advance understanding of humans and society across a variety of disciplines.","authors":["Lisa P. Argyle","Ethan C. Busby","Nancy Fulda","Joshua Gubler","Christopher Rytting","David Wingate"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"new","date":"2022-09-14","first_seen":"2022-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2209.06899","pdf_url":"https://arxiv.org/pdf/2209.06899","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["LLM仿真","算法保真度","社会调查"],"reason":"直接提出用GPT-3模拟人类子群体，并与真实调查数据对照，验证算法保真度，高度…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":4,"question":"大型语言模型（如GPT-3）能否通过条件化生成，准确模拟特定人类子群体的态度和反应分布，从而作为社会科学研究中的人类被试替代品？","design":"使用GPT-3模型，通过输入真实调查参与者的社会人口背景故事（来自ANES等大型调查）作为条件，生成“硅样本”虚拟被试，然后让这些虚拟被试完成与人类相同的任务（自由联想、投票预测、封闭式问题），比较硅样本与人类样本的响应分布。","baseline":"2012、2016、2020年美国国家选举研究（ANES）和Rothschild等人的“Pigeonholing Partisans”数据中真实人类参与者的调查回答。","findings":"GPT-3的算法偏差并非单一宏观属性，而是细粒度且与人口统计特征相关，通过适当条件化可精确模拟多种人类子群体的响应分布。硅样本与人类样本在态度、观念和社会文化背景的复杂交互模式上高度一致，表明GPT-3具有较高的“算法保真度”。","reliability":"论文未讨论","relevance":"该研究直接探索用LLM替代人类被试进行仿真实验，并与真实调查数据严格对照，验证了算法保真度，高度契合研究者对LLM人类仿真可靠性及基准对照的关注，值得精读原文。","inspiration":"借鉴其“硅采样”方法：用真实个体的多维人口背景作为条件提示，生成虚拟被试并测量其态度/行为，再与人类基准数据对比以评估仿真效度。｜可迁移至消费者信心调查或政策偏好预测，例如模拟不同收入、教育、地域群体对通胀预期或税收政策的反应。｜以GPT-4为被试，输入来自美国消费者财务调查（SCF）的家庭人口与财务背景，生成虚拟消费者，询问其未来一年通胀预期，以密歇根大学消费者调查的微观数据作为人类基准，比较分布与相关性。"}},{"id":"2208.10264","version":5,"title":"Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies","zh_title":"使用大语言模型模拟多个人类并复现人类被试研究","abstract":"We introduce a new type of test, called a Turing Experiment (TE), for evaluating to what extent a given language model, such as GPT models, can simulate different aspects of human behavior. A TE can also reveal consistent distortions in a language model's simulation of a specific human behavior. Unlike the Turing Test, which involves simulating a single arbitrary individual, a TE requires simulating a representative sample of participants in human subject research. We carry out TEs that attempt to replicate well-established findings from prior studies. We design a methodology for simulating TEs and illustrate its use to compare how well different language models are able to reproduce classic economic, psycholinguistic, and social psychology experiments: Ultimatum Game, Garden Path Sentences, Milgram Shock Experiment, and Wisdom of Crowds. In the first three TEs, the existing findings were replicated using recent models, while the last TE reveals a \"hyper-accuracy distortion\" present in some language models (including ChatGPT and GPT-4), which could affect downstream applications in education and the arts.","authors":["Gati Aher","Rosa I. Arriaga","Adam Tauman Kalai"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2022-08-18","first_seen":"2022-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2208.10264","pdf_url":"https://arxiv.org/pdf/2208.10264","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人类实验复现","行为经济学"],"reason":"直接复现经典人类实验，用LLM模拟被试并与真实人类数据对照，评估仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":5,"question":"如何系统评估语言模型在模拟人类行为时的忠实程度与系统性扭曲？","design":"提出图灵实验（TE）方法，使用GPT等语言模型，通过零样本提示模拟具有不同姓名和性别称谓的多样化被试样本，在最后通牒博弈、花园路径句、米尔格拉姆电击实验和群体智慧四个经典实验中施加相应刺激，测量接受/拒绝、语法判断等结果变量。","baseline":"对照各经典实验已有的真实人类被试研究结果。","findings":"在前三个TE中，近期模型成功复现了已有发现；在群体智慧TE中，部分模型（包括ChatGPT和GPT-4）表现出“超准确性扭曲”，即模拟的群体估计过于准确，偏离了真实人类群体的典型误差模式。","reliability":"论文指出，零样本要求难以完全保证，因为预训练语料可能已包含相关实验数据；此外，仅用姓名和性别称谓模拟多样性可能不足以捕捉真实人群差异。","relevance":"该研究直接复现经典人类实验，用LLM模拟被试并与真实人类数据对照，评估仿真偏差，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"借鉴其通过姓名和称谓简单操控被试身份以模拟多样性的设计，以及用经典实验范式作为基准测试LLM行为复现能力的方法。｜可迁移到行为经济学中的最后通牒博弈、信任博弈等实验，检验LLM是否能复现真实人类的公平偏好或互惠行为。｜以GPT-4为被试，模拟不同姓名（暗示种族/性别）的个体在最后通牒博弈中的响应，处理为不同的提议金额，结果变量为接受/拒绝，对照真实人类实验的元分析数据，评估LLM是否复现已知的公平偏好及群体差异。"}}]}