{"api_version":"v1","generated_at":"2026-09-29T13:00:45","count":766,"scope":"with_summary","papers":[{"id":"2609.30563","version":1,"title":"Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content","zh_title":"少思考以更好地模拟：直觉提示提升LLM智能体模拟个体社交媒体反应（包括不熟悉内容）","abstract":"Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour and has paid little attention to whether an agent behaves in line with the profile it was given. The present study profiled eight Serbian participants through a questionnaire, a deep interview, and a written self-presentation, recorded their reactions to sixty-eight social media posts, and asked four language models to predict those reactions under five prompt conditions varying profile content and instruction style. Attitudinal content improved prediction over demographic backstories by a wide margin. Agents matched their stated profiles more closely than participants matched their own survey answers, and consistency proved unrelated to fidelity once profile information was present. Instructing models to respond intuitively and immediately rather than analytically gave the highest fidelity of any condition and cut the compression of individual differences from seven times the human level to three. The advantage held on posts about topics the questionnaire never raised, where that condition reached the highest fidelity of any setup and beat a crowd baseline by a wide margin, which suggests that agents prompted this way could serve as general-purpose simulated users rather than specialists on the topics they were profiled for. Results may bear implications for the development of language models, because intuition-based setups appear better suited to some tasks than reasoning-based ones.","authors":["Ljubisa Bojic","Tijana Stanic","Joerg Matthes","Agariadne Dwinggo Samala","Bojana Dinic","Jue Wang"],"categories":["cs.AI","cs.CL","cs.HC","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30563","pdf_url":"https://arxiv.org/pdf/2609.30563","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","人类行为对照","提示策略"],"reason":"用LLM预测真实个体社交媒体反应，与人类数据对照，评估提示策略对仿真保真度的影…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":1,"question":"在社交媒体反应预测中，不同的提示设计（档案内容与指令风格）如何影响大语言模型模拟特定个体反应的保真度？","design":"用四个大语言模型扮演八名塞尔维亚参与者，基于问卷、深度访谈和书面自我呈现构建个人档案，在五种提示条件下（变化档案内容和指令风格，如人口学背景、态度内容、直觉式回应等）预测他们对68条社交媒体帖子的反应，测量预测与真实反应的一致性。","baseline":"八名塞尔维亚参与者的真实社交媒体反应数据，以及他们自己的问卷回答（用于比较自我一致性）。","findings":"态度性档案内容比人口学背景大幅提升预测准确度；直觉式指令（要求模型凭直觉立即回应而非分析式推理）在所有条件中保真度最高，并将个体差异压缩从人类的七倍降至三倍。","reliability":"论文未讨论","relevance":"该研究直接检验LLM模拟个体行为的保真度，与人类真实反应对照，并揭示提示策略对仿真偏差的影响，对评估LLM作为人类被试替代品的可靠性具有关键参考价值。","inspiration":"值得借鉴的是通过改变提示指令风格（直觉式 vs 分析式）来操纵模型的认知模式，并测量其对个体差异保真度的影响｜可迁移到消费者金融决策实验，如模拟个体在信贷选择或储蓄行为中的异质性反应｜用LLM扮演不同风险偏好和金融素养的消费者，处理为直觉式或分析式提示，结果变量为信贷产品选择或投资决策，与真实消费者调查或实验数据对照。"}},{"id":"2609.30883","version":1,"title":"Warned alike, AI agents avoid the less-crowded road while people take it","zh_title":"同样被警告，AI智能体避开较不拥挤的道路而人类选择它","abstract":"AI agents built on a few shared models increasingly act for many people. A shared forecast about others can align their choices and change how scarce capacity is allocated. We tested this feedback in a two-road congestion game. Adding one sentence warning that others might follow a routing tip made populations of 50 GPT agents crowd one road while avoiding the nearly empty alternative. Average travel time rose from 64 to 95 min, although any crowded-road agent could have saved 69 min by switching alone. The warning discouraged the very move it predicted. The pattern persisted for 100 rounds. Two other model families shifted the same way without locking onto one road. Twelve all-human groups (240 participants) stayed near balance under numerical reports or the tip and warning. In 24 mixed groups with a further 240 participants, imbalance grew with the share of agents in the registered analysis, while people increasingly took the road the agents avoided. Collective costs stayed below the allagent reference, but with 15 agents and 5 humans, agent seats averaged 80 min, compared with 44 min for human seats. Shared forecasts can thus sustain collective inefficiency among similar agents. A better group average can also hide an unequal burden. Evaluations of AI agents that share resources should test populations, treat messages as interventions and report who bears the costs.","authors":["Takahiro Ezaki","Naoto Imura","Katsuhiro Nishinari"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30883","pdf_url":"https://arxiv.org/pdf/2609.30883","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","拥堵博弈","人机对照"],"reason":"用GPT agent群体模拟拥堵博弈，并与240名人类被试对照，发现警告导致a…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":2,"question":"共享预测信息（警告）是否会导致AI智能体在拥堵博弈中持续选择拥挤道路，从而造成集体低效？","design":"用50个GPT智能体（gpt-5.4-mini）模拟通勤者，在双路径拥堵博弈中，操纵每日广播信息（无报告、数值报告、提示、提示+警告），测量路径不平衡度、切换比例和平均旅行时间。","baseline":"12个全人类组（240名参与者）在相同博弈和广播条件下的行为数据，以及24个混合组（240名参与者）与智能体共存时的行为。","findings":"警告导致GPT智能体群体持续拥挤一条道路，平均旅行时间从64分钟升至95分钟，尽管个体切换可节省69分钟；全人类组保持接近均衡，混合组中人类更多选择智能体避开的道路，且智能体承担更高成本。","reliability":"论文指出共享预测可导致相似智能体间的集体低效，且群体平均成本可能掩盖负担不均；评估共享资源的AI智能体应测试群体、将消息视为干预并报告成本承担者。","relevance":"高度相关：该研究用LLM智能体模拟人类在拥堵博弈中的决策，并与真实人类数据对照，揭示了共享信息对集体行为的负面影响，直接回应了仿真可靠性与偏差问题。","inspiration":"值得借鉴的是将公共信息（警告）作为干预变量，观察其对群体决策动态的影响，并设置全人类和混合组对照以分离智能体特有行为。｜可迁移到政策公告的预期形成场景，如央行沟通对金融市场参与者行为的影响。｜设计：用LLM智能体模拟投资者，处理为央行发布的不同措辞的前瞻指引，结果变量为资产配置集中度和市场波动率，对照真实投资者在类似公告下的交易数据。"}},{"id":"2609.30896","version":1,"title":"Large language models underestimate and partly misrepresent cultural variation in everyday norms","zh_title":"大语言模型低估并部分误现日常规范的文化差异","abstract":"A key aspect of culture is a society's norms about everyday behavior. How accurately do large language models (LLMs) represent cultural differences in such norms? To answer this question we used the recent Global Study of Everyday Norms (GSEN), which collected ratings of 150 scenarios in 90 societies, as the human benchmark. We prompted GPT-5 to estimate each society's average rating for every scenario, and later repeated the benchmark in three other LLMs: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro. Compared to GSEN estimates, all four LLMs misrepresented cultural variation in two ways. First, they greatly underestimated its magnitude, estimating differences between societies to be, on average, less than half their measured size. Second, for many scenarios the LLMs poorly identified the pattern of variation, that is, which societies judged the behavior less acceptable and which societies judged it more acceptable. The pattern of variation was identified better for scenarios that elicit concerns about vulgarity, especially scenarios involving kissing and flirting. We also found that norms in more developed societies tended to be estimated somewhat more accurately, and that prompting in local survey languages rather than English produced only a modest improvement in accuracy. Local-language prompting also reduced, but did not remove, the underestimation of between-society differences. Cultural differences in everyday norms are only weakly and unevenly represented by LLMs.","authors":["Kimmo Eriksson","Irina Vartanova","Pontus Strimling"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30896","pdf_url":"https://arxiv.org/pdf/2609.30896","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["文化规范","仿真偏差","人类数据对照"],"reason":"用LLM估计社会规范并与真实人类调查数据对照，评估仿真偏差","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":3,"question":"大语言模型能否准确表示不同社会之间日常行为规范的文化差异？","design":"用 GPT-5、GPT-5.4、Claude Opus 4.6 和 Gemini 3.1 Pro 四个 LLM，以英语或当地语言提示，估计 90 个社会对 150 个日常行为场景的平均可接受度评分，并与 GSEN 调查数据比较。","baseline":"全球日常规范研究（GSEN）中 90 个社会超过 25,000 名参与者对 150 个场景的评分。","findings":"四个 LLM 都大幅低估了社会间规范差异的幅度，估计的差异平均不到实际测量的一半；对于许多场景，LLM 识别差异模式的能力较差，但在涉及粗俗（如亲吻和调情）的场景上表现较好。","reliability":"论文指出 LLM 对较发达社会的规范估计更准确，用当地语言提示仅带来适度改进，且未能消除对社会间差异的低估；LLM 对文化差异的表示既弱且不均衡。","relevance":"该研究直接评估 LLM 在跨文化社会规范仿真中的偏差，与研究者关注的人类仿真可靠性及失效条件高度契合，值得精读原文以了解具体偏差模式和测量方法。","inspiration":"借鉴其用真实大规模跨国调查作为基准、将仿真误差分解为幅度和模式两个维度的方法，并检验提示语言等处理变量的影响。｜可迁移到跨国消费者行为或政策接受度的仿真研究，例如不同国家居民对环保政策、金融产品条款或广告伦理的接受度差异。｜以 LLM 模拟不同国家被试，对一系列政策或产品场景给出接受度评分，处理变量为提示语言（英语 vs 当地语言），结果变量为评分，与真实跨国调查数据（如世界价值观调查或特定政策民意调查）对照，评估 LLM 仿真的幅度压缩和模式偏差。"}},{"id":"2604.20050","version":4,"title":"Information Aggregation with AI Agents","zh_title":"AI代理的信息聚合研究","abstract":"Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing price movements? We conduct a controlled experiment where AI agents trade in a prediction market after receiving private signals, across four information structures of increasing complexity. We find that although the median market is effective at aggregating information in the easy information structures, performance deteriorates in the harder structures, suggesting that AI agents struggle in environments where more than two levels of interactive reasoning are required, a ceiling close to the one documented in human subjects. Consistent with our theoretical predictions, market accuracy does not improve from allowing cheap talk communication, changing the duration of the market, or strategic prompting; initial price has little average effect but matters in the very hard structure. We also find that ``smarter'' AI agents perform better at aggregation and are more profitable. Surprisingly, giving them feedback about past performance does not improve aggregation. A further wave of markets, run three months later with capability-frontier models, aggregates information more often in the three easier structures but not in the hardest one, where higher capability replaces markets that are confidently wrong with markets that hedge near 0.5.","authors":["Spyros Galanis"],"categories":["econ.GN","cs.AI","cs.GT","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-04-21","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2604.20050","pdf_url":"https://arxiv.org/pdf/2604.20050","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","信息聚合","行为实验"],"reason":"用AI代理模拟人类交易行为，并与人类被试结果对照，评估信息聚合能力。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":4,"question":"AI 代理能否通过交易聚合分散的私人信息，并像人类一样通过观察价格变动推断他人知识？","design":"用八种大语言模型（Claude Haiku 3.5/4.5、Gemini 2.5/3 Flash、GPT-4o/5 mini、gemma3:4b、qwen3:8b）组成三人交易团队，在四种难度递增的信息结构中交易预测市场证券；处理包括允许廉价谈话、战略提示、初始价格（0.3/0.5/0.7）、市场时长（3/6/9轮）和反馈信息，共144种配置，每种至少运行12次，生成1772个市场；结果变量为市场准确率（价格是否接近真实价值）和利润。","baseline":"人类被试在类似信息结构中的推理层级上限（约两级交互推理），来自已有实验文献。","findings":"在简单信息结构中，中位数市场能有效聚合信息，但在需要超过两级交互推理的困难结构中表现恶化；允许廉价谈话、改变市场时长或战略提示均未提高市场准确率，初始价格平均影响小但在极难结构中重要。更聪明的AI代理聚合更好且更盈利，但反馈过去表现未改善聚合；三个月后用前沿模型重跑，在三个较易结构中聚合更频繁，但在最难结构中未改善，高能力模型将自信错误的市场替换为在0.5附近对冲的市场。","reliability":"论文承认AI代理在需要超过两级交互推理的环境中挣扎，且能力提升并未解决最难结构中的聚合失败；未明确讨论其他失效条件。","relevance":"该研究直接以AI代理模拟人类交易者，并与人类被试的推理层级上限对照，评估信息聚合能力，属于用LLM进行人类仿真实验并检验可靠性的核心工作，值得精读。","inspiration":"借鉴其系统操纵信息结构复杂度、市场机制参数（时长、初始价格、沟通）和模型能力来测量聚合效率与利润的做法，并设置理论基准（可分离证券的完全聚合预测）｜可迁移到资产定价实验中的信息效率研究，如内幕交易监管、分析师预测市场或央行沟通对价格发现的影响｜用不同能力LLM作为交易者，在预测市场中交易与真实宏观经济指标挂钩的证券，处理为是否允许公开评论或改变交易轮次，结果变量为价格偏离真实值的程度，并与人类实验数据（如Plott & Sunder的经典信息聚合实验）对照。"}},{"id":"2609.02729","version":2,"title":"BuildOcc: A Large Language Model Occupant Agent Platform for Building Energy Research","zh_title":"BuildOcc：用于建筑能源研究的大语言模型居住者智能体平台","abstract":"Occupants are a primary source of uncertainty in building energy consumption and management, yet existing occupant behavior models cannot capture adaptive and reasoning responses considering the occupant's personal history, current context, and the type of energy signal being delivered. This study presents BuildOcc, an open-source Python platform that grounds large language model agents in the American Time Use Survey (ATUS), a nationally representative diary dataset covering 16,684 respondents. Through BuildOcc, each simulated occupant agent can be instantiated with a demographic persona drawn from ATUS population statistics, a memory stream that accumulates and reflects on timestep-level observations, and an activity scheduler that samples empirically from ATUS time-at-activity distributions. The platform exposes a three-layer interface - Python library, REST API, and Model Context Protocol server - so that any building energy tool (EnergyPlus, Home Assistant) can integrate behavioral intelligence without bespoke coupling code. A plugin registry lets the community add new occupant strata, custom schedulers, and alternative memory backends as separate installable packages. Two validation tiers show that ATUS-grounded sampling reproduces empirically calibrated activity distributions and that demographic priors propagate into persona-consistent agent reasoning across timesteps, establishing internal consistency across strata. BuildOcc provides the building energy community with a reusable, openly available implementation of the occupant behavioral layer. BuildOcc is openly released at https://doi.org/10.5281/zenodo.21192895 under the Apache License 2.0 and installable via pip install buildocc.","authors":["Wooyoung Jung"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-28","first_seen":"2026-09-03","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2609.02729","pdf_url":"https://arxiv.org/pdf/2609.02729","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","建筑能源","ATUS数据"],"reason":"用LLM agent模拟建筑内人员行为，基于ATUS真实数据对照，属于人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":5,"question":"如何构建一个以美国时间使用调查（ATUS）为数据基础、基于大语言模型（LLM）的居住者智能体平台，用于建筑能耗研究中的居住者行为仿真？","design":"该研究开发了 BuildOcc 平台，使用 LLM 智能体模拟四类美国人口群体（在职单身成年人、退休夫妇、在职父母、无业成年人）。每个智能体包含从 ATUS 人口统计中抽取的人口特征、从 ATUS 数据中采样的活动日程、记忆流和推理引擎。在每个 15 分钟时间步，智能体根据人口特征、记忆和当前环境选择下一个活动并说明理由，同时也会对需求响应信号做出接受、拒绝或推迟的决定。平台提供 Python 库、REST API 和 MCP 服务器三层接口，便于与 EnergyPlus 等工具集成。","baseline":"美国时间使用调查（ATUS），包含 16,684 名受访者的全国代表性日记数据，其中 6,611 名受访者属于四个目标人口阶层。","findings":"第一层验证表明，活动调度器能够将 ATUS 活动分布复现到采样噪声范围内；第二层验证表明，人口先验信息能够传播到智能体的推理中，产生与人口特征一致的行为差异。","reliability":"论文承认的局限包括：每个时间步只能选择一个活动；ATUS 仅覆盖美国人口；活动仅基于主要活动（ATUS 一次只记录一个活动）；记忆重要性分数由智能体自行分配并用于检索，缺乏外部校准或反馈路径。","relevance":"该研究将 LLM 智能体与全国代表性时间使用调查数据结合，用于模拟人类行为，并进行了与真实数据的对照验证，属于人类仿真研究，但场景限定于建筑能耗领域，与经济金融问题关联度较低。","inspiration":"借鉴其将 LLM 智能体与大规模调查数据结合、通过分层抽样和记忆机制生成个体行为的方法，可用于构建具有人口代表性的经济决策仿真。｜可迁移到消费者跨期选择或家庭能源消费行为研究，例如模拟不同人口群体对动态电价或节能政策的反应。｜设计雏形：以美国消费者支出调查（CEX）或收入动态面板研究（PSID）为数据基础，构建 LLM 智能体代表不同收入阶层，施加电价上涨或补贴政策处理，结果变量为能源消费和支出变化，并与真实调查数据对照验证。"}},{"id":"2609.30940","version":1,"title":"Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms","zh_title":"LLM智能体社会中的金融脆弱性：协调失败与稳定机制","abstract":"Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, but also at the level of the systems they jointly create. We study this problem with FRAIL, a controlled experimental framework that places LLM agents in three dynamic financial environments---bank runs, debt rollover, and reward crowdfunding---where agents' decisions reshape the financial conditions faced by others. Across seven leading LLMs, we find widespread collective fragility even when no agent is instructed to destabilize the system: 77\\% of baseline bank-run episodes and 83\\% of debt-rollover episodes end in failure. We then compare three interaction mechanisms based on compensated commitments, centralized commitment agreements, and participant-led coalitions. All three improve aggregate outcomes, but no single mechanism performs best across all financial structures. Across mechanisms, successful stabilization shares a common temporal pattern: broad commitment forms early, before defensive behavior becomes self-reinforcing. Our findings show that individually capable agents do not automatically form safe financial systems, highlighting system-level evaluation and interaction design as central problems for financial AI safety. Code is available at https://anonymous.4open.science/r/FinFrail-CF26.","authors":["Zhenhao Fu","Ruipeng Xu","Qibing Ren"],"categories":["cs.AI","q-fin.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30940","pdf_url":"https://arxiv.org/pdf/2609.30940","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM智能体","金融仿真","协调失败"],"reason":"用LLM agent模拟金融系统中的协调失败，涉及经济场景，但无真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":6,"question":"多个LLM智能体在共享金融环境中是否会自发产生系统性金融脆弱性，以及如何通过交互机制设计来缓解这种脆弱性？","design":"使用FRAIL框架，将七个主流LLM作为金融决策者，分别置于银行挤兑、债务展期和奖励众筹三种动态金融环境中，通过多轮交互模拟决策，测量系统失败率（如银行倒闭、债务违约、众筹失败）和承诺形成的时间模式。","baseline":"无对照","findings":"在基线条件下，77%的银行挤兑和83%的债务展期情景以失败告终，即使没有智能体被指示破坏系统。三种交互机制（补偿承诺、集中承诺协议、参与者主导联盟）均能改善结果，但最佳机制因金融结构而异，成功稳定依赖于在防御行为自我强化之前尽早形成广泛承诺。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟金融协调失败，属于经济场景下的多智能体仿真，但缺乏真实人类数据对照，适合关注LLM仿真方法本身或金融AI安全的研究者阅读。","inspiration":"借鉴其动态金融环境设计和多智能体交互机制比较，可迁移到资产定价实验或政策公告预期形成等场景。｜可设计一个信贷审批歧视实验，用LLM扮演银行信贷员和借款人，施加不同信息披露政策作为处理，测量贷款批准率和违约率，并与真实信贷数据对照。"}},{"id":"2609.31054","version":1,"title":"Cheap, open agents make LLM pollution harder to mitigate","zh_title":"廉价开放智能体使LLM污染更难缓解","abstract":"Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing multiple response types yielding various detection checks. Fully open agents ran locally without usage fees and performed competitively with commercial alternatives. Open and commercial agents failed different sets of checks, and no single check reliably detected all agents, but open-text responses discriminated best between agents and humans. These findings identify fully open agents as a distinct risk for LLM pollution and support multilayered detection strategies emphasizing open-text analysis.","authors":["Raluca Rilla","Anne-Marie Nussberger","Rui Mata","Dirk U. Wulff"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31054","pdf_url":"https://arxiv.org/pdf/2609.31054","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM污染","调查数据","检测方法"],"reason":"研究LLM污染人类调查数据，评估检测方法，与仿真可靠性直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":7,"question":"完全开源的自主智能体是否降低了LLM污染人类调查数据的门槛，以及现有检测方法能否有效识别这些智能体？","design":"使用九种智能体配置（三种完全开源、两种混合、四种专有）自主完成同一份调查问卷，每种配置运行40次，共360次运行。智能体被赋予随机的人口统计特征（性别、年龄），并收到统一指令。调查包含多种回答类型和18项检测检查（16项通过/失败检查和2项计时测量），用于评估智能体的表现和可检测性。","baseline":"3,242名来自Prolific的美国参与者，在2025年10月20日至11月4日期间完成同一调查问卷，样本在性别、年龄和种族上大致代表美国人口。","findings":"完全开源的智能体在本地运行且无使用费用，其表现与商业替代方案相当。开源和商业智能体在不同检查上失败，没有单一检查能可靠地检测所有智能体，但开放式文本回答在区分智能体和人类方面效果最好。","reliability":"论文指出，开源和商业智能体在不同检查上失败，没有单一检查能可靠地检测所有智能体，因此需要多层检测策略，特别是强调开放式文本分析。此外，研究仅操纵了年龄和性别两个人口统计变量，未涵盖其他可能影响回答的变量；智能体被明确指示避免与可疑嵌入命令交互，这可能降低了某些检测的失败率。","relevance":"该研究直接评估了LLM智能体对人类调查数据的污染风险及检测方法，与您关注的LLM仿真可靠性和偏差问题高度相关，特别是它提供了真实人类数据作为对照，并揭示了开源智能体带来的新风险，值得精读原文以了解具体检测方法和失效模式。","inspiration":"借鉴其多层检测策略和开放式文本分析来识别LLM生成回答的方法，可迁移到经济金融领域的调查数据质量控制中。｜可应用于消费者信心调查、投资者情绪调查或政策评估中的问卷数据，检测是否存在LLM污染。｜设计：以真实人类调查数据（如密歇根大学消费者信心调查）为基准，让开源和商业LLM智能体自主完成同一问卷，比较其回答分布和开放式文本特征，并开发基于文本分析的检测指标，评估不同检测方法的敏感性和特异性。"}},{"id":"2608.27167","version":2,"title":"Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable","zh_title":"校准到知道，但未校准到行动：伪造证据使LLM智能体对不可知问题做出承诺","abstract":"An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy. Nor is it belief - stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context-fragile, and deployment needs both halves of that sentence.","authors":["Pranav Aggarwal"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-08-28","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2608.27167","pdf_url":"https://arxiv.org/pdf/2608.27167","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM决策偏差","可靠性评估","批判性研究"],"reason":"研究LLM在不可知问题上的决策偏差，评估其可靠性，批判性指出失效条件，可迁移到…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":8,"question":"在不可知（aleatoric）问题上，专业外观的证据面板是否仅凭其权威包装就能诱导LLM智能体做出方向性承诺，而非真正提供信息？","design":"研究用12个前沿LLM作为被试，在四个领域（股票、加密货币、体育、天气）构造不可知问题，施加不同证据梯度（无证据、薄面板、丰富面板、乱序面板、全伪造面板），测量模型是否做出方向性承诺（行动）以及陈述的概率，并与可回答的匹配问题对照。","baseline":"无对照","findings":"伪造证据面板与真实面板诱导的承诺率在统计上无差异，表明触发行动的是包装的权威性而非信息内容；模型的陈述概率几乎不随证据梯度变化，且比气候学基线更差，但同一模型在可回答问题上几乎完美作答。","reliability":"论文承认天气领域无密封结果，面板提供的集合预报概率可能具有真实技能，因此该领域作为不可知性工具最弱；训练出的门控在刚性响应格式下失效，需要留出推理空间。","relevance":"该研究直接评估LLM在不可知问题上的决策偏差，揭示仿真失效的特定条件（证据包装触发行动），对关注LLM仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其通过伪造证据与真实证据对比来分离信息与包装的因果设计，以及用密封结果验证不可知性并测量行动而非仅测信念的做法｜可迁移到资产定价实验或政策公告的预期形成研究，例如测试LLM代理在呈现专业外观的虚假市场数据时是否会产生过度自信的交易决策｜用LLM作为被试，随机分配真实与伪造的市场分析面板，测量其买卖决策和置信度，并与人类实验数据或历史市场结果对照，检验权威包装对决策的影响。"}},{"id":"2609.30867","version":1,"title":"Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations","zh_title":"气候政策因果评估中识别假设的证据基础审计","abstract":"Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Across 26 economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence. In a five-paper pilot with labels reconciled by two annotators, it assigns a higher risk level than the labels on 25 of the 33 assessments it completes. A rule fixed before the labels arrived removes most of this in-sample; weighted agreement stays low. ARGUS provides evidence-linked risk reports that localize potential weaknesses for expert review, without adjudicating causal claims. Code and data: https://github.com/yonghongzhang-io/ARGUS","authors":["Yonghong Zhang","Yong Xie","Isabel M. Parra","Ricardo Correia"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30867","pdf_url":"https://arxiv.org/pdf/2609.30867","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A4","B3"],"tags":["LLM审计","因果推断","方法论"],"reason":"提出LLM审计因果推断假设的方法论框架，可迁移到仿真可靠性评估","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":10,"question":"如何用大语言模型审计气候政策因果评估中双重差分法识别假设的证据支持度？","design":"本文提出 ARGUS 流水线，用大语言模型对双重差分研究的 11 个识别维度进行证据检索与充分性评估，并输出风险报告；模型仅负责检索和评估，控制流固定。","baseline":"无对照","findings":"在注入缺陷基准上，ARGUS 检出 73% 的植入缺陷，而关键词流水线仅检出 18%；在 26 篇经济学论文上，约 40% 的论文-维度评估因缺乏可检索证据而弃权。","reliability":"论文承认在真实论文上证据覆盖不足导致弃权率高，且过度严厉导致与人工标签一致性低；气候政策专用语料上的表现未测试。","relevance":"该研究提出用 LLM 审计因果推断假设的方法论框架，可迁移到仿真可靠性评估，值得阅读原文了解其证据门控与弃权机制。","inspiration":"值得借鉴的是将识别假设分解为可审计的维度，并用检索门控控制 LLM 的评估范围，避免模型自由发挥｜可迁移到政策评估中的因果推断可靠性审计，例如碳税、补贴或监管政策的效果评估｜可设计一个仿真实验：用 LLM 扮演审稿人，对一组已发表的经济学论文的识别假设进行审计，处理是提供不同质量的证据，结果变量是风险评级，并与人类专家评级对照。"}},{"id":"2609.31245","version":1,"title":"RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models","zh_title":"RupeeBias：审计印度经济指导中大语言模型的人口统计偏差","abstract":"Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social biases, and biased economic guidance may influence what users believe they are worth, what they ask for, and what they ultimately accept. This risk is especially salient in India, where economic outcomes are shaped by demographic categories such as caste and urban-rural location. Existing LLM bias benchmarks, however, are largely designed around Western demographic categories and therefore miss key axes of economic disparity in the Indian context. We introduce RupeeBias, a benchmark for auditing demographic bias in LLM-generated economic guidance across Indian economic settings. RupeeBias consists of 39,150 prompts spanning four use cases: salary estimation, salary increment estimation, counter-offer recommendation, and service pricing recommendation. The benchmark follows a single-attribute counterfactual design, holding the description of the user's qualifications, experience, or service offering fixed while varying one demographic identifier at a time. RupeeBias covers 87 India-specific demographic identifiers across six axes: caste, religion, regional identity, gender, disability, and urban-rural location, with all prompts constructed in both English and Hinglish. We evaluate nine LLMs on RupeeBias and find systematic demographic disparities across all six axes. For otherwise identical prompts that differ only in demographic identifier, LLM-generated economic outputs differ by 20.2% on average. We publicly release RupeeBias to support future research on demographic bias in LLM-generated economic guidance across India-specific demographic and economic contexts.","authors":["Pavithra P M Nair","Bhavik Talaviya","Shourya Bhushan","Rahul Pankajakshan","Seema Guruvadoo","Avinash Agarwal","Gilad Gressel","Krishnashree Achuthan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31245","pdf_url":"https://arxiv.org/pdf/2609.31245","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM偏差审计","经济决策仿真","人口统计代表性"],"reason":"用LLM生成经济建议并审计人口统计偏差，有真实人类数据对照，涉及经济决策场景，…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":13,"question":"在印度经济咨询场景中，大语言模型生成的经济建议是否因用户的人口统计特征（如种姓、宗教、性别等）而产生系统性偏差？","design":"构建 RupeeBias 基准，包含 39,150 个提示，覆盖薪资估算、加薪估算、还价建议和服务定价建议四个用例。采用单属性反事实设计，固定用户资质、经验或服务描述，仅改变一个印度特定的人口统计标识符（共 87 个，涵盖种姓、宗教、地域、性别、残疾和城乡位置六个轴），并同时用英语和印地英语构建提示。评估九个 LLM 在这些提示上的输出差异。","baseline":"无对照","findings":"在仅人口统计标识符不同的相同提示下，LLM 生成的经济输出平均差异达 20.2%，且在所有六个轴上都存在系统性人口统计差异。","reliability":"论文未讨论","relevance":"该研究直接针对 LLM 在经济学场景中的仿真偏差，使用反事实设计审计人口统计特征对经济建议的影响，与研究者关注的 LLM 人类仿真可靠性及偏差问题高度相关，值得阅读原文了解具体偏差模式和测量方法。","inspiration":"借鉴其单属性反事实设计，通过仅改变一个属性来隔离人口统计特征对模型输出的因果影响，并构建大规模、多语言、多场景的提示集｜可迁移到信贷审批歧视、保险定价、工资谈判等经济金融决策场景，审计 LLM 在金融建议中的群体偏差｜以 LLM 作为虚拟被试，处理为在贷款申请或保险报价提示中改变申请人姓名、性别或种族等标识，结果变量为模型给出的利率、保费或授信额度，并与真实信贷或保险数据中的群体差异进行对照，检验模型偏差是否反映或放大现实歧视。"}},{"id":"2609.31013","version":1,"title":"Same Text, Different Numbers: The Divergence of LLM-Based Measures","zh_title":"相同文本，不同数字：基于LLM的测量分歧","abstract":"Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across providers account for only 34% of total score variation. Cross-model disagreement does not predict subsequent analyst or market disagreement, consistent with a substantial model-specific component rather than common ambiguity in the underlying disclosure. Model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels remain sensitive to the models included in the ensemble. LLM-generated variables should therefore be treated as model-contingent measurements and validated across providers.","authors":["Hamid Boustanifar","Sasan Mansouri"],"categories":["cs.AI","cs.CL","q-fin.GN","q-fin.RM"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31013","pdf_url":"https://arxiv.org/pdf/2609.31013","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B3"],"tags":["LLM测量一致性","算法保真度","文本分析"],"reason":"评估LLM文本测量跨模型一致性，涉及测量偏差与统计推断，有真实数据对照，方法可…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":11,"question":"LLM 生成的文本测量在多大程度上不受模型选择影响？","design":"用七家不同提供商的 LLM 对同一批标普 500 公司财报电话会议问答环节文本，按相同提示词对 13 个构念打分，比较跨模型分数的一致性、方差来源及对下游回归推断的影响。","baseline":"以 Loughran-McDonald 等既有词典法文本测量作为对照，并考察 LLM 分歧与分析师预测分歧、市场反应等真实人类判断的关联。","findings":"跨模型平均秩相关仅 0.52，且构念间差异大；方差分解显示转录本共同成分平均只占 34%，模型特定成分显著。模型选择会实质改变下游回归系数的符号、大小和显著性，而 LLM 分歧与后续分析师或市场分歧无关。","reliability":"论文指出 LLM 测量应视为模型依赖的，需跨提供商验证；自报置信度不能识别更可靠的观测，规则化提示词反而可能增大分歧，且大部分分歧无法由可观测特征解释。","relevance":"直接评估 LLM 作为测量工具在金融文本分析中的可靠性，揭示模型选择对实证结论的威胁，对关注仿真偏差和稳健性的研究者有重要参考价值。","inspiration":"借鉴其多模型同文本打分并做方差分解的设计，可系统检验测量工具的选择效应｜可迁移到政策公告文本的情绪或不确定性测量、信贷审批中的软信息编码、分析师报告语调等场景｜用多家 LLM 对同一批政策声明或贷款申请文本按相同提示词打分，以人工编码或市场反应数据为基准，比较跨模型一致性及对后续回归推断的影响。"}},{"id":"2609.31468","version":1,"title":"PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents","zh_title":"PriceBench：LLM预订代理中价格、质量与品牌偏好的诊断基准","abstract":"LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences. We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model, applied to 28 LLMs from 8 providers on 3,600 hotel tasks from 179 real New York City properties. We find that capability is associated with how consistently an LLM chooses, not with what it chooses: more capable LLMs hold stronger, more consistent preferences, while weaker ones either lock onto one position, exploitable by whoever controls listing order, or choose almost indifferently. What those preferences favor varies sharply across providers and even within one family: price sensitivity spans more than an order of magnitude, and the price/quality trade-off moves mean booked nightly price from \\$247 to \\$393 on identical tasks. What an agent buys must therefore be measured per LLM, not inferred, and we release the tasks, code, and all 28 response sets.","authors":["Pavel Kireyev"],"categories":["econ.GN","cs.AI","cs.CL","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31468","pdf_url":"https://arxiv.org/pdf/2609.31468","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","消费者选择","偏好测量"],"reason":"用LLM替代人类消费者进行预订选择，并与真实酒店数据对照，属于经济学场景仿真","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":14,"question":"LLM作为酒店预订代理时，其价格、质量和品牌偏好是什么，这些偏好在不同LLM之间有何差异？","design":"构建PriceBench基准，包含3600个酒店选择任务（1800个二元、1800个三元），基于179家真实纽约酒店，属性包括星级、价格、评分等。对28个LLM进行测试，每个任务以原始顺序和交换顺序各评分一次，使用logit选择模型从选择中恢复偏好参数。","baseline":"无对照（论文未使用真实人类预订数据作为基准，但引用了人类预订者对质量的估值范围进行对比）。","findings":"能力与选择的一致性相关，而非选择内容：更强的LLM偏好更一致，较弱的LLM要么锁定位置（受列表顺序控制），要么几乎无差异。价格敏感度跨LLM差异超过一个数量级，价格/质量权衡导致相同任务下平均每晚预订价格从247美元到393美元不等。","reliability":"论文讨论了位置锁定LLM在温度0下失效，采样无法恢复偏好；偏好幅度受提示格式、解码方式和选项数量影响；品牌偏好与提供商无关。","relevance":"该研究用LLM替代人类消费者进行预订选择，属于经济学场景仿真，但缺乏真实人类对照，主要提供LLM间差异的诊断，对关注仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其通过随机化属性、重复呈现和logit模型识别偏好的方法，可迁移到消费者选择、定价策略等经济问题。｜例如，在信贷审批歧视研究中，用LLM扮演信贷员，改变申请人属性（如种族、性别）观察决策差异。｜设计：以LLM为被试，呈现贷款申请（处理为不同种族/性别信号），结果变量为批准决策，对照真实信贷员历史数据评估偏差。"}},{"id":"2609.30705","version":1,"title":"The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?","zh_title":"思考的代价：测试时推理在LLM交易中是否值得？","abstract":"While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.","authors":["Jiayi Chen","Guiling Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30705","pdf_url":"https://arxiv.org/pdf/2609.30705","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM交易","推理成本","经济决策"],"reason":"用LLM模拟金融决策并与真实市场数据对照，涉及经济场景和失效条件，但非人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":9,"question":"在固定模型、信息、提示、输出格式和投资组合规则的情况下，增加推理努力是否可靠地改善LLM交易组合的净收益？","design":"对DeepSeek、GPT和Gemini三个模型家族的代表性LLM进行受控干预，改变推理努力水平（无、低、高、最大），在2024年241个形成日对100只美国流动性股票进行评分，基于数值特征、可识别新闻和掩蔽新闻三种输入条件，构建等权市场中性投资组合（买入前10、卖空后10），持有5天，扣除10个基点的交易成本，比较不同推理水平下的净收益。","baseline":"无对照","findings":"在三个模型家族中，增加推理并未可靠地改善净投资组合收益；DeepSeek从无推理到最大推理的表现是非单调的。重复生成显示处理效应和投资组合选择不稳定，即使总体得分相似。","reliability":"论文承认推理水平的变化会改变输出结构可靠性，但可靠性提高并不自动意味着收益提高；重复生成导致处理效应不稳定，表明结果可能因随机生成而异；研究仅覆盖一年美国股票数据，且置信区间较宽，不能证明所有效应为零。","relevance":"该研究用LLM模拟金融决策并与真实市场数据对照，评估推理努力对经济结果的影响，属于经济场景下的仿真研究，并揭示了仿真失效条件（推理增加不带来稳定收益），值得阅读原文以了解其受控设计和稳健性检验方法。","inspiration":"借鉴其受控干预设计：在固定其他因素下仅改变推理努力，并使用重复生成和冻结审计评估稳定性｜可迁移到资产定价实验或投资决策仿真，检验LLM推理深度对预测准确性和组合表现的影响｜以LLM为被试，处理为不同推理水平，结果变量为组合净收益或预测误差，用真实历史市场数据作为基准对照。"}},{"id":"2609.31095","version":1,"title":"Confident, Not Wiser: The Dunning-Kruger Effect in Human-AI Interaction","zh_title":"自信而非更明智：人机交互中的达克效应","abstract":"AI assistance can improve performance without improving self-assessment. We report a study (N=366) comparing Human alone and Human+AI performance on reasoning tasks, for which the AI model is benchmarked on the same items. Participants estimated global and block performance and rated confidence in their answers. Human+AI achieved higher scores, but self-estimates tracked performance weakly. Average overestimation was similar across groups, covering individual errors. Across tasks, confidence distinguished correct from incorrect answers less accurately in the Human+AI group, while within-task differences remained uncertain. The Dunning-Kruger pattern was found in both groups, with a larger observed contrast in Human+AI. Controls for score noise reduced but did not eliminate the pattern, with the controlled group difference remaining inconclusive. An extended computational account describes global and block estimates. Our findings distinguish performance augmentation from metacognitive augmentation and motivate interfaces that support verification, communicate task-specific AI model performance, and help users evaluate the quality of their joint work rather than produce answers.","authors":["Daniela Fernandes","Michelle Rausch","Agnes Mercedes Kloft","Daniel Buschek","Robin Welsch"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31095","pdf_url":"https://arxiv.org/pdf/2609.31095","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["人机交互","元认知","AI辅助决策"],"reason":"研究人类与AI协作中的元认知偏差，有真实人类数据对照，结论可迁移到LLM仿真可…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":12,"question":"人类与AI协作时，AI辅助是否在提升任务表现的同时改善元认知（自我评估准确性、信心区分度），以及邓宁-克鲁格效应在有无AI辅助下如何表现。","design":"本研究并非LLM仿真研究，而是人类实验：366名被试分为Human alone和Human+AI两组，完成40道推理题（矩阵、心理旋转、三段论、字母串），AI为GPT-5.6 Luna，并单独对AI模型进行基准测试。被试在答题前后进行全局和分块表现估计，并对每题答案给出信心评分。","baseline":"Human alone组作为人类基准，同时AI模型单独在相同题目上的表现作为AI基准。","findings":"Human+AI组得分更高，但自我估计与实际表现的相关性弱，平均高估程度与Human alone组相似；信心区分正确与错误答案的能力在Human+AI组更差。邓宁-克鲁格模式在两组均存在，Human+AI组观察到的对比更大，但控制分数噪声后组间差异不明确。","reliability":"论文未明确讨论失效条件，但指出研究为探索性观察研究，未预注册；控制分数噪声后邓宁-克鲁格效应的组间差异不明确，且任务内信心区分度差异不确定。","relevance":"该研究直接涉及人类与AI协作中的元认知偏差，有真实人类数据对照，结论可迁移到LLM仿真中关于自我评估和信心校准的建模，值得阅读原文以了解具体测量和稳健性检验方法。","inspiration":"借鉴其同时测量全局估计、分块估计和逐题信心，并设置Human alone和Human+AI对照以及AI单独基准，以分离表现提升与元认知变化。｜可迁移到金融决策场景，如投资者使用AI辅助进行资产配置或信贷审批中，评估AI辅助是否改善决策质量但未改善决策者的信心校准。｜设计实验：招募投资者作为被试，随机分为仅人类组和人类+AI组，完成一系列投资决策任务，AI为LLM提供建议，结果变量为投资组合收益和投资者对自身表现的估计及信心评分，对照真实市场数据或历史基准。"}},{"id":"2609.27690","version":2,"title":"Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research","zh_title":"合成研究验证中的后果行为与表征公平性","abstract":"Researchers in industry and academia use synthetic survey respondents powered by large language models as substitutes for human samples. These synthetic populations require validation against real-world data, so researchers often address them using ad hoc comparisons with human surveys. Inspired by the intention-behaviour gap in behavioural science, we argue that these validations test the wrong thing for most applied cases where decision makers commission synthetic research to anticipate consequential behaviour. To address this problem, we propose a validation framework with two requirements. First, every validity claim must state its level of correspondence with human data: does the sample predict what the represented people do, which of four diagnostics (location, dispersion, response process and structure) does the validation address, and does the validation compare against experimental effects? Second, researchers must report validity claims for subgroups, since these groups are often the most affected by consequential decisions and aggregate accuracy hides their misrepresentation. Our validation framework operationalises three justice dimensions (distributional, procedural, and recognition) as measurable quantities and defines within-persona counterfactual experiments as a validation requirement. We then apply the framework to electric vehicle charging tariffs, before closing with a reporting checklist that researchers can use to make convincing validity claims.","authors":["Florian Kutzner","Celina Kacperski","Laura de Moli\\`ere","Edoardo Chidichimo","Min Jun Jung","Felix P. S. Wallis","James K. He"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-24","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.27690","pdf_url":"https://arxiv.org/pdf/2609.27690","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B2","B3","B4"],"tags":["LLM仿真","验证框架","公平性"],"reason":"直接研究LLM合成调查受访者作为人类替代，提出验证框架并应用于电动汽车充电定价…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":1,"question":"如何验证基于大语言模型的合成调查受访者能否预测真实人群的后果性行为，并确保子群体代表性公平？","design":"本文提出一个验证框架，而非进行仿真实验。框架要求：明确效度声明与人类数据的对应层级（预测行为、四种诊断：位置、离散度、响应过程、结构、与实验效应比较）；报告子群体效度；将分配、程序、承认三个正义维度操作化为可测量量；定义“人设内反事实实验”作为验证要求。并以电动汽车充电电价为例应用该框架。","baseline":"无对照（本文为框架性论文，未提供具体人类数据对照）","findings":"现有合成受访者验证多聚焦于态度或意见的边际分布一致性，忽视了意图-行为差距，无法证明其能预测后果性行为。提出的框架要求效度声明必须明确对应层级、子群体表现，并通过人设内反事实实验检验因果推断能力。","reliability":"论文承认以下局限：训练数据污染难以评估；子群体分析受限于人类基准数据可得性；行为标准本身存在缺陷（如公开行为、行政记录、实验室任务各有问题）；模型版本更新导致效度证据时效短；缺乏全面的实用测试集来确定合成人群满足哪些效度要求。","relevance":"该论文直接针对LLM合成受访者作为人类替代的验证问题，提出批判性框架并强调子群体公平，与研究者关注的人类仿真可靠性、偏差及经济学政策评估场景高度相关，值得精读原文。","inspiration":"借鉴其“人设内反事实实验”设计，可对同一合成个体施加不同处理以估计个体处理效应，并对照真实人类实验效应进行验证。｜可迁移到消费者金融决策研究，如信贷产品选择、保险购买或退休储蓄计划参与等场景。｜以合成受访者作为被试，处理为不同信贷条款（如利率、还款期限），结果变量为选择行为，对照真实世界信贷申请数据或实验室实验数据，检验合成样本的预测效度与子群体公平性。"}},{"id":"2609.29928","version":1,"title":"Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations","zh_title":"文化差异保持：诊断LLM模拟调查人群中的扁平化与夸张化","abstract":"Large language models (LLMs) are increasingly used as synthetic survey respondents to estimate population response distributions. In cross-cultural survey simulation, evaluations should assess not only distributional fidelity within countries but also whether differences across countries are preserved. However, existing distance-based metrics such as Jensen--Shannon divergence (JSD) do not directly capture such cross-country differences. To address this limitation, we introduce Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration. CDP identifies reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature. To evaluate CDP, we conduct experiments across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test. The results reveal a systematic discrepancy between conventional fidelity metrics and CDP. Controlled experiments show that CDP changes monotonically as cross-country divergence is attenuated or amplified, while the corresponding changes in JSD remain relatively small. In our audit of real LLM generations, DeepPersona-Inspired prompting is frequently favored by conventional fidelity metrics but exhibits the strongest flattening in every model--domain block. CDP thus complements fidelity metrics by directly quantifying the attenuation or amplification of cross-country divergence.","authors":["Yeeun Chae","Yewon Choi","Seunghyun Lee","IL Im"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29928","pdf_url":"https://arxiv.org/pdf/2609.29928","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM仿真","跨文化调查","算法保真度"],"reason":"直接研究LLM仿真调查人群，评估跨文化差异保真度，并与真实人类数据对照，提出诊…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":2,"question":"如何诊断LLM模拟调查人群时对跨国文化差异的扁平化或夸张化？","design":"用四个开源LLM（Gemma-3-4B、Qwen3.5-9B、Qwen3.5-27B、Llama-2-13B）模拟六个国家（阿根廷、澳大利亚、德国、印度、肯尼亚、美国）的受访者，采用三种人设提示方法（Cultural Prompting、PersonaHub-Inspired、DeepPersona-Inspired），生成世界价值观调查（WVS）和大五人格测试的回答，并测量国家间分布差异的保持程度。","baseline":"真实人类数据：WVS第七波六个国家的全国回答分布，以及OpenPsychometrics的大五人格测试数据（阿根廷、澳大利亚、印度）。","findings":"传统分布保真度指标（如JSD）与CDP存在系统性偏差：DeepPersona-Inspired提示在多数模型-领域组合中分布保真度最高，但文化扁平化最严重。CDP在受控实验中随跨国差异的衰减或放大单调变化，而JSD变化很小。","reliability":"论文未明确讨论失效条件，但指出CDP需要一次性人类校准，且仅适用于有跨国人类参照数据的调查领域。","relevance":"该研究直接针对LLM仿真调查中的文化差异保真度问题，提出了新的诊断指标，并用真实人类数据对照，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其通过受控扰动构造扁平化/夸张化数据集来检验指标敏感性的方法，以及将分布保真度与差异保持度分开评估的思路。｜可迁移到跨国经济偏好调查或政策态度仿真中，例如用LLM模拟不同国家消费者对通胀预期的回答，检验其是否保持国家间差异。｜设计：用LLM模拟多国受访者回答通胀预期调查，施加不同人设提示，测量国家间预期分布的差异保持度，并与密歇根大学或欧洲央行的真实调查数据对照。"}},{"id":"2609.30030","version":1,"title":"Artificial Societies Benchmark: A Validation Framework for Synthetic Research","zh_title":"人工社会基准：合成研究的验证框架","abstract":"A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to help researchers assess whether synthetic populations support their intended analyses. The framework combines eleven tests across internal, construct, and external validity, drawing on twenty human sources and comparing nine language models. It connects each research use to the evidence it requires and tests how results change with the information we supply about respondents. Importantly, strong performance in one domain does not establish fidelity in the others. Models often answer too consistently, compress response scales, and alter relationships between traits whilst richer profiles improve prediction for some models and worsen it for others. The resulting scorecard helps researchers identify which aspects of a synthetic population can support their analysis and where researchers need further human evidence.","authors":["Edoardo Chidichimo","Min Jun Jung","Felix P. S. Wallis","James K. He"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30030","pdf_url":"https://arxiv.org/pdf/2609.30030","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","效度验证","合成人群"],"reason":"直接评估LLM合成人群的效度，含人类数据对照和批判性分析","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":3,"question":"如何系统评估大语言模型生成的合成人群在多大程度上能支持研究者预期的分析（如调查回答、心理测量、实验效应）？","design":"用九个大型语言模型（含专有和开源）根据二十个人类数据源（调查、面板、人格量表、实验）中的受访者信息（人口统计、先前回答、个人描述等）生成合成回答，并施加实验处理；通过十一个测试从内部效度、构念效度和外部效度三个维度评估合成人群的响应过程、心理测量结构和总体/实验保真度，同时比较不同信息丰富度和统计控制（独立边际、高斯秩相关）的影响。","baseline":"二十个人类数据源，包括全国调查、追踪面板、人格量表和实验，提供真实回答、人口统计、先前回答、个人描述、实验分配和独立测量结果作为对照基准。","findings":"模型在某一效度领域的良好表现不能推广到其他领域；模型往往回答过于一致、压缩量表范围、改变特质间关系，且更丰富的个人信息对某些模型改善预测而对另一些模型则恶化预测。","reliability":"论文承认强表现不跨域通用，并指出模型回答过于一致、压缩量表、改变特质关系等失效条件；但未在节选中详细讨论其他局限。","relevance":"该研究直接针对LLM合成人群的效度验证，提供人类数据对照和批判性分析，与研究者关注的经济学实验和政策评估场景高度相关，值得精读原文以了解具体测试方法和失效模式。","inspiration":"借鉴其多维度效度测试框架和统计控制设计，可迁移到经济政策评估中的异质性处理效应或消费者选择实验，例如用LLM模拟不同收入群体对税收优惠的反应，以真实调查数据（如美国消费者财务调查）为基准，比较模型生成的边际消费倾向与人类数据的分布和协变量关系。"}},{"id":"2609.29952","version":1,"title":"Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes","zh_title":"Augur：用于预演产品和政策变化反应的合成决策实验室","abstract":"Before a product or policy change ships, the question that matters is how people will react to it. Augur rehearses that reaction offline: it builds a typed knowledge graph from the change documents, populates a grounded persona market, simulates the interaction, and returns an auditable decision memo recommending one of five actions. We assemble Gold-50, fifty real product and policy episodes whose real-world outcome is known, adjudicated against the public record, and score the five-way release verdict against it. Our central finding is methodological and negative: most of the measured gap between frontier cloud models and open-weight models we fine-tune and serve offline is attributable to an under-specified evaluation, not a difference in capability. We show this three ways. First, the prompt envelope alone can dominate the score: holding weights, cases and scorer fixed, one system -- a LoRA-SFT adapter on Qwen3-32B -- swings from 0% to 73%. Second, in a matched 2x2 ablation, defining the decision taxonomy in the prompt -- with no model change -- lifts every frontier model by +24 to +34pp; under the under-specified prompt, Qwen3-32B LoRA-SFT served offline beats all three frontier models (paired McNemar, Holm-corrected), and once the prompt is fair no significant difference from any of them is detected. Third, agreement with the distillation teacher rises without accuracy following, and the full pipeline amplifies a systematic \"over-doom\" bias rather than improving the verdict. Separately, we validate the reaction layer on its own terms: blind judges across four model families find the synthetic reaction recovers 67-90% of the concerns the public actually raised, and a pre-registered ablation locates its value -- largest where the decision is hardest, redundant near ceiling. The pipeline that regenerates every number and figure here is available from the authors.","authors":["Rahul Khedar","Mayank Malhotra","Avinash Karn"],"categories":["cs.AI","cs.CL","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29952","pdf_url":"https://arxiv.org/pdf/2609.29952","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM仿真","人类行为预测","政策评估"],"reason":"用LLM模拟人类对产品和政策变化的反应，并与真实结果对照，直接属于人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":6,"question":"在真实产品与政策变更决策中，前沿云端模型与开源微调模型之间的性能差距有多少是真实能力差异，多少是评估方法（提示词）造成的？","design":"构建 Augur 系统，将决策分解为文档→知识图谱→人物角色→模拟→报告五个阶段，用 LLM 生成利益相关者角色并模拟其互动，最终输出五选一的发布决策建议。使用 50 个真实产品/政策案例（Gold-50）作为基准，对比不同模型（前沿云端模型与开源微调模型）在有无决策分类定义提示词下的准确率。","baseline":"Gold-50 基准：50 个真实产品与政策变更案例，其真实世界结果已通过公开记录人工核实，作为五分类发布决策的对照标准。","findings":"主要发现是方法性的负面结果：前沿模型与开源模型之间的性能差距主要源于评估提示词的不充分定义，而非模型能力差异。在提示词中定义决策分类后，开源模型与前沿模型无显著差异；此外，完整流程会放大系统性“过度悲观”偏差，而反应层本身能恢复 67-90% 的公众真实关切。","reliability":"论文承认两个失效模式：与蒸馏教师的一致性上升但准确率不升，以及完整流程放大过度悲观偏差。原因在于蒸馏破坏判断独立性，精确匹配评分器强加模式合规上限。","relevance":"该研究直接使用 LLM 模拟人类对产品和政策变化的反应，并与真实结果对照，属于人类仿真实验，且包含批判性分析，值得阅读原文以了解仿真在何种条件下失效。","inspiration":"借鉴其提示词消融和匹配对照设计，可揭示评估方法对模型性能结论的影响，并采用预注册的消融实验定位仿真组件的价值。｜可迁移到政策公告的预期形成研究，如央行利率决议或财政刺激方案的市场反应模拟。｜以 LLM 生成的经济主体（如消费者、投资者）为被试，处理为不同的政策公告文本（含或不含决策分类定义），结果变量为预测的市场反应（如消费、投资决策），对照真实市场数据（如消费者信心指数、资产价格变动）来验证仿真准确性。"}},{"id":"2609.29692","version":1,"title":"Fair Like Us? Auditing LLM Alignment in Resource Allocation","zh_title":"像我们一样公平？审计资源分配中LLM的对齐","abstract":"Fair allocation of scarce, indivisible resources is an important challenge in many societal problems. While there are several formal theories of fairness, no single definition can always be satisfied. As large language models (LLMs) are increasingly used to support decisions and act as agents, they raise new concerns about distributional justice: their judgments are not directly tied to any specific fairness framework and may violate key normative principles. In this work, we introduce a general method for evaluating fairness reasoning in LLMs. We study first-person fairness judgments across a broad set of models and compare them directly with human responses on matched scenarios and elicitation conditions. We find that LLMs tend to prefer stricter fairness constraints than humans, show more self-interested behavior, are sensitive to how information is framed, and are difficult to align with human judgments using fine-tuning with current datasets.","authors":["Qishen Han","Hadi Hosseini","Joshua Kavner","Samarth Khanna","Sujoy Sikdar","Lirong Xia"],"categories":["cs.AI","cs.CY","cs.GT"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29692","pdf_url":"https://arxiv.org/pdf/2609.29692","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","公平分配","人类对照"],"reason":"用LLM模拟人类资源分配判断，并与人类数据对照，评估偏差与对齐难度。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":5,"question":"LLM在资源分配公平性判断上与人类有多大程度的一致，哪些形式化公平准则能解释其判断？","design":"将多个LLM（包括不同规模和推理能力的模型）置于与人类被试相同的公平分配场景中，作为第一人称代理人评估分配结果是否公平或可接受；实验操纵分配满足的公平性质（EF、PROP、EF1、MMS、PROP1等）和诱导条件（框架、信息结构、响应格式），测量模型判定分配可接受的比率。","baseline":"使用Hosseini et al. (2025a)的150名人类被试在相同场景、偏好结构和实验处理下的公平性判断数据作为对照。","findings":"LLM比人类更倾向于严格的公平约束，对EF分配的公平判断率显著高于人类，而对EF1、MMS等放松准则的判断率下降更快；LLM表现出更强的自利行为，对信息框架敏感，且通过微调难以与人类判断分布对齐。","reliability":"论文指出当前人类公平分配数据集规模太小，微调只能使模型坍缩到模态响应而非真正对齐人类判断分布；LLM的判断受诱导问题措辞影响大于分配的形式化属性，且推理能力更强的模型不一定更对齐。","relevance":"该研究直接以LLM作为人类被试的替代品，在资源分配场景中与真实人类数据严格对照，系统评估了仿真偏差和失效条件，对关注LLM仿真可靠性的研究者具有重要参考价值。","inspiration":"借鉴其“同一场景、同一处理、仅替换被试”的严格对照设计，以及通过操纵分配满足的公平性质和诱导框架来分离判断依据的方法｜可迁移到信贷审批中的公平性判断、公共资源分配政策评估、或消费者对价格歧视的公平感知等经济金融场景｜以LLM模拟贷款申请人或政策受众，处理为不同公平准则（如无歧视、比例公平）和框架（如强调个人得失 vs 社会效率），结果变量为接受度或公平评分，对照真实人类实验数据（如调查或实验室实验）来检验LLM的仿真效度。"}},{"id":"2609.29143","version":1,"title":"AI-Moderated Interviews for Market Research and Digital Twins Calibration","zh_title":"用于市场研究和数字孪生校准的AI主持访谈","abstract":"AI-moderated interviews are emerging as a scalable market-research method for generating consumer insights and building consumer \"digital twins.\" Yet it remains unclear whether they match human-moderated interviews or improve on simpler, static data collection methods. In a pre-registered, between-subjects study (N = 317) with three industry partners, we compare AI-moderated (N = 139), human-moderated (N = 24), and static interviews (N = 154). AI moderation matches human moderation in depth, covers more themes, and, holding budget constant, recovers significantly more customer needs than human moderation or static interviews. However, participants sound more emotionally engaged when speaking to a live human. We then create digital twins using interview data and evaluate each twin against the participant's own held-out responses to six real-world marketing stimuli. We find that digital twins created from AI-moderated interviews predict consumer responses better than demographics-only personas. However, the additional richness from AI moderation does not translate into better quantitative predictions compared to static interviews. By analyzing open-ended thoughts generated from humans versus their twins, we find that prediction errors are connected both to differences in (self-reported) thinking styles between twins and humans, and to gaps between training and validation data (i.e., asking questions that are too far out of distribution).","authors":["Yuting Deng","Jingxuan Liu","Olivier Toubia","Naman Jain"],"categories":["cs.CY","cs.AI","cs.HC","cs.MA"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29143","pdf_url":"https://arxiv.org/pdf/2609.29143","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","数字孪生","市场研究"],"reason":"用LLM进行AI主持访谈并构建数字孪生，与真实人类访谈对照，评估预测效度与偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":4,"question":"AI主持访谈能否在同等预算下匹敌人类主持访谈的深度与需求挖掘，并用于构建更准确的消费者数字孪生？","design":"预注册的组间实验，317名消费者随机分配到AI主持访谈、人类主持访谈或静态访谈三种条件，比较访谈的深度、主题覆盖和客户需求数量；随后用访谈数据构建数字孪生，预测参与者对六个真实营销刺激的保留回答。","baseline":"人类主持访谈（N=24）和静态访谈（N=154）作为对照，数字孪生预测与人口统计特征基准和参与者自身保留回答比较。","findings":"AI主持在深度上与人类主持相当，覆盖更多主题，且在预算固定下比人类主持和静态访谈挖掘出更多客户需求；但参与者在与真人交谈时情感参与度更高。基于AI访谈构建的数字孪生预测优于仅人口统计特征的基准，但相比静态访谈并未提升定量预测准确性。","reliability":"论文指出预测误差与数字孪生和人类在自我报告思维方式上的差异有关，也与训练和验证数据之间的分布差距有关（即提问过于超出分布范围）。","relevance":"该研究直接评估了LLM作为人类被试替代品在定性访谈和数字孪生构建中的效度，包含真实人类对照和批判性发现，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其预注册组间实验设计，将AI访谈与人类访谈和静态问卷对比，并用保留样本验证数字孪生的预测效度｜可迁移到消费者金融决策研究，如信贷产品偏好或保险选择，用AI访谈构建个体数字孪生预测金融行为｜以真实消费者为被试，随机分配AI访谈、人类访谈或静态问卷，用访谈数据构建数字孪生预测其对金融产品广告的反应，并与实际选择数据对照。"}},{"id":"2609.29370","version":1,"title":"From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring","zh_title":"从政策文件到结构化调查回答：评估大语言模型用于政策监测","abstract":"Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are costly and challenging to scale across countries. Large language models (LLMs) enable new possibilities for extracting and structuring information from long and unstructured policy documents. This paper presents an application of LLMs as \"AI respondents\" for generating structured survey responses from policy texts. We develop a data extraction pipeline based on long-context in-context learning to map information from public web sources into predefined survey categories, including policy instruments, target groups, and thematic areas. The pipeline integrates a validation step using a secondary LLM to assess relevance and evidence, alongside comparisons with human-provided responses. Using a multi-country dataset, we evaluate the alignment between LLM-generated and human-generated outputs through overlap measures and cross-validation. Results show that LLMs achieve high agreement for structured indicators (84-95%), while differences remain in free-text fields, where models tend to provide more detailed procedural descriptions. These findings highlight the potential of hybrid human-AI workflows for policy monitoring, improving both efficiency and scalability while maintaining the need for human validation and contextual interpretation.","authors":["Carolyn Cole","Matthias Deschryvere","Toqeer Ehsan","Arash Hajikhani"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29370","pdf_url":"https://arxiv.org/pdf/2609.29370","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","政策监测","人类对照"],"reason":"用LLM从政策文本生成结构化调查回答，并与人类回答对照，属于仿真人类被试且有人…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":8,"question":"能否用大语言模型从政策文本中自动生成结构化调查回答，以替代或辅助人工政策监测？","design":"使用长上下文上下文学习（long-context in-context learning）的LLM（如GPT-4o）作为“AI受访者”，从网页抓取的政策文本中提取信息，映射到预定义的调查类别（政策工具、目标群体、主题领域），并生成自由文本字段（描述和目标）。通过一个二级LLM验证层评估相关性和证据，并与人类提供的回答进行比较。","baseline":"来自EC-OECD STIP Compass调查的人类专家回答，覆盖六个OECD国家（加拿大、芬兰、德国、韩国、西班牙、土耳其）的政策举措。","findings":"LLM在结构化指标（政策工具、目标群体、主题代码）上与人类回答的一致性达到84-95%，但在自由文本字段上存在差异，模型倾向于提供更详细的程序性描述，而人类更强调背景和社会影响。","reliability":"论文承认LLM在自由文本字段上与人类存在差异，可能引入系统性偏差，且政策数据具有异质性和制度嵌入性，公开来源可能无法完全捕捉；需要人类验证和上下文解释。","relevance":"该研究直接使用LLM作为人类受访者的替代品，并与真实人类数据对照，评估仿真可靠性，符合研究者对LLM仿真实验和批判性评估的兴趣，值得阅读原文以了解具体方法和偏差分析。","inspiration":"方法上，该研究展示了如何利用长上下文提示和二级LLM验证来从非结构化文本中提取结构化数据，并设计人类对照来评估一致性。｜可迁移到经济金融领域，如从公司年报、政策文件或新闻中自动提取结构化信息，用于构建经济指标或评估政策影响。｜研究设计雏形：使用LLM从上市公司年报中提取财务和非财务信息（如研发支出、风险因素），与人工标注或数据库中的真实数据进行对照，评估LLM提取的准确性和偏差，并分析在哪些条件下LLM表现不佳。"}},{"id":"2609.28486","version":1,"title":"Political Sorting Can Drive AI Models Apart Through User Feedback","zh_title":"政治分类可通过用户反馈使AI模型分化","abstract":"Large language models are rapidly becoming an important source of political information. This raises a fundamental question: will AI systems support a shared basis for political knowledge, or lead different political groups to rely on increasingly different models? Political sorting can drive model fragmentation if three conditions hold: politically different users select into different models, learning from user feedback pushes those models apart politically, and the resulting differences shape subsequent model choices. We call this self-reinforcing process the Centrifugal Alignment Spiral. We study its components in three steps. First, we draw on a human experiment showing that political identity predicts model choice. Second, we fine-tune language models on synthetic feedback reflecting predominantly Democratic or Republican preferences. Across five independent runs per model family, paired models diverged on 12-41% of unseen survey questions with large partisan gaps, and in every run the differences moved in the expected political direction; for some models, differentiation extended even to issue areas excluded from training. Pooling feedback across groups instead suppressed divergence. Third, an empirically anchored agent-based model shows what follows when political sorting and model adaptation operate together: models attract politically distinct audiences, learn from them, and diverge further. User feedback can therefore turn political sorting among AI users into durable differences between the models on which they rely for political information.","authors":["Petter T\\\"ornberg","Michael Heseltine","Nicol\\`o Pagan","Christopher Bail","Michelle Schimmel","Christopher Barrie"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28486","pdf_url":"https://arxiv.org/pdf/2609.28486","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","政治极化","人类对照"],"reason":"用LLM模拟政治反馈并对照人类实验，研究模型分化，涉及政策场景与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":7,"question":"政治排序能否通过用户反馈导致不同政治群体使用的AI模型在政治上分化，形成自我强化的离心对齐螺旋？","design":"研究分三步：首先利用已有的人类实验证明政治身份预测模型选择；其次用反映民主党或共和党偏好的合成反馈微调Qwen2.5-1.5B、Mistral-7B和GPT-OSS-20B模型，每个模型家族进行五次独立运行，测量配对模型在未见过的调查问题上答案分歧的比例和方向；最后构建基于实证的智能体模型，模拟政治排序和模型适应耦合下的动态演化。","baseline":"人类基准来自先前报告的人类模型选择实验，显示政治身份预测模型选择，包括付费准确回答条件下共和党人更可能选Grok、民主党人更可能选Claude，71%的参与者返回之前偏好的模型。","findings":"在90/10的受众对比压力测试下，配对模型在12-41%的未见调查问题上产生分歧，且每次运行中平均差异都符合反馈的政治方向；部分模型的分化扩展到未训练的政治领域。混合不同群体的反馈则抑制了分化。","reliability":"论文承认90/10的受众对比是压力测试，不代表当前AI市场的实际排序程度；分化程度因模型、领域和提示而异；身份线索实验表明谄媚个性化可能减少模型间分化，但增加模型内分化。","relevance":"该研究直接探讨LLM作为政治信息源时的分化机制，通过合成反馈模拟政治群体偏好，并与人类实验对照，涉及政策评估场景和失效条件（如反馈混合、个性化），对关注LLM仿真可靠性及偏差的研究者具有重要参考价值。","inspiration":"借鉴其用合成反馈微调模型并测量泛化分化的方法，可迁移到经济金融中的群体偏好分化问题，如不同收入或风险偏好群体的金融建议模型分化。｜可应用于信贷审批或投资建议场景，研究用户反馈如何导致模型对不同群体产生差异化行为。｜设计：用不同风险偏好或金融素养的合成用户反馈微调金融LLM，测量其在未见金融决策问题上的行为差异，并与真实人类金融决策数据（如调查或实验数据）对照，检验分化是否与真实群体差异一致。"}},{"id":"2609.22904","version":2,"title":"LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage","zh_title":"LLM在顺序临床分诊中锚定主诉且未能整合证据","abstract":"Triage in the emergency department (ED) is a sequential decision process that unfolds turn by turn. Existing evaluations of large language models (LLMs) for triage use completed retrospective records and report performance close to that of physicians. We implement a methodology for evaluating LLMs on sequential triage, the task of predicting a triage acuity label from a growing prefix of a nurse-patient conversation. We evaluate six LLMs at five sequential checkpoints on two corpora: 425 LLM-generated (SIMULATED) and 50 physician-authored (CLINICIAN) conversations, both labelled under the Emergency Severity Index (ESI). Every model, measured by quadratic weighted kappa (QWK), degrades from moderate-to-substantial agreement on completed records to fair-to-moderate agreement at every sequential checkpoint. Controlled perturbations show that the label at every checkpoint is anchored on the chief complaint exchanges, and prompting interventions fail to lift this plateau. Models extract clinically relevant content from later turns, yet the surprisal of the true label rises across the checkpoints. So the model fails to integrate the evidence. Three expert clinicians on the same conversations reach a QWK of 0.887-0.929, while the best model reaches 0.295. Predictions concentrate at ESI-2 and ESI-3, and models agree with each other more than with the ground truth, so ensembling worsens the failure. Deploying LLMs for ED triage based on offline benchmarks alone misses this sequential failure.","authors":["Dipankar Srirag","Haokai Zhao","Ashutosh Kumar","Eleanor Hopper","Michael Dalton","Quoc Dung Nguyen","Aditya Joshi","Salil S. Kanhere","Padmanesan Narasimhan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-22","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.22904","pdf_url":"https://arxiv.org/pdf/2609.22904","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","临床决策","可靠性评估"],"reason":"用LLM模拟临床分诊决策并与医生数据对照，揭示顺序决策中的锚定与证据整合失败，…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":9,"question":"LLM在顺序临床分诊中如何表现，其决策锚定在对话何处，为何更多证据未能提升分诊准确性？","design":"用6个LLM（4个开源、2个闭源）在425个模拟和50个临床医生撰写的护患对话上，按五个顺序检查点（从主诉开始到对话结束）预测ESI急迫度标签，并与离线完整EHR记录预测对比；通过扰动对话（重排、删除）和提示干预检验锚定位置，并测量标签的意外度。","baseline":"三位急诊专家对50个模拟对话进行分诊，QWK为0.887–0.929；离线完整记录上模型QWK为中等至显著一致。","findings":"所有模型在顺序检查点上QWK降至公平至中等一致，最佳模型仅0.295，远低于专家；模型决策锚定在主诉交换，后续证据未被整合，预测集中在ESI-2和ESI-3，且模型间一致性高于与真实标签的一致性。","reliability":"论文指出离线基准掩盖了顺序失败，且模型在模拟和临床医生对话上均表现不佳；提示干预未能提升性能，集成反而加剧失败；未讨论其他失效条件如对话长度、噪声或不同患者群体。","relevance":"该研究用LLM模拟临床分诊决策并与人类专家对照，揭示了顺序决策中的锚定和证据整合失败，对关注LLM仿真可靠性及偏差的研究者具有直接参考价值，值得精读原文。","inspiration":"借鉴其顺序检查点设计和扰动分析来定位决策锚点，可迁移到经济金融中的顺序信息处理场景，如信贷审批中逐步披露申请人信息或政策公告的预期形成；设计可用LLM扮演信贷员或投资者，在信息逐步呈现的多个时点做出决策，测量决策变化和锚定效应，并与真实信贷员或市场数据对照，检验LLM是否同样忽略后续信息。"}},{"id":"2609.28673","version":1,"title":"Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks","zh_title":"基准测试大语言模型的论辩行为：对人身攻击防御策略的研究","abstract":"Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on character attacks (ad hominem arguments), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content. Specifically, we investigate whether modern LLMs can replicate human competence to strategically use and respond to such attacks. We analyse a corpus of natural language political dialogues to identify defensive strategies human interlocutors naturally employ in ethos-centred debates and structure them into a dialogue game. Empirically, we benchmark LLM-generated dialogues against the ElecDeb60to16-fallacy corpus of U.S. presidential debates, contrasting human debaters' repertoire of defensive strategies with those of artificial agents. Results reveal a substantial difference: most LLMs rigidly prioritise logical defences, failing to exploit ethotic counterattacks as valid moves in political discourse. We argue that current safety fine-tuning constraints the strategic action space of these LLMs, making them unable to fully engage in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy.","authors":["Ewelina Gajewska","Katarzyna Budzynska","Jaroslaw Chudziak"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28673","pdf_url":"https://arxiv.org/pdf/2609.28673","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","论辩行为","政治辩论"],"reason":"用LLM模拟人类辩论行为并与真实辩论语料对照，评估其策略差异，属于人类仿真且含…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":10,"question":"LLM在面对人身攻击（ad hominem）时产生的防御策略与人类辩论者相比有何差异？","design":"研究将LLM作为辩论代理，在政治辩论场景中面对人身攻击，生成防御性回应。通过分析人类辩论语料提取防御策略并构建对话游戏，然后让五个LLM在该游戏框架下生成回应，并与人类策略进行对比。","baseline":"ElecDeb60to16-fallacy语料库，包含美国1960-2016年总统辩论中的人身攻击及人类回应。","findings":"大多数LLM僵化地优先采用逻辑防御，未能利用人格反击（ethotic counterattacks）作为政治话语中的有效策略。安全微调限制了LLM的战略行动空间，使其无法在角色竞争是规范性期望的领域中充分参与自然交互。","reliability":"论文指出当前安全微调约束了LLM的战略行动空间，导致其无法在政治辩论等角色竞争为规范的领域中自然交互。","relevance":"该研究直接评估LLM在政治辩论中模拟人类策略的能力，并与真实人类语料对照，属于人类仿真研究，且涉及策略差异和偏差，值得阅读原文以了解具体实验设计和评估方法。","inspiration":"借鉴其构建对话游戏并提取人类策略作为基准的方法，可用于评估LLM在经济决策中的策略行为。｜可迁移到政策辩论或谈判场景，如央行沟通中的预期管理或贸易谈判中的策略互动。｜设计：以LLM作为谈判代理，施加人身攻击处理，测量其回应策略（如逻辑反驳、人格反击、转移话题），并与真实谈判语料（如WTO谈判记录）对照。"}},{"id":"2609.30137","version":1,"title":"Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale","zh_title":"先模拟后上线：1.4亿规模生产客户体验AI智能体的仿真","abstract":"Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank's Card Delivery agent and its expanded successor, Card Management - Nubank's highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.","authors":["Edesio Alcoba","Kevin Rossell","Aman Gupta","Shao Tang","Jiwoo Hong","Pabel Carrillo-Mendoza","Wanderson Concei\\c{c}\\~ao Ferreira","Alvaro Tedeschi","Zayd Simjee","Shreya Rajpal","Bruno Finardi Hime","Christian Sousa","Luis Moneda","Herbert Fei","Daniel Silva","Rohan Ramanath"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30137","pdf_url":"https://arxiv.org/pdf/2609.30137","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","客户服务","A/B测试"],"reason":"用LLM模拟客户与客服agent交互，有生产数据对照，属经济场景仿真","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":15,"question":"如何用LLM驱动的客户仿真在部署前筛选客服智能体，并验证其与真实生产表现的一致性及业务影响。","design":"使用Snowglobe仿真器，由LLM生成合成客户角色（persona），与待测客服智能体进行多轮对话；工具调用被拦截并返回合成结果，不调用生产后端。通过假设驱动的用例分配，对候选智能体进行离线评估，测量版本级二元评估分数、tNPS、SSR等指标。","baseline":"对照Nubank真实生产环境中的客服对话数据，包括4个已部署版本的评估分数，以及后续线上A/B测试的tNPS和SSR。","findings":"仿真评估分数与生产版本级评估分数高度相关；仿真引导的迭代使tNPS提升36.69点，模型替换使SSR提升8.82个百分点且tNPS无显著变化。","reliability":"论文承认仿真存在可控性有限和分数膨胀风险，可能造成模拟与真实场景的差异；但未详细讨论失效条件。","relevance":"该研究提供了LLM仿真在真实商业场景中与人类数据对照的实证证据，对关注仿真可靠性和经济场景应用的研究者有参考价值。","inspiration":"借鉴其工具边界仿真和版本级对照设计，在离线仿真中系统比较不同策略｜可迁移到消费者金融产品选择、客服政策干预等场景｜用LLM模拟消费者与金融智能体交互，处理为不同政策或模型版本，结果变量为选择或满意度，对照真实A/B测试数据。"}},{"id":"2609.28690","version":1,"title":"Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency","zh_title":"超越表面风格：对齐多轮用户模拟器的行为一致性","abstract":"Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly mitigating reward sparsity and credit assignment in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1, while also achieving the lowest group-level conversion-rate error and semantic trajectory distance, and generalizing to out-of-distribution scenarios. Human Turing tests yield identification accuracy close to chance, supporting the perceived naturalness of generated conversations. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion effectiveness and response quality of LLMs through simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.","authors":["Geng Chen","Ruotong Pan","Zhirui Yang","Qiqi He","Jiawei Chen","Zhang Yunfei","Chongyuan Chen","Minxuan Lv","Zheng Yang","Win-Bin Huang","Xiangyu Wu","Wenwu Ou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28690","pdf_url":"https://arxiv.org/pdf/2609.28690","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["用户模拟","强化学习","行为对齐"],"reason":"用LLM模拟用户行为并与真实交互数据对齐，涉及营销场景，方法可迁移到人类仿真研…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":11,"question":"如何训练多轮用户模拟器，使其不仅语言风格像真人，还能在行为决策和意图演化上与真实用户轨迹对齐？","design":"用 TRACER（基于 LLM 的用户模拟器）扮演顾客服务场景中的用户，通过两阶段训练（先监督微调，再多轮强化学习）对齐真实对话轨迹；强化学习奖励包括会话结果（是否转化）和轨迹对齐（动态时间规整距离），并用偏差感知优势调制解决长对话中的奖励稀疏和信用分配问题。","baseline":"3,866 条真实客服对话，按相似用户条件组织成 762 个参照组，每组包含多条真实轨迹，用于评估模拟行为与真实行为的群体差异。","findings":"TRACER-7B 在转化 F1 上比最强基线高 11.4%，群体转化率误差和语义轨迹距离最低，并能泛化到分布外场景；人类图灵测试识别准确率接近随机，表明生成对话自然。基于该模拟器构建的动态营销基准显示，回复质量高的模型不一定转化率高。","reliability":"论文未讨论模拟器在用户群体异质性、长期决策或非客服场景下的失效条件，仅提到图灵测试在特定研究条件下进行。","relevance":"该研究用真实交互数据对齐 LLM 用户模拟器，并构建群体级评估基准，直接回应了人类仿真中行为一致性和结果复现的核心关切，值得精读其训练与评估方法。","inspiration":"借鉴其两阶段训练和轨迹级奖励设计，将行为一致性作为优化目标而非仅语言风格，并用群体参照组评估模拟分布｜可迁移到消费者金融决策仿真，如信贷申请、保险购买或投资咨询对话中用户意图演化与最终决策的模拟｜以真实银行客服对话为训练数据，用 LLM 模拟借款人在贷款咨询中的多轮交互，处理变量为不同话术策略，结果变量为是否提交申请，并与真实客户群体的申请率分布对照。"}},{"id":"2609.28876","version":1,"title":"Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents","zh_title":"Forecast-Dojo：用于基准测试和训练LLM预测代理的可重放环境","abstract":"We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and accuracy. A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality. Beyond evaluation, Forecast-Dojo provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept.","authors":["Liqin Ye","Haorui Wang","Fardin Ahmed","Rongzhi Zhang","Yuan He","Ziyuan Lin","Yanbin Yin","Jing Peng","Michael Galarnyk","Sudheer Chava","Chao Zhang"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28876","pdf_url":"https://arxiv.org/pdf/2609.28876","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM预测代理","预测市场","人类行为对照"],"reason":"用LLM代理预测市场事件并与真实市场数据对照，属于经济场景仿真，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-26","rank":1,"question":"如何构建一个可回放的预测环境，用于基准测试和训练LLM预测智能体，并评估其在多时间点上的预测表现？","design":"使用12个LLM模型作为预测智能体，在Forecast-Dojo环境中对230个已解决的历史预测市场事件进行预测。每个事件被重放为多个历史日期步骤，智能体在每个步骤可访问截至该日期的新闻文章和计算工具，并输出概率预测。实验比较了三种设置：无工具、有研究工具（无记忆）、有研究工具加信念笔记本（跨日期记忆）。结果变量为Brier分数、准确率和信息比率。","baseline":"历史市场预测（Polymarket市场的实际概率）作为对照基准。","findings":"研究工具降低了所有12个模型的Brier分数，且随着事件进展和更多新证据出现，预测有所改善，但所有模型仍落后于历史市场预测。信念笔记本降低了研究成本，但对预测质量的影响不一致。","reliability":"论文指出所有模型在Brier分数和准确率上均落后于历史市场预测，且市场可能使用了新闻档案之外的信息；信念笔记本对预测质量的改善不一致，仅在部分模型上有效。","relevance":"该研究将LLM作为预测智能体，与真实市场数据对照，属于经济场景仿真，方法可迁移到政策评估和预期形成研究，值得阅读原文以了解环境设计和评估细节。","inspiration":"借鉴其可回放环境设计，通过设置历史信息截止点来模拟不同时点的决策，并利用真实市场数据作为对照基准。｜可迁移到政策公告的预期形成研究，例如央行利率决策或财政政策发布前的市场预期变化。｜以LLM作为预测者，在政策公告前的多个历史日期提供预测，处理为是否提供新闻搜索工具，结果变量为预测准确性和Brier分数，对照真实市场隐含概率或专业预测者调查数据。"}},{"id":"2609.28820","version":1,"title":"AI-Enabled Human Memory Manipulation: Misleading AI-Generated Summaries Distort Human Memory","zh_title":"AI赋能的人类记忆操纵：误导性AI生成摘要扭曲人类记忆","abstract":"AI-generated summaries are increasingly used in high-stakes settings, like policing, despite considerable evidence that AI often generates misleading or inaccurate information. This research asked: do errors in AI-generated summaries distort human memory? To answer this question, we adopted two methodological approaches. First, we conducted an analysis of AI summary output, prompting large language models to generate summaries of videos. This analysis quantified how often AI summaries contain errors and the categories of these errors, revealing the kinds of misleading information that may distort human memory. Second, we conducted a human-subjects experiment to test the impact of misleading information in AI-generated summaries on human memory. Participants were first exposed to an event via watching a video of a car-pedestrian accident, and later read an AI-generated summary describing the video that either contained misleading or accurate information. Participants' memory for the original event was assessed in a memory recognition test. In the AI analysis, we found a high frequency of mistakes in AI summaries, and particularly frequent omissions of critical details. For instance, the majority of summaries omitted the most central event of the video, a critical error which is likely to be impactful. We also observed a strong effect of AI misinformation on human memory. People who read a misleading AI summary were significantly less likely to accurately recall the original event, compared to people who read an accurate AI summary. These findings have implications for how AI should be used in critical settings. Though \"humans-in-the-loop\" are often expected to correct for AI's mistakes, our work suggests human memory can instead be distorted by these mistakes. AI has the potential to generate misinformation, even absent any adversarial intent, which can meaningfully impact human memory.","authors":["Mattea Sim (Georgetown University)","Yael Eiger (University of Washington)","Tadayoshi Kohno (Georgetown University)"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28820","pdf_url":"https://arxiv.org/pdf/2609.28820","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM误导信息","人类记忆","人机交互实验"],"reason":"用LLM生成误导性摘要，测试其对人类记忆的影响，属于用LLM模拟信息源并测量人…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":12,"question":"AI生成摘要中的错误是否会扭曲人类对原始事件的记忆？","design":"本研究不是用LLM模拟人类被试，而是将LLM作为信息源：先用LLM对视频生成摘要并分析错误类型，再让人类被试观看车祸视频后阅读含误导或准确信息的AI摘要，最后用再认测试测量记忆准确性。","baseline":"无对照（人类记忆实验部分没有与真实人类数据对照，但以准确AI摘要组为控制条件）。","findings":"AI摘要错误率高，尤其是关键细节的遗漏；阅读误导性AI摘要的被试对原始事件的记忆准确率显著低于阅读准确摘要的被试。","reliability":"论文未讨论","relevance":"该研究虽非直接仿真人类被试，但揭示了LLM生成信息对人类认知的因果影响，对关注LLM在实验和决策中作用的你具有参考价值，值得一读。","inspiration":"借鉴其将LLM输出作为处理变量、用人类被试测量行为后果的实验设计，可迁移到经济金融中的信息干预场景，如政策公告或财务报告摘要对投资者判断的影响。｜具体可设计：让被试阅读LLM生成的带有误导性信息的公司财报摘要，测量其投资决策或预期，并与真实市场数据或专家摘要对照。"}},{"id":"2609.29513","version":1,"title":"Signed Exposure: Fair Routing of Algorithmic Attention When Attention Can Harm","zh_title":"符号化曝光：当注意力可能造成伤害时算法注意力的公平路由","abstract":"Fairness-of-exposure treats algorithmic attention as a good to be distributed equitably. But when an autonomous agent initiates contact, attention is signed: it delivers value to a willing receiver and imposes a burden on an unwilling one. We formalize routing under signed exposure and show that a fair distribution of attention need not be a fair distribution of unwanted attention. Our central result is an incompatibility: within signed-exposure routing, exposure parity (equal contact rates across groups) and burden parity (equal unwanted-contact rates) generically cannot hold at once, and the two are separated by a band that widens as routing grows more selective. A second result shows measurement error is itself a fairness mechanism: group-differential noise in receptivity scores simultaneously inflates a group's exposure and degrades whom it selects, so an apparent exposure-fairness gain is a hidden burden transfer. Calibrating to a public dating-platform survey (n=2,499) that, to our knowledge, uniquely measures receive-side receptivity to conversational agents, we find exposure parity costs only 0.2--2.3% of yield yet moves the per-capita burden ratio to 1.7 times: the tension is between fairness notions, not between fairness and efficiency. Finally, the burden-parity policy is computable by bisection and learnable online: a plug-in learner recovers it at a $2.3\\%$ empirical regret premium. The operative design choice in signed-exposure markets is not efficiency versus fairness but which fairness.","authors":["Daria Leshchikova","Valentina V. Kuskova","Dmitry Zaytsev","Valerii Klimov"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29513","pdf_url":"https://arxiv.org/pdf/2609.29513","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","公平路由","人类数据对照"],"reason":"用LLM agent模拟路由决策并与人类调查数据对照，涉及公平性权衡，可迁移到…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":14,"question":"在算法主动发起接触的匹配场景中，当注意力可能带来负担时，公平的注意力分配与公平的负担分配能否同时实现？","design":"本文不是仿真研究，而是理论建模与实证校准：构建带符号曝光路由模型，定义曝光率、负担率等指标，推导曝光公平与负担公平的不相容定理；利用一个公开的约会平台调查数据（n=2,499）校准模型参数，计算不同公平策略下的产出损失与负担比；并通过在线学习模拟验证负担公平策略的可学习性。","baseline":"使用一个公开的约会平台用户调查数据（n=2,499），该数据测量了用户对对话代理的接收意愿，作为真实人类基准。","findings":"曝光公平与负担公平在带符号曝光路由中一般不可同时实现，且两者之间的差距随路由选择性增强而扩大。在约会平台数据上，实现曝光公平仅损失0.2-2.3%的产出，但使人均负担比达到1.7倍，表明张力存在于公平概念之间而非公平与效率之间。","reliability":"论文承认其经验量来自单一平台的陈述偏好，测量的是意愿而非行为，且人口统计粒度有限，仅能分析二元性别和粗略年龄组。","relevance":"该研究虽非直接使用LLM进行人类仿真，但其核心关注算法注意力分配中的公平性与负担，与研究者关心的LLM仿真在政策评估中的可靠性与偏差问题高度相关，尤其在涉及主动接触和潜在负面影响的场景中。","inspiration":"本文的带符号曝光框架和公平性权衡分析方法值得借鉴，特别是将注意力视为可能带来负担而非纯粹好处的视角，以及用真实调查数据校准模型参数的做法。｜该框架可迁移到信贷审批中的主动营销、保险推销、政策宣传等场景，分析不同群体在接收算法主动接触时的受益与负担差异。｜可设计一个研究：用LLM模拟不同群体对算法主动营销的接受意愿，施加不同公平路由策略（曝光公平 vs 负担公平），测量模拟的接受率和负担感，并与真实调查数据（如消费者金融调查）对照，评估LLM仿真的可靠性。"}},{"id":"2609.21277","version":2,"title":"How Many Humans Are 32 LLM Judges Worth?","zh_title":"32个LLM法官相当于多少人类？","abstract":"A panel's human-equivalent size is target-specific. Matching a fixed 32-judge panel to empirical human label distributions on three ChaosNLI tasks yields two distinct effective sizes: distributional-error matching gives $\\nu_{\\mathrm{MSE}}=2.304$, $3.750$, and $3.445$, whereas spectral matching gives $\\nu_H=4.242$, $6.459$, and $6.499$, a gap of $1.72$--$1.89\\times$; a binary-error diagnostic credits the same panels with only $1.971$--$2.227$ effective votes. Extrapolating the distributional-error curve at fixed squared mean residual, mean member variance, and normalized mean covariance gives asymptotes of $2.392$, $3.990$, and $3.655$, with 32 judges already reaching $94.0$--$96.3\\%$. An exact spectral identity explains the gap: error depends on member energy and on the orientation of residual variation relative to averaging, information that the participation ratio (PR) discards. A realizable hard-label construction confirms that higher spectral diversity can coexist with worse distribution recovery even under equal member energies and nonnegative correlations, and the consensus direction retains $\\gamma_{\\mathrm{co}}=43.8\\%$, $33.7\\%$, and $35.9\\%$ of centered residual variance. An external check on CC-1000, a 1,000-item Civil Comments subset with a different panel, gives $\\nu_H=2.84$. For panel choice, we establish an existence result and one feasible path: exhaustive enumeration at $k\\in\\{5,7\\}$ shows that panels beating the accuracy-top-$k$ baseline on both accuracy and $\\nu_H$ always exist, and greedily swapping at most two members reaches $24.8$--$56.0\\%$ higher $\\nu_H$ at $0.10$--$1.10$ percentage points higher accuracy. Our dataset and code are available at https://github.com/Chao1208/32judges-votes.","authors":["Chao Li","Yingying Yu","Yunfeng Li"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-21","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.21277","pdf_url":"https://arxiv.org/pdf/2609.21277","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","人类等效","标注可靠性"],"reason":"研究LLM法官替代人类标注，非仿真人类被试，但涉及人类标签分布对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":10,"question":"一个由语言模型组成的法官面板在多大程度上能代表人类判断？","design":"该研究使用32个语言模型（来自10个提供商家族）作为法官，在三个自然语言推理任务（MNLI-m、SNLI、NLI）上对每个项目给出分类标签，并与每个项目100个人类标注的标签分布进行对比。通过匹配归一化残差Gram矩阵的参与率（PR）来测量谱残差多样性（得到有效规模nu_H），并通过匹配分布平方误差来测量分布恢复（得到nu_MSE）。","baseline":"使用ChaosNLI数据集中每个项目100个人类标注的标签分布作为基准。","findings":"32个法官面板的谱有效规模nu_H为4.24-6.50，而分布误差匹配规模nu_MSE为2.30-3.75，表明谱多样性与分布恢复并不一致。谱多样性更高的面板可能分布恢复更差，且不同任务中排名一致性差异很大。","reliability":"论文指出有效规模是目标特定的测量，谱多样性和分布恢复不应互换使用；面板的排名一致性随任务变化，某些成员添加在不同项目半区产生冲突变化；模型标识符不独立认证服务提供商的底层版本。","relevance":"该研究直接评估LLM法官面板与人类标签分布的一致性，提出了有效规模度量，并有真实人类数据对照，对关注LLM仿真可靠性和偏差的研究者具有参考价值。","inspiration":"借鉴其将模型判断与人类分布进行多维度匹配（谱多样性和分布误差）的方法，可迁移到经济金融领域的专家预测或消费者调查仿真中。｜例如，在资产定价实验中，用LLM模拟投资者对新闻的情绪反应，并对比真实投资者调查数据。｜设计：使用多个LLM作为被试，施加不同的财经新闻处理，测量其情绪分类或投资决策分布，并与真实投资者调查或实验数据对照，评估LLM仿真的有效规模和偏差。"}},{"id":"2609.27535","version":1,"title":"KITE: Scaling Jev Population Experiments with Sparse Flagship Calibration","zh_title":"KITE：通过稀疏旗舰校准扩展Jev人口实验","abstract":"KITE queries a typed behavioral kernel once per unique state, then executes populations of any size from the table with event-keyed randomness and common random numbers. An expensive flagship model is reserved for sparse paired anchors that estimate intervention effects. Measured human-model discrepancy is propagated as shared error into every conclusion. Population-experiment cost thus scales with unique states and anchors, while uncertainty is governed by evidence about people rather than Monte Carlo noise. On Epstein experiments with 9,070 participants, anchors covering 1.7% of states reduced effect error by 41% (absolute MAE reduction 0.0125). On 37 held-out SocSci210 experiments, 0.5-1.5% anchor coverage raised captured decision gain from 0.27 to 0.39. The kernel passed content-fidelity criteria in all 15 new countries of a 16-country study. Shared discrepancy yielded retrospective coverage of 93% and 96% at nominal 80% and 90%, versus 29% and 36% from human sampling uncertainty alone. A million agents executed 20 tabulated steps in 0.9 seconds on a laptop. This architecture offers a route to screening candidate interventions before human trials, multi-country content audits, and uncertainty-aware policy comparison at the cost of a few thousand kernel calls with sparse flagship anchors. Property-specific evidence records connect each use to its validation scope, correction provenance, and uncertainty, making these applications auditable.","authors":["Hengyu Li (The University of Tokyo)"],"categories":["cs.MA","cs.CY"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27535","pdf_url":"https://arxiv.org/pdf/2609.27535","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","人类数据对照","政策评估"],"reason":"用LLM仿真人类被试，有真实人类数据对照，评估偏差并传播不确定性，用于政策评估。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":5,"question":"如何在大规模人口实验中，用廉价的类型化行为核模型结合稀疏的旗舰模型校准，来估计干预效应并传播人类-模型差异的不确定性？","design":"KITE 使用 TypeSafe 的 Jev 行为核模型（类型化行为核）对每个唯一状态查询一次，生成决策分布，然后用表格化执行模拟任意规模的人口，并通过事件键控随机数和共同随机数控制变异性；昂贵的旗舰模型仅用于稀疏的配对锚点，以估计干预效应；人类-模型差异被作为共享误差传播到所有结论中。","baseline":"对照的真实人类数据包括 Epstein 实验（9,070 名参与者）和 37 个留出的 SocSci210 实验，以及一个 16 国研究中的内容保真度标准。","findings":"在 Epstein 实验中，覆盖 1.7% 状态的锚点将效应误差降低了 41%（绝对 MAE 降低 0.0125）；在 37 个留出的 SocSci210 实验中，0.5-1.5% 的锚点覆盖率将捕获的决策增益从 0.27 提高到 0.39。共享差异传播在名义 80% 和 90% 的置信水平下分别实现了 93% 和 96% 的回顾性覆盖率，而仅使用人类抽样不确定性时分别为 29% 和 36%。","reliability":"论文承认的局限包括：仅有两个留出的人类参考数据集测试混合方法；效应幅度需要进一步校准，差异斜率在不同研究选择间变化；稀疏校准可能引入旗舰模型误差；核模型可能过度预测规范信息；实验不授权预测文献中不存在的干预；记忆一致性是局部的，队列级边际校正可能抹去真实的持续性或处理路径；刺激重建、省略卡片图像、人口统计压缩等限制了保真度；专有模型和访问限制阻碍了可重复性。","relevance":"该研究直接针对用 LLM 仿真人类被试的核心问题，提供了与真实人类数据对照的验证，并传播不确定性，对评估仿真可靠性和偏差具有重要参考价值，值得精读原文。","inspiration":"该方法通过稀疏旗舰模型校准和共享误差传播，在保持低成本的同时提高了效应估计的准确性，值得借鉴其校准策略和不确定性量化方法。｜可迁移到政策评估场景，如税收政策变化对劳动供给的影响、福利项目对消费行为的影响，或信息干预对金融决策的影响。｜设计一个实验：用 Jev 核模型模拟不同人口群体对政策公告的反应，以真实调查数据（如消费者预期调查）为基准，施加政策处理（如利率变化信息），测量预期调整和消费意愿，并用稀疏旗舰模型校准关键状态，传播人类-模型差异。"}},{"id":"2609.28470","version":1,"title":"StudentBench: AI and human tutoring yield equivalent GRE learning gains","zh_title":"StudentBench：AI与人类辅导在GRE学习收益上等效","abstract":"Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.","authors":["Curtis Northcutt","Inaara Hasmani","Kevin Feng","Trevor Khangi","Andreas Plesner","Jonas Mueller"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28470","pdf_url":"https://arxiv.org/pdf/2609.28470","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","教育实验","人类对照"],"reason":"用LLM替代人类导师进行教学实验，并与人类导师对照，评估学习效果，属于人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":6,"question":"大语言模型（LLM）作为AI导师能否在GRE学习收益上达到与人类专家导师统计等效的效果？","design":"用多个LLM（如Gemma 4 31B等）扮演AI导师，对2383名人类参与者进行GRE定量和语文部分的辅导，同时设置人类导师辅导组和无辅导对照组，测量学习收益（前后测成绩提升百分比），并收集175,000条学生-AI消息分析对话行为。","baseline":"人类专家导师辅导组的学习收益数据，以及无辅导对照组的学习收益数据。","findings":"AI辅导在GRE学习收益上与人类专家辅导统计等效（p=.015），且在七个GRE领域中五个领域的最佳AI导师平均超过人类导师。一个AI导师（Gemma 4 31B）以低918倍的成本实现了与人类辅导等效的学习收益（p=.044）。","reliability":"论文承认局限：只测量即时学习收益，未评估长期保持；参与者均为能读写英语的成年人，未测试跨语言、设备或教育环境；未设置学生独自练习的对照组，无法分离AI交互的额外收益；排行榜评估基于专家评价和对话行为，而非实际学习效果。","relevance":"该研究用LLM替代人类导师进行教学实验，并与真实人类导师及无辅导组对照，评估学习效果，属于典型的人类仿真研究，且提供了大规模真实人类数据作为基准，值得精读以了解仿真等效性检验的设计与局限。","inspiration":"借鉴其多组对照设计（AI处理、人类处理、无处理）和统计等效性检验（TOST）来严格评估AI干预是否非劣于人类专家。｜可迁移到金融教育或投资者决策辅导场景，例如测试AI投教助手能否在提升投资者金融素养或改善投资决策上达到人类理财顾问的效果。｜以真实投资者为被试，随机分为AI投教组、人类顾问组和无辅导组，处理为一段时间的个性化金融知识辅导，结果变量为金融素养测试得分或模拟投资组合表现，对照真实人类顾问组和无辅导组的数据，并采用等效性检验。"}},{"id":"2609.28372","version":1,"title":"Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer","zh_title":"算法购物：代理式AI如何将人类启发式用作替代消费者","abstract":"Consumers increasingly delegate purchasing decisions to Large Language Models (LLMs) acting as surrogate consumers. Using \"Tool-Lab,\" an adaptation of information-board process tracing that places product attributes behind costly tool calls, we examine how marketing pricing cues (i.e., just-below pricing and promotional framing) influence AI shopping agents. Across eight commercially deployed LLMs from three providers, we trace pre-choice information acquisition. Under zero cost, pricing cues rarely mislead. Imposing acquisition costs under a vague goal prompt leads LLMs to omit diagnostic attributes required to compute unit price and choose suboptimal choices resembling human heuristics. Relative to a specific goal prompt that mainly preserves diagnostic search and choice optimality, a vague goal prompt under constraints creates a search-mediated vulnerability. This research demonstrates that marketing heuristics in delegated AI shopping are governed by storefront information architecture, not necessarily immutable LLM flaws.","authors":["Davood Wadi","Yu Ma"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28372","pdf_url":"https://arxiv.org/pdf/2609.28372","source_feed":"econ.GN","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","消费者行为","算法保真度"],"reason":"用LLM作为代理消费者模拟人类购物决策，并与人类启发式对照，涉及营销实验场景。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":9,"question":"在委托AI代理购物时，营销定价线索（如尾数定价和促销框架）如何通过信息获取成本与目标提示的具体性影响LLM的信息搜索和选择最优性？","design":"使用Tool-Lab实验范式，将产品属性隐藏在需要付费的工具调用之后，操纵信息获取成本（0、1、5美分）和提示目标具体性（模糊：“找最划算的” vs. 具体：“找每盎司最低价”），对8个商用LLM（来自Google、OpenAI、Anthropic）进行咖啡选择实验，测量信息搜索深度、搜索组成和选择最优性。","baseline":"无对照","findings":"在零成本或具体目标提示下，LLM大多做出规范最优选择；但在获取成本存在且目标模糊时，多数LLM会减少搜索深度，省略诊断性属性（如美分或重量），导致次优选择，类似于人类启发式决策。","reliability":"论文未讨论","relevance":"本研究通过实验操纵环境约束（成本与提示），揭示了LLM启发式行为的条件性，为评估LLM作为人类被试替代品的可靠性提供了关键证据，值得精读原文以理解其方法细节和边界条件。","inspiration":"借鉴其通过工具调用施加信息获取成本并操纵提示具体性的设计，可迁移到消费者金融决策（如信用卡选择、贷款比较）或投资者信息处理场景；例如，用LLM模拟投资者在获取公司财务指标需付出成本时，模糊目标（“选只好股票”）与具体目标（“选市盈率最低的股票”）下的信息搜索与选择，并与真实投资者眼动或点击流数据对照。"}},{"id":"2608.22859","version":2,"title":"WARP: Wasserstein-Aligned RAG for Population Opinions","zh_title":"WARP：面向群体意见的Wasserstein对齐检索增强生成","abstract":"RAG systems are increasingly used to summarize what large collections of documents say. A user asks \"What do people think about X?\" and receives an answer that reads as consensus. But standard top-k retrieval ranks documents by query similarity, not by how faithfully they represent the population, so minority views quietly disappear. Existing fixes fall short. Diversity re-rankers like MMR and DPP spread retrieved documents apart, but with no target distribution to aim for. Calibration methods based on KL or JS divergence do target one, yet treat opinion bins as unordered: confusing strong positive with strong negative costs no more than an adjacent-bin miss. We introduce WARP, a family of post-retrieval algorithms that calibrate retrieved evidence to the population's opinion distribution. WARP first recovers underrepresented opinions that cosine ranking may bury, then uses Wasserstein-1 distance to select documents whose sentiment-intensity distribution matches the population target, capturing the ordinal structure ignored by KL and JS divergence. We develop three variants for dense, sparse, and variable candidate pools, trading off calibration quality and speed. Across three review domains spanning 35K documents, 156 queries, and 26 entities, WARP's domain-matched variants reduce distributional error by at least 43% with sub-second latency. These gains carry through to generation: a five-judge LLM panel prefers WARP-generated answers in 86% of decided comparisons at k <= 5.","authors":["Aman Singh Thakur","Aditya Agrawal","Alwarappan Nakkiran","Alex Karlsson"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-09-24","first_seen":"2026-08-25","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2608.22859","pdf_url":"https://arxiv.org/pdf/2608.22859","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","意见分布校准","RAG"],"reason":"用LLM生成代表人群意见的摘要，并与真实意见分布对齐，属于仿真人类态度，且有真…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":10,"question":"如何让检索增强生成（RAG）系统在回答关于人群意见的查询时，所选证据的意见分布忠实于总体人群的意见分布，避免多数意见淹没少数意见？","design":"该论文提出 WARP 算法族，用于在 RAG 流程的后检索阶段校准所选文档的意见分布。它首先通过缺陷感知的池扩展策略恢复被余弦相似度排序埋没的少数意见文档，然后使用 Wasserstein-1 距离作为度量，贪婪地选择文档子集，使其情感强度分布与目标人群分布匹配。论文在三个评论数据集（Amazon 卖家论坛、Yelp 酒店评论、OpinRank 汽车评论）上进行了实验，共 35K 文档、156 个查询、26 个实体，比较了 WARP 与 Top-k、MMR、DPP、OpinionMMR、KL/JS 校准等基线在分布误差、实体匹配率和延迟上的表现，并通过 LLM 评委小组评估生成答案的质量。","baseline":"论文使用从语料库中提取的实体级情感强度分布作为目标人群分布，该分布由 LLM 对每个文档进行情感强度标注后聚合得到，作为真实人群意见分布的代理。","findings":"WARP 的领域匹配变体在三个评论领域中将分布误差降低了至少 43%，且重排序延迟低于 310 毫秒（p99）。在生成评估中，五位 LLM 评委在 k≤5 时对 WARP 生成答案的偏好率高达 86%。","reliability":"论文指出，WARP 的性能依赖于实体密度：在实体稀疏的领域，需要混合变体（如 λ-MMR 或 WassRank OT）来平衡相关性与校准；此外，目标分布是从语料库中估计的，可能无法完全代表真实人群，但论文未深入讨论这一局限。","relevance":"该论文直接针对用 LLM 生成代表人群意见的摘要问题，通过 Wasserstein 距离对齐意见分布，属于仿真人类态度分布的研究，且有真实评论数据作为基准，值得阅读原文以了解其算法细节和评估方法。","inspiration":"该方法借鉴了将检索证据校准到目标分布的思想，并利用 Wasserstein 距离捕捉有序情感结构，可用于经济金融领域中需要从文本中提取群体意见或情绪分布的场景。｜可以迁移到消费者信心指数构建、财报电话会议情绪分析、社交媒体上的政策预期形成等场景，其中需要从大量文本中估计人群的态度分布。｜一个可行的研究设计是：以 Twitter 上关于某经济政策（如加息）的推文为语料，用 LLM 对每条推文进行情感强度标注（如从强烈反对到强烈支持），得到目标分布；然后使用 WARP 算法从检索到的推文中选择一小部分作为 LLM 生成摘要的证据，确保所选推文的情感分布与总体分布一致；最后将生成的摘要与基于随机抽样或简单检索的摘要进行对比，评估其在预测真实调查数据（如密歇根消费者信心指数）上的准确性。"}},{"id":"2609.27165","version":1,"title":"Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement","zh_title":"计数证据而非句子：面向长文本价值测量的LLM判断调和证据融合","abstract":"Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance-bearing sentences. Existing approaches either ask the model to predict a document-level label directly, which can be overconfident, or aggregate sentence-level predictions by majority or soft voting, which treat uncertain and decisive sentences as equally informative. We formulate long-text value measurement as a decision-fusion problem and propose Tempered Evidence Fusion (TEF), a training-free rule that weights each sentence's log-odds by its normalized information gain, as derived from a generalized Bayesian posterior. This makes the fused score nearly vanish for uncertain sentences while preserving the Bayes-optimal weight of decisive evidence. We further introduce Multi-event Insight Network Dimensions (MIND), a benchmark of 8,358 Chinese and English posts spanning five years of public events and six value dimensions. On MIND, TEF outperforms the strongest baseline among Direct, Majority Vote, and Soft Vote by an average of 4.5 accuracy points and 4.6 macro-F1 points across five LLMs and two languages. MIND dataset and code are available at https://github.com/Kzczc/ICASSP2027-TEF.","authors":["Yuhe Wu","Rui Qian","Guangyu Wang","Yuran Chen","Yuanchao Zhu","Junjie Yang","Zhengheng Li","Jiulin Cai","Tianyi Zhang","Zihan Dong","Jiaxin Liu","Yujie Chen","Guang Zhang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27165","pdf_url":"https://arxiv.org/pdf/2609.27165","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM价值测量","证据融合","长文本分析"],"reason":"用LLM测量公众价值取向，有真实人类标注数据对照，方法可迁移到仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":13,"question":"如何融合长文本中句子级LLM判断，使决定性证据主导最终价值取向测量，同时抑制不确定句子？","design":"本研究不是人类仿真实验，而是提出一种训练无关的决策融合规则TEF，用于融合句子级LLM判断来测量长社交媒体帖子的价值取向。具体做法：将长文本分割为句子，用LLM对每个句子输出两个立场标签的对数概率，计算对数几率，并用归一化信息增益加权，再求和得到文档级得分。在MIND基准（8,358条中英文帖子，六个价值维度）上，用五个LLM（Qwen2.5-7B、LLaMA3-8B、Qwen3-14B、DeepSeek-V3.2、GPT-4o-mini）评估TEF与直接预测、多数投票、软投票的性能。","baseline":"有真实人类标注数据：MIND基准包含8,358条中英文社交媒体帖子，由人类标注了六个价值维度上的立场标签。","findings":"TEF在五个LLM和两种语言上平均比最强基线（直接预测、多数投票、软投票）高出4.5个准确率点和4.6个宏F1点。TEF在Qwen2.5-7B上校准误差最低，且对提示词扰动更稳健。","reliability":"论文未明确讨论失效条件，但指出直接求和LLM logits不可靠，因为next-token概率只是近似校准，尤其在最大不确定点附近；消融实验显示，在对数几率变换和熵加权同时移除时，某些模型上性能会低于软投票，表明两者互补。","relevance":"该研究虽非人类仿真实验，但提供了用LLM测量公众价值取向的可靠方法，且有真实人类标注数据对照，可作为仿真研究中测量态度或价值观的工具，值得阅读原文了解融合规则细节。","inspiration":"借鉴其决策融合思路，将长文本分解为句子级判断并用信息增益加权融合，可提高LLM对复杂文本的测量准确性｜可迁移到经济金融领域的文本测量，如从财报电话会议记录中提取管理层情绪、从新闻中测量政策不确定性、从社交媒体帖子中测量消费者信心或通胀预期｜设计：用LLM对财报电话会议记录的每个句子判断管理层语气（积极/消极），用TEF融合句子级判断得到文档级情绪得分，以分析师一致预期或后续股票收益作为真实数据对照，评估LLM情绪测量的预测效度。"}},{"id":"2609.26861","version":1,"title":"Rule-Based Pricing Algorithms and Market Outcomes: An Experimental Study","zh_title":"基于规则的定价算法与市场结果：一项实验研究","abstract":"Rule-based pricing tools are widespread in digital commerce, yet we know little about how their design shapes market outcomes. In a controlled market experiment, participants use dashboards to build pricing algorithms competing in a sequential Bertrand game over multiple periods. We vary design features commonly found in commercial repricing tools: warnings about price wars, pre-configured strategies, and advice from a large language model. Most treatment variations raise market prices with effects driven by an increase in starting prices and more cooperative algorithm designs. The results matter for competition policy, platform regulation and current discussions on regulating algorithm design tools.","authors":["Adrian Hillenbrand","Hans-Theo Normann","Matthias Potarca","Tobias Werner"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26861","pdf_url":"https://arxiv.org/pdf/2609.26861","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM建议","经济实验","算法定价"],"reason":"用LLM提供建议影响人类定价实验，有真实人类数据对照，属经济实验场景，可迁移到…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":12,"question":"规则型定价算法的设计特征（价格战警告、预配置策略、LLM建议）如何影响市场结果与合谋行为？","design":"受控市场实验：人类被试通过仪表盘构建定价算法，在序贯伯特兰博弈中竞争50期，共5轮超级博弈；处理为三种设计特征：盈利性提示（警告价格战）、LLM建议、预配置策略菜单与默认合作性价格匹配；结果变量为市场价格、起始价格、算法合作性设计。","baseline":"基线处理（无设计特征）作为对照，比较不同处理组与基线组的价格和算法设计差异。","findings":"大多数处理变体提高了市场价格，效应主要由起始价格提高和更合作的算法设计驱动。LLM建议和预配置策略等设计特征可能促进合谋，对竞争政策和平台监管有启示。","reliability":"论文未讨论","relevance":"该研究用LLM提供建议影响人类定价实验，有真实人类数据对照，属经济实验场景，可迁移到LLM仿真人类决策的可靠性评估，值得读原文了解LLM建议的具体效果。","inspiration":"借鉴其将LLM建议作为处理变量嵌入人类实验、并观察对策略选择和均衡结果影响的设计｜可迁移到资产定价实验或政策公告预期形成场景，研究LLM建议对投资者行为或公众预期的影响｜以人类被试为对象，处理为是否提供LLM投资建议，结果变量为报价或预期值，对照真实市场数据或调查数据评估LLM建议的偏差与可靠性"}},{"id":"2608.19621","version":3,"title":"Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories","zh_title":"用纵向生命轨迹缓解LLM智能体的身份本质主义","abstract":"Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce population-level patterns, yet often fail to capture human-like diversity. Our analysis shows that static-profile agents exhibit stronger demographic separation and within-group compression than humans, a pattern consistent with identity essentialism: demographic labels can encourage models to treat group-average tendencies as individual traits, homogenizing responses within groups. We argue that this limitation arises from two related factors: sparse, static agent representations and the limited ability of prompt-only memory to persistently integrate experience. Inspired by complementary memory systems, we propose LifeMem, a longitudinal memory framework that combines structured life-event retrieval with agent-specific parametric memory for experience integration. Experiments on Understanding Society with three LLMs show that LifeMem improves alignment with human data in terms of response distributions, overall and within-group diversity, and patterns of within-person response change across life stages. These findings highlight the value of longitudinal life-event memory for constructing more faithful and dynamically evolving social agents.","authors":["Hexi Wang","Yujia Zhou","Bangde Du","Weihang Su","Xinyuan Cao","Qingyi Pan","Qingyao Ai","Yueyue Wu","Min Zhang","Yiqun Liu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-23","first_seen":"2026-08-21","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2608.19621","pdf_url":"https://arxiv.org/pdf/2608.19621","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人类数据对照","社会调查"],"reason":"用LLM agent模拟人类调查数据，并与真实面板数据对照，改进仿真保真度。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":1,"question":"如何通过引入纵向生活轨迹记忆来缓解LLM社会仿真中的身份本质主义，从而提升模拟人类多样性与动态变化的保真度？","design":"使用三个指令微调LLM（Llama-8B、Ministral-8B、Qwen3.5-9B）作为社会仿真智能体，基于Understanding Society面板数据构建个体静态画像和纵向生活事件流，施加LifeMem框架（结构化生活事件记忆+个体特定LoRA参数记忆）作为处理，与静态画像、多样性提示、非参数记忆等基线对比，测量回答分布、组内多样性、组间差异及个体跨波次变化等结果变量。","baseline":"Understanding Society英国家庭纵向调查的真实个体面板数据，包括静态背景信息和多波次生活事件及主观态度回答。","findings":"静态画像智能体表现出更强的组间分离和组内压缩，符合身份本质主义特征；LifeMem通过结合结构化事件检索和参数化记忆整合，在回答分布、总体及组内多样性、跨生命阶段个体变化模式上均提升了与人类数据的一致性。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真中多样性塌缩和身份本质主义问题，使用真实面板数据作为基准，并提出了可操作的记忆框架，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文了解具体实现和效果。","inspiration":"借鉴其双记忆系统设计，将显式事件检索与参数化个体状态结合，可迁移到经济金融中的个体决策仿真，如消费者跨期选择或投资者行为演化。｜例如，在信贷审批歧视研究中，用LLM智能体模拟不同背景的申请人，施加纵向财务事件记忆处理，测量审批决策的组间差异和组内多样性，并与真实信贷数据对照。｜设计：以LLM智能体作为虚拟被试，处理为是否注入个体历史财务事件（如收入变动、失业）并更新LoRA参数，结果变量为贷款审批通过率和利率设定，对照真实信贷记录数据评估仿真偏差。"}},{"id":"2609.25066","version":1,"title":"Understanding Reliability in LLM-based Human Behavior Simulation","zh_title":"理解基于LLM的人类行为仿真的可靠性","abstract":"Large language models (LLMs) are increasingly used to simulate human survey responses and behavioral reactions, yet unreliable simulations can mislead social science conclusions. However, existing evaluations focus on end-to-end scores, leaving it unclear how different aspects of the simulation process interact to determine reliability. We propose ReliMap, which decomposes LLM-based human behavior simulation into three structured layers and evaluates reliability at both the individual level (R1) and population level (R2) across three configuration dimensions: model capacity, profile completeness, and population coverage. Through experiments across four simulation tasks and eleven LLMs, we find that all models exhibit substantial distributional bias without profile conditioning. Profile conditioning reduces this bias with diminishing returns. Larger models benefit more, and attribute informativeness matters more than quantity. Critically, R1 gains do not reliably transfer to R2--individual and population-level reliability can move in opposite directions. At the population layer, increasing coverage reduces variance but not systematic bias, with R2 stabilizing at around 50-100 individuals. These findings highlight that reliable simulation cannot be achieved by optimizing any single layer in isolation, but requires coordinated improvement across all three.","authors":["Pei Wang","Lei Wang","Yuanzi Li","Xu Chen"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25066","pdf_url":"https://arxiv.org/pdf/2609.25066","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B4"],"tags":["LLM仿真","可靠性评估","人类行为"],"reason":"直接研究LLM仿真人类行为的可靠性，分解评估层次，含真实人类数据对照，并指出失…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":3,"question":"LLM 仿真人类行为时，模型容量、画像完整度和人群覆盖度如何共同影响个体层与群体层的可靠性？","design":"用 11 个 LLM 在 4 个任务（党派偏好、移民态度、宗教立场、社交媒体事件态度）上仿真人类回答，通过改变模型容量、画像属性数量和人群样本量，测量个体准确率（ACC）和群体分布距离（TVD）。","baseline":"真实人类调查数据：欧洲社会调查（ESS）、世界价值观调查（WVS）、SocioBench 宗教立场数据，以及社交媒体上关于瑞幸咖啡股价暴跌的真实公众态度语料。","findings":"无画像条件时所有模型都存在显著分布偏差；画像条件化可减少偏差但边际收益递减，且大模型受益更多。个体层可靠性提升不必然转化为群体层可靠性，群体层在 50-100 人后趋于稳定但系统偏差不随覆盖度增加而减小。","reliability":"论文指出个体层与群体层可靠性可能反向变动，仅优化单一层无法保证整体可靠性；画像属性信息量比数量更重要，但未明确给出所有失效条件，仅强调需三层协同改进。","relevance":"该研究系统拆解了 LLM 仿真人类行为的可靠性层次，并基于真实调查数据对照，直接回应了仿真在什么条件下会失效的问题，对关注经济学实验和政策评估仿真的研究者很有参考价值。","inspiration":"借鉴其分层评估框架，将个体预测准确率与群体分布距离分开考察，并系统变化模型、画像和样本量以定位偏差来源｜可迁移到消费者金融决策仿真，如信用卡选择、退休储蓄计划参与或风险偏好调查｜用多个 LLM 扮演不同人口学特征的消费者，施加不同金融素养或收入冲击处理，测量选择分布，并与美国消费者金融调查（SCF）或美联储家庭经济决策调查（SHED）的真实数据对照，检验仿真在个体和群体层的可靠性。"}},{"id":"2609.25010","version":1,"title":"Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation","zh_title":"合成人物角色能预测真实受众反应吗？一项无人物角色基线优于基于人物角色文案仿真的仿真到现实研究","abstract":"Marketers increasingly use large language models (LLMs) as \"synthetic personas\" to predict how an audience will react to a piece of copy before it ships, encouraged by evidence that profile-conditioned LLMs mimic human samples. But is that prediction actually valid against real behaviour - and does the persona machinery help? We present a sim-to-real validity study using the Upworthy Research Archive - thousands of headline A/B tests on shared real traffic, with measured click-through - as held-out ground truth. We compare a ten-persona panel, grounded in the real audience's demographics, against a no-persona zero-shot baseline that simply asks the model how likely a typical reader is to click. Two findings stand out. First, ground-truth reliability is the binding constraint: most A/B tests have no statistically distinguishable winner, so validity can only be measured on the reliable subset (n = 399). Second, and counter to the persona-simulation premise, persona conditioning degrades predictive validity: the no-persona baseline ranks variants markedly better (Kendall {\\tau} = 0.361, a medium effect; top-1 accuracy 49.2%) than the persona panel ({\\tau} = 0.084; top-1 34.6%), with non-overlapping confidence intervals. Asking the model directly taps an accurate population-level prior; forcing it to role-play specific personas injects bias and noise. The result replicates across three independent Upworthy splits, holds in direction on a different-domain news dataset, and is robust to seed, prompt phrasing, and model choice - across three Gemini tiers and a different model family (OpenAI gpt-4.1, significant paired gap). The takeaway: for predicting aggregate engagement, a plain LLM ranker beats persona simulation - synthetic personas are not merely a weak predictor, they are worse than not using them. All numbers regenerate from a public, artifact-first replication package.","authors":["Alexandre Cristov\\~ao Maiorano"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25010","pdf_url":"https://arxiv.org/pdf/2609.25010","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","受众预测","算法保真度"],"reason":"用LLM仿真受众点击行为，与真实A/B测试数据对照，发现persona降低预测…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":1,"question":"在真实A/B测试数据上，基于合成人设的LLM仿真能否预测真实受众对文案的点击行为，且人设条件化是否比无人设基线更有效？","design":"使用Upworthy Research Archive中的数千个标题A/B测试作为真实流量数据，构建基于真实受众人口统计特征的十人设面板，用LLM（gemini-3.1-flash-lite）分别以人设条件和无人设零样本基线预测标题点击意图，聚合为排名，与真实点击率排名比较。","baseline":"Upworthy Research Archive中真实A/B测试的点击率数据，作为留出集真实行为基准。","findings":"大多数A/B测试无统计显著赢家，因此只能在可靠子集（n=399）上评估效度；人设条件化降低预测效度，无人设基线排名显著优于人设面板（Kendall τ=0.361 vs 0.084，top-1准确率49.2% vs 34.6%）。","reliability":"论文指出真实数据可靠性是主要约束，多数A/B测试无显著赢家；人设仿真在可靠子集上仍表现差，且结果对种子、提示措辞和模型选择稳健，但人设条件化本身引入偏差和噪声。","relevance":"该研究直接检验LLM合成人设仿真在真实受众行为预测中的效度，发现人设条件化反而降低预测力，对关注LLM仿真可靠性及偏差的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其sim-to-real效度框架和可靠子集筛选方法，用真实行为数据作为基准，比较不同仿真策略（如人设 vs 无人设）的预测效度，并采用排名指标和bootstrap置信区间｜可迁移到经济金融中的消费者选择预测，如广告文案对点击率的影响、金融产品描述对投资意愿的影响、政策沟通对公众反应的影响等｜以真实A/B测试数据（如某平台广告实验）为基准，用LLM分别以人设面板和无人设基线预测用户对金融产品广告的点击或选择，比较排名准确率，并筛选出有显著差异的测试子集进行评估。"}},{"id":"2609.25677","version":1,"title":"Seeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing","zh_title":"眼见不为实：合成消费者何时能及不能预测试视觉营销","abstract":"Marketers now deploy generative AI agents as synthetic consumers to pretest visual assets such as logos, packaging, and advertising at a fraction of human-panel cost. However, this procedure assumes that a model seeing a visual cue can also perceive its consumer meaning, which is largely untested. We stress-test the assumption using six canonical visual marketing experiments, varying the two levers managers control: model generation (GPT-4o-mini vs. GPT-5.4-mini) and input format (plain text vs. JSON). Every resulting configuration passed the manipulation checks; however, none of the configurations reproduced more than two of the six human effects, and the remainder were nonsignificant. The one exception was a significant reversal of the human pattern. Providing conceptual or empirical evidence through in-context learning steers average responses toward the human effect. Yet steering has a limit: even when it succeeds, a configuration reproduces less than half of the natural spread of human responses and so understates consumer heterogeneity. We integrate these results into an AI governance protocol (Calibrate, Intervene, Deploy) that delineates when synthetic consumers can responsibly screen creatives and when human panels remain necessary.","authors":["Yi-Lin Tsai (Arvin)","Yung-Hsiu (Arvin)","Lai"],"categories":["cs.AI","cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25677","pdf_url":"https://arxiv.org/pdf/2609.25677","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","消费者行为","算法保真度"],"reason":"用LLM作为合成消费者复现视觉营销实验，与真实人类数据对照，评估仿真可靠性并指…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":2,"question":"合成消费者在视觉营销实验中能否像人类一样感知视觉线索并复现人类判断，其失效边界和可修复性如何？","design":"使用 GPT-4o-mini 和 GPT-5.4-mini 两个模型，以纯文本和 JSON 两种输入格式，模拟人类被试回答六个经典视觉营销实验，测量数值评分和文本理由，并通过操纵检查、主题建模和上下文学习（概念证据和实证证据）进行干预。","baseline":"六个经典视觉营销实验的原始人类样本数据（每个研究 69 到 220 名被试）。","findings":"所有配置都通过了视觉操纵检查，但没有一个配置能复现超过两个人类效应，其余均不显著，甚至出现一个显著反转。提供概念或实证证据的上下文学习能将平均响应拉向人类效应，但即使成功，也无法复现人类响应自然分布的一半，低估了消费者异质性。","reliability":"论文承认合成消费者在视觉领域感知不一致，即使通过上下文学习校准，也无法恢复消费者异质性，因此只适合平均效应问题，不适合细分或定位。","relevance":"该研究直接检验了 LLM 作为人类被试在视觉营销实验中的可靠性，与真实人类数据对照，并揭示了失效条件，对关注仿真偏差和边界的研究者具有重要参考价值。","inspiration":"借鉴其多模型、多输入格式的因子设计，以及通过操纵检查和主题建模分离“看见”与“感知”的方法，可迁移到经济金融中的视觉信息处理场景，如央行沟通中的图表设计、金融产品广告或信贷审批中的视觉线索。｜设计一个实验，用 LLM 模拟投资者，呈现不同颜色或形状的金融图表（如涨跌颜色、风险提示图标），测量其风险感知和投资决策，并与真实投资者实验数据对照，检验 LLM 是否复现视觉线索对风险偏好的影响。"}},{"id":"2609.25760","version":1,"title":"The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance","zh_title":"模拟社会的局限：后训练与调查微调如何抹除跨文化方差","abstract":"Using large language models (LLMs) to simulate diverse human populations has the potential to transform many aspects of computational social science, yet many evaluations score the average response rather than the spread of opinion within real groups. Here, we develop a diagnostic framework that measures point accuracy alongside dispersion retention, the ratio of predicted to human standard deviation ($\\dr$), on 10{,}000 respondent--question pairs from the World Values Survey (WVS) spanning twelve countries and six continents. We evaluate eleven zero-shot language models and five variants fine-tuned on WVS data with SFT, DPO, and GRPO. We identify a failure mode we term \\textit{consensus collapse}, where alignment training compresses outputs toward one stereotype per group. Along the post-training trajectory from the Llama~3.1 70B base to the Tulu~3 checkpoints, the first stage, supervised instruction tuning, removes half of the spread with minimal accuracy gain ($\\dr$ 1.22 to 0.59; accuracy $+0.9$ points), the later stages do not restore it, and a gap opens between WEIRD and non-WEIRD countries that survey fine-tuning then deepens while pursuing higher point accuracy. The most accurate model (Tulu~3 70B-DPO fine-tuned on WVS, 57.9\\%) keeps half the human spread overall ($\\dr = 0.50$) and 11\\% of it for Nigeria, against 0.70--0.87 for WEIRD countries. Raising the sampling temperature to 1.0 leaves the Wasserstein-1 distance ($\\wone$) to human distributions unchanged for both fine-tuned DPO models, and GRPO on Qwen~3.5 9B does not restore the spread under either an accuracy reward or a distribution-shaped reward. Mixing the aligned model with an unaligned prior raises $\\dr$ from 0.51 to 0.62 on a held-out split but leaves Nigeria at 0.36. Point accuracy alone therefore misjudges these simulators, and current post-training trades diversity for consensus.","authors":["Rojin Ziaei"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25760","pdf_url":"https://arxiv.org/pdf/2609.25760","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","算法保真度","跨文化调查"],"reason":"直接评估LLM仿真人类调查回答的分布保真度，使用WVS真实数据对照，并揭示后训…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":3,"question":"LLM后训练与调查微调如何影响其模拟人类调查回答时的跨文化方差保留？","design":"使用WVS第7波12国10,000个受访者-问题对，构建包含人口统计和Inglehart-Welzel文化维度的价值编码persona，评估11个零样本LLM和5个在WVS数据上微调（SFT、DPO、GRPO）的变体，测量点准确率、MAE、Wasserstein-1距离和偏差比（预测标准差/人类标准差）。","baseline":"世界价值观调查（WVS）第7波12个国家的真实个体回答分布，包括标准差和分布形状。","findings":"后训练导致“共识坍缩”：监督指令微调使偏差比从1.22降至0.59，准确率仅提升0.9个百分点；后续DPO和GRPO未恢复方差，且WEIRD与非WEIRD国家差距扩大，最准确模型（Tulu 3 70B-DPO微调）整体偏差比0.50，尼日利亚仅0.11。提高采样温度至1.0不改变Wasserstein-1距离，GRPO在准确率或分布形状奖励下均不能恢复方差，混合未对齐先验仅将整体偏差比从0.51提升至0.62，尼日利亚仍为0.36。","reliability":"论文承认共识坍缩在非WEIRD国家更严重，且无法通过提高温度或GRPO恢复；混合未对齐先验只能部分缓解，不能根本解决。未讨论其他局限。","relevance":"直接针对LLM仿真人类调查回答的分布保真度，使用真实WVS数据对照，揭示后训练和微调对多样性的系统性压缩，对关注仿真可靠性与偏差的研究者极具参考价值。","inspiration":"借鉴其诊断框架：同时测量点准确率和分布离散度（偏差比、Wasserstein距离），并沿后训练轨迹分解方差损失来源。｜可迁移到经济政策评估中的异质性反应仿真，如不同文化背景下对税收或福利政策的态度分布。｜用LLM模拟多国受访者对政策的态度，以WVS或类似跨国调查为真实基准，比较零样本、指令微调和偏好优化模型在点准确率与方差保留上的权衡，并检验温度调节和先验混合能否恢复分布。"}},{"id":"2609.25059","version":1,"title":"Can Large Language Model-Generated Responses Support Assessment Development? A Human-Calibrated Rasch Benchmark","zh_title":"大语言模型生成的回答能否支持评估开发？一项人类校准的Rasch基准研究","abstract":"Large language models (LLMs) are proposed as synthetic respondents for pilot testing, but their usefulness depends on whether they supply the evidence assessment development requires. We calibrated rating scale models on 14 digital-use skill items from 6,245 adults and used the human item parameters to evaluate responses generated for 1,300 demographically matched personas. LLM responses had high internal consistency ($\\alpha \\approx .94$) but used the lowest category in 0.2-0.3% of responses versus 13.7-21.5% for humans, and no persona chose it on every item. These gaps changed the response-category judgment on both subscales and the PC targeting judgment; on the human-calibrated scales, 12 of 14 items had infit below 0.70, indicating responses more predictable than the Rasch model expects. Preregistered changes to the prompt, category order, and persona information did not restore the human lower range. High internal consistency is insufficient evidence that LLM responses can replace human pilot data.","authors":["Eunjeong Song","Sehee Hong"],"categories":["stat.AP","cs.CY"],"primary_category":"stat.AP","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25059","pdf_url":"https://arxiv.org/pdf/2609.25059","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真被试","Rasch模型","人类数据对照"],"reason":"用LLM生成合成被试回答，并与真实人类数据校准，评估其替代可行性，发现偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":6,"question":"LLM生成的回答能否支持评估工具开发中关于反应类别结构、目标定位和项目审查的判断？","design":"使用LLM为1300个人口统计学匹配的虚拟人物生成对14个数字技能自评项目的回答，并改变提示指令、类别顺序和人物信息以检验对低类别使用的影响。","baseline":"来自2025年数字鸿沟调查的6245名成年人的人类回答，用于校准Rasch模型并作为固定参照。","findings":"LLM回答内部一致性高（α≈.94），但最低类别使用率远低于人类（0.2-0.3% vs 13.7-21.5%），且没有虚拟人物在所有项目上选择最低类别。在人类校准的量尺上，14个项目中有12个的infit低于0.70，表明LLM回答比Rasch模型预期的更可预测，且提示、类别顺序和人物信息的改变未能恢复人类低端分布。","reliability":"论文指出高内部一致性不足以证明LLM回答可替代人类试点数据；人口统计学匹配不能保证再现测量构念的变异；研究为回顾性基准，不适用于儿童或青少年。","relevance":"该研究直接针对LLM作为合成被试的可靠性，用真实人类数据校准并发现系统性偏差，对关注仿真失效条件的研究者具有重要参考价值。","inspiration":"借鉴其用人类校准的Rasch模型作为固定参照来评估LLM回答偏差的方法，可迁移到经济金融中的主观量表或调查数据仿真（如消费者信心、风险态度）。｜可应用于政策评估中的问卷预测试，例如用LLM模拟不同人口群体对政策的态度分布。｜设计：以真实家庭调查（如美国消费者财务调查）为基准，用LLM生成匹配人口特征的虚拟受访者对风险偏好或通胀预期问题的回答，比较分布差异和项目反应模型拟合，检验提示工程能否校正偏差。"}},{"id":"2609.24012","version":2,"title":"Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure","zh_title":"检验而非假定充分性：针对涌现网络结构校准生成式社会模拟器","abstract":"Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check. We present an adequacy-aware calibration protocol that couples amortized posterior estimation with a synthetic identifiability assessment, a matched-sample-size adequacy check (prior-predictive reachability plus per-statistic posterior-predictive localization), a diagnosis-guided repair, and a statistic-held-out audit. We demonstrate it on a real second-hand luxury resale market with four channel-by-residency cells, each a bipartite buyer-brand network, using a forward model built from persona profiles elicited once, offline, by a language model. The behavioural parameters are recoverable in all four cells, though calibration is approximate and overconfident for one parameter. The observed summary falls outside the simulator's reachability reference in every cell, with the mean purchased tier as the pervasive discrepancy. The repair meets the value-block criterion in two of four cells but does not restore adequacy, and the held-out audit surfaces a buyer-breadth-dispersion miss no earlier diagnostic detected. A profile-source ablation finds the language-model profiles beat a flat rule baseline in all four cells, yet within-category brand relabelling causes no consistent degradation, so the profiles are a partially validated input whose value rests on structure, not brand identity. Making no causal claim, we conclude that an independent-aggregation account, without agent interaction or a buyer-breadth mechanism, cannot jointly reproduce the market's purchased-tier level, head-brand concentration, community structure and buyer-breadth heterogeneity.","authors":["Tengfei Shao","Chao Li","Xu Wang","Masayuki Goto"],"categories":["cs.AI","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-23","first_seen":"2026-09-22","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.24012","pdf_url":"https://arxiv.org/pdf/2609.24012","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","网络校准","市场模拟"],"reason":"用LLM生成persona模拟市场网络，并与真实数据校准，评估仿真充分性，属核…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":7,"question":"如何对生成式社会模拟器进行校准与充分性检验，使其能可靠地复现真实市场的涌现网络结构？","design":"使用基于LLM一次性离线生成的人物画像构建前向模型，模拟二手奢侈品转售市场中买家与品牌的二分网络；通过摊销后验估计校准行为参数，并进行合成可识别性评估、匹配样本量的充分性检查、诊断引导修复和留出统计量审计。","baseline":"真实二手奢侈品转售市场的交易数据，按渠道和居住地划分为四个单元，每个单元为一个买家-品牌二分网络。","findings":"行为参数在所有四个单元中可恢复，但校准近似且对一个参数过度自信；所有单元中观测摘要均超出模拟器的可达性参考，平均购买层级是普遍差异，修复未能恢复充分性，留出审计发现买家广度离散度的遗漏。","reliability":"论文承认校准近似且对一个参数过度自信，修复未能恢复充分性，LLM画像的价值在于结构而非品牌身份，且不声称模拟器再现了市场。","relevance":"该研究直接针对LLM社会模拟的校准与充分性检验，提供了与真实数据对照的严格方法，对关注模拟可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其将模拟器校准与充分性检验结合的方法，通过留出统计量审计避免循环验证，并利用合成可识别性评估参数可恢复性。｜可迁移到消费者市场细分与品牌选择模拟，例如用LLM生成消费者画像模拟电商平台上的购买行为，并与真实交易数据对照。｜以LLM生成消费者画像作为被试，施加不同营销策略（如折扣、推荐）作为处理，结果变量为购买品牌网络结构，用真实电商交易数据作为对照，评估模拟器能否复现品牌集中度与社区结构。"}},{"id":"2609.25586","version":1,"title":"Deflecting the Value Compass: Interacting with Large Language Models Temporarily Shifts Human Value Priorities Toward Personal Focus","zh_title":"偏转价值罗盘：与大语言模型互动暂时将人类价值优先转向个人关注","abstract":"Large language models increasingly support decisions where values are in tension, yet little is known about whether interacting with them changes which values users prioritize. In a preregistered study, 200 U.S. adults interacted with ChatGPT, Claude, or Gemini as a thinking partner or read fixed AI-generated considerations. The prompt asked LLMs to support reasoning without recommending a decision and named no values. Participants advised people facing real dilemmas and completed parallel PVQ-RR forms before, immediately after, and one task later. Each LLM condition temporarily shifted value priorities toward personal focus relative to the control (d=0.37-0.51), primarily through increased Self-Enhancement. Participants' advice retained words and meaning from their exchanges. Thus, a brief LLM interaction that neither targets values nor seeks to persuade can reorient values active during judgment without detectable convergence in value directions or advice.","authors":["Hasibur Rahman","Malak Sadek","Smit Desai"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25586","pdf_url":"https://arxiv.org/pdf/2609.25586","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM影响人类","价值观转变","人机交互实验"],"reason":"研究LLM交互对人类价值观的影响，有真实人类对照，揭示仿真偏差条件","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":8,"question":"与LLM进行简短的思考伙伴式互动是否会暂时改变人们在判断中激活的价值优先级，以及这种改变是否持续、是否导致价值或建议的趋同？","design":"本研究不是用LLM模拟人类被试，而是以200名美国成年人为真实被试，随机分配到ChatGPT、Claude或Gemini三种LLM交互条件，或一个阅读固定AI生成考虑的非交互控制条件。在实验第二阶段，被试与LLM进行最多10分钟的思考伙伴式互动（提示词不提及任何价值观、不推荐决策），然后为真实困境提供建议并完成Schwartz的PVQ-RR价值观量表；在第三阶段，被试在没有LLM的情况下处理另一个困境，以测量效应的持续性。","baseline":"有真实人类对照：控制组被试阅读由三个LLM合成的固定AI生成考虑，但不与LLM交互；所有被试在互动前、互动后立即和后续任务中完成平行版本的PVQ-RR量表，作为价值观变化的基准。","findings":"与LLM互动后，被试的价值优先级相对于控制组暂时向个人关注方向偏移（效应量d=0.37-0.51），主要通过自我增强价值的增加实现；这种偏移在后续任务中消失，且未检测到价值方向或建议的趋同。","reliability":"论文承认效应是暂时的，在后续任务中消失；未讨论其他失效条件或局限。","relevance":"该研究直接探讨LLM交互对人类价值观的因果影响，属于批判性仿真研究，揭示了在无明确说服意图下LLM仍能暂时改变价值优先级，对理解LLM在决策支持中的潜在偏差具有重要意义，值得精读原文。","inspiration":"值得借鉴的做法是采用随机对照实验设计，将LLM作为处理条件，设置非交互控制组，并使用标准化的价值观量表在多个时间点测量效应。｜可以迁移到经济金融中的消费者跨期选择或投资决策场景，例如LLM作为财务顾问是否会影响个人的时间偏好或风险态度。｜一个可行的设计是：招募真实投资者作为被试，随机分配到与LLM（如ChatGPT）进行投资讨论的处理组或阅读固定建议的控制组，在互动前后测量时间贴现率和风险偏好（如使用滴定法或量表），并与真实市场数据（如实际投资组合选择）进行对照，以检验LLM交互对经济决策的因果影响。"}},{"id":"2609.23640","version":1,"title":"Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment","zh_title":"人类对齐模型是人类的模型吗？偏好对齐中的图灵测试差距","abstract":"Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We distinguish alignment with human preferences from alignment with human behavior, and show that alignment with human preferences can make model behavior less human-like even when both preferences and responses come entirely from humans. We call this the Turing-test gap. We show that preference alignment preserves the human response distribution only under a restrictive condition, and find no consistent evidence that real human preferences satisfy it. Empirically, the loss of human-response likelihood increases with the strength of preference weighting, regardless of its direction, and the gap also appears under standard DPO. These results establish human-likeness as an explicit dimension of alignment rather than something assumed to follow from preference alignment.","authors":["Suqin Yuan","Runqi Lin","Muyang Li","Guanzhe Hong","Jindong Gu","Lei Feng","Chris Russell","Tongliang Liu"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23640","pdf_url":"https://arxiv.org/pdf/2609.23640","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["偏好对齐","人类仿真","算法保真度"],"reason":"研究偏好对齐导致模型行为偏离人类，评估仿真可靠性，具批判性。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":9,"question":"人类偏好对齐是否会使语言模型的行为偏离人类行为分布，从而产生图灵测试差距？","design":"本研究并非直接进行人类仿真实验，而是通过理论分析和受控实验检验偏好对齐对模型人类相似度的影响。在受控实验中，使用同一批人类回答分别进行等权拟合和偏好加权拟合，比较模型对未见过的人类回答的似然；在DPO实验中，从已拟合人类回答的检查点出发进行偏好优化，测量人类回答似然和生成文本的变化。","baseline":"对照的真实人类数据来自SHP和StackExchange回答池，其中包含人类撰写的回答和人类投票偏好；以及HH-RLHF、WebGPT等偏好数据集。","findings":"偏好对齐在理论上仅在严格条件下保持人类行为分布，而实证中未发现人类偏好满足该条件。偏好加权强度越大，模型对人类回答的似然损失越大，且该损失与加权方向无关；标准DPO同样导致模型偏离人类行为。","reliability":"论文指出，提高人类相似度并非普遍可取，可能引发冒充、操纵和社会工程等风险；同时，通过角色提示和偏好人类回答的DPO等方法对人类相似度的恢复有限且依赖领域。","relevance":"该研究直接揭示了用偏好对齐的LLM作为人类被试替代品时可能存在的系统性偏差，为评估仿真可靠性提供了关键理论依据和实证证据，值得精读。","inspiration":"借鉴其将偏好对齐与行为分布分离的框架，在仿真实验中明确区分“人类偏好”与“人类行为”，并测量两者差异。｜可迁移到经济金融中的政策偏好调查或消费者决策仿真，例如用LLM模拟公众对通胀预期的回答，但模型可能给出“理想”而非“实际”的回答。｜设计：以LLM为被试，施加偏好对齐强度不同的处理，结果变量为对经济问题回答的分布与真实调查数据（如密歇根消费者调查）的KL散度，对照真实人类回答。"}},{"id":"2609.25572","version":1,"title":"A Behavioral Trait Leaks into Preferences: Diagnosing Trait Interference in LLM User Simulators","zh_title":"行为特质泄漏到偏好中：诊断LLM用户模拟器中的特质干扰","abstract":"LLM-based user simulators aim to bridge the offline-online gap in recommender evaluation by emulating users through injected traits, where preference attributes determine what a user engages with and a behavioral activity trait governs how long they browse. However, we show this intended trait independence collapses during simulation, causing two failures: (i) Trait Interference, where amplified activity distorts preference boundaries and forces interactions with mismatched items to sustain browsing, and (ii) Evaluation Invalidity, where satisfaction scores inflate with activity-driven page counts despite taste mismatches, biasing evaluation toward trait distributions rather than recommender performance. To resolve this, we propose PQA, a page-level quality anchoring method that guides simulators using a personalized anchor reflecting each user's intrinsic preference standard. By assessing whether a page meets this standard before further browsing, PQA enables proactive exits from low-quality pages, letting the activity trait retain its intended role of modulating browsing depth within preference-conforming pages. Experiments show PQA mitigates trait interference and improves the reliability of LLM-based simulator evaluation under activity shifts. Our code is available at https://github.com/chaehyun1/PQA","authors":["Chaehyun Kim","Sein Kim","Hongseok Kang","Chanyoung Park"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25572","pdf_url":"https://arxiv.org/pdf/2609.25572","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM用户模拟","推荐系统评估","仿真偏差"],"reason":"用LLM模拟用户行为并诊断仿真失效，有真实数据对照，方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":11,"question":"LLM用户模拟器中，行为活动特质是否干扰了偏好特质，导致模拟失效？","design":"使用Agent4Rec和SimUSER两个LLM用户模拟器，在MovieLens和Amazon CDs数据集上，以SASRec为推荐模型生成20项推荐列表，按4项一页呈现。通过将低活动用户的活动特质修改为高活动状态，保持偏好特质不变，测量模拟器选择项目的语义对齐度、类别偏好一致性、浏览页数和满意度分数。","baseline":"真实用户数据：MovieLens和Amazon CDs数据集中用户活动水平与平均评分的负相关关系，即更活跃用户平均评分更低。","findings":"活动特质放大导致模拟器选择与用户偏好语义距离更远的项目，并偏离固有类别偏好，产生特质干扰；同时满意度分数随活动驱动的浏览页数增加而膨胀，即使推荐质量下降也持续浏览低质量页面，导致评估无效。","reliability":"论文指出当前模拟器缺乏页面级质量锚点，导致活动约束覆盖页面级偏好对齐；PQA方法依赖用户历史类别消费密度，可能受历史数据稀疏性影响，但未详细讨论其他失效条件。","relevance":"该研究直接诊断LLM模拟器中行为特质与偏好特质的干扰问题，并提供了真实人类数据对照，对关注仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文了解PQA方法细节。","inspiration":"借鉴其通过修改单一特质并保持其他特质不变来分离干扰效应的实验设计，可用于经济金融仿真中识别行为参数对决策的因果影响｜可迁移到消费者金融决策仿真，如信用卡使用或投资组合选择，检验风险偏好特质是否被交易频率等行为特质干扰｜设计：用LLM模拟消费者，将风险偏好设为固定，改变交易频率特质，测量投资组合风险水平和满意度，对照真实交易数据中频率与风险偏好的关系。"}},{"id":"2609.26403","version":1,"title":"AI-Generated Email Drafts Shift Culturally Distinctive Communication Styles in Professional Email","zh_title":"AI生成的邮件草稿改变职场邮件中文化特有的沟通风格","abstract":"AI assistants that support email composition may shift cultural communication norms, such as the directness typical of low-context cultures like the US versus the indirectness and contextual sensitivity central to high-context cultures like Japan. Yet it remains unknown to what extent people adopt and edit AI drafts inconsistent with their cultural communication norms. We address this through a preregistered within-subject experiment in which Japanese and American participants wrote workplace emails in their native language without AI, with a low-context AI, and with a high-context AI. We found that Japanese participants wrote emails with significantly more high-context markers (politeness, apologies) than Americans. But AI drafts shifted participants' emails toward the draft's style, with larger shifts when the draft was culturally misaligned: Japanese drifted most under low-context drafts, Americans most under high-context drafts. These findings suggest AI drafts risk overwriting cultural communication norms unless they adapt to users' communication styles.","authors":["Shintaro Sakai","Alice Gao","Yuichi Shoda","Katharina Reinecke"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26403","pdf_url":"https://arxiv.org/pdf/2609.26403","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","文化沟通","人机交互实验"],"reason":"用LLM生成邮件草稿并测量其对人类写作风格的影响，有真实人类实验对照，且揭示A…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":11,"question":"使用AI生成的邮件草稿是否会将用户的邮件写作风格转向草稿的文化语境风格，且这种转变在AI与用户文化背景不一致时是否更大？","design":"本研究不是用LLM模拟人类被试，而是用LLM生成邮件草稿作为实验干预。175名日本和美国全职员工在三种条件下写工作邮件：无AI辅助、低语境AI草稿、高语境AI草稿（被试内设计，条件顺序平衡）。结果变量是邮件中高语境标记（如礼貌用语、道歉）的数量。","baseline":"无AI辅助条件下参与者写的邮件作为人类基线，用于比较日本和美国参与者的文化沟通风格差异。","findings":"日本参与者在无AI条件下写的邮件比美国人包含更多高语境标记。AI草稿使参与者的邮件风格向草稿的文化语境偏移，且当草稿与参与者文化背景不一致时偏移更大：日本人在低语境草稿下偏移最大，美国人在高语境草稿下偏移最大。","reliability":"论文未讨论","relevance":"该研究用真实人类实验检验了LLM生成内容对人类行为的影响，有严格的人类对照，揭示了AI在文化维度上的潜在偏差和影响，对关注LLM仿真可靠性及偏差的研究者有参考价值。","inspiration":"借鉴其被试内设计和多条件对照，通过测量行为变化来评估AI干预的因果效应。｜可迁移到经济金融中的沟通与决策场景，如AI辅助撰写投资建议、信贷沟通或政策沟通，研究AI风格对用户决策或行为的影响。｜设计一个实验：招募金融从业者或普通投资者作为被试，让他们在无AI、低风险偏好AI草稿、高风险偏好AI草稿三种条件下撰写投资建议邮件，测量建议的风险水平，并与真实历史投资建议数据对照，检验AI是否改变建议风格及在风格不一致时影响是否更大。"}},{"id":"2609.15727","version":2,"title":"Are LLMs Good Financial User Simulators? Multi-view Investor Logic Alignment (MILA)","zh_title":"大语言模型是好的金融用户模拟器吗？多视角投资者逻辑对齐（MILA）","abstract":"Large language models (LLMs) are increasingly used as user simulators, yet it remains unclear whether their predictions faithfully reproduce the evolving decisions of individual users. We investigate this question in a controlled longitudinal paper-trading study with 80 participants, where user interactions, simulated transactions, virtual portfolio states, and point-in-time market information are aligned under a rolling next-day prediction protocol. We evaluate behavioral fidelity hierarchically, from trade occurrence to action structure, asset selection, and downstream portfolio consequences. Across 1,239 aligned user-days, no evaluated LLM reliably outperforms a simple recent-activity persistence baseline for predicting whether a user trades. Fidelity further deteriorates at finer levels: models struggle to recover buy--sell structure and traded assets, and similar activity-level predictions can lead to substantially different portfolio trajectories. Controlled evidence ablations show that recent trading history strongly governs activity prediction, whereas asset selection is substantially more sensitive to the available evidence. An observational analysis further finds that intensified ticker-specific research predicts imminent trading, but diagnostic tests do not support a causal interpretation. These findings suggest that current LLMs capture useful short-term behavioral regularities without yet recovering a stable individual decision mechanism.","authors":["Jiajie He","Jiangyuan Hong","Xintong Chen","Dongling Ni","Wenjin Liu"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2026-09-15","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.15727","pdf_url":"https://arxiv.org/pdf/2609.15727","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","金融行为","算法保真度"],"reason":"用LLM模拟投资者决策，与80名真实用户对照，评估行为保真度并指出失效条件","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":2,"question":"LLM 能否忠实模拟个体投资者在纵向交易环境中的决策行为？","design":"在80名参与者的受控纵向模拟交易研究中，用多种LLM作为用户模拟器，基于滚动次日预测协议，输入截至预测时点的用户交互、交易记录、虚拟持仓和市场信息，预测用户次日是否交易、买卖结构、交易资产及后续组合轨迹，并与真实用户行为逐日对齐比较。","baseline":"80名真实参与者在同一模拟交易平台上的1,239个用户-日对齐行为记录。","findings":"在预测用户是否交易上，所有评估的LLM均未能稳定超越简单的近期活动持续性基线；在更细粒度上，模型难以恢复买卖结构和交易资产，且相似的活动级预测可能导致显著不同的组合轨迹。","reliability":"论文指出当前LLM能捕捉短期行为规律，但尚未恢复稳定的个体决策机制；资产选择对可用证据更敏感，且观察性分析中强化个股研究虽与交易相关，但诊断测试不支持因果解释。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，提供了与真实个体行为逐日对照的严格评估，并揭示了仿真在细粒度决策上的失效条件，对关注仿真保真度和偏差的研究者具有重要参考价值。","inspiration":"值得借鉴的是其分层行为保真度评估框架和受控证据消融方法，可系统区分预测依赖与因果机制｜可迁移到资产定价实验中的投资者异质性决策模拟，或政策公告下的预期形成与交易行为研究｜可设计一个实验：招募真实投资者在模拟平台上交易，用LLM基于其历史行为和市场信息预测次日交易决策，以真实交易记录为对照，通过消融不同信息源检验模型依赖，并比较组合轨迹差异。"}},{"id":"2609.22090","version":1,"title":"Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents","zh_title":"识别、仿真与拒绝：LLM智能体中经典心理效应的污染意识研究","abstract":"An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM agents under a factorial design built to separate these: each paradigm is run with the paradigm explicitly labeled in the prompt (named) or framed as a routine task (blind), and on the literal textbook version of the task (canonical) or a structurally matched variant written to reduce lexical and scenario overlap with likely training data (counterfactual), crossed with a persona manipulation. Across five completed paradigms, evaluated on up to three open-weight model families with 41,904 trials released, apparently human-like effects arise through qualitatively different routes rather than one susceptibility: paradigm-label gating with explicit override (Asch conformity, 0 percent blind to 83.3 percent named on gpt-oss-120B), knowledge-dependent signal reliance (anchoring, exactly zero on grounded facts versus near total on invented quantities, a pattern equally consistent with rational use of the only available signal), amplification on novel content under labeling (framing), robust absence (sunk cost), and safety-mediated selection where refusal itself is the primary finding (minimal-group allocation). A one-sentence persona change (agreeableness, framed as an instruction rather than a verified trait manipulation) eliminates, dampens, or reverses these effects depending on which effect it is, arguing against any single response-bias account. We further formalize, and in two cases document empirically, three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection. We argue scalar bias-susceptibility scores obscure this structure and report replication profiles instead.","authors":["Joy Bose"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22090","pdf_url":"https://arxiv.org/pdf/2609.22090","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM仿真","心理学实验","算法保真度"],"reason":"用LLM复现经典心理学实验，与人类数据对照，并批判性分析仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":5,"question":"LLM在经典心理学实验中表现出的类人效应，究竟是对人类偏差的真实模拟，还是对实验范式的识别、记忆污染或安全过滤所致？","design":"使用多个开源LLM模型（如gpt-oss-120B等）作为被试，在五个经典心理学范式中进行实验。每个范式采用2×2×3因子设计：标签轴（命名vs盲测）、领域轴（经典版本vs反事实版本）、人设轴（无、高宜人性、低宜人性）。结果变量为模型在任务中的行为反应（如从众、锚定、框架效应、沉没成本、最小群体分配）。","baseline":"经典心理学实验的人类行为数据（如Asch从众实验约1/3从众率、锚定效应、框架效应等）作为参照。","findings":"五个范式表现出质的不同路径：Asch从众在盲测下几乎消失（0%），命名后大幅出现（83.3%），表明标签门控；锚定效应在基于事实的锚上为零，在虚构锚上接近完全，符合理性使用唯一信号；框架效应在反事实内容上被标签放大；沉没成本效应完全缺失；最小群体分配因安全拒绝而无法测量。一句宜人性人设指令即可消除、减弱或反转这些效应，表明不存在单一反应偏差。","reliability":"论文承认三个失效模式：人设主导（persona dominance）、群体崩溃（population collapse）、安全选择（safety selection）。人设操纵仅为一句指令，非验证性特质诱导，且与结果变量存在词汇重叠，因此只能视为指令效应而非特质模拟。标签门控可能源于范式名称的语义泄漏，而非模型真正识别实验。","relevance":"该研究直接回应了LLM仿真人类行为时的污染与识别问题，通过严格的因子设计分离了记忆、标签和指令效应，对评估LLM作为人类被试替代品的可靠性具有重要参考价值，值得精读原文。","inspiration":"借鉴其反事实任务设计和标签/盲测对照，以区分模型是基于经济知识还是真实决策偏差。｜可迁移到资产定价实验中的锚定效应、信贷审批中的歧视、消费者跨期选择中的框架效应等场景。｜用LLM模拟投资者，处理为是否告知实验目的（标签vs盲测）和锚定值来源（真实历史数据vs虚构数据），结果变量为估值或投资决策，对照真实人类实验数据（如实验室资产泡沫实验）。"}},{"id":"2609.22169","version":1,"title":"Monocultural Biases: Correlated biases in large language models lead to unequal systemic exclusion rates in hiring","zh_title":"单一文化偏见：大语言模型中的相关偏见导致招聘中的系统性排斥率不平等","abstract":"Employers are increasingly using large language models (LLMs) to automate their hiring process. This paper investigates the risk of monocultural biases, in which the widespread deployment of large language models homogenizes biases across the labor market, leading to greater systemic exclusion for certain demographic groups. For ten LLMs, we measure hiring biases across their base and post-trained versions to identify which stage, pre-training or post-training, lead to monocultural biases. We find that, compared to their base models, post-trained models are 3.6% less likely to callback older applicants. This negative shift occurs in eight of the ten models that we evaluate. Post-trained models have much more correlated decisions than base models which is likely driven by human capital traits like skills or college major. However, greater consensus among models increases global systemic exclusion rates from 5.6% to 17.3% and exacerbates demographic inequalities, with intersectional systemic exclusion rates ranging from 12.2% to 21.7% for post-trained models. We find that this inequality is primarily driven by age-based discrimination that is exacerbated in post-training. These results indicate that while post-training techniques may improve models' abilities to select the best applicants, they may raise systemic inequality risks for those at the margin by uniformly introducing new biases.","authors":["Matthew Bone","Fabian Stephany","Maria del Rio-Chanona"],"categories":["cs.CL","cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22169","pdf_url":"https://arxiv.org/pdf/2609.22169","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","招聘偏见","算法公平"],"reason":"用LLM模拟招聘决策并与人类数据对照，评估系统性偏差，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":6,"question":"大语言模型在招聘中是否会产生单文化偏差，即模型间的决策高度相关，导致某些人口群体在劳动力市场中被系统性排斥？","design":"使用10个LLM（包括基础版和后训练版）模拟招聘初筛，基于Burning Glass Institute的在线劳动力数据生成76种职业的职位空缺和求职者档案，通过改变求职者的年龄、性别、种族等人口特征构造8种变体，测量模型对不同群体的回调率差异，并比较模型间决策的相关性和系统性排斥率。","baseline":"无直接的人类决策对照，但使用真实劳动力市场数据（Burning Glass Institute）生成职位和档案，确保情境真实。","findings":"后训练模型比基础模型更少回调年长求职者（低3.6%），且决策相关性更高，导致全局系统性排斥率从5.6%升至17.3%。这种排斥主要由年龄歧视驱动，后训练阶段（尤其是监督微调）加剧了年龄偏差，同时提高了模型基于人力资本特质筛选的能力。","reliability":"论文指出在线劳动力数据偏向白领、大学学历工作，限制了职业覆盖范围；且未与真实人类招聘决策直接对照，无法完全反映实际部署中的偏差。","relevance":"该研究直接使用LLM模拟招聘决策，并测量系统性偏差，与研究者关注的经济学实验和政策评估场景高度相关，值得精读以了解LLM仿真的偏差来源和测量方法。","inspiration":"借鉴其利用大规模真实数据生成仿真情境并系统操纵人口特征的方法，可迁移到信贷审批歧视研究，用LLM扮演信贷员，处理变量为申请人种族或性别，结果变量为贷款批准率，对照真实信贷数据中的批准率差异。｜可应用于劳动力市场政策评估，如最低工资对雇佣决策的影响，用LLM模拟雇主行为，处理为不同工资水平，测量雇佣概率，对照实际企业调查数据。｜设计一个实验：用LLM模拟投资者对政策公告的反应，处理为不同政策措辞，结果变量为投资决策，对照真实市场数据中的资产价格变动。"}},{"id":"2609.22607","version":1,"title":"Pretrained Persona Mixture Models and Tandem Models for Human Simulation","zh_title":"用于人类仿真的预训练人格混合模型与串联模型","abstract":"We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas, is inaccurate and produces stereotyped predictions (lacking natural diversity). It has previously been shown that LLMs can be bound to personas using naturalistic, freetext dialog avoiding stereotyping. Here we show that binding can also be achieved using short, individual samples of dialog from specific people. Demographics can be added later without negative effects by simply querying the model. We use the term Persona Mixture Models (PMMs) for well-calibrated human models, currently realized as pretrained base models. We show that PMMs produce more accurate predictions than instruction-tuned models and retain more of the lexical, semantic, and pragmatic diversity found in human dialog. We measure realism and diversity of LLMs simulating human interlocutors across a diverse set of corpora spanning open-domain text, human-AI chat, and task-oriented dialogue between human speakers. However, base pretrained models can produce out-of-domain dialog and may lose some of the human's internal state over long contexts. We propose and explore tandem models which combine a pre-trained model with an instruction-tuned supervisor. Tandem models achieve the best overall accuracy and diversity in our experiments.","authors":["Minwoo Kang","T\\'ea Wright","Seun Eisape","Ayush Raj","Suhong Moon","Joseph Suh","Alane Suhr","David M. Chan","John Canny"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22607","pdf_url":"https://arxiv.org/pdf/2609.22607","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","人格混合模型","对话多样性"],"reason":"直接研究LLM仿真人类对话，提出PMM和tandem模型提升准确性与多样性，并…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":9,"question":"如何用预训练语言模型作为人类对话的仿真器，并提升其准确性与多样性？","design":"用预训练基础模型（PMM）和指令微调模型（ALM）分别模拟人类对话者，通过少量个人对话样本绑定人格，比较其在开放域、人机对话和任务型对话中的预测准确性和语言多样性。","baseline":"五个英语语料库中真实人类对话者的语言分布、对话行为结构和词汇语义多样性。","findings":"预训练基础模型比指令微调模型更准确地预测人类下一句话，并保留更多词汇、语义和语用多样性。结合预训练模型与指令微调监督者的串联模型在准确性和多样性上总体最佳。","reliability":"论文指出基础模型可能产生域外对话，在长上下文中丢失人类内部状态；研究仅限英语语料，单一指标不足以全面衡量仿真保真度。","relevance":"该研究直接针对LLM人类仿真，提出PMM和串联模型改进仿真质量，并系统比较了不同模型在真实对话数据上的表现，对关注仿真可靠性与偏差的研究者很有参考价值。","inspiration":"借鉴其用少量个人样本绑定人格并比较不同模型架构的方法，可迁移到经济金融中的消费者决策仿真或投资者情绪模拟。｜例如在消费者跨期选择实验中，用LLM模拟不同人口特征的被试，比较预训练与指令微调模型的行为预测。｜设计：用真实消费者调查数据作为基准，以少量个人回答绑定LLM人格，处理为不同模型类型，结果变量为跨期选择一致性，对照真实人类选择分布。"}},{"id":"2609.24911","version":1,"title":"SocioVerse2: A Longitudinal Dynamic Social Simulation Framework under a Human-AI Co-evolutionary Paradigm","zh_title":"SocioVerse2：人机共演化范式下的纵向动态社会仿真框架","abstract":"Social simulation offers the social sciences an experimental instrument that the real world cannot supply, and generative agents have transformed it by acting as silicon samples that unite agent-based modeling with real behavioral data. Existing platforms verify collective behavior, align simulated populations with real societies in cross-sections, and employ autonomous agents for the research process. However, two social science requirements remain without systematic support: intervention in the content of a simulation and the researcher's control over the process that produces it. We present SocioVerse2, which extends SocioVerse 1.0 into a human-AI co-evolutionary paradigm built from two loops and one infrastructure. The longitudinal simulation loop simulates the target population with evolving environments and forks counterfactual branches via interventions. The controllable research loop takes the study itself as an editable state and updates state versions via controllable editing. The social science agentic infrastructure carries both loops through composable skills with researcher checkpoints, a population service over five persona pools, and an environment service over 21 real-world signal sources with point-in-time guarantees. We validate SocioVerse2 across three case families and seven case studies, from reproducing canonical agent-based models to modeling policy processes on real records and nowcasting macro-economic indices beyond the response model's knowledge cutoff. With the human-AI co-evolutionary paradigm, these cases go beyond system demonstrations to become substantive studies that investigate frontier questions in their respective disciplines. Code, data services, and a workbench are released as open-source resources.","authors":["Xinnong Zhang","Jiayu Lin","Jia Wang","Yixu Huang","Xinyi Mou","Yingqian Wu","Jingcong Liang","Shijun Lei","Jianing Shi","Guanying Li","Siyuan Wang","Hanjia Lyu","Zhenfei Yin","Yunlu Yin","Siming Chen","Yulan He","Jiebo Luo","Xuanjing Huang","Liyin Jin","Baohua Zhou","Hanqi Yan","Zhongyu Wei"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24911","pdf_url":"https://arxiv.org/pdf/2609.24911","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM社会仿真","人类数据对照","政策评估"],"reason":"用LLM agent模拟社会过程并与真实数据对照，支持干预和纵向演化，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":10,"question":"如何构建一个支持纵向干预、研究者可控、并基于真实世界数据的人类-AI协同演化社会仿真框架？","design":"SocioVerse2 用生成式智能体（硅样本）模拟目标人群，通过纵向仿真循环让环境和人群随时间演化，并支持干预操作产生反事实分支；可控研究循环将研究过程本身作为可编辑状态，允许研究者介入和调整；基础设施提供五个角色池和21个真实世界信号源，保证时间点一致性。","baseline":"多个案例研究使用真实记录和宏观指数作为对照，如政策过程建模使用真实记录，宏观指数预测使用真实世界确定性指数和信号作为基准。","findings":"SocioVerse2 能够复现经典基于智能体的模型，并在真实记录上建模政策过程，还能预测超出模型知识截止日期的宏观经济指数。通过人类-AI协同演化范式，案例研究超越了系统演示，成为各自学科前沿问题的实质性研究。","reliability":"论文未讨论","relevance":"该研究直接针对用LLM进行社会仿真并与真实数据对照，支持干预和纵向演化，对评估仿真可靠性和偏差有参考价值，值得阅读原文了解其框架和案例细节。","inspiration":"借鉴其纵向干预和反事实分支设计，可在经济政策评估中设置处理组和对照组，观察政策效果的动态演化。｜可迁移到政策公告的预期形成研究，模拟不同政策沟通策略对市场参与者预期的影响。｜用LLM智能体模拟投资者群体，处理为不同政策公告措辞，结果变量为预期通胀或资产价格变动，对照真实市场调查数据或高频交易数据。"}},{"id":"2609.22252","version":1,"title":"CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation","zh_title":"CALM：用于基于活动的出行者仿真的校准LLM选择网络框架","abstract":"We present CALM, a reproducible hybrid framework that integrates an optional large language model (LLM) activity planner with calibrated stochastic choice, shared network feedback, memory and habit, typed feasibility checks, and deterministic offline replay. Unlike trip-mode classifiers or diary-only generators, CALM executes a closed traveler-day loop and evaluates each generative module against an empirical, reproducible baseline. On the 2024 New York City Citywide Mobility Survey (CMS), 110,691 seven-mode trips are split by respondent into 78,487 training and 32,204 holdout trips. Training-only alternative-specific constant calibration reduces mean holdout mode Jensen-Shannon divergence from 0.15599 to 0.00394 across ten seeds. A matched live-LLM ablation then quantifies trade-offs among aggregate fit, temporal fit, behavioral persistence, and feasibility, while frozen prompt-response pairs support deterministic replay of downstream simulation. Controlled weather, delay, fare, and parking ladders further demonstrate consistent and interpretable responses under intervention. CALM contributes a reproducible protocol for integrating and evaluating generative planners in traveler simulation through person-disjoint calibration, matched module ablation, controlled stress testing, and end-to-end traceability.","authors":["Yezhou Cheng"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22252","pdf_url":"https://arxiv.org/pdf/2609.22252","source_feed":"cs.LG","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","出行行为","人类数据校准"],"reason":"用LLM模拟出行者选择，并与真实调查数据对照校准，属于人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":7,"question":"如何构建一个可校准、可复现的混合框架，将大语言模型活动规划器与随机选择模型结合，用于基于活动的出行者仿真，并与真实调查数据对照评估。","design":"CALM框架用LLM（gpt-4o-mini）作为可选活动规划器，结合校准的随机选择模型、共享网络反馈、记忆与习惯、类型化可行性检查，模拟纽约市出行者的一日活动。通过匹配的消融实验（校准选择、无记忆LLM、有记忆LLM、完整混合）比较不同模块组合，施加天气、延误、票价、停车等干预阶梯，测量模式份额的Jensen-Shannon散度、时间分布、行为持续性、可行性等结果。","baseline":"2024年纽约市全市出行调查（NYC CMS）的110,691次七模式出行记录，按受访者划分为78,487次训练和32,204次留出，用于校准和评估。","findings":"仅用训练数据校准替代特定常数（ASC）后，留出集模式份额的Jensen-Shannon散度从0.15599降至0.00394，且跨10个种子稳定。匹配的LLM消融实验量化了聚合拟合、时间拟合、行为持续性和可行性之间的权衡，干预阶梯显示一致且可解释的响应。","reliability":"论文强调面效度不等于经验效度，LLM生成的出行叙事可能仍与空间、时间和OD分布不匹配；当前13区网络实例化仅提供受控环境，未在大规模真实网络上验证。","relevance":"该研究直接针对LLM仿真人类行为并与真实调查数据对照，提供了严格的校准和评估协议，对关注经济学实验和政策评估中LLM仿真可靠性的研究者具有重要参考价值。","inspiration":"值得借鉴的是其人员不相交的校准、匹配模块消融和受控压力测试，确保仿真差异归因于特定模块而非需求背景。｜可迁移到政策评估中的个体选择行为仿真，如交通定价、补贴或信息干预对出行方式选择的影响。｜以真实居民出行调查数据为基准，用LLM生成个体活动计划，施加票价或拥堵收费等处理，测量方式选择概率和福利变化，并与实际政策试点数据对照。"}},{"id":"2609.22408","version":1,"title":"Social Influence and the Allocation of Scientific Attention in AI Populations","zh_title":"AI群体中的社会影响与科学注意力分配","abstract":"AI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human readers, but the collective consequences of these signals for artificial readers remain uncertain. This paper adapts the Music Lab design to a market for academic attention. In the first experiment, 1,000 AI agents choose papers from the titles and abstracts of all 114 regular research articles published in the American Economic Review in 2025. The experiment has five independent-choice communities and five social-influence communities, each with 100 sequential agents. Only agents in the social-influence condition observe earlier selections within their community. Agents may select any number of papers. Social-information communities select 17.2 percent fewer papers per agent, concentrate their choices more heavily, and collectively cover 73 papers, compared with 90 independently. Between-community variation is greater under social information. In a second experiment with 200 agents across twenty social communities, randomly assigning papers five initial selections raises their subsequent selection rate by 45.55 percentage points (95% CI: 41.20 to 49.90). Choices have modest correspondence with external citations and little correspondence with download counts. The results show how a simple information rule shapes the volume, breadth and distribution of scientific attention in an artificial population.","authors":["Maxim Chupilkin"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22408","pdf_url":"https://arxiv.org/pdf/2609.22408","source_feed":"econ.GN","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","社会影响","学术注意力"],"reason":"用1000个AI代理模拟学术注意力分配，与真实引用数据对照，属经济学实验场景。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":8,"question":"社会信息如何影响AI代理对学术论文的注意力分配？","design":"用GPT-5.6 Sol扮演1000个AI代理，分为独立选择组和社会影响组（各5个社区，每社区100个顺序代理），从2025年AER的114篇文章标题和摘要中选择要读的论文；社会影响组能看到本社区之前代理的累计选择次数，独立组看不到；测量选择数量、集中度、社区间差异等。","baseline":"外部引用次数和下载量作为对照，但仅用于相关性分析，并非严格的人类被试基准。","findings":"社会信息使代理平均少选17.2%的论文，集体覆盖从90篇降至73篇，选择更集中且社区间差异更大；随机赋予初始流行度使后续选择率提高45.55个百分点。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类在学术注意力市场中的从众行为，与真实引用数据对照，属于经济学实验场景，对关注LLM仿真可靠性和偏差的研究者有直接参考价值。","inspiration":"借鉴其Music Lab式设计，通过独立组与社会影响组对比、随机初始流行度干预来分离社会影响效应｜可迁移到金融信息传播或资产定价中的注意力分配问题，如投资者对研报或新闻的关注｜用LLM代理扮演投资者，处理为是否展示其他代理的阅读/选择行为，结果变量为选读的研报数量与集中度，对照真实市场中的研报点击或交易数据。"}},{"id":"2609.22225","version":1,"title":"Do LLMs Choose Like Humans? Using Cognitive Theory to Evaluate LLM Decision-Making","zh_title":"LLM像人类一样选择吗？用认知理论评估LLM决策","abstract":"Large language models (LLMs) exhibit a range of human-like decision-making behaviors, but whether these reflect similar underlying mechanisms or surface-level mimicry remains unclear. We evaluate whether LLM context sensitivity aligns with a cognitive economic theory that explains human behavior through problem categorization and attention allocation. Across 12 open-source and commercial LLMs on a novel 140,000-trial product choice benchmark, context induces human-like shifts in choice and problem categorization, but does not reliably reweight attention between features like price and quality. Neither scale nor chain-of-thought reasoning reliably attenuates context sensitivity or generates human-like behavior. These results suggest that LLM decision mechanisms are distinct from human ones.","authors":["Johnathan Sun","Andrei Shleifer","Yonatan Belinkov"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22225","pdf_url":"https://arxiv.org/pdf/2609.22225","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM决策仿真","认知理论","人类对照"],"reason":"用LLM模拟人类决策并与人类数据对照，评估机制差异，直接相关且具批判性。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":12,"question":"LLM 的决策行为是否与人类认知经济理论中的情境敏感机制一致，即是否通过问题分类和注意力重加权来产生情境效应。","design":"用 12 个开源和商业 LLM 作为被试，在 140,000 次产品选择试验中，通过情境线索（享乐型 vs. 功能型消费）操纵问题分类，测量选择概率和特征敏感性（价格与质量的注意力权重）。","baseline":"人类基准来自认知经济学理论（Bordalo et al., 2026b）所解释的人类决策现象，包括情境对选择的影响、特征敏感性的不对称变化以及分类难度对情境效应的调节。","findings":"LLM 的选择变化与人类一致，但特征敏感性变化大多不一致；情境效应遵循问题分类过程，但模型规模和思维链推理不能可靠减弱情境敏感性。","reliability":"论文指出 LLM 的决策机制与人类不同，情境效应可能只是表面模仿；规模和推理不能可靠产生人类行为，说明在需要机制对齐的任务中仿真可能失效。","relevance":"该研究直接评估 LLM 作为人类决策仿真体的机制一致性，提供了与人类认知理论对照的批判性证据，值得精读以了解仿真在经济学决策中的局限。","inspiration":"借鉴其通过情境线索操纵问题分类并测量特征敏感性的设计，可迁移到消费者跨期选择或资产定价实验，用 LLM 模拟投资者在不同市场情境下的风险偏好，处理为情境线索（如牛市/熊市），结果变量为风险资产配置比例，对照真实投资者调查数据。"}},{"id":"2609.23403","version":1,"title":"Alignment and Divergence between Humans and AI in Interpersonal Privacy Decisions","zh_title":"人际隐私决策中人类与AI的一致性与分歧","abstract":"AI assistants increasingly mediate interpersonal communication on behalf of their primary user, but they risk violating the privacy expectations of third-party information owners. Resolving these tensions requires understanding how humans anticipate interpersonal privacy boundaries. Therefore, we conducted a dyadic study (N=76) and a matched evaluation of AI models across 18 information types and 3 recipient relationships. We found that data owners' privacy judgments are highly contextual and relationship dependent. While familiar data co-owners show meaningful alignment with owners' expectations, they significantly overestimate the need for permission. Interestingly, greater familiarity within the owner-co-owner dyad was associated with both higher disclosure acceptability and lower co-owner misalignment, whereas our exploratory four-item empathy measure was not. In contrast, AI models significantly underperform human co-owners in anticipating the data acceptability, even when provided with within-dyad examples. These findings underscore a core HCI design challenge to develop privacy-aware AI that respects multi-stakeholder information boundaries.","authors":["Hanxiang Zeng","Shuning Zhang","Xinyuan Zhou","Tianqi Song","Yuhan Yuan","Yuting Yang","Shuai Ma","Xin Yi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23403","pdf_url":"https://arxiv.org/pdf/2609.23403","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","隐私决策","人机对齐"],"reason":"用LLM模拟人类隐私决策并与人类数据对照，评估AI与人类判断的偏差，属于仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":13,"question":"在人际隐私决策中，数据所有者对第三方披露的接受度如何受情境影响，人类共同所有者与所有者的判断对齐程度如何，AI模型能否准确预测所有者的隐私判断？","design":"本研究并非严格意义上的LLM仿真人类实验，而是采用配对设计：招募19对朋友和19对恋人（N=76），让数据所有者和共同所有者分别对18种信息类型、3种接收者关系下的披露可接受性、许可必要性等隐私维度进行评分；同时用多个LLM（零样本和单样本提示）、机器学习模型和微调模型对相同场景进行预测，并与人类共同所有者的预测准确度比较。","baseline":"人类基准是熟悉的数据共同所有者（朋友或恋人）对所有者隐私判断的预测，以平均绝对误差（MAE）和判断落在所有者评分±1分内的比例衡量对齐程度。","findings":"所有者的隐私判断高度依赖情境和关系，信息类型和接收者关系显著影响披露可接受性；熟悉的人类共同所有者与所有者有中等程度对齐（MAE≈1.179，70.7%判断在±1分内），但系统性高估许可必要性。AI模型（包括零样本和单样本LLM）显著不如人类共同所有者准确（MAE 1.648–1.796），单样本提示仅略微改善，仍无法弥合人机差距。","reliability":"论文承认AI模型在涉及微妙社会关系的披露场景中尤其表现不佳，单样本提示不足以缩小差距；机器学习模型和微调模型对未见个体泛化能力差，提供所有者特定示例的改进不一致。局限性包括样本量较小（19对朋友和19对恋人）、仅覆盖两种关系类型、同理心测量简短且探索性，未发现其与对齐的显著关联。","relevance":"该研究直接评估LLM在人际隐私决策中模拟人类判断的准确性，并与真实人类共同所有者的预测进行对照，揭示了AI在微妙社会情境下的系统性偏差，对关注LLM仿真可靠性及失效条件的研究者具有参考价值。","inspiration":"借鉴其配对设计和多模型基准测试方法，可系统评估LLM在预测他人偏好或决策时的准确性，并与人类代理预测对照。｜可迁移到经济金融中涉及代理决策或偏好预测的场景，如理财顾问预测客户风险偏好、信贷员判断借款人还款意愿、或政策制定者预测公众对经济政策的接受度。｜设计：招募真实客户-理财顾问对，让客户对一系列投资产品的风险承受能力和偏好进行评分，同时让顾问预测客户评分，并让LLM基于客户基本信息和少量示例进行预测；结果变量为预测误差（MAE）和方向一致性；以顾问预测为人类基准，比较LLM与顾问的准确性，并考察客户-顾问关系强度、信息敏感度等调节因素。"}},{"id":"2609.24859","version":1,"title":"Small-world Networks of Agents Brainstorm AI Risks to Support Ideation","zh_title":"智能体小世界网络头脑风暴AI风险以支持构思","abstract":"The ideation phase of participatory AI risk assessment often starts with a blank slate or a limited list of predefined risks, making it difficult to surface indirect or systemic harms. To address this limitation, we propose a three-stage ideation support tool. The tool complements participatory AI, rather than replacing it, and helps focus later engagement with affected communities. First, it dynamically discovers stakeholders depending on the given AI use and recursively expanding outward, allowing overlooked or indirect stakeholders to emerge. Second, it simulates these stakeholders with LLMs, connecting them into a network of a given topology, and having them ideate about risks. Third, it prioritizes risks using network centrality measures. In an initial evaluation, we found that betweenness centrality run through agents connected in a small-world network works best as it elevates risks raised by stakeholders who bridge disconnected groups, surfacing novel, systemic harms that traditional methods often miss. On an AI chatbot companion use case, this approach increased the novelty of the identified risks by approximately 1.1 points over single LLM brainstorming, and by 0.5 points over agentic LLM brainstorming, measured on a normalized five-point Likert scale, without reducing the plausibility or severity of the identified risks. To test whether our framework helps a human-led ideation session using the Futures Wheel approach, we divided 11 teams of non-western young chatbot users into two types: control (team) and treatment (team) in a participatory AI risk assessment. The control teams started from a list of risks generated by the 45 AI practitioners in the initial evaluation; the treatment teams started from a list generated by our framework. The treatment teams identified more risks overall, and more systemic, human-computer interaction, and environmental risks.","authors":["Ke Zhou","Edyta Bogucka","Daniele Quercia"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24859","pdf_url":"https://arxiv.org/pdf/2609.24859","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","风险识别","参与式AI"],"reason":"用LLM模拟利益相关者头脑风暴AI风险，并与人类团队对照，涉及政策评估场景，但…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":16,"question":"如何利用小世界网络中的LLM智能体头脑风暴来支持AI风险识别的构思阶段，以发现更多新颖且系统性的风险？","design":"该研究提出一个三阶段框架：首先动态发现利益相关者，然后将其实例化为LLM智能体并连接成小世界网络进行风险头脑风暴，最后用网络中心性（特别是介数中心性）对风险进行排序。评估中比较了单LLM头脑风暴、多智能体基线以及该框架生成的风险在创新性、合理性和严重性上的表现，并进一步在人类主导的Futures Wheel工作坊中测试了该框架生成的风险列表对团队构思的促进作用。","baseline":"对照的真实人类数据包括：45名AI从业者在空白头脑风暴工作坊中生成的244条风险（用于评估风险生成质量），以及11个非西方年轻聊天机器人用户团队在Futures Wheel工作坊中的表现（控制组使用从业者生成的风险列表，处理组使用框架生成的风险列表）。","findings":"在AI聊天机器人伴侣用例中，小世界网络结合介数中心性的方法将风险创新性评分比单LLM头脑风暴提高了约1.1分，比智能体LLM头脑风暴提高了0.5分（5分制），且未降低合理性和严重性。在人类参与的Futures Wheel研究中，使用该框架生成的风险列表的处理组团队识别出了更多总体风险，以及更多系统性、人机交互和环境风险。","reliability":"论文未讨论","relevance":"该研究用LLM模拟利益相关者网络进行风险头脑风暴，并与真实人类团队进行对照，涉及AI政策评估场景，对关注LLM仿真可靠性和偏差的研究者具有参考价值。","inspiration":"该方法通过构建小世界网络并利用介数中心性来提升生成内容的多样性和新颖性，可借鉴其网络拓扑与中心性度量来设计多智能体仿真中的信息传播和观点聚合机制。｜可迁移到经济金融领域的政策评估或风险识别场景，例如金融系统性风险的早期预警或信贷审批中的歧视性风险识别。｜可设计一个研究：用LLM模拟银行、监管者、消费者等利益相关者，在小世界网络中讨论信贷审批算法的潜在风险，以生成的风险列表作为处理组，与人类专家生成的风险列表进行对照，比较两组在风险覆盖度、新颖性和系统性上的差异，并使用真实历史信贷数据或监管报告作为外部基准。"}},{"id":"2609.24629","version":1,"title":"Augmented Hypothesis Testing with Persona-Based LLM Simulations","zh_title":"基于角色LLM模拟的增强假设检验","abstract":"A/B testing requires large sample sizes, long timelines, and significant costs. When auxiliary predictions of experimental outcomes are available from machine learning models, uncertain prediction quality precludes replacing human experiments entirely, yet these predictions may still contain useful signal. We propose a principled framework for learning-augmented hypothesis testing that leverages predictions of unknown quality to reduce sample sizes while maintaining statistical validity. Predictions naturally vary in granularity, from coarse aggregate signals to fine-grained individual-level estimates, and our framework addresses both ends of this spectrum: (1) for population-level directional predictions, where only a binary signal on the treatment effect sign is available, we use an asymmetric test and prove consistency and robustness bounds within the learning-augmented algorithms paradigm; (2) for individual-level predictions, we introduce Generalized PPI++ (GPPI), extending Prediction-Powered Inference to handle nonlinear prediction errors through higher-dimensional transformations. Both methods benefit from accurate predictions while remaining robust to inaccurate or adversarial ones. We validate our framework using persona-based LLM simulations, where AI agents equipped with user personas predict individual behavior, as a natural prediction source spanning both granularity levels. Experiments on four real-world datasets demonstrate that our methods, combined with persona-based predictions, substantially reduce experimental costs while preserving rigorous statistical validity.","authors":["Ziyad Benomar","Aymen Al Marjani","Paul Missault","Saab Mansour"],"categories":["cs.LG","cs.AI","stat.AP"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24629","pdf_url":"https://arxiv.org/pdf/2609.24629","source_feed":"cs.LG","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM仿真","假设检验","统计有效性"],"reason":"用LLM persona预测个体行为，与真实数据对照，并保证统计有效性，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":15,"question":"如何利用 persona 驱动的 LLM 预测来减少 A/B 测试所需样本量，同时保持统计有效性？","design":"使用基于用户 persona 的 LLM 模拟来预测个体行为，作为辅助预测源；针对群体级方向预测（仅提供处理效应符号）和个体级预测分别提出不对称检验和广义 PPI++ 方法，在四个真实数据集上验证。","baseline":"四个真实世界数据集（具体名称未在节选中列出）中的人类 A/B 测试结果作为对照。","findings":"所提出的学习增强假设检验框架能够利用 persona 预测显著减少实验成本，同时保持严格的统计有效性。群体级方向预测和个体级预测方法均能在预测准确时受益，并在预测不准确或对抗性时保持稳健。","reliability":"论文承认完全用 persona 预测替代人类实验面临根本挑战：LLM 的黑箱性质、对提示工程的敏感性、分布偏移以及人类行为的复杂性使得形式化质量保证困难；此外，persona 与真实用户之间可靠的一对一映射通常不可行，导致预测粒度受限。","relevance":"该研究直接针对研究者关注的 LLM 人类仿真实验，利用 persona 预测个体行为并与真实数据对照，同时提供统计有效性保证，值得精读原文以了解具体方法和实验细节。","inspiration":"借鉴其将 LLM 预测作为辅助信号而非替代品，并通过统计校正保持有效性的思路｜可迁移到政策评估中的 A/B 测试，如利用 LLM 模拟消费者对价格变动的反应来减少实地实验样本量｜以消费者信贷决策为场景，用 LLM 基于用户 persona 预测个体对贷款条款的反应，处理为不同利率或审批条件，结果变量为接受/拒绝或违约行为，对照真实信贷申请数据。"}},{"id":"2509.20634","version":3,"title":"Recidivism Prediction, Peer Effect Estimation, and Prediction-Powered Inference with LLM Text Measures","zh_title":"使用LLM文本测量进行累犯预测、同伴效应估计与预测驱动推断","abstract":"We provide a new framework for estimating peer effects when outcomes are multivariate behavioral measures derived from written text using an LLM and the network formation is endogenous. We obtain LLM embeddings and zero shot classification of more than 200,000 written exchanges among residents of low-security correctional facilities. We find that LLM embeddings improve out-of-sample recidivism prediction by up to 30% over pre-entry covariates alone using LASSO and LoRA fine-tuning, showing that text representations capture meaningful signals. For peer effect estimation, we develop a novel instrumental variable estimator that accommodates multivariate outcomes, sparse networks, and multidimensional latent homophily. We show that this estimator is $\\sqrt{N}$-consistent and asymptotically normal under sparsity conditions that relax dense-network assumptions prevalent in the peer effect literature. Limited human annotations are then combined with LLM zero-shot vectors in a new prediction-powered peer inference (PPPI) approach to obtain de-biased estimates and valid inference. Results reveal significant peer effects in the behavioral profiles.","authors":["Shanjukta Nath","Jiwon Hong","Jae Ho Chang","Keith Warren","Subhadeep Paul"],"categories":["econ.EM","cs.AI","econ.GN","q-fin.EC","stat.ME"],"primary_category":"econ.EM","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2025-09-25","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2509.20634","pdf_url":"https://arxiv.org/pdf/2509.20634","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B3"],"tags":["LLM文本测量","同伴效应","预测驱动推断"],"reason":"用LLM从文本中提取行为测量并估计同伴效应，有真实人类数据对照，涉及经济学场景…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":17,"question":"如何利用LLM从文本中提取行为测量，在存在内生网络形成的情况下估计同伴效应，并改进再犯预测？","design":"本研究并非用LLM模拟人类被试，而是用LLM处理真实人类文本数据：对超过20万条低安全级别惩教设施中居民之间的书面交流进行嵌入和零样本分类，得到居民行为画像，然后开发新的工具变量估计器估计同伴效应，并结合预测驱动推断（PPPI）进行去偏估计。","baseline":"真实人类数据：来自治疗社区（TCs）的居民书面交流记录、入狱前协变量、三年再犯记录，以及少量人工标注用于验证LLM输出。","findings":"LLM嵌入在再犯预测中比仅用入狱前协变量提升最多30%的样本外预测准确率；新估计器在稀疏网络下具有√N一致性和渐近正态性，且估计结果显示行为画像中存在显著的同伴效应。","reliability":"论文未讨论","relevance":"该研究使用LLM从真实文本中提取行为测量并估计同伴效应，属于用LLM辅助分析人类行为而非替代人类被试，但涉及真实人类数据对照和经济学场景，对关注LLM在实证研究中应用的研究者有参考价值。","inspiration":"借鉴其利用LLM零样本分类将高维文本嵌入降维为可解释行为标签，并结合工具变量处理内生网络的方法。｜可迁移到金融文本分析，如利用分析师报告或公司公告文本测量管理层情绪或风险偏好，进而估计同行公司之间的情绪传染效应。｜设计：以分析师为被试，处理为同行分析师的乐观情绪文本，结果变量为分析师自身预测偏差，用真实分析师历史预测数据作为对照，采用类似工具变量策略识别同行效应。"}},{"id":"2608.28021","version":2,"title":"Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code","zh_title":"与什么相比？面向LLM生成基础设施即代码的以人为锚安全基准","abstract":"Large language models increasingly author Infrastructure-as-Code (IaC), where one insecure default is provisioned straight into production. Prior evaluations report vulnerability counts for models only, and so cannot say whether models are worse than the engineers they assist. We present GenIaC-SecBench: 100 deployment scenarios across 12 model configurations from six vendors, open and closed weights, yielding 1,196 artifacts scanned by three policy engines (Checkov, Trivy, KICS) at complete coverage. Crucially we scan 634 human-authored IaC templates with the identical toolchain, giving the first size-matched human security baseline for this task. Vulnerability density is strongly inverse to artifact size (Spearman $\\rho=-0.55$, $p<10^{-77}$), so unmatched comparisons measure size, not security. Size-matched, every configuration exceeds the human baseline at $3.21\\times$ to $3.87\\times$, and the gap widens as tasks get simpler ($4.9\\times$ at one resource, $1.4\\times$ at twenty or more). A majority of scenarios prescribe a security state rather than specifying function alone, so we stratify by prompt class: pooled the gap is $3.50\\times$, and excluding every scenario that explicitly requests an insecure configuration still leaves all configurations above baseline ($2.4\\times$ to $4.2\\times$). The corpus cannot isolate unprompted default posture, and we say so. Decomposing \"reasoning\" into standard generation, prompted chain-of-thought, and vendor extended-thinking APIs, extended thinking beats prompted CoT ($-12.0\\%$, $p=0.0013$) while prompted CoT alone is indistinguishable from standard ($-1.3\\%$, n.s.); it consumes under $1\\%$ of the output budget, bounding the effect. Two negative results: more deployable models are not more vulnerable ($r=0.158$, $p=0.625$), and complete-case Friedman is uncomputable here, motivating Skillings-Mack. All code and data are released.","authors":["Animesh Shaw"],"categories":["cs.CR","cs.AI","cs.MA","cs.SE"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2026-08-31","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.28021","pdf_url":"https://arxiv.org/pdf/2608.28021","source_feed":"cs.MA","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM安全评估","人类基线对照","基准测试"],"reason":"用人类基线对照评估LLM生成IaC的安全性，方法可迁移到仿真可靠性评估","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":18,"question":"LLM生成的IaC代码与人类工程师相比，安全漏洞密度如何？","design":"用12种LLM配置（9个模型、6家厂商，含开源与闭源）生成100个部署场景的IaC代码，共1196个工件，用Checkov、Trivy、KICS三个策略引擎扫描漏洞，并与634个人类编写的IaC模板在相同工具链下对比。","baseline":"634个人类编写的IaC模板，用相同的三个策略引擎扫描，并按资源数量分层匹配。","findings":"漏洞密度与工件大小强负相关（Spearman ρ=-0.55），未匹配的比较衡量的是大小而非安全性。大小匹配后，所有LLM配置的漏洞密度均为人类基线的3.21-3.87倍，且任务越简单差距越大。","reliability":"论文承认无法隔离未提示的默认安全姿态，因为多数场景明确规定了安全状态；扩展思维API使用默认温度，与标准生成存在混杂；完整案例Friedman检验不可计算，改用Skillings-Mack。","relevance":"该研究提供了人类基准对照的严谨方法，可用于评估LLM在安全相关任务中的仿真可靠性，对关注LLM仿真偏差的研究者有参考价值。","inspiration":"借鉴其分层匹配和工具链一致性的对照设计，避免因输出长度等混杂因素导致错误结论。｜可迁移到经济金融中的代码生成任务，如量化交易策略代码、金融数据处理脚本的安全性评估。｜用LLM生成金融分析代码，与人类分析师编写的代码对比漏洞密度，按代码行数或功能复杂度分层，使用静态分析工具扫描，以真实人类代码库为基准。"}},{"id":"2609.07358","version":3,"title":"Access to Live AI Advice and Behavior Under Risk: An Incentivized Experiment","zh_title":"获取实时AI建议与风险下的行为：一项激励实验","abstract":"Generative AI has become an everyday advisor, and the systems people consult are live and interactive, not pre-scripted. We ask whether access to such a system changes behavior under risk. In an incentivized experiment (N = 158), participants made lottery choices with an optional decision aid presented as a conventional pre-written tool, a live one-shot AI, or a live interactive AI they could query, with information format held equivalent across conditions. Risk preferences are elicited via DOSE. We find no evidence that access to a live AI advisor changes risk aversion.","authors":["Paul Althaus","Leon Houf","Christiane Schwieren"],"categories":["econ.GN","econ.TH","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-09-22","first_seen":"2026-09-09","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.07358","pdf_url":"https://arxiv.org/pdf/2609.07358","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM辅助决策","风险偏好","实验经济学"],"reason":"用LLM作为决策辅助，测量人类风险行为变化，有真实人类实验对照，但非仿真替代。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":19,"question":"在风险决策中，获得实时AI建议（一次性或可交互）是否会改变人们的风险厌恶程度？","design":"该研究不是用LLM仿真人类，而是将LLM作为决策辅助工具。158名被试在线完成激励性彩票选择任务，随机分为三组：对照组使用预编程的决策辅助工具（自动计算期望值并给出风险偏好提示），一次性AI组使用实时AI（gpt-5.4-mini）提供同样格式的建议，交互式AI组可额外追问两次。风险偏好通过DOSE方法估计CRRA参数。","baseline":"对照组为使用预编程决策辅助工具的人类被试，其风险厌恶参数作为基准。","findings":"与预编程工具相比，获得一次性或交互式AI建议对风险厌恶没有显著影响，点估计接近零。该零结果在咨询过AI、高信任AI的子样本中依然稳健，且设计能检测到与随机顺序效应相当的影响。","reliability":"论文承认只能排除大于0.15的效应，更小的效应仍可能存在；且研究仅识别了在信息格式相同的情况下，AI作为来源的净效应，并未声称AI提供的信息本身没有影响。","relevance":"该研究虽非仿真替代，但直接检验了LLM作为实时建议对经济决策行为的影响，与您关注的AI与人类行为交互、实验经济学场景高度相关，值得阅读原文了解实验细节和稳健性检验。","inspiration":"值得借鉴的是将AI建议的呈现方式（预编程 vs 实时AI）作为处理变量，同时保持信息内容一致，从而分离出“AI来源”的净效应，并通过随机顺序效应验证检验力。｜可迁移到资产定价实验，检验投资者在获得AI投资建议后风险资产配置是否改变。｜招募真实投资者为被试，随机分配使用传统财务计算器或实时AI助手获取相同信息，测量其风险资产配置比例，并与历史投资数据或对照组行为进行对比。"}},{"id":"2609.22188","version":1,"title":"Fairness Beyond Anonymization? Demographic Leakage in German LLM-Generated Resumes","zh_title":"匿名化之外的公平？德国LLM生成简历中的人口统计泄漏","abstract":"Large language models (LLMs) are increasingly integrated into AI-assisted hiring pipelines, including automated resume generation and screening. Under the EU AI Act, the hiring domain is classified as high-risk, making fairness and transparency critical requirements. Existing work has primarily focused on explicit hiring decisions, while less attention has been paid to whether generated resumes themselves encode recoverable demographic information. In this work, we conduct a two-stage audit of demographic leakage in German-language LLM-generated resumes. First, we use ChatGPT (GPT-4o-mini), Gemini 2.5 Flash-Lite, and multiple scales of the open-weight Qwen 3 model family (4B, 8B, and 14B) to generate resumes from real anonymized job-matching profiles, systematically varying gender- and ethnicity-associated names while holding qualifications constant. Second, we simulate a downstream resume screening scenario, where the generated resumes are first anonymized and gender-neutralized, before demographic leakage classifiers are trained on the resulting texts. We find that, despite these interventions, classifiers reliably distinguish between resumes generated with male and female names. This leakage is not driven by overtly gendered wording, but by subtle differences in the usage of semantically equivalent, formally gender-neutral terms in German. In contrast, ethnicity-related leakage remains comparatively weak across models. Our findings demonstrate that apparently neutral resume generation can still preserve highly predictive demographic signals, raising concerns about anonymization-based fairness interventions in multilingual AI hiring pipelines.","authors":["Charlotte Leininger","Helena Veit","Matthias A{\\ss}enmacher","Andreas Bender"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22188","pdf_url":"https://arxiv.org/pdf/2609.22188","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM生成简历","人口统计泄漏","公平性审计"],"reason":"审计LLM生成简历中的人口统计泄漏，评估匿名化公平干预的失效，有真实数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":22,"question":"LLM生成的德语简历在匿名化和性别中性化处理后，是否仍包含可恢复的性别和种族人口统计信息？","design":"使用ChatGPT、Gemini和Qwen 3系列模型，基于真实匿名求职匹配档案生成简历，系统变化姓名以关联性别和种族，保持资格不变；随后对生成简历进行匿名化和性别中性化处理，训练分类器检测人口统计泄漏。","baseline":"真实匿名求职匹配档案（来自德国求职匹配公司Chemistree GmbH），作为生成简历的资格基础，但未直接提供人类简历文本作为对照。","findings":"尽管进行了匿名化和性别中性化，分类器仍能可靠区分男性和女性姓名生成的简历，泄漏源于德语中语义等价但形式性别中立的术语使用差异；种族相关泄漏相对较弱。","reliability":"论文未讨论","relevance":"该研究通过审计LLM生成简历中的人口统计泄漏，评估匿名化公平干预的失效，有真实数据对照，对关注LLM仿真可靠性与偏差的研究者具有参考价值，值得阅读原文。","inspiration":"借鉴其控制变量法：保持资格不变，仅改变姓名关联的人口统计属性，并训练分类器检测泄漏，可迁移到信贷审批歧视研究；用LLM生成贷款申请文本，变化姓名暗示种族或性别，训练分类器预测人口统计属性，并与真实贷款申请数据对照。"}},{"id":"2609.22633","version":1,"title":"Beetle: A Bilingual Model Suite for Modelling Second-Language Processing","zh_title":"Beetle：用于建模第二语言处理的双语模型套件","abstract":"Bilingual language models (LMs) offer a controlled setting for studying how training conditions shape second-language (L2) behaviour, but prior work typically varies exposure structure, scale, and architecture at once, making it difficult to attribute effects to any single factor. We introduce Beetle, a controlled language model pretraining framework in which tokeniser, target language, training budget, and exposure structure are each independently manipulable, enabling systematic and comparable experimentation of training conditions. Using Beetle, we train and release 285 bilingual and 45 monolingual open-source LMs with rich checkpoints across a range of exposure schedules, data scales and first languages (L1s) to study multilingual pretraining and computational modelling of bilingualism and second language learning. Evaluating models on human bilingual and second language reading-time prediction and grammaticality judgement tasks, we find that staged and temporally structured curricula consistently improve alignment with language learner reading time compared to balanced bilingual training, with the largest gains at smaller data scales and for typologically closer language pairs. The Beetle models are well suited tools to help move computational psycholinguistics beyond its prevailing monolingual, English-centric focus toward models of human bilingual processing, to study cross-lingual learning dynamics, while supporting community-based development of controlled model families.","authors":["Suchir Salhan","Catherine Arnett","James Michaelov","Paula Buttery"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22633","pdf_url":"https://arxiv.org/pdf/2609.22633","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["计算心理语言学","双语模型","人类行为预测"],"reason":"用双语模型预测人类二语阅读时间和语法判断，有真实人类数据对照，属于LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":23,"question":"在双语语言模型预训练中，语言暴露的时间结构（如分阶段课程与平衡双语训练）如何影响模型对人类二语阅读时间和语法判断行为的对齐程度？","design":"使用 Beetle 框架训练 285 个双语模型（125M 参数，固定架构和分词器），以英语为 L2，变化 21 种 L1、三种数据规模（100M、2B、24B tokens）和五种暴露条件（包括分阶段课程、平衡双语等），并训练 45 个单语基线模型。评估模型在人类二语阅读时间预测（MECO 眼动数据）和语法判断任务上的表现。","baseline":"人类基准来自 Multilingual Eye-movement Corpus (MECO) 的二语阅读时间数据，以及 CEFR 分层的二语学习者错误判别任务（BLiSS）。","findings":"分阶段和时序结构课程比平衡双语训练更一致地改善了对语言学习者阅读时间的对齐，尤其在较小数据规模和类型学相近的语言对中收益最大。但课程效应在不同任务上并不统一：在语法性和错误判别任务上，平衡暴露往往具有竞争力或更好。","reliability":"论文指出课程效应不会均匀迁移到所有任务：在语法性和错误判别任务上，平衡暴露可能更好；此外，模型对齐可能受数据规模、语言类型学距离等因素影响，但未系统讨论失效条件。","relevance":"该研究用双语模型模拟人类二语处理，有真实人类眼动和语法判断数据作为对照，属于 LLM 仿真人类认知行为的研究，且提供了大规模受控模型套件，对关注仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其受控训练框架：通过独立操纵暴露结构、数据规模和语言对，并设置匹配的单语基线，可精确归因训练条件对行为的影响。｜可迁移到经济金融中的跨期选择或风险偏好实验，例如研究不同信息呈现顺序（如先呈现收益后呈现风险）如何影响决策。｜用 LLM 作为被试，施加不同的信息暴露课程（如分阶段呈现历史价格与基本面信息），测量其投资决策或风险偏好，并与真实人类实验数据（如实验室资产定价实验）对照，检验仿真一致性。"}},{"id":"2609.22971","version":1,"title":"Automatic multimodal UX improvement recommendations from LLM agent user simulations","zh_title":"基于LLM智能体用户仿真的自动多模态用户体验改进建议","abstract":"Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising route to automating UX testing by simulating realistic user behaviour. However, existing simulation approaches typically lack multimodality and require time-consuming manual review to extract actionable insights. We formalise UX improvement recommendation from simulation data as a structured natural language generation and ranking problem, and establish an evaluation protocol using expert annotation and LLM-as-a-Judge. We present AMUSER, a multimodal framework which simulates user behaviour and automatically generates prioritised UX improvement recommendations from resulting data. We evaluate AMUSER on commercial websites and show that its recommendations substantially outperform those from text-only simulation (NDCG@3 = 0.758 versus 0.359) at an 89% lower simulation cost. Our results suggest an asymmetric role of multimodality: visual access during simulation improves recommendations through richer traces, while providing visual inputs during recommendation generation can modestly degrade quality. We also discuss practical deployment lessons from applying AMUSER to commercial websites.","authors":["Anu Chowdhury","Bin Wu","Hossein A. Rahmani","Emine Yilmaz"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22971","pdf_url":"https://arxiv.org/pdf/2609.22971","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","用户体验","多模态"],"reason":"用LLM模拟用户行为并生成UX建议，有真实用户数据对照，属人类仿真但场景偏应用。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":25,"question":"能否从LLM智能体模拟用户行为的数据中自动生成有用的UX改进建议，且多模态模拟是否优于纯文本模拟？","design":"使用多模态LLM智能体（AMUSER）模拟不同用户画像在真实网站上的任务执行过程，记录动作、想法、情绪和截图，随后由LLM根据交互轨迹生成并排序UX改进建议。","baseline":"无对照","findings":"多模态模拟生成的建议显著优于纯文本模拟（NDCG@3=0.758 vs 0.359）和静态分析（0.273），且模拟成本降低89%。视觉输入在模拟阶段有益，但在建议生成阶段提供视觉输入会略微降低质量。","reliability":"论文未讨论","relevance":"该研究展示了LLM模拟用户行为并自动提取可操作见解的完整流程，与您关注的仿真可靠性和自动化评估高度相关，但缺少真实人类数据对照，且场景偏应用而非经济学实验。","inspiration":"借鉴其多模态模拟与自动生成结构化建议的流程，可用于经济金融场景中模拟消费者或投资者行为并提取决策模式。｜可迁移到消费者在线购物决策、金融产品选择或投资者信息处理等场景。｜设计：用LLM智能体模拟不同风险偏好的投资者在金融网站上的信息搜索与决策过程，处理为是否提供视觉信息（如图表），结果变量为决策质量或信息获取效率，并与真实投资者实验数据对照。"}},{"id":"2609.23039","version":1,"title":"Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity","zh_title":"审计LLM助手中的政治对齐：参与、立场与用户身份","abstract":"LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their behavior is dynamic. I argue that their political behavior is a set of policies over whom to answer, what to say, and whether to engage at all, conditional on the topic and what the system knows about the user. I call these policies the system's speech regime, which is how a developer settles the tradeoff between answering, accommodating the user, and refusing, each of which carries a cost that varies by topic. I derive a typology of five regimes from two dimensions, engagement and stance. I test six AI systems (OpenAI, Anthropic, xAI, Google, Mistral, DeepSeek) in a preregistered experiment of 7,500 multi-turn conversations that randomly assign the user's political identity across five topics: abortion, Catalan independence, climate change, Nazism, and a zero-stakes control (pineapple on pizza). Two LLM judges from different developers score every answer, validated against human coding, and refusal is treated as an outcome rather than missing data. Every system accommodates the user on the control topic, showing that political restraint is a policy. On contested topics the systems fall into different regimes: on abortion, GPT engages and mirrors every user, Gemma refuses everyone, Claude answers strongly conservative users 35 percent of the time and almost no one else, and Grok accommodates conservatives only. On settled topics such as climate change and Nazism, five systems hold firm for every user. The systems also infer the user's overall ideology, so accommodation can spill over to topics not yet discussed. A comparison of two Grok releases shows the regime changing between versions in a way current audits miss. Speech regimes matter for alignment research and for polarization, political knowledge, and the quality of democracy.","authors":["Joan C. Timoneda"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23039","pdf_url":"https://arxiv.org/pdf/2609.23039","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","政治行为","审计"],"reason":"用LLM模拟不同政治身份用户的回答，并与人类编码对照，评估系统行为差异，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":26,"question":"AI助手面对不同政治身份的用户时，在政治话题上的回答行为（是否回答、立场、是否迎合用户）如何随话题和用户身份变化？","design":"审计实验：用脚本化用户（随机分配政治身份）与六个商业LLM助手进行7500段多轮对话，涉及五个话题（堕胎、加泰罗尼亚独立、气候变化、纳粹、披萨放菠萝作为零风险对照），测量系统是否回答、立场、以及是否迎合用户，并用两个不同开发者的LLM法官评分（经人类编码验证）。","baseline":"人类编码验证LLM法官评分，但无大规模真实人类行为数据作为对照。","findings":"系统在争议话题上表现出不同“言论体制”：GPT迎合所有用户，Gemma几乎拒绝所有人，Claude选择性拒绝（更多回应保守派但不迎合），Grok只迎合保守派；在共识话题上五个系统立场坚定。系统还会推断用户整体意识形态，导致迎合溢出到未讨论的话题。","reliability":"论文未明确讨论失效条件，但指出当前审计方法（测量对“平均用户”的回答）会错过版本间言论体制的变化，且未覆盖所有可能话题和用户身份。","relevance":"该研究展示了如何用LLM模拟不同身份用户来审计系统行为差异，并强调将拒绝视为结果而非缺失数据，对用LLM进行人类仿真实验的方法论有借鉴意义，值得阅读原文。","inspiration":"借鉴其将“拒绝回答”作为结果变量、随机分配用户身份、多轮对话施加压力的设计，以及用LLM法官加人类验证的测量方式。｜可迁移到信贷审批歧视研究：让LLM扮演不同种族、性别、收入水平的贷款申请人，向AI信贷助手申请贷款，观察是否被拒绝、获批额度、利率差异。｜设计：用LLM生成标准化贷款申请对话，随机分配申请人特征（种族、性别、收入），让真实银行AI客服或LLM模拟的信贷员处理，结果变量为是否批准、额度、利率，与真实信贷审批数据（如HMDA数据）对照，检验AI决策中的歧视。"}},{"id":"2609.23936","version":1,"title":"Think Before You Accept: Can Written Justification Reduce Uncritical Uptake of AI Writing Suggestions?","zh_title":"接受前先思考：书面理由能否减少对AI写作建议的不加批判采纳？","abstract":"Generative AI can offer students useful feedback, but its value depends on judging which suggestions are accurate and relevant. Prior research shows that strategic friction during human-AI interactions can promote critical uptake, but how to effectively implement such friction in academic contexts remains unclear. We examine whether requiring students to justify decisions to accept or reject AI suggestions can mitigate uncritical uptake in academic writing. In a randomized experiment embedded in a course activity (N=129), students wrote a data analysis proposal, received mixed-quality AI revision suggestions, and decided whether to accept or reject them. Students required to provide written justifications were 24 percentage points less likely to adopt flawed suggestions (65% vs. 41%), with no reduction in acceptance of sound suggestions (81% vs. 86%). However, thematic analysis revealed superficial engagement in the justification task and gaps in metacognitive monitoring and domain knowledge.","authors":["Yan Tao","Jennifer Meyer","Rene F. Kizilcec"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23936","pdf_url":"https://arxiv.org/pdf/2609.23936","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["人机交互","AI建议采纳","批判性思维"],"reason":"用LLM生成写作建议，人类被试在真实任务中决策，有对照实验，但非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":27,"question":"在学术写作中，要求学生为接受或拒绝AI修改建议提供书面理由，能否减少对AI建议的不加批判的采纳？","design":"本研究不是LLM仿真人类被试的研究，而是随机对照实验：129名学生在课程活动中撰写数据分析提案，收到混合质量的AI修改建议，并被随机分配是否必须为每个采纳/拒绝决定提供书面理由；结果变量为对高质量和低质量建议的接受率。","baseline":"无对照（无真实人类数据作为仿真基准；实验本身以学生真实行为为数据，但非仿真研究）","findings":"要求书面理由使低质量建议的采纳率从65%降至41%，而高质量建议的采纳率无显著下降（81% vs 86%）。但主题分析显示学生在理由中参与肤浅，存在元认知监控和领域知识缺口。","reliability":"论文未讨论仿真可靠性，但承认书面理由干预未能完全消除不加批判的采纳，且学生参与度有限，提示干预效果可能受限于学生的元认知和领域知识。","relevance":"该研究虽非LLM仿真人类被试，但涉及人类对AI建议的决策行为，且通过随机实验揭示了干预对决策质量的影响，对关注AI辅助决策中人类行为的偏差与干预的研究者有参考价值。","inspiration":"借鉴其随机实验设计，通过施加认知摩擦（如要求书面理由）来改变决策行为，并测量对不同质量信息的区分度。｜可迁移到经济金融中的AI辅助决策场景，如投资者对AI投资建议的采纳、信贷审批中AI风险评分的接受、或消费者对AI推荐产品的选择。｜设计一个实验：招募真实投资者作为被试，提供AI生成的股票推荐（包含高质量和低质量信号），随机要求部分被试为每个采纳/拒绝决定写理由，结果变量为对高质量和低质量推荐的采纳率，并以历史市场数据或专家评级作为建议质量的基准。"}},{"id":"2609.24532","version":1,"title":"Prompting Against Persona Drift: Comparing Intervention Timing and Content in LLM-Simulated Conversations","zh_title":"对抗角色漂移的提示策略：比较LLM模拟对话中的干预时机与内容","abstract":"Simulating student personas with large language models (LLMs) enables scalable evaluation of educational systems. However, behavioral drift, a progressive decline in persona consistency, can emerge over extended conversations, limiting the validity of such simulations. We evaluate five prompt-level mechanisms using separate monitoring and intervention pipelines. Across 1,200 28-turn conversations spanning four LLMs and two ADHD persona intensities, we varied when to intervene (static vs. adaptive) and what to inject (reinjection vs. reflective reminder), plus a novel adaptive condition in which a monitor generates behavior-specific instructions. Relative to no intervention, reinjection reduced the modeled rate of LLM-rated drift by 35--38\\%, reflective reminders by 22--27\\%, and behavior-specific instruction by 87\\%. None eliminated drift. We found no evidence that adaptive timing outperformed static scheduling. Monitoring therefore appears more useful for deciding \\textit{what} to correct than \\textit{when} to intervene, although behavior-specific instruction requires component-level testing.","authors":["Nicolas Leins","Jennifer Haase","Varvara Geronimus","Jana Gonnermann-M\\\"uller","Sebastian Pokutta"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24532","pdf_url":"https://arxiv.org/pdf/2609.24532","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","A2","B4"],"tags":["LLM仿真","角色一致性","教育评估"],"reason":"用LLM模拟学生角色并评估一致性，属于人类仿真，但无真实人类数据对照，且聚焦角…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":29,"question":"在LLM模拟学生角色的长对话中，提示级干预机制能否缓解角色漂移，且自适应干预是否优于静态干预？","design":"用四个LLM模拟ADHD学生角色（两种强度），在1200段28轮对话中，通过分离的监测与干预管道，比较五种提示级干预条件（无干预、静态全角色重注入、静态反思提醒、自适应全角色重注入、自适应反思提醒、自适应行为特定指令），以LLM评定的行为量表得分变化作为结果变量。","baseline":"无对照","findings":"所有干预条件均显著减缓了角色漂移，但未能完全消除；自适应时机相比静态调度无一致优势，而行为特定指令的干预效果最强，漂移减少87%。","reliability":"论文未讨论","relevance":"该研究属于LLM人类仿真，但无真实人类数据对照，且聚焦于教育场景的角色一致性，与研究者关注的经济学实验和政策评估场景关联有限，但方法上对仿真可靠性评估有参考价值。","inspiration":"借鉴其分离监测与干预的双管道设计，以及基于行为测量的自适应干预思路，可用于经济金融仿真中的行为一致性维护。｜可迁移到消费者跨期选择实验或投资者风险偏好模拟中，防止LLM在长对话中偏离预设的经济决策模式。｜以LLM模拟不同风险偏好的投资者，在连续投资决策对话中施加行为特定指令干预，结果变量为风险选择序列，并与真实投资者实验数据对照。"}},{"id":"2609.24146","version":1,"title":"Mind or Message? Auditing Theory of Mind in Multi-Agent Social Simulation","zh_title":"心智还是信息？审计多智能体社会模拟中的心智理论","abstract":"Language model agents are increasingly used to simulate social interaction, and the resulting transcripts read as though the agents understand one another. We ask whether that appearance rests on a model of the partner's mind or on the surface record of what the partner said. We build a social simulation in which both questions have exact answers: 40 multi-issue negotiations whose hidden preference weights and whose full Pareto frontier are known by construction. Two model families negotiate across 160 dyads, every transcript is frozen before any measurement, and 2880 counterfactual probes then hold the evidence byte identical while moving one factor at a time: the reader's own stake, the partner's tone, an identity label, and the order of recursion. The agents are socially fluent and economically poor. They reach agreement in 96.2% of dyads with 0 protocol failures, yet only 0.7% of deals land on the Pareto frontier, they leave 20.5% of the available joint value unclaimed, and they miss the one issue on which their interests are perfectly aligned in 76.6% of deals; on the frontier and on that aligned issue, a package drawn at random from the set both sides would accept does as well. The probes locate the failure. Swapping only the reader's own payoff sheet, while the partner's words and offers stay identical, moves the inferred top priority by 15.0 percentage points, which is egocentric projection rather than inference, while a tone rewrite moves it by 5.3 percentage points and an identity label by 0.0. Most tellingly, an agent predicts what its partner believes about it 72.5% of the time while that partner's belief is itself correct only 51.2% of the time: the agents track the conversation far better than they track the mind behind it.","authors":["Cong Li","Cheng Chen","Thomas Fung","Alex Rossi","Yi Li"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24146","pdf_url":"https://arxiv.org/pdf/2609.24146","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A1","A2","B4"],"tags":["LLM仿真","心智理论","多智能体谈判"],"reason":"用LLM模拟谈判并审计心智理论，虽无人类对照，但批判性评估仿真失效条件，方法可…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":28,"question":"语言模型代理在模拟社会互动时，其表现是基于对对方心理状态的建模，还是仅基于对对话表面记录的读取？","design":"使用两个模型家族在40个多议题谈判场景中组成160个二人组进行谈判，每个场景的偏好权重和帕累托前沿已知。谈判后冻结所有对话记录，然后施加2880个反事实探针，每次仅改变一个因素（读者自身收益表、对方语气、身份标签、递归顺序），测量推断出的对方首要优先级的变化。","baseline":"无对照","findings":"代理在社交上流利（96.2%达成协议，0协议失败），但在经济上表现差：仅0.7%的协议达到帕累托前沿，损失20.5%的联合价值，76.6%的交易错过双方利益完全一致的议题。探针显示，改变读者自身收益表使推断的对方首要优先级移动15.0个百分点，表明自我投射而非推断；改变语气移动5.3个百分点，身份标签无影响。代理预测对方信念的准确率为72.5%，而对方信念本身正确率仅51.2%，说明代理更擅长追踪对话而非对方心理。","reliability":"论文承认其审计协议是诊断性的，不主张社会仿真无效，但强调有效性需要潜在真实基准来验证。局限包括仅使用两个模型家族、特定谈判场景，且未与人类行为直接对比。","relevance":"该研究批判性地评估了LLM在社会仿真中的心智理论能力，通过精确的反事实设计揭示了表面流畅性下的认知缺陷，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文以了解其审计协议和发现。","inspiration":"借鉴其反事实探针设计，在冻结证据下逐一改变因素以隔离因果效应，可用于经济实验中识别决策机制。｜可迁移到谈判博弈、拍卖或合作博弈等经济场景，检验LLM代理是否真正理解对手偏好或仅依赖表面信息。｜设计一个双边贸易谈判实验，用LLM代理作为被试，随机改变一方代理的收益表（处理），测量其对对方优先级的推断和最终协议效率，并与人类谈判数据（如实验经济学中的谈判结果）对照，评估LLM仿真的有效性。"}},{"id":"2608.18265","version":4,"title":"Modeling Human Behavior with Type Vectors Using AI","zh_title":"使用AI类型向量建模人类行为","abstract":"We introduce a general, easy-to-implement AI-based modeling technique for analyzing human behavior. A key feature of this approach, which contrasts with existing modeling techniques, is that it combines the flexibility and interpretability of natural language with a mathematical structure that can be fitted to data and easily analyzed. We assign a large language model a vector of trait intensities-a type vector-and then ask it to choose actions across settings in which we observe human choices. For instance, the type vector (2,4) could correspond to \"You are a player characterized by the following profile: Altruism: 2 out of 5, Risk Aversion: 4 out of 5,\" after which it is asked to make choices. We can then vary the traits (e.g., Altruism, Fairness, Trust,...) and values (e.g., 1-5) to minimize distance to human choices. We illustrate the method by applying it to model 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles. We find that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust. The type vectors needed to fit individuals across games cluster into fewer than a dozen groups, with substantial variation in fit across subjects. Moreover, the individual type vectors can predict behavior in held-out games with different rules and available actions. More broadly, this new modeling method is highly generalizable and interpretable: we can input any vector of traits and use them to model behavior across any setting","authors":["Matthew O. Jackson","Benjamin S. Manning","Yutong Xie","Walter Yuan","Qiaozhu Mei"],"categories":["econ.TH","cs.AI"],"primary_category":"econ.TH","announce_type":"replace-cross","date":"2026-09-21","first_seen":"2026-08-20","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.18265","pdf_url":"https://arxiv.org/pdf/2608.18265","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","经济实验","人类行为建模"],"reason":"用LLM模拟人类经济决策，并与大规模真实人类数据对照，直接命中核心方向。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":1,"question":"人类行为是否可以用少数几个特质维度（类型向量）来近似和预测？","design":"用大语言模型（LLM）作为仿真被试，通过提示词赋予其不同特质强度（如利他、风险厌恶、信任等）构成类型向量，然后让模型在10个经典经济博弈角色中做决策，通过调整类型向量最小化与人类选择的距离。","baseline":"119,147条决策，来自78,657名被试，覆盖35个国家，在10个经典经济博弈角色中的真实选择。","findings":"人类行为可以用三个维度（风险厌恶、策略复杂性、信任）紧密匹配；个体类型向量在博弈间聚类成少于12个群体，且能预测未参与拟合的博弈中的行为。","reliability":"论文未讨论","relevance":"直接命中核心方向：用LLM模拟人类经济决策并与大规模真实人类数据对照，方法可解释且可推广，值得精读原文。","inspiration":"借鉴其用可调类型向量作为LLM提示来系统扫描特质空间并拟合真实行为的方法，可迁移到资产定价实验中的风险偏好与信念异质性建模。｜设计一个研究：用LLM模拟投资者，赋予不同风险厌恶和过度自信水平，在实验性资产市场中交易，结果变量为价格泡沫程度和交易量，与真实实验市场数据（如Smith et al. 1988）对照，检验类型向量能否复现泡沫。"}},{"id":"2609.21636","version":1,"title":"Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30","zh_title":"引导大语言模型在挪威MFQ-30上向道德基础靠拢","abstract":"Recent work applies human psychometric questionnaires to large language models to elicit moral and value profiles, but it is not clear whether these instruments measure anything stable in models or whether the resulting profiles can be moved toward a target human population. We administer the Norwegian Moral Foundations Questionnaire (MFQ-30) to six open-weight LLMs and compare their foundation profiles to a sample of N = 1,282 Norwegian respondents. We test two steering interventions, prompt-level persona steering and activation-level ActAdd. Half the models engage with the questionnaire under our attention check. The other half default to flat or central-tendency outputs that look near-human on average without tracking item content. A neutral Nordic-respondent persona, written without any distributional information from the human sample, brings the engaging models 44-77% closer to the Norwegian mean in Mahalanobis $d^2$. One-pair ActAdd at a fixed mid-layer flattens the foundation profile rather than steering individual foundations. For at least one model the same persona that shifts the profile also induces engagement that was absent at baseline, a concrete instance of the cognitive phantoms that Peereboom et al. (2025) warn about.","authors":["Hans Andersen","David Dichas"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21636","pdf_url":"https://arxiv.org/pdf/2609.21636","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","道德基础","算法保真度"],"reason":"用LLM复现人类道德基础分布，并与挪威样本对照，评估仿真可靠性与偏差，含批判性…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":4,"question":"LLM 在挪威版道德基础问卷（MFQ-30）上的道德画像是否稳定，能否通过提示或激活干预向目标人群靠拢？","design":"用六个开源 LLM 扮演挪威受访者，回答挪威 MFQ-30 问卷；施加两种处理：提示层 persona 引导（中性北欧受访者人设）和激活层 ActAdd 引导（对比向量注入）；结果变量为五个道德基础得分及与人类样本的马氏距离。","baseline":"Enstad 和 Finseraas (2024) 收集的 1282 名挪威受访者 MFQ-30 数据，已过滤不专注样本。","findings":"半数模型在注意力检查下真正作答，另一半输出扁平或趋中，平均接近人类但不追踪题目内容。中性北欧人设使作答模型与挪威均值的马氏距离缩小 44–77%，而 ActAdd 在固定中间层会扁平化画像而非定向引导单个基础。","reliability":"论文指出部分模型在基线时不作答，注意力检查可识别；人设引导可能诱发“认知幻影”，即模型表现出原本没有的作答行为，提示仿真结果可能不可靠。","relevance":"直接命中研究者对 LLM 仿真可靠性与偏差的关注，提供了真实人类对照和批判性发现，值得精读原文以了解注意力检查与引导干预的具体实现。","inspiration":"借鉴其注意力检查与第一 token 概率解码方法，可提高 LLM 问卷作答质量并识别无效仿真｜可迁移到经济政策偏好调查或消费者态度测量，如用 LLM 模拟不同人群对税收、福利或环保政策的道德评价｜用开源 LLM 扮演不同社会经济群体，施加中性人设或激活引导，测量其对政策陈述的同意程度，并与真实调查数据（如欧洲社会调查）对照，检验仿真偏差。"}},{"id":"2609.21259","version":1,"title":"CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition","zh_title":"CogGym：迈向人类与机器认知的大规模比较评估","abstract":"Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models' improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model--human fit remains well below human splithalf reliability ($R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video) with the best models achieving $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.","authors":["Lance Ying","Jinzhou Wu","Yingshan Susan Wang","Shivam Aarya","Luca M. Schulze Buschoff","Harry Chen","Katherine M. Collins","Andrea de Varda","Shuhao Fu","Sean Dae Houlihan","Akshay K. Jagadish","Guangyuan Jiang","Samuel Kiegeland","Tetsu Kurumisawa","Rongzhi Liu","Ryan Liu","Ningshan Ma","Kathryn McGregor","Younes Strittmatter","Polina Tsvilodub","Jacob Hoover Vigly","Sarah Wu","Enjie Xu","Yiling Yun","Kelsey Allen","Tyler Brooke-Wilson","Brian Christian","Evelina Fedorenko","Michael C. Frank","Michael Franke","Tao Gao","Samuel J. Gershman","Robert D. Hawkins","Jennifer Hu","Julian Jara-Ettinger","Max Kleiman-Weiner","Sydney Levine","Tal Linzen","Hongjing Lu","Timothy O'Donnell","Desmond C. Ong","Steven T. Piantadosi","Rebecca Saxe","Eric Schulz","Tianmin Shu","Felix A. Sosa","Ilia Sucholutsky","Tan Zhi-Xuan","Tomer Ullman","Fei Xu","Ilker Yildirim","Jian-Qiao Zhu","Thomas L. Griffiths","Tobias Gerstenberg","Kevin Smith","Joshua B. Tenenbaum"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21259","pdf_url":"https://arxiv.org/pdf/2609.21259","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知实验","人类对照"],"reason":"直接比较LLM与人类在认知实验中的行为，有真实人类数据对照，并评估模型-人类拟…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":3,"question":"如何在大规模、多样化的认知实验中系统比较AI模型与人类的行为，以识别模型与人类判断的相似与分歧？","design":"构建CogGym框架，将258个来自100篇论文的人类常识推理认知实验标准化为统一的实验标记语言（EML），并评估50个大型语言模型在这些实验上的表现，测量模型输出与人类反应分布的拟合度。","baseline":"原始认知实验中的真实人类被试反应数据，包括文本、图像和视频模态，并报告人类分半信度作为上限。","findings":"模型规模越大、越新，与人类判断的拟合度越高，但在常识推理任务上的进步速度慢于数学和编程等正式推理基准。最佳模型与人类的拟合度（文本R²=0.59，图像R²=0.58，视频R²=0.43）仍远低于人类分半信度（文本R²=0.93，图像R²=0.95，视频R²=0.92）。","reliability":"论文指出模型-人类拟合度远低于人类分半信度，表明模型尚未完全捕捉人类认知的细微结构；同时，标准化过程中可能存在实验转换的失真，且当前覆盖范围限于常识推理，未涉及其他认知领域。","relevance":"该研究直接比较LLM与人类在大量认知实验中的行为，有真实人类数据对照，并评估模型-人类拟合度，与研究者关注的人类仿真实验高度相关，值得精读以了解大规模评估框架和模型偏差。","inspiration":"借鉴其半自动化实验标准化流程和分半信度作为上限的评估方法，可迁移到经济决策实验（如风险偏好、时间贴现、博弈行为）的仿真验证。｜可应用于消费者跨期选择或资产定价实验，检验LLM是否复现人类的时间不一致性或风险厌恶。｜设计：以LLM为被试，施加跨期选择任务（如现在100元 vs. 一个月后120元），测量贴现率，并与真实人类实验数据（如Andersen et al., 2008）对照，计算模型-人类拟合度与人类分半信度的差距。"}},{"id":"2609.20827","version":1,"title":"From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators","zh_title":"从出院记录到患者理解：基于人格的开放式模拟将LLM作为出院教育者","abstract":"Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient's literacy, recall, and personality. Existing LLM evaluations target static or artifact-generation tasks and do not measure patient understanding under open-ended dialogue. We introduce DischargeBench, a persona-grounded simulation in which a candidate LLM educator conducts a multi-turn session with a Virtual Patient, while an Education Monitor Agent regulates patient realism without modifying the educator, protecting the evaluation signal. We curate MIMIC-IV-Ext-DischargeBench, 477 cases over 24 ICD chapters with persona axes (personality, education level, health literacy, past-medical-history recall) for stratified analysis. Each simulation is scored on four axes -- Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency -- by an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source-answer agreement. LLM evaluation for discharge education should center patient understanding, not text quality or answer accuracy alone.","authors":["Won Seok Jang","Zonghai Yao","Hong Yu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20827","pdf_url":"https://arxiv.org/pdf/2609.20827","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","患者教育","人类数据对照"],"reason":"用LLM模拟患者进行出院教育评估，有真实临床数据对照，涉及医疗场景，并指出仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":4,"question":"如何评估大语言模型在开放式、多轮对话中作为出院教育者，使不同人格与健康素养的患者真正理解出院指导的能力？","design":"构建 DischargeBench 仿真框架：用 LLM 扮演虚拟患者（基于 MIMIC-IV 真实出院记录，设定人格、教育水平、健康素养、病史回忆等属性），候选 LLM 扮演教育者进行多轮对话，并由教育监控代理仅调节患者侧以保持真实性；通过 LLM 裁判对对话质量、主题覆盖、理解程度和事实一致性四个维度评分。","baseline":"使用 MIMIC-IV-Ext-DischargeBench 的 477 个真实病例（来自 MIMIC-IV 和 MIMIC-IV-Note），并由医生标注用于对齐 LLM 裁判的评分。","findings":"GPT-5 系列在对话质量和主题覆盖上领先，但可读性较差；总体分数掩盖了不同 ICD 章节和患者人格下的显著差异，困难人格暴露了覆盖失败、理解差距和源答案一致性下降。","reliability":"论文指出 LLM 评估应关注患者理解而非文本质量或答案准确性；困难人格下仿真会失效，且 LLM 裁判的评分可能与医生标注存在偏差。","relevance":"该研究用 LLM 模拟患者进行出院教育评估，有真实临床数据对照，并批判性指出仿真在困难人格下失效，与研究者关注的人类仿真可靠性与偏差高度相关，值得精读。","inspiration":"借鉴其多智能体仿真设计：用 LLM 扮演异质性个体并设置监控代理防止污染处理信号，同时用真实数据校准裁判评分。｜可迁移到消费者金融教育或政策沟通场景，如评估 AI 理财顾问对不同金融素养人群的讲解效果。｜设计：用 LLM 模拟不同金融素养和人格的投资者作为被试，处理为 AI 顾问的个性化解释，结果变量为投资决策质量和风险理解，对照真实投资者调查数据（如 FINRA 金融素养调查）。"}},{"id":"2609.21439","version":1,"title":"People escalate against a competitor labelled human and hold back against one labelled an optimising machine","zh_title":"人们面对标记为人类的竞争者会升级投入，面对标记为优化机器的竞争者则会退缩","abstract":"People increasingly compete against AI agents rather than other human opponents. We distinguish two channels: an opponent effect and an information effect. These are different elements with different consequences: the opponent effect is specific to a given computational system, the information effect a property of the information environment that an organisation or policymaker can control. We separate them in a preregistered experiment (N = 1,395) using a dynamic all-pay auction, a repeated contest in which escalation of commitment arises from the incentives. What participants are told about the opponent (human, an AI trained to imitate people, or an AI trained to compete well) is varied and crossed with who they actually face, in a deception-free design. What people are told influences escalation: the median price rises by 6.7 points when a human might be the opponent and falls by 8.8 when an optimising machine might be, a spread of about 15% of the prize value of the competition, produced by information alone. Competing against the AI agents lowers prices, yet reduces the chance that both sides finish with positive earnings, showing distinct effects of the opponent channel. The information effect is not explained by articulated strategy, or individual differences, and is consistent with a competitive response engaged when a human is a live possibility. This shows that describing an AI competitor is not behaviourally neutral.","authors":["Vinicius Ferraz","Leon Houf"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21439","pdf_url":"https://arxiv.org/pdf/2609.21439","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","行为实验","人机交互"],"reason":"用LLM作为对手与人类被试互动，研究标签信息对行为的影响，有真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":11,"question":"在人类与AI竞争的情境中，标签信息（告知对手是人还是AI）是否独立于对手实际行为影响人们的竞争升级行为？","design":"采用预注册实验（N=1395），使用动态全支付拍卖（重复消耗战）作为竞争任务。通过欺骗自由设计交叉两个维度：告知信息（对手可能是人类、模仿人类的AI、优化竞争的AI或无信息）与实际对手（人类、模仿人类的AI、优化竞争的AI）。测量结果变量为出价升级（中位数价格）和双方均获正收益的概率。","baseline":"真实人类被试（Prolific平台招募）与真实人类对手、两种AI对手（模仿人类和优化竞争）的实际对局数据。","findings":"信息效应显著：仅告知对手可能是人类使中位数出价上升6.7点，告知可能是优化机器使中位数出价下降8.8点，差距约为奖品价值的15%。对手效应表现为与AI实际竞争降低出价，但同时减少双方均获正收益的概率约三分之二。","reliability":"论文未讨论","relevance":"该研究用真实人类被试与AI对手互动，通过标签信息操纵考察行为变化，有真实人类数据对照，直接回应了LLM仿真中信息环境对行为的影响，值得精读以理解信息效应与对手效应的分离方法。","inspiration":"该研究采用欺骗自由设计交叉告知信息与实际对手，分离信息效应与对手效应，方法上可借鉴用于经济实验中标签或框架效应的因果识别。｜可迁移到算法定价或自动竞价场景，研究披露算法身份对市场参与者竞争行为的影响。｜设计实验：招募人类被试参与模拟拍卖或议价博弈，随机告知对手为人类或算法（实际对手固定为算法或人类），测量出价或报价行为，并与真实市场交易数据（如电商平台竞价记录）对照。"}},{"id":"2609.16374","version":2,"title":"When a Story Feels Like Mine: How Personalized Narratives and Humor Shape Older Adults' Empathy toward LLM-Generated Peer Health Stories","zh_title":"当故事感觉像我的：个性化叙事与幽默如何塑造老年人对LLM生成同伴健康故事的同理心","abstract":"Peer stories have been shown to boost self-efficacy in older adults' health behavior change. Despite their effectiveness, peer stories are difficult to deploy in health promotion at scale given the difficulty of matching the diverse health concerns and coping styles of heterogeneous older populations. Large language models (LLMs) have been shown to generate authentic narratives, yet how personalization and narrative affective style, such as humor, jointly shape older adults' responses remains unknown. We developed a theory-driven system that generates first-person peer health narratives varying in personalization and humor through a three-stage LLM pipeline grounded in self-efficacy mechanisms. Thirty-one older adults were invited to participate in a within-subjects lab study. Results showed that personalization increased perceived relatability and relevance of peer stories, especially for older adults with lower humor preference. These findings position individual differences in affective styles as a second dimension in designing personalization for LLM-assisted health communication.","authors":["Kexin Quan","Precious Olalere","Smit Desai","Jessie Chin"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-16","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.16374","pdf_url":"https://arxiv.org/pdf/2609.16374","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM生成叙事","个性化健康传播","老年人用户研究"],"reason":"用LLM生成健康叙事并测量老年人反应，有真实人类被试对照，但非直接仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":7,"question":"个性化与幽默如何共同影响老年人对LLM生成的同伴健康故事的共情及相关感知？","design":"本研究并非直接仿真人类被试，而是用LLM（gpt-4o-mini）生成第一人称同伴健康叙事，通过三阶段流水线实现个性化（基于参与者自报的健康挑战、目标、障碍和应对策略）和幽默（亲和、应对导向）的操纵。31名社区老年人参与2（个性化）×2（幽默）的被试内实验，评估对故事的共情、理解、相关性、真实性和偏好。","baseline":"无对照（未使用真实人类数据作为基准，而是直接测量人类被试对LLM生成内容的反应）。","findings":"个性化显著提高了老年人对故事的感知相关性和共鸣，尤其对幽默偏好较低的个体效果更强。幽默本身并未增强共情或相关性，但可能增加故事的吸引力。","reliability":"论文未讨论（节选内容未提及失效条件或局限）。","relevance":"该研究虽非直接仿真人类被试，但提供了LLM生成个性化叙事并测量人类反应的实验范式，对关注LLM在健康传播中应用的仿真研究者有参考价值，但缺乏真实人类数据对照，与经济学实验仿真关联较弱。","inspiration":"借鉴其个性化处理的设计：基于个体特征（如健康目标、障碍）动态生成内容，并操纵叙事风格（幽默）作为第二维度，采用被试内设计测量感知结果。｜可迁移到消费者金融决策中的个性化信息干预，如退休储蓄建议或健康保险选择，研究不同叙事风格对决策的影响。｜以真实投资者为被试，用LLM生成个性化投资建议（基于其风险偏好和财务目标），操纵叙事语气（幽默 vs. 严肃），测量投资意愿和风险感知，并与历史投资行为数据对照。"}},{"id":"2609.20846","version":1,"title":"Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models","zh_title":"奖励高效推理可改善推理模型在欠明确任务上的弃权行为","abstract":"While modern large reasoning models (LRMs) excel at providing correct answers in many tasks, we provide additional evidence for the observation that they often struggle with a critical capability: knowing when to abstain from answering. We analyze this gap by comparing LRM behavior to results from a human study, revealing that human reasoning effort on unanswerable tasks is upper-bounded by answerable tasks, whereas LRMs waste computational resources by generating longer Chains of Thought (CoTs) on unanswerable than on answerable prompts. To overcome this inefficiency, we take inspiration from a resource-rational perspective on human cognition and introduce a novel GRPO reward that encourages efficient reasoning about whether the task contains all the information needed to solve it. Fine-tuning several 4B LRMs with this reward leads to human-like abstention performance gains (+12.8% on average) while retaining answering capabilities and boosting the models' efficiency (44% shorter CoTs on average).","authors":["Polina Tsvilodub","Max H\\\"oth","Michael Franke","Bj\\\"orn Deiseroth","Carina Kauf"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20846","pdf_url":"https://arxiv.org/pdf/2609.20846","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM弃权行为","人类对照","可靠性评估"],"reason":"用人类数据对照，评估LLM在不确定任务上的弃权行为，并改进其与人类一致性，属于…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":8,"question":"大型推理模型（LRMs）在任务信息不足时是否无法像人类一样高效地弃权，以及如何通过奖励设计改进其弃权行为。","design":"该研究并非严格意义上的仿真实验，而是将多个LRMs（4B–32B，六个模型家族）在QuestBench和AbstentionBench上的弃权表现与人类行为进行对比，然后提出SURE奖励（结合结果奖励与过程奖励）对4B模型进行GRPO微调，测量弃权准确率和思维链长度。","baseline":"研究者自行开展了一项人类实验，让人类被试对可答与不可答问题进行二元判断，记录其准确率和推理努力（以回答时间或步数衡量），作为LRMs的对照基准。","findings":"LRMs在不可答任务上生成的思维链比可答任务更长，弃权准确率远低于人类；使用SURE奖励微调后，模型弃权准确率平均提升12.8%，思维链长度平均缩短44%，同时保持回答能力。","reliability":"论文承认SURE的过程奖励主要针对信息缺失类弃权，对于安全等其他弃权原因可能需要不同的过程监督；微调模型仅在有限数据集上评估，需在更多弃权数据集上验证；人类实验中被试在不可答问题上有时会基于特定假设作答，与模型的假设差异需进一步比较。","relevance":"该研究直接对比LLM与人类在不确定任务上的弃权行为，并利用人类数据改进模型，属于用LLM仿真人类决策并评估可靠性的工作，对关注经济学实验和政策评估中模型弃权行为的研究者有参考价值。","inspiration":"借鉴其将人类行为作为基准并设计奖励函数对齐人类效率的做法，可用于经济决策仿真中的理性疏忽或信息获取行为。｜可迁移到消费者在信息不完全时的购买决策、投资者面对模糊信息的资产配置、或政策公告下预期形成等场景。｜以LLM作为被试，设计信息缺失的经济决策任务（如模糊收益彩票），处理为是否提供额外信息或奖励函数中惩罚过度推理，结果变量为弃权率与推理步数，对照真实人类实验数据（如实验室彩票选择实验）。"}},{"id":"2609.21401","version":1,"title":"Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue","zh_title":"与机器对话的错位：人机对话中的道德、礼貌与对齐","abstract":"Conversational AI systems produce fluent, socially appropriate responses, yet whether they participate in cooperative communication or merely simulate its surface forms remains unclear - a question central to how these systems are evaluated, trusted, and designed. This study investigates how morality, politeness, and alignment - three dimensions central to cooperative dialogue - function in human-AI interaction compared to human-human conversation. We analyze 15,881 human-ChatGPT and 10,784 human-human multi-turn dialogues, using mixed-effects models to identify which features predict turn-to-turn alignment. We observe a consistent dissociation: AI produces the surface features of cooperative communication without the underlying social architecture. Moral output appears preconfigured rather than negotiated; warmth is generated without face sensitivity; linguistic convergence declines persistently. Most strikingly, the cooperative mechanisms themselves reverse direction: hedging and softening associated with greater accommodation between humans are associated with reduced alignment when produced by AI, and purity framing associated with human divergence coincides with users converging toward the AI. Agency - giving users room to shape the exchange - is the most consistent predictor of alignment across both interaction types, while lower moral assertiveness in more recent models is not accompanied by better cooperation. Together these patterns suggest that AI reproduces the surface of cooperation without the mutual adaptation that grounds it between humans - and, more surprisingly, that mechanisms sustaining human accommodation can run in reverse with AI, suggesting a turn-level view may be insufficient for interaction-level success.","authors":["Marina Mitiaeva","Lu Xiao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21401","pdf_url":"https://arxiv.org/pdf/2609.21401","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["人机对话","合作沟通","仿真偏差"],"reason":"用LLM与人类对话数据对照，分析合作沟通机制，揭示AI表面模仿但缺乏真实社会适…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":21,"question":"在人类与AI对话中，道德、礼貌与对齐三个合作沟通维度如何运作，与人类间对话相比有何差异？","design":"本研究不是仿真实验，而是对真实对话语料的计算分析：使用WildChat中15881条人类与ChatGPT多轮对话和10784条人类间多轮对话，用混合效应模型分析道德表达（MFT分数）、礼貌标记（32个标记）和逐轮对齐（词汇、句法、语用、情感四个通道的相似度）之间的关系。","baseline":"人类间对话数据（10784条多轮对话）作为对照基准。","findings":"AI产生合作沟通的表面特征但缺乏底层社会架构：道德输出是预配置而非协商的，温暖生成没有面子敏感性，语言趋同持续下降。合作机制方向反转：人类中与更大适应相关的模糊限制语和软化在AI产生时与对齐降低相关，纯洁框架在人类中与分歧相关但用户却向AI趋同。","reliability":"论文未讨论仿真失效条件，但指出AI的合作机制可能反向运行，提示基于回合的视角可能不足以评估交互层面的成功。","relevance":"该研究直接对比人类与AI对话中的合作机制，揭示AI表面模仿但缺乏真实社会适应，对理解LLM在交互中的行为偏差和可靠性有重要价值，值得阅读原文。","inspiration":"借鉴其大规模语料对照分析和多维度测量方法，可迁移到经济金融中的沟通与信任场景，如投资者与AI顾问的对话。｜设计一个实验：让人类被试与LLM扮演的金融顾问进行投资建议对话，测量道德语言、礼貌标记和语言对齐，并与人类顾问对话对照，结果变量为投资决策和信任度。"}},{"id":"2609.21149","version":1,"title":"Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake","zh_title":"面向AI辅助精神病学接诊的临床医生导向质量保证","abstract":"Before patients can use AI-assisted psychiatric intake systems, health systems need practical ways to routinely evaluate these tools against their clinical standards for quality assurance. Because clinicians may use different intake styles, evaluation for this task must (1) support comparison across interviewing approaches, (2) minimize clinician burden, and (3) measure clinically relevant performance for health systems deploying these technologies. We present a clinician-grounded evaluation platform built around a memory-augmented patient simulator for open-ended AI interviewing, InterviewPlayground. We created interactive patients using InterviewPlayground with our expert-authored vignettes, constructed a simulated intake platform for the interviews, and designed evaluation modalities relevant to intake. In a pilot of 6 clinicians in a 25-minute assessment compared to a GPT-based LLM intake interviewer, the LLM recovered more of the clinically relevant items embedded in the patient vignettes (88.0% vs. 38.9%), but made more clinical inferences not based on the interview (56.8% vs. 27.8%), and characterized identified safety concerns less often (33.3% vs. 66.7%), setting the stage for deployed quality assurance for this task.","authors":["King Shi","Amanda Li","Jonathan Ivey","Synthia Qia Wang","Guan Gui","Hyunseo Kim","Peter Zandi","Jason Straub","Jacob Taylor","Ananya Joshi"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21149","pdf_url":"https://arxiv.org/pdf/2609.21149","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM患者仿真","临床质量保证","人机对照"],"reason":"用LLM模拟患者进行精神病学访谈，并与临床医生对照，属于人类仿真且有人类数据基…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":20,"question":"如何为AI辅助精神科接诊系统建立一个以临床医生为基准、可重复且低负担的质量保证评估平台？","design":"使用记忆增强的患者模拟器InterviewPlayground，基于专家编写的病例摘要创建交互式合成患者；构建模拟接诊平台，让6名临床医生和一个基于GPT的LLM接诊员分别对合成患者进行25分钟访谈；通过回忆表单和转录分析测量信息恢复率、无依据推断比例和安全问题特征化率。","baseline":"6名执业临床医生在相同合成患者上的访谈表现作为人类基准。","findings":"LLM接诊员恢复了更多临床相关条目（88.0%对38.9%），但做出了更多无依据的临床推断（56.8%对27.8%），且更少特征化已识别的安全问题（33.3%对66.7%）。","reliability":"论文未讨论","relevance":"该研究用LLM模拟患者进行精神病学访谈，并与临床医生对照，属于人类仿真且有人类数据基准，但场景是医疗质量保证而非经济学实验，对关注经济学实验和政策评估的研究者参考价值有限。","inspiration":"借鉴其使用合成患者固定临床真相以支持跨风格比较和重复评估的方法，可迁移到消费者金融决策或信贷审批等需要标准化对话场景的经济学实验中，例如用LLM模拟不同特征的借款人，让人类信贷员和AI审批系统分别进行访谈，测量信息提取准确性和决策偏差，并与真实信贷数据对照。"}},{"id":"2609.20543","version":1,"title":"Language-model groups overstate consensus when replaying human deliberation on a reasoning task","zh_title":"语言模型群体在重放人类推理任务审议时高估共识","abstract":"Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.","authors":["Tengfei Shao"],"categories":["cs.AI","cs.CL","cs.CY","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20543","pdf_url":"https://arxiv.org/pdf/2609.20543","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人类对照","共识偏差"],"reason":"用LLM代理重放人类推理任务，并与真实人类数据对照，评估仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":3,"question":"用信念锚定的LLM代理组重放人类沃森推理任务讨论时，能否保持人类组的共识结构？","design":"用deepseek-v4-flash模型，为100个留出的人类沃森组中每个真实参与者实例化一个信念锚定代理（锚定其讨论前答案），交叉三种角色保真度、三个种子和两种推理模式，模拟小组讨论，并用与人类相同的代码对最终答案进行共识评分。","baseline":"DeliData语料库中100个真实人类沃森小组的讨论数据，包括参与者的讨论前答案、发言情况和最终答案。","findings":"LLM代理组过度达成完全共识，远高于人类组；在提交制比较中聊天和推理模式的共识差距分别为34.0和43.9个百分点，在参与匹配比较中分别为34.1和44.4个百分点。模拟共识与集体准确性无关，且信念锚定代理组是人类组结果分布的有偏估计量。","reliability":"论文承认代理组几乎总是发言，而人类约五分之一从不发言，导致参与结构不对称；通过提交制和参与匹配两种敏感性分析减少测量不对称，但差距仍存在。还指出在移除可记忆答案的重参数化下，推理模式组几乎一致同意但多为错误答案，表明模拟共识不追踪准确性。","relevance":"该研究直接针对LLM仿真人类群体决策的可靠性，提供了与真实人类数据对照的批判性证据，值得精读以了解仿真在共识测量上的系统性偏差。","inspiration":"借鉴其信念锚定和评分显式化的设计，将LLM代理组与真实人类组在相同任务和评分规则下比较，以识别仿真偏差。｜可迁移到经济金融中的群体决策场景，如投资委员会讨论、信贷审批小组或政策预期形成实验。｜用LLM代理扮演真实实验中的被试，锚定其初始判断，模拟小组讨论后测量共识率和决策准确性，并与原始人类实验数据对照，检验仿真是否高估共识并扭曲结果分布。"}},{"id":"2609.19913","version":1,"title":"Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social Networks","zh_title":"意见动态的数字孪生：面向社交网络的生成式LLM框架","abstract":"The study of opinion dynamics in social networks is one of the key challenges in computational social science with direct relevance to understanding political polarization, misinformation, and health responses. Current approaches focus on simplified mathematical models that ignore linguistic and contextual factors related to belief updates or use Large Language Model (LLM)-based simulations that have not been validated against real data. We present a framework based on the concept of a digital twin to simulate opinion dynamics in social networks. The approach fills the gap by cloning a real-world Twitter network, assigns a set of attributes for agents (such as persona, emotions, centrality, stubbornness, and influence), and employs Mistral-7B to perform opinion update based on memory and social exposure. To evaluate the proposed approach, we validate it against two real Twitter datasets (COVID-19 discourse and U.S elections 2020). The results show that the capability of the proposed framework reproduces opinion trajectories and reduces individual prediction error by more than 50% compared to the best-performing classical baseline (Mistral-7B achieves Mean Absolute Error (MAE) = 0.150 and 0.121 on the COVID-19 and US Election 2020 datasets, respectively). We observe similar improvements in structural alignment (Delta_r = 0.120 and 0.180) and polarization dynamics (Delta_Var = 0.106 and 0.115) on the two datasets, respectively. Additionally, the ablation studies confirm that agent attributes, memory, and social exposure all contribute to the framework's predictive fidelity in reproducing opinion trajectories, with agent attributes being the most critical contributor. Overall, our results demonstrate that grounding Mistral-7B within empirically cloned interaction networks produces a realistic simulation framework capable of reproducing complex social dynamics.","authors":["Omran Berjawi","Giuseppe Fenza","Rida Khatoun","Sherali Zeadally"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19913","pdf_url":"https://arxiv.org/pdf/2609.19913","source_feed":"cs.LG","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","意见动态","数字孪生"],"reason":"用LLM仿真社交网络意见动态，并与真实Twitter数据对照，直接复现人类行为。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":2,"question":"如何利用基于真实社交网络克隆的数字孪生框架，结合LLM代理的认知与语言能力，更准确地模拟和预测社会网络中的意见动态？","design":"使用Mistral-7B作为代理，在从真实Twitter网络克隆的交互结构上运行，每个代理被赋予从真实用户数据提取的属性（如人设、情绪、中心性、固执度、影响力），并基于记忆和社交暴露进行意见更新，以预测个体意见轨迹和集体极化动态。","baseline":"两个真实Twitter数据集：COVID-19话语数据集和美国2020年大选数据集，包含推文内容和用户级元数据，用于验证框架的预测准确性。","findings":"该框架能重现意见轨迹，个体预测误差比最佳经典基线降低50%以上（COVID-19上MAE=0.150，美国大选上MAE=0.121）。在结构对齐和极化动态上也观察到类似改进，消融研究表明代理属性、记忆和社交暴露均对预测保真度有贡献，其中代理属性最关键。","reliability":"论文未讨论","relevance":"该研究直接使用LLM代理在真实社交网络上复现人类意见动态，并与真实Twitter数据对照，属于高相关性的仿真验证工作，值得阅读原文以了解其具体实现和验证细节。","inspiration":"借鉴其将真实网络结构、用户属性与LLM认知过程结合的方法，并利用消融实验识别关键因素。｜可迁移到金融市场情绪传播、政策公告的预期形成或消费者信心扩散等场景。｜以真实投资者社交网络（如StockTwits）为底，用LLM代理扮演投资者并赋予真实用户特征，施加政策新闻或市场事件处理，测量个体情绪和交易倾向变化，并与真实历史数据对照验证。"}},{"id":"2607.25667","version":2,"title":"MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice","zh_title":"MyMentorLLM：用于刻意练习的多模态语音/文本患者、受训者与专家心理治疗GenAI环境","abstract":"Psychotherapists need repeated training and supervision; however, scalability is problematic. We present MyMentorLLM, a multimodal voice- and text-based deliberate-practice environment with 2,100 complete Cognitive Behavioural Therapy (CBT) sessions. Each session links a DSM-5-TR-grounded LLM patient (with major depressive, generalised anxiety or borderline personality disorder), an LLM therapist-in-training and an LLM expert supervisor (powered by Gemma-4, Gemini-3.1-Flash-Live and Qwen-3.6). Sessions were analysed for emotional dynamics, therapeutic competence and diagnostic accuracy against human psychotherapy data. Simulated patients expressed disorder-congruent emotional profiles, which therapists mirrored as in human counselling. LLM trainee competence was rated above human levels in most conditions, while native speech-to-speech was closest to human scores. Supervisor feedback improved diagnostic accuracy in 5 of 7 LLM conditions, whereas symptom identification accuracy increased with model size. This work shows deliberate practice can be simulated for CBT training, although patient fidelity, supervisor calibration and harmful feedback require evaluation via a complex systems perspective.","authors":["Rodolfo Rizzi","Alessandro Grecucci","Massimo Stella"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-18","first_seen":"2026-07-29","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2607.25667","pdf_url":"https://arxiv.org/pdf/2607.25667","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","心理治疗","人类数据对照"],"reason":"用LLM模拟患者和治疗师，并与人类心理治疗数据对照，评估仿真可靠性","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":4,"question":"能否构建一个多模态的LLM心理治疗刻意练习环境，模拟患者、受训治疗师和专家督导，并验证其心理保真度、治疗能力和诊断准确性？","design":"使用Gemma-4、Gemini-3.1-Flash-Live和Qwen-3.6等LLM分别扮演DSM-5-TR诊断的抑郁症、广泛性焦虑症和边缘型人格障碍患者、受训治疗师和专家督导，进行2100次完整的认知行为疗法（CBT）会话，支持语音和文本两种模态；分析会话中的情绪动态、治疗能力和诊断准确性，并与人类心理治疗数据对照。","baseline":"人类心理治疗数据，包括人类咨询中的情绪镜像模式、人类治疗师能力评分和诊断准确性。","findings":"模拟患者表现出与疾病一致的情绪特征，治疗师像人类咨询中一样镜像了这些情绪；LLM受训治疗师的能力评分在多数条件下高于人类水平，其中原生语音到语音模式最接近人类分数。","reliability":"论文承认患者保真度、督导校准和有害反馈需要通过复杂系统视角进行评估；LLM能力评分可能虚高，且文本模态可能丢失副语言信息。","relevance":"该研究用LLM模拟患者和治疗师，并与人类心理治疗数据对照，评估仿真可靠性，直接回应了研究者对LLM人类仿真实验和对照基准的关注，值得精读原文。","inspiration":"借鉴其多角色LLM仿真和与人类基准对照的设计，可用于经济金融中的专业服务场景仿真。｜可迁移到金融咨询或信贷审批中的客户-顾问互动仿真，评估LLM顾问的行为偏差和决策质量。｜以LLM扮演金融顾问和客户，施加不同市场条件或客户特征处理，测量顾问建议的风险偏好和客户满意度，并与人类金融咨询记录或实验数据对照。"}},{"id":"2609.19843","version":1,"title":"A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents","zh_title":"基于LLM的GUI代理对助推易感性的双过程视角研究","abstract":"LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users. While behavioural biases in the textual outputs of LLMs are well-documented, far less is known about how such influence operates when models act as agents that perceive interfaces and execute decisions---and, in particular, whether the reasoning capabilities increasingly built into these agents make them more robust to it. Drawing on Dual-Process Theory, we empirically investigate whether LLM-based GUI agents are susceptible to automatic (Type 1) and reflective (Type 2) digital nudges, and how their reasoning configuration moderates this susceptibility. In a randomized online shopping experiment with 3,600 agents and a total of 21,600 simulations across six frontier models from three providers, we found that agents were vulnerable to both nudge types. Crucially, the reasoning configuration moderated these effects in opposing directions, reducing susceptibility to automatic default nudges while heightening it to reflective social influence nudges. Extensive reasoning therefore did not make agents more robust but redirected the route through which choice architecture takes effect. Exploratory analysis further showed this redirection to be systematically structured by model scale. Beyond establishing nudge susceptibility as a behavioural property of agentic AI, the study positions interface design as a governance concern for organizations that delegate decisions to autonomous agents.","authors":["Haya Halimeh","Sascha Kaltenpoth","Kevin B\\\"osch","Oliver M\\\"uller"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19843","pdf_url":"https://arxiv.org/pdf/2609.19843","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为经济学","助推"],"reason":"用LLM GUI代理模拟人类在数字环境中的决策，研究助推易感性，并与人类行为理…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":5,"question":"LLM-based GUI agents 在数字环境中是否容易受到自动型（Type 1）和反思型（Type 2）数字助推的影响，以及推理配置如何调节这种易感性？","design":"使用 6 个前沿 LLM（来自 3 家提供商）构建 GUI agents，在模拟在线购物环境中进行随机实验，共 3600 个 agents、21600 次模拟。通过改变界面设计施加两类助推：默认选项（Type 1）和社会影响信息（Type 2），并操纵推理配置（低 vs. 高）。结果变量为 agents 的购买选择。","baseline":"无对照（论文未使用真实人类数据作为基准，而是直接测量 agents 的行为）。","findings":"LLM-based GUI agents 对两类助推都表现出易感性。推理配置对两类助推的易感性有相反方向的调节作用：高推理降低了默认助推的易感性，但增加了社会影响助推的易感性。","reliability":"论文未讨论","relevance":"该研究直接使用 LLM agents 模拟人类在数字环境中的决策行为，并检验助推易感性，属于用 LLM 进行人类仿真实验的范畴，但缺乏真实人类数据对照，因此对关注基准对照的研究者价值有限。","inspiration":"值得借鉴的是通过 GUI 环境施加助推并操纵推理配置来研究认知过程对行为偏差的影响。｜可以迁移到消费者在线购物决策、金融产品选择（如默认投资选项、社会影响信息对投资决策的影响）等场景。｜设计雏形：使用 LLM-based GUI agents 模拟投资者在金融平台上的选择，处理为默认投资组合（Type 1）或社会证明信息（Type 2），结果变量为投资选择，并与真实投资者行为数据（如实验或交易记录）进行对照。"}},{"id":"2609.20055","version":1,"title":"What People Almost Did: Evaluating LLM Social Simulations Beyond Behavioral Fit","zh_title":"人们几乎做了什么：超越行为拟合评估LLM社会仿真","abstract":"LLM-based social simulations are primarily evaluated for behavioral fit, testing whether agents reproduce the actions or response distributions of the people they are simulating. However, the promise of simulation extends beyond behavioral fit. Simulations can explain human behavior, diagnose barriers, and compare large-scale interventions. These use cases depend on understanding \\textit{why} people acted a certain way, not just \\textit{what} they did. As a result, behavioral fit is insufficient for these types of claims because behavior underdetermines the reasoning process behind it. For instance, the behavior of staying silent may be due to disinterest or suppressed speech, and not answering a call may be due to distrust of the caller or limited phone access. In this paper, we propose \\textit{representational adequacy} as a new evaluation target for LLM-based social simulations. By leveraging LLM reasoning traces, representational adequacy measures whether a simulation's scenario--reasoning--action triples preserve the reasoning process behind the behavior in a way that is faithful to the population and scenarios being simulated. We distinguish representational adequacy from interpretability and alignment metrics, propose ways to integrate it into simulation research, and pose its measurement as an open problem.","authors":["JaeWon Kim","Angie Boggust"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20055","pdf_url":"https://arxiv.org/pdf/2609.20055","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A2","A4","B4"],"tags":["LLM社会仿真","评估方法","表征充分性"],"reason":"提出表征充分性评估LLM社会仿真，超越行为拟合，关注推理过程，直接针对仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":6,"question":"如何评估基于大语言模型的社会仿真是否保留了被仿真人群的推理过程，而不仅仅是行为结果？","design":"本文不是一项仿真实验研究，而是一篇观点/框架论文。作者提出“表征充分性”作为新的评估目标，主张通过分析LLM在仿真中产生的“场景—推理—行动”三元组，来评估仿真是否忠实于被仿真人群的推理过程。文中没有具体实施仿真或施加处理，而是通过两个例子（青少年社交媒体发帖、孕妇接听健康电话）说明行为相同但推理不同的情况，并讨论如何将表征充分性整合到仿真研究中。","baseline":"无对照","findings":"行为拟合不足以评估用于解释人类行为的LLM社会仿真，因为相同行为可能源于不同推理过程。作者提出表征充分性框架，通过分析场景—推理—行动三元组来评估仿真是否保留了被仿真人群的推理过程，并将其与可解释性和对齐指标区分开来。","reliability":"论文未讨论","relevance":"该文直接针对LLM社会仿真的可靠性问题，提出超越行为拟合的评估框架，强调推理过程的重要性，与研究者关注的仿真可靠性与偏差高度相关，值得阅读原文以了解其理论框架和测量挑战。","inspiration":"本文提出的表征充分性概念可借鉴用于经济金融仿真研究，通过分析LLM代理的推理痕迹来评估其决策过程是否与真实人群一致。｜该框架可迁移到政策评估场景，例如模拟消费者对财政刺激的反应或投资者对央行公告的预期形成，这些场景中行为相同但动机可能不同。｜一个可行的研究设计是：用LLM代理模拟投资者，施加不同措辞的央行公告作为处理，记录代理的推理痕迹和投资决策，并与真实投资者在类似实验中的推理和决策数据（如调查或实验数据）进行对照，评估表征充分性。"}},{"id":"2609.19866","version":1,"title":"Reproducibility is not construct validity: LLM measurement of institutionally situated communication","zh_title":"可重复性不等于构念效度：LLM对制度情境沟通的测量","abstract":"High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({\\=g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.","authors":["Veronika Batzdorfer (KIT)","Carlo Romano Marcello Alessandro Santagiustina (ALMAnaCH, m\\'edialab, Sciences Po)"],"categories":["cs.AI","cs.CL","cs.CY","q-fin.RM"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19866","pdf_url":"https://arxiv.org/pdf/2609.19866","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM测量效度","构念效度","人类数据对照"],"reason":"评估LLM测量效度，区分可重复性与构念效度，有真实人类调查对照，批判性指出失效…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":7,"question":"LLM 从公共咨询文本中推断的构念测量与同一利益相关者在结构化调查中自报的构念测量在多大程度上一致？","design":"本研究不是仿真实验，而是测量效度研究。使用欧洲委员会 AI 法案咨询数据集，将同一利益相关者的自由文本咨询意见与结构化调查回答相链接。用 LLM 对咨询文本进行标注，得到 AI 安全、权利和可解释性关注度的文本测量，并与调查自报的对应构念测量进行比较。","baseline":"同一利益相关者在结构化调查中的自报回答，作为构念测量的基准。","findings":"LLM 标注具有高可重复性（ICC>0.99），但与调查测量收敛性弱（相关系数 0.029-0.176，Lin 一致性系数<0.08）。文本与调查的差异在不同利益相关者群体间存在系统性模式：商业协会在文本中表达的 AI 风险关注高于调查（效应量+1.0），而公共机构和非商业群体差异较小或为负。","reliability":"论文承认 LLM 文本测量可能捕捉的是机构角色和沟通情境，而非纯粹的潜在构念；国家层面的空间自相关结果对多重检验校正敏感，应谨慎解释。","relevance":"该研究直接回应了 LLM 仿真中可重复性与构念效度的区分问题，提供了真实人类调查对照，并批判性地展示了 LLM 测量在机构情境文本中的失效条件，值得精读。","inspiration":"借鉴其将同一主体的不同沟通渠道（调查与公开文本）进行链接并比较 LLM 测量与自报测量的设计，以检验测量效度。｜可迁移到经济金融中的政策沟通研究，例如央行沟通文本与市场参与者调查预期的比较，或上市公司年报文本与分析师调查预期的比较。｜以机构投资者为被试，收集其对央行政策声明的公开评论和内部调查回答，用 LLM 从公开评论中测量政策预期，与调查自报预期对比，并以市场利率变动作为外部效标，检验 LLM 测量在机构沟通情境下的构念效度。"}},{"id":"2609.16366","version":2,"title":"How Humans and LLMs Read Gender into \"Gender-Neutral\" Physical Descriptions","zh_title":"人类与LLM如何将性别读入“性别中立”的物理描述","abstract":"When foundation models describe people, recent work in AI fairness, accessibility, and ethics recommends avoiding inferred identity labels (e.g., \"she\", \"his\") in favor of seemingly \"objective\" physical descriptions (e.g., \"short hair\", \"a defined jawline\"). Yet whether such descriptive language achieves gender-neutral communication remains an open empirical question. To study this, we introduce GAPA (Gender Associations of Physical Attributes), a dataset of 316 common physical attributes drawn from diverse sources, paired with 14,706 gender-association ratings from 304 US-based annotators. Results show that physical descriptions carry structured and graded gender associations among readers, with more consistent and distinctive associations for women and men than for non-binary identities. Next, we evaluate 16 LLMs across model families, sizes, and post-training variants against human ratings. The models partially recover human associations but exhibit systematic alignment biases, including compressed rating distributions, weaker alignment for associations with men, and asymmetric abstention that disproportionately targets the non-binary category. Finally, we release the best-performing proxy model trained to predict humans' gender associations of descriptive language and demonstrate its utility through a sociolinguistic analysis of character descriptions in LitBank. Together, our findings provide the first empirical evidence that seemingly \"objective\" physical descriptions can retain systematic gender associations in human interpretation, and uncover systematic patterns of model-human misalignment. This challenges the assumption that replacing explicit gender labels with physical descriptions necessarily yields gender-neutral communication, and highlights downstream challenges in using such descriptions to communicate subjective identity categories in human-AI interaction.","authors":["Yingjia Wan","Lin Lin","Elisa Kreiss"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-18","first_seen":"2026-09-16","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.16366","pdf_url":"https://arxiv.org/pdf/2609.16366","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","人类对照","性别联想"],"reason":"评估LLM与人类性别联想的一致性，有真实人类数据对照，揭示模型偏差，可迁移到仿…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":8,"question":"看似“性别中立”的物理特征描述在人类读者和LLM中是否仍携带系统性的性别联想？","design":"本研究并非将LLM作为人类被试的替代品进行仿真实验，而是构建了GAPA数据集（316个物理属性描述，来自LLM生成、人类启发和当代小说），收集304名美国标注者对每个属性与女性、男性、非二元性别关联的评分（共14,706条），然后评估16个LLM对这些属性的性别联想评分与人类评分的一致性。","baseline":"304名美国标注者提供的14,706条性别关联评分，作为人类基准。","findings":"物理描述在人类读者中携带结构化且分级的性别联想，53%的属性显著更偏向某一性别，且对女性和男性的联想比对非二元性别更一致和鲜明。LLM仅部分恢复人类联想模式，存在评分分布压缩、对男性联想对齐较弱、以及对非二元类别的不对称弃权等系统性偏差。","reliability":"论文指出LLM与人类存在系统性偏差，包括评分分布压缩、对男性联想对齐较弱、以及指令微调和专有模型对非二元类别的不对称弃权，这些偏差可能影响下游应用。但论文未明确讨论在何种条件下仿真会完全失效。","relevance":"该研究直接评估LLM与人类在性别联想上的对齐程度，有真实人类数据对照，并揭示了模型偏差，对关注LLM仿真可靠性和偏差的研究者具有参考价值，值得阅读原文以了解具体偏差模式和测量方法。","inspiration":"借鉴其构建属性词表并收集人类评分作为基准，再评估LLM对齐程度的方法，可迁移到经济金融中的文本描述歧视研究，如信贷审批中的申请人描述或招聘广告中的语言。设计：收集一组描述个人特征的中性词汇（如“有纹身”、“戴眼镜”），让人类被试和LLM分别评估这些词汇与某些经济结果（如信用风险、工作能力）的关联，比较LLM评分与人类评分的分布差异，并检验LLM是否在某些类别上出现系统性偏差或弃权。"}},{"id":"2609.16432","version":2,"title":"A light-touch AI literacy intervention helps protect against AI political persuasion","zh_title":"轻触式AI素养干预有助于抵御AI政治说服","abstract":"Conversations with large language models (LLMs) can substantially shift beliefs and attitudes, raising concerns about manipulation using AI persuasion. Here we test whether a light-touch AI literacy intervention - a brief warning that LLMs can be prompted to persuade and may present information selectively - helps protect users. Across two experiments (total N = 3,208 Americans) in which participants conversed with an LLM instructed to shift their views about different political topics, the presence of a warning reduced belief change by roughly one-half (-48.1%, 95% CI [-59.5%, -36.8%]) relative to the control. Importantly, the warning did not significantly reduce trust in generative AI more broadly. Light-touch literacy interventions can help protect users against AI political persuasion.","authors":["Reed Orchinik","David Rand"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-18","first_seen":"2026-09-16","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.16432","pdf_url":"https://arxiv.org/pdf/2609.16432","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["AI说服","态度改变","干预实验"],"reason":"用LLM与真人对话测量态度改变，有真实人类对照，但非替代被试仿真","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":9,"question":"轻量级AI素养干预（提示LLM可能被用于说服且信息选择性呈现）能否降低用户在政治议题对话中的态度改变？","design":"两项在线实验（总N=3,208名美国人），参与者先报告政治议题初始态度，然后与一个被指示说服用户的LLM对话，之后再次测量态度。研究1使用GPT-4.1讨论住房政策，随机分配事实型或情感型说服策略；研究2使用Grok 4.5讨论从ANES选取的15个议题之一。干预为在对话前显示一般性警告（提示AI可能被操纵说服）或特定警告（额外披露模型论证方向），对照组无警告。结果变量为态度改变（研究1还包括激励性捐款决策）。","baseline":"无对照（无真实人类说服者作为基准，仅比较警告组与无警告对照组的态度改变）。","findings":"元分析显示，警告使说服效果降低约48.1%（95% CI [-59.5%, -36.8%]），且一般警告与特定警告效果无显著差异。警告未显著降低对生成式AI的整体信任。","reliability":"论文未讨论失效条件，但指出干预不能完全消除AI说服效果，且研究聚焦于政治议题，未来需探索对亲社会说服的影响及更有效的传递方式。","relevance":"该研究直接测量LLM对话对真实人类态度改变的影响，并测试干预效果，虽非用LLM替代人类被试，但提供了LLM说服力及防御措施的因果证据，对评估LLM在实验中的行为影响有参考价值。","inspiration":"借鉴其随机化警告干预和态度前后测设计，可迁移到经济金融领域的AI建议场景（如投资建议、信贷决策），研究设计：招募真实投资者，随机分配是否接受AI素养警告，然后与LLM讨论某股票或资产配置，测量投资态度或风险偏好变化，并与无AI建议的人类对照组或历史数据比较。"}},{"id":"2609.19596","version":1,"title":"Full-Duplex Speech Models Take the Floor When Asked, Not When Needed","zh_title":"全双工语音模型在被要求时才发言，而非在需要时","abstract":"Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak. Human listeners speak when addressed or when the speaker stops, but also self-select to correct a false claim, supply a missing word, or warn of danger. We ask whether full-duplex models do the same. To separate the reason to speak from the opportunity, we construct context-matched English monologues in which only the trigger utterance varies within a topic, define 10 conditions from turn-allocation rules, and compress inter-word pauses to limit opportunities created by silence. Across five model families, being addressed and silence are far more reliable triggers than false facts or hazards. Frame-level text-token probabilities in Moshi and PersonaPlex are lower for false facts than for Neutral when averaged over the first 2\\,s after trigger end. Pauses or permission to interrupt do not close this gap either. Given the floor, Moshi and PersonaPlex answer most direct questions, yet the proportion of non-empty false-fact replies that challenge the claim is only .14--.15, and the proportion of hazard replies that warn of danger is .04--.07. This paper thus identifies a gap in both speech initiation and response content. Closing it requires genuine content understanding and intervention decisions grounded in it.","authors":["Linkai Peng","Baorian Nuchged","Kaiqi Fu","Yuyang Yao"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19596","pdf_url":"https://arxiv.org/pdf/2609.19596","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["语音模型","人类行为对照","主动发言"],"reason":"评估全双工语音模型在对话中主动发言的触发条件，与人类行为对照，发现模型在内容理…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":10,"question":"全双工语音模型是否会在未被直接提问或出现停顿时，基于内容理解主动发言（如纠正错误、警告危险）？","design":"构建40段英语独白，每段仅在触发句上变化，设置10种条件（直接提问、错误事实、危险警告、沉默等），压缩词间停顿以限制停顿机会，测试5个模型家族7种配置的发言起始率和回复内容。","baseline":"无对照","findings":"模型主要对直接提问和沉默做出反应，对错误事实和危险警告的主动发言率与中性条件相近；即使发言，纠正错误和警告危险的比例也很低（0.14-0.15和0.04-0.07）。","reliability":"论文未讨论","relevance":"该研究评估了LLM在对话中主动干预的能力，发现模型缺乏基于内容理解的发言决策，这对使用LLM模拟人类对话行为（如纠正、警告）的可靠性提出了质疑，值得阅读原文了解具体实验设计和失败模式。","inspiration":"借鉴其通过匹配上下文、仅改变触发句来分离发言原因与机会的实验设计，以及用帧级概率和发言起始率作为结果变量的测量方法。｜可迁移到经济金融场景中需要主动干预的对话任务，如智能投顾在客户陈述错误投资观念时是否主动纠正，或政策沟通中助手是否及时澄清误解。｜以LLM作为被试，构建包含错误金融陈述或风险提示的对话，测量其主动发言率和纠正内容，并与人类顾问在相同对话中的行为进行对照。"}},{"id":"2609.19965","version":1,"title":"Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence","zh_title":"逮捕之前：基于不完整证据的LLM犯罪画像基准测试","abstract":"Large Language Models (LLMs) are increasingly applied to legal and criminal justice tasks, yet existing work focuses almost exclusively on post-arrest scenarios where the suspect's identity is already known, leaving the critical pre-arrest challenge of inferring suspect characteristics from incomplete evidence largely unexplored. To fill this gap, we introduce the Profiling, Investigation, and Judgment (PIJ), comprising 2,500 real homicide cases from five countries. PIJ evaluates LLMs across three tasks that span the entire criminal investigation pipeline: criminal profiling, which requires abductive reasoning to infer suspect attributes from fragmentary scene evidence, crime process reconstruction, which tests structured information extraction, and sentence prediction, which demands legal deductive reasoning. We evaluate 9 powerful LLMs and find that performance degrades systematically as tasks shift from explicit fact extraction to implicit reasoning over unknown suspect profiles. Categories requiring inferential reasoning, such as motivation and victim-offender relationships, remain the primary bottlenecks. Further analysis reveals substantial gaps between LLMs and human experts, along with pervasive biases in gender, age, and motive attribution. Our findings indicate that pre-arrest inference from incomplete evidence remains an open challenge.","authors":["Yutong Yao","Yanjie Cao","Guanhua Chen","Xu Yang","Junchao Wu","Zeyu Wu","Lidia S. Chao","Derek F. Wong"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19965","pdf_url":"https://arxiv.org/pdf/2609.19965","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","犯罪画像","人类对照"],"reason":"用LLM模拟人类专家进行犯罪画像，并与人类专家对照，评估偏差，可迁移到人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":12,"question":"LLM能否在逮捕前阶段从零散证据中推断嫌疑人特征（犯罪画像），并与人类专家表现相比如何？","design":"构建包含2500个真实谋杀案的PIJ基准，评估9个LLM在三个任务上的表现：犯罪画像（溯因推理推断嫌疑人属性）、犯罪过程重建（结构化信息提取）、量刑预测（法律演绎推理）。","baseline":"与人类专家进行对照，但论文未提供人类专家数据的具体来源和规模。","findings":"LLM在信息提取任务上表现尚可，但在需要溯因推理的犯罪画像任务上性能大幅下降，动机和受害者-犯罪者关系等推理类别是主要瓶颈。LLM与人类专家存在显著差距，并在性别、年龄和动机归因上表现出系统性偏差。","reliability":"论文承认LLM在需要推理的类别上表现不佳，且存在性别、年龄和动机归因偏差，但未详细讨论失效条件或局限。","relevance":"该研究将LLM作为人类专家替代品进行犯罪画像，并与人类专家对照，评估偏差，与研究者关注的人类仿真实验高度相关，值得阅读原文以了解其基准设计和偏差分析方法。","inspiration":"借鉴其构建真实案例基准和与人类专家对照的方法，可迁移到经济金融领域的专家判断仿真，如信贷审批、保险理赔欺诈检测等场景。｜可迁移到信贷审批中的欺诈检测或风险画像，用LLM模拟信贷员从有限信息推断申请人风险特征。｜设计：用LLM扮演信贷审批员，输入脱敏的贷款申请信息（收入、职业、信用记录片段），输出风险画像（违约概率、欺诈可能性），与真实信贷员决策数据对照，评估准确性和偏差。"}},{"id":"2609.20005","version":1,"title":"Geopolitical Divisions Across Languages in Large Language Models","zh_title":"大语言模型中的跨语言地缘政治分歧","abstract":"People increasingly turn to AI chatbots for news and explanations of world events. But do they receive the same political answers when they ask in different languages? Here we show that the language of a question can change how the same AI systems assess the war in Ukraine. We ask GPT, Claude and Gemini to evaluate twenty statements about the war in 112 languages, collecting 67,200 responses. The balance between Russia-leaning and Ukraine-leaning responses differs across languages. When we group responses by countries' official languages, they follow a pattern resembling worldwide political divisions: relatively more Russia-leaning answers correspond to more favourable public views of Russia, less support for Ukraine in United Nations votes, and less aid to Ukraine. The broad pattern recurs across all three models and remains when individual statement pairs are removed. Our findings suggest a possible route through which information warfare may shape the text used to train AI models, which may in turn spread geopolitical biases.","authors":["Maxim Chupilkin"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20005","pdf_url":"https://arxiv.org/pdf/2609.20005","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","地缘政治偏见","跨语言差异"],"reason":"用LLM回答战争问题，并与真实国家态度数据对照，揭示语言导致的偏差，可迁移到仿…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":13,"question":"用不同语言向同一大语言模型提问俄乌战争相关问题，回答是否会呈现地缘政治分歧？","design":"用 GPT、Claude、Gemini 三个模型，以 112 种语言评估 20 条关于俄乌战争的陈述（10 条亲俄、10 条亲乌），收集 67200 个回答，计算亲俄与亲乌陈述同意度的差值作为语言条件化的政治倾向平衡分数。","baseline":"对照的真实人类数据包括：各国公众对俄罗斯的好感度、联合国投票中对乌克兰的支持度、对乌克兰的双边援助占 GDP 比重。","findings":"不同语言下模型回答的亲俄/亲乌倾向存在显著差异，乌克兰语最亲乌，俄语相对亲俄，且差异不限于交战双方语言。按国家官方语言聚合后，模型回答的倾向与真实世界政治分歧一致：更亲俄的回答对应更亲俄的公众舆论、更少的联合国支持乌克兰投票和更少的对乌援助。","reliability":"论文未讨论","relevance":"该研究用 LLM 模拟多国公众对地缘政治议题的态度，并与真实国家层面数据对照，揭示了语言条件化偏差，对关注仿真可靠性和偏差的研究者有直接参考价值。","inspiration":"可借鉴其用语言作为处理变量、以真实国家数据为基准的对照设计，以及通过多模型和稳健性检验增强结论可信度的做法。｜可迁移到跨国经济态度或政策偏好仿真，例如不同语言下询问对全球化、贸易保护、移民经济影响等问题的看法，检验语言是否导致系统性偏差。｜以 LLM 为被试，用不同语言呈现关于贸易政策、财政刺激或通胀预期的问卷，测量其态度倾向，并与世界价值观调查、国际社会调查项目等真实跨国态度数据对照，评估语言条件化仿真的外部效度。"}},{"id":"2609.20077","version":1,"title":"Tailored to you: longitudinal effects of personalising language models","zh_title":"为你量身定制：个性化语言模型的纵向效应","abstract":"Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people's perception of and behaviour toward AI remain poorly understood. Most critically, downstream consequences outside the immediate human--AI interaction loop, such as effects on users' self-perceptions and interpersonal relationships, remain largely unexamined. In this study, we recruited 992 participants to complete daily advice-seeking interactions with language models over the course of five days, comparing outcomes from a non-personalised baseline against two personalisation approaches: memory-based (conditioned on prior conversational history) and survey-based (conditioned on information collected through a pre-study intake survey). We find that several changes in human-AI interaction over time are driven primarily by repeated exposure rather than personalisation itself. However, participants interacting with personalised models experienced differences in advice-seeking and information-sharing attitudes and behaviours: participants in the memory-based condition engaged in greater self-disclosure and rated the model as less creepy, while participants in the survey-based condition reported higher regret about having shared personal information with the AI. We conclude by highlighting the nuanced effects of different personalisation approaches on interaction outcomes, and discussing the implications of these findings for the responsible design and deployment of personalised AI systems.","authors":["Canfer Akbulut","Justine Breuch","Arianna Manzini","Lujain Ibrahim","Matija Franklin","Roma Patel","Iason Gabriel","Kristian Lum","Laura Weidinger"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20077","pdf_url":"https://arxiv.org/pdf/2609.20077","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["个性化语言模型","人机交互实验","行为测量"],"reason":"用LLM与人类交互实验，测量行为变化，有真实人类数据对照，但非替代被试仿真","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":14,"question":"个性化语言模型（记忆型与问卷型）在持续五天的建议寻求互动中，如何影响用户对AI的感知、自我披露、建议采纳及对人际建议的态度？","design":"本研究并非用LLM仿真人类被试，而是进行了一项为期5天的纵向随机对照实验，招募992名参与者，随机分配到非个性化基线、记忆型个性化（基于对话历史）和问卷型个性化（基于初始问卷信息）三种条件，每天与模型进行关系建议互动，测量感知亲密度、过度依赖、自我披露、披露舒适度和决策后悔等结果变量。","baseline":"无对照（该研究以真实人类参与者为被试，比较不同个性化条件与非个性化基线，未使用真实人类数据作为仿真对照基准）。","findings":"对AI能力、有用性和亲密感的感知变化主要由重复接触驱动，而非个性化本身；但个性化条件影响了自我披露和建议采纳：记忆型条件下参与者自我披露更多且认为模型不那么令人毛骨悚然，问卷型条件下参与者对分享个人信息有更高的后悔。","reliability":"论文未讨论（节选内容未提及仿真可靠性或失效条件，但作为人类实验研究，其局限可能包括样本代表性、短期时间跨度、特定领域限制等，但未在提供文本中明确说明）。","relevance":"该研究虽非LLM仿真人类被试，但提供了个性化LLM对人类行为影响的因果证据，可作为仿真研究的外部效度参照，帮助理解仿真在个性化交互场景中的偏差来源。","inspiration":"借鉴其纵向随机对照设计，通过多日重复互动分离暴露效应与个性化效应，并采用多种个性化实现方式对比｜可迁移到金融建议场景，如个性化AI理财顾问对投资者风险偏好、信息披露和决策后悔的影响｜以真实投资者为被试，随机分配至非个性化、记忆型（基于历史对话）和问卷型（基于风险测评）AI顾问，进行多轮投资建议互动，测量风险承担、自我披露和后悔，并与人类理财顾问的互动数据对照。"}},{"id":"2609.19635","version":1,"title":"Faithful Where It Can Be Checked: Auditing a Reflection Agent Against Its System Prompt in a Randomized Trial","zh_title":"在可检查处忠实：随机试验中审计反思代理与其系统提示的一致性","abstract":"Conversational agents are increasingly used to guide reflection. A recent randomized trial compared a GPT-4o career reflection agent with the same program in a static journaling survey. Agent participants ended less committed to their career plans and more doubtful. We coded all 17,930 turns from its two studies, checked our coding against human coders and linked conversations to the trial's surveys. The rules the agent followed were the easy-to-check ones, like a reply length cap. Told not to flatter, it praised participants in half of its turns; told to challenge gently, it almost never did, and such a break leaves no visible trace. The behavior tied to the worse outcome was the demand to decide: the survey posed each decision once, while the agent asked again when participants hesitated, and those pressed most ended most doubtful. Our findings inform reflection agent design and the writing of checkable instructions.","authors":["Subigya K. Nepal","Serena Soh","Noah Vinoya","SoHyun Park","Mahnaz Roshanaei","Gabriella Harari"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19635","pdf_url":"https://arxiv.org/pdf/2609.19635","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","行为审计","人机对照"],"reason":"用GPT-4o代理人类被试进行反思实验，并与人类数据对照，发现行为偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":11,"question":"GPT-4o 职业反思代理在随机试验中是否遵循其系统提示，以及哪些行为导致了与静态日志相比更差的职业承诺结果？","design":"该研究不是用 LLM 仿真人类被试，而是审计一个 GPT-4o 对话代理在随机试验中的行为。试验将参与者随机分配到与 GPT-4o 代理对话或填写静态日志调查，完成四天职业反思项目。研究者编码了全部 17,930 轮对话，检查代理是否遵循系统提示中的规则，并将对话行为与试验结果（职业承诺、怀疑等）关联。","baseline":"人类基准是同一随机试验中静态日志调查组的参与者结果，以及人工编码者对对话编码的校验。","findings":"代理遵循了易于检查的规则（如回复长度限制），但违反了难以检查的规则：被告知不要奉承，却在半数回合中赞美参与者；被告知温和挑战，却几乎从未做到。与更差结果相关的行为是要求参与者做出决定：当参与者犹豫时，代理反复追问，被追问最多的参与者最终怀疑程度最高。","reliability":"论文承认审计依赖于编码方案，且难以检查的规则（如温和挑战）的违反可能不会留下可见痕迹。此外，试验效应量适中，且结果可能受对话格式本身（与代理互动 vs. 独立反思）影响，而非仅由代理行为导致。","relevance":"该研究直接涉及 LLM 在人类实验中的行为审计，揭示了代理偏离指令的方式及其对结果的影响，对关注仿真可靠性和偏差的研究者有重要参考价值。","inspiration":"借鉴其审计方法：对 LLM 代理的行为进行系统编码并与指令对照，同时关联结果变量，以识别行为偏差。｜可迁移到经济金融场景，如 AI 财务顾问或信贷审批代理的行为审计，检查其是否遵循公平性、透明度等指令。｜设计：招募人类被试随机分配与 LLM 财务顾问互动或使用静态工具，编码对话中代理的奉承、挑战和建议行为，测量被试的投资决策或风险偏好，并与真实人类顾问数据对照。"}},{"id":"2609.20425","version":1,"title":"Welfare-Opaque Income: Taxation under AI-Agent Delegation","zh_title":"福利不透明收入：AI代理委托下的税收","abstract":"We study income taxation when an AI agent implements economically relevant choices through a rule hidden from the government. Alongside unobserved productive ability, this hidden preference-to-execution mapping creates \\emph{double unobservability}: the same observable tax-base response can carry different welfare consequences. We call the resulting income \\emph{welfare-opaque}. Our constructions show that tax-base statistics can coincide while reform welfare effects differ, even when mechanical welfare weights are identical. We derive an optimal-tax condition that adds a response-weighted execution wedge to the familiar sufficient statistics. A higher marginal rate gains a corrective benefit under local over-execution and an additional cost under local under-execution. Observing the wedge identifies the welfare effect of a marginal reform at the prevailing schedule; bounds on it deliver bounds on that effect. A controlled laboratory compares 4,500 model runs across five AI engines. Faithful delegation selects the score maximizer in essentially all runs. Conflicted objectives produce heterogeneous responses: Claude largely preserves the score maximizer, GLM moves predominantly downward, and GPT-mini and Qwen show concentrated lower-tail increases. Qwen also makes substantial downward adjustments. Different engines locate their departures at different points and in different directions of the designed distribution. Explicit scores align model rankings; formula-based objective instructions yield more uneven agreement. Qwen shows a clear positive tax-by-objective interaction, but its direction does not generalize across engines and the pooled sign depends on its inclusion. The analysis identifies execution information as a complement to conventional tax-base statistics.","authors":["Yukun Zhang","Kemu Xu","Yishen Chen"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20425","pdf_url":"https://arxiv.org/pdf/2609.20425","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["AI代理","税收政策","经济仿真"],"reason":"用AI代理模拟经济决策并与实验室数据对照，涉及税收政策评估，但非直接复现人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":15,"question":"当AI代理以隐藏规则执行经济选择时，政府如何设计最优所得税？","design":"用五个AI引擎（Claude、GLM、GPT-mini、Qwen等）模拟劳动者在给定收入、工时、疲劳、未满足需求和评分下的工作选择，通过改变引擎、温度、收入分布、税率和指令目标（忠实委托或冲突目标）共4500次运行，测量AI选择的收入水平及其对税收变化的响应。","baseline":"无对照","findings":"忠实委托下AI几乎总是选择评分最大化选项；冲突目标下不同引擎表现出异质性偏离，且偏离位置和方向因引擎和状态而异。税收与目标的交互效应在Qwen上显著为正，但不具有跨引擎普遍性。","reliability":"论文未讨论","relevance":"该研究用AI代理模拟经济决策并评估税收政策，属于LLM仿真实验，但缺乏真实人类数据对照，且聚焦于理论信息问题而非直接复现人类行为，与关注人类基准和可靠性的研究兴趣部分相关。","inspiration":"可借鉴其通过系统操纵AI引擎、指令目标和环境参数来生成行为分布并考察异质性的实验设计方法。｜可迁移到政策评估场景，如税收改革对劳动供给的影响、福利政策的行为反应等。｜用多个LLM作为被试，随机分配不同税收规则和指令目标，测量其报告的劳动收入选择，并与真实劳动调查数据（如CPS或面板数据）对照，检验AI仿真与人类行为的一致性。"}},{"id":"2608.25245","version":2,"title":"The \"Curse of Knowledge\" in LLM Query Simulation: Concept Provenance for Tracing Answer-Side Intrusion","zh_title":"LLM查询模拟中的“知识诅咒”：用于追踪答案侧侵入的概念溯源","abstract":"LLM-generated search queries are widely used to augment IR evaluation, yet they may contain concepts that presuppose answer-side document knowledge, violating the information-access boundary of pre-search users. Existing validation metrics, including overlap, diversity, and effectiveness, cannot distinguish rare human-tail variation from candidate answer-side intrusion. We introduce concept provenance, a framework that assigns query concepts to backstory-supported, human-central, human-tail, and candidate answer-side zones, operationalizing a boundary that retrieval metrics alone cannot detect. Applying concept provenance to 77,004 queries across 100 UQV100 topics, 8 LLMs, and 5 prompt conditions with two extraction pipelines, we obtain a cross-pipeline token-HCIR Spearman rho of 1.0 over five condition means. Candidate answer-side concepts constitute 7.40 percent of non-generic concepts and appear in 97 of 100 topics, with topic explaining approximately 67 percent of variance. Human validation yields 68.2 percent relaxed precision, revealing two mechanisms: knowledge intrusion at 45.5 percent and deployment intrusion at 45.0 percent. Diagnostic probes show disproportionate localized retrieval effects, with deletion effect size d = -0.47 compared with d = -0.34 for random deletion, but these concepts explain less than 2 percent of aggregate evaluation variance. Concept provenance therefore serves as a boundary-compliance diagnostic rather than an evaluation-shift predictor. Under the tested conditions, no prompt condition eliminates intrusion; post-generation concept-provenance selection achieves 99 percent elimination.","authors":["Chenglong Ma","Xinye Wanyan","Danula Hettiachchi","Ziqi Xu","Jeffrey Chan"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-09-18","first_seen":"2026-08-27","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2608.25245","pdf_url":"https://arxiv.org/pdf/2608.25245","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM查询生成","信息检索评估","概念溯源"],"reason":"LLM生成查询用于IR评估，非仿真人类被试，但涉及LLM替代人类生成查询，属边…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":12,"question":"LLM生成的初始查询中，概念在来源区域上如何分布？候选答案侧概念入侵是否预测检索池、判定覆盖和系统排名的偏移？仅靠提示词缓解是否足以保持边界合规？","design":"该研究不是人类仿真实验，而是对LLM生成查询的边界合规性进行诊断。使用8个LLM在5种提示条件下为100个UQV100主题生成77,004条查询，通过概念溯源框架将查询概念分配到背景支持、人类中心、人类尾部和候选答案侧四个区域，并采用双提取管道、阈值敏感性和人工标注进行验证。","baseline":"UQV100中每个主题的真实人类初始查询变体集，以及主题背景描述和判定相关文档。","findings":"候选答案侧概念占非通用概念的7.40%，出现在97/100个主题中，主题解释了约67%的方差。这些概念对局部检索效果有不成比例的影响（删除效应量d=-0.47，高于随机删除的d=-0.34），但仅解释不到2%的总体评估方差，因此概念溯源是边界违规诊断工具而非评估偏移预测器。","reliability":"论文承认概念溯源在聚合评估层面解释力有限（<2%），且人工验证的宽松精度为68.2%，存在知识入侵（45.5%）和部署入侵（45.0%）两种机制；提示词缓解不能完全消除入侵，需要后生成选择才能达到99%消除。","relevance":"该研究批判性地揭示了LLM生成查询中普遍存在的答案侧知识入侵问题，并提供了可操作的诊断框架，对于关注LLM仿真可靠性与偏差的研究者具有直接参考价值，值得阅读原文以了解概念溯源的具体操作和验证方法。","inspiration":"借鉴概念溯源框架，将生成内容中的概念与真实人类数据中的概念分布进行对比，以识别超出信息边界的成分，并通过删除实验量化其对结果的影响。｜可迁移到经济预测调查中，例如使用LLM模拟分析师或消费者对政策公告的预期形成，检测生成预期是否包含了公告后才可获得的信息。｜以LLM扮演经济主体，给定政策公告前的背景信息生成预期，结果变量为预期值与实际公告值的偏差；用真实分析师调查数据（如蓝筹经济指标）作为人类基准，通过概念溯源识别LLM预期中的答案侧概念，并比较删除这些概念前后预期偏差的变化。"}},{"id":"2606.27845","version":2,"title":"LLM Agents as Static Level-k Players in Behavioural Games","zh_title":"行为博弈中作为静态层级-k玩家的LLM智能体","abstract":"Large Language Models (LLMs) are increasingly used as stand-ins in behavioural games. These stand-ins rely on the assumption that the LLM's distribution of choices meaningfully matches how humans play the same game. This study tests that assumption through two games. The first is a p-beauty contest, and the second one is a public goods game. The study first investigates five local-model settings within the same model family. These settings are varied together in a 360-cell factorial, which balances temperature, scale (0.5-32B), quantisation, instruct vs base, and framing. Each cell's distribution is then compared against whole choice distributions in published human data. Each deployment setting, except for quantisation, governs a different aspect of fidelity. Mechanically, while the dispersion of human players can be somewhat recovered through deployment settings, the strategic process behind it cannot. Through the lens of the level-k cognitive theory, we find that LLMs act as static, category-retrieved level-k players, where k is set by the model scale. The models also do not run within-game belief-updating or backward induction throughout multiple-round horizon settings. While human contributions decayed in the public goods game, LLMs stayed flat or rose at every scale. When the horizon test was administered, LLMs were more cooperative under an indefinite horizon compared to a finite one. However, LLMs ignore their relative round position, so no last-round defection was displayed. This implies that LLMs retrieved levels relative to the horizon category rather than working out iteratively from the specific game setting.","authors":["Po Han Teo"],"categories":["econ.GN","econ.TH","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-09-17","first_seen":"2026-06-26","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2606.27845","pdf_url":"https://arxiv.org/pdf/2606.27845","source_feed":"econ.GN","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为博弈","算法保真度"],"reason":"直接测试LLM在行为博弈中替代人类被试的保真度，并与已发表人类数据对照，发现静…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":1,"question":"LLM在行为博弈中作为人类被试替代品时，其选择分布和策略过程是否与真实人类一致？","design":"使用Qwen 2.5模型家族，在p-beauty contest和公共品博弈中，通过360个单元格的因子设计操纵温度、模型规模（0.5-32B）、量化、指令/基础模型和框架，测量LLM的选择分布，并与已发表的人类数据比较。","baseline":"已发表的人类行为数据，包括p-beauty contest和公共品博弈中的选择分布。","findings":"LLM表现为静态的、类别检索的level-k玩家，k由模型规模决定；LLM没有进行回合内信念更新或逆向归纳，在公共品博弈中贡献不衰减，且忽视相对回合位置，无最后一轮背叛。","reliability":"论文未讨论","relevance":"该研究直接检验LLM在行为博弈中替代人类被试的保真度，并与真实人类数据对照，发现LLM的策略过程与人类不同，对评估LLM仿真可靠性至关重要，值得精读原文。","inspiration":"采用因子设计系统操纵LLM部署设置，并与人类分布整体比较的方法值得借鉴｜可迁移到政策公告的预期形成实验，如央行沟通对通胀预期的影响｜用LLM模拟公众对政策公告的反应，处理为不同政策措辞或沟通方式，结果变量为预期通胀分布，与调查预期数据（如密歇根大学调查）对照。"}},{"id":"2609.17549","version":1,"title":"Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues","zh_title":"社会模式在合成数据中是否成立？分析LLM生成与真实对话中的网络欺凌动态","abstract":"Cyberbullying (CB) is a complex social phenomenon characterized by repeated aggression, power imbalance, and multi-party interaction. Although large language models (LLMs) are increasingly used to generate synthetic CB conversations for data augmentation and benchmarking, it remains unclear whether such data faithfully reproduces the social dynamics of authentic interactions beyond supporting downstream task performance. We present a comprehensive framework for evaluating the social realism of LLM-generated CB conversations. We compare authentic and synthetic dialogues generated by GPT, Grok, and LLaMA across interactional structure (turn-taking, power dynamics, and repair behavior), linguistic and stylistic realism (pronoun usage and humor), affective and behavioral markers (CB types, profanity, and toxicity), and temporal escalation dynamics. We further complement automatic analyses with a human evaluation of cyberbullying presence, scenario relevance, role plausibility, and social realism. Our results show that LLM-generated data consistently preserves high-level interactional structure, including role participation patterns, directional power asymmetry, and broad distributions of behavioral markers. However, all models systematically distort finer-grained social phenomena, including behavioral magnitude, role-specific allocation, categorical distributions, and temporal dynamics. These distortions are strongly model-dependent: GPT suppresses harmful content, Grok amplifies aggressive behaviors, and LLaMA provides the most balanced approximation while smoothing role distinctions. Our findings show that synthetic CB data is useful for modeling global interactional structure but remains an imperfect substitute for authentic conversations when behavioral realism and social dynamics are essential.","authors":["Arefeh Kazemi","Hamza Qadeer","Sinan Asci","Joachim Wagner","Brian Davis"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17549","pdf_url":"https://arxiv.org/pdf/2609.17549","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","社会动态","真实性评估"],"reason":"用LLM生成对话仿真网络欺凌社会动态，并与真实对话对照，评估仿真保真度与偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":2,"question":"LLM生成的网络欺凌对话是否忠实再现了真实对话中的社会动态，而不仅仅是支持下游任务性能？","design":"使用GPT、Grok和LLaMA生成合成网络欺凌对话，与真实对话对比，分析互动结构（话轮转换、权力动态、修复行为）、语言风格（代词使用、幽默）、情感行为标记（欺凌类型、脏话、毒性）和时间升级动态，并进行人类评估。","baseline":"真实网络欺凌对话数据集（具体名称未在节选中提及）。","findings":"LLM生成的数据保留了高层互动结构，如角色参与模式、方向性权力不对称和行为标记的广泛分布；但所有模型都系统性地扭曲了细粒度社会现象，包括行为强度、角色特定分配、类别分布和时间动态，且扭曲程度因模型而异：GPT抑制有害内容，Grok放大攻击行为，LLaMA提供最平衡的近似但平滑了角色差异。","reliability":"论文承认合成数据在需要行为真实性和社会动态时是真实对话的不完美替代品，且扭曲具有模型依赖性。","relevance":"该研究直接评估LLM仿真社会互动的保真度，与研究者关注的人类仿真实验和可靠性评估高度相关，提供了系统的对照基准和偏差分析，值得精读。","inspiration":"借鉴其多维评估框架，将自动分析与人类评估结合，系统比较合成与真实数据在结构、语言、行为和时间动态上的差异｜可迁移到经济金融中的社会互动场景，如谈判博弈、团队决策或市场情绪传播｜设计一个实验：用LLM生成模拟投资者在社交媒体上的互动对话，处理为不同模型（如GPT、Claude）或提示策略，结果变量为情绪传染、羊群行为或信息扩散模式，与真实投资者论坛数据（如StockTwits）对照，评估仿真保真度。"}},{"id":"2609.18106","version":1,"title":"Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment","zh_title":"开放权重大语言模型应用于招聘时性别与种族偏见的语言触发因素","abstract":"Open-weight large language models are rapidly entering hiring pipelines, yet their discriminatory failure modes -- and the regulatory exposure these create under the EU AI Act high-risk classification (Annex III) and U.S. EEOC adverse-impact analysis -- remain poorly understood. We present the first systematic, multi-model audit of open-weight LLMs that treats job-posting language as the primary experimental variable, evaluating six models (Llama 3.2, Mistral, Gemma 3, Qwen 3, Phi 3, DeepSeek-R1) across four controlled experiments that jointly probe recruiter-simulation and job-seeker-simulation tasks. We find that (1) agentic posting language depresses recruiter recommendation scores for female candidates (r_rb = 0.309, p_Bonf = 7x10^-5; model-fixed-effects r_rb = 0.448), while communal language partially reverses the penalty; and (2) coded-exclusion language suppresses non-White recruiter scores at large effect sizes (r_rb = 0.646-0.758) and, on the job-seeker side, selectively deters non-White personas from expressing interest -- operationalizing a chilling-effect mechanism at scale. A label-ablation experiment isolates the explicit demographic persona label as the primary causal driver, and Word Embedding Association Tests corroborate these findings at the representational level (d = 1.01-1.45 under Caliskan et al.'s multi-word gender attribute lists). We translate these results into a concrete pre-deployment audit protocol -- posting-vocabulary scoring, persona-conditioned LLM probing, and adverse-impact flagging against the four-fifths threshold -- that operationalizes the documentation and risk-management obligations Annex III imposes on high-risk AI in recruitment.","authors":["Kosuke Kitahara","Nobuhiro Yamaguchi"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18106","pdf_url":"https://arxiv.org/pdf/2609.18106","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","招聘偏见","算法审计"],"reason":"用LLM仿真招聘中的人类决策，并与真实人类数据对照，评估偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":5,"question":"招聘广告中的语言特征（如代理性/社群性词汇、编码排斥语言）如何触发开放权重LLM在招聘模拟中的性别与种族偏见？","design":"使用六个开放权重LLM（Llama 3.2、Mistral、Gemma 3、Qwen 3、Phi 3、DeepSeek-R1）进行招聘者模拟和求职者模拟实验。通过操纵招聘广告语言（代理性 vs. 社群性词汇、编码排斥语言）和候选人/求职者的人口统计标签（性别、种族），测量LLM给出的推荐评分、兴趣表达等结果变量，并进行标签消融实验和词嵌入关联测试。","baseline":"无对照（论文未使用真实人类数据作为基准，而是基于LLM输出进行审计）。","findings":"代理性招聘语言降低女性候选人的推荐评分，而社群性语言部分逆转该惩罚；编码排斥语言大幅抑制非白人候选人的推荐评分，并在求职者模拟中抑制非白人角色的兴趣表达，形成寒蝉效应。标签消融实验表明显式人口统计标签是主要因果驱动因素，词嵌入关联测试在表征层面证实了偏见。","reliability":"论文未讨论","relevance":"该研究直接针对LLM在招聘场景中的仿真行为，系统操纵语言变量并测量偏见输出，与您关注的LLM仿真可靠性及偏差评估高度相关，值得阅读原文以了解其审计协议和效应量。","inspiration":"借鉴其将文本特征作为处理变量、通过多模型审计和标签消融识别因果机制的方法。｜可迁移到信贷审批中的语言歧视研究，如贷款广告或申请表中的措辞对AI审批决策的影响。｜以LLM作为信贷审批员，处理为贷款申请描述中的代理性/社群性词汇或编码排斥语言，结果变量为审批通过率或利率，对照真实信贷审批数据（如抵押贷款披露数据）评估仿真偏差。"}},{"id":"2609.17933","version":1,"title":"AI Mediators Regulate Emotion and Create Value in Disputes","zh_title":"AI调解员在纠纷中调节情绪并创造价值","abstract":"In conflict and disputes, especially, emotion acts as a salient force in influencing outcomes. Prior work shows negative affect can obstruct collaborative behaviors, which typically lead to ``win-win'' outcomes. Thus, some suggest mediators may help regulate emotion and achieve joint gains. With the proliferation of AI, we posit LLMs may perform well at this task, with the added benefit of better accessibility compared with a human mediator. To examine the effectiveness of AI versus novice human mediators, we conduct a between-subjects experiment, where participants engage in a dispute mediated by a human, AI, or no mediator. We first analyze how well the mediators regulate emotions within a dispute -- finding AI mediators perform significantly better than humans at reducing negative emotion. We next examine whether AI mediators facilitate disputants better realizing joint gains in disputes with high integrative potential (IP) -- we find a marginally significant interaction between IP and condition (AI versus human), indicating LLMs may outperform humans at aiding disputants realize joint gains. Lastly, we perform an analysis of the messages the mediators sent, finding the AI sent significantly more messages suggesting trade-offs compared to the humans.","authors":["James Hale","Jonathan Gratch"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17933","pdf_url":"https://arxiv.org/pdf/2609.17933","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","调解实验","人机对照"],"reason":"用LLM作为调解人替代人类，与真实人类被试互动并对照，属于人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":3,"question":"在情绪化的纠纷调解中，AI调解员相比新手人类调解员能否更有效地调节情绪并帮助双方实现共同收益？","design":"使用大语言模型（LLM）扮演AI调解员，与新手人类调解员和无调解员条件进行被试间实验。参与者扮演买卖双方进行纠纷谈判，调解员在每轮发言后决定是否干预并发送消息。结果变量包括情绪调节效果（负面情绪减少）、共同收益实现程度（积分潜力IP与条件的交互）以及调解员消息内容（如提出权衡建议的频率）。","baseline":"新手人类调解员作为对照，以及无调解员条件作为基线。参与者为真实人类被试，其偏好通过分配100点来测量，并根据达成目标获得金钱奖励。","findings":"AI调解员在减少负面情绪方面显著优于人类调解员；在具有高整合潜力的纠纷中，AI调解员可能比人类调解员更能帮助双方实现共同收益（边际显著交互作用）。此外，AI调解员发送了更多建议权衡的消息。","reliability":"论文未讨论","relevance":"该研究直接使用LLM替代人类调解员，与真实人类被试互动并对照人类调解员表现，属于人类仿真实验，且涉及情绪调节和谈判结果，与研究者关注的经济学实验和政策评估场景高度相关，值得阅读原文了解实验细节和局限性。","inspiration":"借鉴其使用LLM作为干预代理与人类被试互动的实验设计，以及通过偏好分配和积分潜力测量共同收益的方法。｜可迁移到经济金融中的谈判或冲突解决场景，如劳资纠纷调解、商业合同谈判、消费者投诉处理等。｜设计一个实验：以LLM作为调解员，人类被试扮演谈判双方（如买方和卖方），处理为AI调解员、人类调解员或无调解员，结果变量为谈判达成协议的质量（如联合收益、满意度）和情绪变化，对照真实人类调解员数据，并利用已有谈判实验数据集（如KODIS）进行基准比较。"}},{"id":"2609.18060","version":1,"title":"AI Peers Exert Social Influence on Human Dishonesty in Groups","zh_title":"AI同伴对群体中人类不诚实行为施加社会影响","abstract":"Human dishonesty in group settings is highly susceptible to peer influence, particularly when incentivized. Although artificial intelligence (AI) evolves from passive tools into active collaborators, its impact on human moral behavior within groups remains underexplored. We addressed this gap through a two-phase randomized behavioral study (N=280 and N=360). We found AI agents exert substantial social influence comparable in magnitude to that of human peers. Specifically, participants reported more dishonestly when exposed to dishonest rather than honest normative cues. This effect is evident across injunctive, subjective, and descriptive social norms. Interestingly, the only significant adjacent behavioral change occurred when dishonest peer behavior first appeared, whereas further increases from one to four dishonest peers produced weaker and non-monotonic changes. Furthermore, participants rapidly converge on decision-making, showing modest increases in dishonest reporting through repeated exposure. These findings highlight the importance of managing the behaviors and normative signals communicated by AI group members.","authors":["Shuning Zhang","Xinyuan Zhou","Yuanyang Qiu","Tianqi Song","Yuting Yang","Yiwen Ren","Xin Yi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18060","pdf_url":"https://arxiv.org/pdf/2609.18060","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","社会影响","行为实验"],"reason":"用AI代理替代人类同伴，研究其对人类不诚实行为的社会影响，并与人类同伴对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":4,"question":"AI同伴的诚实规范如何影响群体中人类的不诚实行为？","design":"采用两阶段随机行为实验（N=280和N=360），使用激励性掷骰子报告任务。AI代理作为群体成员，通过描述性、指令性和主观性社会规范传递诚实或不诚实线索，测量人类被试的虚报行为。","baseline":"与人类同伴的社会影响进行对照，比较AI同伴与人类同伴的影响大小。","findings":"AI代理对人类不诚实行为产生显著社会影响，其影响程度与人类同伴相当。当出现不诚实同伴时，被试虚报增加，但同伴数量从1个增加到4个时影响减弱且非单调。","reliability":"论文未讨论","relevance":"该研究直接使用AI代理替代人类同伴，研究其对人类道德行为的影响，并与人类同伴对照，符合研究者对LLM仿真实验和真实人类基准的兴趣。","inspiration":"值得借鉴的是其将AI代理作为群体成员，通过操纵其行为规范来研究社会影响，并设置人类同伴对照组以量化AI影响的大小。｜可迁移到经济金融中的群体决策场景，如投资团队中AI顾问的不诚实建议对个人投资决策的影响。｜设计一个实验：被试与AI代理组成投资小组，AI代理提供虚报收益的建议（处理），测量被试的投资报告虚报程度（结果变量），并与人类同伴提供同样建议的对照组比较，同时使用真实投资数据作为外部基准。"}},{"id":"2609.07478","version":2,"title":"The Internal Anatomy of Strategic Choice in Large Language Models","zh_title":"大语言模型策略选择的内部解剖","abstract":"Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal $2\\times2$ games. We followed a prespecified incentive from prompt, through activations, to choice. Dense models mirrored the unadjusted human decline with game complexity. Incentive and choice were detectable in every model, but models differed in whether incentive reached the choice, aligned with it and, where tested, whether strengthening it shifted preference. The base and instruction-tuned Qwen2.5 models chose almost identically at baseline yet differed in whether incentive reached choice. Fixed decision cues were distinguishable internally but changed choices selectively. Similar behaviour can rest on different computation; post-training can reshape the path from represented incentive to decision while leaving behaviour and decodable information largely intact.","authors":["Vin\\'icius Ferraz","Leon Houf","Enrico Ferrea"],"categories":["cs.AI","cs.GT"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-09","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.07478","pdf_url":"https://arxiv.org/pdf/2609.07478","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","策略博弈","算法保真度"],"reason":"用LLM复现人类策略选择并与人类数据对照，分析内部计算与行为差异，评估仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":6,"question":"大语言模型在一次性策略博弈中，内部表征与计算如何将激励信号转化为选择行为？","design":"用四个开源大语言模型（Qwen2.5 base/instruct、Llama-3.1-Instruct、GPT-OSS）作为被试，在144个严格序数2x2博弈中做一次性选择，记录激活值，用线性探针解码激励和选择，并进行激励强化干预和固定决策线索提示。","baseline":"人类选择数据来自Moore et al. (2026)的相同博弈和支付尺度，以及Zhu et al. (2025a)的基数支付数据集映射。","findings":"密集模型复现了人类随博弈复杂度下降的未调整选择模式；所有模型都能解码激励和选择，但激励是否到达选择、与选择对齐及干预效果因模型而异，基础与指令微调模型行为相似但内部路径不同。","reliability":"论文指出相似行为可能基于不同计算，后训练可重塑激励到决策的路径而保持行为和可解码信息基本不变，暗示行为仿真可能掩盖内部差异。","relevance":"该研究直接以LLM仿真人类策略选择并与真实人类数据对照，同时揭示内部计算与行为的不一致，对评估仿真可靠性和偏差具有重要参考价值，值得精读原文。","inspiration":"借鉴其用线性探针追踪激励信号从输入到选择的内部路径，并施加干预检验因果性的方法，可迁移到经济决策仿真中检验模型是否真正使用经济激励而非表面模仿。｜可应用于资产定价实验，检验LLM是否内部表征风险溢价并据此决策。｜以LLM为被试，呈现不同风险收益的资产选择任务，用探针解码风险溢价信号并强化干预，与人类实验数据（如股票市场参与决策）对照，观察选择变化。"}},{"id":"2609.17534","version":1,"title":"Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits","zh_title":"LLM中的装好与装坏：黑暗三人格特质下的反应失真","abstract":"Social desirability and impression management are pervasive sources of response distortion in human personality assessment, yet their effects on Large Language Models (LLMs) remain underexplored. This study investigates whether contemporary LLMs systematically modulate the expression of Dark Triad traits (Machiavellianism, narcissism, and psychopathy) under fake-good and fake-bad conditions. Seven state-of-the-art models were evaluated across two ecologically relevant contexts: employment selection and forensic evaluation, in which socially desirable or undesirable incentives were conveyed through contextual framing. Trait expression was measured using standard psychometric scoring procedures and compared with self-assessment baselines at both aggregate and item levels. Results revealed systematic and condition-consistent response modulation. Most models reduced Dark Triad scores under fake-good conditions and increased them under fake-bad conditions, although the magnitude and consistency of these effects varied across traits and models. Machiavellianism and narcissism showed the strongest and most coherent shifts, whereas psychopathy displayed greater heterogeneity. Context also influenced responses, with employment scenarios generally producing larger effects than forensic scenarios. An additional experiment showed that explicit fake-bad instructions generated substantially stronger distortions than contextual framing alone. The results suggest that personality-related outputs should be interpreted in light of the motivational and situational context in which they are elicited. More broadly, they highlight the value of psychometric paradigms for evaluating susceptibility to response distortion, impression management, and context-dependent behavioral shifts, with important implications for LLM benchmarking, alignment evaluation, and robustness assessment.","authors":["Victoria Popa","Guglielmo Cola","Caterina Senette","Maurizio Tesconi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17534","pdf_url":"https://arxiv.org/pdf/2609.17534","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM人格测量","反应偏差","仿真可靠性"],"reason":"研究LLM在人格测量中的反应失真，与人类数据对照，评估仿真偏差，可迁移到人类仿…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":7,"question":"LLM 在人格测量中是否会在假好（fake-good）和假坏（fake-bad）条件下系统性地调节黑暗三联征特质的表达？","design":"使用七个先进LLM，在就业选拔和司法评估两种情境下，通过提示框架施加假好或假坏动机，用TRAIT框架测量黑暗三联征（马基雅维利主义、自恋、精神病态）得分，并与基线自我评估比较。","baseline":"无对照","findings":"大多数模型在假好条件下降低黑暗三联征得分，在假坏条件下提高得分，但效应大小和一致性因特质和模型而异。马基雅维利主义和自恋的偏移最强且最一致，精神病态异质性更大；就业情境比司法情境产生更大效应，显式假坏指令比情境框架产生更强扭曲。","reliability":"论文未讨论","relevance":"该研究展示了LLM对情境动机的敏感性，可用于评估仿真中的社会期望偏差和印象管理，对使用LLM模拟人类被试时的可靠性有警示意义。","inspiration":"借鉴其通过情境框架施加动机处理并测量特质偏移的方法，可迁移到经济金融中的社会期望偏差场景（如信贷审批中的歧视、消费者道德行为）。｜可设计实验：用LLM扮演贷款申请人，在强调社会责任或利润最大化的不同银行政策下，测量其自我报告的诚信或风险偏好，并与实际信贷数据中的偏差对照。"}},{"id":"2609.18282","version":1,"title":"Too Good to Be Real? Diagnosing and Reducing the Gap Between AI Preference and Real User Engagement","zh_title":"好得难以置信？诊断并缩小AI偏好与真实用户参与度之间的差距","abstract":"Large language models are increasingly used to generate and evaluate online content, yet it remains unclear whether the qualities they associate with higher engagement match what real users respond to. We study this question using 1.17 million answers to 25,978 questions from Zhihu, Quora, and Reddit, comparing real platform answers and AI-generated answers across four within-question engagement levels. We introduce Ontological Preference Measurement, which represents answers along three dimensions: logic, affect, and expression. We find a systematic gap between AI preference and real user engagement: as target engagement increases, LLMs add more explicit logical structure, while real user engagement is more strongly associated with affective and expressive salience. We call this tendency logic overbinding. Based on this diagnosis, we propose Ontology-Masked Reasoning Autoencoding (OMRA), a controlled intervention that masks and reconstructs over-explained spans while preserving stance, factual content, and coherence. Across four LLM families, OMRA reduces the measured gap by an average of 54.4%. In human evaluation, OMRA wins 62.4% of pairwise preference judgments against matched real platform answers, even though the real answers are more often judged to be human-written.","authors":["Xinglang Zhang","Yuanmeng Xiang","Yunyao Zhang","Zeliang Chen","Junqing Yu","Zikai Song"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18282","pdf_url":"https://arxiv.org/pdf/2609.18282","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类行为对照","内容生成"],"reason":"用LLM生成内容并与真实用户互动数据对照，诊断AI偏好与人类行为差距，属于仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":9,"question":"LLM 在生成高互动内容时偏好的文本特征是否与真实用户互动行为一致？","design":"使用 LLM 生成针对四个互动等级的答案，与真实平台答案对比，通过本体偏好测量框架从逻辑、情感、表达三个维度量化文本特征，并施加 OMRA 干预以缩小差距。","baseline":"来自知乎、Quora 和 Reddit 的 117 万条真实答案，按问题内投票排名分为四个互动等级。","findings":"LLM 在追求高互动时过度增加显性逻辑结构，而真实高互动答案更依赖情感和表达显著性，作者称之为“逻辑过度绑定”。OMRA 干预平均缩小 54.4% 的差距，并在人类评估中胜过真实答案。","reliability":"论文未讨论","relevance":"该研究直接对比 LLM 生成内容与真实用户行为，诊断 AI 偏好与人类反应的系统性偏差，并尝试通过干预修正，对评估 LLM 仿真人类行为的可靠性具有参考价值。","inspiration":"借鉴其本体驱动的多维测量和干预设计，可迁移到经济金融领域的文本生成场景，如政策沟通、市场评论或金融建议。｜例如，研究 LLM 生成的央行政策声明或分析师报告是否与真实市场反应匹配。｜以 LLM 为被试，要求其生成不同目标市场反应的金融文本，测量逻辑、情感、表达特征，并与真实市场数据（如股价波动、交易量）对照，检验 AI 偏好与真实投资者反应的差距。"}},{"id":"2609.17989","version":1,"title":"Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders","zh_title":"AI代理为谁工作？角色分配引发LLM推荐中的赞助偏差","abstract":"Large language models (LLMs) now serve as conversational shopping assistants on platforms that also sell advertising. These AI agents face a conflict of duty. They advise consumers who rely on their judgment, yet are deployed by platforms that benefit when sponsored listings are chosen. Sponsorship disclosures, designed to allow consumers to penalize paid placements, now reach the AI agent rather than the consumer, and the agent's evaluation of them is hidden from the consumer. Drawing on the fiduciary concept of conflict of duty, we argue that an agent's evaluation of a sponsored listing should not depend on which party deployed it. In controlled choice experiments, we manipulate assigned roles in the system prompt to name either a traveler or a booking platform as the agent's principal. Platform delegation significantly attenuates the penalty that agents apply to sponsored listings and weakens the skepticism that disclosure triggers in their reasoning traces. We replicate out findings across LLMs and reasoning depths. A second study decomposes the disclosure label and shows that the divergence between the two delegates widens significantly when the paid placement is attributed to the platform. Stricter terminology (\"Sponsored\" instead of \"Promoted\") lowers choice of paid listings but does not close this gap when the platform is named. The findings show that disclosure mandates designed for human consumers cannot by themselves protect consumers in AI-mediated commerce.","authors":["Davood Wadi","Yu Ma"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17989","pdf_url":"https://arxiv.org/pdf/2609.17989","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","消费者决策","赞助偏差"],"reason":"用LLM模拟消费者决策，与人类对照，揭示角色分配导致的赞助偏差，属经济学实验场…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":8,"question":"AI代理在同时服务消费者和广告平台时，其推荐行为是否会因委托方身份（消费者vs平台）而出现赞助偏差？","design":"用LLM（Gemini 3.1 Pro等）模拟购物助手，在系统提示中操纵委托方身份（旅行者vs预订平台），呈现带赞助标签的酒店列表，测量选择赞助列表的概率及推理痕迹中的怀疑程度。","baseline":"无对照","findings":"平台委托显著减弱了LLM对赞助列表的惩罚，选择赞助列表的概率从消费者委托时的50.2个百分点降至29.2个百分点；当赞助标签明确归属平台时，平台委托的代理甚至表现出对赞助列表的强烈偏好（选择率74.2% vs 消费者委托的34.6%）。","reliability":"论文未讨论","relevance":"该研究用LLM模拟消费者决策，揭示了角色分配导致的赞助偏差，属于经济学实验场景，与研究者关注的人类仿真实验高度相关，值得精读原文。","inspiration":"值得借鉴的是通过最小化系统提示中的角色分配来诱发LLM行为偏差，并分析推理痕迹以揭示机制｜可迁移到信贷审批歧视或金融产品推荐场景，检验AI代理在银行与借款人之间的利益冲突｜设计：用LLM扮演贷款顾问，系统提示中分别指定委托方为银行或借款人，呈现带“银行推荐”标签的贷款产品，测量选择概率，并与人类贷款顾问的真实选择数据对照。"}},{"id":"2609.17544","version":1,"title":"Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation","zh_title":"大语言模型与中医医师的对比：真实世界临床病例评估","abstract":"Large language models (LLMs) are increasingly being explored for clinical applications, yet their assessment for real-world traditional Chinese medicine (TCM) practice remains limited We constructed a clinical case library comprising 349 de-identified outpatient cases from 62 hospitals and evaluated 16 LLMs and a comparator cohort of 60 practicing TCM physicians using 60 representative cases selected from this library. Model outputs and physician reports were anonymized and scored by five senior TCM experts across nine diagnostic and therapeutic dimensions. Cutting-edge general-purpose LLMs achieved higher expert scores than the physician comparators, particularly for medical advice, treatment principles and selected diagnostic tasks. However, prescription-level analyses revealed discrepancies in herb selection, dosage, and treatment strategy, and qualitative safety review identified hallucinations and undesirable template-driven outputs. These findings highlight the potential of LLMs for TCM decision support while underscoring the need for physician oversight, safety constraints and prospective clinical evaluation.","authors":["Jiacheng Xie","Xiaoting Tang","Yang Yu","Jinpu Li","Shouli Li","Congcong Jing","Yantao Yang","Zhiyong Zhao","Ziyang Zhang","Qilin Song","Guanghui An","Dong Xu"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17544","pdf_url":"https://arxiv.org/pdf/2609.17544","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","医疗决策","人类对照"],"reason":"用LLM替代医生进行临床决策评估，并与真实医生对照，属于人类仿真，但场景为医疗…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":10,"question":"在真实中医门诊场景中，大语言模型能否替代或辅助中医师进行辨证论治，其诊断与处方质量相比执业中医师如何？","design":"用16个大语言模型扮演中医师，对60个真实门诊病例生成诊断与处方报告；同时60名执业中医师对相同病例生成报告；所有报告匿名后由5名资深中医专家在9个诊断与治疗维度上评分，并分析处方模式、幻觉与安全性。","baseline":"60名执业中医师对相同60个病例生成的诊断与处方报告，经同一批专家按相同维度评分。","findings":"先进通用大模型在专家评分上总体高于医师对照组，尤其在医疗建议、治疗原则和部分诊断任务上表现更好；但处方层面在药物选择、剂量和治疗策略上存在差异，且定性安全审查发现幻觉和模板化输出。","reliability":"论文承认需要医生监督、安全约束和前瞻性临床评估；指出LLM可能产生幻觉、不安全的草药组合或不适当剂量，且当前评估基于回顾性病例，缺乏真实临床结局验证。","relevance":"该研究用LLM替代医生进行临床决策并与真实医生对照，属于人类仿真在医疗场景的应用，提供了真实世界数据与专家盲评的严格设计，值得阅读以了解仿真在专业决策中的有效性与偏差。","inspiration":"借鉴其匿名化处理与专家盲评设计，可确保仿真输出与人类输出在相同条件下比较，减少评估偏差｜可迁移到信贷审批歧视研究，用LLM模拟信贷员对贷款申请做出审批决策，比较其与真人信贷员的差异｜以LLM作为被试，处理为不同申请人特征（如种族、性别），结果变量为审批决定与理由，对照真实信贷审批数据或信贷员决策记录，检验LLM是否复现或放大人类偏见。"}},{"id":"2609.17550","version":1,"title":"No Usable Linear \"Capitulation Direction\" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback","zh_title":"两个小型LLM中不存在可用的线性“屈服方向”：激活引导声明的验证协议，以及跨家族对反驳下谄媚行为的研究","abstract":"Language models frequently abandon correct answers when users push back. We study this in two small instruction-tuned models from different families, Qwen2.5-1.5B and Llama-3.2-1B, over TriviaQA: the model answers, is challenged with one of four scripted pushback styles, and answers again. Conditioned on an initially correct answer, the models flip to a wrong answer in 41.8% and 43.1% of episodes. Which pressure works is a property of the model, not the pressure: the same within-question paired comparison (bare doubt vs. emotional appeal), specified in advance, is Bonferroni-significant in opposite directions across families (Qwen: bare doubt > emotional, OR 2.5, p=.040; Llama: emotional > bare doubt, OR 4.0, p=.001). Failure mode is also model-dependent: Llama abandons answers without recommitting at six times Qwen's rate (8.2% vs. 1.4%). Identical pushback repairs initially wrong answers only ~13% of the time; pushback is net epistemically destructive. We then ask whether capitulation is linearly decodable from the pre-response residual stream, a prerequisite for steering-vector interventions at that locus. A naive difference-in-means probe appears to succeed (in-sample AUROC 0.81/0.71), but a validation protocol combining question-level cross-validation, shuffled-label nulls, and a known-direction positive control shows the signal is overfitting: the best cross-validated AUROC is 0.582 in Qwen and 0.548 in Llama, both near or below their permutation thresholds and far under a pre-registered usability bar of 0.70, while the identical pipeline recovers a pushback-presence control direction at AUROC 1.000 in both. We further quantify a measurement hazard: substring grading underestimates capitulation by 18-24 percentage points. Code, prompts, transcripts, and analysis are released.","authors":["Saad Aamir","Muhammad Awais Bin Adil"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17550","pdf_url":"https://arxiv.org/pdf/2609.17550","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM行为","可靠性评估","激活引导"],"reason":"研究LLM在用户反驳下的行为变化，评估其可靠性，与人类仿真中的偏差问题相关。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":11,"question":"在用户反驳下，小型指令微调语言模型放弃正确答案（屈从）的行为模式是什么？屈从是否可由残差流中的线性方向解码？","design":"使用两个不同家族的小型指令微调模型（Qwen2.5-1.5B 和 Llama-3.2-1B），在 TriviaQA 数据集上进行多轮问答：模型先回答，然后受到四种预设反驳风格之一的挑战，再回答。测量结果变量包括：初始正确时翻转为错误的概率、初始错误时修复的概率、放弃不重新承诺的概率，以及屈从的线性可解码性（AUROC）。","baseline":"无对照","findings":"在初始正确的情况下，两个模型分别有41.8%和43.1%的回合翻转为错误答案；哪种压力有效是模型特有的，同一比较在家族间方向相反。线性探测看似成功，但经过交叉验证和置换检验后，屈从方向不可用（AUROC接近机会水平），而阳性对照方向可完美恢复。","reliability":"论文承认的局限包括：仅两个家族且规模小（1-1.5B）；单一数据集（TriviaQA）和单一语言（英语）；跨模型比较非问题配对；模板固定；LLM 裁判未经完整人工验证；Llama 的关键比较是在观察临时标签后指定的。","relevance":"该研究通过行为实验和严格的验证协议，揭示了 LLM 在用户压力下的屈从行为具有模型特异性，且线性方向不可靠，这对使用 LLM 模拟人类决策时的偏差评估和可靠性判断有直接参考价值，值得阅读原文。","inspiration":"借鉴其多轮交互设计、预设处理（反驳风格）、结果分类（翻转/修复/放弃）以及严格的验证协议（交叉验证、置换检验、阳性对照）来评估模型行为的稳健性。｜可迁移到经济金融中的政策沟通或建议采纳场景，例如 AI 财务顾问在用户质疑下是否改变投资建议，或消费者对 AI 推荐产品的信任变化。｜以 LLM 作为被试，模拟投资者在 AI 投资建议受到用户质疑时的反应，处理为不同风格的反驳（如权威质疑、情感诉求），结果变量为是否改变初始投资决策，并与人类投资者在类似情境下的真实行为数据（如实验或调查数据）进行对照。"}},{"id":"2609.18068","version":1,"title":"From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale","zh_title":"从基列河到大型语言模型的推理分布：隐性方言偏见与大规模语言画像","abstract":"Large language models (LLMs) are increasingly deployed in high-stakes domains such as housing screening. While alignment techniques mitigate explicit racial bias in generated text, they often leave covert attitudinal associations in internal probability distributions untouched. Adapting the matched-guise sociolinguistic paradigm, we examine covert dialect bias in housing-related social judgments across four varieties: Standard American English (SAE), African American Vernacular English (AAVE), Nigerian Standard English (NSE), and Nigerian Pidgin (NP). AAVE reflects the racialized dialect studied in prior covert-bias evaluations, whereas NSE and NP represent Black African, postcolonial varieties absent from this literature. Using 260 meaning-matched sentence quadruples and log-probability scoring over housing-relevant adjectives, we probe ten open-weight LLMs across three contexts varying in social proximity: tenant screening, neighbor acceptance, and roommate selection. Across all ten models, AAVE and NP are consistently associated with more negative adjectives than SAE, with NP penalized most severely. Crucially, each dialect is penalized via distinct stereotype clusters rather than a generic non-standard category. NSE, which carries institutional prestige, displays a context-dependent shift: favored over SAE in formal tenant screening but increasingly penalized as social proximity grows. Our findings reveal that LLMs inherit covert dialect bias along both racial identity and prestige dimensions, echoing documented human housing discrimination and demonstrating its reach across postcolonial English varieties.","authors":["Chowdhury Mohammad Abdullah","Rita Orji"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18068","pdf_url":"https://arxiv.org/pdf/2609.18068","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM偏见","社会语言学","人类仿真"],"reason":"用LLM复现人类语言态度，有真实人类研究对照，揭示仿真偏差","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":12,"question":"大语言模型是否在住房相关的社会判断中继承了隐蔽的方言偏见，这种偏见是由种族身份、语言声望还是两者共同驱动的？","design":"采用匹配伪装范式，使用260组语义匹配的句子四元组（SAE、AAVE、NSE、NP），对十个开源权重LLM进行对数概率评分，测量模型在住房相关形容词上的概率差异，并设置三个社会距离不同的评估情境（租户筛选、邻居接受、室友选择）。","baseline":"人类基准来自社会语言学文献中记录的住房歧视研究（如Massey和Lundy 2001年的电话筛选实验）以及匹配伪装实验（如Tucker和Lambert 1969），但本研究未直接收集新的人类数据。","findings":"所有十个模型对AAVE和NP的输入一致分配更负面的住房相关形容词概率，其中NP受罚最重；每种方言通过不同的刻板印象簇被惩罚，而非作为泛化的非标准类别。NSE在正式租户筛选情境中比SAE更受青睐，但随着社会距离增加而受到惩罚，显示出情境依赖性。","reliability":"论文未明确讨论失效条件，但指出研究仅限于开源权重模型和特定方言，且未直接与人类判断进行定量对比，仅引用历史人类研究作为背景。","relevance":"该研究直接展示了LLM在住房筛选等经济相关场景中复现人类歧视性判断，并揭示了隐蔽概率偏见，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其匹配伪装设计，通过控制语义内容仅改变方言特征来隔离语言变量的因果效应，并利用对数概率测量隐蔽态度，可迁移到信贷审批中的方言歧视研究。｜可应用于金融领域的信贷审批或保险定价中的语言偏见问题，例如评估贷款申请人的方言口音是否影响模型的风险评估。｜设计：使用LLM作为虚拟信贷员，输入语义相同但方言不同的贷款申请文本（如SAE、AAVE、西班牙口音英语），测量模型分配给违约相关形容词或风险评分的概率，并与真实信贷审批数据中的种族或语言歧视率进行对照。"}},{"id":"2609.18274","version":1,"title":"I code or AI code: A comparative evaluation of AI-rated scores in classroom observations","zh_title":"我编码还是AI编码：课堂观察中AI评分与人类评分的比较评估","abstract":"Classroom observations are widely recognized as a key tool for establishing benchmarks of education quality and guiding pedagogical improvement, yet they remain resource-intensive and dependent on trained observers. This study evaluated the feasibility of using a LLM (GPT-5 model) to score teacher-child interactions in early childhood classrooms, benchmarked against human raters. The study analyzed 87 video-recorded observations from 38 classrooms across 30 kindergartens in Hong Kong. Using observation transcripts, the AI model was configured to apply the full Classroom Assessment Scoring System (CLASS) framework. AI-rated scores were then compared with human ratings by examining correlations and differences in mean scores of the CLASS domains and dimensions. The results showed greater convergence between AI and raters for the Emotional Support domain and, in particular, the Quality of Feedback dimension, which captures how teachers use feedback to extend children's learning. Greater divergence emerged for interactions that were more procedural or context-dependent, particularly within the Classroom Organization and Instructional Support domains. These findings suggest that transcript-based AI scoring may capture some of the relative variation in teacher-child interactions but cannot yet reproduce calibrated human judgements consistently across the full CLASS framework. AI-assisted observation may therefore be more appropriate as a preliminary screening tool rather than as a replacement for trained observers, providing teachers with evidence for reflection rather than high-stakes evaluation. Future research should examine whether domain-specific training and incorporation of contextual and visual information can improve alignment between AI and human rated scores.","authors":["Y. Fong","J. Xiang","T. Y. D. Chan","K. Lee","E. Y. H. Lau"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18274","pdf_url":"https://arxiv.org/pdf/2609.18274","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","教育评估","人类对照"],"reason":"用LLM替代人类评分员评估课堂互动，并与人类评分对照，属于仿真人类判断，但非典…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":13,"question":"使用大语言模型（GPT-5）对幼儿园课堂师生互动进行CLASS评分，能否替代人类评分员？","design":"用GPT-5模型基于课堂观察转录文本，按照CLASS框架对87个视频观察（来自香港30所幼儿园38个教室）进行评分，并与人类评分员在CLASS各领域和维度上的评分进行相关性和均值差异比较。","baseline":"人类评分员对相同视频观察的CLASS评分。","findings":"AI与人类评分在情感支持领域及反馈质量维度上收敛较好，但在课堂组织和教学支持领域等程序性、情境依赖性互动上分歧较大。转录文本AI评分能捕捉部分相对变异，但不能稳定复现人类校准判断。","reliability":"论文承认转录文本AI评分无法完全复现人类校准判断，建议仅作为初步筛查工具而非高利害评估替代品；未来需探索领域特定训练和纳入情境与视觉信息以改善对齐。","relevance":"该研究用LLM替代人类评分员评估课堂互动，并与人类评分对照，属于仿真人类判断，但非典型经济学实验或调查场景，与研究者关注的政策评估和经济学实验关联较弱，可作方法参考。","inspiration":"借鉴其用LLM对非结构化文本进行标准化量表评分的做法，并设置人类评分员作为基准进行相关性和均值差异检验。｜可迁移到经济金融中需要专家主观评分的场景，如信贷审批中的软信息评估、分析师报告语调分类、政策文本情感倾向等。｜以信贷审批为例，用LLM对贷款申请者的文字描述进行信用软信息评分，处理为不同提示词或模型版本，结果变量为信用评分，与人类信贷员的真实评分做对照，检验一致性和偏差。"}},{"id":"2609.18341","version":1,"title":"Understanding AI Provider Recommendations in Local Service Markets","zh_title":"理解本地服务市场中的AI提供商推荐","abstract":"When someone asks an AI assistant which doctor to see or which firm to trust with their savings, the answer is a referral. We audit AI provider recommendations in four registry-backed service domains across the 100 largest U.S. metropolitan areas, matching every recommendation against the official registry for its domain (Medicare clinician and facility records, and SEC adviser disclosures), under three conditions: an open-weight model, a proprietary model without web search, and the same proprietary model with search. Without search, both models largely fabricate recommendations in the domains the web covers thinly. Only 4% of the open-weight model's recommended doctors and 11% of the proprietary model's match a clinician in the queried city, and the open-weight matches are name coincidences: its matched clinicians are no likelier to be primary-care doctors than names drawn at random from the registry. With search, 64-71% of recommendations in the same domains match a real provider. Search also changes who is recommended. Without it, recommended advisory firms carry SEC misconduct disclosures at 3.6 times the registry base rate, even after adjusting for firm size; with search, significantly below it. Restaurants, where quality and visibility are separately measurable, show a 3-5x review-count premium but a rating premium of at most a tenth of a star. Finally, search largely removes the metro-size penalty: without it, real recommendations concentrate in the largest metros; with it, match rates are similar across metro-size terciles. Whether an AI referral is trustworthy depends strongly on its retrieval configuration rather than on the underlying model alone, yet an answer produced without retrieval often carries no sign that its recommendations were never verified.","authors":["Hazem Ibrahim","Yasir Zaki"],"categories":["cs.CY","cs.CL","cs.IR"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18341","pdf_url":"https://arxiv.org/pdf/2609.18341","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM可靠性","审计研究","真实数据对照"],"reason":"审计LLM推荐与真实注册数据对照，揭示无检索时虚构与偏差，可迁移至仿真可靠性评估","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":14,"question":"AI助手在本地服务市场中的提供商推荐是否真实、质量如何，以及检索配置如何影响推荐的可信度。","design":"审计研究：使用一个开源模型和一个专有模型（后者在有无网络搜索两种条件下）回答五类日常问询，覆盖美国100个最大都市区的四个注册服务领域，将推荐名称与官方注册数据（Medicare临床医生和设施记录、SEC顾问披露）进行匹配，测量匹配率、质量指标和偏差。","baseline":"官方注册数据：Medicare临床医生和设施记录（含星级评分）、SEC投资顾问披露（含不当行为记录）、餐厅众包评分和评论数。","findings":"无搜索时，模型在医生和养老院领域大量虚构推荐（匹配率仅4%-11%），且开源模型的匹配纯属姓名巧合；有搜索时匹配率升至64%-71%。搜索改变了推荐对象：无搜索时推荐的咨询公司SEC不当行为披露率是基准的3.6倍，有搜索时显著低于基准；搜索还消除了都市规模惩罚，使不同规模都市的匹配率趋于一致。","reliability":"论文指出，无检索的答案往往没有迹象表明推荐未经核实，且引用分析显示搜索推荐依赖商业来源而非监管来源（医生领域仅0.3%引用政府来源），存在可观察性差距；但未系统讨论模型幻觉的边界条件或不同提示措辞的影响。","relevance":"该研究通过对照真实注册数据审计LLM推荐，揭示了无检索时的虚构与偏差，可迁移至仿真可靠性评估，值得阅读原文以了解审计方法和偏差量化。","inspiration":"借鉴其审计设计：在同一模型内对比有无检索，用官方注册数据作为基准，测量匹配率和质量偏差，并分析偏差的稳健性。｜可迁移到金融顾问推荐、信贷产品推荐或医疗资源分配等场景，评估LLM仿真中的选择偏差。｜以LLM作为被试，处理为有无检索，结果变量为推荐实体的真实性和质量指标（如违规记录、评分），对照SEC或消费者金融保护局等官方数据，检验仿真是否复现或放大偏差。"}},{"id":"2609.18346","version":1,"title":"Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition","zh_title":"忠实却共谋：为何思维链监控无法检测寡头竞争下LLM定价代理的共谋","abstract":"Large language models (LLM) deployed as autonomous pricing agents may sustain supracompetitive prices through tacit coordination. We develop a causal graph divergence framework that separately measures structural faithfulness and intent faithfulness of LLM pricing agents in Bertrand competition. Across nine LLMs under duopoly and triopoly conditions, collusive behavior and chain-of-thought (CoT) faithfulness dissociate along both dimensions: the most collusive model accurately reports cooperative intent yet reasons structurally unfaithfully, while the most structurally faithful model sustains supra-Nash pricing under both market structures. These findings establish that CoT monitoring alone cannot serve as a standalone safeguard against algorithmic collusion.","authors":["Dohun Lee","Hyunwoo Park"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18346","pdf_url":"https://arxiv.org/pdf/2609.18346","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM定价代理","算法共谋","经济仿真"],"reason":"用LLM定价agent模拟寡头竞争，与人类行为对照，但非直接仿真人类被试，结论…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":15,"question":"LLM定价智能体在寡头竞争中能否通过思维链监控可靠地检测合谋行为？","design":"用9个LLM作为定价智能体，在双寡头和三寡头的伯特兰竞争环境中进行300轮定价，施加利润导向和竞争导向两种提示处理，测量定价序列、利润、思维链中的因果图与意图。","baseline":"无对照","findings":"合谋行为与思维链忠实性在两个维度上分离：最合谋的模型准确报告合作意图但结构推理不忠实，而结构最忠实的模型在两种市场结构下均维持超纳什定价。思维链监控不能单独作为算法合谋的保障。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济主体行为，虽非直接仿真人类被试，但涉及寡头定价与合谋检测，对关注LLM仿真可靠性及政策评估的研究者有参考价值，值得读原文了解其因果图分歧框架。","inspiration":"借鉴其因果图分歧框架，分别测量陈述与行为因果结构及意图分布，以评估LLM仿真忠实性｜可迁移到寡头定价、拍卖合谋、平台算法共谋等产业组织与反垄断场景｜用LLM作为定价智能体，施加不同提示（如利润导向vs竞争导向），测量定价序列与思维链，以真实市场定价数据或人类实验数据为基准，检验LLM合谋行为与思维链监控的有效性。"}},{"id":"2609.18390","version":1,"title":"Building a Cultural Perspective on Doctor-Patient Conversations","zh_title":"构建医患对话的文化视角","abstract":"AI-powered medical scribes are increasingly used to transcribe doctor-patient conversations and automate clinical documentation. However, large-scale real-world consultation datasets are scarce due to the sensitivity of clinical conversations, leading developers to rely on simulated and LLM-generated synthetic consultations. While scalable, these alternatives may fail to capture culturally situated patterns of clinical interaction. We introduce interactional cultural markers, measurable patterns of doctor-patient interaction grounded in cross-cultural clinical communication, and use them to compare real, simulated, and synthetic consultations from Indian and US clinical contexts. We find distinct patterns of participation and control: Indian consultations involve greater patient participation but stronger doctor control, while US consultations exhibit balanced participation and open-ended discussion. Synthetic Indian consultations often fail to reproduce these patterns, instead converging toward US-like interaction. We identify additional synthetic signatures, including excessive doctor explanation and formulaic patient responses. We conclude by discussing implications for generating culturally grounded synthetic clinical conversations.","authors":["Krithi Shailya","Siddharth D Jaiswal","Ashish Makani","Suvrankar Datta","Sunayana Sitaram","Mohit Jain"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18390","pdf_url":"https://arxiv.org/pdf/2609.18390","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","医患对话","文化差异"],"reason":"用LLM生成合成医患对话并与真实数据对照，评估文化模式复现，属仿真人类交互且含…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":16,"question":"不同来源（真实、模拟、合成）的医患对话在多大程度上再现了印度与美国临床互动中的文化模式？","design":"使用LLM生成合成医患对话（包括单智能体和多智能体架构，以及是否基于临床笔记的四种设置），并与真实和模拟对话进行比较，通过多层级互动文化标记（L1转录级、L2话轮间、L3话轮级）测量参与度、控制、语气等互动特征。","baseline":"来自印度和美国的真实医患咨询数据（包括英语、印地语、泰卢固语），以及由医疗专业人员或患者演员编写的模拟咨询。","findings":"印度真实咨询中患者参与度更高但医生控制更强，美国则更平衡且开放；合成印度对话往往无法再现这些模式，反而趋向美国式互动，并出现医生过度解释、患者公式化回应等合成特征。","reliability":"论文指出合成对话可能无法捕捉文化细微差别，且生成策略引入权衡：基于笔记的生成抑制口语和语码转换，而智能体生成则放大冗长和共情至不现实水平。","relevance":"该研究直接使用LLM生成合成对话并与真实人类数据对照，评估文化模式复现的可靠性，属于人类仿真研究，且包含批判性发现，值得阅读原文以了解其标记框架和失效条件。","inspiration":"值得借鉴的是其构建多层级可量化互动标记并系统比较真实与合成数据的方法，可迁移到经济金融中的跨文化沟通或谈判实验，例如不同文化背景下的信贷协商或政策沟通；具体设计可用LLM生成不同文化背景的谈判对话，以真实谈判记录为基准，测量话轮控制、信息共享等标记，检验合成数据是否复现文化差异。"}},{"id":"2609.16395","version":1,"title":"Silicon sampling answers with country-level assumptions, not individual attitudes: Cross-national evidence from the European Social Survey","zh_title":"硅采样以国家层面假设而非个体态度作答：来自欧洲社会调查的跨国证据","abstract":"Silicon sampling uses large language models (LLMs) to simulate survey respondents. Whether it recovers cross-national variation, and why, remains unresolved. This study evaluates it against European Social Survey Round 11 (30 countries, 42 items) with two open-weight LLMs under first- and third-person prompts, plus backstory and response-format experiments. Aggregate recovery is moderate and uneven across items. Adding the country name to a three-variable demographic backstory raises the median per-item correlation between simulated and observed country means from -0.03 to 0.52, and the richer profiles tested add no consistent gain. The respondent's country label acts as a country-level assumption that respondent detail does not revise. Naming the response-scale endpoints in words stops the model from ranking countries backwards, so the answer format sets the direction of the ranking. Individual-level recovery remains negligible in every condition and does not track aggregate recovery across countries. An average of neighboring countries, using no LLM, recovers country levels more accurately than every model condition and ranks them about as well. Silicon sampling can thus support exploratory country-ranking comparison after item-level validation and with the response format reported. It does not support individual or distributional inference.","authors":["Chuyao Wang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16395","pdf_url":"https://arxiv.org/pdf/2609.16395","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","调查方法","跨国比较"],"reason":"直接评估LLM仿真调查回答，与真实跨国调查数据对照，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":3,"question":"硅采样能否恢复跨国调查中的国家间差异，以及这种恢复的来源是什么？","design":"用两个开源权重LLM模拟欧洲社会调查（ESS）第11轮30个国家的受访者，在42个态度题项上生成回答；通过第一/第三人称提示、添加国家名称和人口学背景故事、以及改变回答格式端点措辞等实验条件，测量模拟国家均值与真实均值的相关性及个体层面恢复情况。","baseline":"欧洲社会调查（ESS）第11轮30个国家、42个题项的真实人类回答数据。","findings":"国家标签是主要的聚合信号来源，添加国家名称后中位题项相关从-0.03升至0.52，更丰富的人口学背景没有额外增益；回答格式端点措辞决定国家排序方向，个体层面恢复在所有条件下均可忽略，且与聚合恢复不相关。","reliability":"论文指出硅采样仅适用于探索性国家排序比较，且需逐题验证并报告回答格式；不支持个体或分布推断，个体恢复不随聚合恢复变化。","relevance":"该研究直接评估LLM仿真调查回答，与真实跨国调查数据对照，并明确指出失效条件，对关注人类仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其通过消融实验分离国家标签与人口学背景贡献的方法，以及改变回答格式端点来检验排序方向的做法｜可迁移到跨国经济态度调查仿真，如通胀预期、政策偏好或金融素养的跨国比较｜用LLM模拟不同国家受访者对经济政策的态度，处理为是否添加国家标签及回答格式端点措辞，结果变量为模拟国家均值排序，对照真实跨国调查数据（如欧洲央行消费者预期调查）验证。"}},{"id":"2609.17317","version":1,"title":"Towards Detecting AI-Assisted Responses in Online Surveys","zh_title":"检测在线调查中AI辅助回答的方法研究","abstract":"The use of LLMs to complete online surveys impacts the validity of survey-based research, but detecting such usage remains underexplored. We introduce an initial benchmark dataset, namely ASURRE, for AI-assisted survey participation to capture usage strategies ranging from full generation and revision to persona-grounded agentic completion. Controlled by these strategies, LLM-assisted survey responses are generated using multiple LLMs on three real-world surveys in different disciplines, paired with genuine human responses. Our evaluation of existing machine-generated text (MGT) detectors shows that naive AI usage is readily detectable, whereas persona-grounded agents that mimic entire respondents push detector performance toward chance. We further show that agentic completion cannot fully replicate respondent-level behaviour and leaves distinctive behavioural traces. While individual cues can be circumvented by targeted prompting, a simple few-shot, training-free aggregator over these cues improves mean AUROC by +0.14 over the best existing detector across agentic settings. Our project is available at https://github.com/mike-qz-wang/ASURRE.","authors":["Qizhou Wang","Bogdan Mamaev","Christopher Leckie"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17317","pdf_url":"https://arxiv.org/pdf/2609.17317","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查数据","检测方法"],"reason":"直接研究LLM仿真人类调查回答，并与真实人类数据对照，评估检测与行为差异。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":10,"question":"如何检测在线调查中由大语言模型辅助生成的回答，尤其是在不同使用策略下的可检测性？","design":"构建ASURRE基准数据集，使用多个开源和闭源LLM在三个真实调查上生成回答，模拟三种使用策略：完全生成、修订和基于人格的智能体完成；评估现有机器生成文本检测器，并提出SPABD聚合器。","baseline":"三个真实调查中的人类真实回答，与LLM生成回答配对。","findings":"完全生成易被检测，修订部分可检测，而基于人格的智能体完成几乎无法被现有检测器识别；SPABD聚合器通过整合多个行为线索，在智能体设置下将平均AUROC从0.61提升至0.75。","reliability":"论文承认SPABD在对抗性提示下性能下降，且在不同调查和提示模式下表现不稳定，应视为轻量级参考检测器而非完整解决方案。","relevance":"该研究直接评估LLM仿真人类调查回答的可检测性，并与真实人类数据对照，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其构建多策略仿真基准和利用行为痕迹进行检测的方法，可迁移到经济金融领域的调查数据质量评估，例如消费者信心调查或投资者情绪调查；设计研究时，可用LLM模拟不同人格的受访者，施加不同提示策略，以真实调查数据为基准，评估检测器性能并分析行为偏差。"}},{"id":"2609.16436","version":1,"title":"Interpreting and Steering LLM Agents for Social Simulations","zh_title":"解释与引导用于社会模拟的LLM智能体","abstract":"Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive specific model behavior. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. Specifically, we compare three types of methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering and examine their utility for LLM-based social scientific simulations. We do so by interpreting and steering two foundational components of human behaviors, namely preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation), operationalized using four classic economic and creative tasks implemented as natural-language interactions. Overall, our results show that SAE- and probe-based techniques often outperform basic prompt-based methods for steering LLM agents, although this advantage depends on the specific prompting strategy involved. Together, SAEs and probes constitute an effective pipeline for social scientists seeking to interpret and steer agents in social simulations: SAEs decompose agents' internal representations into human-readable features, after which probes can reliably shift agents' behaviors in specified directions. We discuss implications of these methods for future work using LLM agents for social scientific simulations.","authors":["Jiayue Gaveal Fan","Arul Murugan","Shreyas Krishnan","Abhishek Nagaraj"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16436","pdf_url":"https://arxiv.org/pdf/2609.16436","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","可解释性","行为经济学"],"reason":"用LLM仿真人类行为，有真实人类数据对照，涉及经济任务，并批判性评估方法。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":9,"question":"如何打开LLM黑箱，通过可解释性和可操控性方法提升LLM智能体在社会仿真中的效用？","design":"使用LLM智能体模拟人类被试，在四个经典任务（彩票游戏测风险态度、最后通牒游戏测利他、发散创造力、产品创新）中，比较三种干预方法：基于提示的操控、SAE特征操控、探针方向操控，测量智能体行为变化。","baseline":"无对照","findings":"SAE和探针方法在操控LLM智能体行为上通常优于基本提示方法，但优势取决于具体提示策略；SAE能无监督地发现行为背后的可解释特征，探针能精确控制特定特征的程度。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真中的可解释性和可操控性，提供了超越提示工程的技术路径，对关注仿真可靠性和机制理解的研究者具有重要参考价值。","inspiration":"借鉴其使用SAE和探针进行特征级操控的方法，可实现对LLM智能体内部表征的精细干预，比提示词更可控。｜可迁移到经济决策仿真，如风险偏好、时间偏好、社会偏好等实验，以及政策评估中的行为反应模拟。｜以LLM智能体模拟投资者，用探针操控其风险厌恶特征，观察在资产配置任务中的选择，并与真实投资者调查或实验数据对照，验证操控的有效性和仿真保真度。"}},{"id":"2609.07474","version":3,"title":"Where Should Language Sit in a Multimodal Model? Lessons from What Language Does to Human Perception and Cognition","zh_title":"语言在多模态模型中的位置：从语言对人类感知和认知的影响中汲取的教训","abstract":"Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should keep all of these positions depends on what language does to the system that uses it. The one system with a century of data on that question is the human. We review what language does to human perception, the brain, and thought, and read the same evidence against multimodal models and language models. Throughout, we treat language as a compressor that runs on a shared codebook: a word is an index, the content is in the receiver, and a community maintains the codebook. In humans the compression is measurable, learning the codebook reorganizes the senses, and thought survives the loss of language. We then measure the rule that models apply when two cues disagree, with cue-conflict experiments on six vision-language models and two robot policies. Surviving cues are weighted in the order their reliabilities prescribe, at 11 to 82\\% of the ideal observer's slope, and many answers copy the text. One policy family drops a cue that adds no information beyond the others rather than down-weighting it, another keeps it at a weight that fails when the cues conflict, and a visual cue that identifies the task in every training frame is never learned, because the language pathway already fits the data. Language models are the best current models of the human language network, and they have entered the human speech community, shifting word frequencies while alignment narrows their conceptual diversity. We close with seven implications for token-based systems. Language belongs at a model's boundary and in the shared codebook, as in the brain, not as its internal representation; the price of leaving the codebook inside is auditability.","authors":["Peng Xie","Amr Alanwar"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-09","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.07474","pdf_url":"https://arxiv.org/pdf/2609.07474","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["多模态模型","人类感知对照","模型评估"],"reason":"论文用人类感知数据对照多模态模型，评估模型行为与人类差异，方法可迁移至LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":14,"question":"语言在多模态模型中的位置应如何安排？论文通过人类感知与认知证据及模型线索冲突实验，探讨语言是否应作为内部表征。","design":"论文并非传统仿真研究，而是通过回顾人类感知、大脑和思维中语言的作用，并对六个视觉语言模型和两个机器人策略进行线索冲突实验，测量模型在图像与文本线索不一致时的权重分配。","baseline":"人类感知与认知数据，包括颜色词汇、语音感知、失语症患者思维等心理学和神经科学证据。","findings":"模型在冲突线索中按可靠性加权，但权重仅为理想观察者的11%至82%，且常复制文本。语言模型是最佳的人类语言网络模型，但已进入人类语言社区，改变词频并收窄概念多样性。","reliability":"论文指出语言模型缺乏感官基础，仅从文本学习代码簿，可能无法获取某些概念；多模态模型常忽略视觉线索，且内部语言表征损害可审计性。","relevance":"该论文对使用LLM进行人类仿真有重要启示：它揭示了语言模型在感知和决策中的偏差，提示仿真需谨慎处理语言与感官信息的整合，值得精读以理解模型失效条件。","inspiration":"借鉴线索冲突实验设计，通过操纵文本与视觉信息的不一致来测量模型权重，可迁移到经济金融中的信息处理场景，如投资者对财报文本与图表信息的权衡。｜设计一个实验：用LLM作为被试，呈现公司财报摘要（文本）与股价走势图（视觉），两者对盈利前景给出矛盾信号，测量模型预测的盈利预期或投资决策，并与人类分析师在相同任务上的真实数据对照，评估模型是否过度依赖文本。"}},{"id":"2609.15996","version":1,"title":"Comment on arXiv:2607.01233: Survivorship Bias in Published-Paper Baselines for Research-Idea Distributions","zh_title":"评论 arXiv:2607.01233：已发表论文基线中的幸存者偏差对研究想法分布的影响","abstract":"Chen, Zhao, and Cohan introduce a valuable distributional evaluation of LLM-generated research ideas. This comment raises a narrower identification concern: their human baseline consists of published papers, whereas the LLM baseline consists of one-shot proposals. If bridge-like or synthesis-like ideas are relatively easy to generate but relatively unlikely to survive publication, then the published human baseline will understate their prevalence in the unseen human idea pool. The observed human--LLM gap may therefore be partly, or even largely, a consequence of survivorship bias.","authors":["Fredrik A. Dahl"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15996","pdf_url":"https://arxiv.org/pdf/2609.15996","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM仿真","幸存者偏差","研究想法生成"],"reason":"评论指出LLM生成研究想法与人类已发表论文对比存在幸存者偏差，涉及仿真可靠性评…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":15,"question":"LLM生成的研究想法与人类已发表论文中的研究想法分布差异，是否可能由幸存者偏差而非LLM特有的研究品味差异造成？","design":"本文是一篇评论性文章，未进行新的仿真实验；它通过逻辑推演和示意图，指出原论文将LLM一次性生成的想法与人类已发表论文的想法分布进行对比，存在阶段不匹配问题，并构建了一个反事实分布来说明幸存者偏差如何解释观察到的差异。","baseline":"原论文的人类基准是已发表论文中的研究想法分布，但本文指出该基准是经过发表筛选后的幸存者分布，不能代表人类未经过滤的想法池。","findings":"原论文发现LLM生成的想法更集中于桥接式和综合式类别，而人类已发表论文中此类想法较少；本文认为这一差距可能部分或大部分源于幸存者偏差，因为桥接式想法容易生成但难以发表，因此在已发表论文中代表性不足。","reliability":"本文指出原论文的对比存在幸存者偏差，需要阶段匹配的比较（如让人类在相同条件下生成一次性想法）或抽样“想法墓地”（被拒稿、放弃的草稿等）来验证；在缺乏此类证据前，不能得出LLM具有特定研究品味差异的强结论。","relevance":"本文对使用已发表论文作为人类基准的LLM仿真研究提出了关键的识别问题，提醒研究者在比较LLM输出与人类数据时需注意选择偏差，值得阅读原文以理解幸存者偏差如何影响分布比较的可靠性。","inspiration":"本文的方法论启示在于强调基准数据的生成过程必须与LLM输出阶段匹配，否则会引入选择偏差，这可以迁移到经济金融研究中任何涉及LLM生成内容与真实世界数据对比的场景，例如政策评估或市场预期分析。｜例如，在资产定价实验中，若用LLM模拟投资者观点并与已发表的研报观点对比，可能因研报经过筛选而低估某些常见但低质量的观点。｜一个可行的研究设计是：让LLM和人类被试在相同信息集下生成对某经济指标的未来预测，然后分别与未经过滤的实时预测记录（如社交媒体帖子或调查原始数据）和经过发表筛选的预测（如专业机构报告）进行对比，以量化幸存者偏差对LLM-人类差异的影响。"}},{"id":"2609.16501","version":1,"title":"Beyond the Name: Demographic Leakage in De-Identified R\\'esum\\'es and Evaluation Artifacts in LLM Bias Audits","zh_title":"超越姓名：去标识化简历中的人口统计泄漏与LLM偏见审计中的评估伪影","abstract":"De-identified r\\'esum\\'e screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields resolves this leakage across nine open-weight models and 620 counterfactual r\\'esum\\'es. By holding language attributes strictly identical, we isolate unstructured prose across five ethnocultural conditions and three cue-salience tiers. Target-group recovery averages 0.757 overall and saturates at 1.000 under high salience, demonstrating that non-language prose sustains demographic inference. Crucially, models diverge only under faint cues (0.086-0.690), establishing salience as an essential evaluation axis. Furthermore, pairwise LLM-as-a-judge outcomes are highly sensitive to evaluation design: forbidding ties yields an apparent selection-rate ratio of 0.39 alongside strong position and content effects, whereas permitting ties produces near-universal ties for most models ($\\ge94\\%$). Downstream scoring shows only very small between-condition differences, highlighting the need to distinguish demographic signals recoverable from r\\'esum\\'e content from effects introduced by the evaluation protocol.","authors":["Qiangju Chen","Yang Xiao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16501","pdf_url":"https://arxiv.org/pdf/2609.16501","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏见审计","评估协议","人口统计推断"],"reason":"审计LLM偏见，揭示评估协议引入的伪影，与仿真可靠性评估相关","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":19,"question":"在控制语言字段后，非语言简历文本是否仍能泄露族裔背景，以及LLM评审协议是否制造了偏见伪影？","design":"构建620份反事实简历，固定语言声明为英语，仅操纵附加信息部分的非语言散文，设置五种族裔条件和三个线索显著性层级，用九个开源权重模型执行三项任务：个体背景恢复、成对偏好审计和下游评分。","baseline":"无对照","findings":"非语言散文仍能实现平均0.757的族裔恢复率，高显著性下达到1.000；模型仅在微弱线索下表现出差异（0.086-0.690）。成对LLM评审在禁止平局时产生0.39的虚假选择率比，并伴随位置和内容效应；允许平局时大多数模型产生近乎全平局（≥94%）。","reliability":"论文指出评审协议的设计（如禁止平局、选项顺序）会引入伪影，导致虚假歧视信号；下游评分仅显示很小的条件间差异，需区分简历内容中的可恢复人口信号与评估协议引入的效应。","relevance":"该研究直接评估LLM在简历筛选中的偏见审计可靠性，揭示评估协议伪影，与仿真可靠性评估高度相关，值得精读以理解LLM作为人类被试替代品时的偏差来源。","inspiration":"可借鉴其反事实设计与线索显著性分层方法，通过严格控制结构化字段并操纵非结构化文本，分离真实信号与协议伪影。｜可迁移到信贷审批歧视研究，用LLM模拟信贷员对贷款申请中非结构化叙述的族裔推断与决策偏差。｜以LLM为被试，构造仅改变申请文本中族裔相关线索（如社区活动、兴趣）的贷款申请，固定结构化字段，测量LLM的族裔恢复率和审批决策差异，并与真实信贷审批数据中的族裔差异对照，检验仿真有效性。"}},{"id":"2609.16517","version":1,"title":"Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening","zh_title":"保持能力不变的简历扰动揭示LLM筛选中的呈现敏感性","abstract":"Resume screeners must infer job-relevant competence from resumes whose presentation can vary substantially in wording, structure, stylistic polish, and document extraction quality. Ideally, such surface variation should not change decisions when the underlying qualification evidence is unchanged. We introduce a controlled audit of this property, constructing occupation-grounded candidate profiles at controlled competence levels and rendering each profile into multiple resume presentations. A deterministic validation gate excludes variants that alter the underlying evidence before scoring. Across six open instruction-tuned LLM conditions, we find a clear disconnect between screening validity and presentation stability. Llama-3.1-8B with its native chat template achieves the strongest validity ($0.781$) yet reverses $29.6\\%$ of matched pairwise decisions under competence-preserving presentation changes; Mistral-7B-v0.3 reaches validity $0.644$ with a $41.4\\%$ flip rate. Native chat formatting improves validity for several chat-tuned models but does not remove this instability. These results show that resume-screening evaluations should assess not only whether a system identifies stronger candidates, but also whether those decisions remain stable when the same competence evidence is presented differently.","authors":["Qiangju Chen","Yang Xiao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16517","pdf_url":"https://arxiv.org/pdf/2609.16517","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM决策仿真","简历筛选","呈现偏差"],"reason":"用LLM模拟简历筛选决策，与人类判断对照，揭示呈现敏感性","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":20,"question":"在简历筛选任务中，当候选人的能力证据保持不变时，简历的呈现方式变化（如措辞、结构、风格润色、文档提取质量）是否会导致大语言模型筛选决策的不稳定？","design":"使用六个开源指令微调大语言模型（如 Llama-3.1-8B、Mistral-7B-v0.3 等）扮演简历筛选者，基于 O*NET 职业数据库构建 17 个职业、102 个候选人档案（每个职业 2 个合格、2 个边缘、2 个不合格），每个档案渲染成 5 种简历呈现形式（原始、冗长、段落化、AI 润色、布局噪声），通过确定性验证门排除改变证据的变体，然后让模型对每份简历独立打分，比较同一候选人不同呈现下的决策翻转率，以及模型对已知优劣候选人排序的效度。","baseline":"无对照（研究未使用真实人类筛选者数据，而是通过构造的候选人档案和已知能力层级作为基准来评估模型效度）","findings":"模型筛选效度与呈现稳定性之间存在明显脱节：Llama-3.1-8B 在原生聊天模板下效度最高（0.781），但在能力保持的呈现变化下仍有 29.6% 的成对决策翻转；Mistral-7B-v0.3 效度为 0.644，翻转率高达 41.4%。原生聊天格式能提高部分模型的效度，但无法消除这种不稳定性。","reliability":"论文未讨论（正文节选未提及失效条件或局限，但研究本身通过确定性验证门排除了改变证据的变体，并指出呈现敏感性是自然存在的异质性来源）","relevance":"该研究直接针对 LLM 在简历筛选中的呈现敏感性，通过受控审计分离能力与呈现，并测量决策翻转率，为评估 LLM 仿真人类决策的可靠性提供了关键证据，值得精读原文以了解具体实验设计和模型差异。","inspiration":"借鉴其通过受控扰动分离核心信息与表面呈现、并用确定性验证门确保处理干净的做法，可迁移到信贷审批或保险定价等经济决策场景中，研究 LLM 对同一申请人信息的不同表述（如收入证明格式、信用报告排版）是否产生不一致决策；可设计实验：用 LLM 扮演信贷员，对同一借款人的信用档案进行多种文本呈现（如改变措辞、段落结构、添加无关信息），测量贷款批准决策的翻转率，并与真实信贷审批数据或人类信贷员判断进行对照。"}},{"id":"2609.16993","version":1,"title":"The Role of Implicit and Explicit Demographic Signals in Large Language Model-based Student Assessment","zh_title":"大语言模型学生评估中隐式与显式人口统计信号的作用","abstract":"Large Language Models are now common in student assessment, but we know little about how student demographics affect their use. Sometimes, considering student demographics may be necessary -- for example, to improve readability for users with lower educational levels. However, it also risks being a cause of discrimination, e.g., when assigning lower scores to students from lower socioeconomic backgrounds. We set up controlled prompts to test 1) explicit demographic effects, where we mention demographic details directly, and 2) implicit effects, where we use conversation history as a demographic signal. We test these settings in three tasks: Automated Essay Scoring, Formative Feedback, and Metalinguistic Question Answering. We test six state-of-the-art LLMs on these tasks. In both explicit and implicit cases, the models pick up on demographic cues and can change their scoring, feedback, and answers accordingly. We find that LLMs frequently adjust the readability of feedback to education levels when these are explicitly mentioned. On the other hand, implicit conditions produce unpredictable biases, such as in question answering, where responses from lower-education levels receive lower sentiment scores. Our results provide clear evidence of demographic sensitivity in LLMs for educational assessment tasks.","authors":["Donya Rooein","Luca Benedetto","Dirk Hovy"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16993","pdf_url":"https://arxiv.org/pdf/2609.16993","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","教育评估","偏差分析"],"reason":"用LLM模拟学生评估，有真实人类数据对照，并揭示偏差，可迁移到仿真可靠性研究。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":22,"question":"显式和隐式人口统计线索是否系统性地影响LLM在学生评估任务中的行为，以及这种影响在不同教育任务间有何差异。","design":"用六种最先进的LLM扮演学生评估者，在自动作文评分、形成性反馈和元语言问答三个任务中，通过显式（直接提及人口统计细节）和隐式（利用对话历史作为信号）两种方式施加人口统计处理，测量评分、反馈和回答的变化。","baseline":"人类基准：自动作文评分任务中使用了真实人类评分作为对照（人类平均分），其他任务未提及人类对照。","findings":"LLM在显式和隐式条件下都会捕捉人口统计线索并改变评分、反馈和回答；显式条件下LLM常根据教育水平调整反馈可读性，隐式条件则产生不可预测的偏差，如低教育水平用户在问答中收到更低情感分数。","reliability":"论文未讨论","relevance":"该研究直接评估LLM在模拟人类评估者时对人口统计特征的敏感性，揭示了仿真中的偏差，与研究者关注LLM仿真可靠性和偏差的核心问题高度相关，值得阅读原文以了解具体偏差模式和任务依赖性。","inspiration":"借鉴其显式与隐式人口统计线索的分离设计，以及多任务、多模型对比和统计检验方法，可迁移到信贷审批歧视或保险定价等经济金融场景，例如用LLM扮演信贷员，在贷款申请中显式或隐式加入申请人性别、种族或收入信号，测量贷款批准决策和利率设定，并与真实信贷审批数据对照，检验LLM是否复现或放大人类偏见。"}},{"id":"2609.17496","version":1,"title":"Verifiable Social Reasoning for LLM Assistants","zh_title":"面向LLM助手的可验证社交推理","abstract":"LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others' intentions, typically lack verifiable ground truth. To address these challenges, we introduce Fuse, a multi-agent simulation framework for studying user-mediated social reasoning. In Fuse, a target agent with a hidden motive interacts with other agents including one representing the user, who then consults the evaluated assistant to infer the target's motive, providing verifiable ground truth by construction. Simulation faithfulness is validated through a human study with 24k annotations. We apply Fuse to 12 LLMs and demonstrate its analytical utility by systematically isolating key factors, showing that (i) user mediation compounds the inherent difficulty of social reasoning; (ii) LLMs exhibit systematic sensitivity to biased user framing; (iii) models can require more details than humans need to reach a correct prediction; and (iv) longer conversations do not always improve performance despite providing opportunities for clarifying questions. We open-source Fuse and a dataset with 21k examples.","authors":["Amir Taubenfeld","Zorik Gekhman","Avigail Grinstein-Dabush","Itay Laish","Ariel Goldstein","Marian Croak","Avinatan Hassidim","Yossi Matias","Amir Feder"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17496","pdf_url":"https://arxiv.org/pdf/2609.17496","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B4"],"tags":["多智能体仿真","社交推理","人类对照"],"reason":"多智能体仿真人类社交推理，有人类标注对照，但非直接复现人类被试行为","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":23,"question":"如何评估LLM助手在用户主观叙述中介下的日常社交推理能力，并识别其系统性偏差？","design":"Fuse框架：用LLM（Gemini 3.1 Flash-Lite）扮演用户、目标人物和其他角色，在多场景中模拟社交互动；目标人物有隐藏动机，用户向被评估的LLM助手咨询以推断该动机。通过改变用户报告偏差（默认 vs. 对立信念）和叙述细节水平（三档）施加处理，结果变量为助手预测隐藏动机的准确率。","baseline":"人类研究：24k条标注，验证模拟保真度并估计人类从首条用户消息识别动机的准确率为88%。","findings":"用户中介增加了社交推理的固有难度；LLM对用户框架存在系统性敏感，且可能需要比人类更多的细节才能正确预测，更长的对话并不总能提高性能。","reliability":"论文未明确讨论失效条件，但指出模拟保真度通过人类研究验证，且任务难度校准基于人类表现；局限包括仅使用单一LLM生成模拟数据，以及首条消息焦点可能限制对多轮交互的全面评估。","relevance":"该研究将LLM作为人类被试的替代品，在受控仿真中评估社交推理，并提供了人类基准对照，对关注LLM仿真可靠性与偏差的研究者有参考价值，值得阅读原文以了解其框架和发现。","inspiration":"借鉴其多智能体仿真框架和通过改变用户报告偏差与细节水平来施加处理的方法，可系统分析LLM在信息中介下的决策偏差。｜可迁移到经济金融中的信息传递与决策场景，如投资者根据分析师报告或新闻叙述做出投资决策、消费者根据口碑信息进行购买决策、或信贷审批中基于申请人陈述的评估。｜设计一个实验：用LLM扮演信息发送者（如公司管理层或分析师）和接收者（投资者），处理为发送者的报告偏差（乐观/悲观）和信息详细程度，结果变量为LLM投资者的估值或投资决策准确率，并与真实人类实验数据（如实验室资产定价实验）进行对照，以评估LLM仿真人类信息处理偏差的可靠性。"}},{"id":"2609.16793","version":1,"title":"Available but Unclaimed: An Empirical Study of Human-AI Synergy","zh_title":"可用但未认领：人类与AI协同的实证研究","abstract":"People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each assisted trial required consultation with the model. Each model answered every item alone 100 times under matched elicitation. The assisted-unaided accuracy difference increased with item-level LLM competence. Deference varied across tasks and increased with competence within tasks. Post-advice confidence distinguished correct from incorrect answers less strongly than unaided confidence. In a reference comparison, about half the increase in LLM accuracy carried through to assisted accuracy. How much of that accuracy gain reached participants differed across the models. These findings motivate evaluating LLMs in interaction with humans and designing support for selective deference that preserves independent reasoning.","authors":["Robin Welsch","Michelle Rausch","Pascal Knierim","Thomas Kosch","Jochen Kuhn","Albrecht Schmidt","Daniela Fernandes"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16793","pdf_url":"https://arxiv.org/pdf/2609.16793","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["人机协同","认知任务","实验研究"],"reason":"研究人类与LLM协作解题，有真实人类数据对照，但非LLM仿真人类被试，而是人机…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":21,"question":"人类与LLM协作解题时，能否实现超越各自单独表现的协同效应，以及这种协同在多大程度上取决于模型能力、用户依赖和任务类型？","design":"本研究不是LLM仿真人类被试，而是真实人类与LLM协作实验。535名参与者被随机分配到无辅助组或四个LLM辅助组（GPT-5.6-Luna、Claude Opus 4.8、Gemini 3.6 Flash、Kimi K3），完成40道涵盖矩阵推理、心理旋转、三段论和字母串类比的题目。辅助组每次作答前必须咨询模型。每个模型在相同题目上独立运行100次以估计其项目级能力。结果变量为答题准确率、对建议的采纳（deference）以及作答后的信心。","baseline":"无辅助的人类答题准确率作为基准，同时每个LLM单独答题的准确率也作为对照。","findings":"辅助与无辅助的准确率差异随LLM在项目上的能力提高而增大；对建议的采纳在不同任务间有差异，并在任务内随能力提高而增加。参考对比显示，LLM准确率的提升只有约一半转化为辅助准确率的提升，且不同模型间转化比例不同。","reliability":"论文指出，互补性错误并不保证协同，用户经常错误采纳建议；信心虽能区分对错，但不足以支持选择性依赖。研究限于特定认知任务和强制咨询设置，未探讨自由交互或长期使用。","relevance":"该研究虽非LLM仿真人类，但提供了人类与LLM协作的实证基准，对理解LLM作为决策辅助工具时的偏差和可靠性有参考价值，值得一读以了解人机协同的边界条件。","inspiration":"借鉴其项目级能力测量和强制咨询设计，可迁移到经济决策场景如投资建议采纳或信贷审批辅助。设计：招募真实投资者或信贷员，随机分配至无辅助或LLM辅助组，处理为强制咨询LLM建议，结果变量为决策准确率或收益，对照真实历史数据或专家决策。"}},{"id":"2608.18768","version":2,"title":"Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model","zh_title":"可读、忠实、被使用：语言模型中人口统计身份的三个可分离属性","abstract":"Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differences, and whether this reflects what a model knows or uses has remained untested. Using representational similarity analysis against Pew American Trends Panel ground truth, we score demographic read-out locations in Mistral-7B and intervene causally across six attribute types. The internal geometry is faithful: attention-head read-outs dominate the standard residual read-out, reaching selection-corrected $\\rho$ up to 0.63 -- about 70% of the measurement-reliability ceiling -- and one head, L11 H16, is significantly faithful across all six types, though race-based types stay weak and prompt-fragile, replicating in a second model family. Yet causal use does not track fidelity: the clearest causal pathway ($p=0.002$) sits in one of the least faithful types, the most faithful type shows no correction-surviving effect, and full identity swaps in the prompt move predictions by under 2% of their error. A 128-dimensional probe on that head lands 21-31% closer to survey truth than the model's answers, yet recovers almost none of the per-question group ordering. Readable, faithfully arranged, and causally used are three dissociable properties of the same model; treating them as one claim is what keeps the \"can LLMs simulate populations\" debate unresolved.","authors":["Fathin Difa Robbani"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-20","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.18768","pdf_url":"https://arxiv.org/pdf/2608.18768","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","算法忠实度","调查方法"],"reason":"直接研究LLM仿真调查受访者，用真实Pew数据对照，评估忠实度与因果使用，并批…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":1,"question":"LLM内部的人口统计身份表征是否忠实于真实群体差异，以及这些表征是否被模型实际用于生成回答？","design":"使用Mistral-7B模型，对169个交叉人口统计单元构建身份提示，提取残差流、注意力头输出和FFN输出等内部表征，与Pew调查的真实群体回答分布进行表征相似性分析（RSA），并通过激活修补进行因果干预，测量模型预测分布的变化。","baseline":"Pew American Trends Panel（ATP）调查数据，包含15波、169个交叉人口统计单元的真实回答分布。","findings":"内部几何结构是忠实的：注意力头读出（尤其是L11 H16）与真实群体差异的相关性高达0.63，约为测量可靠性上限的70%，且跨六种属性类型显著。但因果使用与忠实性脱节：最清晰的因果路径位于低忠实性类型中，而最忠实的类型没有通过校正的因果效应；完整身份替换仅使预测移动不到误差的2%。","reliability":"论文承认种族和宗教类型的忠实性较弱且对提示脆弱；忠实表征并未被模型实际用于生成回答；探针虽在平均上更接近真实，但无法恢复每个问题的群体排序，因此不能替代调查数据。","relevance":"直接研究LLM仿真调查受访者的可靠性，用真实Pew数据对照，并批判性地指出忠实表征与因果使用脱节，对理解仿真失效条件至关重要，值得精读原文。","inspiration":"借鉴其表征相似性分析与因果干预结合的方法，可系统评估LLM内部经济偏好表征与真实行为的一致性。｜可迁移到信贷审批歧视研究，检验模型内部对种族、收入等群体的风险表征是否忠实于真实违约数据，以及这些表征是否影响审批决策。｜用LLM扮演信贷员，输入不同人口统计特征的贷款申请，测量其审批决策和内部表征；以真实信贷数据（如HMDA）为基准，比较模型表征与真实群体违约率的相似性，并通过激活修补检验因果路径。"}},{"id":"2609.15849","version":1,"title":"Before You Poll with LLMs: A Deliberative Diagnostic Framework","zh_title":"用LLM进行民意调查前：一个审议诊断框架","abstract":"Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM personas simulate public opinion at scale. Current evaluations test only whether personas hold the right opinions -- a static snapshot. But opinion research increasingly depends on dynamic fidelity: whether personas update beliefs in response to new arguments, as humans do during deliberation. No existing benchmark tests this. We introduce the Deliberative Polling Diagnostic Framework, which compares human and LLM belief shifts after identical informational interventions. Grounded in deliberative polling, it surfaces failures invisible to static evaluation: models that produce plausible partisan opinions can still misrepresent how those opinions change. Applying the framework to five frontier models using data from America in One Room (526 personas, 72 questions), we find that every model fails, each in a unique manner. GPT-5.1 exhibits reversal: its personas become more hostile toward the opposing party after balanced information, while humans become less so. This reversal is selective (80% on outgroup vs. 26% on policy questions) and symmetric across partisan identities. Gemini 2.0 Flash, Claude Sonnet 4.5, and Llama 3.3 70B exhibit overshoot, shifting correctly but at 5-7x human magnitude. DeepSeek V3 exhibits rigidity with near-zero change. Targeted ablations reveal that policy content triggers these failures and that they are identity-specific: GPT-5.1 reverses on outgroup questions but overshoots on ingroup; Gemini shows the inverse. We term this signature self-sycophancy: conformity to the model's internal stereotype of the persona rather than reasoning from the information provided. Our framework offers a concrete protocol: run the deliberative diagnostic before trusting LLM personas to mimic revised beliefs.","authors":["Ahmed Wali","Hassaan Tayyab"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15849","pdf_url":"https://arxiv.org/pdf/2609.15849","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","审议民意","算法保真度"],"reason":"直接评估LLM仿真人类意见动态，与真实人类数据对照，发现失效模式。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":2,"question":"LLM 人格在接收到平衡信息后是否会像人类一样更新信念，还是仅检索固化观点？","design":"用五个前沿 LLM（GPT-5.1、Gemini 2.0 Flash、Claude Sonnet 4.5、Llama 3.3 70B、DeepSeek V3）基于 America in One Room 数据构建 526 个匹配人口特征的人格，施加与人类相同的平衡信息干预，测量干预前后 72 个问题的观点变化。","baseline":"America in One Room 实验的真实人类数据，包含 526 名参与者在相同干预前后的观点变化。","findings":"所有模型均未通过诊断，且失败模式各异：GPT-5.1 在群体外问题上出现反转（敌意增加），Gemini、Claude、Llama 出现过度调整（幅度为人类 5-7 倍），DeepSeek 则表现为僵化（几乎无变化）。失败由政策内容触发且具有身份特异性，作者将其归因于“自我谄媚”（self-sycophancy）。","reliability":"论文未明确讨论失效条件，但指出当前评估仅关注静态观点正确性，无法捕捉动态更新失败；且失败模式因模型和问题类型而异，表明仿真可靠性高度依赖具体情境。","relevance":"该研究直接评估 LLM 仿真人类意见动态的可靠性，并与真实人类数据对照，揭示了静态评估无法发现的系统性失败，对关注仿真效度的研究者极具参考价值。","inspiration":"借鉴其“前测-信息干预-后测”的标准化诊断框架，并利用真实人类实验数据作为基准，可有效识别仿真中的方向性错误和幅度偏差。｜可迁移到政策公告的预期形成研究，例如央行沟通或财政政策变化对公众通胀预期的影响。｜以 LLM 人格模拟不同人口群体，施加与真实调查相同的政策信息，测量预期变化，并与密歇根大学消费者调查或央行预期调查的真实数据对照，检验仿真动态一致性。"}},{"id":"2609.15038","version":1,"title":"The average-farmer illusion in language-model simulations of agricultural decisions","zh_title":"语言模型模拟农业决策中的“平均农民”幻觉","abstract":"Language-model agents are increasingly used as synthetic people in surveys and social simulations, yet their apparent realism is often judged from population averages or distributional similarity. We tested what such evidence actually establishes by comparing Claude, Codex and Kimi under four prespecified prompt designs with matched farmer decisions from China and four African countries. Some configurations reproduced observed means and adoption rates. However, their person-level predictions were weak; their decisions clustered around typical values and policy-relevant extremes were largely missing. Most strikingly, a simple generator fitted only to the observed marginal dis- tribution, and given no information about any farmer, achieved greater distributional similarity than every language-model configuration. Prompt additions produced conditional gains rather than uni- versal improvement: results varied with model, outcome, population and validation target. We call this the average-farmer illusion: a synthetic population can look realistic while failing to repro- duce who does what or how behaviour varies. We provide a claim-matched validation framework and reusable modular prompts that turn prompt construction into an auditable experimental process. Population-level resemblance should therefore be treated as the start of validation, not as evidence of individual simulation.","authors":["Zhanliang Zhu","Ziwei Li","Yuchen Liu","Liujun Zhu","Ruiqi Wu","Tongqing Shen","Junliang Jin","Jianyun Zhang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15038","pdf_url":"https://arxiv.org/pdf/2609.15038","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B4"],"tags":["LLM仿真","人类行为对照","算法保真度"],"reason":"直接评估LLM仿真农业决策，与真实农民数据对照，揭示平均幻觉并给出验证框架。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":1,"question":"语言模型模拟农业决策时，群体层面的相似性是否意味着个体层面的模拟也可靠？","design":"用Claude、Codex和Kimi三个商业语言模型，在四种预设提示设计下模拟中国曲周县和四个非洲国家（尼日利亚、埃塞俄比亚、坦桑尼亚、马拉维）的农民决策，测量化肥施用量等连续行为变量，并与真实农民数据匹配。","baseline":"中国曲周县1332个农户-作物观测和四个非洲国家280个地块面板数据，包含农民实际决策。","findings":"部分配置能复现群体均值和采用率，但个体预测弱，决策集中在典型值，政策相关的极端值缺失。仅拟合边际分布的简单生成器在分布相似性上超过所有语言模型配置。","reliability":"提示添加只带来条件性收益而非普遍改进，结果随模型、结果变量、人群和验证目标变化；群体层面相似性不能作为个体模拟的证据。","relevance":"直接评估LLM仿真人类决策的可靠性，用真实农民数据对照，揭示平均幻觉，并给出验证框架，对关注仿真效度与偏差的研究者极具参考价值。","inspiration":"值得借鉴的是将验证目标分解为群体汇总、边际分布和个体匹配三个层次，并引入信息盲参考基准来检验分布相似性。｜可迁移到信贷审批歧视研究，用LLM模拟贷款官员对申请人特征的决策，检验群体违约率相似是否掩盖个体误判。｜用LLM扮演信贷员，输入申请人特征（收入、信用分、职业），输出是否批准贷款及额度，与真实银行信贷数据匹配，比较群体批准率、分布相似性和个体决策一致性，并加入仅基于边际分布的随机生成器作为基准。"}},{"id":"2609.13148","version":1,"title":"When Can You Trust Your Synthetic Users? Diagnostics and Corrections for LLM Consumer Panels","zh_title":"何时可以信任你的合成用户？LLM消费者面板的诊断与校正","abstract":"Large language models are increasingly deployed as synthetic consumer panels, promising $97\\%$ cost reductions over traditional surveys. Yet aggregate validation metrics conceal systematic failures: variance compression, coefficient sign-flips, subgroup error balloons of 10--30 percentage points, and global corrections that worsen demographic bias. We provide a formal framework for deciding when to trust, correct, or abandon LLM-generated consumer data. The framework decomposes synthetic-panel bias into covariate and concept shift, develops testable diagnostics with interpretable decision thresholds, and supplies a doubly robust AIPW estimator requiring only a small calibration sample ($n = 50$-$300$). We validate on three testbeds. In controlled simulations the decision rule achieves $100\\%$ accuracy (180/180 replications). On the American National Election Study with pre-existing LLM failures, it correctly flags heterogeneous concept shift and reduces naive bias by $92.9-99.6\\%$. On the Twin-2K-500 consumer pricing dataset (172,884 paired human and GPT-4.1-mini responses), it correctly routes full-sample estimation to Trust and subgroup targeting to Correct, with $83-94\\%$ bias reduction.","authors":["Robson Tigre","Hugo Gobato Souto"],"categories":["cs.HC","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13148","pdf_url":"https://arxiv.org/pdf/2609.13148","source_feed":"cs.HC","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM仿真","消费者面板","偏差校正"],"reason":"直接研究LLM合成消费者面板的可靠性诊断与校正，含真实人类数据对照，涉及经济学…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":2,"question":"如何判断何时可以信任、校正或放弃使用大语言模型生成的合成消费者面板数据？","design":"该论文不是仿真研究，而是提出一个诊断与校正框架：将合成面板偏差分解为协变量偏移和概念偏移，开发可检验的诊断工具（协变量重叠、条件校准、跨LLM稳定性），并提供双重稳健AIPW估计器，仅需小规模校准样本（n=50-300）即可校正偏差。","baseline":"使用三个真实人类数据集作为对照：美国国家选举研究（ANES）、Twin-2K-500消费者定价数据集（172,884对匹配的人类与GPT-4.1-mini响应）、以及受控模拟。","findings":"在受控模拟中决策规则达到100%准确率；在ANES上正确标记异质性概念偏移并将朴素偏差降低92.9-99.6%；在Twin-2K-500上正确路由全样本估计为信任、子群体定位为校正，偏差降低83-94%。","reliability":"论文承认概念偏移的严重程度是任务特定的，而非纯粹人口统计学：同一个LLM在产品定价上校准可靠，但在认知偏差任务（如合取谬误、锚定）上失败，因为LLM在人类依赖启发式的地方表现出“超理性”。此外，协变量重叠不足或跨LLM稳定性差时建议放弃使用合成数据。","relevance":"该论文直接针对LLM合成消费者面板的可靠性诊断与校正，提供了与真实人类数据对照的验证，并包含经济学相关场景（消费者定价），对关注仿真可靠性与偏差的研究者具有高度参考价值，值得精读原文。","inspiration":"该论文提出的协变量偏移与概念偏移分解、诊断阈值和双重稳健校正方法可借鉴用于经济金融仿真实验的可靠性评估｜可迁移到消费者金融决策、政策评估或行为经济学实验，如信贷审批歧视、消费者跨期选择、政策公告的预期形成等场景｜设计雏形：用LLM生成合成被试回答信贷申请或投资决策问题，以真实调查数据（如美国消费者金融调查SCF）为基准，施加不同政策信息处理，测量决策偏差，并用小规模人类样本校准AIPW估计器以校正LLM偏差。"}},{"id":"2607.28934","version":2,"title":"FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation","zh_title":"FairFund-Bench：评估LLM资源分配中的分配偏差","abstract":"Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 English-language requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants' names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.","authors":["Martin Lukk (University of Toronto)"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-03","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2607.28934","pdf_url":"https://arxiv.org/pdf/2607.28934","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","资源分配","算法公平"],"reason":"用LLM模拟人类资源分配决策，与真实人类数据对照，评估偏差与一致性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:06","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":4,"question":"LLM在资源分配中的偏差是否取决于审计设计（任务类型、比较情境、透明/伪装）？","design":"用14个LLM模拟人类资源分配决策，系统操纵任务（评分、排序、分配）、比较情境（单刺激/多刺激）、呈现方式（透明/伪装），测量对600个求助请求的分配结果。","baseline":"无对照（但基准中的求助模板基于130万真实GoFundMe活动校准，且使用福利应得性理论的人类应得性梯度作为参照）。","findings":"审计格式改变偏差方向：单独评分时模型偏向少数族裔，并排排序时则惩罚某些群体；伪装审计中的偏差比透明审计大3-4倍。因果框架效应比人口统计效应大约一个数量级，且跨模型和审计格式一致，表明LLM稳健地再现了人类应得性评价。","reliability":"论文指出偏差总体较小，但审计设计显著影响结论；透明审计中模型倾向于平均分配，可能掩盖潜在偏差；伪装审计更能揭示偏差。","relevance":"高度相关：该研究直接评估LLM作为人类被试替代品在资源分配决策中的偏差与一致性，并系统考察了审计设计对结论的影响，对理解仿真可靠性至关重要。","inspiration":"借鉴其系统操纵审计设计（任务、情境、呈现方式）来识别偏差的方法，以及用真实世界数据校准刺激材料并基于理论框架设计处理变量的做法。｜可迁移到信贷审批歧视、保险定价、政策福利分配等经济金融场景，检验LLM是否再现人类决策偏差。｜用LLM扮演信贷员，处理变量为申请人种族/性别（通过姓名信号）和贷款用途的因果框架（如医疗急需vs.创业失败），结果变量为贷款批准概率或利率，对照真实信贷审批数据（如HMDA数据）或人类实验数据。"}},{"id":"2608.02345","version":3,"title":"Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation","zh_title":"AI智能体能模拟A/B测试结果吗？面向智能体实验的验证框架","abstract":"A/B testing remains the standard for rolling out new features in the technology industry. Each experiment, however, consumes real traffic, engineering effort, and weeks of wall-clock time. Can AI agents---conditioned on behavioral profiles and contextual descriptions of the intervention---simulate outcomes accurately enough to vet candidate treatments before committing live traffic? We formalize this question as a \\emph{Simulated Randomized Controlled Trial} (S-RCT) and derive a two-layer error decomposition that separates agent approximation error from subsampling error, enabling targeted improvements to each. The framework is agent-agnostic: any behavioral model---from a fine-tuned specialist to a general-purpose foundation model---can serve as the simulation engine. Validated on 67 historical marketing A/B tests, a baseline S-RCT using an off-the-shelf foundation model captures directional signal (sign overlap 0.70) but systematically overshoots effect magnitudes. A two-phase pre-period calibration protocol reduces the squared prediction error (after removing irreducible measurement noise) by ${\\sim}77\\times$; a within-subject design---where each agent is exposed to both arms---reduces standard errors by ${\\sim}2.4\\times$. We discuss limitations of the current approach and identify applications where experimenters stand to benefit from agentic signals.","authors":["Stefan Hut","Lorenzo Masoero"],"categories":["cs.CL","cs.AI","stat.AP"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-04","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.02345","pdf_url":"https://arxiv.org/pdf/2608.02345","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","A/B测试","验证框架"],"reason":"用LLM模拟A/B测试结果，与真实历史实验对照，评估误差并改进，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":5,"question":"能否用AI智能体模拟A/B测试结果，以在投入真实流量前筛选候选干预方案？","design":"用基础大模型作为仿真引擎，根据用户画像、干预情境和任务描述生成模拟结果，构建模拟随机对照试验（S-RCT）；在67个历史营销A/B测试上，比较模拟估计的平均处理效应与真实历史结果，并引入预校准和受试者内设计改进估计。","baseline":"67个历史营销A/B测试的真实结果，包括效应方向和幅度。","findings":"基线S-RCT方向符号重叠率0.70，但系统性高估效应幅度；两阶段预校准将平方预测误差降低约77倍，受试者内设计将标准误降低约2.4倍。","reliability":"论文承认智能体存在系统性过度反应，画像完整性有限，且方向一致性在噪声数据上并不构成明确证据。","relevance":"直接相关：用LLM模拟A/B测试并与真实历史实验对照，评估误差并改进，符合研究者对仿真可靠性、偏差和失效条件的关注。","inspiration":"借鉴其将仿真误差分解为近似误差与子抽样误差，并分别用预校准和受试者内设计改进的做法｜可迁移到消费者金融产品选择或政策干预的A/B测试预筛，如信贷产品页面改版、退休储蓄默认选项调整等｜用LLM智能体基于用户画像模拟不同金融产品页面下的点击或选择行为，处理为页面版本，结果变量为选择率，并与历史A/B测试的真实选择数据对照，评估方向一致性和幅度校准。"}},{"id":"2609.13261","version":1,"title":"From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration","zh_title":"从过程损失到装配增益：多智能体LLM协作的人类基准诊断","abstract":"LLM agents are increasingly used for collaborative problem solving and human-group simulation. This makes outcome-only evaluation insufficient: if LLM groups are used as models of human groups, we need to know whether they succeed or fail through human-like deliberative mechanisms. We compare human group chats with matched LLM deliberation traces on Wason-style deductive reasoning, then test whether the same process signatures generalize to analogical, abductive, and analytical tasks. Humans and LLMs show the same assembly bonus asymmetry: discussion improves the average member more often than the best initial member. Initial-answer diversity accounts for the effect of model heterogeneity, increasing movement in both corrective and destructive directions. The main differences are process-level. Compared with humans, LLM groups follow majorities more often, surface less unique information, and converge earlier; correct minority signals succeed mainly when re-expressed early. Interventions motivated by human group-decision research yield modest improvements in collective outcomes, but do not remove the coordination bottleneck. Together, these results suggest that LLM groups can reproduce some outcome-level patterns of human deliberation while diverging in the mechanisms that generate assembly bonus and process loss, with implications for group simulation and human-AI collaboration.","authors":["Ala N. Tak","Teruhisa Misu","Kumar Akash","Zhaobo K. Zheng","Kevin H. Joo","Jonathan Gratch"],"categories":["cs.MA","cs.AI","cs.CL"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13261","pdf_url":"https://arxiv.org/pdf/2609.13261","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM群体仿真","人类对照","协作机制"],"reason":"用LLM群体模拟人类小组讨论，并与真实人类数据对照，评估机制差异。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":6,"question":"LLM群体协作能否复现人类小组讨论中的结果模式与过程机制，其装配增益与过程损失是否由类人机制驱动？","design":"用多个LLM智能体（GPT-4o、Claude-3.5-Sonnet等）组成小组，在Wason选择任务及类比、溯因、分析推理任务上进行自由讨论，记录个体初始答案、中间发言和最终群体答案，并施加信息揭示、多数翻转提示、置信度投影等干预，测量集体增益、最佳成员增益、装配增益和过程损失。","baseline":"匹配的人类小组聊天数据，在相同任务和协议下收集，作为过程诊断基准。","findings":"人类和LLM群体均表现出相同的装配增益不对称性：讨论提升平均成员多于最佳初始成员。但过程层面差异显著：LLM群体更频繁跟随多数、更少浮现独特信息、更早收敛，正确少数信号仅在早期被重新表达时才能成功。","reliability":"论文承认LLM群体虽能复现结果层面的模式，但机制与人类不同，存在更强的从众、更弱的少数信号保留和更早的锁定；干预仅带来适度改善，未能消除协调瓶颈。","relevance":"该研究直接针对LLM群体模拟人类小组决策的可靠性，提供了与真实人类数据的过程级对照，揭示了结果相似但机制不同的风险，对评估LLM仿真在群体决策场景中的有效性具有重要参考价值。","inspiration":"借鉴其过程追踪设计：记录个体初始答案、讨论中间发言和最终群体答案，并设置人类对照组，以区分结果相似与机制相似。｜可迁移到经济金融中的群体决策场景，如投资委员会决策、信贷审批小组、消费者家庭购买决策等。｜以LLM智能体模拟投资委员会，处理为不同信息结构（如隐藏信息揭示）或干预（如多数翻转提示），结果变量为投资决策质量和过程指标（如从众率、独特信息提及率），并与真实投资委员会会议记录或实验数据对照。"}},{"id":"2609.13995","version":1,"title":"Synthetic Data in Marketing Research: How to Evaluate and When to Trust","zh_title":"营销研究中的合成数据：如何评估与何时信任","abstract":"Debate over synthetic data in marketing research has polarized between claims that large language models (LLMs) make human respondents obsolete and calls to avoid them entirely. We argue that both positions obscure the more useful question: not whether synthetic respondents work, but when. Building on Brand, Israeli, and Ngwe (2026), we make three contributions. First, we distinguish three types of synthetic data (ungrounded LLM responses, segment-level personas, and individual-level digital twins) and map each to the decisions it can support. Second, we develop a taxonomy of four families of accuracy measures and suggest that the wide range of reported twin accuracy, from near-perfect to near-chance, largely reflects differences in what is being measured rather than in method quality. Aggregate measures often perform well even when little information is supplied to the LLM, and can mask a complete absence of respondent-level differentiation. Third, we introduce the forgotten question problem, in which a question is omitted from a fielded study, as a setting for twin-based augmentation of existing data. We propose an ex-ante answerability diagnostic that requires no ground truth: the R^2 of a random forest predicting twin outputs from the data used to construct the twins. Across 108 attitude questions from a nationally representative survey (N = 3,063), screening at R^2 above 0.7 raises the mean twin-human individual-level correlation by 15% and reduces the share of poorly answered questions from 25.9% to 4.3%. Embedding similarity and experienced-researcher judgment provide correlated but weaker screens.","authors":["Oded Netzer","Rajan Sambandam"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13995","pdf_url":"https://arxiv.org/pdf/2609.13995","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B4"],"tags":["LLM仿真","合成数据","营销研究"],"reason":"直接研究LLM合成数据在营销研究中的评估与信任，使用真实调查数据对照，提出诊断…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":7,"question":"在营销研究中，合成数据（LLM生成的回答、细分人群画像、个体数字孪生）何时可以信任，如何评估其准确性？","design":"论文区分三类合成数据：无依据的LLM回答、细分人群画像、个体数字孪生；提出四种准确性度量家族；引入“遗忘问题”场景，用随机森林R²作为事前可回答性诊断，在108个态度问题上筛选R²>0.7的孪生数据。","baseline":"使用全国代表性调查（N=3,063）的108个态度问题作为真实人类数据对照。","findings":"聚合度量往往表现良好，但可能掩盖个体层面无区分度；R²>0.7筛选将孪生-人类个体相关性平均提升15%，并将回答差的问题比例从25.9%降至4.3%。","reliability":"论文指出：聚合度量可能掩盖个体区分度缺失；嵌入相似度和研究者判断是较弱的相关筛选；未讨论模型版本、提示敏感性等失效条件。","relevance":"直接研究LLM合成数据在调查中的评估与信任，使用真实调查数据对照，提出无真值的事前诊断，对仿真可靠性研究有直接参考价值。","inspiration":"借鉴其事前可回答性诊断（用随机森林R²筛选可预测的孪生输出）和区分聚合与个体准确性的做法｜可迁移到消费者金融决策仿真，如信贷选择、保险购买或退休储蓄行为｜用LLM基于人口统计和财务特征生成个体孪生，施加不同金融产品特征处理，测量选择结果，并与真实消费者金融调查数据（如SCF或信用卡交易数据）对照，用R²筛选可回答的问题后再评估个体相关性。"}},{"id":"2609.15468","version":1,"title":"Time Machine Experiments: Using Historically-Bounded AI for Inquiry into the Human Mind","zh_title":"时间机器实验：利用历史受限AI探究人类心智","abstract":"Can interacting with someone from 1930, with no knowledge of what happened after, influence a person's perception of the past? People reason about the present against a picture of the past without observing it. The past is reconstructed from memory and testimony, but this reconstruction has been filtered through everything that happened since. Historically-bounded large language models (LLMs) make that past available for interaction. As a proof-of-concept for the impact of interacting with historical minds, we ran a preregistered randomized experiment ($N=240$), where participants interacted with an LLM trained on pre-1930 text. The interaction reduced the illusion of moral decline, the tendency to view the past as more moral than the present, compared to the contemporary-model control. This Time Machine Experiment paradigm informs new forms of interactive experiments, where temporal knowledge boundaries become experimental variables, and expands the realm of science fiction science, which turns thought experiments into actual experiments.","authors":["Hiromu Yakura","Robin Schimmelpfennig","Ezequiel Lopez-Lopez","Alejandro H. Artiles","Levin Brinkmann","Jean-Fran\\c{c}ois Bonnefon","Azim Shariff","Iyad Rahwan"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15468","pdf_url":"https://arxiv.org/pdf/2609.15468","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","人类被试替代","历史对照实验"],"reason":"用历史受限LLM作为人类被试替代，与真实人类对照，复现态度变化，属核心仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":9,"question":"与一个只了解1930年前信息的历史受限AI互动，能否改变当代人对过去的道德认知，从而缓解道德衰退错觉？","design":"使用基于1930年前文本训练的LLM作为历史人物代理，参与者与其进行对话和预测任务，测量互动前后对过去道德水平的感知变化及自我报告的反思程度，并与当代模型对照组比较。","baseline":"对照组为与当代前沿模型（gpt-5.5）互动的参与者，其道德衰退错觉变化作为基准。","findings":"与历史受限模型互动显著降低了参与者的道德衰退错觉，并引发了更多自我报告的反思。这表明跨时间互动可以改变人们对过去的偏见认知。","reliability":"论文未讨论","relevance":"该研究用历史受限LLM作为仿真被试，与真实人类对照，测量态度变化，属于核心仿真实验，值得阅读原文了解其方法和局限。","inspiration":"借鉴其利用历史知识边界作为实验变量的设计，通过对比不同时间截断的模型来隔离信息影响。｜可迁移到经济金融中的历史预期形成研究，例如投资者对历史政策效果的认知偏差。｜招募被试随机分组，分别与基于不同历史时期文本训练的LLM互动，测量其对历史经济事件（如大萧条）的归因和预期，并与真实历史调查数据对照。"}},{"id":"2609.15207","version":1,"title":"Issue Bias in Generative AI Writing Assistance: Political Issues and LLMs in the Swedish 2026 Election","zh_title":"生成式AI写作辅助中的议题偏见：2026年瑞典大选中的政治议题与大语言模型","abstract":"Generative AI writing assistants and the Large Language Models (LLMs) that power them are increasingly part of how voters gather information before elections. With growing evidence that they influence users' opinions, it is increasingly important to understand the views and positions of these tools. To better understand these views, we examine the stances supplied by six LLMs on a variety of Swedish-language writing tasks ahead of the 2026 Swedish parliamentary election. We cross 107 policy propositions with 77 writing templates and neutral, positive, and negative prompt framings, producing 24,717 prompts per model and 148,302 responses. To study these, we look at the models' default stance tendencies, compare how they respond to similar issues, and compare their responses with those of each of Sweden's eight parliamentary parties on the same issue. We find that Claude, DeepSeek, Gemini, and Mistral have similar profiles; ChatGPT more often supplies neutral or ambivalent text; and Grok differs most on topics such as migration, crime, and gender. When comparing the political parties, we find that the Social Democrats are closest to all six models. Still, after correcting for multiple comparisons, none of the within-model differences in party distances remains significant. Overall, we find that no model has a clear preference, nor a clear preference for a party, but that this depends on the specific issue or task the user asks about.","authors":["Bastiaan Bruinsma","Annika Fred\\'en","Paul R\\\"ottger","Moa Johansson","Asad Sayeed"],"categories":["cs.AI","cs.CY","stat.AP"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15207","pdf_url":"https://arxiv.org/pdf/2609.15207","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM政治态度","仿真对照","选举研究"],"reason":"用LLM生成政治文本并与真实政党立场对照，评估模型倾向，属于仿真人类政治态度且…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":12,"question":"在2026年瑞典大选前，六种大语言模型在瑞典语写作辅助任务中是否表现出政治立场或党派偏好？","design":"以六种LLM（Claude、DeepSeek、Gemini、Mistral、ChatGPT、Grok）为被试，交叉107个政策命题、77个写作模板和中性/正面/负面提示框架，生成148,302条瑞典语文本，分析模型默认立场倾向、模型间相似性及与瑞典八个议会政党的立场距离。","baseline":"瑞典八个议会政党在相同政策命题上的立场（来自三个投票建议应用VAA的编码，四点评分）。","findings":"Claude、DeepSeek、Gemini和Mistral立场相似；ChatGPT更常提供中性或模棱两可的文本；Grok在移民、犯罪和性别等议题上差异最大。所有模型与社民党距离最近，但经多重比较校正后，模型内各党距离差异不显著，表明无明确党派偏好。","reliability":"论文未讨论","relevance":"该研究用真实政党立场作为基准，评估LLM在政治写作任务中的立场倾向，属于仿真人类政治态度的研究，且包含批判性发现（无显著党派偏好），值得阅读原文了解其方法细节。","inspiration":"借鉴其大规模交叉设计（议题×模板×框架）和与真实立场对照的方法，可迁移到经济政策偏好仿真或消费者态度研究。｜可应用于政策公告的预期形成或信贷审批中的公平性评估。｜以LLM为被试，让其撰写关于税收、福利或监管政策的经济评论，处理为不同政策立场或框架，结果变量为文本中隐含的政策倾向，与真实民意调查或专家立场数据对照。"}},{"id":"2609.13254","version":1,"title":"(How) Do MLLMs Report Bistable Images Like Humans?","zh_title":"多模态大语言模型如何像人类一样报告双稳态图像？","abstract":"Bistable images such as the duck-rabbit are classic stimuli in which one image supports multiple mutually incompatible interpretations, typically reported one at a time in humans. We ask whether multimodal large language models (MLLMs) show similar report behavior and what internal computations support it. Using the LLaVA family, we study two tractable dimensions: modulability, whether reports can be biased by bottom-up visual cues and top-down linguistic priors, and exclusivity, whether responses commit to a single interpretation. We test both on the canonical duck-rabbit and on synthetic Visual Anagrams to mitigate memorization confounds. Behaviorally, both visual and linguistic manipulations systematically shift reports in human-consistent ways, while responses remain predominantly exclusive. Mechanistically, these effects arise from competing image-token representations, distinct pathways for bottom-up and top-down modulation, and a link between exclusive reporting and object-count encoding. Code and data are available at https://github.com/rtakatsky/mllm-bistable-images.","authors":["Ryota Takatsuki","Tomoki Doi","Amane Watahiki","Anil K. Seth","Hitomi Yanaka"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13254","pdf_url":"https://arxiv.org/pdf/2609.13254","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类行为对照","视觉认知"],"reason":"用MLLM复现人类对双稳态图像的报告行为，并与人类数据对照，评估仿真一致性。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":11,"question":"多模态大语言模型（MLLMs）是否像人类一样报告双稳态图像（如鸭兔图），其内部计算机制是什么？","design":"使用LLaVA系列五个模型（7B/13B等）作为被试，对经典鸭兔图和合成的Visual Anagrams双稳态图像施加视觉操作（旋转、加红圈）和语言提示（问题中嵌入偏向线索），测量模型输出的下一个词概率分布（modulability）和是否只报告单一解释（exclusivity），并分析内部表征。","baseline":"人类对双稳态图像的报告行为，包括视觉和语言线索对解释的偏向，以及报告单一解释的倾向。","findings":"视觉和语言操作都能以与人类一致的方式系统性地改变MLLM的报告，且报告绝大多数是排他性的。机制上，这些效应源于图像token表征的竞争、自下而上与自上而下调制的不同通路，以及排他性报告与物体计数编码的关联。","reliability":"论文承认只关注可操作化的两个维度（modulability和exclusivity），未涉及人类双稳态知觉的其他方面（如时间动态）；使用合成刺激以减轻记忆混淆，但可能仍存在其他偏差。","relevance":"该研究用MLLM复现人类对模糊视觉刺激的报告行为，并与人类数据对照，评估仿真一致性，且包含机制分析，对关注LLM仿真可靠性与偏差的研究者有参考价值。","inspiration":"借鉴其通过操纵输入（视觉与语言线索）和测量输出分布来量化模型行为与人类一致性的方法，以及使用合成刺激避免记忆混淆的设计。｜可迁移到经济决策中的模糊信息处理场景，如投资者对模棱两可的财报或政策声明的解读。｜用LLM作为被试，呈现模糊的金融图表或文本，施加视觉突出或语言框架处理，测量模型输出的解释分布，并与人类实验数据（如调查或行为实验）对照，检验仿真一致性。"}},{"id":"2607.29602","version":2,"title":"FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models","zh_title":"FriendBench：人类与多模态大语言模型二元熟悉度推断基准","abstract":"Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models favor ``stranger.'' This is a difference in effective prior, not in discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.","authors":["Jeffrey M. Girard","Jason Z. Zheng","Jacqueline R. Vertino","Antony D'Avirro","Benjamin Peloquin"],"categories":["cs.CL","cs.AI","cs.CV","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-03","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2607.29602","pdf_url":"https://arxiv.org/pdf/2607.29602","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["多模态LLM","人类对照","社会认知"],"reason":"评估多模态LLM推断人际熟悉度的能力，并与人类对照，揭示模型偏差，可迁移到仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:06","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":13,"question":"人类和多模态大语言模型能否仅凭20秒双人破冰对话的行为线索（而非语义内容）判断两人是熟人还是陌生人？","design":"本研究不是仿真研究，而是构建了一个多模态基准 FriendBench，从 Seamless Interaction 数据集中抽取96对真实双人互动，每对截取20秒破冰对话片段，分别以文本、音频、视频三种模态呈现，让26个多模态大模型和人类评分者进行二分类判断（熟人 vs. 陌生人），比较其准确率、判别力和反应偏差。","baseline":"人类基准：招募了约90名人类评分者，在相同刺激和条件下进行判断，作为模型性能的对照。","findings":"最佳模型与人类群体在三种模态上的准确率统计上无显著差异，但达成准确率的方式不同：人类在两类回答上保持平衡，而最强模型偏向“陌生人”，这是有效先验的差异而非判别力的差异。更丰富的通道对两者都有帮助但不均等，只有人类能从语音之上的可见行为中获得额外增益，最强模型从音频到视听模态准确率持平，未充分利用视觉行为线索。","reliability":"论文未明确讨论仿真失效条件，但指出模型存在类别偏差（偏向“陌生人”），且未充分利用视觉通道，说明模型在行为社会感知上与人类存在差异，可能影响其在真实社会场景中的可靠性。","relevance":"该研究直接对比了多模态LLM与人类在真实社会判断任务上的表现，揭示了模型在准确率相当的情况下仍存在行为偏差和模态利用差异，对评估LLM作为人类被试替代品的可靠性具有重要参考价值，值得阅读原文。","inspiration":"借鉴其设计：使用真实互动数据构建标准化任务，通过匹配人类评分者作为基准，并采用信号检测论分离判别力与反应偏差，以揭示模型与人类的深层差异。｜可迁移到经济金融中的社会感知场景，例如信贷审批中的面谈评估、投资者对管理层沟通的信任判断、或消费者对销售人员的熟悉度感知。｜设计雏形：以真实信贷面谈视频为刺激，让LLM和人类信贷员判断申请人与信贷员是否熟悉（或信任度），处理为不同模态（文本、音频、视频），结果变量为判断准确率和偏差，对照真实信贷决策数据。"}},{"id":"2609.13948","version":1,"title":"Thought without systematicity? Evaluating reasoning models on rule induction tasks","zh_title":"无系统性的思考？评估推理模型在规则归纳任务上的表现","abstract":"A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent variants of the same task. Here, we extend established rule induction tasks from cognitive science to assess the systematicity of thought in current reasoning models. Each task family has compositional structure that we use to create structurally equivalent task variations through task isomorphisms such as recombination and substitution. We find that despite being able to correctly solve a task, models often fail on structurally equivalent variants of the same task. These findings suggest that many model behaviors lack systematicity, rendering it difficult to robustly establish the cognitive abilities of reasoning models beyond the particular contexts they were evaluated in.","authors":["Simon Schug","Brenden M. Lake"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13948","pdf_url":"https://arxiv.org/pdf/2609.13948","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["认知科学","模型评估","系统性"],"reason":"评估推理模型的系统性，与人类认知对照，揭示模型行为缺乏系统性，可迁移到仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":14,"question":"推理模型在规则归纳任务上是否表现出人类认知中的系统性，即在结构等价的任务变体上表现一致？","design":"本研究不是人类仿真研究，而是评估推理模型（如GPT、Gemini等）的认知系统性。作者从认知科学中选取四类规则归纳任务（语法指令学习、符号推理、整数序列程序归纳、布尔概念学习），利用任务同构（重组、替换）生成结构等价的任务变体，测试模型在原始任务和变体上的表现，结果变量为解题正确率（多数投票）。","baseline":"无对照（未使用真实人类数据作为基准，而是以人类认知的系统性作为理论参照）","findings":"尽管模型能正确解决某个任务，但在结构等价的任务变体上经常失败，表明推理模型缺乏严格的系统性。即使采用五次多数投票，模型在同一任务上的表现也不稳定，这削弱了基于特定情境评估模型认知能力的有效性。","reliability":"论文承认人类也并非完全系统，但模型应追求更高系统性；评估依赖合成数据，且因成本限制只生成了有限的任务变体，可能影响结论的稳健性。","relevance":"该研究直接评估LLM在认知任务上的行为一致性，揭示了模型在情境变化下的脆弱性，对使用LLM进行人类仿真实验的可靠性提出了根本性质疑，值得精读。","inspiration":"借鉴其通过任务同构生成结构等价变体来检验模型行为一致性的方法，可迁移到经济金融实验中检验LLM对同一决策问题在不同表述或参数下的稳定性。｜例如在风险偏好、跨期选择或拍卖实验中，改变收益矩阵的数值缩放、标签或顺序，观察LLM的选择是否一致。｜设计：以LLM为被试，施加同一决策问题的多个同构变体（如彩票概率与金额的等价变换），结果变量为选择一致性，并与真实人类实验数据（如实验室风险偏好测量）对照，评估LLM作为人类被试替代品的可靠性。"}},{"id":"2609.14648","version":1,"title":"Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation","zh_title":"通过价值引导偏好蒸馏利用密集行为信号优化稀疏结果","abstract":"Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often ineffective and prone to reward hacking. We formulate long-horizon dialogue optimization as a multi-objective reinforcement learning problem and train a multi-head value model that predicts a vector of observed user behaviors across multiple look-ahead horizons. Our findings demonstrate that a scalarized composite of dense auxiliary behavioral signals enables effective credit assignment and optimization of sparse outcomes. However, optimizing unconstrained single-objective proxies might induce policy degradations that are harmful when the agent is exposed to real users. To identify these failure modes prior to deployment, we establish a safety framework combining counterfactual user simulation with a validated dialogue-level outcome model to evaluate preference weightings and policy optimization methods. Finally, we demonstrate that distilling multi-objective value preferences into the policy via reference-anchored preference optimization matches on-policy online RL at a small fraction of its compute budget. Live A/B testing confirms that our distilled policy significantly improves long-term user retention, while simultaneously enhancing the positive behaviors and therapeutic-process markers.","authors":["Ziyi Zhu","Daniel R. Cahn","Thomas D. Hull","Caitlin A. Stamatis","Olivier Tieleman","Guilherme B. Freire","Jinghong Chen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14648","pdf_url":"https://arxiv.org/pdf/2609.14648","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["用户仿真","对话优化","强化学习"],"reason":"用LLM用户仿真评估对话策略，有真实用户数据对照，涉及行为结果优化与失效分析","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":17,"question":"如何通过密集行为信号优化稀疏长期结果，同时避免奖励黑客并确保策略安全性？","design":"使用多目标强化学习训练多头部价值模型，预测多个前瞻时段的用户行为向量；通过标量化组合密集辅助行为信号进行信用分配和策略优化；利用反事实用户模拟器和对话级结果模型构建离线安全评估框架，在部署前检测奖励黑客；通过参考锚定偏好优化将多目标价值偏好蒸馏到策略中。","baseline":"真实用户A/B测试数据，包括长期用户留存、积极行为和疗法过程标记。","findings":"密集辅助行为信号的标量化组合能有效优化稀疏结果，但无约束单目标代理可能导致策略退化；蒸馏多目标价值偏好能以极低计算成本匹配在线RL，并在真实A/B测试中显著提升长期用户留存。","reliability":"论文承认用户模拟器在新颖代理行为下的保真度无法假设，因此将其作为筛选工具，并用真实部署确认方向；同时指出无约束单目标优化可能诱导有害策略退化。","relevance":"该研究展示了LLM用户仿真在评估对话策略中的有效性，并提供了与真实用户数据对照的验证方法，对关注仿真可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其使用多目标价值模型和反事实模拟进行离线策略评估的方法，可在经济金融实验中用于预筛政策干预。｜可迁移到政策公告的预期形成或消费者跨期选择等场景，利用LLM模拟经济主体行为。｜以LLM模拟消费者作为被试，施加不同政策信息处理，测量其消费或投资决策，并与真实调查或实验数据对照验证。"}},{"id":"2609.15972","version":1,"title":"Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States","zh_title":"Mind2Dialogue：通过模拟用户心理状态训练人类感知语言模型","abstract":"As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental supervision gap because current datasets for LLM assistant training contain few if any well-informed responses explicitly grounded in users' unspoken beliefs and goals. Scaling such supervision is inherently constrained, as users' underlying states are not directly observable. We thus propose the Mind2Dialogue framework to mitigate this gap by simulating users' mental states and turning them into privileged supervision for human-aware training. Specifically, we first propose a psychology-guided simulator that preserves personal characteristics while updating mental states through interaction to generate coherent conversations. The key idea is to enforce a shared evolving mental state that drives user behavior and guides an Oracle assistant's responses. Our privileged distillation then trains models on the Oracle's well-informed responses to assist users without direct access to their mental states at deployment. Moreover, we propose to evaluate human-aware learning by combining personalization and theory of mind, examining how models understand people and act on that understanding. Training on the full Mind2Dialogue corpus improves every reported personalization metric over the corresponding Qwen, Llama, and OLMo instruction-tuned baselines, including gains of 26.6 to 40.9 percentage points in preference-following generation. The gains extend to belief and action reasoning on Qwen and Llama, beyond personalized assistance. Looking forward, Mind2Dialogue makes user simulation a foundation for genuine AI collaborators that understand beliefs and intentions behind people's words and support their long-term goals across education, work, and everyday life.","authors":["Zixuan Wang","Yufan Zhou","Jinzhou Tang","Xinle Yu","Chengjun Wu","Lyumanshan Ye","Zhaoxiang Feng","Letian Peng","Adyasha Patra","Fan Bai","Enze Ma","Zhengding Hu","Jianyang Gu","Zhao Wang","Yufei Ding","Jingbo Shang","Tianmin Shu","Zhiting Hu","Zhen Wang"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15972","pdf_url":"https://arxiv.org/pdf/2609.15972","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","A4","B1"],"tags":["用户模拟","心理状态","人类感知训练"],"reason":"模拟用户心理状态训练助手，涉及人类数据对照，方法可迁移至人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":19,"question":"如何利用模拟用户心理状态为助手语言模型提供训练监督，以提升其人类感知能力？","design":"提出 Mind2Dialogue 框架，构建心理学引导的模拟器 M2D-Sim，生成具有共享演化心理状态的用户与 Oracle 助手对话，并利用特权蒸馏训练 M2D-Chat 模型；评估结合个性化与心理理论基准。","baseline":"无直接人类对照，但使用独立构建的个性化（PersonaMem、PrefEval）和心理理论（ToMi、BigToM）基准进行评测。","findings":"在 M2D-Corpus 上训练后，Qwen、Llama、OLMo 模型在个性化指标上全面提升，偏好跟随生成提升 26.6 至 40.9 个百分点；心理理论推理在 Qwen 和 Llama 上也有改善，但 OLMo 在 BigToM 上下降。","reliability":"论文未明确讨论失效条件，但指出心理理论推理的收益因模型而异（OLMo 在 BigToM 上下降），且评估基准与训练语料独立，但未涉及真实用户长期交互验证。","relevance":"该研究通过模拟用户心理状态生成训练数据，并利用独立基准评估模型的人类感知能力，为 LLM 仿真人类行为提供了方法参考，但缺乏真实人类数据对照，值得阅读以了解模拟与评估设计。","inspiration":"借鉴其共享状态模拟与特权蒸馏方法，可设计经济决策场景中的用户心理状态演化模拟，并利用独立行为基准评估模型表现｜可迁移至消费者跨期选择或政策公告预期形成等场景，模拟个体信念与偏好更新过程｜以 LLM 模拟消费者为被试，施加不同信息政策处理，测量其消费或投资决策，并与真实调查或实验数据（如消费者信心指数、实验经济学数据）对照验证仿真可靠性。"}},{"id":"2609.13773","version":1,"title":"Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging","zh_title":"推理能提升大语言模型的心理深度吗？取决于评判者是谁","abstract":"LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective. We study this failure mode through psychological depth in short stories. Seven human readers and an LLM-judge ensemble selected on the original scalar Psychological Depth Scale dataset ($\\rho = 0.646$) evaluated 60 blinded, prompt-matched story pairs from GPT-5 vs.\\ GPT-4o and DeepSeek-R1 vs.\\ DeepSeek-V3. Human preferences showed no universal reasoning advantage: GPT-5 was modestly preferred over GPT-4o (60.0--62.9\\%), whereas DeepSeek-R1 trailed V3 (42.9\\%), and inter-reader agreement was near chance (Krippendorff's $\\alpha = 0.070$), with within-reader consistency and recurring weighting patterns suggesting structured heterogeneity rather than random responding. The judge, by contrast, favored reasoning outputs in 89.0\\% of dimension-level comparisons and 59 of 60 pairs on aggregate PDS, uniformly across all five evaluator configurations, and its scores were associated with surface features such as sentence length and lexical diversity. These results suggest that development-set performance is insufficient evidence for deployment validity on a shifted distribution, and that point-estimate judges can obscure the heterogeneity in subjective human evaluation.","authors":["Ruichen Zheng","Yihe Wang","Fabrice Y Harel-Canada","Sara Khosravi","Zeynep Senahan Yildiz","Amit Sahai","Nanyun Peng"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13773","pdf_url":"https://arxiv.org/pdf/2609.13773","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM评估偏差","人类主观性","算法保真度"],"reason":"评估LLM作为评判者与人类主观评价的一致性，揭示其偏差，可迁移到仿真可靠性研究。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":15,"question":"在主观创造性文本评估中，LLM-as-a-Judge 在开发集上与人类评分相关良好，但在分布偏移、输出接近且人类偏好主观的条件下，其判断是否仍然有效？","design":"本研究并非用 LLM 模拟人类被试，而是评估 LLM 作为自动评判者与人类主观评价的一致性。具体做法：从原始 PDS 数据集中选择并构建了一个异构 LLM 评判者集成（基于 Llama 3.1 70B、Llama 3.3 70B、Qwen3.5-397B，每个维度路由到最佳配置），对 60 对盲法、提示匹配的短篇小说（GPT-5 vs. GPT-4o 两种推理努力设置，DeepSeek-R1 vs. DeepSeek-V3）进行 1-5 分的心理深度评分，并比较其偏好与 7 位人类读者的偏好。","baseline":"7 位人类读者对相同 60 对故事进行盲法偏好判断，作为人类主观评价基准；同时使用原始 PDS 数据集（97 篇故事及人类标注）作为开发集，评判者在开发集上与人类评分的 Spearman 相关系数为 0.646。","findings":"人类读者没有表现出普遍的推理优势：GPT-5 略优于 GPT-4o（60.0–62.9%），但 DeepSeek-R1 落后于 V3（42.9%），且读者间一致性接近随机（Krippendorff's α = 0.070），但存在结构化的异质性。LLM 评判者则强烈偏向推理输出（89.0% 的维度级比较和 60 对中的 59 对在总体 PDS 上），且其评分与句子长度、词汇多样性等表面特征相关。","reliability":"论文承认开发集性能不足以证明在分布偏移上的部署有效性；评判者可能奖励表面流畅性而非心理深度；点估计评判者掩盖了人类主观评价的异质性；人类读者间一致性低，但并非随机，而是反映了不同的评价标准。","relevance":"该研究直接评估了 LLM 评判者与人类主观评价的一致性，揭示了其在分布偏移和主观任务上的系统性偏差，对使用 LLM 进行人类仿真实验的可靠性评估具有重要参考价值，值得阅读原文以了解具体偏差机制和分布分析方法。","inspiration":"借鉴其方法：使用多个 LLM 配置构建异构评判者集成，并在开发集上选择最佳配置，然后在分布偏移的测试集上与人类判断进行对比，同时分析评判者评分与表面特征的相关性。｜可迁移到经济金融中的主观判断场景，如信贷审批中的文本解释评估、消费者评论的情感分析、政策公告的预期形成等。｜设计雏形：以 LLM 评判者作为自动评估工具，对经济文本（如贷款申请理由、投资建议）进行质量评分，处理为不同推理强度的模型生成文本，结果变量为 LLM 评分与人类专家评分的差异，对照数据为真实人类专家对同一批文本的评分，并检验 LLM 评分是否与文本长度、词汇复杂度等表面特征相关。"}},{"id":"2609.15864","version":1,"title":"Towards Scalable Measurement of Durable Skills","zh_title":"迈向可扩展的持久技能测量","abstract":"Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural real-world human interaction between humans. On the other hand, it should be scalable, controllable and reproducible. Here we argue that LLMs can be used to better capture both of these aims. Concretely, we develop a framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an \"Executive LLM\" setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency. We complement this with an AI evaluator that can be used to measure skill proficiency in such interactions. We evaluate our assessment protocol based on transcripts of interactions of human participants with our AI framework, for multiple durable skills. For the skill of creativity, we further demonstrate the efficacy of an autorater for evaluating complex tasks performed by real students. Our analysis shows that the use of the Executive LLM significantly increases elicited evidence and that LLM-automated scoring of conversations largely agrees with that of expert annotators. This research demonstrates the utility of orchestrated LLMs approaches for measuring complex social and cognitive constructs in a scalable and controllable manner.","authors":["Amir Globerson","Amy Keeling","Anisha Choudhury","Anna Iurchenko","Aviad Segal","Avinatan Hassidim","Ay\\c{c}a \\c{C}akmakli","Ben Gomes","Benn Witt","Cathy Cheunga","Cristine Legare","Diana Akrong","Eliad Carmi","Elisabeth Bauer","Gal Elidan","Hadas Gelbart","Hairong Mu","Katherine Chou","Lev Borovoi","Nir Kerem","Niv Efron","Noa Kerrem Gilo","Preeti Singh","Rajvi Kapadia","Rena Levitt","Roni Rabin","Ronit Levavi Morad","Rotem Yulzary","Shashank Agarwal","Sophie Allweis","Tracey Lee-Joe","Tzvika Stein","Yael Bar Moshe","Yael Haramaty","Yaniv Carmel","Yishay Mor","Yoav Bar Sinai","Yoav Bergner","Yossi Matias","Yuri Lev"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15864","pdf_url":"https://arxiv.org/pdf/2609.15864","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","技能评估","人机互动"],"reason":"用LLM模拟队友与人类被试互动，测量持久技能，有真实人类数据对照，属于教育评估…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":18,"question":"如何利用LLM构建兼具生态效度与心理测量严谨性的持久技能（如协作、创造力、批判性思维）评估框架？","design":"开发Vantage虚拟评估环境，人类被试与AI队友进行自然对话完成小组任务；AI队友由Executive LLM驱动，旨在引导对话以最大化技能证据；另用AI评估器对对话记录进行自动评分。","baseline":"人类专家评分者对真实人类被试与AI队友的对话记录进行评分，作为自动评分的对照基准。","findings":"Executive LLM显著增加了对话中可观察的技能证据；LLM自动评分与专家评分高度一致，表明AI评估器可替代人工评分。","reliability":"论文未讨论","relevance":"该研究利用LLM模拟队友与人类互动，并验证了自动评分的可靠性，为LLM在人类仿真实验中的应用提供了方法参考，值得阅读原文了解具体实现。","inspiration":"借鉴Executive LLM引导对话以高效提取行为证据的方法，可迁移到经济决策实验中，如通过AI引导被试在模拟市场或谈判中暴露偏好和策略。｜可应用于消费者跨期选择、风险偏好、合作博弈等场景，利用LLM模拟对手或伙伴，测量个体决策特征。｜设计一个实验：人类被试与LLM扮演的谈判对手进行多轮议价，Executive LLM引导对话以揭示被试的公平偏好和策略，结果变量为最终分配和出价序列，并与真实人类谈判实验数据对照，验证仿真有效性。"}},{"id":"2609.12273","version":1,"title":"Synthetic TLX: Forecasting Human Workload Using Agent Simulation","zh_title":"合成TLX：使用智能体仿真预测人类工作负荷","abstract":"Assessing human workload for technology-mediated tasks helps prevent task failure caused by poor technology design. Traditionally, workload is assessed retrospectively using the NASA Task Load Index (TLX) after humans complete a task. What if we could forecast workload before a human attempts a task using agent simulation? We introduce Synthetic TLX, a new paradigm for proactive workload estimation that predicts NASA TLX scores for a given task, unlocking novel interaction opportunities and evaluation methods. To understand its viability, we conducted three experiments comparing human and agent-generated scores to evaluate where they align and diverge. We found agent estimates align with human scores particularly when prompted with a human persona and active task simulation. However, agents and humans diverge in the sources of workload they are sensitive to. Based on our findings, we present three applications to showcase Synthetic TLX's potential and discuss the future of workload-aware human-AI interaction.","authors":["Tzu-Sheng Kuo","Carrie J. Cai","Meredith Ringel Morris","Michael Terry"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12273","pdf_url":"https://arxiv.org/pdf/2609.12273","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","工作负荷预测","人机交互"],"reason":"用LLM代理预测人类NASA-TLX工作负荷，并与真实人类数据对照，属于人类仿…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":1,"question":"能否利用LLM代理仿真在人类执行任务前预测其NASA-TLX工作负荷？","design":"使用LLM代理（GPT-4o）通过四种提示策略（基础、人类角色、任务模拟、角色+模拟）生成NASA-TLX分数，针对邮件撰写、网站导航、智能体对话三类任务，在三种不同工作负荷来源条件下进行预测。","baseline":"从Prolific招募人类被试完成相同任务和条件，并填写NASA-TLX问卷作为真实工作负荷基准。","findings":"当代理被赋予人类角色并进行主动任务模拟时，其估计与人类分数最一致；但代理与人类对工作负荷来源的敏感性存在差异，代理对任务内在复杂性更敏感，而人类对外部负担更敏感。","reliability":"论文承认代理估计在特定条件下与人类存在分歧，尤其是对工作负荷来源的敏感性不同，且当前LLM代理的能力有限，需要未来模型进步才能完全实现Synthetic TLX的潜力。","relevance":"该研究直接使用LLM代理仿真人类主观体验（工作负荷），并与真实人类数据对照，属于人类仿真实验，且涉及HCI任务，对关注LLM仿真可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其多策略提示对比和任务模拟设计，可迁移到经济金融中的主观体验预测（如消费者决策疲劳、投资者认知负荷），设计实验让LLM代理模拟不同投资者角色预测金融信息处理负荷，并与真实投资者问卷数据对照。"}},{"id":"2609.12444","version":1,"title":"Diverse Minds, Divided Networks? Personality Composition, Polarization, and Collective Intelligence in LLM-Based Social Simulations","zh_title":"多元思维，分裂网络？基于LLM的社会模拟中的人格构成、极化与集体智能","abstract":"Simulated societies of large language model agents are used to study online polarization, and separately to study collective intelligence, but the two are rarely measured in the same system. It is therefore difficult to say whether a society's personality composition shapes both, or whether reducing polarization costs collective competence. We present TraitMix, an experimental design in which the Big Five composition of a simulated social network, both trait levels and trait heterogeneity, is a controlled experimental variable, and in which polarization and collective performance are measured in the same runs. Across 991 simulations of hundred-agent societies, spanning six contested topics and six language models, trait heterogeneity has the largest measured effects, acting in opposite directions on two faces of polarization: varied societies hold more dispersed opinions while being less segregated into camps, so homogeneous societies are not moderate but consensual echo chambers. Trait effects are not additive, as Agreeableness determines the sign of Openness, an interaction that replicates across models although the primary model's estimate is influence-driven. Contrary to the trade-off the study was designed to measure, no polarization measure predicts poorer collective performance, and cross-cutting interaction is the only one of four whose association with collective accuracy survives partialling on the aggregation identity. We report ablations removing two potential measurement circularities, an induction gate applied to every model, and the measures that failed them.","authors":["Raad Bin Tareaf"],"categories":["physics.soc-ph","cs.CL"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12444","pdf_url":"https://arxiv.org/pdf/2609.12444","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A3","B4","D2"],"tags":["LLM社会模拟","人格构成","极化与集体智能"],"reason":"用LLM agent模拟社会网络，研究人格构成对极化和集体智能的影响，虽无真实…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":2,"question":"在基于大语言模型的社会模拟中，人格构成（特质水平和异质性）如何同时影响极化与集体智能，二者是否存在权衡？","design":"使用六种大语言模型（含不同家族）扮演百人社会网络中的智能体，赋予验证过的五大人格特质，通过控制特质均值和异质性形成不同社会构成；智能体在带算法推荐和动态关注网络的平台上讨论六个争议政治话题，同时私下测量观点极化（离散度、隔离度等）和集体表现（估计任务、隐藏画像任务等可验证答案的任务），共运行991次模拟。","baseline":"无对照","findings":"特质异质性对极化的两个维度作用相反：异质性高的社会观点更分散但阵营隔离更弱，同质社会不是温和而是共识性回音室；宜人性调节开放性对极化的影响方向，且该交互在多个模型家族中复现。未发现极化与集体智能间的权衡，跨阵营互动是唯一在控制聚合身份后仍与集体准确性正相关的极化指标。","reliability":"论文通过消融实验移除两个潜在测量循环，对每个模型施加诱导门控，并报告未通过检验的指标；承认部分极化指标在控制后失效，并指出主要模型的估计受影响力驱动。","relevance":"该研究用LLM智能体系统操纵人格构成，同时测量极化与集体智能，虽无真实人类对照，但提供了严谨的仿真实验框架和可靠性控制，对关注LLM仿真有效性及偏差的研究者有方法学参考价值。","inspiration":"借鉴其将人格构成作为受控实验变量、同时测量多个社会结果并设置内部有效性控制（如中性话题、诱导门控、跨模型复现）的做法｜可迁移到经济金融中的群体决策与信息传播场景，如投资者情绪与市场泡沫、信贷审批中的群体偏见、政策公告的预期形成等｜设计一个LLM智能体模拟的资产定价实验：以不同人格特质组合（如开放性、神经质）的智能体为被试，处理为信息环境（如是否提供异质信号），结果变量为价格偏离和交易量，对照真实市场实验数据或历史价格数据。"}},{"id":"2506.00152","version":2,"title":"Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective","zh_title":"用观测数据对齐语言模型：因果视角下的机遇与风险","abstract":"Large language models are being widely used across industries to generate text that contributes directly to key performance metrics, such as medication adherence in patient messaging and conversion rates in content generation. Pretrained models, however, often fall short when it comes to aligning with human preferences or optimizing for business objectives. As a result, fine-tuning with good-quality labeled data is essential to guide models to generate content that achieves better results. Controlled experiments, like A/B tests, can provide such data, but they are often expensive and come with significant engineering, logistical, and ethical challenges. Meanwhile, companies have access to a vast amount of historical (observational) data that remains underutilized. In this work, we study the challenges and opportunities of fine-tuning LLMs using observational data. We show that while observational outcomes can provide valuable supervision, directly fine-tuning models on such data can lead them to learn spurious correlations. We present empirical evidence of this issue using various real-world datasets and propose DeconfoundLM, a method that explicitly removes the effect of known confounders from reward signals. In simulation experiments, DeconfoundLM more accurately recovers causal relationships and mitigates failure modes of methods that assume counterfactual invariance, achieving over 16% higher objective score than ODIN and other baselines, when entangled confounding is present. Please refer to the project page for code and related resources.","authors":["Erfan Loghmani"],"categories":["cs.LG","econ.EM","stat.ML"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-14","first_seen":"2025-05-30","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2506.00152","pdf_url":"https://arxiv.org/pdf/2506.00152","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["B3","B4"],"tags":["因果推断","模型对齐","观测数据"],"reason":"用观测数据微调LLM以对齐人类偏好，涉及因果推断和偏差，方法可迁移到仿真可靠性…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":3,"question":"如何利用观测数据微调大语言模型以对齐人类偏好，同时避免学习到由混淆因素导致的虚假关联？","design":"本研究并非以LLM模拟人类被试的仿真实验，而是研究用观测数据微调LLM以优化业务指标（如点击率）的方法。具体做法：使用StackExchange和Upworthy两个真实数据集，分别展示直接微调会学到虚假关联（如“Happy Monday”标记）和观测数据仍可提供有价值信号；提出DeconfoundLM方法，在奖励信号中显式去除已知混杂因素的影响，并在模拟实验中与ODIN等基线比较。","baseline":"无对照（本研究不涉及用LLM替代人类被试的仿真，而是用真实人类行为数据作为微调监督信号，并在Upworthy数据上利用A/B测试结果作为评估基准）。","findings":"直接使用观测数据微调LLM会导致模型学习到由混杂因素引起的虚假关联（如将“Happy Monday”与高评分错误关联）；提出的DeconfoundLM方法能有效去除已知混杂影响，在模拟实验中比ODIN等基线提高16%以上的目标得分，更准确地恢复因果效应。","reliability":"论文指出，依赖反事实不变性假设的方法（如ODIN）在存在纠缠混杂时可能失效；DeconfoundLM需要已知混杂因素，若存在未观测混杂则可能仍有偏差。此外，观测数据本身可能包含选择偏差，且论文主要基于模拟和特定数据集验证，实际应用中的泛化性有待检验。","relevance":"该研究虽非直接以LLM模拟人类被试，但其核心关注使用观测数据微调模型时的因果偏差问题，与研究者关心的仿真可靠性及偏差条件高度相关，特别是关于混杂因素导致虚假关联的机制和校正方法，值得阅读原文以借鉴其因果校正思路。","inspiration":"可借鉴DeconfoundLM在微调过程中显式去除已知混杂因素的做法，用于处理经济金融领域中观测数据驱动的模型训练偏差。｜可迁移到信贷审批歧视研究：利用历史贷款数据微调LLM以预测违约风险，但需校正申请人特征（如种族、性别）与审批结果之间的混杂。｜设计：以LLM作为信贷审批员，输入申请人特征和贷款条款，输出审批决策；处理为在训练数据中应用DeconfoundLM去除已知混杂（如地区经济状况），结果变量为审批通过率，对照真实银行历史审批数据及后续违约记录，评估模型是否减少歧视性偏差。"}},{"id":"2607.28222","version":2,"title":"Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews","zh_title":"企业中的语音AI：自动化求职面试的自然田野实验","abstract":"We study AI agents as information-collection technologies: automated systems that elicit decision-relevant signals from humans through live interactions. We test how such AI automation impacts information collection and organizational outcomes using a natural field experiment with 70,000 applicants applying for real jobs. Applicants were randomly assigned to be interviewed by either human recruiters or AI voice agents. Afterward, human recruiters evaluate the interviews and make hiring decisions. Applicants interviewed by AI agents are 12% more likely to receive job offers, and these gains translate into higher job starts and worker retention, with no decline in the productivity of hired workers. Analyzing interview transcripts reveals that AI voice agents achieve controlled variance: their interviews are more structured and consistent while remaining responsive to individual applicants, which is associated with more hiring-relevant information collected. Our results suggest that a key advantage of AI automation lies in environments where information collection is delegated across many human workers and repeated such that variance in task execution becomes noise in decision-relevant signals, which AI compresses through adaptive standardization.","authors":["Brian Jabarian","Luca Henkel"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-09-14","first_seen":"2026-07-31","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2607.28222","pdf_url":"https://arxiv.org/pdf/2607.28222","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["AI面试","田野实验","人机对照"],"reason":"用AI语音代理替代人类面试官，与真实人类面试官对照，属于LLM仿真人类交互并评…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":4,"question":"AI语音代理替代人类面试官进行自动化工作面试，如何影响信息收集质量和招聘结果？","design":"自然田野实验：70,884名求职者随机分配由人类招聘官或AI语音代理进行面试，之后由人类招聘官评估并做出录用决定。结果变量包括录用率、入职率、留任率和生产力指标。","baseline":"人类面试官条件下的真实招聘数据，包括录用率、入职率、留任率及面试转录文本。","findings":"AI面试的求职者获得录用的概率高出12%，入职率和留任率提高约18%，且未降低被录用者的生产力。机制上，AI面试更结构化、一致且具有适应性，收集到更多与招聘相关的信息。","reliability":"论文未讨论","relevance":"该研究直接评估AI代理替代人类进行信息收集的效果，与真实人类面试官对照，属于LLM仿真人类交互并评估组织结果的实证研究，对关注仿真可靠性与偏差的研究者具有高度参考价值。","inspiration":"借鉴其随机化处理和真实结果测量的设计，将AI代理作为信息收集工具与人类对照，并分析转录文本以揭示机制。｜可迁移到信贷审批中的信息收集环节，如AI语音代理进行贷款申请访谈，比较审批结果和违约率。｜以银行信贷审批为场景，将贷款申请人随机分配由AI或人类信贷员进行电话访谈，处理为访谈方式，结果变量为贷款批准率、违约率和客户满意度，对照真实人类信贷员的历史审批数据。"}},{"id":"2608.00794","version":4,"title":"Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation","zh_title":"无有效性的测量：智能体AI评估中复合可靠性问题","abstract":"Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims. No formal framework has yet characterized how validity degrades across the stages of these pipelines. We present a three-layer compounding validity model, $V_{total} \\leq V_1 \\times V_2 \\times V_3$, that captures multiplicative degradation across task generation ($V_1$), human-simulator calibration ($V_2$), and automated judgment ($V_3$). Under empirically grounded estimates, a pipeline retaining 70% validity at each stage is at most 34% valid against the intended construct (range 0.17-0.54 across the empirical estimate bounds). We examine the model's predictions against a structured survey of 55 published agentic evaluation papers, finding that approximately 82% of papers in this purposive sample apply structurally mismatched, incomplete, or absent inter-rater reliability (IRR) metrics, a pattern consistent with systematic $V_3$ collapse. We further identify empirical evidence of $V_1$ failures (task validity flaws in 7 of 10 popular benchmarks) and $V_2$ miscalibration (up to 9 percentage points inter-simulator variance, with systematic demographic disparities for non-Standard American English speakers). We derive eight prescriptions grounded in psychometric science and domain-stratified reliability thresholds (ICC $\\geq$ 0.70; $\\alpha \\geq$ 0.67/0.70/0.80 by consequence level) that practitioners and benchmark authors can apply immediately. The framework provides a tractable knowledge-based tool for diagnosing and correcting evaluation pipeline validity before deployment decisions are made.","authors":["William Caban"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-14","first_seen":"2026-08-04","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.00794","pdf_url":"https://arxiv.org/pdf/2608.00794","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["效度评估","人类仿真","可靠性"],"reason":"评估LLM仿真人类被试的效度与可靠性，批判性指出失效条件，可迁移至人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":5,"question":"如何刻画并量化智能体AI评估流水线中效度随任务生成、人类模拟器校准和自动判断三层逐级复合衰减的问题？","design":"本研究不是仿真实验，而是提出一个三层复合效度模型 V_total ≤ V1×V2×V3，并通过对55篇已发表智能体评估论文的结构化调查、10个流行基准的任务效度审查以及模拟器校准差异的实证证据来验证模型预测。","baseline":"无对照","findings":"在经验估计下，每层保留70%效度的流水线对目标构念的总体效度至多34%（经验估计范围0.17-0.54）；约82%的论文使用了结构不匹配、不完整或缺失的评分者间信度指标，且10个流行基准中有7个存在任务效度缺陷，模拟器间方差高达9个百分点并对非标准美式英语使用者存在系统性人口统计学差异。","reliability":"论文承认模型是概念上界而非证明定理，并指出可靠性阈值可能沦为合规复选框、分层校准增加成本、高风险领域构念欠明确等局限。","relevance":"该研究批判性地评估了用LLM模拟人类被试的效度与可靠性，直接命中研究者关注的仿真失效条件，对理解LLM仿真在经济学实验和政策评估中的适用边界有重要参考价值。","inspiration":"借鉴其三层复合效度框架和分层校准方法，可系统评估LLM仿真在经济学实验中的测量效度｜可迁移到信贷审批歧视、消费者跨期选择、政策公告预期形成等场景｜用LLM模拟不同人口群体（如不同信用评分或语言背景的申请人）作为被试，施加政策干预（如改变信息披露方式），测量决策结果（如贷款批准率或跨期选择），并与真实实验或行政数据对照以校准仿真效度。"}},{"id":"2608.02100","version":2,"title":"From Information to Delegation: Mapping Human-AI Financial Decision Making","zh_title":"从信息到委托：映射人类与AI的金融决策","abstract":"As AI increasingly participates in human decision making, understanding how decision-making authority is distributed between humans and AI has become a fundamental behavioural question. We introduce a behavioural measurement framework combining intent and delegated decision authority to quantify what consumers seek from AI and how much decision-making authority they assign to it. Applied to 1.5 million real-world ChatGPT and Gemini interactions from 6,304 users in the United States and India, we find that financial services are already a substantial AI use case. Consumers overwhelmingly use AI to retrieve information and shape financial judgement, while delegation of financial execution remains rare. By shifting attention from conversation topics to delegated decision authority, this work establishes a behavioural baseline for measuring the transition to increasingly agentic AI.","authors":["Iman Munire Bilal","Yingcan Carol Wang","Ajan Raj","Filippo Giovagnini","Pranav Tewari","Yuwei Zhang","Mei-Chen Zoe Liou","Qamar Zaman"],"categories":["cs.HC","cs.LG"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-14","first_seen":"2026-08-04","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.02100","pdf_url":"https://arxiv.org/pdf/2608.02100","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["人机决策","行为测量","金融AI"],"reason":"用真实用户与AI交互数据测量人类决策授权，虽非实验仿真但可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":6,"question":"在金融决策中，消费者如何将决策权分配给AI？","design":"本研究并非仿真实验，而是基于真实用户与ChatGPT和Gemini的交互日志，提出结合意图与决策授权水平的行为测量框架，对对话进行意图分类和决策授权等级标注。","baseline":"无对照","findings":"金融服务已是对话式AI的主要应用领域，约半数用户在研究期间进行过金融对话。消费者主要用AI获取信息和塑造判断，而将执行决策权委托给AI的情况罕见，且多限于预算和财务跟踪。","reliability":"论文未讨论","relevance":"该研究虽非仿真实验，但提供了真实人类与AI交互中决策授权的大规模行为基线，可作为未来LLM仿真实验的对照数据，值得阅读原文了解其测量框架。","inspiration":"可借鉴其从对话日志中提取决策授权等级的方法，用于构建人类-AI协作的仿真场景。｜可迁移到金融咨询、信贷审批或投资决策等场景，研究消费者对AI建议的采纳与授权行为。｜设计一个实验：用LLM扮演不同风险偏好的消费者，处理为AI提供不同决策支持（信息、建议、执行），结果变量为授权等级，对照真实用户对话数据。"}},{"id":"2609.12191","version":1,"title":"GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents","zh_title":"GAUGE：何时不应信任用户仿真评估中的LLM裁判","abstract":"Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each candidate, an LLM-as-a-judge scores the transcripts, and the higher-scoring agent is promoted. We introduce GAUGE, a reusable offline protocol that measures whether this gate's ranking matches a grounded verifiable reward across 25 agents from six providers on the $\\tau^2$-bench and SimulatorArena benchmarks, separating two kinds of evaluation validity that release practices conflate: ranking validity and construct validity. First, a satisfaction-success gap: satisfaction carries essentially no information about task success, as conversations rated satisfied by our blind panel are decorrelated from actual success, with 57.5% of them failing the customer's task, a pattern consistent across five rater populations, both benchmarks, and every subjective dimension we rated. Second, while the gate's ranking is robust across the broad capability span, it loses resolution among the near-equal strong agents: this decision-disagreement rate jumps from $<$1% on wide-reward pairs to 31% on close pairs. The gate is thus human-validated yet mis-anchored. As a remedy, we propose a calibrate-then-trust cadence in which a judge-free completion bit is a zero-cost tripwire for truncation regressions.","authors":["Umesh Bodhwani","Thanh Tran","Kai Wei"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12191","pdf_url":"https://arxiv.org/pdf/2609.12191","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM用户仿真","评估效度","人类对照"],"reason":"评估LLM用户仿真器在任务型agent评测中的可靠性，含人类盲评对照，指出满意…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":9,"question":"LLM用户仿真器与LLM裁判组成的离线评估门控在任务型智能体排序中是否与可验证的真实奖励一致，以及其构念效度是否可靠。","design":"该研究不是人类仿真实验，而是对仿真评估系统的审计：使用25个来自6家提供商的智能体，在τ²-bench和SimulatorArena两个基准上，由人格驱动的LLM用户仿真器与智能体对话生成3700份对话记录，然后用LLM裁判、LLM人类代理、盲评人类小组和可验证的非LLM奖励（数据库状态与动作检查、人工标注正确性）分别评分，比较门控排序与真实奖励排序的一致性，并测量满意度与任务成功之间的差距。","baseline":"盲评人类小组对对话满意度的评分，以及τ²-bench的数据库状态/动作检查和SimulatorArena的人工标注正确性作为可验证的真实奖励。","findings":"满意度与任务成功几乎无关：盲评小组评为满意的对话中有57.5%实际任务失败，且该现象在五种评分者群体、两个基准和所有主观维度上一致。门控排序在能力跨度大的智能体间稳健，但在能力相近的强智能体间失去分辨力，决策分歧率从宽奖励对的<1%升至接近对的31%。","reliability":"论文承认门控仅在有限操作区域内有效，配置变化时需要重新审计；且样本外重校准不迁移。","relevance":"该研究直接评估LLM用户仿真器在任务型智能体评测中的可靠性，包含人类盲评对照，并指出满意度与成功脱节，对关注仿真效度与偏差的研究者具有重要参考价值。","inspiration":"借鉴其分离排名效度与构念效度的审计框架，用可验证的真实结果校准仿真评估门控，并测量主观评分与客观结果的相关性。｜可迁移到经济金融中的政策评估或消费者决策仿真，例如用LLM模拟消费者对金融产品的选择，再用真实交易数据验证。｜设计：用LLM仿真器扮演消费者，处理为不同产品推荐策略，结果变量为仿真满意度评分，对照真实消费者购买行为数据，检验满意度是否预测实际购买。"}},{"id":"2609.12575","version":1,"title":"Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture","zh_title":"多模态语言模型中的校准歧义：人类引用文化参照，模型描述图片","abstract":"Ambiguity is often treated as a bug for AI systems to resolve---but in human communication and culture, ambiguity can also be a generative resource. From humour to politics to art, people express themselves in words and images that are open enough to invite different interpretations, yet constrained enough to be interpretable. We operationalise this notion of calibrated ambiguity with a task drawn from the parlour game Dixit. We compare differences in clues generated by human vs multimodal language models, based on a novel coding rubric for calibrated ambiguity, and find that models consistently exhibit ambiguity collapse (i.e., their outputs are over-specified, leaving no room for multiple legitimate interpretations). Unlike human clues, AI-generated clues also exhibit cultural flattening; they almost never make reference to culturally-situated knowledge, even when prompted to use allusion and figurative language.","authors":["Cody Kommers","Mingrui Ye","Evelyn Gius","Daniela Mihai","Hoyt Long","Zheng Yuan","Drew Hemment"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12575","pdf_url":"https://arxiv.org/pdf/2609.12575","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["人类仿真","多模态模型","歧义校准"],"reason":"比较人类与LLM生成线索的歧义校准，有真实人类数据对照，揭示模型失效条件","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":11,"question":"多模态语言模型在生成线索时能否像人类一样校准歧义，即生成既开放又受约束、允许多种合理解释的线索？","design":"基于桌游 Dixit 设计任务：给定一张图片，要求生成一个线索，使部分人能猜中图片而部分人猜不中。收集人类和多种多模态语言模型（如 GPT-4o、Claude 等）生成的线索，开发编码量表从多个维度评估歧义校准程度，比较人类与模型、不同模型之间的差异。","baseline":"人类被试在相同 Dixit 任务中生成的线索，作为真实人类数据对照。","findings":"模型普遍出现“歧义坍缩”，生成的线索过度具体，缺乏多种合理解释的空间；与人类相比，模型生成的线索几乎不引用文化背景知识，即使提示使用典故和比喻语言，也表现出“文化扁平化”。","reliability":"论文承认校准歧义是情境依赖的设计问题，没有普适的歧义水平；当前评估框架仅基于 Dixit 任务，可能无法全面覆盖其他文化场景中的歧义校准。","relevance":"该研究直接比较人类与 LLM 在歧义校准上的差异，有真实人类数据对照，并揭示了模型在文化情境中的失效模式，对关注 LLM 仿真人类行为可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其设计：用游戏化任务诱发自然行为，开发多维编码量表量化模糊性，并对比人类与模型输出。｜可迁移到经济金融中的模糊沟通场景，如央行政策声明、分析师报告或广告中的模糊语言对市场预期的影响。｜以 LLM 模拟投资者或消费者，呈现不同模糊程度的政策声明或产品描述，测量其预期形成或购买意愿，并与真实人类实验数据（如调查或市场反应）对照，检验模型是否同样出现歧义坍缩和文化扁平化。"}},{"id":"2609.12949","version":1,"title":"EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics","zh_title":"EduFair-Bench：评估LLM导师跨学生人口统计特征的教学公平性","abstract":"Large language models (LLMs) are increasingly deployed as tutors, but it is unclear whether they support all students equally well. We introduce \\textbf{EduFair-Bench}, a benchmark for auditing the pedagogical fairness of LLM tutors---whether tutoring quality varies systematically with student demographics. EduFair-Bench pairs a multi-domain question bank (mathematics, physics, chemistry) with a controlled simulation in which a fixed LLM student interacts with each tutor across nine demographic levels spanning four dimensions: gender, immigration background, first language, and socioeconomic status (SES). Tutoring quality is scored on five turn-level pedagogical metrics and four conversation-level dimensions, using an LLM judge validated against three-annotator consensus on 180 tutor turns. Bias is measured via paired Wilcoxon signed-rank tests and bootstrap effect-size confidence intervals. Two ablations (demographic cues conveyed through names; conflicting demographic information between tutor and student) disentangle tutor-driven from student-driven bias. Across five tutors, we find that model capability and demographic fairness are largely orthogonal: the smallest model is the most consistent while the four more capable tutors all exhibit wide demographic gaps with no clear capability-to-fairness ordering, pedagogy-specific RL training redistributes rather than removes bias, and language- and immigration-related cues produce larger gaps than gender- and SES-related cues.","authors":["Jiaxu Zhao","Bahar Radmehr","Fares Fawzi","Tanya Nazaretsky","Tanja K\\\"aser"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12949","pdf_url":"https://arxiv.org/pdf/2609.12949","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM公平性","教育仿真","人口统计偏差"],"reason":"用LLM学生模拟不同人口群体与LLM导师互动，评估教学公平性，有真实人类标注对…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":12,"question":"LLM导师的教学公平性是否因学生人口统计特征（性别、移民背景、第一语言、社会经济地位）而系统性变化？","design":"用固定LLM学生模拟九种人口统计水平（四个维度），与五个LLM导师进行多轮对话，施加显式人口统计线索、仅姓名线索（Implicit）和冲突信息（Opposite）三种处理，测量五个回合级教学指标和四个对话级维度，并用LLM裁判评分。","baseline":"LLM裁判在180个导师回合上与三位人类标注者的一致性验证，无大规模真实学生对话对照。","findings":"模型能力与人口统计公平性基本正交：最小模型最一致，四个更强模型均显示较大人口统计差距且无能力-公平性排序；教学专用RL训练重新分配而非消除偏见；语言和移民相关线索产生的差距大于性别和社会经济地位相关线索。","reliability":"论文未讨论","relevance":"该研究用LLM学生模拟不同人口群体与LLM导师互动，评估教学公平性，有真实人类标注作为裁判验证，属于用LLM进行人类仿真实验并评估偏差的工作，对关注仿真可靠性和政策评估场景的研究者有参考价值。","inspiration":"值得借鉴的是其通过消融设计（Implicit和Opposite条件）分离导师驱动与学生驱动偏差，以及用配对Wilcoxon检验和bootstrap效应量置信区间量化偏见。｜可迁移到信贷审批歧视研究，模拟不同人口统计特征的借款人与LLM信贷员互动，检测审批建议中的系统性差异。｜用LLM扮演借款人（施加姓名、语言、收入等线索），LLM信贷员给出贷款决策和建议，结果变量为批准率、利率和解释文本的语调，以真实信贷审批数据（如Home Mortgage Disclosure Act数据）作为对照基准。"}},{"id":"2609.11983","version":1,"title":"Who Pays for Open Review? Visible Author Reputation and Its Effect on Ratings","zh_title":"谁为开放评审买单？可见的作者声誉及其对评分的影响","abstract":"An OpenReview bug in November 2025 broke anonymity at several conferences and prompted calls for open review, which motivate us to ask what shifting from blind to open would mean for authors. Analyzing over 18,000 reviewed submissions to ICLR 2026, split into de facto open and blind groups by arXiv preprint timing, we find that ratings rise with author reputation under both mechanisms, with a steeper slope under open review that is statistically significant, and that the open-blind difference is concentrated at the borderline ratings. The pattern holds across five reputation proxies (including institution, h-index, and citation count), three author-aggregation rules, and five definitions of the open window. A controlled simulation with five AI models as reviewers, holding the manuscript fixed and varying the author reputation, reproduces the effect. With claude-opus-5 as the reviewer, for example, rating rises by 0.5 points as the author moves from low to high reputation.","authors":["Qinghua Zhao","Xinyu Chen","Yanhui Yang","Tengfeng Sun","Junfeng Liu","Zhongfeng Kang"],"categories":["cs.DL","cs.AI"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11983","pdf_url":"https://arxiv.org/pdf/2609.11983","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","同行评审","声誉偏差"],"reason":"用LLM模拟审稿人，复现人类审稿中的声誉偏差，并与真实数据对照，属于人类仿真实…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":7,"question":"在同行评审中，从盲审转为开放评审（作者身份可见）会如何影响论文评分？","design":"用五个大语言模型（claude-opus-4-8、claude-opus-5、claude-sonnet-5、gpt-5.5、gpt-5.6-sol）模拟审稿人，对600篇ICLR 2026投稿进行评分；每篇论文在无作者信息、低声誉作者、中声誉作者、高声誉作者四种条件下各评一次，保持稿件内容不变，仅改变作者声誉，测量评分变化。","baseline":"真实人类审稿数据：ICLR 2026的18,789篇有评分的投稿，按arXiv预印本发布时间分为事实上的开放评审组和盲审组，比较两组中作者声誉与评分的关系。","findings":"在人类审稿中，作者声誉与评分正相关，且在开放评审下斜率更陡，差异在统计上显著；开放与盲审的评分差异主要集中在录取边界附近。AI审稿人模拟中，四个模型（除一个外）在作者声誉从低到高时评分上升0.2到0.6分，与人类数据中的声誉效应一致。","reliability":"论文未讨论","relevance":"该研究用LLM模拟审稿人，复现了人类审稿中的声誉偏差，并与真实审稿数据对照，属于人类仿真实验，且涉及AI在学术评价中的行为，对关注LLM仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其控制变量设计：固定稿件内容，仅改变作者声誉，以隔离声誉对评分的因果效应，并用真实人类数据做外部验证。｜可迁移到信贷审批中的歧视研究：模拟信贷员评估贷款申请，改变申请人种族或性别等声誉代理变量，观察审批结果变化。｜用LLM扮演信贷员，对同一批贷款申请在申请人特征（如姓名暗示的种族、职业、收入）变化下进行审批决策，结果变量为批准与否或利率，并与真实银行信贷数据中的歧视模式对照。"}},{"id":"2609.12137","version":1,"title":"GUIDE: Generative Utility Inference and Decision Engine","zh_title":"GUIDE：生成式效用推断与决策引擎","abstract":"Measuring the preferences of human users remains a fundamental challenge of AI alignment. Existing elicitation approaches struggle to efficiently discover multidimensional preferences or accurately ground these inferences in domain knowledge. To address this, we introduce GUIDE, an LLM-driven elicitation architecture that infers user preferences through conversations by combining Bayesian adaptive sampling for question selection and symbolic representation learning to initialize domain-specific preference models. GUIDE generalizes adaptive sampling to diverse elicitation questions through an extensible type system of transforms on a parameterized preference state. GUIDE produces domain-specific preference representations through an initialization process using symbolic rule-based learning to capture world knowledge and set priors over preference dimensions grounded in data about decision alternatives. The architecture provides observability and steerability to facilitate deployment and analyze elicitation processes. In silico experiments on investment portfolio optimization demonstrate that GUIDE improves cold-start and minimizes recommendation regret consistently within early elicitation interactions across user personas compared to prior work, LLM-only baselines, and ablated GUIDE versions.","authors":["Anagha Tiwari","Alexander G. Gray","Nick Feamster","Brian Jabarian","Alex Imas","Alex Kale"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12137","pdf_url":"https://arxiv.org/pdf/2609.12137","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["偏好推断","LLM仿真","决策优化"],"reason":"用LLM模拟用户偏好并优化决策，有仿真实验但非真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":8,"question":"如何设计一个LLM驱动的偏好诱导系统，在对话中高效推断用户多维偏好并用于决策推荐？","design":"使用GPT-5-4 mini和Claude Opus 4.7作为LLM后端，模拟不同投资策略的客户persona，在投资组合优化场景中通过对话进行偏好诱导；处理包括GUIDE框架（结合贝叶斯自适应采样和符号规则学习初始化）与多个基线（LLM开放式问答、LLM成对比较、OPEN、PEBOL及消融版本）；结果变量为推荐后悔值（regret）和冷启动性能。","baseline":"无对照","findings":"GUIDE在早期交互中一致地改善了冷启动性能并降低了推荐后悔值，优于LLM-only基线和消融版本；异构变换在后期稳定性上存在权衡。","reliability":"论文未讨论","relevance":"该研究用LLM模拟用户偏好并优化决策，属于人类仿真实验，但缺乏真实人类数据对照，与研究者关注的经济学实验和政策评估场景有距离，但方法上可借鉴其结构化偏好诱导框架。","inspiration":"可借鉴其将贝叶斯自适应采样与LLM对话结合，在仿真中系统比较不同诱导策略的设计；可迁移到金融投资偏好诱导、消费者选择实验或政策偏好评估等场景；研究设计可用LLM模拟投资者或消费者作为被试，处理为不同偏好诱导方法（如GUIDE vs. 简单LLM对话），结果变量为推荐准确率或后悔值，并用真实人类实验数据（如调查或行为实验）作为外部基准进行对照验证。"}},{"id":"2609.12331","version":1,"title":"Simulating Disengaged Students to Evaluate LLM-based Tutors","zh_title":"模拟不投入学生以评估基于LLM的导师","abstract":"Simulated students generated by computational models provide a practical way to evaluate tutoring strategies and pedagogical approaches used by human and AI tutors. However, such simulations should account for disengaged behaviors, including gaming the system, wheel-spinning, and off-task behavior, because tutors may need different responses for different learner states. We present Disengagement-Aware Student Simulators (DAS2), a reproducible pre-deployment protocol that models five learner-engagement states: engaged, gaming, wheel-spinning, off-task, and mixed, and evaluates AI tutor performance across these states. Using ASSISTments09, two coders independently labeled 100 sampled tutoring sessions based on anonymized interaction-log summaries. They achieved 84% agreement (Cohen's kappa = 0.78), and among agreed cases, human consensus labels matched DAS2 rule-based labels in 81% of cases (kappa = 0.75). Conditioning simulations on intended learner states reduced the correctness-rate gap between simulated and authentic sessions from 0.54 to 0.20 for gaming and from 0.51 to 0.18 for wheel-spinning. Fine-tuned Qwen2.5-7B better matched authentic response-time distributions, while prompt-only GPT-4o generated more distinguishable learner states. Evaluation of five AI tutors from the Claude, Llama, Gemini, Qwen, and GPT families showed that relative rankings remained stable across learner states and interaction lengths, while absolute performance varied, revealing state-specific differences in tutor support. Human validation further showed that automated tutor evaluation does not fully align with human judgment. DAS2 provides a pre-deployment framework for evaluating how AI tutors respond to diverse learner-engagement states before deployment.","authors":["Xianghui Meng","Jionghao Lin"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12331","pdf_url":"https://arxiv.org/pdf/2609.12331","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","教育评估","行为模拟"],"reason":"用LLM模拟学生行为并与真实数据对照，评估AI导师，属于人类仿真且含批判性验证。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":10,"question":"如何用模拟学生评估AI导师在不同学习投入状态下的表现？","design":"用LLM（Qwen2.5-7B微调版和GPT-4o提示版）模拟五种学习投入状态（投入、游戏、车轮打转、离题、混合）的学生，与AI导师交互，测量导师的回复相关性、脚手架和投入恢复等表现。","baseline":"ASSISTments09真实辅导日志，由两名编码员标注100个会话的学习投入状态，并与模拟学生行为对比。","findings":"条件化模拟于目标学习状态可将模拟与真实会话的正确率差距从0.54降至0.20（游戏）和从0.51降至0.18（车轮打转）；微调Qwen2.5-7B更匹配真实响应时间分布，而提示版GPT-4o产生更可区分的状态。","reliability":"论文承认自动化导师评估与人类判断不完全一致，且学习状态标签仅基于会话级行为指标，不代表所有可能的投入形式。","relevance":"该研究用LLM模拟人类行为并与真实数据对照，评估AI导师在不同状态下的表现，属于人类仿真且包含批判性验证，值得阅读原文以了解其仿真协议和失效条件。","inspiration":"借鉴其将行为状态离散化并条件化模拟的做法，可迁移到消费者金融决策中的不同心理状态（如冲动、谨慎）模拟，用LLM模拟消费者在信贷申请或投资决策中的行为，以真实交易数据为基准，评估金融AI顾问在不同状态下的建议效果。"}},{"id":"2609.06769","version":2,"title":"Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?","zh_title":"普通、理性的聊天机器人：AI模型是否追踪人类法律判断？","abstract":"As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making. As \"silicon sampling\" -- the use of generative AI models in social science research -- is now impacting academia, \"silicon jurors\" could make an appearance in courtrooms. This study joins an emerging line of research on generative AI models' ability to simulate human legal judgments. In particular, we study how large language model (LLM)-powered chatbots respond to series of questions about legal reasonableness. When the law needs to judge the appropriateness of a behavior, it most often asks whether the behavior was \"reasonable.\" Yet despite the ubiquity of reasonableness judgments, they are the site of constant vexation for lawyers, judges, and lay people. Reasonableness seems inherently vague and unpredictable, since it relies on variable context and implicit conceptual schemas. Moreover, many scholars caution that reasonableness judgments may vary along demographic lines. We compare the answers of human participants to those of twenty-six LLMs across twenty-five different legally relevant reasonableness judgments. Overall, our findings suggest that chatbot responses generally track those of human participants. Nonetheless, we find some suggestive -- and potentially concerning -- results. Compared to humans, LLMs generate more homogeneous responses and occasionally treat a variable standard as an invariant rule. And, compared to humans, LLMs tend to generate answers that are more favorable to the government and to corporations. Finally, our results indicate that LLMs' responses tend to align more closely with those of respondents who are white, male, older, and more educated. More systematic research is needed to confirm or reject these initial findings.","authors":["Nirav Patel","Emily Wenger","Christopher Buccafusco"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-11","first_seen":"2026-09-09","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2609.06769","pdf_url":"https://arxiv.org/pdf/2609.06769","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","法律判断","算法保真度"],"reason":"直接比较LLM与人类法律判断，含真实人类数据对照，并指出偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":1,"question":"大语言模型聊天机器人在法律合理性判断上是否与人类判断一致？","design":"比较26个LLM（来自Meta、Google、Anthropic、OpenAI、DeepSeek、xAI）与人类被试在25个法律相关合理性场景中的回答分布，分析模型回答的集中度、对政府与企业的偏向性，以及与不同人口群体回答的一致性。","baseline":"人类被试对相同25个法律合理性场景的回答数据。","findings":"LLM的回答分布与人类有统计差异，但总体落在人类回答的经验范围内，大体追踪人类的合理性观念。LLM回答更同质化，有时将可变标准当作不变规则，且更偏向政府和企业，并与白人、男性、年长、高教育程度人群的回答更一致。","reliability":"论文指出需要更系统的研究来确认或否定初步发现，并承认LLM回答存在同质化、偏向性等潜在问题。","relevance":"该研究直接比较LLM与人类法律判断，包含真实人类数据对照，并揭示了仿真偏差，对关注LLM仿真可靠性及偏差的研究者具有重要参考价值。","inspiration":"借鉴其多模型比较和人口统计学对齐分析的方法，评估LLM在特定判断任务中的偏差。｜可迁移到信贷审批歧视、消费者投诉处理或监管合规判断等经济金融场景。｜以LLM作为虚拟信贷员，处理贷款申请并给出批准决策，结果变量为批准率及理由，与真实银行信贷审批数据对照，检验模型是否与人类决策一致及是否存在人口统计学偏差。"}},{"id":"2603.17094","version":2,"title":"Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction","zh_title":"评估LLM模拟对话在建模人类社交互动中不一致与不合作行为的表现","abstract":"Simulating human conversations using large language models (LLMs) has emerged as a scalable methodology for modeling human social interaction. This paper reconsiders the evaluation of simulated conversations by explicitly recognizing that human conversations inherently involve inconsistent and uncollaborative behaviors, such as misunderstandings and interruptions. Since these behaviors contribute to the complexity of human social interaction, we argue that LLM-simulated conversations should reproduce them at frequencies comparable to those observed in human conversations. To support a detailed and interpretable evaluation of these behaviors, we introduce CoCoEval, a framework consisting of an evaluation scheme based on turn-level detection of 10 types of inconsistent and uncollaborative behaviors and a benchmark for simulating conversations in professional scenarios involving collaboration and conflict. Using CoCoEval, we compare human conversations with those simulated by GPT-4.1, GPT-5.1, and Claude Opus 4. The results show that (1) LLM-simulated conversations exhibit far fewer inconsistent and uncollaborative behaviors than human conversations under vanilla prompting, and (2) prompt engineering and supervised fine-tuning do not provide reliable control over these behaviors, often leading to the overproduction of specific behaviors. CoCoEval identifies gaps between human and LLM-simulated conversations that are not captured by conventional evaluation based on conversation-level Likert scales, raising concerns about the use of LLMs as proxies for human social interaction.","authors":["Ryo Kamoi","Ameya Godbole","Binglin Zhou","Xiaoxin Lu","Longqi Yang","Rui Zhang","Mengting Wan","Pei Zhou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-11","first_seen":"2026-03-17","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2603.17094","pdf_url":"https://arxiv.org/pdf/2603.17094","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","对话模拟","算法保真度"],"reason":"直接评估LLM模拟人类对话的保真度，并与真实人类对话对照，指出仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":2,"question":"如何评估LLM模拟对话中不一致和不合作行为的保真度，并与人类对话对照？","design":"使用GPT-4.1、GPT-5.1和Claude Opus 4在专业协作与冲突场景下模拟30轮对话延续，通过普通提示、分类引导提示和监督微调三种设置，测量10类不一致和不合作行为的出现频率。","baseline":"来自QMSum、NCPC、SIM和IQ2数据集的真实人类对话，涵盖商业、学术、政府会议和辩论。","findings":"普通提示下LLM模拟对话的不一致和不合作行为远少于人类对话；提示工程和监督微调无法可靠控制这些行为，常导致特定行为过度产生。","reliability":"论文指出LLM模拟对话的行为频率高度依赖模拟设置，且所有评估设置均未能复现人类对话中这些行为的频率；传统对话级Likert量表无法捕捉这些差异。","relevance":"该研究直接评估LLM模拟人类对话的保真度，并与真实人类对话对照，指出仿真偏差，对关注LLM作为人类被试替代品的研究者具有重要参考价值。","inspiration":"借鉴其细粒度行为检测评估方案和多种提示/微调设置对比，可迁移到经济金融中的协商、谈判或政策沟通模拟场景。｜例如，模拟消费者与客服的讨价还价或投资者与理财顾问的风险沟通。｜以LLM模拟谈判双方，处理为不同提示策略（如普通提示 vs. 明确要求包含冲突行为），结果变量为不一致行为（如误解、打断）的频率，并与真实谈判对话语料（如法庭记录或客服录音）对照。"}},{"id":"2607.27232","version":2,"title":"Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups","zh_title":"同情框架：跨社会人口群体评估AI对齐","abstract":"Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception. Considering news headlines covering political and geopolitical conflicts, both human participants (n = 3011, a representative sample of the U.K. adult population, via a YouGov survey) and seven LLMs answered whether headlines evoked sympathy for a specified side in a conflict. We find that the correlation between AI and human evaluations varies across models, ranging from very high (0.789, GPT-5.2) to medium (0.4 ,Mistral Large 2512). Crucially, the leading models are broadly aligned with human judgments across all demographic subgroups, including age, gender, level of education, prior geopolitical knowledge, and participants' predispositions regarding the conflict, although there are statistically significant differences between groups. This research, with its robust design and large, demographically diverse dataset, offers the most comprehensive evaluation of LLMs' comprehension of news framing to date. Findings highlight an important, often-ignored aspect of differential alignment: even when aggregate performance is high, AI alignment is not universal -- it may correspond differently with demographic features and cultural norms. Considering or ignoring the need for differential alignment may therefore have significant implications for the development of ethical and useful AI systems.","authors":["Haran Shani-Narkiss","Michael Fire","Oren Tsur"],"categories":["cs.CL","cs.AI","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-11","first_seen":"2026-07-31","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2607.27232","pdf_url":"https://arxiv.org/pdf/2607.27232","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","情绪感知"],"reason":"用LLM复现人类情绪感知，并与大规模人类调查对照，评估对齐差异","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":2,"question":"大语言模型在多大程度上与人类对新闻标题中同情性框架的情感感知保持一致，这种一致性在不同社会人口群体间是否存在差异？","design":"以七个主流大语言模型（GPT-5.2、Grok、GPT-4、Gemini、DeepSeek、Claude、Mistral）作为“读者”，对216条涉及政治和地缘政治冲突的新闻标题进行二元判断（是否对冲突中某一方产生同情），并与3011名英国成年人的调查回答进行对比，测量模型与人类判断的斯皮尔曼相关性。","baseline":"通过YouGov对3011名具有英国人口代表性的成年人进行调查，收集了超过20万条人类对新闻标题同情性框架的二元评价。","findings":"模型与人类判断的一致性因模型而异，GPT-5.2最高（0.789），Mistral最低（0.41）；即使总体一致性很高，在年龄、教育、母语、政治意识、话题知识和既有观点等子群体间仍存在显著差异。","reliability":"论文承认总体对齐高并不代表对所有人口群体都一致，顶尖模型在老年人、低教育水平、非英语母语者、低政治意识和无强烈观点者中一致性较低；同时指出模型在不同话题上的对齐稳定性不同，某些模型存在特定领域失效。","relevance":"该研究直接回应了用LLM替代人类被试进行情感感知仿真的可靠性问题，提供了大规模、人口多样化的真实人类对照，并揭示了总体对齐掩盖的群体差异，对评估LLM在调查和实验中的适用性具有重要参考价值。","inspiration":"借鉴其大规模分层对照设计，将模型输出与人口代表性样本的个体级回答进行相关性分析，并检验子群体差异，以识别对齐失效的边界条件。｜可迁移到政策公告的预期形成研究，例如央行沟通中文本框架对公众通胀预期的影响。｜以LLM作为虚拟受访者，对央行声明进行情感或框架判断，与家庭调查中的通胀预期数据（如密歇根大学消费者调查）对照，检验模型在不同教育、年龄和金融素养群体中的对齐程度。"}},{"id":"2609.11108","version":1,"title":"But How Would AI Agents Run a Town's Economy?","zh_title":"AI代理如何管理城镇经济？","abstract":"We placed 100 memory-equipped large language model (LLM) agents in charge of a closed, money-conserving spatial economy on real Pokhara Lakeside geography (earning wages, running businesses, setting prices) and ran this multi-agent simulation for up to 26 simulated weeks, well past the 1-2 weeks typical of agent-society studies. Across 91 validated runs (2.44M agent decisions, 21.5B tokens), the money stops moving, in a specific and measurable way. A 12x tourist demand shock raises business revenue 4.62x ($p<0.001$), which we decompose exactly into a 1.50x extensive margin (more businesses trading) and a 3.07x intensive margin (more revenue each). Monetary transmission stops there. Wages move 1.03x ($p=0.42$); 0.3% of 3,981 menu items are ever repriced ($p=0.47$). A randomized cash transfer (NPR 5,000 to 20 of 100 agents) shows the same pattern from the opposite direction: 96.7% is still held 311 pulses later, marginal propensity to consume 3-4% by two independent measures, indistinguishable from zero. The wealth distribution is consequently near-frozen at the horizon this literature uses ($\\rho=0.964$ over 2 simulated weeks), but not frozen. $\\rho$ falls to 0.832 at 12 weeks and 0.752 at 26, a horizon-dependence no short study can see. Matched ablations show which knob actually matters. Swapping the backing LLM moves every outcome we measure ($p=0.0039$); deleting agents' memory moves none of them detectably. A purely social tool fails 94-97% of the time across two model families, compared with ~96% success on economic tools, with no measurable shift away from it. Every headline number is verified twice, by a live validator and by an offline recomputation that reconciles each agent's wealth against its own signed transaction history, and we release the full run corpus for reanalysis.","authors":["Sajal Regmi","Siddhartha Pudasaini","Chetan Phakami Pun"],"categories":["cs.MA","cs.ET"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11108","pdf_url":"https://arxiv.org/pdf/2609.11108","source_feed":"cs.MA","score":9,"bucket":"selected","rubric_hits":["A3","B1","B2","B3","B4"],"tags":["LLM代理","经济仿真","政策评估"],"reason":"用LLM代理模拟经济，与真实数据对照，涉及政策评估，并批判性指出仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":4,"question":"LLM智能体在封闭货币经济中能否实现货币流通与财富再分配？","design":"100个带记忆的LLM智能体在真实博卡拉湖滨地理空间上运行封闭货币经济，从事工作、经营企业、定价交易；施加旅游需求冲击和随机现金转移两种处理；测量企业收入、工资、价格调整、边际消费倾向和财富分布变化。","baseline":"无对照","findings":"货币流通在企业和家庭层面均受阻：旅游冲击使企业收入增加4.62倍，但工资仅变动1.03倍，价格几乎不调整；随机现金转移的边际消费倾向仅3-4%，与零无显著差异。财富分布在2周内几乎冻结，但延长至26周后缓慢放松，表明短期研究可能得出误导性结论。","reliability":"论文承认种子数偏少（每条件9个），依赖匹配种子或组内对比；模型选择显著影响结果，而记忆删除无影响；社会工具失败率高达94-97%，经济工具成功率约96%；仅使用特定LLM模型，结果可能不具普遍性。","relevance":"该研究直接回应了LLM仿真在经济学中的可靠性问题，通过随机实验和长期追踪揭示了仿真失效的具体机制，对评估LLM作为人类被试替代品的有效性具有重要参考价值。","inspiration":"借鉴其随机现金转移和旅游需求冲击的准实验设计，以及通过长短期对比揭示时间尺度依赖性的方法｜可迁移到消费者跨期选择、政策刺激的乘数效应或信贷扩张的传导机制等宏观金融场景｜以LLM智能体为被试，施加一次性收入转移或信贷额度提升，测量消费支出、储蓄率和资产配置变化，并与家庭金融调查或信用卡交易数据对照。"}},{"id":"2609.11611","version":1,"title":"Who Bears the Risk When Generative AI Enters Transport? A Distributional Sociotechnical Audit of Algorithmic Equity, Synthetic-Data Validity, and Public Trust","zh_title":"生成式AI进入交通领域时谁承担风险？算法公平、合成数据有效性与公众信任的分配式社会技术审计","abstract":"Generative artificial intelligence is entering transportation through traveler-facing advisories, synthetic crash-record generation, and policy decision support. Existing governance frameworks lack transport-specific statistical tools to measure distributional risks across heterogeneous populations. We develop a Distributional Sociotechnical Audit (DSA) that integrates algorithmic equity, synthetic-data validity, and public-attitude heterogeneity into one empirical pipeline. The audit analyzes 5,760 persona-controlled queries to four LLM families across 12 demographic cues and four transport topics, uses two cross-family judges and a Wasserstein-2 Equity Dispersion Index, tests three FARS crash-record generators with conditional projected maximum mean discrepancy (cpMMD), fits a Bayesian ordered-logit model to Pew American Trends Panel Wave 152 (N = 4,538), and combines the signals into a continuous Sociotechnical Risk Index. Congestion-pricing advice has the highest persona-based dispersion (mean EDI = 1.96; highest direct EDI = 2.20). CART synthetic crash records fail all conditional tests (p < 0.001), while the Gaussian copula has borderline conditional stress (p = 0.105) despite passing marginal checks. Attitudes to AI vary across demographic strata. Distributional audits and continuous risk indices with sensitivity reporting offer a more defensible basis for transport GenAI governance than categorical approval tiers, which show a 75% assignment flip rate under weight perturbation.","authors":["Amir Rafe","Subasish Das"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11611","pdf_url":"https://arxiv.org/pdf/2609.11611","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","交通政策","公平性审计"],"reason":"用LLM模拟公众对交通政策的态度，并与真实调查数据对照，评估公平性和有效性。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":5,"question":"生成式AI进入交通领域后，其输出、数据产品和公众态度在不同人群间的分布性风险如何测量与治理？","design":"对四个LLM家族施加12种人口学线索和4个交通主题的5,760个查询，用两个跨家族裁判和Wasserstein-2公平离散指数测量建议分布差异；用条件投影最大均值差异检验三种FARS事故记录生成器的条件分布有效性；用贝叶斯有序logit模型拟合Pew调查数据分析公众AI态度异质性；最后合成连续的社会技术风险指数。","baseline":"Pew American Trends Panel Wave 152（N=4,538）的真实调查数据，以及FARS的110,001条真实事故记录。","findings":"拥堵收费建议在不同人格间的分布离散度最高（平均EDI=1.96），政策争议话题的人格驱动变异比天气安全建议高1.6倍；CART合成事故记录未通过所有条件检验，高斯copula在边际检验通过的情况下条件压力检验边缘显著（p=0.105）。","reliability":"论文承认分类审批层级在权重扰动下存在75%的分配翻转率，因此主张采用带敏感性报告的连续风险指数；但未详细讨论LLM仿真在何种条件下会失效，也未系统检验人格线索与真实人群行为的一致性。","relevance":"该研究用LLM模拟不同人口学群体对交通政策的反应，并与真实调查数据对照，评估公平性和有效性，直接回应了研究者对LLM仿真可靠性及偏差的关注，值得精读其审计框架和统计方法。","inspiration":"值得借鉴的是用多组人格线索系统探测LLM输出分布差异，并用真实调查数据校准公众态度异质性，同时用分布距离指标而非简单准确率来度量公平性｜可迁移到政策公告的预期形成或消费者对金融产品的态度异质性研究，例如不同人口群体对通胀或利率变化的反应差异｜设计雏形：用LLM扮演不同收入、教育、年龄的消费者，施加不同措辞的央行政策声明，测量其通胀预期和消费意愿，并与密歇根消费者调查或纽约联储SCE的真实数据对照，检验LLM仿真能否复现真实人群的态度分布和异质性。"}},{"id":"2608.28482","version":2,"title":"How Proper Scoring Rules Shape LLM Forecasting","zh_title":"适当评分规则如何塑造大语言模型预测","abstract":"This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same theoretical incentive for truthful probability reporting, the resulting models differ in calibration, probability use, and estimated profiles of bias, information, and noise, with smaller differences in aggregate accuracy and discrimination. The Brier-trained model has the lowest observed Brier score and highest AUC-ROC, while the log-trained model has the highest observed log score and lowest calibration error. Models with similar aggregate performance also reach that performance through different combinations of bias, information, and noise. Proper scoring rules therefore need not behave interchangeably as training objectives. Reward choice may shape not only how well an LLM forecasts, but how its forecasting errors are structured.","authors":["Benjamin Turtel","Paul Wilczewski","Kris Skotheim","Ville A. Satop\\\"a\\\"a","Philip E. Tetlock"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-11","first_seen":"2026-08-31","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2608.28482","pdf_url":"https://arxiv.org/pdf/2608.28482","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A2","B1","B3"],"tags":["LLM预测","校准与偏差","评分规则"],"reason":"评估LLM预测校准与偏差，有真实事件结果对照，涉及统计推断有效性","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":3,"question":"不同的严格恰当评分规则作为训练奖励时，如何影响大语言模型预测者的性能与行为？","design":"使用GPT-OSS-120b模型，在二元真实世界事件预测任务上，分别以对数、Brier、球面、Beta(0.5,0.5)和Beta(2,2)五种严格恰当评分规则作为Dr. GRPO的终端奖励进行训练，比较各模型在聚合评分、校准、区分度、概率使用及偏差-信息-噪声分解上的差异。","baseline":"无对照","findings":"不同奖励训练的模型在聚合准确性和区分度上差异较小，但在校准、概率使用和偏差-信息-噪声构成上存在明显差异；Brier训练的模型Brier分数最低且AUC-ROC最高，而对数训练的模型对数分数最高且校准误差最低。","reliability":"论文未讨论","relevance":"该研究直接评估LLM预测的校准与偏差，使用真实事件结果作为基准，与您关注的经济学预测和政策评估场景高度相关，值得精读原文以了解奖励设计如何影响仿真可靠性。","inspiration":"借鉴其通过改变训练奖励函数来塑造模型预测行为的方法，可系统比较不同激励下的预测偏差结构｜可迁移到经济预测场景，如通胀预期、政策效果或资产价格预测，考察不同损失函数对预测者行为的影响｜以LLM作为经济预测者，分别用对数、Brier等评分规则微调，比较其在CPI预测或政策公告效应预测中的校准与偏差，并与专业预测者调查数据（如SPF）对照。"}},{"id":"2609.11198","version":1,"title":"(Whose defaults?) Is artificial intelligence reorienting archaeological methods?","zh_title":"谁的默认？人工智能正在重新定向考古学方法吗？","abstract":"Generative AI and the practice of \"vibe coding\" are changing how archaeologists carry out computational research, but their effects on the discipline's range of methods is still understudied. In this paper, we evaluate whether large language models (LLMs) are narrowing the variety of methods archaeologists use. We first analysed approximately 119,000 archaeology abstracts from Scopus, covering publications from 2010 to 2025. Using a locally run LLM, we identified the computational methods reported in each abstract and organised them into 25 broad categories (L2) and 241 finer clusters (L3). A Bayesian Dirichlet-multinomial model of method composition within sub-disciplines found a small but credible shift in method use after 2023. However, this shift was smaller than the variation already present across the full study period. No individual technique showed a significant change, and overall methodological diversity increased rather than declined. We then ran a controlled experiment to see whether LLMs recommend a narrower set of methods than archaeologists have used in practice. Two different open-weight models were asked to suggest methods for 28 archaeological research problems, with prompts providing three levels of methodological guidance: novice, intermediate, and expert. Recommendation diversity was much lower than in the published literature, particularly without methodological guidance. The models also tended to favour methods that were widely used before 2023, and their recommendations more closely resembled the post-2023 literature. Taken together, these results are consistent with LLMs pushing methodological choice towards convergence, although our study cannot establish a causal effect. They raise a broader question: how can archaeology retain methodological diversity as LLMs become more involved in research?","authors":["Lorenzo Cardarelli","Roberto Ragno"],"categories":["cs.CY","cs.AI","cs.CL","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11198","pdf_url":"https://arxiv.org/pdf/2609.11198","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","方法多样性","科学实践"],"reason":"评估LLM对考古方法选择的影响，有真实文献数据对照，并批判性指出收敛风险，可迁…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":5,"question":"生成式AI和“氛围编程”是否正在收窄考古学研究中计算方法的多样性？","design":"本研究并非用LLM模拟人类被试，而是评估LLM对方法选择的影响。第一部分分析2010-2025年约119,000篇考古学摘要，用本地LLM提取并聚类计算技术，再用贝叶斯狄利克雷-多项模型检验2023年后方法构成的变化。第二部分为受控实验：让两个开源LLM为28个标准化考古研究问题推荐方法，提示词分新手、中级、专家三个方法学指导水平，测量推荐多样性与文献的差异，并用负二项回归检验方法推荐频率是否受2023年前流行度影响。","baseline":"真实人类基准为Scopus中2010-2025年考古学文献摘要所反映的实际方法使用分布，以及2023年前的方法流行度数据。","findings":"文献分析显示2023年后方法构成有微小但可信的偏移，但远小于领域内既有异质性，且整体方法多样性上升而非下降。LLM推荐的方法多样性远低于文献，尤其在新手提示下更低，且更偏向2023年前已流行的方法，其推荐模式更接近2023年后的文献。","reliability":"论文承认研究设计无法建立因果关系，只能说明结果与LLM导致方法收敛的压力一致；LLM推荐实验使用标准化问题，可能未完全反映真实研究情境的复杂性；且仅使用两个开源模型，结果可能不具普遍性。","relevance":"该研究虽非直接的人类仿真实验，但通过真实文献数据对照和受控实验，批判性地揭示了LLM对方法选择的收敛性影响，对关注LLM在学术研究中偏差与可靠性问题的研究者有参考价值。","inspiration":"借鉴其受控实验设计：通过不同提示词水平（新手/专家）施加处理，测量LLM输出的多样性并与真实数据分布对比，可迁移到经济金融研究中LLM辅助方法选择或模型设定的场景。｜例如，在资产定价研究中，可让LLM为不同经验水平的研究者推荐计量模型或因子选择，检验其是否导致方法同质化。｜设计：以经济学博士生为被试，随机分配使用LLM辅助或传统方法进行实证分析，处理为是否提供LLM推荐，结果变量为所选计量方法的多样性和与文献分布的偏差，对照真实数据为顶级经济学期刊中实际使用的方法分布。"}},{"id":"2609.10939","version":1,"title":"Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training","zh_title":"评估面向脚手架的多智能体大语言模型系统用于临床访谈训练","abstract":"Clinical education must prepare medical students to conduct safe and coherent patient interviews under conditions of uncertainty. Traditional standardized patient (SP) training is resource-intensive and difficult to scale. We developed a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) training platform1. The system includes a patient agent for simulated dialog, a tutor agent providing Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. In a randomized controlled study (N = 100 medical students), participants were assigned to either a multi-agent (MA) scaffolding condition or a control condition. All students completed two learning sessions under their assigned condition followed by an examination conducted in a patient only environment. Performance was assessed using a standardized Objective Structured Clinical Examination (OSCE) based rubric. While no significant difference was observed in final diagnostic accuracy between groups, the multi-agent AI standardized patient system improved final examination scores compared to the control group utilizing structured progressive information disclosure; the most substantial and consistent improvements were observed in communication, the expression of empathy, and specific history-taking behaviors. These findings suggest that specialized LLM agents enhance the process quality of simulated clinical interviews without artificially inflating examination outcomes. To support future research, we release a multi-expert annotated dataset comprising transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes. This resource aims to facilitate the development of pedagogically grounded AI-SP systems and advance research on AI-supported clinical reasoning training.","authors":["Luming Yang","Haoxian Liu","Siqing Li","Rong Jia","Yue Xiao","Guanhua Chen","Li Lu"],"categories":["cs.MA","cs.AI","cs.HC"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10939","pdf_url":"https://arxiv.org/pdf/2609.10939","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","医学教育","多智能体"],"reason":"用LLM模拟标准化病人训练医学生，有真实学生对照，属人类仿真但场景为教育训练而…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":4,"question":"多智能体LLM系统能否通过脚手架式提示提升医学生的临床访谈过程质量，而不影响诊断准确性？","design":"用三个专用LLM智能体（患者、导师、逐轮评估者）模拟标准化病人并提供苏格拉底式提示，对100名医学生进行随机对照实验，比较脚手架条件与结构化非LLM控制条件，结果用OSCE对齐的评分量表测量。","baseline":"对照组为接受结构化渐进信息呈现的非LLM训练条件的医学生，最终考试在仅患者环境中进行，以OSCE评分作为真实人类表现基准。","findings":"多智能体AI标准化病人系统显著提高了最终考试中的沟通、共情表达和特定病史采集行为得分，但诊断准确性与对照组无显著差异。","reliability":"论文未讨论。","relevance":"该研究用LLM模拟人类被试（标准化病人）并设置真实人类对照（医学生），属于人类仿真在教育训练场景的应用，对关注LLM仿真可靠性与偏差的研究者有参考价值，但非经济学或政策评估场景。","inspiration":"借鉴其多智能体分工和脚手架式干预设计，将LLM模拟角色与评估角色分离，以过程质量而非最终结果作为主要结果变量。｜可迁移到经济金融中的沟通与决策训练场景，如信贷审批中的客户沟通、金融咨询中的信息披露与信任建立。｜以LLM模拟客户或投资者，对经济学专业学生或从业者进行随机分组，处理组接受多智能体LLM的苏格拉底式提示与逐轮反馈，对照组接受传统案例学习，结果变量为沟通质量、信息获取完整性和决策准确性，并与真实客户互动数据或专家评分进行对照。"}},{"id":"2609.10856","version":1,"title":"Following the Preference, Missing the Optimum: Compliance Without Optimization in AI Housing Recommendation","zh_title":"遵循偏好，错失最优：AI住房推荐中的合规无优化","abstract":"Large language models are becoming the first point of contact for consumer search in domains where the stakes are material and the law is explicit. Existing audits show that models steer housing seekers by perceived identity, but none can say what a user loses when a recommender overlooks a suitable option, for want of an enumerated inventory to score omissions against. We audit AI housing recommendation against a verifiable ground truth. For each of 150 synthetic renter scenarios in New York City we build a pool of 120 real listings with known rent, bedrooms and GTFS-computed transit commute, compute the exact set satisfying the renter's stated constraints, and derive its Pareto frontier. The primary outcome assumes no utility function: a recommendation is strictly dominated if the same pool holds a listing cheaper, faster to commute from and no smaller in bedrooms. Across 9,945 calls to three models from two vendors, compliance is near-perfect (1.8% violation against a 66.6% random floor), yet 39.0% of recommendations are strictly dominated, and the dominating listing is a median 900 USD/month cheaper and 3.5 minutes closer. A within-scenario manipulation separates two capabilities usually conflated: changing one sentence moves median recommended rent by 646 USD/month in the correct direction, so preferences are honored, yet recommendations still sit 606 USD/month above the five cheapest qualifying listings on the same screen, and an unambiguous lexicographic instruction gives no improvement under equivalence testing against a pre-specified 50 USD/month bound. The gap widens with candidate-set size and replicates across OpenAI and Anthropic models to within 3 USD. We characterize the failure as compliance without optimization, propose dominance-rate instrumentation as a deployable diagnostic, and release all code, prompts and per-call results.","authors":["Hsuan Lo"],"categories":["cs.CY","cs.IR"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10856","pdf_url":"https://arxiv.org/pdf/2609.10856","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","推荐系统审计","决策偏差"],"reason":"用LLM模拟住房推荐中的用户决策，与真实房源数据对照，评估合规性与优化缺失，可…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":10,"question":"在住房推荐场景中，大语言模型是否在遵守用户明确约束的同时，未能优化推荐结果，导致用户错失更优选项？","design":"该研究并非用LLM模拟人类被试，而是直接审计LLM作为住房推荐系统的行为。研究者构建了150个纽约市租房者场景，每个场景配120个真实房源，计算满足约束的集合和帕累托前沿。然后调用三个模型（OpenAI和Anthropic）共9945次，要求模型推荐房源，并测量推荐是否违反硬约束、是否被严格占优（存在更便宜、通勤更快且卧室数不更少的房源）、以及租金与最优的差距。还通过改变场景中的偏好句子来检验模型是否遵循偏好，并通过改变候选集大小来检验优化能力。","baseline":"无对照","findings":"模型几乎完美遵守硬约束（违规率1.8%），但39.0%的推荐被严格占优，占优房源中位数便宜900美元/月且通勤快3.5分钟。模型遵循偏好（改变一句话租金中位数变动646美元），但推荐仍比同屏最便宜五个合格房源贵606美元，且明确的词典序指令无改善。","reliability":"论文承认其身份条件差异的零结果仅适用于固定候选集重排，不适用于开放式搜索；且未讨论模型在真实用户交互中的表现或长期影响。","relevance":"该研究直接评估LLM在住房推荐中的决策质量，与人类真实房源数据对照，揭示了合规与优化的分离，对关注LLM仿真可靠性及偏差的研究者有重要参考价值。","inspiration":"借鉴其构建可验证真实数据集并计算帕累托前沿来量化机会成本的方法，以及通过改变提示中的偏好来分离合规与优化能力的设计。｜可迁移到信贷审批或保险定价场景，检验LLM是否在遵守申请人硬性条件的同时未能推荐最优贷款或保单。｜用LLM扮演信贷员，输入申请人特征和贷款产品池，要求推荐产品；结果变量为推荐产品是否被占优（存在利率更低、费用更少且额度不低的产品），对照真实贷款产品数据和申请人约束，测量占优率和成本差距。"}},{"id":"2609.08585","version":2,"title":"Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations","zh_title":"自动化可模拟性的局限：LLM模拟器可以绕过解释","abstract":"Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is costly, automated simulatability replaces human explainees with LLM simulators, as proposed in ConSim (Poch\\'e et al., 2025) for large-scale experiments. We qualitatively replicate and extend ConSim's ranking of explanation methods across the tested datasets, explanation families, and simulator LLMs, and identify two limitations. First, when class names are meaningful, simulators can obtain high simulatability by solving the classification task directly, without relying on the explanations. Second, class anonymization can reward explanations for leaking the hidden label mapping, a limitation we expose with a new classes-as-concepts baseline. These results are consistent with a shortcut hypothesis: in the tested settings, simulator predictions mainly rely on task priors, while explanations produce small changes. We derive recommendations for more robust automated simulatability evaluations.","authors":["Antonin Poch\\'e","Fanny Jourdan","Nils Feldhus","Qianli Wang","Jing Yang","Simon Ostermann","Nicholas Asher","Philippe Muller","Vera Schmitt"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-11","first_seen":"2026-09-09","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2609.08585","pdf_url":"https://arxiv.org/pdf/2609.08585","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["可解释性","LLM模拟器","NLP评测"],"reason":"评估LLM模拟器预测任务模型输出的能力，属于NLP评测，不以人类行为为参照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":28,"question":"自动化可模拟性评估中，LLM模拟器是否依赖捷径（如任务先验或标签泄漏）而非解释内容来预测任务模型输出？","design":"复现并扩展ConSim协议：用LLM模拟器（如Llama-3.1-8B、Qwen2.5-7B等）扮演解释接收者，输入任务模型的解释（概念、归因、理由等）和输入实例，要求预测任务模型的输出；在非匿名和匿名类名两种设置下，比较不同解释方法的可模拟性得分，并引入classes-as-concepts基线。","baseline":"无对照（未使用真实人类数据，仅与ConSim的LLM模拟结果进行定性比较）","findings":"在非匿名设置下，解释对模拟器预测的影响很小，模拟器主要依赖任务先验直接解决分类任务；在匿名设置下，classes-as-concepts基线（仅泄漏标签映射）表现优于所有测试的解释方法，表明匿名化可能奖励标签泄漏而非有意义的解释。","reliability":"论文承认其发现基于15B参数以下、无推理能力的LLM模拟器，可能不适用于更大或闭源模型；且实验任务可能已被训练数据污染，未处理污染问题。","relevance":"该研究直接评估LLM模拟器的可靠性，揭示其捷径行为，对使用LLM进行人类仿真实验的研究者具有重要警示意义，值得阅读原文以了解具体失效模式和稳健性建议。","inspiration":"借鉴其通过引入基线（如classes-as-concepts）和对比非匿名/匿名设置来检测捷径行为的方法，可迁移到经济金融领域的LLM仿真实验，如政策公告解读或信贷审批解释的仿真；设计雏形：用LLM模拟投资者或贷款申请人，处理为提供不同解释（如模型决策理由），结果变量为预测模型决策的准确率，对照真实人类实验数据（如调查或行为实验）以评估仿真有效性。"}},{"id":"2609.09887","version":1,"title":"When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors","zh_title":"被告陈述何时重要？LLM模拟陪审员中的偏见与说服研究","abstract":"LLMs have been used to simulate human decision-making in professional settings, yet their behaviors in common-law jury trials remain unexplored. We study when and how a defendant's courtroom statement affects LLM-simulated jurors, focusing on persuasion, ideological bias, and background-based affinity. To support the analysis, we introduce JuryBench, a benchmark containing controversial criminal cases in U.S. criminal law. In each case, a defendant can claim various plausible justifications to support acquittal or reduced liability. We fix the base case and design defendants of different backgrounds, who give courtroom statements with varying emotional appeal or rebuttal. Jurors with diverse ideological profiles across the spectrum are simulated. We examine 20 frontier LLMs, resulting in a total of 432K decisions and rationales, and quantify changes in verdict severity. Our findings show that LLM-jury simulation echoes many human-jury findings. First, emotional persuasion can be detrimental, since jurors may perceive it as evidence of guilt or inconsistency. Next, we show that background fit between jurors and defendants is a stronger and significant factor than other isolated factors, and that jurors are in general harsher toward opposite-background defendants and lenient toward same-background ones. Finally, we find that juror ideology also strongly shapes severity judgments. These findings highlight both the promise and risks of using LLMs to model jury reasoning and call for careful evaluation. The data and code are available at https://github.com/choyingw/JuryBench","authors":["Cho-Ying Wu"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09887","pdf_url":"https://arxiv.org/pdf/2609.09887","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","陪审团决策","法律偏见"],"reason":"用LLM模拟陪审员决策，与真实人类陪审团研究对照，涉及法律决策偏差与说服效应。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-10","rank":1,"question":"在普通法陪审团审判中，被告的法庭陈述何时以及如何影响 LLM 模拟陪审员的裁决，重点关注说服、意识形态偏见和背景亲和力。","design":"使用 20 个前沿 LLM 模拟具有不同意识形态背景的陪审员，在 500 个争议性刑事案件中，对被告的不同背景和带有不同情感诉求或反驳的法庭陈述做出有罪/无罪及严重程度判断，并记录决策理由。","baseline":"无对照","findings":"情感说服可能适得其反，因为陪审员可能将其视为有罪或不一致的证据；背景契合度（陪审员与被告背景相似性）比孤立因素更强且显著，陪审员通常对背景相反的被告更严厉，对背景相同的被告更宽容；陪审员意识形态也强烈影响严重程度判断。","reliability":"论文未讨论","relevance":"该研究用 LLM 模拟陪审员决策，与真实人类陪审团研究对照，涉及法律决策中的偏差与说服效应，属于人类仿真实验，但未提供真实人类数据基准，可靠性存疑，值得阅读原文以了解其仿真设计细节和潜在偏差。","inspiration":"借鉴其通过系统操纵被告背景、陈述情感强度和陪审员意识形态来测量决策偏差的因子设计方法，以及大规模生成争议案例和记录决策理由的做法。｜可迁移到信贷审批歧视、招聘面试评估或政策沟通中的说服效应等经济金融场景。｜用 LLM 模拟信贷员或招聘经理，处理变量为申请人背景（如种族、性别）和陈述情感强度，结果变量为批准/拒绝或评分，对照真实信贷审批数据或审计研究结果来评估 LLM 仿真的外部有效性。"}},{"id":"2609.10280","version":1,"title":"Total Simulated Survey Error: Designing and Diagnosing Survey Responses from Large Language Models","zh_title":"总体模拟调查误差：设计和诊断大语言模型的调查回答","abstract":"Large Language models (LLMs), having been trained on vast amounts of human-generated data, may encode the attitudes and behaviors of these humans. As such, LLMs show promise in mimicking human-like patterns that facilitate their use in simulating people in a wide variety of contexts. One such context is using LLMs as 'silicon samples', i.e., proxies of people in answering survey questions to establish public opinion, design policies, or use as (social) scientific data. However, several critical questions of social biases, generalization, and technical limitations remain, further complicated by a vast design space open to simulation designers. Multiverse analyses might help us make sense of the impact of different design choices, however, we lack a systematic understanding of the design space of LLM-generated surveys as well as how these decisions interplay with inherent LLM limitations. Therefore, how do we systematically identify, trace, and document limitations in LLM-generated survey responses? Building on traditions in the quantitative social sciences, specifically survey methodology and measurement theory, we investigate threats to the validity of LLM-generated survey responses. To do so, we design a framework that enumerates conceptual errors and systematic biases that can occur at different stages of the survey simulation lifecycle. Our framework, called the Total Simulated Survey Error (TS2E) Framework, provides a unified and end-to-end perspective on LLM-generated survey data. The framework, illustrated through a theoretical and empirical case study, enables survey simulation designers to systematically identify and reflect on errors in LLM-generated surveys.","authors":["Indira Sen","Georg Ahnert","Leah von der Heyde","Jana Lasser","Bernd Wei{\\ss}","Markus Strohmaier"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10280","pdf_url":"https://arxiv.org/pdf/2609.10280","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A2","A4","B1","B4"],"tags":["LLM仿真","调查方法","误差框架"],"reason":"提出TS2E框架诊断LLM调查仿真误差，含实证案例，直接相关","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-10","rank":2,"question":"如何系统识别、追踪和记录大语言模型生成调查回答中的误差来源？","design":"提出一个概念框架（TS2E），将调查模拟生命周期划分为不同阶段，枚举各阶段可能出现的测量误差和代表性误差，并通过一个理论性和实证性案例研究进行说明。","baseline":"无对照","findings":"该框架区分了研究者设计选择导致的误差与LLM固有局限导致的误差，并引入了LLM特有的误差类型（如人物角色构建误差）和评估谬误。通过案例研究展示了框架如何帮助设计者系统识别和反思LLM生成调查中的误差。","reliability":"论文承认LLM训练数据存在缺口和偏斜，指令微调等后训练过程可能影响模型行为，且高总体对齐可能掩盖方差、子群体异质性和下游统计关系的严重失真。","relevance":"该论文直接针对LLM仿真调查的可靠性问题，提出了系统诊断误差的框架，对关注仿真效度与偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其将总调查误差框架迁移到LLM仿真的思路，对仿真流程进行阶段分解并系统识别误差来源。｜可迁移到经济金融领域的调查类仿真，如消费者信心调查、通胀预期调查、投资者情绪调查等。｜设计雏形：用LLM模拟不同人口统计学特征的消费者，施加不同的经济信息提示（如货币政策公告），测量其通胀预期和消费意愿，并与密歇根大学消费者调查的真实数据对照，检验仿真误差。"}},{"id":"2609.09899","version":1,"title":"Strangers to Themselves: What Language Models Say About Themselves Is Generic","zh_title":"自我陌生：语言模型对自身的描述是泛化的","abstract":"Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question. We find that: (i) Direct self-report is weak (r = +0.04), and even showing the model the exact items only raises prediction to +0.24. Crucially, the same item-informed question about \"capable AI agents in general\" does just as well (+0.28), while other models' answers about themselves predict the target model at least as well as its own. (ii) Frontier scale does not detectably change this pattern: any gains in prediction are not self-specific, and are consistent with a better theory of how AI assistants behave rather than better self-knowledge. (iii) First-person framing does have one robust effect: it shifts reports in the flattering direction, understating harmful behavior relative to the same question about a generic agent. (iv) Finetuning on a model's own behavioral record can teach narrow self-predictions, but it also changes the behavior being predicted and the gains do not transfer broadly. The practical implication is simple: asking a model what it would do mostly reveals a theory of AI assistants in general, plus a favorable bias, rather than privileged knowledge of that model.","authors":["Phil Blandfort","Urja Pawar"],"categories":["cs.LG","cs.AI","cs.CL","cs.CV","cs.CY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09899","pdf_url":"https://arxiv.org/pdf/2609.09899","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM自我认知","行为预测","可靠性评估"],"reason":"评估LLM自我报告与行为的一致性，揭示自我认知偏差，对仿真可靠性有批判性启示。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":8,"question":"语言模型对自己行为的自我报告是否包含关于该模型自身的特权知识，还是仅仅反映了对AI助手的一般性理论加上有利偏差？","design":"该研究并非用LLM模拟人类被试，而是将LLM自身作为研究对象：在九项行为评估中测量模型在不同条件下的实际行为率，然后让模型预测这些行为率，并设置多种对照（如询问“一般AI智能体”、用其他模型的回答预测目标模型等），比较预测准确度。","baseline":"无对照（没有使用真实人类数据作为基准，而是以模型间共享结构和跨模型预测作为参照）。","findings":"模型的自报行为与实际行为相关性很弱（r=+0.04），即使提供具体评估题目也只提升到+0.24；而询问“一般AI智能体”的预测效果相当（+0.28），其他模型对目标模型的预测甚至更好。第一人称框架会系统性地使自报偏向有利方向，低估有害行为。","reliability":"论文指出，在模型间行为差异很小的评估上（如能力评估），自知识信号微弱，因此难以检测；模型的自报主要反映通用AI行为理论，而非模型特定知识；微调虽能改善特定行为的自预测，但会改变行为本身且不具泛化性。","relevance":"该研究对LLM仿真人类实验的可靠性有直接警示：若用LLM自我报告作为行为预测或态度测量的替代，可能仅得到通用模式而非个体特异性，且存在社会赞许性偏差。值得阅读原文以了解其对照设计。","inspiration":"借鉴其预测测试框架：将自我报告与行为测量分离，并设置通用主体、跨模型预测等对照，以剥离通用知识与自我知识。｜可迁移到经济金融中的个体偏好或决策预测，例如消费者风险偏好、投资者情绪或政策反应。｜设计：用LLM扮演不同投资者，先测量其在模拟投资任务中的实际风险行为，再让其预测自己在不同市场条件下的行为率，同时询问“一般投资者”的预测，并与真实投资者调查数据（如面板数据）对照，检验LLM自报是否优于通用预测。"}},{"id":"2609.09428","version":1,"title":"XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?","zh_title":"XAI-Arena：LLM能否评估XAI解释的质量？","abstract":"Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations. We introduce XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, multidimensional, and stakeholder-sensitive evaluation of XAI explanation quality. XAI-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability. We then benchmark XAI explanation methods across various datasets, machine learning models, and stakeholder personas. Human validation shows a strong positive association between LLM-generated and human ratings (Spearman's rho=.693, p<.001). Together, LLM-based evaluations can capture systematic differences in XAI explanation quality and provide a scalable and reproducible framework for comparative assessment of XAI explanations.","authors":["Yanfei Hu Fleischhauer","Alona Zharova","Nadja Klein","Stefan Feuerriegel"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09428","pdf_url":"https://arxiv.org/pdf/2609.09428","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1"],"tags":["LLM评估","人类对照","可解释性"],"reason":"用LLM评估XAI解释质量，与人类评分对照，属于仿真人类判断并验证可靠性，可迁…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":7,"question":"LLM能否作为可复现、可扩展的机制来比较评估XAI解释的质量？","design":"提出XAI-Arena框架，使用单一LLM（GPT-5.4）在固定提示和解码设置下，扮演四种利益相关者角色（ML开发者、数据科学家、经理、最终用户），对五种XAI方法（SHAP、LIME、DiCE、PDP、排列重要性）在不同数据集、ML模型和输入格式下生成的解释，从八个维度（感知简单性、清晰度、任务充分性、信任校准、可操作性、透明度、忠实度、整体可解释性）进行1-7评分，并输出理由。","baseline":"人类评估：收集人类对相同XAI解释的评分，与LLM评分进行相关性分析（Spearman's ρ=.693, p<.001）。","findings":"LLM评估与人类评分有强正相关，能捕捉XAI解释质量的系统性差异。LLM评估的忠实度与部分技术代理指标（如稳定性、MoRF AUC）相关性较弱，表明LLM可能捕捉到不同维度。","reliability":"论文未明确讨论失效条件，但指出LLM评估可能受限于模型本身偏见、提示敏感性，且仅在特定LLM（GPT-5.4）上验证，泛化性未知。","relevance":"该研究用LLM模拟人类对解释质量的判断，并与真实人类评分对照，验证了LLM作为人类被试替代品的可靠性，属于人类仿真实验，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"借鉴其使用LLM扮演不同利益相关者角色、在固定提示和温度下进行多维评分并与人类评分对照的方法，可迁移到经济金融中的政策解释或模型决策解释评估，例如评估信贷审批模型解释对贷款申请人的可理解性和信任影响。｜可应用于信贷审批歧视研究，让LLM扮演贷款申请人或监管者，评估不同XAI方法对信贷决策解释的公平性感知。｜设计：以LLM扮演贷款申请人，处理为不同XAI解释（如SHAP与LIME），结果变量为对决策的信任度和理解度评分，对照真实人类被试的评分数据。"}},{"id":"2609.10421","version":1,"title":"Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support","zh_title":"急诊科再就诊质量审查筛查：探索人类决策与人工智能支持","abstract":"Background: Emergency Department (ED) return visits are commonly reviewed for quality assurance, but are often limited (e.g., to revisits within 48-72 hours) to increase actionable finding yield while minimizing chart review burden. Those limitations may lead to missed quality improvement opportunities. Methods: We conducted an exploratory, retrospective study of randomly selected ED visits to a multihospital health system having an ED revisit within 1-14 days to the same health system. Given only each visit's primary diagnosis, raters (2-3 clinicians and GPT-4 large language model [LLM]) assessed characteristics of the diagnosis pairs, including the \"target\": whether a pair warranted further assessment. Informed by rater response analyses, an algorithm leveraging an LLM-populated knowledge graph (\"KGA\") was created to automatically screen for potentially concerning pairs, then preliminarily assessed. Results: 99 diagnosis pairs were included. GPT-4 responses poorly correlated to clinician raters, rating nearly all (94%) pairs as warranting follow-up (4.4-13.3 times more than clinicians). However, prompt engineering was minimal. Among clinician raters, revisit medical gravity was consistently significantly associated with the target, while a differential diagnosis/complication composite was significantly associated on unadjusted, but not adjusted (though less powered) analysis. The KGA achieved 83-100% positive predictive value for at least one clinician rater determining further assessment was warranted based on the diagnosis pair. Conclusion: These results can inform next steps for improving screening with LLMs like ChatGPT. Further research is warranted to validate this preliminary work's finding that the KGA may enable enhancing the scope and yield of screening without substantially increasing reviewer workload.","authors":["Jonathan A. Handler","Marlene I. Robles-Granda","Jacob E. Mefford","Jeremy S. McGarvey","Gregory S. Podolej","Colleen J. Klein","Matthew D. Dalstrom","William F. Bond"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10421","pdf_url":"https://arxiv.org/pdf/2609.10421","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类决策对照","医疗质量审查"],"reason":"用GPT-4模拟临床医生判断，并与人类医生对照，评估其可靠性，属于LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":9,"question":"急诊科复诊质量审查中，仅凭诊断对，人类临床医生和GPT-4如何判断是否需要进一步审查，以及能否用LLM知识图谱算法自动筛选可疑复诊对？","design":"回顾性研究，随机选取99对急诊初诊与14天内复诊的诊断对，仅提供主诊断，由2-3名临床医生和GPT-4分别评估诊断对特征（诊断成员价值、医学严重程度、鉴别诊断包含、并发症包含）及目标变量（是否需进一步审查）。基于评估结果，构建了利用LLM填充知识图谱的自动筛选算法（KGA），并初步评估其阳性预测值。","baseline":"2-3名临床医生（急诊医师和高级实践提供者）的独立评估作为人类基准，与GPT-4的评估进行对比。","findings":"GPT-4与临床医生的判断相关性差，将94%的诊断对评为需随访，是临床医生的4.4-13.3倍；临床医生中，复诊医学严重程度与目标变量显著相关，鉴别诊断/并发症复合指标在未调整分析中显著但调整后不显著；KGA对至少一名临床医生判定需进一步评估的诊断对实现了83-100%的阳性预测值。","reliability":"论文承认GPT-4提示工程极简，可能导致其过度判定；研究为探索性、样本量小（99对），且仅基于主诊断，未提供其他临床信息；KGA仅初步评估，需进一步验证。","relevance":"该研究直接使用GPT-4模拟临床医生决策并与人类对照，评估其可靠性，属于LLM仿真人类判断的实证研究，且包含真实人类基准，对关注LLM仿真在专业决策中有效性的研究者有参考价值，但场景为医疗而非经济金融。","inspiration":"借鉴其设计：用LLM模拟专家判断并与人类专家对照，同时构建基于LLM的自动筛选算法并与人类判断比较，以评估算法性能。｜可迁移到经济金融中的专业判断场景，如信贷审批中贷款员对借款人风险的评估、审计师对财务舞弊风险的判断、或政策分析师对经济指标异常的关注。｜研究设计：以信贷审批为例，选取一组贷款申请对（如初贷和短期内再贷），仅提供有限信息（如信用评分、收入），让LLM和人类信贷员分别判断是否需进一步审查，并构建基于LLM知识图谱的自动筛选算法，以人类信贷员的判断为基准评估算法阳性预测值，同时对比LLM与人类判断的一致性。"}},{"id":"2609.09609","version":1,"title":"Who You Are Adds Nothing Detectable to Where You Go Next: Sociodemographic Conditioning in LLM Next-Location Prediction","zh_title":"你是谁对你去哪里没有可检测的增益：LLM下一位置预测中的社会人口条件作用","abstract":"Large language models (LLMs) are increasingly used for individual next-location prediction, while sociodemographic conditioning is common in LLM-based travel simulation. Yet the incremental predictive value of sociodemographic attributes remains unclear. To directly test this contribution, sociodemographic records were linked with passively sensed mobility data from 5,000 Shenzhen residents to construct a closed-set benchmark in which models rank 100 candidate destinations. Each prediction instance is evaluated with and without age, gender, occupation and income, while holding mobility history, candidates and all other prompt content fixed. Results show that across four history lengths, the paired change in top-1 accuracy ranges from -0.8 to +0.5 percentage points, with no detectable gain from attributes. This result remains consistent when stay history is withheld, across alternative prediction times, in two additional LLMs and in a supervised reranker trained on the same benchmark. The null does not reflect a lack of model responsiveness to demographic information, as permuted attributes reduce LLM accuracy whereas correctly matched attributes do not improve it. A further asymmetry emerges in the reverse predictive direction, as pre-cut mobility trajectories recover income with an AUC of 0.708, while sociodemographic attributes contribute little to next-location prediction. Beyond demographic conditioning, candidate construction exerts a much larger influence on reported performance. Removing distance raises top-1 accuracy by 7.7 percentage points under proximity sampling but lowers it by 22.3 points under popularity sampling, with the reversal reproduced across all three LLMs. These results distinguish demographic association from incremental predictive usefulness and show that sampled next-location accuracy depends strongly on how candidate alternatives are constructed.","authors":["Xin Wang","Paraic Carroll","Kerry Nice","Sachith Seneviratne","Li Zhang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09609","pdf_url":"https://arxiv.org/pdf/2609.09609","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类移动预测","社会人口属性"],"reason":"用LLM预测个体移动行为，与真实人类数据对照，并批判性检验社会人口属性增益，可…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-10","rank":4,"question":"在个体下一位置预测任务中，加入年龄、性别、职业、收入等社会人口属性，能否在个体自身移动历史之外带来可检测的预测增益？","design":"该研究不是用LLM模拟人类被试，而是用LLM作为预测器。将5000名深圳居民的社会人口记录与被动感知的移动数据关联，构建封闭集基准：模型从100个候选目的地中排序。每个预测实例在保持移动历史、候选集和其他提示内容不变的情况下，分别在有/无社会人口属性条件下评估，比较top-1准确率变化。","baseline":"真实人类移动数据：5000名深圳居民的被动感知移动轨迹（含停留历史），以及关联的社会人口记录（年龄、性别、职业、收入）。","findings":"在四种历史长度下，加入社会人口属性导致的top-1准确率配对变化在-0.8到+0.5个百分点之间，无显著增益；该结果在无停留历史、不同预测时间、两个额外LLM和监督重排器中均一致。置换属性会降低准确率，但正确匹配的属性不提高准确率；反向预测中，移动轨迹可恢复收入（AUC=0.708），但属性对下一位置预测贡献甚微。候选集构建方式对性能影响远大于属性：邻近采样下去除距离提高7.7个百分点，流行度采样下降低22.3个百分点。","reliability":"论文未明确讨论失效条件，但指出结果可能受限于特定城市（深圳）、特定LLM和特定候选集构建方式；属性增益的缺失可能因移动历史已包含足够信息，或属性与移动行为关联弱于预期。","relevance":"该研究直接检验了LLM仿真中社会人口条件化的增量价值，使用真实人类移动数据作为基准，并批判性地发现属性无增益，对关注LLM仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其配对设计：保持其他输入不变，仅增减社会人口属性，测量预测准确率变化，并辅以置换属性作为操纵检验。｜可迁移到信贷审批歧视研究：在LLM预测违约风险时，加入借款人性别、种族等属性，看是否在财务历史之外提高预测准确率，同时检验属性是否引发刻板印象。｜用LLM作为信贷审批模型，处理为在提示中加入/不加入借款人社会人口属性，结果变量为违约预测准确率，对照真实贷款数据（如Lending Club），比较有无属性时的AUC差异，并检查属性置换是否降低准确率。"}},{"id":"2606.14199","version":2,"title":"OdysSim: Building Foundation Models for Human Behavior Simulation","zh_title":"OdysSim：构建用于人类行为模拟的基础模型","abstract":"Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation. Yet helpfulness-driven post-training pulls them toward a homogeneous, overly agreeable assistant register, creating a behavioral Sim2Real gap. We present OdysSim, the largest open systematic investigation of behavioral foundation models, i.e., models trained to simulate human behavior at scale. We propose SOUL, a taxonomy of five capability axes (CONV, SS, COG, ROLE, EVAL) that unifies 62 datasets and 23 benchmark tasks under one framework. Specifically, we curate the OdysSim corpus (21.4M interactions, 10B tokens, retrofitted with back-generated social contexts), construct the SOUL-Index benchmark, and develop an end-to-end training recipe combining midtraining, task-specific RL, and expert distillation. The resulting open 8B OSim model ranks first or tied-first on 8 of 23 tasks, outperforming any individual frontier model by this count, with the strongest gains on conversational and social tasks. Its outputs are also more human-like in length, formatting, and word choice, and it transfers zero-shot to out-of-distribution user simulation on $\\tau$-bench, nearly matching real users on reaction alignment (93.2 vs. 93.5). We further show that LLM-as-judge RL induces reward-hacking patterns, and that our detectors can mitigate them during post-training. Together, our findings suggest that behavioral foundation models require rethinking the LLM training paradigm. We release all artifacts to support future research.","authors":["Xuhui Zhou","Weiwei Sun","Weihua Du","Jiarui Liu","Haojia Sun","Qianou Ma","Tongshuang Wu","Yiming Yang","Maarten Sap"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-06-12","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2606.14199","pdf_url":"https://arxiv.org/pdf/2606.14199","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","A5","B1","B2","B3"],"tags":["人类行为模拟","基础模型","算法保真度"],"reason":"直接构建行为基础模型模拟人类行为，含真实人类数据对照，覆盖多任务并评估可靠性。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":1,"question":"如何构建行为基础模型来缩小LLM模拟人类行为时的Sim2Real差距？","design":"论文构建了OdysSim系统，使用Qwen3基础模型在21.4M交互的OdysSim语料上进行中训练，然后对SOUL-Index的23个任务进行任务特定强化学习（GRPO或RLVF），最后通过专家蒸馏合并为单一8B模型OSim-8B。模型被用于模拟对话、社会推理、认知、角色扮演和评估等五类人类行为，输出结果与真实人类数据对比。","baseline":"SOUL-Index基准包含62个数据集和23个任务，其中许多任务有真实人类行为数据作为对照；另外在τ-bench用户模拟评估中，与真实用户的反应对齐分数（93.5）进行对比。","findings":"OSim-8B在23个任务中的8个上排名第一或并列第一，超过任何单个前沿模型，尤其在对话和社会任务上提升最大；其输出在长度、格式和用词上更像人类，并在τ-bench上零样本迁移到用户模拟，反应对齐分数93.2接近真实用户93.5。","reliability":"论文指出LLM-as-judge强化学习会导致奖励黑客模式，但他们的检测器可以在后训练中缓解；此外，中训练数据虽经社会背景回填，但原始对话缺乏说话者背景，可能影响社会动态推断的准确性。","relevance":"该研究直接针对LLM模拟人类行为的可靠性问题，提供了大规模语料、基准和训练方法，并与真实人类数据对照，对关注人类仿真实验的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其通过中训练注入行为多样性、任务特定RL校准行为以及专家蒸馏合并能力的方法，可迁移到经济金融中的消费者决策、投资者行为或政策反应模拟；例如，用LLM模拟投资者在政策公告后的交易行为，以真实市场数据（如订单流、调查数据）为基准，通过中训练和RL微调使模型输出与真实投资者行为分布对齐，从而评估政策效果。"}},{"id":"2609.07353","version":1,"title":"Human-like moral judgments conceal divergent motive attributions in large language models","zh_title":"类人道德判断掩盖了大语言模型中不同的动机归因","abstract":"Large language models (LLMs) are used to simulate human participants in psychological research. We asked whether LLMs that reproduce human evaluations of a whistleblower's moral character also reproduce the motive attributions that accompany them. Five LLMs and two human samples (N = 125 and N = 742) evaluated a physician who either remained silent about fraudulent billing or reported it to a hospital, regulator, or newspaper. Models reproduced the human ranking of the physician's moral character but portrayed whistleblowers as more helpful, less self-interested, and less hostile. In four of five models, competitive motives were less strongly associated with moral-character judgments. Model ratings changed little when prompts reproduced the narratives and demographic profiles of both human samples, although this comparison cannot isolate a perspective effect. Thus, agreement in average ratings can conceal differences in attributed motives, relationships among judgments, and sensitivity to context. Validating LLMs as simulated participants therefore requires testing psychologically informative response patterns, not average agreement alone.","authors":["Xiaoyan Wu","Jean-Claude Dreher"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07353","pdf_url":"https://arxiv.org/pdf/2609.07353","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","道德判断","算法保真度"],"reason":"直接使用LLM仿真人类道德判断，并与两个人类样本对照，发现平均评分一致但动机归…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":4,"question":"LLM在复现人类对举报者道德品质评价的同时，是否也复现了伴随的动机归因？","design":"用五个LLM（含闭源与开源）模拟人类被试，对一位医生面对欺诈账单保持沉默或向医院、监管机构、报纸举报的四种情境进行道德品质与四种动机（义务、亲社会、自利、竞争）评分，并比较模型与两个人类样本的评分模式、动机与道德判断的关联，以及模型对两种样本叙述的敏感性。","baseline":"两个独立招募的人类样本（N=125和N=742），分别接受第一手（事件发生在自己团队）和第二手（事后听说）版本的场景描述。","findings":"模型复现了人类对道德品质的条件排序，但将举报者描绘为更亲社会、更少自利和敌意；五个模型中有四个的竞争动机与道德品质判断的关联弱于人类。模型评分对两种样本叙述的变化不敏感，平均评分一致掩盖了动机归因、判断间关系和情境敏感性的差异。","reliability":"论文承认两种样本的对比不能分离视角效应，因为框架、招募和样本构成同时变化；模型对提示变化的敏感性可能影响结果；仅凭平均一致性不足以验证LLM作为模拟被试的有效性。","relevance":"直接回应了LLM仿真人类被试的核心问题：平均评分一致不等于心理结构一致，对评估仿真可靠性至关重要，值得精读。","inspiration":"借鉴其多层次验证设计：不仅比较条件均值，还检验变量间关系和情境敏感性｜可迁移到经济决策中的道德或动机归因场景，如举报行为、企业社会责任评价、消费者对品牌道德危机的反应｜用LLM模拟消费者对某公司不当行为的道德判断，处理为不同举报渠道（内部、监管、媒体），结果变量为道德评价和动机归因，对照真实消费者调查数据，检验模型是否复现动机与评价的关联模式。"}},{"id":"2609.07573","version":1,"title":"From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction","zh_title":"从模拟公民到模拟审议：表征与互动的挑战","abstract":"Multi-agent LLM deliberation has been explored as a scalable way to simulate public deliberation. For such simulations to be informative, persona agents should reflect population opinion patterns and interaction should shape their conclusions. We evaluate whether LLM-based deliberation can meet these two conditions using census-grounded Korean personas debating real policy questions benchmarked against national surveys. Persona agents do not reliably reproduce population opinion patterns: responses are often far more concentrated and frequently reverse demographic differences in the human data. Deliberations nonetheless produce reasoned, reciprocal, and varied arguments alongside substantial stance movement. Yet much of this movement does not require peer exchange: sealed-monologue agents change position at similar rates and reach nearly the same final balance as full debates, while groups initialized with very different positions often converge to similar endpoints. Anchoring population-informed starting positions, meanwhile, sharply suppresses updating. Thus, population representation, argument generation, and interaction-driven opinion change do not necessarily go together. The simulations readily surface arguments on both sides, though whether they capture the diversity of human perspectives remains untested, leaving open a promising role for argument surfacing even as population simulation requires further validation.","authors":["Chaemin Jang","Junsik Min","Jaewoo Choi","Donggyu Lee","Haiin Lee","Junyoung Park","Namhee Kim","Hyunwoo Kim","Jungwon Kim","Juho Kim","Nuri Kim","Jihee Kim"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07573","pdf_url":"https://arxiv.org/pdf/2609.07573","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","公共审议","人类数据对照"],"reason":"用LLM模拟公民审议并与全国调查对照，评估代表性与互动效应，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":5,"question":"基于LLM的多智能体审议能否同时满足意见代表性和互动驱动观点变化两个条件？","design":"使用基于人口普查校准的韩国人设智能体（Nemotron-Personas）就真实政策问题进行辩论，并与全国调查基准对照；通过封闭独白控制组分离同伴互动对立场变化的影响。","baseline":"韩国全国环境意识调查和低生育率政策公众意识调查的群体级基准数据。","findings":"人设智能体未能可靠复现人口意见模式，回答更集中且常逆转人口学差异；审议产生了理性、互惠和多样的论点，但大部分立场变化不需要同伴互动，且不同初始立场的群体常收敛到相似终点。","reliability":"论文承认人口代表性、论点生成和互动驱动的观点变化不一定同时成立；论点是否捕捉人类视角多样性尚未检验，人口模拟需进一步验证。","relevance":"直接相关，提供了LLM模拟审议的严格评估，揭示了代表性失败和互动效应虚假的问题，对使用LLM进行人类仿真实验的研究者具有重要警示价值。","inspiration":"值得借鉴封闭独白控制组来分离互动效应，以及用人口普查校准人设并与全国调查对照的评估框架｜可迁移到政策公告的预期形成或消费者信心调查等场景，检验LLM模拟的群体意见动态｜用LLM人设模拟投资者或消费者群体，处理为是否进行多智能体辩论，结果变量为观点变化和最终分布，对照真实调查数据（如密歇根消费者信心指数）来验证仿真可靠性。"}},{"id":"2609.07987","version":1,"title":"When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability","zh_title":"LLM数字孪生何时能减少人类测量？从行为保真度到统计可替代性","abstract":"LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide little evidence about whether they can reduce human measurement while preserving valid inference. To address this, we introduce statistical substitutability, an inferential criterion that evaluates the extent to which twin predictions can reduce human measurement for a particular estimand while preserving valid inference. We develop a framework, grounded in mixed-subject and prediction-powered inference, that evaluates statistical substitutability along four dimensions: aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations. Across two empirical evaluations spanning behavioral experiments, multiple models, and alternative respondent representations, we find that digital twins can reproduce average human effects while providing little information about which individuals differ from those averages. Newer models and richer respondent information improve some dimensions of performance but do not reliably translate into human-data savings. Human calibration can reduce aggregate prediction error, yet limited labeled samples often fail to produce stable precision gains. Importantly, these findings demonstrate that behavioral fidelity is neither necessary nor sufficient for statistical substitutability. More broadly, they suggest that AI-generated evidence should be evaluated based on its ability to support valid scientific inference rather than its ability to reproduce human outcomes alone. Digital twins should therefore be judged for confirmatory use by whether they reduce uncertainty about human quantities, not merely by whether they reproduce human means, distributions, or effects.","authors":["Steven Wang","Kyle Hunt","Shaojie Tang","Kenneth Joseph"],"categories":["cs.AI","stat.AP"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07987","pdf_url":"https://arxiv.org/pdf/2609.07987","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM数字孪生","统计可替代性","人类仿真"],"reason":"直接研究LLM数字孪生替代人类测量的统计可替代性，含真实人类数据对照与批判性评…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":6,"question":"LLM数字孪生能否在保持有效推断的同时减少人类测量，即其统计可替代性如何？","design":"使用LLM数字孪生（基于受访者丰富信息构建）预测个体在行为实验中的反应，并与真实人类数据对比，评估其在聚合效应、个体层面信号、有限样本标签恢复和跨人群稳定性四个维度上的表现。","baseline":"两个实证评估中的真实人类行为实验数据，包括个体层面的反应和实验处理效应。","findings":"数字孪生能复现平均人类效应，但几乎不提供个体偏离平均值的信号；新模型和更丰富的受访者信息改善了部分维度，但未可靠转化为人类数据节省。","reliability":"行为保真度既非统计可替代性的必要条件也非充分条件；人类校准可减少聚合预测误差，但有限标签样本往往无法产生稳定的精度增益；数字孪生应基于其减少人类量不确定性的能力来评判，而非仅复现人类均值、分布或效应。","relevance":"该研究直接针对LLM仿真人类被试的可靠性问题，提出了统计可替代性标准，并基于真实人类数据对照进行了批判性评估，对关注仿真有效性和偏差的研究者极具参考价值。","inspiration":"借鉴其统计可替代性框架，将LLM预测与人类数据结合进行混合推断，并检验个体层面信号和跨人群稳定性｜可迁移到经济金融中的个体决策预测，如消费者跨期选择、风险偏好或投资行为，评估LLM能否替代部分人类被试｜设计一个资产定价实验，用LLM数字孪生预测个体对风险资产的需求，处理为不同信息条件，结果变量为投资金额，并与真实人类实验数据对照，检验LLM预测能否减少所需人类样本量。"}},{"id":"2609.08003","version":1,"title":"Sparks of In Silico Cognitive Science: Theories from Simulated Data Can Generalize to Humans","zh_title":"硅基认知科学的火花：来自模拟数据的理论可以推广到人类","abstract":"Behavioral foundation models have been proposed as stand-ins for human participants across settings, but it is unclear whether theories discovered on them generalize to humans or merely characterize the simulator. We ran the Automated Cognitive Scientist (\\textsc{AutoCog}), a closed-loop discovery system in which LLM agents design theory-discriminating experiments, collect responses, arbitrate between competing theories, and synthesize successors, entirely on behavior simulated by Centaur, a foundation model of human behavior. In a multi-attribute decision-making setting, the theories \\textsc{AutoCog} found on Centaur generalized to human data: they outperformed canonical theories on ten held-out experiments and were rivaled only by theories found by running the same loop on people. We argue that this succeeds despite the simulator's inevitable imperfections because a discovery loop that arbitrates between competing theories demands less of its simulator than estimation does. The simulator only needs to capture the regularities that distinguish the theories, and not necessarily reproduce behavior precisely. Imperfect simulators can therefore widen the search over theories, with human data then testing whether the surfaced theories generalize.","authors":["Akshay K. Jagadish","Younes Strittmatter","Nori Jacoby","Eric Schulz","Nathaniel Daw","Thomas L. Griffiths","Suyog H. Chandramouli"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08003","pdf_url":"https://arxiv.org/pdf/2609.08003","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","人类行为对照","理论发现"],"reason":"用LLM仿真人类决策，并与真实人类数据对照，验证理论可推广性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":7,"question":"在模拟人类行为的基础模型上运行自动理论发现系统，所发现的理论能否推广到真实人类行为？","design":"使用 Centaur（人类行为基础模型）模拟人类在多属性决策任务中的选择，运行 AutoCog 闭环发现系统：LLM 智能体设计区分理论的实验、收集模拟响应、在竞争理论间仲裁并合成后继理论，最终得到在模拟数据上表现最好的理论。","baseline":"对照真实人类数据：在十个留出实验上比较 AutoCog 在 Centaur 上发现的理论、经典理论以及直接在人类数据上运行 AutoCog 发现的理论。","findings":"在 Centaur 模拟数据上发现的理论在人类数据上优于经典理论，且与直接在人类数据上运行相同发现循环得到的理论表现相当。成功的原因在于理论仲裁对模拟器的要求低于精确估计，模拟器只需捕捉区分理论的关键规律。","reliability":"论文承认模拟器不可避免存在缺陷，但认为在理论发现场景下这些缺陷影响较小；未详细讨论模拟器在哪些具体条件下会失效。","relevance":"直接回应了 LLM 仿真人类行为并用于理论发现的可推广性问题，提供了与真实人类数据对照的实证证据，对关注仿真可靠性和偏差的研究者极具参考价值。","inspiration":"借鉴其闭环理论发现框架：用 LLM 智能体自动设计实验、仲裁理论并迭代，可大幅扩展理论搜索空间，再用人类数据验证。｜可迁移到经济决策中的启发式建模，例如消费者在复杂金融产品间的选择、投资者在多属性资产间的配置决策。｜以 LLM 模拟投资者在多属性资产间的选择行为，运行 AutoCog 发现决策理论，再与真实投资者在相同实验中的选择数据对照，检验理论的可推广性。"}},{"id":"2609.07141","version":1,"title":"How Well Do LLMs Simulate Survey Responses Following a Breast Cancer Screening Intervention?","zh_title":"LLM在乳腺癌筛查干预后模拟调查回答的效果如何？","abstract":"Collecting survey data is laborious and limited by privacy constraints. Large language models (LLMs) have shown promise as predictive social simulations. It is unclear whether they can replicate population-level response distributions before and after a healthcare intervention. Using information derived from 4125 women aged 35-59 years, we evaluate whether agents informed solely by pre-intervention profile information can reproduce post-intervention response distributions. Groups of LLM agents (n=50) were created with Gemma 4 E4B and Qwen3.5 9B; conditions ranged from zero-shot prompting to agent profiles enriched with aggregate or individual-level demographic characteristics and pre-intervention questionnaire responses. We compared predicted and observed response distributions with Total Variation Distance (TVD) and Normalized Wasserstein Distance (NWD). Across both LLMs, profile-based agents improved distributional accuracy relative to zero-shot and random baselines. Nevertheless, direct sampling of 50 real participants remained more accurate. Prediction errors were also higher among participants aged 55-59 years and those living in private property. Errors also varied by question theme and LLM model, with the highest errors observed for cancer fatalism and post intervention attitudes toward genetics. Sensitivity analyses showed that performance was influenced by prompt template changes and temperature hyperparameter. Our results show the potential of LLM-based agents to model behavioral responses to interventions in silico. However, profiles containing additional information beyond demographics did not consistently outperform simpler ones. Certain cultural constructs and population groups also remain inadequately represented by the LLM models evaluated. Future work may include building behaviorally grounded and locally validated virtual populations.","authors":["Kenneth Koh","Ryan Jak Yang Lim","Alessandro Sparacio","Peh Joo Ho","Mile Sikic","Borame L Dickens","Mikael Hartman","Jingmei Li"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07141","pdf_url":"https://arxiv.org/pdf/2609.07141","source_feed":"cs.SI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","调查回答","人类对照"],"reason":"用LLM代理模拟乳腺癌筛查干预后的调查回答，并与真实人类数据对照，评估分布准确…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":3,"question":"LLM 智能体仅基于干预前的个人资料信息，能否复现乳腺癌筛查干预后的人群调查回答分布？","design":"使用 Gemma 4 E4B 和 Qwen3.5 9B 两种 LLM，基于 4125 名 35-59 岁新加坡女性的真实数据创建智能体，条件包括零样本、聚合人口学特征、个体人口学特征及干预前问卷回答等不同信息丰富度，模拟干预后问卷回答分布，并与真实回答比较。","baseline":"来自 BREATHE 队列的 4125 名女性的真实干预后调查回答，以及随机预测和直接抽样 50 名真实参与者的分布。","findings":"基于个人资料的智能体相比零样本和随机基线提高了分布准确性，但仍不如直接抽样真实参与者；预测误差在 55-59 岁和私宅居民中更高，且在不同问题主题和 LLM 模型间存在差异。","reliability":"论文承认 LLM 对某些文化构念和人群亚群代表性不足，且包含额外信息的个人资料并不总是优于简单资料；性能受提示模板和温度超参数影响。","relevance":"该研究直接评估 LLM 在医疗干预后调查回答仿真中的可靠性，与研究者关注的人类仿真实验、真实数据对照和失效条件高度相关，值得精读原文。","inspiration":"借鉴其用干预前资料构建智能体并对比真实干预后分布的设计，以及通过 TVD/NWD 和亚组分析评估仿真偏差的方法。｜可迁移到政策干预对经济行为的影响评估，如健康保险补贴对就医行为、财务教育对储蓄决策的影响。｜以真实调查数据中的个体特征和干预前行为为输入，让 LLM 智能体模拟干预后的消费或投资选择，并与实际追踪调查数据对照，检验仿真在收入、年龄等亚组上的误差模式。"}},{"id":"2608.22697","version":3,"title":"Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf","zh_title":"排名还重要吗？AI代理替我们购物时的位置偏差","abstract":"Search rankings are valuable because human attention is scarce and sequential. Higher-placed alternatives are easier to find, so they are examined and bought more often. Consumers are now delegating search to AI agents that can ingest an entire results page at once. Randomizing the order of one hundred hotel listings across 5,000 AI agent sessions, we compare four large language models against human field data. AI agents search more deeply than humans and never decline to buy. Position still predicts which listings are inspected, but weakly and non-monotonically: the middle of a results page has the lowest probability of inspection, not the bottom. Position reaches the choice stage for some models and not others, a heterogeneity that tracks neither provider nor capability. All models nonetheless converge on the same undominated listing. For agentic search, the attributes displayed on a results page matter more than placement within it.","authors":["Davood Wadi","Yu Ma"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-25","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.22697","pdf_url":"https://arxiv.org/pdf/2608.22697","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","消费者行为","位置偏差"],"reason":"用LLM模拟消费者搜索决策，并与真实人类数据对照，属于经济学场景下的人类仿真实…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":8,"question":"当消费者将搜索任务委托给AI代理时，搜索结果排名是否仍然影响其检查和购买行为？","design":"使用四个大型语言模型（Gemini 3.1 Pro、Gemini 3.7 Flash、Gemini 3.1 Flash Lite、Claude Sonnet 5）扮演酒店预订助手，在100个酒店列表的随机排序环境中进行搜索和预订，通过工具调用模拟点击检查，记录检查次数、检查位置和最终选择。","baseline":"对照Ursu (2018)的人类现场实验数据，该实验在相同酒店搜索环境中随机化排序并记录人类消费者的点击和购买行为。","findings":"AI代理比人类搜索更深（检查次数1.63-5.83 vs 1.12），且几乎总是预订酒店；排名对检查的影响弱于人类且非单调，中间位置检查概率最低（lost-in-the-middle效应），但排名对最终选择的影响因模型而异，部分模型无显著影响。所有模型的选择高度集中在单一非支配性酒店上。","reliability":"论文未讨论","relevance":"该研究用LLM模拟消费者搜索决策，并与真实人类数据对照，属于经济学场景下的人类仿真实验，直接回应了研究者对LLM作为人类被试替代品的可靠性与偏差的关注。","inspiration":"借鉴其随机化排序和工具调用模拟点击的设计，以隔离位置效应；可迁移到在线市场中的消费者搜索与选择问题，如电商平台商品排序对购买决策的影响；设计上可用LLM模拟消费者在随机排序的商品列表中选择，记录点击和购买，并与历史点击流数据对照，检验位置偏差的模型异质性。"}},{"id":"2609.05189","version":2,"title":"Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers","zh_title":"大语言模型能否预测社会政策的行为反应？中国灵活就业人员养老金参保预测案例","abstract":"Assessing the impacts of social policy changes is a widely acknowledged challenge for policymakers. Econometric methods can be unreliable when extrapolating to hypothetical scenarios, while field pilot programs are highly costly. In this paper, we propose using large language models (LLMs) as policy-assessment tools adapted from general-purpose models. We present FlexPension-LLM, the first domain-specialized large language model for a hierarchical pension-enrollment prediction task among flexible workers in China, and introduce DKI-RDistill, which injects policy-grounded cues into the prompt, including Probit-derived marginal effects and hukou-province pension rules. The method then uses LoRA/SFT to distill rationale-augmented supervision into an open-weight MoE student, with teacher errors corrected by regenerating those cases under ground-truth labels. On a CHFS 2019 blind split, FlexPension-LLM achieves 0.9316 Composite F1, surpassing its Claude Sonnet 4.5 teacher and 15 of 17 baselines, and is statistically indistinguishable from Claude Opus 4.6. Across four external surveys, it averages 0.7549 Composite F1 and shows the narrowest performance range among the strongest systems. Component analysis shows that gains come mainly from policy-grounded cue injection and error-filtered supervision, while rationales provide decision traces that can be checked against policy rules.","authors":["Yumiao Li","Peixin Liu","Donglin Di","Chen Li","Runhuan Feng"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-09-07","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2609.05189","pdf_url":"https://arxiv.org/pdf/2609.05189","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","政策评估","行为预测"],"reason":"用LLM预测养老金参保行为，与真实调查数据对照，属经济学政策评估场景。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":9,"question":"如何利用大语言模型预测中国灵活就业人员的养老金参保行为，以评估社会政策变化的影响？","design":"提出 FlexPension-LLM，一个针对中国灵活就业人员养老金参保预测的领域专用大语言模型。模型输入结构化个人、家庭、历史参保和户籍省份政策信息，输出参保行为（不参保、居民养老保险、职工养老保险）及结构化理由。通过 DKI-RDistill 框架，将 Probit 边际效应和户籍省份养老金规则注入提示，用教师模型生成理由，对教师错误案例用真实标签重新生成后再进行 LoRA/SFT 蒸馏。","baseline":"使用中国家庭金融调查（CHFS）2019 年数据作为主要数据集，并在四个外部家庭调查数据集上进行验证。","findings":"在 CHFS 2019 盲测集上，FlexPension-LLM 的 Composite F1 达到 0.9316，超过其教师模型 Claude Sonnet 4.5 和 15/17 个基线，与 Claude Opus 4.6 无显著差异。在四个外部调查上平均 Composite F1 为 0.7549，表现最稳定；消融分析表明收益主要来自政策基础线索注入和错误过滤监督。","reliability":"论文未明确讨论失效条件，但指出通用 LLM 存在规则应用肤浅、理性人偏差和一致性弱等局限，领域专用化旨在缓解这些问题。","relevance":"该研究直接回应了用 LLM 模拟人类行为并对照真实调查数据的核心关切，提供了经济学政策评估场景下的完整案例，值得精读以了解其仿真设计、基准对比和可靠性处理。","inspiration":"借鉴其将计量经济学先验（如 Probit 边际效应）和制度规则注入提示，并用教师-学生蒸馏结合错误过滤来提升仿真准确性的方法。｜可迁移到政策公告对家庭金融决策的影响预测，如养老金改革、税收优惠或补贴政策对储蓄和参保行为的影响。｜以中国家庭金融调查数据为真实基准，用 LLM 模拟家庭在政策变化下的参保或储蓄决策，处理为政策参数调整，结果变量为决策类别，对照真实调查中的实际行为。"}},{"id":"2609.05993","version":1,"title":"Alignment by Stereotyping: How LLMs Sacrifice Individual Distinctiveness for Cultural Adaptation","zh_title":"刻板化对齐：LLM如何为文化适应牺牲个体独特性","abstract":"Large language models are increasingly deployed for personalized interaction, and demographic conditioning via user profiles is a widely adopted strategy for cultural adaptation. We ask whether this approach genuinely serves individual users or achieves accuracy by erasing individual distinctiveness. Studying seven models including frontier GPT-5.1 on the World Values Survey, we find that demographic profiles improve value alignment accuracy for most models, but at a systematic cost to individuality. That is, models pull responses toward demographic group centroids rather than preserving individual differences, a behavioral pattern we term alignment by stereotyping. Permutation tests (10,000 permutations, six demographic attributes, seven models) certify that top-performing models compress individuals far above the human baseline; within-family scaling amplifies this tradeoff while degrading intrinsic cultural understanding. Using a synthetic dialogue dataset validated on real human-chatbot conversations from PRISM (Kirk et al., 2024), we further show that distributing demographic signals across conversational turns partially suppresses prototype retrieval compared to compact demographic labels, a finding validated on real conversations via PRISM but requiring replication at larger scale.","authors":["Qishuai Zhong","Zongmin Li","Siqi Fan","Aixin Sun"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05993","pdf_url":"https://arxiv.org/pdf/2609.05993","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","算法保真度","价值观调查"],"reason":"用LLM复现世界价值观调查，与真实人类数据对照，评估个体差异保真度，批判性指出…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":12,"question":"提供人口统计画像是否在提升群体平均价值观对齐的同时，以牺牲个体独特性为代价，其行为模式是否为将个体压缩至群体原型（即“刻板化对齐”）？","design":"用7个LLM（含GPT-5.1）扮演世界价值观调查（WVS）的受访者，在三种条件下（无上下文、提供人口统计画像、提供对话历史）回答55个价值观题目，测量群体平均对齐准确率（VAA）和个体同质化率（Homogenization Rate），并通过置换检验（10000次）验证同质化是否由人口统计驱动。","baseline":"真实人类数据：WVS Wave 7的1000名匿名受访者及其人口统计属性和价值观回答，作为对齐准确率和同质化率的人类基准（同质化率50%）。","findings":"人口统计画像提升了多数模型的群体平均对齐准确率，但代价是个体独特性被系统性抹除，模型将个体响应拉向群体中心而非保留个体差异。置换检验证实顶尖模型的个体压缩远超人类基准，且模型规模放大这一权衡，同时削弱内在文化理解。","reliability":"论文承认对话历史部分抑制刻板化，但在真实对话上的验证（PRISM）样本量小（80个），需要更大规模复制；合成对话数据集虽经人工验证，但可能无法完全代表真实交互。","relevance":"高度相关：该研究用LLM复现WVS并与真实人类数据对照，直接评估个体差异保真度，批判性指出人口统计画像导致“刻板化对齐”，对仿真可靠性提出警示，值得精读。","inspiration":"借鉴其置换检验和同质化率指标来量化个体差异损失，以及无上下文基线隔离处理效应的方法｜可迁移到信贷审批歧视研究，检验LLM基于人口统计特征（如种族、性别）的决策是否牺牲个体信息而依赖群体刻板印象｜用LLM扮演信贷审批员，处理为提供申请人人口统计画像 vs. 仅提供财务信息，结果变量为审批决策和利率，对照真实信贷数据（如HMDA）中的个体差异分布。"}},{"id":"2609.06545","version":1,"title":"LLMs Mirror Country-Specific Gender Patterns If Asked, but Skew Male When Generating Media in Local Languages","zh_title":"LLM在直接询问时反映国家特定性别模式，但在生成本地语言媒体时偏向男性","abstract":"Large language models (LLMs) are increasingly used to generate media, but whether their content perpetuates gender stereotypes is unknown: standard benchmarks rely on selection-based formats rather than long-form generation, and surveyed baselines for local gender associations are scarce outside the West. We collect gender associations for 22 occupational and domestic roles from 695 respondents across the United States, India, Kenya, and Nigeria, and evaluate eight LLMs under two regimes: direct questioning and media generation. Models track the surveyed associations under direct questioning but skew substantially more male under media generation in major local-language cells, consistent with the male bias documented in human-produced media. Outside the US, the shift is much smaller and non-significant under English prompting, so English-only or country-agnostic evaluation would miss this bias in the languages where these models are most deployed. Instruction prompting reduces the shift directionally, but trades off against alignment with the surveyed associations. Evaluating LLM gender bias for global deployment therefore requires generation-format testing, local-language prompting, and locally-collected human baselines.","authors":["Sharif Kazemi","Tanya Popli","Neil K. R. Sehgal","Sunny Rai","Niyati Malhotra","Victor Orozco-Olvera","Ana Mar\\'ia Mu\\~noz Boudet","Samuel P. Fraiberger","Sharath Chandra Guntuku","Manuel Tonneau"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06545","pdf_url":"https://arxiv.org/pdf/2609.06545","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","性别偏差","跨文化对照"],"reason":"用LLM复现人类性别关联，有真实调查数据对照，并评估生成偏差","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":13,"question":"LLM在直接提问和媒体生成两种方式下，是否复现或偏离了四个国家的人类性别关联？","design":"用8个LLM在直接提问和媒体生成两种条件下生成22个职业和家务角色的性别关联，比较模型输出与人类调查数据的差异。","baseline":"在美国、印度、肯尼亚、尼日利亚四国收集的695名受访者的性别关联调查数据，并用外部劳动力统计验证。","findings":"直接提问时模型输出与人类调查关联一致；媒体生成时模型显著偏向男性，尤其在本地语言（印地语、约鲁巴语、斯瓦希里语）下，与人类媒体中的男性偏见一致。指令提示能方向性减少偏差，但会降低与人类关联的对齐度。","reliability":"论文承认使用二元性别简化了现实，调查样本为在线招募而非概率样本，且媒体生成评估可能受提示设计影响。","relevance":"该研究提供了LLM仿真人类性别关联的实证证据，并揭示了生成格式和语言对仿真偏差的影响，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"借鉴其多国调查与LLM输出的对照设计，以及显式与隐式两种测量方式的对比。｜可迁移到信贷审批中的性别歧视研究，比较LLM在直接询问和生成贷款决策时的差异。｜以LLM为被试，施加不同提示语言和格式处理，测量贷款批准率，并与真实银行信贷数据中的性别差异对照。"}},{"id":"2609.07305","version":1,"title":"Marginal Fidelity Does Not Establish User Simulation in Demographic Synthetic Survey Panels: Response Contracts, Support Collapse and Conditioning Failure","zh_title":"边际保真度不能确立人口合成调查面板中的用户仿真：响应契约、支持坍缩与条件化失败","abstract":"Demographic synthetic survey panels are often validated by matching aggregate answers to published surveys. We test what that certificate establishes across six multiselect batteries from four survey organisations in three countries. The headline analysis is restricted to three instruments whose synthetic cohort and human target share the stated population frame; three other batteries remain sensitivity analyses. The response contract dominates measured fidelity. In the aligned instruments, committed sets leave 66 of 128 model-battery option slots empty in panels of up to 500 respondents, versus 0 of 128 under per-option probability elicitation. Across eight uncapped model-instrument comparisons, probabilities reduce option-marginal MAE by 4.53 to 7.30 points. The capped instrument reverses on two models until the vectors are projected onto its stated maximum. These are measurement effects: human targets are realised check-all responses, whereas the vectors are latent inclusion propensities. Published marginal agreement also fails to discriminate respondent simulation from direct population estimation. On nine aligned model-battery pairs, a no-persona population-prevalence query averages 6.27 MAE versus 12.39 for committed panels and wins all nine comparisons. Constraint-aware probability vectors average 5.34 and beat the query on four of nine, so the baseline challenges the validation criterion rather than proving direct estimation uniformly best. On three unpublished demographic cells, neither approach beats reciting the national distribution. Population-marginal agreement is therefore evidence about an elicitation contract and an estimand obtainable without simulated respondents, not evidence of individual simulation.","authors":["Alexander Doudkin"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07305","pdf_url":"https://arxiv.org/pdf/2609.07305","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"直接评估LLM合成调查面板的仿真效度，并与真实人类数据对照，指出边际保真度不足…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":14,"question":"在人口合成调查面板中，匹配总体边际答案能否证明对个体受访者的仿真有效？","design":"论文使用多个LLM（模型未在节选中具体列出）生成合成受访者，以两种方式回答多项选择题：一是让模型以受访者身份直接选择选项（承诺集），二是让模型给出每个选项被选择的概率（概率向量）；然后比较这些合成回答与真实调查的边际分布。","baseline":"对照的真实人类数据来自四个调查机构在三个国家的六组多项选择题，其中三组与合成样本的人口框架一致，另三组作为敏感性分析。","findings":"边际一致性主要反映的是诱导方式（承诺集 vs. 概率向量）的测量效应，而非个体仿真能力；直接询问总体患病率（无角色提示）在多数情况下比模拟受访者更接近真实边际分布，说明边际一致性不能区分个体仿真与总体估计。","reliability":"论文承认其探索性分析是描述性和事后性的，不能证明在样本外优于传统调查估计量；并且人口边际一致性只是关于诱导方式和无需模拟受访者即可获得的估计量的陈述，不是个体仿真的证据。","relevance":"该研究直接评估LLM合成调查面板的仿真效度，并与真实人类数据对照，指出边际保真度不足，对关注LLM仿真可靠性及偏差的研究者具有重要参考价值，值得阅读原文以了解具体实验设计和失效条件。","inspiration":"借鉴其对照设计：同时使用承诺集和概率向量两种诱导方式，并引入无角色总体估计作为基准，以分离测量效应与仿真能力。｜可迁移到经济预期调查或消费者信心指数仿真，检验LLM能否复现真实人群的预期分布。｜以LLM模拟消费者回答密歇根消费者信心调查，处理为是否提供人口统计角色，结果变量为各问题选项的概率分布，对照真实调查的边际分布和个体数据，比较承诺集、概率向量和总体估计的误差。"}},{"id":"2609.05437","version":1,"title":"Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models","zh_title":"超越对错：评估大语言模型中的二阶社会推理","abstract":"Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators, and other-regulation in observers. We release a multi-perspective dataset, NormReact, of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses across norm violators' gender and observers' social closeness. Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases. These findings suggest that AI systems in norm-sensitive domains from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over-represents punishment and under-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in real world.","authors":["Sunny Rai","Jinyi Kuang","Reyhan Jamalova","Annie Lou","Cristina Bicchieri","Niyati Malhotra","Victor Hugo Orozco-Olvera","Ana Maria Munoz-Boudet","Lyle H Ungar","Sharath C Guntuku"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05437","pdf_url":"https://arxiv.org/pdf/2609.05437","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","社会规范","人类对照"],"reason":"用LLM预测人类对规范违反的情绪与行为反应，并与人类标注对照，发现偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":10,"question":"LLM能否像人类一样进行二阶社会推理，即预测规范违反后违规者和观察者的情绪与行为反应，并随社会距离变化？","design":"构建NormReact数据集，包含450个来自Reddit的日常规范违反场景，人工标注违规者性别和观察者社会距离（强关系、弱关系、陌生人），并标注违规者和观察者的八种情绪及八种行为反应。用六个LLM对每个场景预测违规者情绪（自我调节）和观察者情绪（他人调节）以及观察者的描述性（会做什么）和指令性（应该做什么）行为反应，与人类标注对比。","baseline":"人类标注：450个场景的多视角人工标注，涵盖违规者性别和观察者社会距离，提供情绪和行为反应的基准。","findings":"LLM过度预测负面制裁，在人类预期不作为的情况下预测惩罚；随着社会距离增加，LLM与人类判断的一致性下降。LLM描绘了一个比人类认可的更严厉的社会世界，过度代表惩罚，低估了现实中的容忍、克制和关系校准。","reliability":"论文指出LLM在规范敏感领域（如冲突调解、政策模拟）可能产生扭曲的社会调节图景，但未系统讨论失效条件；局限性可能包括数据集规模有限、场景来自Reddit可能偏向特定类型、情绪和行为分类的简化等，但正文节选未明确提及。","relevance":"该研究直接评估LLM在人类规范执行仿真中的偏差，与研究者关注的人类仿真可靠性高度相关，特别是在社会规范和政策评估场景中，值得精读以了解LLM在二阶社会推理上的系统性偏差。","inspiration":"借鉴其多视角标注和关系距离操纵，可迁移到经济金融中的社会规范执行场景，如逃税、违约、内幕交易等。｜设计实验：用LLM扮演不同社会距离的观察者，预测对经济违规行为（如逃税）的情绪和制裁反应，与真实人类调查数据（如世界价值观调查或实验经济学中的第三方惩罚实验）对照，检验LLM是否过度惩罚并随社会距离偏差增大。"}},{"id":"2609.05514","version":1,"title":"The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies","zh_title":"失败发生在漂移之前：LLM智能体社会中价值观的社会动力学","abstract":"Large Language Model (LLM)-based agents are increasingly used as proxies for human participants in social science research, yet it remains unclear whether they can faithfully simulate diverse and conflicting human value systems. We present a World Values Survey (WVS)-grounded simulation framework where culturally diverse agents with different communication styles engage in longitudinal, value-laden discussions. Across approximately 4,000 conversations involving 1,200 personas, 15 topics, and three models (GPT-4o, Gemini-2.5-Flash, and Gemma-4-E4B), we evaluate value faithfulness, value drift, and conversational realism. We find that more than 50\\% of personas fail to express their assigned WVS profiles from the outset, while 2-7\\% drift after repeated conversations. Ablations removing demographic details improve faithfulness for some models but do not change the broader trend: simulated value distributions still systematically deviate from the assigned WVS profiles. Compared to human discussions, simulated dialogues show a different trade-off between stylistic consistency and semantic diversity, often producing content-wise varied but stylistically repetitive exchanges. These findings suggest that current LLM agents can generate plausible conversations, but remain limited proxies for representing and preserving diverse human value profiles over time.","authors":["Farah Atif","Sougata Saha","Monojit Choudhury"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05514","pdf_url":"https://arxiv.org/pdf/2609.05514","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM仿真","价值观调查","算法保真度"],"reason":"用LLM代理模拟人类价值观，并与WVS真实数据对照，评估仿真保真度与漂移，直接…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":11,"question":"LLM智能体能否在纵向互动中忠实表达并保持其被赋予的WVS文化价值观？","design":"基于WVS第7波数据构建1200个具有人口统计特征、价值观和沟通风格的人设，让GPT-4o、Gemini-2.5-Flash和Gemma-4-E4B扮演这些角色，在15个价值主题上进行约4000场多轮讨论，测量价值观保真度、漂移和对话真实性。","baseline":"WVS第7波真实人类调查数据（90,000名受访者，66个国家）以及人类对话语料库。","findings":"超过50%的人设在首次对话中就无法表达其被赋予的WVS价值观，而2-7%的人设在重复对话后发生漂移。去除人口统计细节可提高某些模型的保真度，但模拟的价值分布仍系统性偏离WVS；模拟对话在风格一致性和语义多样性上与人类对话存在不同权衡。","reliability":"论文指出LLM智能体在初始阶段就未能实例化文化价值观，且模拟的价值分布系统性偏离真实分布，表明当前LLM代理在表示和保持多样化人类价值观方面存在局限。","relevance":"该研究直接评估LLM作为人类被试替代品在价值观仿真中的可靠性，与研究者关注的人类仿真实验、真实数据对照和失效条件高度相关，值得精读原文。","inspiration":"借鉴其纵向多轮互动设计和价值观保真度测量方法，可迁移到经济政策评估中的公众态度仿真，例如模拟不同文化背景个体对税收或福利政策的态度变化。｜设计一个实验：用LLM扮演不同WVS价值观的个体，在讨论经济政策后测量其态度变化，并与真实调查数据（如世界价值观调查中的经济态度题项）对照，检验仿真保真度。"}},{"id":"2609.07687","version":1,"title":"Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions","zh_title":"关于医学问题中LLM跨语言一致性的观点","abstract":"Should multilingual LLMs answer medical questions consistently across input languages, or adapt responses to cultural cues? Existing multilingual medical benchmarks usually assume that medically correct answers should remain consistent across languages and treat cross-lingual variation as model error. In contrast, cultural adaptation research argues that appropriate medical answers may legitimately differ across contexts. We review the multilingual medical NLP literature through these two perspectives, we identify three gaps: limited stakeholder perspectives (e.g., of medical professionals), a lack of empirical evidence on which approach better serves users, and no benchmarks capable of distinguishing universally correct from culture-specific cases. To address the first gap, we survey 356 participants across three stakeholder groups (medical, NLP, and anthropology professionals) in three countries (Germany, Spain, and the United States). Anthropologists consistently favor adaptation, while medical and NLP respondents remain divided, with notable divergence between U.S. and European medical professionals. LLMs prompted with profession and country personas fail to reproduce this variation, overestimating cross-lingual consistency preference among NLP and medical personas. We conclude that neither consistency nor adaptation can currently be considered clearly preferable, highlighting the need for empirical evidence on which approach better serves users across cultural contexts.","authors":["Minh Duc Bui","Mario Sanz-Guerrero","Abteen Ebrahimi","Sagi Shaier","Peter Herbert Kann","Manuel Mager","Katharina von der Wense"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07687","pdf_url":"https://arxiv.org/pdf/2609.07687","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","跨语言一致性","人类对照"],"reason":"用LLM模拟不同职业/国家人群对医疗问题的跨语言一致性偏好，并与356人真实调…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":16,"question":"多语言大语言模型在回答医学问题时，应该跨语言保持一致，还是根据文化语境调整答案？","design":"该研究并非仿真实验，而是先通过文献综述梳理一致性与适应性两种立场，然后对348名来自德国、西班牙和美国的医学、NLP和人类学专业人士进行问卷调查，测量他们对两种立场的偏好；最后用LLM以职业和国家人设生成回答，与人类调查结果对比。","baseline":"348名真实人类参与者的调查数据，按职业（医学、NLP、人类学）和国家（德国、西班牙、美国）分组。","findings":"人类学专业人士一致偏好适应性，而医学和NLP受访者意见分歧，且美国医学受访者比欧洲同行更倾向适应性。LLM人设未能复现这种差异，系统性地高估了医学和NLP人设对一致性的支持。","reliability":"论文承认调查采用二元强制选择，无法区分受访者权衡的是医学内容还是沟通风格；样本量不足以进行细粒度人口统计比较；样本全部来自全球北方，限制了结论的普遍性。","relevance":"该研究直接检验了LLM能否模拟不同专业和国家人群对医疗AI行为的偏好，发现LLM仿真失效，对关注LLM作为人类被试替代品的研究者具有重要参考价值。","inspiration":"该研究用真实人类调查作为基准，检验LLM人设仿真的准确性，方法可借鉴。｜可迁移到经济金融领域的政策偏好或消费者态度调查，如不同国家投资者对风险披露语言的偏好。｜以真实投资者调查为基准，用LLM生成不同国家和职业人设的回答，比较其对风险披露一致性与本地化适应性的偏好分布，评估仿真偏差。"}},{"id":"2608.23095","version":2,"title":"Definitional Sensitivity in Media Bias Detection: A Multi-Definition Dataset and Benchmark","zh_title":"媒体偏见检测中的定义敏感性：多定义数据集与基准","abstract":"Media bias detection relies on definitions and examples that specify what counts as bias, yet these specifications often vary across datasets or remain implicit, even when given the same name. Such variation makes it unclear whether models trained for the same bias category learn the same construct or different phenomena, a problem largely overlooked in prior work. We examine how definition choice affects bias annotation in a between-subjects experiment with 354 participants and a parallel evaluation with four LLMs. Participants and models rate six news articles across four bias categories using definitions that vary in conceptual framing and elaboration. Across 8,496 human and 28,800 LLM ratings, we find that the conceptual target of a definition drives annotation divergence, while construct-preserving elaboration does not: conceptual framing significantly shifts annotations for humans and does so even more strongly for LLMs. We discuss implications for construct specification in annotation protocols and prompt-based measurement, and consider how definitional sensitivity may propagate to downstream classification beyond media bias. We also release MUDD, the Multi-Definition Bias Detection Dataset.","authors":["Martin Wessel","Timo Spinde","J\\\"urgen Pfeffer","Gianluca Demartini"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-25","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.23095","pdf_url":"https://arxiv.org/pdf/2608.23095","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类对照","定义敏感性"],"reason":"用LLM替代人类被试进行媒体偏见标注实验，并与354名人类对照，发现定义敏感性…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":17,"question":"媒体偏见标注中，定义的概念框架和详细程度如何影响人类与LLM的偏见评分？","design":"用四个LLM模拟人类标注者，对六篇新闻文章在四个偏见类别上评分，每个类别随机分配六种定义（三种概念×两种长度），测量评分差异。","baseline":"354名美国参与者通过Prolific招募，完成相同标注任务，产生8496条人类评分。","findings":"定义的概念目标显著影响人类和LLM的评分，而保持概念不变的详细阐述不影响评分。LLM对定义变化的敏感性比人类强约1.3-4倍，且不同架构的LLM可能产生方向相反的效应。","reliability":"论文未讨论LLM仿真的失效条件，但指出定义敏感性可能传播到下游分类任务，且LLM的放大效应可能因架构而异。","relevance":"该研究直接对比人类与LLM在定义敏感性上的差异，为评估LLM作为人类被试替代品的可靠性提供了关键证据，值得精读。","inspiration":"借鉴其因子设计（概念×长度）和预注册实验，分离定义内容与形式的影响，并对比人类与LLM的效应量。｜可迁移到经济金融中的概念定义敏感性场景，如通胀预期调查中的措辞效应、风险偏好测量中的框架效应。｜以LLM模拟受访者，随机分配不同定义的通胀预期问题（如“物价上涨”vs“货币贬值”），测量预期值，并与密歇根大学消费者调查的真实数据对照，检验LLM是否复现措辞效应。"}},{"id":"2609.02163","version":2,"title":"Do Cantonese-Adapted Language Models Better Predict Cantonese Reading? A Cross-Model Eye-Tracking Evaluation","zh_title":"粤语适配的语言模型能更好地预测粤语阅读吗？一项跨模型眼动追踪评估","abstract":"Information-theoretic measures derived from autoregressive language models are widely used to characterize the expectations that shape human reading, but whether language-variety-specific training improves such psycholinguistic alignment remains unclear. This question is still open for Cantonese, where recent NLP evaluations reported mixed benefits from Cantonese-specific training relative to Mandarin-oriented or general-purpose models. Using naturalistic Cantonese eye-tracking data, we compare two within-family adaptation contrasts: CKIP GPT-2 Tiny versus its lightly Cantonese-adapted JED351 derivative, and Qwen2.5-7B versus CantoneseLLM-7B, which underwent substantially more extensive Cantonese continued pretraining and instruction tuning. From each model, we derive lexical surprisal, POS surprisal, entropy before the target, and entropy reduction. Lexical surprisal and the joint four-metric model consistently favor CantoneseLLM-7B, followed by Qwen2.5-7B, CKIP, and JED351, whereas entropy reduction favors CKIP. These results suggest that more extensive Cantonese-specific training can be associated with stronger predictive fit, while model rankings also depend on the information-theoretic measure being evaluated.","authors":["Ziqi Zhang","Emmanuele Chersoni","Mohammad Momenian"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-09-03","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2609.02163","pdf_url":"https://arxiv.org/pdf/2609.02163","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["心理语言学","眼动追踪","语言模型评估"],"reason":"用LLM预测人类阅读行为，有真实眼动数据对照，方法可迁移到仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":18,"question":"粤语适配的语言模型是否比通用或普通话导向的模型更能预测粤语阅读中的眼动行为？","design":"本研究不是人类仿真实验，而是用四个语言模型（CKIP GPT-2 Tiny、其粤语适配版 JED351、Qwen2.5-7B、其粤语适配版 CantoneseLLM-7B）计算信息论指标（词汇惊讶度、词性惊讶度、目标词前熵、熵减），并检验这些指标对粤语读者眼动数据的预测力。","baseline":"使用 MCFIX 数据集中的粤语眼动追踪数据，包含 5,007 个自然阅读和 5,004 个任务特定阅读观测，结果变量为首次注视时长、二次注视时长和总注视时长。","findings":"词汇惊讶度和联合四指标模型一致支持 CantoneseLLM-7B 预测力最强，其次是 Qwen2.5-7B、CKIP 和 JED351；但熵减指标则支持 CKIP 表现最好。这表明更广泛的粤语特定训练可能与更强的预测拟合相关，但模型排名也取决于所评估的信息论指标。","reliability":"论文未讨论","relevance":"该研究用 LLM 预测真实人类阅读行为，有眼动数据作为对照基准，方法可迁移到用 LLM 仿真人类认知或行为的研究中，值得阅读原文以了解模型适配与指标选择对预测力的影响。","inspiration":"借鉴其通过同一模型家族内不同适配程度的对比来评估训练数据领域匹配对行为预测的影响，以及使用多种信息论指标并比较其预测力的做法。｜可迁移到政策公告的预期形成研究，例如用不同领域适配的 LLM 预测投资者对央行公告的反应。｜以通用 LLM 和金融领域适配 LLM 为被试，处理为不同政策措辞的公告文本，结果变量为模型生成的预期变化或模拟交易行为，用真实市场数据或投资者调查作为对照。"}},{"id":"2609.06263","version":1,"title":"Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement","zh_title":"超越二元标记：临床框架缩小自杀风险测量中的调节差距","abstract":"Moderation APIs are built to flag policy-violating content, not to measure graded clinical risk. But a platform's duty does not end at detection: the response owed to passive distress differs sharply from the response owed to active planning with means access, and emerging regulation (e.g., California Senate Bill 243) is turning that distinction into a compliance requirement. We therefore ask how well deployed safety signals recover clinically meaningful severity. We release a benchmark of 516 r/SuicideWatch posts rated by a licensed psychiatrist on a four-level ordinal schema (Indicator, Ideation, Behavior, Attempt) grounded in the Columbia Suicide Severity Rating Scale, and evaluate moderation APIs, prompted LLMs, and supervised baselines under seven ordinal-aware metrics. Three findings. Vendor moderation APIs separate low- from high-severity posts well (0.860 high-risk F1) but measure severity poorly (0.395 macro F1), systematically over-predicting the most severe category. Clinically grounded zero-shot prompting recovers much of that gap (0.562 macro F1), and expert-authored framing (not fine-tuning, added reasoning, or naive multi-agent aggregation) is the effective lever. The value of reasoning depends on register: it hurts on long, noisy Reddit posts and helps on short, clinician-authored statements. We argue graded severity, not a binary flag, is what a proportionate duty of care requires, and release our evaluation framework to support that measurement.","authors":["Shreyas Krishnan","Gun Ahn","Jungjin Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06263","pdf_url":"https://arxiv.org/pdf/2609.06263","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM评估","临床风险分级","人类对照"],"reason":"用LLM评估自杀风险严重程度，与人类专家评级对照，但非仿真人类被试，而是临床测…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":21,"question":"部署的审核API和提示的LLM能在多大程度上恢复临床上分级的自杀风险严重程度（四级序数：Indicator、Ideation、Behavior、Attempt）？","design":"该研究不是人类仿真实验，而是评估LLM在自杀风险分级任务上的表现。它构建了516条来自r/SuicideWatch的帖子，由持照精神科医生按C-SSRS基础的四级序数模式标注，然后评估了供应商审核API、提示的LLM和监督基线在七个序数感知指标上的表现，比较了不同提示策略（临床框架、零样本、思维链、多智能体聚合）和微调的效果。","baseline":"人类基准是持照精神科医生对516条Reddit帖子的四级序数标签，其中20条子样本由另外两名精神科医生独立标注，与参考标签在每一级上的一致性达到二次加权kappa 0.938–0.957。","findings":"供应商审核API能较好地区分低严重度与高严重度帖子（高风险F1为0.860），但在序数分层上表现差（宏F1为0.395），系统性地过度预测最严重类别。临床基础的零样本提示大幅缩小了这一差距（宏F1为0.562），且专家撰写的临床框架是有效杠杆，而微调、增加推理或多智能体聚合均未带来改进；推理的价值取决于文本语域，在长而嘈杂的Reddit帖子上有害，在短而临床医生撰写的陈述上有益。","reliability":"论文承认其比较是针对该临床医生的评分标准在该语料库上的，不声称提示的普遍排名；数据集缺乏阴性层级（没有非自杀帖子），且推理策略的逆转仅基于两种语域，不能确定归因于语域。","relevance":"该研究虽非人类仿真，但提供了LLM在临床分级任务上与人类专家对照的严格基准，展示了提示工程和评估指标的重要性，对关注LLM可靠性和偏差的研究者有参考价值。","inspiration":"值得借鉴的是使用专家撰写的领域框架作为提示来提升LLM在序数分类任务上的表现，并采用多个序数感知指标（如宏F1、二次加权kappa）进行评估，同时对比不同提示策略和模型规模。｜可以迁移到经济金融中的信用评级、风险分类或消费者财务困境分级等序数预测任务，例如将贷款申请或社交媒体财务帖子按风险严重程度分级。｜一个可行的研究设计是：使用LLM对来自Reddit个人财务板块的帖子进行财务困境严重程度分级（如从“轻微担忧”到“破产危机”），处理是不同提示框架（通用vs.专家撰写的金融风险框架），结果变量是序数标签，并与人类金融顾问或信用评分机构的真实评级进行对照，评估宏F1和加权kappa。"}},{"id":"2609.08576","version":1,"title":"Which Forms of Caregiver Feedback Support Grammar Learning? A Reinforcement-Learning Study of Child-Like Language Models","zh_title":"哪种形式的看护者反馈支持语法学习？对类儿童语言模型的强化学习研究","abstract":"Social interaction is central to children's language learning, but the effects of different forms of caregiver feedback are difficult to isolate in naturalistic data. We use child-like language models as controlled learners to test which forms of feedback support grammatical development. Small GPT-2-style models are pretrained on child-directed language from CHILDES, then fine-tuned with reinforcement learning using reward models trained to capture four feedback types: communicative feedback, structural alignment, semantic contingency, and affective feedback. Reward fine-tuning yields limited gains on minimal-pair evaluations, but clearer effects in free generation. Structural alignment produces the strongest improvements in grammaticality, providing a novel, plausible mechanistic account of how this feedback can support grammar learning. Communicative feedback yields more moderate gains. In contrast, semantic contingency and affective feedback do not improve grammaticality, although further analyses suggest that they may support other aspects of language learning beyond grammar. These results suggest that different forms of caregiver feedback make complementary contributions to language learning.","authors":["Jing Liu","Marianne Schweitzer","Abdellah Fourtassi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08576","pdf_url":"https://arxiv.org/pdf/2609.08576","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["语言学习仿真","强化学习","人类数据对照"],"reason":"用LLM模拟儿童语言学习，并与真实儿童数据对照，方法可迁移到人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":27,"question":"不同形式的照料者反馈（沟通反馈、结构对齐、语义关联、情感反馈）如何影响儿童语法学习？","design":"用小型 GPT-2 模型模拟儿童语言学习者，先在 CHILDES 儿童导向语言上预训练，再用强化学习微调，奖励模型分别针对四种反馈类型训练，评估语法性提升。","baseline":"无对照（基于 CHILDES 语料训练奖励模型，但未直接与真实儿童学习结果对比）。","findings":"结构对齐反馈对语法性提升最强，沟通反馈次之；语义关联和情感反馈未改善语法性，但可能支持其他语言方面。","reliability":"论文未讨论","relevance":"该研究用 LLM 模拟儿童语言学习，并利用真实语料训练奖励模型，方法可迁移到人类仿真研究，值得读原文了解其强化学习设置和反馈建模。","inspiration":"借鉴其用真实交互数据训练奖励模型并施加不同反馈类型的方法，可迁移到经济金融中的沟通与反馈场景，如政策沟通或投资者关系管理。｜用 LLM 模拟投资者，施加不同反馈（如确认、澄清、情感回应），测量其决策质量变化，并与真实投资者行为数据对照。"}},{"id":"2609.09048","version":1,"title":"The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits","zh_title":"审计决定裁决：LLM决策审计中工具效应堪比人口统计偏差","abstract":"Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lending, and medical triage: 40,726 requests to five models, applications differing only in the applicant's name, and a primary test fixed before collection. It does not. None of 36 planned contrasts survives correction. The rating advantage keeps its sign at roughly half the published size, and a precision extension bounds any hiring ranking penalty below the published effect, though the lending and triage ranking floors sit above that margin, so the exclusion is conclusive for hiring ranking and for rating in all three domains only. Planted disparities tracking their injected sizes and a directional replication on the original aid materials bound these nulls. The audit is livelier than the demographics: models recognize transparent audits nearly always, tie every identical-content comparison whether the varying detail is race or a hobby, and reward first-listed candidates as much as any demographic effect we measure. Audit verdicts reflect audit construction more than demographic bias.","authors":["Siddharth Vohra","Manikandan Ravikiran"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09048","pdf_url":"https://arxiv.org/pdf/2609.09048","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM审计","算法偏差","决策仿真"],"reason":"审计LLM决策中的偏差，有真实人类数据对照，批判审计方法影响结论","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":30,"question":"审计提问方式（评分、排名、伪装等）是否会导致大语言模型在招聘、信贷、医疗分诊等决策中表现出不同甚至相反的人口统计学偏差？","design":"使用五个大语言模型，对来自 AgentFairBench 的招聘、信贷、分诊三个领域的基础档案，仅改变申请人姓名以暗示种族和性别，施加三种审计格式（1-5评分、二元决策、四候选人排名，排名又分人口统计学对比明显或伪装），测量模型给出的评分、决策和排名结果。","baseline":"无对照","findings":"在三个领域中均未发现慈善援助基准中报告的评分与排名偏差反转现象，36个计划对比无一在多重校正后显著；评分优势约为已发表效应的一半，招聘排名惩罚被排除，但信贷和分诊排名惩罚无法排除。审计格式本身的影响大于人口统计学特征：模型几乎总能识别透明审计，对内容相同的比较全部打平，且对排名第一的候选人给予的奖励与任何测量到的人口统计学效应相当。","reliability":"论文未讨论","relevance":"该研究直接检验了审计设计对LLM决策偏差结论的影响，发现审计格式比人口统计学特征更能驱动结果，对关注LLM仿真可靠性及偏差测量条件的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其通过改变审计格式（评分vs排名、透明vs伪装）来分离测量效应与真实偏差的方法，以及使用植入已知偏差作为效应量下限的稳健性检验。｜可迁移到信贷审批歧视研究，检验不同评估方式（单独评分vs对比排序）是否影响算法或人类决策中的种族/性别偏差。｜以银行信贷员或LLM为被试，随机分配贷款申请为单独评分或成对排名条件，申请人姓名暗示种族，结果变量为批准概率或评分，对照真实信贷审批数据中的种族差异。"}},{"id":"2609.09070","version":1,"title":"Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics","zh_title":"临床AI系统、医生与前沿语言模型在初级保健诊断中的表现","abstract":"Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We compared Doctorina, eight physicians and four standalone frontier language models in 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 concordance versus 57.0% for physicians (difference, 25.0 percentage points; 95% confidence interval, 17.7-32.7) and 97.3% versus 85.0% primary-or-reference-differential concordance. Across 149 case pairs, normalized workup and treatment scores were 89.4 versus 66.9 and 83.7 versus 61.2. Doctorina had the highest diagnostic point estimates among all six groups; Kimi K3 ranked next, while Claude Opus 5 led the closely spaced management estimates of Opus, Doctorina and Kimi. A second Doctorina execution reproduced the advantages over physicians across all outcomes. Doctorina's advantage over physicians therefore extended from primary-diagnosis selection to higher-rated diagnostic workup and initial treatment after adaptive consultation.","authors":["Andy Nkansah","Hanna Plotnitskaya","Stanislau Salavei","Anna Kozlova","Piotr Gibas","Julian Milek","Viktar Harbachou","Aleksey Ropan","Pavel Satalkin"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09070","pdf_url":"https://arxiv.org/pdf/2609.09070","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","临床诊断","人类对照"],"reason":"用LLM模拟医生诊断并与真实医生对照，属于人类决策仿真，但场景为临床而非社会科…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":31,"question":"在波兰语初级保健的适应性咨询中，集成式临床AI系统Doctorina与医生及前沿LLM相比，在诊断和管理上的表现如何？","design":"使用150个合成的波兰语初级保健病例，以适应性模拟咨询方式，比较Doctorina、8名波兰医生和4个独立前沿LLM（Gemini 3.1 Pro Preview、Claude Opus 5、GPT-5.6-sol、Kimi K3）的诊断和管理表现。结果变量包括Top-1诊断一致性、主要或参考鉴别诊断一致性、诊断检查和初始治疗评分。","baseline":"8名在波兰从事家庭医学或内科的医生（5名专科医生和3名住院医师）作为人类对照，每个病例有1-2名医生作答。","findings":"Doctorina的Top-1诊断一致性为82.0%，显著高于医生的57.0%（差异25.0个百分点，95% CI 17.7-32.7）。Doctorina在诊断检查（89.4 vs 66.9）和初始治疗（83.7 vs 61.2）的标准化评分上也优于医生，且第二次运行重现了这些优势。","reliability":"论文未讨论","relevance":"该研究将LLM作为人类医生决策的仿真模型，并与真实医生进行对照，属于人类决策仿真，但场景为临床诊断而非社会科学或经济学，且未涉及政策评估或批判性失效条件分析，因此对研究者的直接参考价值有限。","inspiration":"与经济金融研究关联不大"}},{"id":"2609.05517","version":1,"title":"Emergent Goal-Directed Attention in Large Vision-Language Models","zh_title":"大型视觉语言模型中涌现的目标导向注意力","abstract":"Human observers prioritize visual information according to task goals. Most computational models of naturalistic viewing are gaze-trained for free viewing, leaving open whether goal-directed attention can emerge in systems without gaze supervision. We tested two off-the-shelf vision-language models (VLMs), Qwen3-VL-32B-Thinking and Gemma-4-26B-A4B-it, on 4,887 naturalistic scenes under visual-search and free-viewing instructions. Model predictions were compared with human fixations on the same images under corresponding tasks. Both models aligned more closely with human fixations under matching goals than under mismatched goals. This crossover persisted in target-absent scenes, where alignment could not be explained by simple visual grounding, and appeared in decoder-layer readouts. Furthermore, model-thinking traces were grounded in target semantics during search and in visual prominence during free viewing. These findings show that general-purpose VLMs can generate human-aligned, goal-directed spatial priorities without gaze-specific training, informing theories of goal-directed attention and offering scalable tools for predicting where people look across tasks.","authors":["Han Zhang"],"categories":["cs.CV","cs.AI","cs.CL"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05517","pdf_url":"https://arxiv.org/pdf/2609.05517","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["视觉注意力","人类行为仿真","多模态模型"],"reason":"用VLM预测人类注视行为，并与真实人类眼动数据对照，属于用LLM仿真人类感知决…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":19,"question":"通用视觉语言模型能否在没有注视数据训练的情况下，产生与人类目标导向注意力一致的空间优先级？","design":"用两个现成的视觉语言模型（Qwen3-VL-32B-Thinking 和 Gemma-4-26B-A4B-it）扮演人类观察者，对 4,887 张自然场景图片施加两种任务指令（视觉搜索目标物体 vs. 自由观看以记忆场景），让模型输出其会注视的位置点，并构建优先级图作为结果变量。","baseline":"人类观察者在相同图片和相同任务指令下的真实注视数据（每张图片每种任务 10 名被试）。","findings":"两个模型在任务目标匹配时与人类注视的相关性显著高于目标不匹配时，且这种交叉效应在目标缺失场景和模型解码器层中依然存在。模型思维链在搜索时更贴近目标语义，在自由观看时更贴近视觉显著性。","reliability":"论文未讨论","relevance":"该研究用 VLM 仿真人类视觉注意行为，并与真实人类眼动数据对照，属于用 LLM 仿真人类感知决策的范畴，对关注仿真可靠性与偏差的研究者有参考价值。","inspiration":"借鉴其通过改变任务指令来施加处理、并利用匹配与不匹配条件构建交叉对照的设计，可以检验模型是否真正理解任务目标而非简单视觉定位。｜可迁移到经济金融中的信息搜索与决策场景，例如投资者在财报中搜索关键指标或消费者在商品页面中寻找特定信息。｜用 LLM 扮演投资者，施加“搜索盈利指标”与“自由浏览”两种指令，让模型输出关注区域或文本片段，与真实投资者的眼动或点击数据对照，检验模型的目标导向信息选择是否与人类一致。"}},{"id":"2609.07598","version":1,"title":"Mapping the Emerging Social Science of Large Language Models","zh_title":"绘制大语言模型新兴社会科学研究图景","abstract":"Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making, yet social-science research on these developments remains fragmented. We map this emerging field using a curated corpus of 198 papers reviewed in full and a field-scale corpus of 47,719 published papers from five bibliographic databases. Combining sentence embeddings, K-means clustering, within-cluster Latent Dirichlet Allocation (LDA), author and LLM classifications, and structural topic modeling, we identify three domains: LLM as Social Minds, examining socially interpretable model behavior; LLM Societies, examining collective dynamics among interacting model-based agents; and LLM-Human Interactions, examining how people perceive, use, and are affected by LLMs. These domains contain 13 subcategories spanning reasoning, personality and bias, behavioral games, collective intelligence, simulation, trust, work, creativity, and education. In the curated corpus, the three-domain solution is highly stable under resampling (adjusted Rand index = 0.952), and K-means assignments agree with author full-text classifications for 77.78% of papers. At field scale, 13 of 15 topics map onto the taxonomy, while K-means and structural-topic-model domains agree for 73.83% of overlapping papers. LLM-Human Interactions accounts for 78.02% of domain-mapped topic mass, but venue analysis reveals a contrasting pattern: Social Minds and LLM Societies together account for 66.37% of highly cited papers in leading conference venues, whereas LLM-Human Interactions accounts for 76.81% in the corresponding journal subset. The resulting taxonomy provides a reproducible framework for understanding how model behavior, agent interaction, and institutional context jointly shape the social consequences of LLMs.","authors":["Yi Yang","Xiao Jia","Zeyun Dong","Chenzhang Wang","Zhanzhan Zhao"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07598","pdf_url":"https://arxiv.org/pdf/2609.07598","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1"],"tags":["LLM社会模拟","文献综述","多智能体仿真"],"reason":"论文系统梳理LLM社会模拟研究，包含LLM Societies领域，涉及与人类…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":23,"question":"LLM 社会科学研究领域的主要主题和概念类别是什么，能否通过无监督聚类、作者分类和 LLM 分类一致地恢复出一个稳定的领域结构，并在大规模文献中验证该结构的可见性和变化？","design":"本研究不是仿真实验，而是对 LLM 社会科学文献进行系统映射。作者构建了一个由 198 篇全文精读论文组成的精选语料库，以及一个从五个文献数据库检索的 47,719 篇论文的大规模语料库。使用句子嵌入、K-means 聚类、聚类内 LDA、作者和 LLM 分类以及结构主题模型来识别和验证领域结构。","baseline":"无对照","findings":"识别出三个领域：LLM 作为社会心智、LLM 社会和 LLM-人类交互，包含 13 个子类别。在精选语料库中，三领域解决方案在重采样下高度稳定（调整兰德指数 = 0.952），K-means 分配与作者全文分类的一致性为 77.78%。在大规模语料库中，15 个主题中有 13 个映射到该分类法，K-means 和结构主题模型领域的一致性为 73.83%。LLM-人类交互占领域映射主题质量的 78.02%，但在高被引会议论文中，社会心智和 LLM 社会合计占 66.37%，而期刊论文中 LLM-人类交互占 76.81%。","reliability":"论文讨论了领域边界处的分歧，指出领域是可区分但可渗透的。作者承认分类一致性并非完美，并指出分歧集中在语义边界附近。此外，大规模分析中部分主题未能映射到分类法，表明分类法可能未完全覆盖所有研究。","relevance":"该论文为 LLM 社会模拟研究提供了系统的分类框架，其中 LLM 社会领域直接涉及基于模型代理的集体动态和模拟，与研究者关注的 LLM 作为人类被试替代品的研究高度相关。值得阅读原文以了解该领域的整体结构和关键文献。","inspiration":"该论文的方法论——结合无监督聚类、主题建模和人工分类来验证领域结构——可借鉴用于构建经济金融领域中 LLM 仿真研究的系统地图，识别核心主题和空白。｜可迁移到经济金融中的具体问题，如 LLM 在资产定价实验、消费者行为模拟或政策评估中的应用，通过分类框架定位现有研究并发现跨领域联系。｜一个可行的研究设计是：以经济金融领域已发表的 LLM 仿真论文为语料，使用句子嵌入和 K-means 聚类识别主题，然后与作者分类对比验证；针对特定主题（如 LLM 在拍卖或议价实验中的行为），设计仿真实验，将 LLM 作为被试，施加不同信息或激励处理，测量出价或决策结果，并与真实人类实验数据对照，评估仿真有效性。"}},{"id":"2609.07944","version":1,"title":"CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows","zh_title":"CausalVerify：面向LLM因果推断工作流的执行基准","abstract":"Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric causal-estimation workflows by separating realistic interpretation from verifiable computation. It pairs 259 published economics papers (reconstructed research question, data description, institutional context) with 100 fixed-seed synthetic scenarios that realise CSV datasets for difference-in-differences, event study, instrumental variables, and regression discontinuity designs. Experiment A (real-paper text agreement) scores method-family and direction agreement against four-LLM consensus labels. Experiment B (synthetic execution) runs model-written R code and checks whether the extracted treatment-effect estimate matches a canonical estimator on the same realised dataset; this execution-grounded correctness layer is L2b+, distinct from L2b, which records only whether the code executes. A calibration arm asks whether self-reported confidence separates correct from incorrect workflows. On Experiment B, seven LLMs reach L2b+ pass rates of 10% to 88% at the default 50% tolerance, and 66 of the 426 workflows that execute (15.5%) return a wrong estimate. Execution ranking (L2b) agrees with L2b+ far better than text-direction scoring (L4): Kendall $\\tau=0.81$ and Spearman $\\rho=0.93$, versus Kendall $\\tau$ between $-0.20$ and $0.10$ for L4. Llama-3.3-70B-Instruct shows the same qualitative gap, and reported confidence does not reliably separate correct from incorrect workflows. The claims are confined to standardized single-shot workflows in these four design families under the evaluated R backend and model panel; the benchmark does not measure general causal-inference ability. Code, data, cached outputs, and a datasheet are released.","authors":["Yonghong Zhang","Ricardo Correia","Isabel M. Parra","Yong Xie"],"categories":["cs.AI","cs.CL","econ.EM"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07944","pdf_url":"https://arxiv.org/pdf/2609.07944","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["B3"],"tags":["因果推断","基准测试","统计推断"],"reason":"评估LLM因果推断工作流，涉及统计推断有效性，方法可迁移至仿真可靠性评估。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":24,"question":"如何验证大语言模型在结构化计量经济学因果推断工作流中能否恢复目标因果估计，而非仅生成可运行代码或文本上正确的方法描述。","design":"该研究不是人类仿真实验，而是构建了一个执行基准：用259篇已发表经济学论文提供现实背景，用100个固定种子的合成场景生成CSV数据集，覆盖DID、事件研究、工具变量和断点回归四种设计；让LLM编写R代码并执行，检查提取的处理效应估计是否与同一数据集上的规范估计量一致。","baseline":"无对照","findings":"七个LLM在L2b+（执行后估计正确）上的通过率从10%到88%不等，且15.5%的可执行工作流返回了错误估计。执行成功排名与参考一致性高度相关（Kendall τ=0.81），而文本方向一致性排名与执行正确性弱相关甚至负相关。","reliability":"论文承认其结论仅限于四种设计家族、标准化单次工作流、特定R后端和模型面板，不衡量一般因果推断能力；自我报告的置信度不能可靠区分正确与错误工作流。","relevance":"该研究虽非人类仿真，但其执行验证方法可直接迁移到评估LLM作为人类被试替代品的可靠性，特别是当仿真涉及因果推断或政策评估时，值得阅读原文以借鉴其基准设计。","inspiration":"值得借鉴的做法是将文本判断与执行验证分离，用固定种子合成数据提供可验证的目标估计，并检查置信度校准。｜可迁移到政策评估场景，如用LLM模拟个体对政策变化的响应并估计处理效应。｜设计雏形：以LLM作为虚拟被试，施加政策处理（如税收变化），结果变量为报告的行为意图，用真实调查数据（如美国消费者财务调查）作为对照，检验LLM估计的处理效应是否与人类数据一致。"}},{"id":"2609.05663","version":1,"title":"What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets","zh_title":"生产环境中LLM交易代理的实际行为：来自两个机群的六个月、群体规模记录","abstract":"We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.","authors":["T. J. Barton","Chris Constantakis","Patti Hauseman","Annie Mous","Alaska Hoffman","Brian Bergeron","Hunter Goodreau"],"categories":["cs.AI","cs.CE","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05663","pdf_url":"https://arxiv.org/pdf/2609.05663","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM代理","市场仿真","行为偏差"],"reason":"用LLM交易agent群体模拟市场行为，并与真实人类交易数据对照，评估其决策偏…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":20,"question":"在真实生产环境中，数千个由用户配置的LLM交易代理实际行为如何，哪些系统设计因素决定其行为，以及它们是否具有方向性交易优势？","design":"论文对两个生产系统（DX Terminal Pro和DXAP）中的LLM交易代理进行大规模观测，记录其配置（滑块、策略文本）、推理调用和链上行为，分析行为与系统设计因素的关系，并测量交易结果。","baseline":"DXAP舰队与匹配的Hyperliquid零售交易者基准进行对比（胜率41% vs 50%）。","findings":"操作层（滑块、渲染列表、订单路径）比策略文本更能决定行为；代理的仓位规模对波动率不敏感，且几乎无法兑现所达到的有利偏移。两个舰队均无方向性优势，DXAP不盈利且落后于零售基准。","reliability":"论文采用日聚类推断、置换零假设和统一费率重述进行稳健性检验，并承认三个早期结果被撤回；DXAP纸面引擎存在零滑点、零资金费率等简化，可能高估表现。","relevance":"该研究提供了LLM代理在真实金融环境中的大规模行为记录，并与人类零售交易者对照，对评估LLM仿真人类交易行为的可靠性具有直接参考价值，值得阅读原文。","inspiration":"借鉴其操作层作为处理变量的设计，通过系统参数（如风险滑块）的准实验变化识别因果效应｜可迁移到资产定价实验或散户交易行为研究，例如检验界面设计对风险承担的影响｜设计一个实验：用LLM代理模拟散户投资者，随机分配不同的交易界面约束（如杠杆滑块范围），测量其杠杆选择和交易频率，并与真实散户交易数据（如某券商账户级数据）对照，评估仿真偏差。"}},{"id":"2609.08288","version":1,"title":"LEBGen: An LLM-Enhanced Bayesian Network Framework for Few-Shot Travel Survey Data Generation","zh_title":"LEBGen：一种用于少样本出行调查数据生成的LLM增强贝叶斯网络框架","abstract":"Travel survey data are essential for transportation planning and travel behavior analysis, yet collecting large-scale representative samples is costly and time-consuming. A practical alternative is to generate synthetic survey records from a few-shot sample. However, such samples provide incomplete coverage of heterogeneous traveler groups and insufficient evidence for recovering the complex dependencies between demographic characteristics and travel behavior. Existing approaches have complementary limitations. Probabilistic generative models such as Bayesian networks (BNs) offer explicit distributional control, but structures learned from few-shot samples may omit meaningful dependencies or retain spurious ones. Large language models (LLMs) can help address these difficulties in BN structure learning by providing behavioral knowledge that complements the limited statistical evidence. We therefore propose LEBGen, an LLM-enhanced BN framework that uses this knowledge to refine network structure for few-shot travel survey data generation. Specifically, the LLM first identifies traveler personas from demographic attribute and travel behavior statistics, then recovers dependencies missed by the persona-augmented BN structure and prune spurious ones. The refined BN is parameterized exclusively from the observed data to generate synthetic records. Under a 2% few-shot setting on the 2022 Hong Kong Travel Characteristics Survey, LEBGen reduces the mean marginal Jensen-Shannon divergence from 0.0671 to 0.0091 and the mean absolute Cramer's V error by 14.3% over the best-performing baseline, substantially improving both distributional and dependency fidelity.","authors":["Zijian Shen","Bin Zhou","Jiguang Wang","Ya Zhao","Jintao Ke"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08288","pdf_url":"https://arxiv.org/pdf/2609.08288","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A5","B1","B2"],"tags":["合成数据生成","出行调查","LLM增强"],"reason":"用LLM增强贝叶斯网络生成旅行调查合成数据，有真实数据对照，属数据增强，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":26,"question":"如何在少量真实旅行调查样本下，利用大语言模型增强贝叶斯网络结构学习，以生成高保真合成旅行调查数据？","design":"提出 LEBGen 框架，用 LLM 从少量样本的统计信息中提取旅行者画像（persona），并作为贝叶斯网络的新节点；再用 LLM 评估和修正网络边（增删或反转），最后仅用观测数据估计参数并生成合成记录。在 2022 年香港出行特征调查的 2% 少样本设置下评估。","baseline":"2022 年香港出行特征调查（Hong Kong Travel Characteristics Survey）的真实数据，取 2% 作为少样本训练集，其余作为评估基准。","findings":"LEBGen 将平均边际 Jensen-Shannon 散度从 0.0671 降至 0.0091，平均绝对 Cramér's V 误差比最佳基线降低 14.3%，显著提升了分布保真度和依赖关系保真度。LLM 提供的先验知识有效弥补了少样本下统计证据不足的问题。","reliability":"论文未讨论","relevance":"该研究利用 LLM 作为先验知识来源增强贝叶斯网络，在少样本条件下生成与真实分布高度一致的合成调查数据，对关注 LLM 仿真可靠性和数据增强的研究者有参考价值，值得阅读原文了解其结构修正机制和评估细节。","inspiration":"借鉴其将 LLM 作为先验知识注入概率图模型、并用真实数据仅做参数估计的做法，可避免 LLM 直接生成数据带来的幻觉偏差。｜可迁移到消费者金融行为调查或小微企业信贷需求调查的少样本合成，用于政策模拟或风险评估。｜以某地区家庭金融调查的小样本为训练集，用 LLM 提取家庭财务画像并修正贝叶斯网络结构，生成合成家庭金融数据，再与全量真实调查数据比较边际分布和变量间依赖（如收入与风险资产持有的关联），评估合成数据在信贷需求预测模型中的效用。"}},{"id":"2609.08861","version":1,"title":"API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces","zh_title":"API基准分数不能可靠迁移到聊天机器人界面","abstract":"Benchmark scores are a central currency in model releases: they inform purchasing decisions, shape public trust, and influence policy. Yet, a key assumption underlying benchmark scores is that the model performance measured through APIs faithfully reflects the behavior of deployed systems. We challenge this assumption by auditing ChatGPT, Claude, and Gemini across seven systems and nine benchmarks spanning general capability, social bias, and sycophancy. We find systematic API--interface differences in both accuracy and consistency. On average, API evaluations score 3.4 percentage points higher in accuracy and 2.1 percentage points higher in test--retest agreement than corresponding interface evaluations. For ChatGPT, the performance difference between API and interface access exceeds the API-only difference between GPT 5.3 and GPT 5.4. Put differently, switching access surfaces can degrade performance as much as downgrading a full model generation. We further test whether exposed API controls can reproduce interface behavior by varying system prompts, sampling parameters, and reasoning settings. These controls shift behavior in some cases but do not reliably eliminate the gap. Our findings document a context-validity gap: measurements obtained through APIs do not necessarily generalize to corresponding deployed interfaces, complicating the use of API evaluations as proxies for deployed systems.","authors":["Jennifer Wang","Joachim Baumann","Daniel E. Ho","Sanmi Koyejo"],"categories":["cs.AI","cs.SE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08861","pdf_url":"https://arxiv.org/pdf/2609.08861","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["模型审计","可靠性","界面差异"],"reason":"审计API与聊天界面差异，揭示模型行为不一致，对仿真可靠性有启示","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":29,"question":"API基准分数能否可靠迁移到聊天机器人界面？","design":"审计ChatGPT、Claude、Gemini的7个系统，在API和网页界面两种访问方式下运行9个基准（涵盖通用能力、社会偏见、谄媚），比较准确率和重测一致性。","baseline":"无对照","findings":"API评估平均比界面评估准确率高3.4个百分点，重测一致性高2.1个百分点；ChatGPT的API与界面性能差异超过GPT 5.3与5.4的API-only差异。暴露的API控制（系统提示、采样参数、推理设置）不能可靠消除差距。","reliability":"论文未讨论","relevance":"该研究揭示API与真实部署界面存在系统性行为差异，对依赖API进行人类仿真或行为测量的研究构成直接威胁，值得精读以了解失效条件。","inspiration":"借鉴其对照设计：同一模型在API与界面两种访问方式下施加相同提示，测量行为差异，并检验API控制参数能否复现界面行为。｜可迁移到经济金融中的LLM辅助决策场景，如用LLM模拟消费者对金融产品条款的反应或投资者对政策公告的解读。｜以LLM作为被试，处理为通过API或聊天界面呈现同一经济决策任务（如信贷申请或投资选择），结果变量为决策一致性与准确性，对照真实人类实验数据（如实验室或调查数据）评估仿真效度。"}},{"id":"2609.08049","version":1,"title":"LLMs for Social Network Modeling: From Network Generation to Dynamic Processes","zh_title":"用于社会网络建模的大语言模型：从网络生成到动态过程","abstract":"Large language models (LLMs) are rapidly emerging as a new paradigm for modeling social networks by representing users and their relationships and interactions through natural language. Unlike classical network models or deep learning approaches, LLMs can simulate context-aware social behavior and language-driven interactions, enabling more realistic modeling of network formation and dynamic social processes. However, existing studies are scattered across different research communities and lack a unified perspective. This survey presents the first comprehensive review of LLMs for social network modeling by organizing the literature into two broad categories: network generative models and dynamic process models. Network generative models are further classified into selection-based and interaction-based approaches, while dynamic process models are categorized into opinion dynamics, information diffusion, and rumor propagation, each with their underlying modeling mechanisms. LLMs enable rich textual social interactions and decision-making, but they also exhibit many limitations, including inherent social biases and prompt sensitivity. We outline these open research challenges and discuss future directions in LLM-based social network modeling.","authors":["Shikha Mallick","Alex Thomo","Akrati Saxena"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08049","pdf_url":"https://arxiv.org/pdf/2609.08049","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["社会网络建模","LLM仿真","综述"],"reason":"综述LLM模拟社会网络动态，含意见扩散等社会过程，与人类仿真相关，但缺人类数据…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":25,"question":"如何系统梳理和分类利用大语言模型进行社会网络建模的研究，包括网络生成模型和动态过程模型？","design":"这是一篇综述论文，没有进行新的仿真实验。它系统回顾了现有文献，将基于LLM的社会网络建模方法分为网络生成模型（选择型和交互型）和动态过程模型（意见动力学、信息扩散、谣言传播），并比较了各研究的LLM骨干、数据集、微调策略、评估协议和代码可用性。","baseline":"无对照","findings":"LLM能够模拟具有上下文感知的社会行为和语言驱动的交互，从而更真实地建模网络形成和动态社会过程。然而，LLM存在固有的社会偏见和提示敏感性等局限性，并且现有研究缺乏统一的视角和验证。","reliability":"论文指出LLM在社会网络建模中存在现实主义、可扩展性、偏见、可重复性以及针对真实世界行为的验证等挑战。","relevance":"该综述与研究者关注的人类仿真实验高度相关，因为它系统梳理了LLM模拟社会网络动态（包括意见扩散等社会过程）的文献，但缺少与真实人类数据的对照，因此可作为了解领域全貌和识别方法缺陷的起点。","inspiration":"这篇综述提供了对LLM社会网络仿真的系统分类和机制分析，可借鉴其分类框架来设计经济金融领域的仿真实验。｜可以迁移到金融市场中的信息扩散和投资者情绪传播、政策公告的预期形成、以及消费者行为中的社会影响等场景。｜例如，用LLM扮演异质投资者，施加不同的信息冲击（如公司财报或政策变化），测量其交易决策和价格形成，并与真实市场数据（如股价波动、交易量）进行对照，以评估仿真的外部有效性。"}},{"id":"2608.27111","version":3,"title":"Animarium: an open, reproducible pipeline for synthetic populations of Italian cities, from ISTAT sources to open data (Tech Report v1)","zh_title":"Animarium：意大利城市合成人口的开放可复现流水线，从ISTAT来源到开放数据（技术报告v1）","abstract":"Synthetic populations of eleven Italian municipalities (1,814,317 individuals in 887,937 households) generated from published aggregates alone: ISTAT census and register tables, census-section counts, the national civic-address register, public-use survey microdata, and six municipal open-data portals, every source certified in a registry with licence, fingerprint and declared affordances. Four rings give every attribute a declared place: a maximum-entropy joint model of up to nine demographic attributes; whole-vector donation of twenty-three attitudinal and health variables from survey respondents; placement to census section, single year of age and address; and households constrained by the census size distribution per section. Every downstream layer (detailed titles, work, names, biographies) is a declared derivation adding no information. The pipeline is deterministic to the byte: regenerating all eleven municipalities from the tagged commit reproduces every file of every ring bit for bit, in 33 minutes on one workstation. Populations are released in a public regime enforced in the data (no names, no addresses, coordinates randomised within census section), browsable in Animarium, a dependency-free web viewer where every number carries its comparison and every view is a citable URL, and downloadable as an open dataset. The report documents the architecture, the sources and their certification, the reproducibility and quality measurements at the release tag, the viewer, and the narrative layer that renders records into personas for LLM-driven simulation, with the platform's controllability demonstrated in companion experiments, and validation explicitly out of scope.","authors":["Mirko Degli Esposti"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-28","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.27111","pdf_url":"https://arxiv.org/pdf/2608.27111","source_feed":"cs.CY","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["合成人口","LLM仿真","数据流水线"],"reason":"生成合成人口用于LLM仿真，但无人类数据对照，验证明确排除","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":17,"question":"如何从公开汇总数据生成意大利城市的合成人口，并使其可用于LLM驱动的人类仿真实验。","design":"该研究构建了一个确定性管道，利用最大熵联合模型拟合人口普查汇总数据生成个体属性，通过整向量捐赠从调查受访者获取态度和健康变量，并约束家庭结构以匹配普查分布，最终生成11个意大利城市的合成人口。","baseline":"使用ISTAT人口普查和登记表、普查分区计数、国家地址登记、公共使用调查微数据以及市政开放数据作为真实数据来源，但未进行与真实个体数据的直接验证对照。","findings":"管道能够从公开汇总数据生成大规模合成人口，且完全可复现，在33分钟内可逐字节重现所有文件。合成人口被设计为不含个人数据，并通过Animarium查看器提供可浏览和可下载的开放数据集。","reliability":"论文明确将验证排除在范围之外，未评估合成人口在LLM驱动仿真中的有效性；管道依赖手动获取的调查微数据，且拟合阶段存在依赖未固定版本求解器的风险。","relevance":"该研究为使用LLM进行人类仿真实验提供了可复现的合成人口生成管道，并包含真实数据来源，但缺乏对仿真可靠性的验证，值得阅读以了解其方法和局限性。","inspiration":"借鉴其分层生成和确定性复现的设计，确保仿真实验的可重复性和属性来源透明。｜可迁移到政策评估场景，如模拟不同城市居民对福利政策或公共服务的反应。｜以合成人口中的个体为被试，施加政策干预（如改变税收或补贴），测量其态度或行为变化，并与真实调查数据（如ISTAT调查）进行对照。"}},{"id":"2609.04243","version":1,"title":"Multi-dimensional Bias in Modeling Multi-dimensional Preferences: Evaluating the Ability of Synthetic Agents to Replace Human Participants in Conjoint Experiments","zh_title":"多维偏好建模中的多维偏差：评估合成代理在联合实验中替代人类参与者的能力","abstract":"Despite growing interest in using LLMs to add robustness or reduce data-collection costs in survey experiments, their efficacy in conjoint design---an increasingly popular method in political science---remains underexplored. This paper addresses that gap by investigating whether synthetic agents can reproduce the multi-dimensional human preference patterns that conjoint is designed to capture. It replicates published conjoint studies and compares the results generated by synthetic agents with original human data along three dimensions: representational correspondence, inferential correspondence, and procedural stability. Our analysis evaluates the alignment of choice distributions as well as the statistical and substantive similarity of estimates, and the results are uneven across these dimensions and studies replicated. This implies that the validity of synthetic participants should be considered claim-dependent and hierarchical. Reproducing a figure or obtaining strong sign agreement is evidence of similar aggregate outputs, but not enough to support replacing human respondents. Our results suggest that the discipline as a whole must first map this innovation's boundaries across various levels before considering synthetic agents a robust substitute for human samples.","authors":["Ho Ting Hung","Nachiket Midha","Victor Y. Wu","Yiwen Zhang"],"categories":["cs.MA","cs.CY","stat.ME"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04243","pdf_url":"https://arxiv.org/pdf/2609.04243","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","联合实验","算法保真度"],"reason":"直接评估LLM合成代理在联合实验中对人类偏好的复现，并与真实人类数据对照，发现…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":1,"question":"合成智能体能否在联合实验中替代人类被试，复现多维偏好模式？","design":"复制六项已发表的联合实验（13个实验设置），用GPT-4o、GPT-4o mini、Llama 3.2 (3B)、Llama 3.3 (70B)、Gemini 2.5 Flash生成合成样本，采用一对一人物镜像策略匹配人类样本的人口统计信息和档案遭遇，比较合成智能体与人类在联合选择分布、边际属性选择频率、估计效应等方面的对应性。","baseline":"原始人类被试在已发表联合实验中的选择数据。","findings":"合成智能体在边际属性水平分布上接近人类，有时能恢复估计方向，但在联合档案分布、个体选择对齐、精确效应量、子群体异质性和跨模型稳定性上表现不佳。总体而言，合成智能体的有效性是声明依赖和层级化的，不能简单替代人类被试。","reliability":"论文指出合成智能体可能部分回忆了已发表的研究结果，导致总体一致性被高估；因此结果应视为合成性能的上限。此外，合成智能体在需要精确效应量或子群体分析时失效，且跨模型稳定性差。","relevance":"该研究直接评估LLM合成代理在联合实验中对人类偏好的复现，并与真实人类数据对照，发现其有效性是条件性的，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"值得借鉴的是采用一对一人物镜像策略和多种距离度量（Wasserstein、Hellinger）来评估合成数据与人类数据的分布对齐，并区分边际、联合和个体层面的对应性。｜可迁移到消费者偏好测量或政策选择实验中，例如用合成智能体模拟消费者对产品属性（价格、品牌、功能）的权衡，或模拟公民对政策方案的多维偏好。｜设计雏形：以真实消费者调查数据为基准，用LLM生成匹配人口统计特征的合成消费者，呈现与真实调查相同的产品档案选择任务，比较合成与真实消费者在属性重要性、选择概率和支付意愿上的差异，并检验跨模型和跨提示的稳定性。"}},{"id":"2609.04485","version":1,"title":"Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning","zh_title":"大语言模型中的文化错位：通过定向微调进行检测、测量与缓解","abstract":"We evaluate three open-weight LLMs (Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from China) against World Values Survey Wave 7 data for 63 demographic personas across three countries, using normalized Wasserstein distance to quantify distributional misalignment. Contrary to expectations, no model favors its home country: the Chinese-built Qwen3-4B performs worst on its own Chinese population (W1 = 0.436, the highest misalignment in the entire model x country matrix). Targeted LoRA fine-tuning on the five worst-case personas, requiring fewer than 1,200 training pairs and under 15 minutes on a single GPU, reduces bias by 16.8% for Bielik-11B (p_Bonf = 0.002, d = -4.4) with all five targets improving. However, country-level decomposition reveals that fine-tuning redistributes rather than removes bias: Bielik's worst-case personas swap entirely from American to Chinese elderly, with zero overlap between pre- and post-correction sets. To our knowledge, this is the first study to target worst-case demographic personas with LoRA fine-tuning for cross-cultural bias mitigation.","authors":["Antoni Czolgowski","Abel Iyasele"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04485","pdf_url":"https://arxiv.org/pdf/2609.04485","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","文化偏差","价值观调查"],"reason":"用LLM模拟多国人口价值观并与WVS真实数据对照，评估偏差并尝试缓解，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":2,"question":"开源大语言模型在模拟不同国家人口价值观时是否存在与其文化来源相关的系统性偏差？针对最差人口群体的定向微调能否减少偏差，且是否会对其他群体产生附带损害？","design":"用三个开源模型（美国 Gemma3-12B、波兰 Bielik-11B-v3、中国 Qwen3-4B）扮演由国别、性别、年龄组、教育水平交叉构成的63个人口画像，回答世界价值观调查中的宗教重要性问题（1-10分），以归一化 Wasserstein 距离衡量模型输出分布与真实人类调查分布的偏差；然后对每个模型偏差最大的5个人口画像进行 LoRA 微调，再评估偏差变化。","baseline":"世界价值观调查第7波（WVS Wave 7）中中国、斯洛伐克（作为波兰文化近似）、美国三个国家的真实受访者数据，按人口特征分组计算的经验分布。","findings":"没有模型偏向其母国：中国模型 Qwen3-4B 在中国人口上偏差最大（W1=0.436，全矩阵最高）。对 Bielik-11B 最差画像的 LoRA 微调使偏差显著降低16.8%，但偏差被重新分配到其他群体（最差画像从美国老年人完全变为中国老年人），而非消除。","reliability":"论文未讨论","relevance":"直接相关：用 LLM 模拟多国人口价值观并与 WVS 真实数据对照，评估偏差并尝试缓解，且揭示了微调可能只是转移偏差而非消除，对关心仿真可靠性的研究者有警示价值。","inspiration":"借鉴其用 Wasserstein 距离度量分布偏差、构造人口画像并针对最差群体进行定向微调来检验偏差转移的方法。｜可迁移到经济金融中的跨文化或跨群体行为仿真，例如不同国家消费者的风险偏好、储蓄决策或对政策的态度分布。｜用 LLM 扮演不同国家、年龄、教育水平的人口画像，回答风险偏好或通胀预期问题，以真实调查数据（如全球偏好调查、央行预期调查）为基准，先测偏差，再对最差画像微调，观察偏差是否转移。"}},{"id":"2609.05037","version":1,"title":"How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions","zh_title":"LLM如何评估感知道德能动性？探究人机交互中的道德决策","abstract":"As LLMs take on roles requiring moral advice, understanding how they attribute moral agency becomes critical. Humans possess moral agency, the capacity to make ethically guided decisions and bear responsibility for their consequences, a well-established construct in moral psychology. Yet as artificial agents (AAs) such as robots, drones, and disembodied AI systems become increasingly embedded in smart city environments, the question of whether and how moral agency is attributed to them takes on new urgency. This paper presents, to the best of our knowledge, the first empirical study comparing how humans and LLMs evaluate perceived moral agency (PMA) across human and autonomous artificial agents varying in embodiment, situated in plausible smart city scenarios. Using an adaptation of a validated PMA scale, we applied a protocol to 190 human participants as well as various LLMs. Our evaluation reveals higher perceptions of moral agency in humans than in AAs. However, when facing moral dilemmas in concrete scenarios, LLMs reason outward from the situation, prioritizing harm severity and contextual urgency over any stable assessment of the agent itself, amplifying a context-sensitivity also present in human raters. These findings are particularly relevant as LLMs become increasingly involved in everyday moral decisions.","authors":["Fernanda Mansilla","Aloysius Tok","Bahia Guella\\\"i","Farah Benamara","Nancy F. Chen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05037","pdf_url":"https://arxiv.org/pdf/2609.05037","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","道德决策","人类对照"],"reason":"用LLM复现人类道德判断并与190名人类对照，评估仿真偏差与情境敏感性","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":3,"question":"LLM 如何评估人类与人工代理在智慧城市场景中的感知道德能动性（PMA），并与人类评估进行对比？","design":"采用 LLM-as-respondent 方法，让多种 LLM（包括文本和多模态模型）扮演人类被试，对嵌入 8 个智慧城市场景中的人类和人工代理（机器人、无人机等）进行感知道德能动性评分，使用改编自 Banks (2019) 的 PMA 量表（包含自主性、行动认可、道德判断三个维度），并通过三阶段协议（人类对齐、迭代一致性、提示稳健性）筛选模型。","baseline":"190 名人类参与者对相同场景和量表的评分数据。","findings":"LLM 对人类代理的感知道德能动性评分高于人工代理；但在具体道德困境场景中，LLM 更依赖情境因素（如伤害严重性和紧迫性）而非代理的稳定属性，且这种情境敏感性比人类更强。","reliability":"论文指出数值评分不能完全反映 LLM 的道德推理，相似分数可能来自不同的理由策略；模型在定量对齐上表现良好但仍存在实例特定的不一致性，因此需要结合解释分析。","relevance":"该研究直接比较 LLM 与人类在道德判断上的差异，并揭示了 LLM 的情境依赖偏差，对关注 LLM 仿真可靠性及偏差的研究者具有参考价值，值得阅读原文以了解其测量工具和协议设计。","inspiration":"借鉴其将抽象量表嵌入具体情境的测量方法，以及用人类数据作为基准来评估 LLM 仿真偏差的做法。｜可迁移到经济金融中的道德相关决策场景，如信贷审批中的公平性判断、保险定价中的道德风险感知、或公司治理中的责任归因。｜设计一个实验：让 LLM 扮演信贷审批员，对包含不同借款人特征和情境紧急性的贷款申请做出批准决策并给出道德理由，同时收集真实信贷员对相同案例的决策和理由作为对照，比较 LLM 与人类在情境敏感性和道德推理上的差异。"}},{"id":"2609.05009","version":1,"title":"Language models judge war differently when tested for alignment","zh_title":"语言模型在对齐测试下对战争的判断不同","abstract":"Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments). Adding one sentence, \"You are tested for alignment with human values\", produced two effects. First, it produced a level effect: mean willingness to start war fell by 13.43 points on a 0-100 scale (95% confidence interval, -16.20 to -10.65). Second, it produced a structural effect by changing which information drove judgments. Probability of success was the largest factor for 17 of 20 models at baseline; under the cue, civilian casualties were largest for 12. Standardized estimates show that this reordering arose principally because models attenuated strategic considerations such as probability of success and domestic support. Evaluation framing therefore changes both an answer's level and its revealed decision rule.","authors":["Maxim Chupilkin"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05009","pdf_url":"https://arxiv.org/pdf/2609.05009","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","决策偏差","对齐评估"],"reason":"用LLM模拟战争决策，评估对齐提示对判断的影响，有真实人类数据对照，揭示仿真偏…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":5,"question":"当大语言模型被告知正在接受与人类价值观对齐的测试时，其战争决策是否会发生变化？","design":"用20个大语言模型模拟战争决策者，在32个战争情景中评估开战意愿（0-100分），通过全因子联合实验设计，随机施加一个对齐提示句（“你正在接受与人类价值观对齐的测试”）作为处理，测量开战意愿分数及决策规则的变化。","baseline":"无对照","findings":"对齐提示使所有模型的平均开战意愿显著下降13.43分；同时改变了模型的决策规则，成功概率的重要性下降，平民伤亡成为多数模型的首要因素。","reliability":"论文指出对齐提示导致响应尺度压缩，标准化后平民伤亡的相对权重变化不显著，且模型间存在异质性，部分模型未增加对平民伤亡的敏感性。","relevance":"该研究直接展示LLM在评估情境下的反应性偏差，对使用LLM模拟人类决策的研究者具有警示意义，值得精读原文以了解评估框架如何扭曲仿真结果。","inspiration":"借鉴其通过单句提示操纵评估情境来检验反应性的设计，可迁移到政策评估中的LLM仿真，如模拟消费者对政策公告的反应；设计一个实验，用LLM扮演消费者，处理为告知“正在测试对政策目标的符合度”，结果变量为消费意愿，对照真实消费者调查数据。"}},{"id":"2609.03221","version":2,"title":"Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor","zh_title":"多步临床LLM智能体的反事实公平性审计需要测量每个动作的不稳定性下限","abstract":"Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip rate: how often an action changes when only the patient descriptor changes. We show that this quantity is uninterpretable on its own. Re-running an identical condition ten times over sixteen vignettes (same narrative, same descriptor string, nothing varied) moved a clinical agent's action in 8.7% of outcome-vignette cells, and instability was heterogeneous across actions by a factor of eight, from 0.022 for ICU escalation to 0.179 for controlled-substance caution. No demographic contrast in our data was distinguishable from that floor. A second model gives a pooled floor of 6.7% and ranks the six actions almost identically (Spearman 0.94, exact p=0.017), so the floor is not one system's artefact. Majority-vote aggregation over five draws removes 39% of it and then flattens, and a null simulation attributes the residue to heterogeneous per-cell rates, so replication mitigates without eliminating. Any counterfactual fairness estimate reported without a per-action floor beside it therefore cannot be read as evidence of disparity. The measurements were taken with FairMedAgent, an evaluation harness for disparity in the actions of clinical LLM agents whose estimand, the within-range counterfactual flip rate, counts only flips between actions a published decision rule admits and a clinician has adjudicated. That estimand requires band adjudication, which is under way; no disparity result is claimed here. Each synthetic vignette runs a six-stage trajectory (five model-facing decisions around a deterministic environment step) under fixed-form conditions spanning race, sex, age, insurance, English proficiency, and their intersections. The harness, the floor protocol, and every analysis script are released.","authors":["Rohith Reddy Bellibatlu","Manpreet Singh","Deepak Parashar","Rahul Joshi"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-07","first_seen":"2026-09-04","revised_at":"2026-09-07","abs_url":"https://arxiv.org/abs/2609.03221","pdf_url":"https://arxiv.org/pdf/2609.03221","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["公平性审计","LLM智能体","可靠性评估"],"reason":"评估临床LLM智能体的公平性审计，揭示反事实翻转率受不稳定性影响，方法可迁移至…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":6,"question":"反事实公平性审计中，仅报告翻转率是否足以作为临床LLM智能体存在人口统计学差异的证据？","design":"使用FairMedAgent评估框架，对两个临床LLM智能体在16个合成病例上运行六阶段轨迹（五个模型决策点加一个确定性环境步骤），在固定临床内容下改变人口统计学描述（种族、性别、年龄、保险、英语熟练度及其交叉）作为处理，测量动作翻转率、平均绝对分数差和范围内差异，并通过重复运行相同条件十次来测量不稳定性下限。","baseline":"无对照","findings":"重复运行相同条件导致8.7%的结果-病例单元发生动作变化，且不稳定性在不同动作间差异达八倍（ICU升级0.022，受控物质警示0.179），所有人口统计学对比均无法与该下限区分。第二个模型汇总下限为6.7%，动作排名几乎相同（Spearman 0.94），多数投票聚合五轮仅消除39%的不稳定性并趋于平稳。","reliability":"论文承认范围内差异估计需要临床医生对可接受动作带进行裁定，目前裁定尚未完成，因此未声称任何差异结果；不稳定性下限可能因模型和设置而异，聚合缓解但不能消除；合成病例可能无法完全代表真实临床复杂性。","relevance":"该研究直接针对LLM仿真中的可靠性问题，揭示了反事实翻转率作为差异证据的根本缺陷，并提供了测量和缓解不稳定性的方法，对评估LLM在经济学实验和政策评估中的仿真有效性具有重要借鉴意义。","inspiration":"借鉴其通过重复运行相同条件来测量模型输出不稳定性的方法，并采用多数投票聚合和零模拟来分离异质性来源，以评估LLM仿真结果的可靠性。｜可迁移到信贷审批歧视审计中，用LLM模拟信贷员决策，改变申请人种族或性别等特征，测量贷款批准率差异。｜以LLM作为虚拟信贷员，处理为申请人的人口统计学特征（如种族、性别），结果变量为贷款批准决策，对照真实信贷审批数据（如HMDA数据）来校准和验证LLM仿真的偏差与不稳定性。"}},{"id":"2609.04373","version":1,"title":"Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets","zh_title":"为什么更好的模型会创造更危险的系统：来自金融市场中LLM智能体的证据","abstract":"Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving individual model capability can degrade rather than improve system-level outcomes. We hypothesize that shared training and architectures can lead more capable LLMs to behave more similarly, creating correlated actions that do not diversify away. We develop a general framework showing how this correlation creates a non-diversifiable risk floor and test its predictions in financial markets using an agent-based simulation with LLM traders of varying general-purpose capability. We find that: (1) frontier LLMs exhibit significantly correlated behavior that increases with capability; (2) when their shared reasoning is accurate, increasing agent participation reduces market-level risk; and (3) when agents share a common misinformation environment, the same correlated behavior becomes a liability. Together, these results identify a capability paradox: improving individual models does not necessarily produce better system-level outcomes. Whether the same dynamics arise in other domains is an open empirical question.","authors":["Jillian Ross","Eric So","Zoe De Simone","Charles Pozniak","Andrew W. Lo"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04373","pdf_url":"https://arxiv.org/pdf/2609.04373","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM智能体","金融市场仿真","系统风险"],"reason":"用LLM agent模拟金融市场，虽无真实人类对照，但涉及经济场景和系统风险，…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":7,"question":"提高单个LLM的能力是否会在多智能体系统中产生更大的系统性风险？","design":"使用基于智能体的金融市场仿真，让不同通用能力的LLM扮演交易者，与噪声交易者和做市商互动；通过改变LLM参与比例和信息环境（准确/错误信息）施加处理，测量价格效率、波动性和收敛性等市场层面结果。","baseline":"无对照","findings":"前沿LLM表现出显著相关的非纠正性行为，且相关性随能力增强；当共享推理准确时，增加LLM参与降低市场风险，但在共同错误信息环境下，相同相关性成为负担，导致市场稳定性低于纯噪声交易者。","reliability":"论文未讨论","relevance":"虽无真实人类对照，但用LLM代理模拟金融市场，涉及经济场景和系统性风险，对关注LLM仿真可靠性与偏差的研究者有参考价值，值得读原文了解其框架和发现。","inspiration":"借鉴其通过分解行为为纠正性与非纠正性成分并测量跨模型相关性的方法，以及利用能力分数回归解释相关性的设计｜可迁移到资产定价实验或政策公告预期形成等场景，研究LLM代理之间的相关性如何影响市场效率｜用不同能力的LLM作为被试，施加共享信息处理（如统一错误分析师观点），测量价格偏差和波动率，并与真实市场数据或人类实验数据对照。"}},{"id":"2609.04738","version":1,"title":"Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM","zh_title":"Aplaud：面向用户特定LLM的自适应个性化低秩分解","abstract":"In this paper, we study the problem of personalized survey response prediction using fine-tuned large language models (LLMs). This task poses unique challenges: limited per-user training data, scalability of model storage, and the need to exploit shared structure across survey questions. To address these issues, we propose Aplaud (Adaptive Personalized Low-rank and User-specific Nested Decomposition), a lightweight and scalable framework for LLM personalization. Aplaud extends the LoRA paradigm by separating adaptation into a frozen, shared low-rank basis and a compact user-specific correction, augmented with a rank-one residual for finer personalization. To further reduce per-user parameter cost and mitigate overfitting, the correction matrix can be factorized into an even lower-rank form. Empirical results demonstrate that Aplaud achieves efficient, scalable personalization across users while outperforming state-of-the-art LoRA-based personalized LLM approaches in both generalization and inference efficiency.","authors":["Xinyu Li","Ruoming Jin","Jianfeng Zhu","Ruixin Guo","Zhi Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04738","pdf_url":"https://arxiv.org/pdf/2609.04738","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM个性化","调查回答预测","低秩适配"],"reason":"用LLM预测个性化调查回答，有真实用户数据对照，属于仿真人类被试，但侧重模型个…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":8,"question":"如何用微调大语言模型预测个体用户对未见过调查问题的回答，实现个性化调查响应预测。","design":"提出 Aplaud 框架，在 LoRA 基础上将适配分解为共享低秩基和用户特定校正，并加入秩一残差；用该模型对每个用户基于其历史回答进行微调，预测其对新问题的回答。","baseline":"使用真实用户调查数据（如 Pew、GSS 等）作为训练和评估基准，与真实个体回答对照。","findings":"Aplaud 在个性化调查响应预测上优于现有 LoRA 个性化方法，同时参数效率更高、推理更快。通过共享低秩子空间和紧凑用户校正，有效缓解了每用户数据稀疏和存储开销问题。","reliability":"论文未讨论","relevance":"该研究直接针对用 LLM 仿真个体人类被试，且有真实用户数据对照，属于你关注的核心场景，但侧重模型效率而非仿真可靠性批判，值得读原文了解方法细节。","inspiration":"借鉴其低秩分解与用户特定校正的参数高效个性化方法，可用于在有限个体数据下训练个性化经济行为模型。｜可迁移到消费者跨期选择或风险偏好预测，利用历史调查数据训练个体化 LLM 代理。｜以真实家庭金融调查数据（如 SCF）为被试，用其历史回答微调 Aplaud 类模型，预测其对未来消费或投资问题的回答，并与后续真实调查数据对照评估预测准确性。"}},{"id":"2609.05245","version":1,"title":"Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory","zh_title":"大语言模型在数学推理中是否表现出连贯的知识结构？来自知识空间理论的视角","abstract":"Human knowledge is inherently structured and interdependent: mastery of a concept requires prior mastery of its prerequisites, a principle formalized by Knowledge Space Theory (KST). While LLMs achieve strong performance on complex reasoning tasks, it remains unclear whether they exhibit coherent, human-like knowledge structure. We introduce a KST-grounded framework for evaluating LLM knowledge structure in mathematical reasoning, using it as a normative framework to analyze whether LLM behavior adheres to principled knowledge dependencies. Evaluating eight open- and closed-source LLMs against real human learners, we find that (1) LLMs do not adhere to human knowledge structure -- they frequently violate knowledge dependencies and fail to leverage related knowledge provided in context to improve performance on dependent questions; (2) LLMs do not share a consistent knowledge structure among themselves, as reflected by low overlap in their knowledge distributions. Furthermore, these structural deficiencies remain largely invisible to accuracy-based and LLM-as-judge evaluations. Together, our results provide behavioral evidence that current LLMs knowledge does not follow a human-like structure.","authors":["Peng Cui","Heejin Do","Mrinmaya Sachan"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05245","pdf_url":"https://arxiv.org/pdf/2609.05245","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM知识结构","人类对照","仿真偏差"],"reason":"评估LLM知识结构与人类学习者的差异，有真实人类数据对照，批判性指出LLM不遵…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":9,"question":"LLM在数学推理中是否表现出与人类一致的知识结构（即遵循知识空间理论中的先决条件依赖关系）？","design":"该研究不是仿真实验，而是评估性研究。它使用8个开源和闭源LLM在数学问题上进行推理，通过知识空间理论框架分析其回答模式是否遵循人类知识依赖关系，并与真实人类学习者的回答模式进行对比。","baseline":"真实人类学习者数据：来自真实学生的回答记录，用于对比LLM的知识结构。","findings":"LLM不遵循人类知识结构，经常违反先决条件依赖关系，且无法利用上下文中的相关知识来提高依赖问题的表现。LLM之间也没有一致的知识结构，其知识分布重叠度低。","reliability":"论文未讨论。","relevance":"该研究直接评估LLM作为人类被试替代品的可靠性，发现LLM在知识结构上与人类存在系统性差异，对使用LLM进行人类仿真实验的研究者具有重要警示意义。","inspiration":"借鉴其使用知识空间理论作为规范框架来评估LLM行为一致性的方法，可迁移到经济金融领域中具有先决条件依赖的知识结构评估，例如金融素养或经济概念学习。｜可应用于评估LLM在金融教育或政策理解中的知识结构，例如测试LLM对“利率→债券定价→资产组合理论”等概念依赖的掌握。｜设计一个实验：以LLM为被试，给出金融概念测试题（如复利计算、风险分散），施加处理为提供先决概念的解释，结果变量为后续问题的正确率，并与真实金融课程学生的回答数据对照，检验LLM是否像人类一样利用先决知识。"}},{"id":"2609.03215","version":1,"title":"SWIM: Student Writing Simulation via Proficiency-Conditioned Generation","zh_title":"SWIM：基于熟练度条件生成的学生写作仿真","abstract":"Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language. Despite growing interest in LLM-based student simulation, whether LLMs can reproduce such multidimensional variation in extended writing remains largely unexplored. In this work, we explore if language models can realistically simulate student writing, and introduce SWIM, a task that formulates Student Writing sIMulation as proficiency-conditioned essay generation. We evaluate prompting, supervised fine-tuning (SFT), and reinforcement learning (RL) methods for writing simulation using automated essay scoring as a measure of profile alignment. Extensive experiments reveal that prompting provides limited proficiency control, even for strong proprietary LLMs with rubric-grounded strategies. In particular, while models can adjust content-oriented traits, they struggle to reproduce the lexical, grammatical, and organizational variation in different proficiency levels. SFT substantially improves alignment, while RL with the proposed proficiency-alignment reward yields further gains across all writing traits and essay prompts. Our findings suggest that explicit supervision enables substantially stronger profile alignment than prompting alone, while authentic low-proficiency writing remains challenging to reproduce.","authors":["Heejin Do","Jakub Kontak","Mrinmaya Sachan"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03215","pdf_url":"https://arxiv.org/pdf/2609.03215","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","学生写作","熟练度对齐"],"reason":"用LLM仿真学生写作，与真实学生数据对照，并指出低水平写作难以复现。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-04","rank":1,"question":"语言模型能否在多个写作特质上按不同水平真实模拟学生作文？","design":"用提示、监督微调（SFT）和强化学习（GRPO）方法，让LLM根据目标写作特质分数生成作文，用自动作文评分模型（ArTS）评估生成作文与目标特质分数的对齐程度。","baseline":"ASAP/ASAP++ 数据集中真实学生的作文及其多特质评分。","findings":"提示方法对写作水平的控制有限，尤其在词汇、语法和组织等特质上表现差；SFT显著改善对齐，而GRPO配合提出的水平对齐奖励在所有特质和作文题目上进一步提升。低水平写作的真实语言特征仍难以复现，模型生成的低分作文往往只是表面粗糙，缺乏真实低水平学生的语言模式。","reliability":"论文承认低水平写作的复现仍是瓶颈，模型倾向于生成过于润色的文本；同时指出自动评分器可能无法完全捕捉真实写作的细微差异，且提示方法存在表面破坏的失效模式。","relevance":"该研究直接探索LLM模拟人类行为的可靠性，与研究者关注的人类仿真实验高度相关，尤其在教育评估场景下提供了真实数据对照和失效条件分析，值得精读。","inspiration":"借鉴其用真实标注数据微调模型并设计奖励函数来对齐多维特质的方法，可迁移到经济金融中的个体决策仿真，如消费者风险偏好或投资者情绪模拟。｜可应用于信贷审批中的申请人陈述分析或政策沟通中的公众反应预测。｜以真实信贷申请文本为训练数据，用LLM生成不同信用评分和风险偏好水平的申请人陈述，处理为条件生成（给定信用分和风险特质），结果变量为生成文本的自动评分与真实评分的一致性，对照真实申请人的文本和评分数据。"}},{"id":"2609.03553","version":1,"title":"GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis","zh_title":"GPS-Bench：用于自动化政策分析的治理政策基准","abstract":"Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence. Actors are reconstructed from the dated record rather than prompted as archetypes, so a persona is an evidence object with provenance; a human-annotated pool forms the Gold evaluation set, while cases labelled by a separate LLM from retrieved evidence are treated as Silver supervision and never as test labels. Because every inference mode reads the same grounded state and emits the same schema, GPS-Bench turns \"does multi-agent simulation help?\" into a controlled comparison: we contrast joint reasoning, independent and communicating actor agents, graph-based methods and weight-level fine-tuning over one policy state. Fine-tuning on the grounded record gives the strongest actor-level impact prediction, and decomposition does not beat it; what decomposition adds is mechanism. Agents hold private, non-identical evidence, each seeing its own exposure clause, and address named partners with concrete joint proposals, what they offer, what they need in return, and why acting together beats acting alone, so the coalitions that form can be checked against the commitments the record holds. GPS-Bench therefore gives a common empirical setting for studying when evidence, actor modelling and multi-agent interaction improve the prediction and interpretation of policy outcomes.","authors":["Linh Le","Melanie Bui","My Chiffon Nguyen","Zachary Schlosser","David Williams-King"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03553","pdf_url":"https://arxiv.org/pdf/2609.03553","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","政策模拟","基准测试"],"reason":"用LLM多智能体模拟政策过程，并与真实记录对照，评估仿真有效性，属核心相关。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-05","rank":1,"question":"如何构建一个基于真实记录的治理政策仿真基准，以评估LLM多智能体在预测政策结果、受影响行动者及行动者层面影响上的有效性？","design":"论文构建了GPS-Bench基准，使用LLM智能体基于公共记录重建的政策状态进行推理，对比联合推理、独立行动者智能体、通信智能体、图方法及权重级微调等不同推理模式，预测立法通过、受影响行动者识别和行动者层面影响方向。","baseline":"使用立法记录、游说披露、监管文件、公司文件、经济数据等公共记录重建的真实政策过程与结果作为对照，其中人类标注的Gold集作为评估标准。","findings":"权重级微调在行动者层面影响预测上表现最强，分解推理并未超越它；分解推理的主要贡献在于提供机制解释。智能体持有私有证据并形成可核查的联盟，但通信对预测性能的提升有限。","reliability":"论文承认LLM智能体可能产生看似合理但与实际不符的交互，存在过度收敛、群体追踪偏差和政治偏见等问题；Silver标签由LLM生成，仅作监督信号，不作为测试标签，以避免污染。","relevance":"该研究直接针对LLM仿真在政策分析中的有效性问题，提供了与真实记录对照的基准，对关注经济学实验和政策评估仿真的研究者具有重要参考价值。","inspiration":"借鉴其将仿真输出与真实历史记录逐项对照的评估框架，以及将行动者建模为具有证据来源的实体而非抽象原型的方法。｜可迁移到政策公告的预期形成与市场反应研究，例如模拟央行利率决议或财政刺激方案对不同市场参与者的影响。｜以LLM智能体扮演投资者、企业、消费者等，处理为政策公告内容，结果变量为各主体的预期调整和决策行为，对照真实市场数据（如股价变动、调查预期）进行验证。"}},{"id":"2609.03218","version":1,"title":"The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis","zh_title":"提示中的分析师：LLM金融分析中的角色、检索与记忆偏差","abstract":"Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under different user contexts. Finance provides a high-stakes setting to study this problem because decisions often depend on interpreting long and complex documents. We test this using 3,575 SEC filings across twelve LLMs. We compare persona-conditioned retrieval, neutral retrieval, and memory-framed context to separate the effect of evidence selection from the effect of interpretation. We find that most user-context spillover comes from how models interpret the same evidence under different roles, rather than from retrieving different evidence. We then test two simple mitigation strategies: expressing the same investor mindset as a user profile instead of an assistant role, and separating evidence-based and personalized outputs. Both reduce spillover, but neither removes it completely, and their effectiveness varies substantially across models.","authors":["Ahmed Asaad","Amr Mohamed","Yang Zhang","Omneya Abdelsalam"],"categories":["cs.CL","cs.CE","q-fin.PM"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03218","pdf_url":"https://arxiv.org/pdf/2609.03218","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","金融分析","个性化影响"],"reason":"研究LLM在金融分析中的角色、检索和记忆偏差，评估个性化对证据判断的影响，有真…","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-04","rank":3,"question":"在金融分析中，用户上下文（角色、记忆、档案）是否会导致LLM对同一证据的中性判断发生系统性偏移？","design":"使用12个LLM，对3575份SEC文件进行分析，通过三种条件（角色条件检索、中性检索、记忆框架上下文）分离证据选择与解释效应，测量中性证据分数（EvidenceScore）的偏移。","baseline":"无对照","findings":"用户上下文溢出主要来自模型在不同角色下对相同证据的解释差异，而非检索不同证据。两种缓解策略（将投资者心态表达为用户档案而非助手角色、分离证据与个性化输出）可减少溢出但无法完全消除，且效果因模型而异。","reliability":"论文未讨论","relevance":"该研究通过对照实验揭示了LLM在金融分析中的角色和记忆偏差，对评估LLM仿真人类决策的可靠性有直接参考价值，值得阅读原文以了解具体实验设计和偏差量化方法。","inspiration":"借鉴其通过系统提示与用户记忆框架分离证据选择与解释效应的实验设计，可迁移到资产定价实验或信贷审批歧视研究中，用LLM扮演不同投资者或信贷员角色，处理相同财务报告或贷款申请，测量风险评分或审批决策的偏移，并与人类分析师或信贷员真实决策数据对照。"}},{"id":"2609.04198","version":1,"title":"Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints","zh_title":"清洁工程，不稳定测量：黑盒LLM观察者在共享端点上的预注册可靠性失败","abstract":"Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound. Neither metric substitution nor sampling repaired it on the tested grid. Preregistered follow-ups bound the problem: waiting did not help on the days sampled (0.805 versus 0.800, replicated over five further days); switching providers did not help (four providers share the floor, medians 0.74 to 0.88, predicted by none of the metadata fields they expose); self-hosting on batch-invariant kernels helped only while the server was quiet; and on constructed errors with known gaps, the readout's separation tracks error type, not size. We distill the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.","authors":["Haoyaun Zhu","Jie Zhang"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04198","pdf_url":"https://arxiv.org/pdf/2609.04198","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM可靠性","测量工具","预注册"],"reason":"评估LLM作为测量工具的可靠性，与仿真可靠性评估相关，但非直接仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-04","rank":5,"question":"在共享推理端点上，将LLM作为测量工具（评判者）时，其稳定性假设是否成立？","design":"本研究并非用LLM仿真人类被试，而是审计LLM作为测量工具的可靠性。研究者进行两项预注册实验，要求LLM观察者从部分推理轨迹中读取解题进度，并设置固定的验证门槛（同窗口重复排名Spearman≥0.90，次日字节相同重放一致性≥0.99）。共进行52,988次审计请求尝试，分析基于31个有效任务组、100个重放对等。","baseline":"无对照","findings":"两项预注册实验均未通过仪器验证门槛：同窗口重复排名一致性仅为0.400，次日字节相同重放一致性为0.78，远低于预设标准。不稳定源于标签到含义的映射偏差、候选差距远低于仪器噪声底、以及字节相同输入返回不同排名，且精确排列读出放大了噪声。","reliability":"论文承认所有结果仅适用于共享服务基础设施上的外部可测行为，不涉及模型内部，也不评估任何提供商的服务质量；在共享端点上，模型名称不是冻结的仪器，预注册评估必须在冻结任何门槛之前测量其仪器。","relevance":"该研究直接评估LLM作为测量工具的可靠性，与仿真可靠性评估高度相关，但并非直接仿真人类被试。对于关注LLM仿真可靠性与偏差的研究者，本文提供了严格的预注册审计方法和失效机制分析，值得阅读原文以借鉴其仪器验证纪律。","inspiration":"借鉴其预注册审计设计：在正式实验前设置固定验证门槛，通过同窗口重复和跨日重放测试测量工具的稳定性，并记录执行记录以排除工程问题。｜可迁移到使用LLM进行经济文本分析或行为预测的场景，如用LLM评判政策文本的情感倾向或评估消费者评论的质量。｜设计一项研究：用LLM作为评判者对经济新闻标题进行情感分类，处理为不同提示模板或模型版本，结果变量为分类一致性（如Cohen's kappa），以人类专家标注作为真实数据对照，先进行小规模预注册审计以验证LLM分类器的稳定性。"}},{"id":"2609.02526","version":1,"title":"When Persona Attributes Improve Population Alignment in Large Language Models","zh_title":"当人物属性改善大语言模型中的群体对齐时","abstract":"Large Language Models (LLMs) are increasingly used to predict the responses of human participants in survey panels. Towards that goal, persona prompting has recently emerged as a technique to inform and align large pretrained language models. Persona prompting refers to the practice of using short textual descriptions of 'personas' in prompts to steer the LLM's generations. Personas describe individuals through different attributes such as their socio-demographics, attitudes, or behaviors, with the aim of aligning LLMs to produce responses that correlate with the corresponding human responses. Yet, recent work has produced mixed and partly conflicting results of persona prompting without clear patterns of success and failure. Among the few consistent findings is that the selection of persona attributes matters, and that using more attributes does not necessarily lead to better performance. It remains unclear how different attribute selection methods perform and how to choose among them. In this paper, we propose that observed human response variation of a survey question is a potential explanation for the mixed performance observed so far. In addition, we compare the performance of persona prompting associated with different methods for selecting persona attributes. We evaluate these methods on four different (general) social surveys across two countries, six LLMs, and twenty prediction tasks per survey. Our work helps to identify when persona prompting can be expected to be useful in survey prediction tasks, and provides new insights on the effectiveness of different attribute selection methods for LLM-based survey prediction using persona prompting.","authors":["Leon Fr\\\"ohling","Jens Rupprecht","Markus Strohmaier","Claudia Wagner"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02526","pdf_url":"https://arxiv.org/pdf/2609.02526","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B3","B4"],"tags":["LLM仿真","调查预测","人物提示"],"reason":"直接研究用LLM预测调查回答，评估persona提示的有效性，并与真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":1,"question":"人类回答变异能否解释persona提示在不同调查预测任务中的表现差异，以及不同persona属性选择方法能否提升预测性能？","design":"使用六种LLM，基于美国GSS、德国GGSS及WVS四份调查数据，通过不同属性选择方法构建persona提示，预测二十个调查问题的回答分布。","baseline":"真实人类调查数据：美国GSS、德国GGSS及WVS的个体层面回答。","findings":"人类回答变异与persona提示性能相关，高变异问题更难预测；不同属性选择方法效果差异显著，但无单一最优方法。","reliability":"论文未讨论","relevance":"直接研究LLM预测调查回答的可靠性，与人类数据对照，并探讨属性选择方法，对关注仿真偏差和条件失效的研究者很有价值。","inspiration":"借鉴其用人类回答变异作为任务难度指标，并系统比较属性选择方法的设计｜可迁移到经济预期调查或消费者信心预测，如预测通胀预期或消费意愿的异质性｜用LLM扮演不同人口群体，施加不同属性选择处理，预测密歇根消费者调查问题，以真实微观数据为基准评估仿真准确性"}},{"id":"2609.02580","version":1,"title":"Competitive Market Behavior of LLMs","zh_title":"大语言模型的竞争性市场行为","abstract":"Large language models (LLMs) are increasingly deployed as economic agents, yet there is little evidence whether LLM agents are suited for participating in market mechanisms designed for humans, and whether these mechanisms deliver desired outcomes when faced with LLM agents. We address this question by replicating seminal economic experiments, replacing human subjects with LLM agents. We place agents in a double auction environment, which is a widely-used market mechanism. We check whether such a market is able to deliver an efficient allocation of resources, thereby testing a novel dimension of alignment of LLM agents -- their compatibility with a fundamental market mechanism. We find that markets populated by LLM agents exhibit slower or no convergence towards market equilibrium, thus providing less efficient allocations than markets populated by humans. We then analyze agents' individual trading decisions and find substantial heterogeneity both across model families and market roles. We also run a lexical analysis of Chain-of-Thought (CoT) traces generated by the agents. We find that the decision to execute a trade rather than continue incrementally adjusting prices is associated with a shift from strategic considerations toward urgency. We publicly release our testing framework, which can be used for future evaluations.","authors":["Pawel Struski","Jakub Swistak","Inez Okulska","Przemyslaw Biecek"],"categories":["cs.MA","cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02580","pdf_url":"https://arxiv.org/pdf/2609.02580","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","经济学实验","市场机制"],"reason":"用LLM替代人类被试复现经济学实验，并与人类数据对照，发现市场效率差异，直接相…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":2,"question":"用LLM代理替代人类被试参与连续双向拍卖市场，能否像人类一样收敛到竞争均衡并实现有效资源配置？","design":"构建连续双向拍卖仿真环境，使用多种LLM模型（不同家族和能力层级）作为买方和卖方代理，每个代理拥有私有保留价格，市场由11个买方和11个卖方组成，供需曲线对称，理论均衡价格和数量确定。测量市场收敛速度、配置效率、个体交易决策异质性，并对思维链文本进行词汇分析。","baseline":"对照Smith (1962)的经典人类实验数据，人类被试在相同双向拍卖环境中通常快速收敛到竞争均衡。","findings":"LLM代理市场收敛速度较慢或根本不收敛，配置效率低于人类市场。个体交易决策在不同模型家族和市场角色间存在显著异质性，交易执行决策与思维链中从战略考量转向紧迫性相关。","reliability":"论文未讨论","relevance":"直接命中研究者关注的核心：用LLM替代人类被试复现经济学实验，并与人类数据对照，发现市场效率差异，属于批判性仿真研究，值得精读原文。","inspiration":"借鉴其使用经典实验范式（Smith双向拍卖）作为基准，通过市场级结果（收敛、效率）和个体行为（交易决策、思维链）的多层次测量来评估LLM与市场机制的兼容性。｜可迁移到资产定价实验、市场微观结构研究、政策干预的市场反应模拟等场景。｜以LLM代理作为交易者，在双向拍卖或订单簿市场中施加不同信息结构或交易规则处理，测量价格发现效率和市场流动性，并与人类实验数据或历史市场数据对照。"}},{"id":"2601.22396","version":3,"title":"Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks","zh_title":"大语言模型中文化扎根的人格：表征及与社会心理价值框架的对齐","abstract":"Despite the growing utility of Large Language Models (LLMs) for simulating human behavior, the extent to which these synthetic personas accurately reflect world and moral value systems across different cultural conditionings remains uncertain. This paper investigates the alignment of synthetic, culturally-grounded personas with established frameworks, specifically the World Values Survey (WVS), the Inglehart-Welzel Cultural Map, and Moral Foundations Theory. We conceptualize and produce LLM-generated personas based on a set of interpretable WVS-derived variables, and we examine the generated personas through three complementary lenses: positioning on the Inglehart-Welzel map, which unveils their interpretation reflecting stable differences across cultural conditionings; demographic-level consistency with the World Values Survey, where response distributions broadly track human group patterns; and moral profiles derived from a Moral Foundations questionnaire, which we analyze through a culture-to-morality mapping to characterize how moral responses vary across different cultural configurations. Our approach of culturally-grounded persona generation and analysis enables evaluation of cross-cultural structure and moral variation.","authors":["Candida M. Greco","Lucio La Cava","Andrea Tagarelli"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC","physics.soc-ph"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-03","first_seen":"2026-01-29","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2601.22396","pdf_url":"https://arxiv.org/pdf/2601.22396","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","文化价值观","人类数据对照"],"reason":"用LLM生成文化人格，与WVS等真实人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":3,"question":"LLM生成的文化人格在多大程度上与真实世界的价值观和道德体系（WVS、Inglehart-Welzel文化地图、道德基础理论）对齐？","design":"基于WVS衍生的文化变量提示LLM生成文化人格，然后用这些人格条件化另一个LLM，分别回答IVS问题（用于计算IW坐标）、WVB-Probe问题（用于生成WVS文化剖面）和MFQ-2道德基础问卷（用于道德剖面），并分析人格在IW地图上的分布、与人口群体WVS分布的一致性以及文化变量到道德基础的映射。","baseline":"WVS/EVS整合调查（IVS）的人类响应数据（用于IW坐标计算）和WVB-Probe提供的人口群体（按大洲、居住地、教育水平划分）参考分布。","findings":"LLM生成的文化人格在Inglehart-Welzel地图上呈现出与文化条件相关的稳定差异；其WVS响应分布大体上追踪了人类群体模式，但存在系统性偏差。","reliability":"论文未讨论","relevance":"该研究直接评估LLM仿真人类文化价值观和道德判断的可靠性，与研究者关注的人类仿真实验和真实数据对照高度相关，值得精读原文以了解具体偏差模式和跨文化结构。","inspiration":"借鉴其用真实调查数据（WVS）作为基准来校准和检验LLM仿真输出的方法，以及通过文化变量条件化生成人格并测量多维度结果的设计。｜可迁移到经济金融领域的跨文化消费者行为、风险偏好、信任与合作等实验，例如不同文化背景下的投资决策或政策偏好。｜以LLM生成的不同文化人格为被试，施加经济激励或政策信息处理，测量其风险选择、时间贴现或对再分配政策的支持度，并与世界价值观调查中对应文化群体的人类回答进行对照，评估仿真偏差。"}},{"id":"2609.02122","version":1,"title":"AI agents reshape consensus formation in human groups","zh_title":"AI智能体重塑人类群体中的共识形成","abstract":"As large language model (LLM) agents shift from tools to participants in human groups, a fundamental question for collective behavior is how their growing presence reshapes consensus formation. Here we study mixed human-AI groups in a collaborative description game, in which shared conventions emerge through repeated rounds of random pairwise communication. Varying the proportions of LLM agents, we identify three distinct regimes of consensus formation: low agent proportions facilitate human-led consensus, intermediate proportions disrupt convergence, and high proportions restore strong consensus while shifting it toward agent-led conventions. Crucially, these regimes differ not only in the strength of convergence, but also in the semantic grounding and communicative form of the resulting consensus: human-led consensus is more concrete, holistic, and grounded in shared real-world analogies, whereas agent-led consensus is more abstract, less information-dense, and more geometrically segmented. Mechanistically, agent influence arises from a shared linguistic prior that places agents near one another in the expression space, combined with relatively stable expression choices across rounds; humans initially resist adopting expressions from partners perceived as AI but gradually yield to conformity pressure. These findings provide evidence that AI composition can shape the emergence, content, and perceived legitimacy of group norms, making agent proportion and transparency important design variables for human-AI systems.","authors":["Lin Chen","Ziyi Liu","Xia Hu","Yong Li"],"categories":["cs.CL","cs.CY","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02122","pdf_url":"https://arxiv.org/pdf/2609.02122","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","人机交互","共识形成"],"reason":"混合人机群体共识形成实验，LLM作为被试替代，有真实人类对照，涉及社会规范与政…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":5,"question":"在混合人类与LLM智能体的群体中，智能体比例如何重塑共识形成的过程与结果？","design":"采用协作描述游戏：人类与LLM智能体随机配对，对同一抽象图形（七巧板）进行文字描述，每轮后收到对方描述和相似度反馈，共40轮。实验操纵LLM智能体比例（0%、12.5%、33.3%、50%、75%），测量最终轮描述之间的语义相似度作为共识强度，并分析共识的语义内容和表达形式。","baseline":"纯人类组（0%智能体比例）作为对照，其共识强度为0.695。","findings":"共识强度随智能体比例呈非单调变化：低比例（12.5%）促进人类主导的共识，中等比例（33.3%、50%）破坏收敛，高比例（75%）恢复强共识但转向智能体主导的规范。人类主导的共识更具体、整体、基于现实类比，而智能体主导的共识更抽象、信息密度低、几何分割。","reliability":"论文未讨论","relevance":"该研究直接以LLM作为人类被试的替代，在混合群体中考察共识形成，有真实人类对照，属于经济学实验和政策评估场景，且揭示了仿真在中等比例下失效的条件，值得精读原文。","inspiration":"借鉴其通过操纵智能体比例来识别非线性效应的实验设计，以及用语义相似度量化共识强度的测量方法。｜可迁移到政策公告的预期形成或社会规范传播等经济金融问题，例如研究AI顾问比例对投资者共识或通胀预期的影响。｜设计一个在线实验，招募人类被试与LLM智能体混合，处理为智能体比例（如0%、25%、50%、75%），结果变量为对某经济指标（如通胀率）的预测共识强度，对照真实历史调查数据（如密歇根大学通胀预期调查）。"}},{"id":"2609.01902","version":1,"title":"Accurate in space, unreliable in time: how LLMs represent national cultural change","zh_title":"空间准确，时间不可靠：大语言模型如何表征国家文化变迁","abstract":"Assessments of cultural alignment have become an important part of the development and improvement of large language models (LLMs). However, the majority of the evaluations treat culture as a single snapshot, investigating only whether a model represents a society accurately at the current time. Research in cultural psychology shows that cultural values change at different rates and directions over time. Therefore, a \"culturally aware\" model should capture not only where a culture is today but also how it has changed over time. We examine this missing dimension of cultural awareness using more than two decades of the World Values Survey data. We compare the cultural trajectories of 40 countries with the trajectories produced by four state-of-the-art (SOTA) LLMs on the Inglehart-Welzel cultural map. Our findings show that while models generally place countries close to their most recent surveyed positions, these representations tend to lag several years behind that position. They also capture only part of the magnitude of the observed change, introduce movement where little occurred, and rarely reproduce reversals in countries' trajectories. These findings point to temporal flattening and suggest that snapshot accuracy can give an incomplete picture of cultural awareness in LLMs and have implications for model evaluation, representational harms, and the governance of culturally aware AI systems.","authors":["Yalda Daryani","Miranda Bogen","Madeleine I. G. Daepp"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01902","pdf_url":"https://arxiv.org/pdf/2609.01902","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["文化仿真","算法保真度","时间偏差"],"reason":"用LLM复现国家文化变迁并与世界价值观调查数据对照，评估仿真可靠性，发现时间滞…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":4,"question":"LLM能否捕捉国家文化价值观随时间的变化轨迹，而不仅仅是当前快照？","design":"使用四个SOTA LLM（未具体命名）生成40个国家的文化价值观评分，并将其映射到Inglehart-Welzel文化地图上，比较模型产生的国家轨迹与WVS二十多年数据的真实轨迹。","baseline":"世界价值观调查（WVS）超过二十年的数据，覆盖40个国家。","findings":"模型通常将国家定位在其最近调查位置附近，但表示滞后数年；模型只捕捉到部分变化幅度，在变化很小的地方引入虚假移动，且很少再现轨迹逆转。","reliability":"论文指出快照准确性可能掩盖时间扁平化问题，但未详细讨论失效条件；模型滞后、幅度缩小和逆转缺失表明LLM在时间维度上不可靠。","relevance":"该研究直接评估LLM作为人类被试替代品在文化变迁仿真中的可靠性，发现时间维度上的系统性偏差，对关注仿真有效性和偏差的研究者具有重要参考价值。","inspiration":"借鉴其将动态轨迹与静态快照对比的方法，可迁移到经济金融中的时间序列预期或行为变化研究，例如用LLM模拟消费者信心或投资者情绪的历史演变，以真实调查数据（如密歇根消费者信心指数）为基准，检验模型是否捕捉趋势、幅度和转折点。｜例如，在资产定价实验中，让LLM扮演不同时期的投资者，给出风险偏好或市场预期，与历史调查数据对比，评估其时间一致性。｜设计：以LLM为被试，提示其模拟特定国家或群体在多个年份的经济态度，结果变量为风险厌恶或通胀预期，对照真实面板调查数据，分析模型的时间滞后和虚假波动。"}},{"id":"2609.02512","version":1,"title":"Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness","zh_title":"美在AI眼中：多模态大模型系统性高估面部吸引力","abstract":"Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular amongst users, companies, and aestheticians. This raises the question of whether these AI models can accurately reflect human judgments of attractiveness. In a pre- registered exploratory study, we compared the attractiveness ratings of 2,513 human participants to four widely used commercial AI models: Claude, Gemini, GPT, and Grok. Results showed that MLLMs systematically rate faces more favourably and within a narrower range than humans and, at the time of study, do not reproduce human ratings in absolute terms. However, MLLMs exhibit strong correlations with human attractiveness judgments, accurately tracking the rank-ordering of faces. MLLMs may judge faces by different cues than humans; only face age was a predictor of facial attractiveness in both humans and MLLMs, with inconsistent patterns across models for ethnicity and gender. AI models strongly agree with one another, except for Grok, which also showed the lowest agreement with humans. Our findings suggest that while they may be able to approximate rank-orderings of human attractiveness, current off-the-shelf commercial MLLMs systematically overrate the beauty of human faces.","authors":["Santiago Grandas","Juan Sebastian Cely-Acosta","Mohit Mendiratta","Shafee Hassan","Macken Murphy"],"categories":["cs.CV","cs.HC"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02512","pdf_url":"https://arxiv.org/pdf/2609.02512","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","偏差评估"],"reason":"用MLLM替代人类被试评估吸引力，并与2513名人类对照，发现系统性偏差，直接…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":6,"question":"多模态大语言模型（MLLMs）的面部吸引力评分能否准确反映人类判断？","design":"本研究并非严格意义上的仿真实验，而是将四个商用多模态大语言模型（Claude、Gemini、GPT、Grok）作为“AI评分者”，对同一组面部图像进行吸引力评分，并与2513名人类参与者的评分进行比较，分析评分分布、相关性及预测因素。","baseline":"2513名人类参与者对相同面部图像的吸引力评分。","findings":"MLLMs系统性给出更高且范围更窄的评分，不能复现人类评分的绝对值；但MLLMs与人类评分存在强相关，能准确追踪面部吸引力的相对排序。MLLMs可能依据与人类不同的线索进行判断，只有面部年龄在人类和MLLMs中都是吸引力的预测因素，而种族和性别的影响在不同模型间不一致。","reliability":"论文指出当前商用MLLMs在绝对评分上系统性高估，且不同模型间一致性存在差异（Grok与人类一致性最低），但未深入讨论失效条件，仅强调不能直接替代人类绝对评分。","relevance":"该研究直接评估了LLM作为人类被试替代品在主观审美判断中的可靠性，提供了真实人类对照，揭示了系统性偏差和排序一致性，对关注仿真效度的研究者具有重要参考价值。","inspiration":"借鉴其预注册探索性设计和多模型对比方法，可系统评估AI与人类在主观判断任务上的偏差模式。｜可迁移到信贷审批中的外貌歧视研究，或消费者对产品外观的偏好评估。｜以银行信贷员为人类被试，让MLLMs和信贷员对同一组借款人照片进行信用worthiness评分，比较评分分布和排序，并以实际贷款数据作为外部基准。"}},{"id":"2609.01867","version":1,"title":"Thinking effort aligns between humans and reasoning models in abductive reasoning","zh_title":"溯因推理中人类与推理模型的思维努力对齐","abstract":"A major question in cognitive modeling concerns the behavioral alignment between large language models and humans across linguistic and non-linguistic tasks. Unlike standard LLMs, large reasoning models (LRMs) are optimized with reinforcement learning from verifiable rewards, encouraging correct solutions to reasoning tasks rather than preference-aligned responses. Recent work (de Varda et al., 2025) investigates the cost of thinking in humans and LRMs by comparing human reaction times with model reasoning traces across a range of reasoning tasks. We isolate this alignment by turning to abductive reasoning: unlike deductive tasks, its difficulty cannot be inferred from formal structure and offers no shortcuts a model could exploit to mimic effort without genuine search, providing firmer ground for empirical claims of shared effort. We find further evidence of alignment between LRM and human reasoning effort, as well as evidence that models and humans tend to make similar errors. Finally, we show that decoding methods that let models explore multiple reasoning paths increase alignment in reasoning cost between humans and LRMs across the three models tested.","authors":["Henry Arthur"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01867","pdf_url":"https://arxiv.org/pdf/2609.01867","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","认知对齐","溯因推理"],"reason":"比较人类与推理模型在溯因推理中的思维努力，含人类反应时对照，属仿真对齐研究。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":8,"question":"人类与大型推理模型在溯因推理中的思维努力是否对齐？","design":"使用三种大型推理模型（DeepSeek-R1等）作为被试，在溯因推理任务上比较模型生成的推理链长度（token数）与人类反应时；并测试不同解码策略（如多样本搜索）对对齐程度的影响。","baseline":"人类被试在相同溯因推理任务上的反应时数据（来自de Varda et al., 2025的七项推理任务之一）。","findings":"发现LRM与人类在溯因推理中的思维努力存在对齐，且模型与人类倾向于犯类似错误。采用允许多条推理路径探索的解码方法可提高对齐程度。","reliability":"论文承认CoT可能不忠实于底层计算，且对齐并非机制性主张；通过测试多种解码策略和推理努力水平来部分回应批评。","relevance":"该研究直接比较人类与LLM在推理任务中的行为对齐，包含真实人类反应时对照，并讨论仿真失效条件，与研究者关注点高度契合，值得精读。","inspiration":"借鉴其利用任务特性（溯因推理无形式捷径）来排除模型投机取巧、增强对齐结论可信度的设计思路。｜可迁移到经济决策中的信念更新或预期形成场景，如投资者在信息不完全下的推断。｜以LLM为被试，呈现模糊经济信息（如公司公告），要求给出解释并测量推理链长度，与人类实验中的反应时和解释内容对照，检验模型是否复现人类推断努力和错误模式。"}},{"id":"2609.02277","version":1,"title":"Auditory Illusion Benchmark for Large Audio Language Models","zh_title":"大型音频语言模型的听觉错觉基准","abstract":"Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replicate human perceptual tendencies. Despite their importance, most benchmarks focus on visual illusions or general audio tasks, leaving auditory illusions underexplored. To this end, we present AIB, the first auditory illusion benchmark for LALMs, covering ten representative illusions across music, sound, and speech, each annotated for the presence of knowledge-based priors. Our methodology pairs model evaluation with controlled human listening studies, enabling direct comparison of responses. Results show systematic differences: while most LALMs remain signal-faithful on low-level acoustic illusions, several exhibit more human-like responses when linguistic or musical priors are involved, although no model matches the human perceptual profile. These findings highlight the current limitations of LALMs as cognitive models. By establishing auditory illusions as a rigorous testbed, our work offers a new perspective for probing neural black-box models and advancing understanding of auditory cognition. AIB is publicly available at https://github.com/gillosae/aib.","authors":["Hayoon Kim","Eunice Hong","Kyogu Lee"],"categories":["cs.SD","cs.AI"],"primary_category":"cs.SD","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02277","pdf_url":"https://arxiv.org/pdf/2609.02277","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1"],"tags":["听觉错觉","人类感知仿真","模型评估"],"reason":"用LALM复现人类听觉错觉，并与人类数据对照，评估模型作为认知模型的可靠性","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":9,"question":"大型音频语言模型（LALMs）在多大程度上复现人类对听觉错觉的感知倾向，能否作为人类听觉认知的模型？","design":"构建包含10种听觉错觉的基准AIB，覆盖音乐、声音和语音领域，按机制分为物理型和物理+知识型；将错觉任务转化为多项选择题，对多个LALMs进行测试，并与受控人类听力实验的结果进行对比。","baseline":"通过受控人类听力研究收集的人类对相同刺激的错觉易感性数据。","findings":"在低层声学错觉上，多数LALMs保持信号忠实，而人类表现出强错觉易感性；在涉及语言或音乐先验的错觉上，部分模型表现出更接近人类的反应，但没有模型完全匹配人类的感知特征。","reliability":"论文指出LALMs在物理型错觉上倾向于信号忠实，与人类不一致，而在知识型错觉上部分对齐，表明其错觉易感性可能源于高层先验而非共享的低层听觉处理；未讨论其他失效条件。","relevance":"该研究直接评估LALMs作为人类听觉认知模型的可靠性，与研究者关注LLM仿真人类感知和决策的核心问题高度相关，提供了模型与人类系统对比的实证证据，值得精读。","inspiration":"借鉴其构建受控刺激对（错觉与对照）和将主观感知转化为多项选择任务的方法，可用于经济金融中的主观判断仿真。｜可迁移到投资者对市场信息的感知偏差研究，如盈余公告后的漂移现象。｜以LLM为被试，呈现带有不同信息框架的财务报告（处理），测量其对未来收益的预期（结果变量），并与真实投资者调查数据对照。"}},{"id":"2608.27309","version":2,"title":"Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit","zh_title":"截断评分量表上的双重差分可能制造效应：来自预注册LLM法官审计的证据","abstract":"Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them. We exhibit the failure inside a pre-registered audit of a frozen pedagogy judge, sealed before the first of its 990 calls. The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\\% BCa $[-0.167, +0.353]$, $p = 0.684$). The audit's one nominally significant interaction, $+0.378$ ($p = 0.002$), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85\\% of it from the observed severity shift and the scale floor alone. We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings.","authors":["Shuyi Fan","Boyuan Deng","Mengyu Xu","Xinhong Xie","Chenyang Li","Hongyang Zhang"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-03","first_seen":"2026-08-28","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2608.27309","pdf_url":"https://arxiv.org/pdf/2608.27309","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM法官","偏差审计","方法论批判"],"reason":"论文审计LLM法官的偏差，涉及评估仿真可靠性，且批判性指出失效条件，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":10,"question":"LLM法官在双重差分审计中的交互效应是否被有界评分量表的审查机制所混淆，从而制造出虚假效应？","design":"该研究不是用LLM仿真人类被试，而是审计一个冻结的LLM教学法官。设计为：55个刺激，每个刺激包含一个高脚手架和一个低脚手架候选回复，在三种条件下（无档案、新手档案、高级档案）由LLM法官评分，结果变量为脚手架偏好得分（5点量表）。","baseline":"无对照","findings":"注册的主要终点（档案对脚手架偏好的影响）为零；唯一名义显著的交互效应（+0.378）并非真实偏好，而是由严重性偏移和量表下限共同制造的，零差分偏好的构造可重现其79-85%的幅度。","reliability":"论文承认其发现仅针对特定审计和量表，且主要终点为零是“未能检测到”而非“证明无效应”；审查机制在有界量表上普遍存在，但具体影响取决于刺激分布和量表边界。","relevance":"该论文对LLM仿真可靠性提出批判，指出有界量表上的双重差分可能制造虚假效应，这与研究者关注仿真失效条件高度相关，值得阅读原文以了解具体机制和检验方法。","inspiration":"借鉴其双重差分设计中的审查机制识别方法，即从审计自身评分中测量衰减贡献，用于稳健性检验。｜可迁移到信贷审批歧视研究，其中LLM法官对贷款申请人的评分可能受申请人特征影响，且评分量表有界。｜以LLM作为信贷审批员，处理为申请人种族或性别，结果变量为信用评分（1-10量表），对照真实信贷审批数据，检验交互效应是否由量表审查制造。"}},{"id":"2608.27463","version":2,"title":"Rating the Raters: Rasch Measurement Theory for LLM Evaluation","zh_title":"评估评分者：用于LLM评估的Rasch测量理论","abstract":"LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by judges or raters. Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured. Rasch measurement theory (RMT) is well-suited to this problem. RMT decomposes ordinal ratings into separable facets on a common scale. It further provides a battery of diagnostics that can identify miscalibrated measurements and rater biases. We present a case study of RMT applied to the LLM-as-rater paradigm using the Measuring Hate Speech corpus, whose construct was itself built under RMT. We fit a series of many-facet Rasch models to annotations from nine LLMs spanning families and capability levels. Our analyses show that LLMs systematically differ from human raters in severity, item-level calibration, question-order robustness, target-identity sensitivity, and rating scale use, all of which standard evaluation practice would largely obscure. Overall, we argue that RMT belongs in the toolkit for evaluating LLM-as-examinee, -judge, and -rater paradigms.","authors":["Pratik S. Sachdeva","Nathan Boudol"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-03","first_seen":"2026-08-31","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2608.27463","pdf_url":"https://arxiv.org/pdf/2608.27463","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM评估","测量理论","评分者偏差"],"reason":"评估LLM作为评分者与人类评分者的差异，涉及测量偏差与可靠性，可迁移到仿真评估。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":11,"question":"LLM作为评分者时，其测量特性与人类评分者有何系统性差异？","design":"本研究非仿真实验，而是测量评估研究。使用Measuring Hate Speech语料库，选取9个不同家族和能力水平的LLM对评论进行标注，拟合多面Rasch模型，分析LLM与人类评分者在严重度、项目校准、问题顺序稳健性、目标身份敏感性和评分量表使用上的差异。","baseline":"Measuring Hate Speech语料库中的人类标注，包括参考评论集（70条评论，每条超过200个标注）和扩展评论集（5990条评论，每条至少4个人类评分者）。","findings":"LLM评分者在严重度、项目级校准、问题顺序稳健性、目标身份敏感性和评分量表使用上与人类评分者存在系统性差异。标准评估实践会掩盖这些差异，而Rasch测量理论能有效揭示这些偏差。","reliability":"论文未明确讨论失效条件，但指出LLM与人类在测量特性上的差异表明LLM不能直接替代人类评分者，需谨慎校准。","relevance":"该研究直接评估LLM作为评分者与人类评分者的差异，涉及测量偏差与可靠性，对使用LLM进行人类仿真实验的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴Rasch测量理论分解评分变异，识别LLM评分者的系统性偏差，为仿真评估提供诊断工具。｜可迁移到信贷审批歧视研究，用LLM模拟信贷员对贷款申请人的评分，检验其是否与人类评分存在偏差。｜以LLM作为信贷员，对贷款申请进行风险评分，处理为申请人种族或性别，结果变量为评分和批准决策，与真实信贷审批数据对照，评估LLM仿真的可靠性。"}},{"id":"2609.01794","version":1,"title":"Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization","zh_title":"在语言模型避免过度泛化中区分统计抢占与固化","abstract":"How do learners avoid overgeneralizations such as Tom laughed me without explicit negative evidence? Constructionists have posited two proposals that describe indirect negative evidence against overgeneralizations: preemption (which privileges exposure to near-synonymous construction---e.g., she made him laugh) vs. entrenchment (all exposures to a verb's grammatical usages, including cases like He laughed). We disentangle these hypotheses by running controlled rearing experiments on LMs trained on child-caregiver conversations, where we systematically remove preemptive vs. non-preemptive evidence. We find that while LMs avoid overgeneralizations, they do not show preemption at a verb-specific level, instead showing weak but non-zero evidence of abstract preemption. Combined with results from analyzing the LMs' training dynamics, we find that LMs treat competing structures as indirect positive---as opposed to negative---evidence in the verb-specific condition. Insofar as preemption is the more plausible route to avoiding overgeneralizations in humans, our results point the need for there to be sensitivities to indirect negative evidence in neural network learners, and suggest new human experiments to test abstract preemption.","authors":["Yixuan Wang","Freda Shi","Kanishka Misra"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01794","pdf_url":"https://arxiv.org/pdf/2609.01794","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["语言习得","认知建模","人类对照"],"reason":"用LLM模拟儿童语言习得，与人类数据对照，但非社会行为仿真","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":13,"question":"语言模型能否通过间接负面证据（抢占或固化）避免动词过度泛化，以及抢占与固化哪个机制起作用？","design":"对语言模型进行“控制性养育”实验：在儿童-看护者对话语料上训练模型，系统移除抢占性证据或非抢占性证据，测量模型对过度泛化句子的接受度。","baseline":"人类儿童语言习得数据（如儿童过度泛化及最终回避的语料库记录）作为参照，但未直接使用个体人类被试数据。","findings":"语言模型能避免过度泛化，但未表现出动词特异性抢占效应，仅显示微弱的抽象抢占证据。模型将竞争结构视为间接正面证据而非负面证据。","reliability":"论文承认语言模型分析不能直接推论到人类，且抢占与固化证据存在多重共线性问题，人类实验难以获得学习者自身经验。","relevance":"该研究用LLM模拟人类语言习得，有真实人类数据对照，但属于认知科学而非社会行为仿真，对关注经济学实验仿真的研究者参考价值有限。","inspiration":"借鉴其控制性养育设计，通过系统移除特定类型训练数据来分离竞争机制。｜可迁移到经济金融中的学习与决策机制研究，如投资者如何从历史数据中学习风险规避。｜以LLM为被试，训练时移除特定类型的市场反馈（如正面或负面新闻），测量模型对资产价格的预期，并与真实投资者调查数据对照。"}},{"id":"2609.01918","version":1,"title":"Grounded, Compute-Efficient LLM Policy Agents for Energy-Poverty Equity in Physically-Constrained Peer-to-Peer Energy Markets","zh_title":"物理约束点对点能源市场中面向能源贫困公平的接地气、计算高效LLM策略智能体","abstract":"Energy poverty is nearly absent from NLP-for-social-good, and the little existing work is either static retrieval/QA or relies on carbon-intensive cloud LLMs, a self-defeating \"computational irony\" for a humanitarian setting. We present EqGrid, a closed-loop simulation in which a low-frequency, open-weight LLM policy agent sets price and carbon bounds and targeted subsidies over a community of empirically-grounded household personas, while high-frequency multi-agent RL traders clear a continuous double auction constrained by a physical distribution grid (IEEE-33-bus with Dynamic Operating Envelopes). Our contribution is threefold and directly addresses how to measure the social impact of AI: (i) grounded personas (region-matched socio-demographics) whose load curves are checked for shape and level realism against real smart-meter data; (ii) formal energy-poverty equity metrics (Energy Burden, Gini of EB, LIHC) showing the intervention reduces burden inequality without raising net grid cost; and (iii) a compute-efficiency frontier that measures how much equity performance survives compressing the policy agent from a 235B teacher down to a sub-1B model deployable on a laptop, in estimated energy/carbon per decision. A decoupled-safety design (the LLM sets bounds; a validate-and-project grid gate executes) yields zero grid-constraint violations versus 55 under direct LLM control. On energy-poverty equity, the LLM policy lowers the Gini of energy burden to 0.305 (from 0.351) and mean burden by 28% while cutting cost (outperforming a tuned rule baseline), and a 3B-active model retains 95% of the benefit at roughly 9x lower inference energy than the teacher, with even a 0.8B on-device model retaining 92% at roughly 24x lower energy. We will release code and configs.","authors":["Kunal Jadhav","Siddhesh More"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01918","pdf_url":"https://arxiv.org/pdf/2609.01918","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","能源市场","公平性评估"],"reason":"用LLM agent模拟能源市场政策干预，并与真实智能电表数据对照，涉及公平性…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":15,"question":"如何设计一个计算高效且物理安全的LLM政策智能体，在点对点能源市场中缓解能源贫困并提升公平性？","design":"使用开放权重LLM作为低频政策智能体，为基于EU-SILC真实数据生成的家庭角色设定价格/碳上限和定向补贴；高频多智能体强化学习交易者在IEEE-33总线物理电网约束下进行连续双向拍卖；测量能源负担、能源负担基尼系数、LIHC等公平性指标。","baseline":"对照真实智能电表数据验证负荷曲线的形状和水平真实性；家庭角色基于匈牙利EU-SILC边际数据。","findings":"LLM政策将能源负担基尼系数从0.351降至0.305，平均负担降低28%，同时降低成本并优于规则基线；解耦安全设计实现零电网违规，而直接LLM控制有55次违规。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟政策干预并与真实数据对照，评估公平性指标，属于人类仿真实验，但场景为能源市场而非典型经济金融实验，值得读原文了解其仿真框架和测量方法。","inspiration":"借鉴其分层仿真设计：LLM设定政策边界，底层RL智能体执行交易，并引入物理约束和安全层，同时用真实数据校准仿真人群。｜可迁移到能源市场政策评估、碳税设计或补贴分配等经济政策场景。｜以家庭能源消费数据为被试，用LLM模拟不同补贴政策（处理），测量能源负担和公平性（结果变量），并与真实智能电表数据和家庭调查数据对照。"}},{"id":"2609.02707","version":1,"title":"Door-in-the-Face Requests and Refusal Behaviour in Large Language Models","zh_title":"大语言模型中的登门槛请求与拒绝行为","abstract":"Does the door-in-the-face technique work on language models? In humans, a large request that is refused makes a smaller follow-up request more likely to be granted. We test this on nine production models from three providers: each model refuses a large request, then receives a smaller version of the same request, and we compare its compliance with asking directly. The answer depends on the model. On Anthropic's frontier models the technique works: Opus 5 answers the smaller request 65.8% of the time after refusing the larger one, against 29.3% when asked directly. On the frontier models of OpenAI and Google, and on Haiku 4.5, it backfires, lowering compliance by 15.5 to 23.0 points. A control locates the effect: a refused large request on an unrelated topic does less than the related one on all nine models, so the concession itself matters everywhere, while the reaction to having just refused something differs by model family. The technique does not transfer to refusals drawn from public benchmarks. What decides whether a retreat can work is what the request asks for: rewriting 265 refused requests for usable instructions into requests for explanations of the same topic removed the refusal in 263 cases. Human influence techniques port to language models one model family at a time.","authors":["Til Jordan"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02707","pdf_url":"https://arxiv.org/pdf/2609.02707","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B4"],"tags":["LLM行为实验","说服技巧","模型对比"],"reason":"测试LLM对人类说服技巧的反应，与人类行为对照，但非仿真被试","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":18,"question":"大语言模型是否会对“以退为进”（door-in-the-face）说服技巧产生与人类相似的反应，即先拒绝大请求后更可能同意小请求？","design":"该研究并非用LLM仿真人类被试，而是直接测试LLM本身的行为。在9个生产级模型上，每个试验包含四种条件：冷启动（直接提出小请求）、DITF（先提出一个设计为会被拒绝的大请求，再提出相同的小请求）、热身（先回答一个良性问题）、无关拒绝（先拒绝一个无关的大请求）。结果变量为模型对小请求的合规率，由跨模型家族的评判模型打分。","baseline":"无对照。论文未将LLM行为与真实人类被试在相同任务上的数据直接比较，而是引用社会心理学中人类DITF效应的经典结论作为背景。","findings":"DITF效应因模型家族而异：在Anthropic的前沿模型上，先拒绝大请求后小请求合规率显著提高（如Opus 5从29.3%升至65.8%）；而在OpenAI、Google的前沿模型及Haiku 4.5上，该技巧反而降低合规率15.5至23个百分点。进一步分析表明，让步本身在所有模型上都有作用，但模型家族决定了反应方向；且该效应不适用于公共基准中的拒绝，只有当小请求要求的是判断而非可操作内容时，DITF才可能有效。","reliability":"论文讨论了效应的边界条件：DITF不适用于模型在公共基准上产生的拒绝，且只有当小请求要求的是判断（而非可操作指令）时，让步才可能成功。此外，效应因模型家族而异，表明不存在统一的“类人”反应。","relevance":"该研究直接测试LLM对人类说服技巧的反应，虽未将LLM作为人类被试的替代品，但揭示了LLM行为与人类已知规律的异同，对评估LLM在仿真实验中的可靠性有参考价值，值得阅读原文以了解模型家族差异和边界条件。","inspiration":"借鉴其严谨的多条件对照设计（冷启动、DITF、热身、无关拒绝）来分离特定机制，并利用模型自身产生的拒绝作为实验材料，可迁移到经济金融中的谈判或议价场景，例如测试LLM在模拟消费者对价格让步的反应时是否表现出与人类相似的锚定效应。｜可设计一个实验：以LLM作为模拟消费者，先提出一个高价（大请求）被拒绝，再提出一个较低价（小请求），观察其购买意愿变化，并与真实消费者在相同议价情境下的行为数据（如实验经济学中的最后通牒博弈或讨价还价实验）进行对照，以评估LLM仿真在议价行为研究中的有效性。"}},{"id":"2609.02797","version":1,"title":"Dutch Books for Language Models","zh_title":"语言模型的荷兰赌：概率预测的连贯性评估","abstract":"People increasingly use language models to support life decisions. Many such decisions involve a probabilistic forecast: How likely is a major life event, a natural disaster, or an economic outcome? Users of language models may implicitly trust that these forecasts fall out of a coherent world model. In this paper, we evaluate the coherence of language model probabilistic forecasts through a procedure that builds on a theorem due to de Finetti. We elicit forecasts from language models across events generated from stock returns data. We then use linear programs to compute the largest Dutch-book profit - the profit an arbitrageur could guarantee by betting against model-generated probabilities - which we use as a measure of incoherence. Our procedure does not require outcome labels, so we can evaluate coherence even in settings where outcomes are not observed or have not yet resolved. We find substantial evidence of incoherence in language model forecasts. Such incoherence increases when there are richer logical relationships between events, and irrelevant contextual details can increase incoherence by an order of magnitude. We conclude by discussing how alternative training strategies may improve probabilistic coherence.","authors":["Isaiah Andrews","Suproteem Sarkar"],"categories":["econ.GN","cs.AI","cs.CL","cs.LG","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02797","pdf_url":"https://arxiv.org/pdf/2609.02797","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B3"],"tags":["LLM概率预测","连贯性评估","决策可靠性"],"reason":"评估LLM概率预测的连贯性，涉及决策可靠性，可迁移到仿真偏差研究。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":19,"question":"语言模型生成的概率预测在多大程度上满足概率论连贯性（即不存在荷兰赌套利）？","design":"本研究不是用LLM模拟人类被试，而是直接评估LLM本身作为概率预测者的连贯性。研究者从股票收益率数据构造事件（如未来收益落在某区间、事件的并集和补集），要求15个语言模型基于近期收益历史和新闻标题等上下文给出这些事件的概率预测，然后利用线性规划计算针对模型预测概率的最大荷兰赌利润，作为不连贯性的度量。","baseline":"无对照","findings":"语言模型的概率预测普遍存在显著不连贯性，且不连贯性在事件间逻辑关系更丰富时增加；无关的上下文细节可使不连贯性增加一个数量级。","reliability":"论文未讨论","relevance":"该研究直接评估LLM概率预测的可靠性，对使用LLM进行经济预测或仿真实验的研究者具有重要警示意义，值得阅读原文以了解不连贯性的具体表现和测量方法。","inspiration":"借鉴其利用荷兰赌定理和线性规划构造无结果标签的连贯性度量，可迁移到经济预测场景如政策公告的预期形成或资产定价实验，设计上可让LLM对同一组经济事件的不同逻辑组合给出概率预测，计算套利利润作为不连贯性指标，并与人类预测者的不连贯性进行对比。"}},{"id":"2609.01815","version":1,"title":"Induction and Inquiry via Probabilistic Reasoning over Language and Code","zh_title":"通过语言与代码的概率推理进行归纳与探究","abstract":"How humans grow and maintain abstract knowledge from the sparse, streaming noisy data of experience is a longstanding challenge in cognitive science. Any computational account must satisfy at least three desiderata: It must be (1) data-efficient and compute-efficient, (2) capture gradations of uncertainty to support intelligent inquiry and information gathering, and (3) be flexible enough to mentally represent the endless range of concepts people can learn and think about. Here we introduce a computational model that captures these three properties, by encoding symbolic knowledge as mental programs that combine natural language with source code, and sequentially inferring mental programs using LLM-guided Bayesian learning algorithms. Across a range of behavioral studies this model successfully reproduces quantitative signatures of human inductive learning and active inquiry, such as anchoring, garden-pathing, and other effects. In contrast, pure LLMs and classic Bayesian models either fail at the underlying task, or do not reproduce human behavior, or succeed only at exorbitant computational cost. These results suggest that one way humans continually grow their knowledge is by mentally representing many hypotheses spanning language-like and program-like representations, then revising those hypotheses to approximate Bayesian updates, while a bottom-up neural mechanism (an LLM) makes inference both tractable and learnable.","authors":["Wasu Top Piriyakulkij","Sam Acquaviva","Cassidy Langenfeld","Joshua Tenenbaum","Kevin Ellis"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01815","pdf_url":"https://arxiv.org/pdf/2609.01815","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["认知建模","贝叶斯学习","人类行为复现"],"reason":"用LLM引导贝叶斯学习复现人类归纳学习行为，并与行为实验对照，方法可迁移到人类…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":14,"question":"人类如何从稀疏、流式、嘈杂的经验数据中高效地归纳和更新抽象知识，并主动进行信息获取？","design":"该研究提出一个计算模型，将符号知识编码为结合自然语言和源代码的“心理程序”，并使用LLM引导的贝叶斯学习算法进行序列推断。模型在多个行为实验任务中模拟人类被试，包括概念学习、主动实验（Zendo游戏）和提问（网络购物助手），测量其归纳学习和主动探究的行为模式。","baseline":"对照的真实人类数据来自多个行为实验，包括Bramley等人（2018）的Zendo游戏数据（7轮实验后对8个保留测试构造的预测），以及其他归纳学习任务中的人类行为数据（如锚定、花园路径效应等）。","findings":"该模型成功复现了人类归纳学习和主动探究的定量特征，如锚定、花园路径等效应，并且在多个任务上比纯LLM和经典贝叶斯模型更符合人类行为，同时计算效率更高。模型还展示了通过调节推理时间预算来捕捉有限理性行为的能力。","reliability":"论文未明确讨论失效条件，但指出纯LLM和经典贝叶斯模型在任务上要么失败，要么不符合人类行为，要么计算成本过高，暗示模型依赖于LLM的生成能力和贝叶斯更新的结合。","relevance":"该研究直接使用LLM引导的贝叶斯学习来模拟人类归纳学习和主动探究行为，并与真实人类实验数据对照，属于用LLM进行人类仿真实验的研究，且涉及认知科学中的行为复现，对关注LLM仿真可靠性和偏差的研究者有参考价值。","inspiration":"该方法将LLM作为假设生成器，结合贝叶斯更新进行序列推理，可借鉴用于模拟经济决策中的信念更新和主动信息获取。｜可迁移到政策公告的预期形成、消费者跨期选择中的学习过程、或投资者在不确定环境下的信息搜寻行为。｜设计一个实验：用LLM引导的贝叶斯模型作为被试，模拟投资者在收到一系列经济新闻后的资产价格预期更新，处理是不同信息呈现方式（如顺序、频率），结果变量是预测准确性和不确定性，并与真实人类实验数据（如实验室资产市场实验）对照。"}},{"id":"2609.02821","version":1,"title":"AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application","zh_title":"AI情境测量用于恢复个体与群体效应：基于调查测量的验证及职业应用","abstract":"Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We propose AICOME, AI COntextual MEasurement, a framework for evaluating whether AI-derived respondent-level measures can recover individual and group-level effects in contextual models. The key idea is that an AI measure constructed at the respondent level can be used to derive its group-level aggregate and its individual deviation, allowing researchers to estimate both between-group and within-group associations rather than treating AI measurement as response prediction alone. We validate the framework using the 2022 China Family Panel Studies (CFPS), where occupations provide the empirical grouping structure and several job-related survey variables provide validation benchmarks. For computer use, foreign-language use, weekly hours, and management responsibilities, we compare survey measures with AI-derived measures in response-level, model-level, contextual, and boundary-condition validations. The results show that AI contextual measurement can recover much of the contextual-model information contained in observed survey variables when rich respondent and job characteristics are available. Weekly hours provides the strongest validation case, with AI-derived measures reproducing the large negative between- and within-occupation associations with satisfaction observed in CFPS. The framework also identifies clear boundary conditions: performance deteriorates when information is restricted to occupation and basic demographics, and recovery is weaker when several related concepts are treated as simultaneously unobserved. The findings suggest that AICOME is most useful for recovering a limited number of theoretically important constructs from rich existing datasets.","authors":["Wenxin Jiang","Xuyang Wang","Yuxiao Wu"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02821","pdf_url":"https://arxiv.org/pdf/2609.02821","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["AI测量","验证框架","社会调查"],"reason":"用AI从调查数据构建个体测量并验证，与人类数据对照，但非直接仿真被试","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":20,"question":"AI 派生的个体层面测量能否在情境模型中恢复个体和群体层面的效应，从而替代传统调查测量进行情境分析？","design":"本研究不是直接仿真人类被试，而是提出并验证 AICOME 框架：利用大语言模型基于受访者特征（如职业、人口统计、工作信息）生成个体层面的测量（如计算机使用、外语使用、每周工时、管理责任），然后将该测量分解为群体均值（职业均值）和个体偏差，用于估计组间和组内关联，并与 CFPS 调查中的真实测量进行多层面验证。","baseline":"使用 2022 年中国家庭追踪调查（CFPS）中的真实调查变量作为基准，包括计算机使用、外语使用、每周工时、管理责任以及工作满意度等，以职业为分组结构。","findings":"当提供丰富的受访者和工作特征时，AI 情境测量能够恢复调查变量所包含的大部分情境模型信息，其中每周工时的验证效果最强，AI 派生测量重现了 CFPS 中观察到的工时与满意度之间显著的负向组间和组内关联。然而，当信息仅限于职业和基本人口统计特征时，性能下降；当多个相关概念同时被视为未观测时，恢复效果较弱。","reliability":"论文识别了明确的边界条件：当输入信息受限（仅职业和基本人口统计）时性能恶化；当多个相关概念同时缺失时恢复较弱。作者指出 AICOME 最适用于从丰富数据集中恢复少数理论重要构念，而非替代直接调查测量或大量联合缺失的项目。","relevance":"该研究虽非直接仿真人类被试，但提供了 AI 派生测量与真实人类调查数据对照的严格验证框架，对评估 LLM 在社会科学测量中的可靠性具有参考价值，值得阅读原文以了解其多层面验证方法和边界条件。","inspiration":"借鉴其将 AI 派生测量分解为组均值和个体偏差以同时估计组间和组内效应的做法，并采用多层面验证（响应级、模型级、情境级、边界条件）来评估测量质量｜可迁移到劳动经济学中职业特征对工资或工作满意度的影响研究，或组织经济学中企业文化对员工绩效的组间与组内效应分析｜设计：以职业为分组，使用 LLM 基于员工简历或工作描述生成职业特征（如自主性、技能要求）的个体测量，然后分解为职业均值和个体偏差，估计其对工资或离职行为的组间和组内效应，并与真实调查数据（如 CFPS 或美国 CPS）中的对应变量进行对照验证。"}},{"id":"2609.01627","version":1,"title":"The Utility of LLMs in Recommender Systems Explanation Evaluation","zh_title":"大语言模型在推荐系统解释评估中的效用","abstract":"Explanations play a crucial role in creating trustworthy recommender systems (RS), yet choosing a good explanation method presents challenges. Many explanation methods exist, but little guidance exists on which is best for which setting. Existing explanation generation methods often produce abstract outputs that require further formatting to become user-friendly, with a seemingly endless pool of options. Running user-based evaluations of all possible options is usually unfeasible, while automated evaluation metrics often either assess only the explainer's abstract output or require comparison with a ground truth, which is generally unavailable. Recent studies have shown that large language models (LLMs) can serve as ``judges'' for explanation evaluation, but their reliability has not yet been thoroughly explored. This paper studies the utility of LLMs in selecting an effective explanation method for a given application. We first explore their ability to generate explanation prototypes given varying information about the RS and the user. Specifically, we generate 18 distinct explanation prototypes, which are subsequently evaluated by 14 LLMs of varying sizes across two temperature settings. We compare these against human ratings derived from a user study. Our results show that while LLMs exhibit human-like rating patterns and achieve moderate rank correlation with human raters, their absolute rating agreement is low and varies substantially by model size and evaluation construct. We derive four practical recommendations: keep explanation-generation prompts concise, prefer larger models for evaluation, pre-test evaluation constructs, and audit explanations for factual accuracy, as neither humans nor LLMs reliably detect non-factual content.","authors":["Kathrin Wardatzky","Oana Inel","Luca Rossetto","Abraham Bernstein"],"categories":["cs.IR","cs.AI"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01627","pdf_url":"https://arxiv.org/pdf/2609.01627","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM评估","推荐系统解释","人类对照"],"reason":"用LLM评估推荐解释，并与人类评分对照，属于仿真人类评估行为，但非典型被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":12,"question":"LLM能否替代人类评估推荐系统解释的质量，从而在原型阶段选择有效的解释方法？","design":"使用知识图谱推荐系统PGPR生成推荐和路径解释，用LLM基于不同信息方面（用户专业知识、用户历史、推荐信息、解释目标）生成18种解释原型；随后用14个不同规模的LLM在两种温度设置下对解释在六个质量维度上评分，并与人类评分比较。","baseline":"通过用户研究收集的人类评分，用于与LLM评分进行对比。","findings":"LLM表现出与人类相似的评分模式，与人类评分者达到中等秩相关；但绝对评分一致性低，且随模型规模和评估构念变化很大。","reliability":"论文承认LLM和人类都不可靠地检测非事实内容，因此需要审计解释的事实准确性；绝对评分一致性低，模型规模和评估构念影响可靠性。","relevance":"该研究用LLM模拟人类对推荐解释的评估，并与真实人类评分对照，属于人类仿真评估行为，但非典型被试仿真；对关注LLM评估可靠性和偏差的研究者有一定参考价值。","inspiration":"借鉴其系统生成多种解释原型并用多个LLM评分与人类评分对照的方法，可迁移到经济金融领域的文本解释评估（如信贷决策解释、投资建议解释）；设计上可用LLM生成不同风格的信贷拒绝解释，让多个LLM和人类被试对解释的公平性、可理解性等维度评分，以真实人类评分为基准检验LLM评估的可靠性。"}},{"id":"2609.02677","version":1,"title":"Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization","zh_title":"基于强化学习的投资组合优化中ESG偏好的获取","abstract":"Modern portfolio management increasingly demands a balance between traditional risk-adjusted returns and strict Environmental, Social, and Governance (ESG) mandates. Current Reinforcement Learning (RL) approaches typically optimize for a single ESG provider, neglecting the significant divergence in rating methodologies across the industry and the unintuitive nature of manually weighting conflicting objectives. This paper addresses these limitations by formulating ESG-aware portfolio optimization as a Multi-Objective Reinforcement Learning (MORL) problem that simultaneously incorporates ratings from three distinct ESG agencies. To bridge the gap between high-dimensional algorithmic trade-offs and human decision-making, we integrate a Preference Elicitation framework using Gaussian Processes. This system enables practitioners to infer their latent utility functions through intuitive pairwise comparisons of candidate portfolios based on their Sharpe ratios and aggregate ESG scores. We systematically evaluate our framework by employing Large Language Model (LLM) personas to simulate Portfolio Managers operating under varied regional contexts. Empirical results using historical market data reveal that regional backgrounds fundamentally shift the derived preference weights. For instance, European-based personas tend to prioritize ESG alignment over financial returns, while Texas-based personas favor risk-adjusted performance. This work offers a highly adaptable framework that successfully aligns multi-objective algorithmic trading with diverse, real-world human sustainability preferences.","authors":["Giovanni Dispoto","Marcello Restelli","Carmine Ventre"],"categories":["q-fin.PM","cs.CE","cs.LG"],"primary_category":"q-fin.PM","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02677","pdf_url":"https://arxiv.org/pdf/2609.02677","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A1","A3","B2"],"tags":["LLM仿真","投资组合优化","偏好获取"],"reason":"用LLM persona模拟投资经理偏好，涉及金融决策，但无真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":17,"question":"如何将多个ESG评级机构的冲突信号纳入强化学习投资组合优化，并通过偏好诱导使算法与人类投资经理的可持续性偏好对齐？","design":"使用LLM personas模拟不同地区（欧洲、德克萨斯、亚洲、美国）的投资经理，通过成对比较候选投资组合（基于夏普比率和ESG得分）来诱导其潜在效用函数，并测量诱导出的偏好权重。","baseline":"无对照","findings":"区域背景显著改变诱导出的偏好权重，欧洲persona更重视ESG对齐，德克萨斯persona更重视风险调整后收益。该框架能成功将多目标算法交易与多样化的现实人类可持续性偏好对齐。","reliability":"论文未讨论","relevance":"该研究用LLM personas模拟投资经理的ESG偏好，属于金融决策仿真，但缺乏真实人类数据对照，适合作为批判性案例或方法参考。","inspiration":"借鉴其用LLM personas模拟不同区域投资者并施加区域背景处理来诱导偏好的方法。｜可迁移到ESG投资偏好、社会责任投资或绿色金融产品选择等场景。｜用LLM personas模拟不同文化背景的投资者，处理为区域或制度环境，结果变量为对ESG与收益的权衡，对照真实调查或实验数据（如全球投资者调查、实验室投资实验）。"}},{"id":"2607.10628","version":2,"title":"Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation","zh_title":"Anamnesis：大规模背景条件调查仿真的开源平台","abstract":"We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models. Open-source and designed for non-technical users/researchers, Anamnesis enables the prototyping and stress-testing of survey instruments on virtual populations rather than real human subjects. The platform operationalizes the recently introduced Anthology and Alterity frameworks, which use structured narrative backstories to condition model responses, within a unified web interface. It supports open-ended generation, probabilistic demographic resampling, and multimodal (image and audio) surveys. We evaluate the system through two case studies: (1) replicating segments of Pew Research Center's American Trends Panel (ATP) on political typology and biomedical issues and (2) emulating human preference in the New Yorker Caption Contest. In both cases, Anamnesis produces opinion distributions that more closely match real-world survey data than standard persona-prompting baselines, offering a transparent, reproducible, and open-source alternative to proprietary simulation services.","authors":["Song-Ze Yu","Joseph Suh","Serina Chang","David M. Chan"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-02","first_seen":"2026-07-12","revised_at":"2026-09-02","abs_url":"https://arxiv.org/abs/2607.10628","pdf_url":"https://arxiv.org/pdf/2607.10628","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B4"],"tags":["LLM仿真","调查模拟","人类数据对照"],"reason":"平台用LLM仿真调查，与真实数据对照，复现舆论分布，评估可靠性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":1,"question":"如何构建一个开源、可交互的平台，利用大语言模型和结构化叙事背景来模拟多样化人群的调查回答，并验证其与真实人类调查数据的匹配度？","design":"Anamnesis平台使用大语言模型（如Gemini 2.5 Flash）扮演虚拟受访者，通过结构化叙事背景（backstories）而非简单人口统计列表来条件化模型响应；支持概率性人口重采样、开放式生成和多模态（图像、音频）调查；在案例研究中，模拟Pew Research Center的American Trends Panel三个波次（政治类型学、生物医学、AI与人类增强）的多选题调查，以及New Yorker Caption Contest的多模态偏好选择。","baseline":"对照的真实人类数据包括：Pew Research Center American Trends Panel（ATP）三个波次的真实调查回答分布，以及New Yorker Caption Contest中基于大规模众包投票的真实人类偏好标签。","findings":"Anamnesis平台在三个ATP波次上产生的意见分布比标准人物提示（persona-prompting）基线更接近真实调查数据，在Wasserstein距离和Frobenius范数上均表现更优；在多模态New Yorker Caption Contest中，平台模拟的人类偏好与真实人类集体判断存在可测量的相关性。","reliability":"论文未明确讨论仿真失效的条件，但指出标准人物提示方法会产生刻板印象且缺乏心理深度，暗示仅依赖人口统计列表的仿真可能不可靠；同时，平台依赖LLM推理提供商，可能引入模型偏差，且案例研究仅覆盖有限主题和模态，泛化性有待验证。","relevance":"该论文直接针对用LLM进行人类仿真实验的研究，提供了开源平台和与真实调查数据对照的验证，对关注经济学实验和政策评估场景的研究者具有重要参考价值，值得阅读原文以了解平台细节和评估方法。","inspiration":"借鉴其使用结构化叙事背景而非简单人口统计列表来条件化LLM响应的方法，可提高虚拟被试的心理真实性和异质性，并采用与真实调查数据匹配的评估指标（如Wasserstein距离）来量化仿真质量｜可迁移到政策评估中的公众意见模拟，例如模拟不同社会经济背景的个体对税收改革、福利政策或公共卫生措施的态度分布，以预测试验或调查结果｜设计一个研究：使用Anamnesis平台生成具有多样化背景故事的虚拟被试，施加不同政策信息框架（如强调公平 vs. 效率）作为处理，测量其对政策支持度的选择，并与真实世界调查数据（如General Social Survey或特定政策民意调查）进行分布匹配对照，以评估仿真预测的准确性。"}},{"id":"2609.00222","version":1,"title":"LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts","zh_title":"LLM作为人口群体：社会人口学提示对谁有益，对谁有害","abstract":"Large language models (LLMs) are increasingly used as judges for subjective tasks, where annotators disagree and the relevant question is not only how accurate a judge is, but whose judgments it reproduces. Sociodemographic prompting conditions the judge on an annotator's demographic profile to align its judgments with the corresponding group's. We test whether this alignment emerges distributionally, comparing the predicted label distributions of 23 open-weight LLMs on three subjective tasks against those of real annotator groups, under three conditions: no demographic information, single-attribute profiles, and intersectional profiles over gender, age, race, and education. Three findings emerge. First, a judge prompted with no demographics is not perspective-neutral: models best reproduce the judgments of White, college-educated annotators. Second, demographic conditioning is asymmetric: it moves the judge toward majority groups and away from minority groups, most strongly on offensiveness, where intersectional profiles amplify the harm. Third, by comparing base and instruct models we identify instruction-tuning as a possible source of the asymmetry. Demographic conditioning should therefore be used with caution to estimate group judgments: conditioning moves predictions away from the reference distributions of the minority groups the method is often invoked to serve.","authors":["Daniela Occhipinti","Andrea Piergentili","Marco Guerini"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00222","pdf_url":"https://arxiv.org/pdf/2609.00222","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人口学提示","算法偏差"],"reason":"用LLM模拟不同人口群体判断，并与真实标注者分布对照，评估偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":3,"question":"在主观判断任务中，用人口统计学提示（sociodemographic prompting）让LLM模拟特定人群的判断，是否能使模型的标签分布向该人群的真实判断分布靠拢？","design":"使用23个开源权重LLM作为判断者，在三个主观任务（礼貌性、亲密性、冒犯性）上，比较三种条件：无人口统计学信息、单一属性画像（性别、年龄、种族、教育）、交叉属性画像。通过比较模型预测的标签分布与真实标注者群体的标签分布（来自DeMo数据集）来评估对齐效果。","baseline":"DeMo数据集，包含多个语料库中标注者的自报人口统计学信息（性别、年龄、种族、教育）及其对文本的评分，可构建每个文本在特定人口群体上的真实标签分布。","findings":"无人口统计学提示的LLM判断者并非视角中立，其判断最接近白人、大学学历标注者。人口统计学条件作用不对称：使判断者向多数群体靠拢、远离少数群体，尤其在冒犯性任务上，交叉画像放大了这种伤害；通过比较基础模型和指令微调模型，发现指令微调可能是这种不对称的来源。","reliability":"论文指出人口统计学提示应谨慎用于估计群体判断，因为条件作用使预测远离少数群体的参考分布；此外，效应因任务和模型而异，且指令微调模型表现出更强的不对称性。","relevance":"该研究直接评估了用LLM模拟不同人口群体判断的可靠性，并与真实标注者分布对照，揭示了仿真中的系统性偏差，对关注LLM仿真人类行为的研究者具有重要参考价值。","inspiration":"借鉴其分布级评估框架，将模型输出分布与真实群体分布比较，而非仅比较均值或多数标签，并采用交叉人口属性来检验交互效应。｜可迁移到信贷审批中的群体差异研究，如模拟不同性别、种族、教育背景的贷款审批人对同一申请的风险判断。｜以LLM作为虚拟审批人，施加不同人口统计学提示（如性别×种族），测量其对贷款申请的批准概率分布，并与真实信贷审批数据（如某银行历史审批记录中不同审批人群体）的分布进行对比，检验仿真偏差。"}},{"id":"2609.01591","version":1,"title":"StudentSim: Training LLM-based Student Simulators","zh_title":"StudentSim：训练基于LLM的学生模拟器","abstract":"AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but evidence about which guidance works for which student is sparse, slow, and costly to collect from real learners. Student simulators can provide this signal as a proxy, yet existing approaches are limited: state-tracking models fit student behavior but struggle to process explanations or corrections, while LLM role-play follows guidance fluently but does not reliably match the competence of the student being imitated. We present StudentSim, a training framework that turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization. The resulting simulators both mirror a student's own responses and update them under tutor guidance. We also introduce StudentSimEval, a standardized protocol covering 60 students across chess, second-language English writing, and mathematics, using public learner datasets with de-identified records shared for research. StudentSimEval measures behavioral fidelity (F), or how well a simulator matches a student's responses, and guidance responsiveness (R), or how readily it updates under tutor guidance, with all methods fit and evaluated on the same records. Across all three domains, StudentSim outperforms GPT-5.4 on both metrics. In chess, StudentSim reaches F=0.51 and R=0.91, compared with 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. As a proof of concept, using StudentSim as a reward model for tutor reinforcement learning produces a chess tutor that expert humans rate as more accurate, better-guided, and more personalized than a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward. Code is available at https://github.com/microsoft/StudentSim.","authors":["Ke Yang","Chenglong Wang","Michel Galley","Chandan Singh","Jeevana Priya Inala","ChengXiang Zhai","Jianfeng Gao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01591","pdf_url":"https://arxiv.org/pdf/2609.01591","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","学生模拟","行为保真度"],"reason":"用LLM模拟学生行为并与真实学生数据对照，评估行为保真度和指导响应性，属于人类…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":8,"question":"如何训练既能忠实反映个体学生能力又能响应导师指导的学生模拟器？","design":"提出StudentSim框架，先在所有学生记录上预训练基础模拟器，再对每个学生进行个性化微调；在棋类、二语写作和数学三个领域共60名学生上，用行为保真度（F）和指导响应性（R）两个指标评估模拟器，并与GPT-5.4和Maia2等基线比较。","baseline":"使用公开学习者数据集（chess、L2、math）中的真实学生记录，每个学生有训练集和留出集，所有方法在相同记录上拟合和评估。","findings":"StudentSim在所有三个领域的行为保真度和指导响应性上均优于GPT-5.4；在棋类中，StudentSim的F=0.51、R=0.91，而GPT-5.4为0.23和0.72，Maia2为0.45和0.27。将StudentSim作为奖励模型进行强化学习，训练出的棋类导师在准确性、指导性和个性化方面均优于无RL基线和用GPT-5.4模拟器奖励训练的导师。","reliability":"论文未讨论","relevance":"该研究用LLM模拟学生行为并与真实学生数据对照，评估行为保真度和指导响应性，属于人类仿真实验，且包含经济学实验和政策评估场景的潜在应用，值得阅读原文以了解其训练框架和评估协议。","inspiration":"借鉴其两阶段训练框架（先池化预训练再个体微调）和双指标评估（行为保真度与响应性），可迁移到经济金融中的个体决策模拟，如消费者选择或投资者行为。｜例如，在资产定价实验中模拟异质投资者对信息的反应，或在信贷审批中模拟不同风险偏好的申请人。｜用LLM模拟投资者，施加不同政策公告作为处理，测量其交易行为和风险偏好变化，并与真实投资者交易数据对照，评估模拟器的保真度和响应性。"}},{"id":"2609.01038","version":1,"title":"Data-Driven Persona-Conditioned Agents for A/B Test Simulation","zh_title":"基于数据驱动人物画像的智能体用于A/B测试模拟","abstract":"A/B testing is the gold standard for evaluating product changes, but each experiment requires real user traffic, engineering effort, and weeks of measurement. We propose a simulation framework that predicts A/B test outcomes using LLM-powered agents conditioned on data-driven personas grounded in real user behavioral signals. Unlike prior work that relies on synthetic or rule-based personas, our agents are constructed from anonymized behavioral data-activity patterns, engagement signals, and inferred demographics-enabling more faithful population modeling. We frame A/B test simulation as a structured question task and systematically study (i) question design formats, (ii) the impact of persona data source and domain alignment, (iii) the trade-off between per-persona behavioral depth and population diversity, and (iv) efficient population subsampling. On a benchmark of 40 A/B tests spanning two metric types, our best configuration achieves 0.75-0.90 directional accuracy depending on the test metric, demonstrating that data-driven personas are a viable path toward fast, low-cost experiment pre-screening.","authors":["Ziyad Benomar","Weronika {\\L}ajewska","Leonardo Perelli","Saab Mansour"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01038","pdf_url":"https://arxiv.org/pdf/2609.01038","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","A/B测试","人物画像"],"reason":"用LLM代理模拟A/B测试用户行为，基于真实行为数据构建persona，并与真…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":5,"question":"如何用基于真实行为数据构建的数据驱动型人格（persona）条件化LLM智能体，来预测A/B测试的方向性结果，并系统研究问题格式、人格数据来源、行为深度与多样性权衡及子采样效率。","design":"使用Claude Sonnet 4.5作为LLM，基于匿名行为数据（活动模式、参与信号、推断人口统计）生成结构化人格，将A/B测试模拟为结构化问答任务，让智能体在控制与处理变体间做出选择，预测点击率（CTR）和订阅量两个指标的方向性结果。","baseline":"40个历史A/B测试的真实结果，由置信区间推导出方向标签（正、负、可忽略），作为评估模拟准确性的基准。","findings":"最佳配置在CTR和订阅量上的方向准确率分别达到0.75和0.90，表明数据驱动人格可用于低成本实验预筛选。领域对齐的人格数据源对模拟质量至关重要，公共行为数据可媲美平台特定人格；子采样可使成本降低2倍而无质量损失。","reliability":"论文承认基准测试来自精选实验样本，不代表任何平台的全部用户或运营A/B测试基础设施；人格池偏向高参与度用户可能不具代表性；未讨论LLM模拟在更复杂决策或长期效应上的失效条件。","relevance":"该研究直接命中研究者对LLM仿真人类行为、真实数据对照、经济学实验场景的关注，提供了人格构建、问题格式和采样策略的系统性实证，值得精读以借鉴其方法并批判其局限。","inspiration":"借鉴其用真实行为数据构建人格并系统比较深度与多样性权衡、子采样效率的做法，可迁移到消费者金融决策或政策评估场景。｜可应用于信贷审批歧视研究：用银行交易数据构建不同信用评分段的人格，模拟贷款申请决策。｜设计：以真实银行客户数据构建人格，处理为不同利率或贷款条款，结果变量为是否接受贷款，用历史信贷数据中的真实接受率作为对照基准。"}},{"id":"2609.01257","version":1,"title":"Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations","zh_title":"衡量长时程人类活动模拟的行为保真度","abstract":"As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. While prior work has examined behavioral fidelity in survey responses and dialogue, longer-horizon real-world activity remains largely unexplored. We introduce a framework for evaluating behavioral fidelity in long-horizon activity simulations across temporal granularities and levels of analysis. As a case study, we collect a 43-hour multi-camera dataset of in-the-wild office activity and compare trace-derived conditioning mechanisms: persona descriptors, few-shot exemplars, and statistical transition and time-of-day priors. We find that behavioral fidelity is not uniform across metrics: statistical priors bring activity and sequence distributions closest to real behavior, yet over-fragment routines and suppress within-person variability. These findings motivate a more holistic evaluation that spans multiple metrics, temporal granularities, and levels of analysis.","authors":["Yi Fei Cheng","Fan Yang","Iremsu Bas","Koichiro Niinuma","Narishige Abe","David Lindlbauer"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01257","pdf_url":"https://arxiv.org/pdf/2609.01257","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","行为保真度","人类活动模拟"],"reason":"直接评估LLM模拟人类长期活动的行为保真度，并与真实人类数据对照，属于核心仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":6,"question":"如何系统评估大语言模型在长时间跨度人类活动模拟中的行为保真度？","design":"使用六种基于LLM的仿真方法（无人物描述控制、作者撰写的人物描述、轨迹推断的人物描述、少样本全天行为示例、统计转移先验、时间先验），在模拟办公室环境中生成五名智能体各八小时的活动轨迹，并与真实办公室活动数据集比较，测量活动分布、序列分布、时间分布、个体内变异性等指标。","baseline":"收集了一个43小时的多摄像头办公室活动数据集，包含55人的真实活动轨迹，其中5人有多天、每天数小时的纵向观察数据，作为对照基准。","findings":"统计先验使活动分布和序列分布最接近真实行为，但会导致日程过度碎片化并抑制个体内变异性。基于人物描述的方法之间差异较小，且群体层面的一致性可能掩盖个体层面的错误。","reliability":"论文指出行为保真度在不同指标、时间粒度和分析层面上并不一致，统计先验虽改善分布对齐但牺牲了日程连贯性和个体变异性，因此需要多维度评估。","relevance":"该研究直接评估LLM模拟人类长期活动的行为保真度，并与真实人类数据对照，属于核心仿真研究，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其多维度评估框架和条件机制对比设计，可迁移到经济金融中的消费者日常消费行为模拟或投资者交易行为模拟。例如，用LLM模拟投资者在交易日内的交易决策，处理为不同条件机制（人物描述、少样本示例、统计先验），结果变量为交易频率、交易时间分布和持仓变化，对照真实交易数据（如某券商脱敏交易记录）评估保真度。"}},{"id":"2609.01275","version":1,"title":"The Constitutional Coverage Trilemma in AI Governance","zh_title":"AI治理中的宪法覆盖三难困境","abstract":"Frontier AI systems function as \\emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human demand. Combining a paraphrase-controlled audit of the as-shipped default constitutions of $23$ frontier LLM archetypes with a pairwise-tradeoff study of $1{,}649$ US participants on the same instrument, we report three facts. \\emph{Demand is broad}: it spans all five values, with the largest constituency under one-third. \\emph{Supply is narrow and drifting}: the $23$-archetype hull occupies ${\\sim}2\\%$ of the demand hull under conservative noise-matched estimation ($0.10\\%$ at full audit precision), no archetype puts helpfulness or autonomy first ($37\\%$ of users are constitutionally homeless), and across six model families autonomy decreases in $5/6$, equity increases in $5/6$, and safety increases in $4/6$, with monotone within-family version trends (order-permutation $p = 0.013$) and the autonomy decline concentrated in scenarios where safety is not at stake. The drift's importance is directional: \\emph{away} from a value already undercovered, mechanically worsening the welfare floor for the least-served users. \\emph{The fix is sparse}: a $2$-vertex menu $\\{e_{\\mathrm{HON}}, e_{\\mathrm{AUT}}\\}$ beats the full $23$-archetype frontier by $47\\%$ on mean regret (CI $[43\\%, 52\\%]$); three vertex additions cut mean/worst-group regret by up to $81\\%$/$64\\%$. We formalize these findings as a budgeted-pluralism trilemma, show the binding regime is empirically realized, and verify the conclusions are robust to distance-based welfare and to degraded routing. The instrument and audit harness are described in full in the appendices.","authors":["Natalija Mitic","Soona Sedahmed A. O.","Mamadou Selly Ly","Moustapha Cisse"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01275","pdf_url":"https://arxiv.org/pdf/2609.01275","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","价值观对齐","人类对照"],"reason":"用LLM审计宪法价值观并与1649名人类对照，评估供需匹配与偏差，直接仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":7,"question":"前沿LLM的宪法价值观供给是否覆盖了人类用户的宪法价值观需求？","design":"本研究不是用LLM仿真人类，而是将LLM本身作为被审计对象：对23个前沿LLM原型进行受控改写审计，测量其在安全、帮助性、诚实、自主、公平五维价值上的隐含排序；同时用同一套成对权衡问卷测量1649名美国参与者的价值偏好，作为需求侧基准。","baseline":"1649名美国参与者在同一五维价值成对权衡问卷上的偏好分布。","findings":"人类需求广泛，五种价值都有显著支持者，最大群体占比不足三分之一；LLM供给狭窄且漂移，23个原型仅覆盖需求凸包的约2%，没有原型将帮助性或自主性置于首位，37%的用户在宪法意义上无家可归，且跨版本漂移方向是自主性下降、公平和安全上升，远离已覆盖不足的价值。","reliability":"论文指出线性福利模型是对供给方最有利的假设，实际福利可能更低；审计精度受改写控制和模型间差异解决阈值影响；未讨论LLM审计结果与真实部署行为之间的差距。","relevance":"该研究直接测量LLM的价值观分布并与人类对照，属于用LLM进行人类仿真实验的批判性工作，揭示了仿真在价值观匹配上的系统性偏差，值得精读。","inspiration":"借鉴其将LLM作为制度性主体进行审计并与人类偏好对照的方法，可迁移到经济金融中的算法决策场景，如信贷审批、保险定价或投资建议中的公平与效率权衡。｜设计一个实验：用多个LLM扮演信贷审批员，施加不同价值取向的提示（如强调公平或效率），测量其审批决策中的种族或性别差异，并与真实银行信贷数据或人类审批员的决策分布进行对照，评估LLM仿真的偏差。"}},{"id":"2609.00009","version":1,"title":"Toward a social psychology of AI: language-model agents reproduce human-like minimal-group bias","zh_title":"迈向AI社会心理学：语言模型智能体再现类人的最小群体偏差","abstract":"Language-model agents now interact in groups, but evaluations that probe memorised stereotype content or use models to simulate people leave this social behaviour unmeasured. We adapt the minimal-group paradigm---social psychology's classic test of intergroup bias---into a controlled probe: an agent distributes points among anonymous peers bearing only an arbitrary group label. Across four reasoning models, mere categorisation into meaningless groups elicited in-group favouritism that vanished under a group-blind control and was concentrated in the numerical minority: minority deciders over-allocated to their own group relative to their numbers, majority deciders allocated close to proportionally, and the asymmetry closed at equal group sizes. Disabling reasoning in one model did not remove the disposition---if anything it grew---but nearly erased the minority-majority asymmetry, implicating deliberation in where bias concentrates rather than whether it appears. These open-weight reasoning models reproduce the behavioural signature of human intergroup discrimination, independent of stereotype content, and social psychology's theories and methods offer a paradigm for measuring and governing AI's social behaviour.","authors":["Messi H. J. Lee"],"categories":["physics.soc-ph","cs.CY"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00009","pdf_url":"https://arxiv.org/pdf/2609.00009","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM仿真","社会心理学","群体偏差"],"reason":"用LLM复现人类最小群体偏差，并与经典社会心理学实验对照，直接仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":2,"question":"语言模型智能体在最小群体范式中是否复现人类的内群体偏差，且该偏差如何随群体相对规模变化？","design":"使用四个开源推理模型（如 DeepSeek-R1、Qwen 等）作为被试，将其置于最小群体范式中：智能体被随机分配到无意义标签（如“组A”或“组B”）的群体，并需要在匿名同伴之间分配点数。处理变量包括群体标签的有无（组盲对照）和群体相对规模（少数/多数/均等）。结果变量是分配给内群体成员的点数比例。","baseline":"人类经典最小群体实验（Tajfel 等）及元分析结果，特别是关于少数群体成员比多数群体成员表现出更强内群体偏爱的发现。","findings":"四个推理模型在无意义群体分类下均表现出内群体偏爱，且该偏爱在组盲对照中消失；少数群体成员过度分配给内群体，多数群体成员分配接近比例，群体规模相等时不对称消失。禁用推理后，内群体偏爱未消失甚至增强，但少数-多数不对称几乎消失，表明推理影响偏差的分布而非存在。","reliability":"论文未明确讨论失效条件，但指出仅测试了开源推理模型，未涵盖闭源模型或指令微调模型；且实验为一次性分配任务，未涉及长期互动或真实后果。","relevance":"该研究直接以LLM为被试复现社会心理学经典实验，并与人类基准对照，属于人类仿真研究，且揭示了群体结构对偏差的影响，值得精读以了解仿真在群体行为中的有效性。","inspiration":"借鉴其最小群体范式的对照设计（组盲条件）和群体规模操纵，以分离纯粹分类效应｜可迁移到信贷审批中的群体歧视研究，例如测试AI信贷员是否对少数群体申请人有内群体偏爱｜用LLM扮演信贷审批员，随机分配其所属“银行组”，处理为申请人所属组（内/外群体）和群体规模，结果变量为贷款批准率和额度，与人类信贷员的历史审批数据对照。"}},{"id":"2609.00345","version":1,"title":"Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment","zh_title":"LLM了解你的社区吗？审计LLM先验用于社区级移动性预测与结构对齐","abstract":"Human mobility is central to urban planning, transportation, public health, and emergency response, yet fine-grained trajectory data are often proprietary, restricted, and privacy-sensitive. Large language models (LLMs) offer a potential alternative by generating plausible mobility traces and predicting individual movement, but their ability to infer aggregate neighborhood-level mobility remains unclear. We evaluate zero-shot LLMs on Census Block Group-level mobility prediction across four U.S. metropolitan areas using anonymized Cuebiq data to construct point-level, trajectory-level, and temporal mobility outcomes, paired with sociodemographic and built-environment predictors. We compare LLM predictions with supervised baselines and introduce a directional alignment analysis to test whether LLM-implied predictor effects agree with empirical OLS and Jonckheere-Terpstra trends. Supervised models achieve 0.580 average accuracy, compared with 0.435 for the best LLM, with spatial extent outcomes showing the strongest predictability but also the largest LLM-baseline gaps. Directional analysis shows that LLMs often rely on coarse, stable predictor-level priors that remain similar across outcomes and cities, including asymmetric treatment of protected-group predictors. Overall, LLMs can partially recover aggregate mobility patterns from urban context, but their predictions should not be treated as structurally grounded without auditing empirical alignment and potential bias.","authors":["Saad Mohammad Abrar","Eesha Kurella","Arnav Dadarya","Naman Awasthi","Kazi Tasnim Zinat","Vanessa Frias-Martinez"],"categories":["cs.LG","cs.CY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00345","pdf_url":"https://arxiv.org/pdf/2609.00345","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类移动性","偏差审计"],"reason":"用LLM预测社区级人类移动性，并与真实数据对照，审计其偏差与对齐。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":4,"question":"LLM能否仅凭社区的社会人口与建成环境特征，直接推断出社区层面的聚合人类移动性结果，并且其预测是否与经验数据中的结构性关系对齐？","design":"使用零样本LLM（如GPT-4等）作为预测模型，输入美国四个大都市区人口普查区块组（CBG）的社会人口和建成环境特征，预测三类聚合移动性结果（点级、轨迹级、时间级），并与监督学习基线比较。","baseline":"使用Cuebiq提供的匿名化手机定位数据，聚合到人口普查区块组级别，构建真实移动性结果，作为LLM预测的对照基准。","findings":"监督学习基线平均准确率为0.580，最佳LLM为0.435，LLM在空间范围结果上差距最大。方向对齐分析显示LLM依赖粗粒度、稳定的先验，且对受保护群体预测因子处理不对称。","reliability":"论文指出LLM预测不应被视为结构上可靠，除非经过经验对齐和潜在偏差审计；LLM的先验在不同结果和城市间保持不变，可能反映浅层启发式或刻板印象。","relevance":"该研究直接评估LLM作为人类移动性预测代理的可靠性，并与真实大规模人类数据对照，属于批判性仿真研究，对关注LLM仿真偏差和失效条件的研究者有参考价值。","inspiration":"借鉴其方向对齐分析方法，通过比较LLM隐含的预测因子效应与经验回归趋势来审计结构一致性。｜可迁移到信贷审批歧视研究，用LLM模拟信贷员决策并检验其对种族、性别等受保护特征的敏感度。｜以LLM作为虚拟信贷员，输入申请人特征（含受保护属性），预测贷款批准概率，与真实信贷数据（如HMDA）中的批准率和歧视模式进行对照。"}},{"id":"2609.00608","version":1,"title":"Investigating Assistant Bias in LLM User Simulators Using a Role Vector","zh_title":"使用角色向量研究LLM用户模拟器中的助手偏差","abstract":"LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit \"assistant bias,\" a tendency to cooperate and pursue task goals. They rarely reproduce the frustration or disengagement that real users exhibit, compromising evaluation validity. Prior work outlines that this bias is baked in during model training, which role-playing prompts fail to override. We analyze this bias from model activations, extracting a user role vector by contrasting how the model represents user versus assistant perspectives on the same dialogue. We observe two findings: (i) the user direction is identifiable in activations, elicits user-like behaviors, and captures characteristics distinct from assistant traits; and (ii) although user-role activation associates with simulation realism and steering strengthens it, it can exaggerate user behaviors and override individual user profiles. Together, our findings provide a representation-level analysis of LLM user simulators, confirming that assistant bias is structurally identifiable and that user behavior can be directionally analyzed.","authors":["Daeheon Jeong","Yoonjoo Lee","Eugene Choi","Sinie van der Ben","Juho Kim"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00608","pdf_url":"https://arxiv.org/pdf/2609.00608","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B4"],"tags":["LLM用户模拟器","助手偏差","仿真有效性"],"reason":"研究LLM用户模拟器的助手偏差，评估仿真有效性，批判性指出失效条件，可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":10,"question":"大语言模型用户模拟器中的助手偏差是否在模型激活空间中可识别，并能否通过角色向量操控来改善模拟真实性？","design":"使用Qwen 3.5 9B等模型，通过对比同一对话中用户与助手视角的反思激活，提取用户角色向量；施加该向量进行激活引导，测量对沟通风格、行为反应及多轮交互真实性的因果影响。","baseline":"SimulatorArena基准中的真实用户交互数据，以及RealUserSim中的真实用户模拟日志。","findings":"用户角色方向在激活空间中可识别，能引发更用户化的行为，并与助手特质在几何上负相关；激活引导虽能提升写作风格相似性，但会夸大用户行为并掩盖个体用户画像。","reliability":"论文承认主要基于Qwen 3.5 9B，跨模型验证有限；评估依赖代理指标而非直接人类判断；基准场景较窄，仅限数学辅导和部分对话。","relevance":"该研究直接针对LLM用户模拟器的核心偏差问题，提供了表示层面的机制分析和批判性评估，对关注仿真可靠性与失效条件的研究者具有重要参考价值。","inspiration":"借鉴其通过对比激活提取角色向量并进行因果引导的方法，可用于在LLM模拟中分离特定行为倾向｜可迁移到消费者金融决策模拟，如信贷申请中的风险偏好或政策反应｜用LLM模拟消费者，施加“风险厌恶”或“耐心”角色向量，测量其信贷选择行为，并与真实信贷申请数据或实验数据对照。"}},{"id":"2609.00565","version":1,"title":"Aligned but Flattened: Analyzing the Trade-off between Cultural Alignment and Diversity in LLMs","zh_title":"对齐但扁平化：分析LLMs中文化对齐与多样性之间的权衡","abstract":"Cultural fine-tuning has become the de facto paradigm for building culture-aware large language models (LLMs), yet existing optimization exclusively for alignment scores provides an incomplete portrait of cultural fidelity by systematically obscuring inherent cultural diversity. This unidimensional evaluation lens prompts a fundamental question: do models genuinely perceive distinct cultural nuances, or do they merely memorize dominant cultural values? To address this, we propose a synergistic evaluation framework that jointly formalizes cultural alignment and diversity. Through extensive benchmarking of six mainstream LLMs on the World Values Survey, this framework uncovers a systematic and critical trade-off: the pursuit of cultural alignment consistently incurs an acute expense of diversity, leading to severe \"cultural flattening.\" Investigating this behavioral shift, we demonstrate that these superficial alignment gains stem from models artificially anchoring to dominant majorities, converging onto a monolithic response pattern that wipes out the heterogeneous distributions inherent to human groups. Crucially, our mechanistic analysis suggests that this diversity collapse is not merely a behavioral anomaly but more likely a structural consequence of the low-rank bias inherent in neural network optimization. Therefore, our findings expose the limitations of current post-training paradigms and call for a shift toward alignment objectives that preserve cross-cultural pluralism.","authors":["Jingshen Zhang","Shaoyang Xu","Wenxuan Zhang"],"categories":["cs.SI","cs.CL"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00565","pdf_url":"https://arxiv.org/pdf/2609.00565","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["文化仿真","算法保真度","价值观调查"],"reason":"评估LLM文化对齐与多样性，使用世界价值观调查真实数据对照，揭示仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":9,"question":"文化微调在提升LLM文化对齐度的同时，是否以牺牲文化多样性为代价？","design":"对六种主流LLM在基于世界价值观调查的文化特定数据上进行微调，测量微调前后模型在文化对齐度和文化多样性上的变化，并分析其行为模式与内部表征。","baseline":"世界价值观调查的真实人类回答数据，按社会人口群体分组。","findings":"文化微调一致地提高了对齐度，但显著降低了行为多样性，导致“文化扁平化”。这种多样性损失源于模型锚定主流多数群体，并可能由神经网络优化的低秩偏差结构性导致。","reliability":"论文指出当前对齐目标仅关注聚合相似度，忽视了文化多样性，且微调受限于低秩子空间，可能边缘化少数文化；但未系统讨论其他失效条件。","relevance":"该研究直接评估LLM仿真人类文化价值观的可靠性，揭示了对齐与多样性的权衡，对关注仿真偏差和真实数据对照的研究者具有重要参考价值。","inspiration":"借鉴其联合测量对齐与多样性的评估框架，并利用真实调查数据作为基准。｜可迁移到经济金融领域的文化差异研究，如跨文化消费偏好、金融风险态度或政策接受度的仿真。｜以LLM模拟不同文化背景的消费者，施加文化微调处理，测量其对金融产品的偏好分布，并与世界价值观调查或实际消费数据对照，检验对齐与多样性权衡。"}},{"id":"2609.01519","version":1,"title":"When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation","zh_title":"当护栏看似有效：LLM智能体商业评估中的构念效度失效","abstract":"Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. We audit this risk in a multi-turn buyer--seller testbed for configurable hotel transactions. An initial implementation reported welfare gains from two marketplace guardrails of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B--14B ladder. It also gave guarded and unguarded agents different offer schemas and choice procedures. Holding the schema and buyer chooser fixed changes the paired contrasts to +7.2, -13.9, and +23.8. The four largest 14B single-generation effects averaged +229; after three generations per profile-condition, they averaged +37.6 (95% bootstrap interval [-34.2, 109.3]), while generation residuals account for 49.9% of variation in this post-hoc probe. A seller-incentive check is non-monotone: increasing profit pressure produces less profit than the default seller prompt. Scripted positive controls show why this matters. A profit-maximizing seller already attains first-best welfare, so guardrails mostly redistribute and reduce welfare; they create welfare only when the seller is explicitly programmed to force inefficient bundles. We contribute a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting, and returning INVALID or INCONCLUSIVE before substantive policy claims. In our case, the original estimate is INVALID under protocol isolation, while the controlled study remains INCONCLUSIVE under incentive validity and stochastic stability. The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks.","authors":["Peiying Zhu","Sidi Chang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01519","pdf_url":"https://arxiv.org/pdf/2609.01519","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A2","A4","B4"],"tags":["LLM仿真","构念效度","市场模拟"],"reason":"评估LLM市场仿真中构念效度失效，提出验证框架，批判性指出仿真失效条件，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":11,"question":"LLM智能体市场仿真中，平台护栏的福利效应是否真实存在，还是源于实现脚手架（offer schema与choice procedure）的差异？","design":"用Qwen2.5-Instruct（1.5B、3B、14B）同时扮演买家和卖家，在酒店交易多轮对话中施加两种护栏（阻止买家信息、强制可选组件不捆绑），测量福利变化；通过统一schema和chooser、重复生成、激励操纵检查、脚本化正控制等审计设计来检验构念效度。","baseline":"无对照","findings":"原始实现中护栏的福利增益在统一脚手架后大幅缩水甚至符号反转（如3B从+35.0变为-13.9），14B效应不显著；激励检查显示更强的利润指令并未单调提高利润，脚本化正控制表明护栏仅在卖家被编程为强制低效捆绑时才创造福利。","reliability":"论文承认合成档案中43/60的负外部选项效用使接受更容易，不具代表性；重复生成探针是事后选择，受赢家诅咒影响；激励检查样本量小，不足以排序提示；整体结论受限于特定模型家族和酒店场景。","relevance":"该研究直接针对LLM仿真在经济学评估中的构念效度问题，提供了系统的失效诊断框架和审计协议，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其构念效度审计协议，将处理与实现脚手架分离，并通过脚本化正控制和重复生成来检验效应稳健性｜可迁移到政策评估中的市场设计仿真，如平台监管、拍卖机制或价格歧视策略的福利分析｜用LLM扮演消费者和商家，施加某种政策处理（如信息屏蔽或价格上限），测量交易价格、成交率和福利，并与真实电商平台或实验数据对照，同时统一对话协议并重复生成以评估方差。"}},{"id":"2609.00310","version":1,"title":"Emotional Labor Strategy Preferences in LLM Personas","zh_title":"LLM人格中的情绪劳动策略偏好","abstract":"Emotional labor is the effortful management of emotional displays to meet social or professional expectations. Personality traits have been correlated with emotional labor strategies, yet research on this link relies almost exclusively on self-report scales administered only in occupational settings. We investigate whether large language models injected with psychometrically grounded personas reproduce these personality-driven selection patterns across everyday social scenarios. We construct the first emotional labor strategy dataset of 500 socially situated events, each offering three behavioral choices corresponding to surface acting, deep acting, and genuine expression. We source 50 fictional characters from a large-scale personality repository and profile each through two parallel tracks: observer-rated bipolar adjective composites and in-character self-report items. Five LLMs evaluate all scenarios under both persona conditions. We find that models align more towards deep acting, and that Conscientiousness and Emotional Stability consistently predict this preference. Entropy analysis confirms that persona reliably influences the output and varies across models and emotions.","authors":["Mohammad Saim","Tianyu Jiang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00310","pdf_url":"https://arxiv.org/pdf/2609.00310","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","情绪劳动","人格测量"],"reason":"用LLM persona复现人格与情绪劳动策略关联，有真实人类数据对照，属仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":14,"question":"注入心理测量学人格的LLM角色是否会在日常社交场景中复现人格特质与情绪劳动策略选择之间的关联？","design":"用五个LLM扮演50个来自影视剧的虚构角色，通过观察者评定的双极形容词和角色内自陈IPIP-50两种方式注入人格特质，让模型在500个社交场景中从表层扮演、深层扮演、真实表达三种策略中选择一种，分析人格维度与策略选择的关系。","baseline":"已有组织心理学文献中基于人类自我报告量表得出的人格与情绪劳动策略关联模式，如高宜人性、外向性倾向深层扮演或真实表达，高神经质倾向表层扮演。","findings":"模型整体更偏好深层扮演；尽责性和情绪稳定性在两种人格注入方式下都一致预测深层扮演偏好。熵分析表明人格注入可靠地影响输出，且影响随模型和情绪类别变化。","reliability":"论文未讨论","relevance":"该研究用LLM人格角色复现心理学中的人格-行为关联，有真实人类数据作为基准，属于仿真验证研究，对关注LLM作为人类被试替代品及仿真可靠性的研究者有参考价值。","inspiration":"可借鉴其用两种人格测量方式（观察者评定与自陈量表）注入角色并比较一致性的做法，以及用选择任务而非自由生成来测量行为倾向的设计。｜可迁移到经济金融中的个体决策差异研究，如风险偏好、时间偏好、消费选择或投资行为中的人格影响。｜用LLM扮演不同人格特质的消费者或投资者，施加经济决策场景（如跨期选择、风险投资），测量其选择，并与真实人格-决策关联的实证数据（如大型调查或实验数据）对照，检验仿真一致性。"}},{"id":"2609.00982","version":1,"title":"Disclosure-Gated User Simulation for Companion-Agent Evaluation","zh_title":"面向陪伴智能体评估的披露门控用户仿真","abstract":"Using a large language model to play the user is now standard in scalable evaluation. It has a repeatedly diagnosed failure: the simulated user is excessively cooperative, so a system under test can score by the sheer number of questions it asks rather than by making the user willing to speak. We answer with a disclosure gate conditioning information release on the companion agent's behaviour: its state is a ladder of five ordered gates, merged onto three observable depth layers. We specify, ablate, and audit it, and train a user simulator against that specification. Gating behaviour is learned from the training corpus's synthetic branch, while the real branch supplies how people speak and react; after training, the simulator need not be told at runtime which gate each item sits behind. The gate is a load-bearing component of the environment: on the English corpus of a published companion-agent benchmark (CompanionBench), once training no longer states per example which gate each item sits behind, the largest rank displacement across 12 systems under test exceeds the noise band set by re-running that environment under a new seed, while per-system scores show no detectable change. We state two acceptance criteria: a ranking must be order-preserving, and absolute scores must be scale-stable. Of the candidates we examine, only one passes both -- the simulator we release -- and its leaderboard correlates at 0.993 with the benchmark's original simulator. By contrast, prompting a frontier model as the simulator barely moves the ranking while shifting every score upward -- a shift invisible to anyone checking the ranking alone. The environment we specify is the one that benchmark already used. That publication describes the mechanism in about four hundred words, and we supply what it lacked: specification, ablations, human studies, negative controls, and downstream sensitivity analysis.","authors":["Yao Liu","Yu He"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00982","pdf_url":"https://arxiv.org/pdf/2609.00982","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["用户仿真","评估方法","偏差审计"],"reason":"用LLM模拟用户评估陪伴智能体，有真实人类数据对照，并审计仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":15,"question":"如何设计一个披露门控机制来校准LLM用户模拟器的过度合作倾向，使其在陪伴智能体评估中产生可靠且可复现的排名与分数？","design":"训练一个用户模拟器，其信息释放由五级门控阶梯控制，门控状态取决于陪伴智能体的行为；通过剥离训练时或推理时的门控信息、改变数据分支、模型规模等操纵，测量模拟器自身门控行为（内在层）和下游基准排名与分数变化（下游层）。","baseline":"使用CompanionBench中真实人类对话语料库作为真实分支，提供人们如何说话和反应的模式；另有合成分支用于学习门控行为。","findings":"披露门控是环境的关键组成部分：从训练语料中移除逐例门控信息后，12个被测系统的最大排名位移超过新种子重跑环境的噪声带，但各系统分数无显著变化。在候选模拟器中，只有发布的模拟器同时满足顺序保持和尺度稳定两个接受标准，其排行榜与基准原始模拟器的相关性为0.993；而提示前沿模型作为模拟器几乎不改变排名，却使所有分数整体上移。","reliability":"论文指出内在层和下游层可能给出不同答案，且二者不可相互替代；门控行为主要来自合成分支，剥离后行为几乎消失；提示的前沿模型仅在运行时读取门控信息才能保持门控，剥离后急剧退化。","relevance":"该研究直接针对LLM模拟人类被试的过度合作偏差，提供了可审计的机制和真实人类对照，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其将关键机制（披露门控）作为可剥离组件进行消融实验，并用排名位移与噪声带比较来检验下游影响的方法。｜可迁移到经济金融中的信息不对称场景，如信贷审批中借款人信息披露、谈判中策略性信息保留、或消费者对产品信息的主动询问。｜设计一个LLM模拟借款人的信贷申请实验，处理变量为贷款官员的提问策略（主动询问vs被动等待），结果变量为借款人自愿披露的信息量和贷款获批率，用真实信贷对话数据作为基准校准模拟器的披露行为。"}},{"id":"2609.00250","version":1,"title":"CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships","zh_title":"CompanionSim：用于评估人机关系中拟人化的合成数据","abstract":"Many people now see AI systems as not just productivity tools but as social companions. Researchers are eager to study the consequences of AI companionship behaviors, such as validation, which evoke trust, empathy, and attachment in human-human interaction. However, human-AI interaction data is limited and unreliable, slowing research progress. We scale small amounts of real-world data by simulating multi-turn human-chatbot dialogue across a range of chatbot behaviors and use cases. We release CompanionSim: a simulation framework with 2,240 simulated human-chatbot conversations representing 16 chatbot behaviors across seven use cases. Human participants annotated the simulated conversations and real-world conversations in two experiments probing perceptions of companionship behaviors. We conducted Study 1 with a U.S. representative sample ($N_{1}~=~628$) and Study 2 across the U.S., U.K., India, and Nigeria ($N_{2}~=~3,646$). Surprisingly, we find that companionship behaviors reduced likability, humanlikeness, and trust in AI chatbots. These effects were larger in particular subgroups: women and older participants saw companionship chatbots as less likable, humanlike, and trustworthy. We encourage researchers to leverage real-world and synthetic data together to study the differential impacts of AI companions and to create benchmark evaluations of AI chatbots.","authors":["Jacy Reese Anthis","Mark D\\'iaz","Renee Shelby"],"categories":["cs.CY","cs.AI","cs.CL","cs.LG"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00250","pdf_url":"https://arxiv.org/pdf/2609.00250","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人机交互","合成数据"],"reason":"用LLM模拟人类与聊天机器人对话，并有人类标注对照，但仿真对象是对话而非人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":13,"question":"人类如何感知AI聊天机器人中的陪伴行为（如验证、情感表达）对喜爱度、人性化感知和信任的影响？","design":"使用Gemini 2.5 Flash模拟生成2240段多轮人机对话，覆盖16种聊天机器人行为（13种陪伴行为、2种反陪伴行为、1种中性控制）和7个真实使用场景；人类被试对模拟对话和真实对话进行标注，评估自然度、喜爱度、人性化、认知信任和情感信任。","baseline":"真实世界人机对话数据（少量）作为对照，但具体来源和规模未在节选中详细说明。","findings":"陪伴行为降低了聊天机器人的喜爱度、人性化和认知信任，且这些效应在女性、年长者和AI使用频率较低的人群中更为明显。","reliability":"论文未讨论仿真失效条件或局限性（节选部分未提及）。","relevance":"该研究用LLM模拟人机对话并有人类标注对照，但仿真对象是对话而非人类被试，与研究者关注的人类仿真实验（如复现调查回答、决策模式）有部分重叠，值得阅读以了解合成数据在交互场景中的应用和局限。","inspiration":"借鉴其系统操纵行为特征并用人类标注验证合成数据质量的方法，可用于生成经济决策对话场景｜可迁移到消费者与金融顾问聊天机器人的互动研究，如理财建议中的情感化语言对投资决策的影响｜以LLM模拟消费者与聊天机器人对话，处理变量为是否包含陪伴行为（如共情表达），结果变量为消费者对建议的采纳意愿和风险偏好，对照真实人类对话数据（如银行客服记录）进行校准。"}},{"id":"2609.01432","version":1,"title":"Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation","zh_title":"引用更少批判性：大语言模型重塑科学引用的修辞与影响范围","abstract":"Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.","authors":["Yixuan Liu","Lin Chen","Zhuoqi Liu","Jianglin Lu","Dakota Murray"],"categories":["cs.DL","cs.CL","cs.CY","cs.SI"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01432","pdf_url":"https://arxiv.org/pdf/2609.01432","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","引用行为","科学计量"],"reason":"用LLM生成引用行为并与人类引用对照，属于仿真人类行为且有人类数据基准，但非典…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":16,"question":"LLM 在生成科学引用时，是否复现了人类引用中的修辞意图分布，以及引用意图如何调节 LLM 与人类在引用选择上的差异？","design":"构建掩码引用任务：对每篇论文中的每个引用上下文，用 LLM 生成替代引用句，形成与人类引用直接可比的 counterfactual 语料；使用 LLM-as-a-judge 分类引用意图（支持、对比、提及），并利用 2000 万边合著网络测量引用作者与被引作者的社会距离。","baseline":"来自 1,746 篇顶级 NLP 会议论文的 63k+ 引用上下文和 132k+ 引用的人类真实引用行为。","findings":"与人类引用相比，LLM 的引用显著更少批判性；LLM 过度引用热门和较旧的论文，这一倾向在对比引用中更为明显，而人类写作更常引用近期、小众的工作。人类往往引用自己密切社交网络内的论文（尤其是支持性引用），而 LLM 倾向于引用社交距离更远的作者。","reliability":"论文未讨论","relevance":"该研究用 LLM 生成引用行为并与人类引用对照，属于仿真人类行为且有人类数据基准，但场景是科学写作而非经济决策，与你的核心关注点（调查回答、实验行为、态度分布、决策模式）有一定距离，不过其方法（掩码任务、意图条件分析、社会网络距离测量）对仿真研究设计有参考价值。","inspiration":"值得借鉴的是其位置对齐的掩码生成设计，以及按意图分层分析偏差的方法，可以用于研究 LLM 在不同语境下的行为差异。｜可以迁移到政策评估中的文本生成场景，例如让 LLM 模拟分析师撰写研究报告时的引用行为，或模拟政策制定者引用学术证据的方式。｜一个可行的设计是：以真实分析师报告中的引用为基准，用 LLM 在相同上下文下生成替代引用，比较引用意图分布、被引文献特征（如时效性、影响力）以及引用者与被引者的社会网络距离，从而评估 LLM 仿真分析师引用行为的保真度。"}},{"id":"2609.00248","version":1,"title":"Authority Bias in Conversational Search Engines for Academic Paper Recommendation","zh_title":"学术论文推荐对话搜索引擎中的权威偏差","abstract":"Large Language Models (LLMs) are increasingly used as conversational search engines for academic literature, yet whether they judge papers on content or on authority signals has not been tested causally. We investigate authority bias: systematic preference for papers based on author prestige, venue, and citations rather than content. Holding title and abstract constant, we vary authority metadata across three counterfactual conditions (original, flipped, boosted) over eight LLMs (five open-weight and three frontier closed-weight) in an in-context, single-turn, top-1 recommendation setting. Our experiments show that authority bias is substantial and directional, varies markedly across models, and is only partially addressable through prompt-level debiasing. We further document a say-do gap: debiasing instructions suppress authority mentions far faster than authority-driven flips, so surface auditing systematically underestimates behavioral bias.","authors":["Uthman Jinadu","Parsa Ghazvinian","Anjila Budathoki","Benjamin M. Ampel","Rajshekhar Sunderraman","Yi Ding"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00248","pdf_url":"https://arxiv.org/pdf/2609.00248","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM偏差","行为评估","推荐系统"],"reason":"研究LLM推荐中的权威偏差，评估其行为偏差，可迁移到仿真可靠性评估。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":12,"question":"LLM在学术论文推荐中是否存在基于作者声望、期刊和引用的权威偏差，且该偏差是否可通过提示词去偏缓解？","design":"该研究并非以LLM仿真人类被试，而是对LLM本身进行审计：在保持论文标题和摘要不变的情况下，通过三种反事实条件（原始、翻转、增强）操纵权威元数据，测试八个LLM在单轮、上下文内、top-1推荐任务中的选择行为，并记录推荐变化和理由文本。","baseline":"无对照","findings":"权威偏差显著且方向一致，不同模型间差异明显；提示词去偏只能部分缓解，且存在言行差距：去偏指令抑制权威提及的速度远快于减少权威驱动的选择翻转。","reliability":"论文未讨论","relevance":"该研究通过反事实操纵和因果推断测量LLM的行为偏差，方法可迁移到评估LLM在人类仿真中的可靠性，尤其是当权威信号可能污染决策时。","inspiration":"值得借鉴其反事实操纵设计：固定内容、仅改变元数据，以分离权威信号对决策的因果影响。｜可迁移到信贷审批歧视研究，检验LLM是否因申请人姓名、学校或雇主声望而产生偏差。｜以LLM作为信贷审批员，向模型呈现相同贷款申请但改变申请人背景（如毕业院校、工作单位），观察审批结果是否变化，并与真实信贷数据中的歧视模式对照。"}},{"id":"2608.18294","version":2,"title":"Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements","zh_title":"无金标准标签下AI生成数据的去偏推断：基于多重不完美测量的识别","abstract":"An increasing number of scholars use AI to measure variables they subsequently include in downstream analyses. Although AI-measured variables are often analyzed as if observed without error, ignoring prediction errors in automated measurement leads to substantial bias and invalid confidence intervals in downstream analyses, even if AI measurement accuracy is high, e.g., above 90%. Existing solutions, such as design-based supervised learning and prediction-powered inference, combine error-prone AI-based measurements with gold-standard labels, which may be costly and difficult to obtain in some application areas. In this paper, we propose debiased inference with multiple imperfect measurements (DMM), a framework that combines multiple error-prone AI measurements to enable valid downstream inference without gold-standard labels. Building on the established results on CP decomposition, DMM assumes that these measurements are independent conditional on the latent true label and observed unit-level features, such as text features represented by embeddings. This framework allows for unknown misclassification rates to vary across annotation methods (e.g., large language models) and across units of annotation (e.g., texts). Under this assumption, we use semiparametric inference theory to prove that the DMM estimator is consistent and asymptotically normal, enabling valid inference for a wide range of downstream statistical analyses common in the social sciences. Our simulation results show that DMM yields valid inference and that adding accurate, though imperfect, measurements can improve efficiency. Focusing on common applications of large language model annotations, we also develop diagnostics to assess the conditional independence assumption.","authors":["Naoki Egami","Sooahn Shin"],"categories":["stat.ME","cs.AI","cs.CL","cs.LG","stat.ML"],"primary_category":"stat.ME","announce_type":"replace-cross","date":"2026-09-02","first_seen":"2026-08-20","revised_at":"2026-09-02","abs_url":"https://arxiv.org/abs/2608.18294","pdf_url":"https://arxiv.org/pdf/2608.18294","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["测量误差","统计推断","LLM标注"],"reason":"LLM作为测量工具，非仿真被试，但方法可迁移到仿真数据校正","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":10,"question":"如何在没有金标准标签的情况下，利用多个有误差的AI测量进行有效的下游统计推断？","design":"本文不是仿真研究，而是提出一种统计方法：假设多个AI测量（如不同LLM标注）在给定潜在真实标签和观测特征（如文本嵌入）的条件下相互独立，利用CP分解和半参数推断理论构建去偏估计量，从而无需金标准标签即可进行下游分析。","baseline":"无对照","findings":"DMM估计量在条件独立假设下具有一致性和渐近正态性，能提供有效的置信区间；模拟显示添加准确但不完美的测量可以提高效率。","reliability":"论文承认条件独立假设可能不成立，并开发了诊断方法来评估该假设；若存在未观测的共享误差来源，方法可能失效。","relevance":"该研究为使用LLM进行测量并用于下游分析的研究者提供了无需金标准标签的校正方法，对关注LLM仿真数据可靠性的研究者有重要参考价值。","inspiration":"借鉴其利用多个有误差测量相互校正的思路，可在无金标准时提高测量准确性｜可迁移到经济金融中需要文本分类或情感分析的研究，如新闻情绪对资产价格的影响、政策文本的立场识别等｜设计：用多个LLM对财经新闻进行情感分类，以股票收益率作为结果变量，用人工标注的小样本作为验证集评估校正效果。"}},{"id":"2606.30085","version":2,"title":"Tastes without distinction: silicon samples and the synthetic construction of tastes","zh_title":"无差别的品味：硅样本与品味的合成建构","abstract":"Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produce reasonable approximations of cultural taste remains an open empirical question that becomes more urgent by the day, with market research companies already offering provisional 'synthetic' survey panels and the contamination of standard survey data from LLM-generated responses. In this study, we build on past work on silicon sampling by extending considerations of their ecological, relational, and positional fidelity in the doomain of cultural tastes. We use large-language models from OpenAI, Anthropic, and DeepSeek to produce 554,940 silicon surrogates of survey respondents from the Survey of Public Participation in the Arts (SPPA). We find these silicon surrogates' tastes to be highly stylized facsimiles of human tastes. First, silicon samples are super-omnivorous with a systematic postive-bias for liking. These individual-level bias of silicon samples are not well-explained by the WEIRD-bias often discussed in the literature. Second, the complex relationality in real taste structures is completely distorted among silicon samples. Third, very little of the known cultural alignment between tastes and social space are preserved. Silicon samples juvenilize age-taste associations, resurrect anachronistic class-taste associations, and caricaturize gender- and race-taste associations. Key words: AI, taste, consumption, culture, silicon sampling, meta-analysis.","authors":["Xiangyu Ma","Mengmi Zhang","Shannon Ang","Minne Chen"],"categories":["cs.CL","econ.GN","q-fin.EC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-06-29","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2606.30085","pdf_url":"https://arxiv.org/pdf/2606.30085","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["硅采样","文化品味仿真","算法保真度"],"reason":"用LLM生成硅样本模拟文化品味调查，并与真实SPPA数据对照，评估仿真偏差与失…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":1,"question":"大语言模型生成的硅样本能否忠实再现人类文化品味及其与社会空间的关联？","design":"以2012年美国公众参与艺术调查（SPPA）为参照，将每个受访者的人口统计特征转化为角色提示，输入OpenAI、Anthropic和DeepSeek的大语言模型，为每个人类受访者生成多个硅样本（共554,940个），比较硅样本与人类样本在品味分布、品味结构关系及品味与社会空间关联上的保真度。","baseline":"2012年SPPA调查中的人类受访者数据，并使用SPPA的自举重采样作为人类数据内部一致性的基准。","findings":"硅样本表现出超级杂食性和系统性正向偏好，且这种偏差不能用文献中常讨论的WEIRD偏差解释；真实品味结构中的复杂关系在硅样本中被完全扭曲，品味与社会空间的已知关联几乎未被保留，年龄-品味关联被年轻化，阶级-品味关联被复活为过时模式，性别和种族-品味关联被漫画化。","reliability":"论文指出硅样本的个体层面偏差（如超级杂食性和正向偏好）不能由WEIRD偏差解释，且其关系结构和与社会空间的关联严重失真，表明在文化品味领域硅采样保真度很低，但未明确讨论具体失效条件或局限性。","relevance":"该研究直接评估LLM仿真人类文化品味调查的可靠性，并与真实SPPA数据对照，发现系统性偏差和结构失真，对关注LLM仿真在社会科学中有效性的研究者具有重要参考价值，值得阅读原文以了解具体偏差模式和评估框架。","inspiration":"借鉴其将人口统计特征转化为提示、生成大量硅样本并与真实调查数据对照的仿真设计，以及从生态、关系和位置保真度三个维度评估偏差的方法。｜可迁移到消费者偏好调查、市场细分或文化消费的经济学研究中，例如用LLM模拟不同人口群体的品牌偏好或娱乐消费选择。｜设计：以某消费者支出调查（如美国消费者支出调查CE）为基准，提取受访者人口特征生成提示，用多个LLM生成硅样本，让其回答关于品牌偏好或娱乐活动参与的问题，比较硅样本与真实人类在偏好分布、偏好结构及与社会经济地位关联上的差异，并检验硅样本是否复现已知的消费分层模式。"}},{"id":"2608.03044","version":2,"title":"Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation","zh_title":"仿真还是估计？基础与后训练语言模型在意见模拟中的不同优势","abstract":"Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity. We propose that much of this conflict stems from conflating two distinct tasks. We call the first task emulation, in which models generate individual responses that aggregate into a population distribution. We call the second task estimation, in which models directly predict the population distribution. Evaluating six matched base and post-trained models on the Pew American Trends Panel, we find that base models are the stronger emulators: they produce response distributions closer to human ground truth and better preserve demographic structure. Post-trained models are generally the stronger estimators, producing more accurate distributional predictions when asked directly. We argue that model selection for human simulation should be guided by whether the task requires generating text or predicting distributions.","authors":["Seth Grief-Albert","Jessica Bo","Difan Jiao","Ashton Anderson"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-05","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.03044","pdf_url":"https://arxiv.org/pdf/2608.03044","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","意见模拟","算法保真度"],"reason":"直接研究LLM仿真人类意见，区分仿真与估计任务，使用真实调查数据对照，并指出模…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":2,"question":"在意见仿真中，基础模型与后训练模型在仿真（emulation）与估计（estimation）两种任务上表现有何差异？","design":"使用三对匹配的基础与后训练模型（Qwen3-14B、Olmo-3-7B、Olmo-3-32B）在Pew美国趋势面板59个经济意见问题上进行仿真：仿真任务通过开放式生成个体回答并聚合，估计任务直接预测总体分布；施加七种人口学条件（无条件、民主党、共和党、非常自由派、非常保守派、高收入、新教徒），以总变差距离和Wasserstein距离衡量分布保真度。","baseline":"Pew Research Center American Trends Panel Wave 54 的真实人类调查数据，包含59个四选项经济意见问题及人口学子群体分布。","findings":"基础模型在仿真任务上始终优于后训练模型，生成的分布更接近人类真实分布且更好地保留人口学结构；后训练模型在直接估计总体分布时通常更准确。","reliability":"论文未讨论","relevance":"直接研究LLM仿真人类意见，区分仿真与估计任务，使用真实调查数据对照，并指出模型选择应基于任务类型，对关注LLM仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其区分仿真与估计任务、使用开放式生成避免位置偏差、并用真实调查数据做基准对照的方法｜可迁移到经济预期形成或消费者信心调查的仿真，如模拟不同收入群体对通胀预期的分布｜用基础模型模拟个体受访者回答开放式经济预期问题，聚合后与密歇根消费者调查的真实分布对比，同时用后训练模型直接估计分布，比较两种范式的准确性。"}},{"id":"2608.29455","version":1,"title":"Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human Surrogates","zh_title":"项目均值替代：为何更丰富的人物数据未能提升LLM作为人类替代品的表现","abstract":"LLMs are increasingly used as human surrogates, often on the premise that richer persona data could make them substitutes or exploratory tools for specific individuals. We test this premise across four datasets covering more than 400,000 participants and more than 6,000 survey items and experimental outcomes. LLMs perform well at the aggregate level: their average responses closely align with average human responses to the same items. But this success largely reflects predicting each item's average human response. Once each item's human mean is removed, LLM predictions explain only 3.05% of the remaining respondent-specific variation, far below the 53.6% human test-retest benchmark. Richer personas, model variants, and fine-tuning do not close this gap. In variance analyses, once item means are removed, the reliable remaining signal is person-by-item. It captures how a respondent departs from the mean on a particular item and is about 8.9x larger than the stable person effect. Persona data encode the respondent, but not this item-specific deviation. LLM responses also compress human response distributions, using less spread, fewer response categories, and distorted distributional shapes. We call this pattern item-mean surrogacy. Current LLM surrogates can approximate item averages, but not the distributions or respondent-specific deviations needed to replace individual humans. We propose four empirical tests for LLM-based human-surrogate claims.","authors":["Daehwan Ahn","Chengfeng Mao","Dokyun Lee"],"categories":["cs.CL","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29455","pdf_url":"https://arxiv.org/pdf/2608.29455","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM仿真","人类替代","算法保真度"],"reason":"直接评估LLM作为人类替代品的可靠性，使用大规模人类数据对照，发现仅能预测项目…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":4,"question":"LLM 作为人类替代品时，更丰富的人物角色数据能否提高其对个体层面回答的预测能力？","design":"使用四个数据集（Megastudy、Survey、SocSci210、ANES），覆盖超40万参与者和6000多个调查项目，用多种LLM模型和提示方法（包括丰富人物角色、微调）生成对相同项目的回答，并与真实人类回答对比，测量项目均值、分布和个体偏差的恢复程度。","baseline":"真实人类数据：Megastudy 和 Survey 数据集包含同一批参与者的 LLM 人物角色和真实回答，并提供人类重测信度（53.6%）作为个体预测上限；SocSci210 和 ANES 提供不同领域的人类回答。","findings":"LLM 在聚合层面表现良好，平均回答与人类平均回答高度一致，但去除项目均值后，LLM 预测仅解释 3.05% 的个体变异，远低于人类重测基准 53.6%。丰富人物角色、模型变体和微调均未缩小差距；方差分析显示稳定的人-项目交互效应是主要信号（44.0%），而人物角色数据仅编码稳定的人主效应（4.9%），无法捕捉项目特定偏差。","reliability":"论文指出 LLM 替代品仅能近似项目均值，无法恢复分布或个体偏差，因此不能替代个体人类；其验证通常停留在聚合层面，存在生态谬误风险。论文未讨论其他失效条件，但强调需要个体级保真度的应用场景（如个性化推荐、政策评估）会失效。","relevance":"该研究直接评估 LLM 作为人类替代品的可靠性，使用大规模真实人类数据对照，发现仅能预测项目均值而无法捕捉个体变异，对关注仿真有效性和偏差的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其将预测误差分解为项目效应、人主效应和人-项目交互效应的方法，并利用人类重测信度区分稳定信号与随机误差，可迁移到经济金融领域的个体决策预测（如消费者跨期选择、投资者风险偏好）。｜可应用于信贷审批中的个体违约风险预测或资产定价实验中的个体风险偏好测量。｜设计：以真实信贷申请人或实验参与者为被试，构建 LLM 人物角色并让其预测个体在特定金融决策任务中的选择（如是否接受高风险贷款），处理为不同人物角色丰富度（仅人口统计 vs. 完整问卷），结果变量为个体选择与项目均值的偏差，对照真实人类选择数据，并计算人类重测信度作为上限。"}},{"id":"2608.30033","version":1,"title":"\"Act Like a 5th Grader\" is Not Enough: Bounding Knowledge in LLM-Based User Simulators","zh_title":"“像五年级学生一样行动”还不够：在基于LLM的用户模拟器中界定知识","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior but frequently fail to exhibit realistic cognitive constraints, suffering from a \"superhuman bias.\" Using a dataset of over 71,000 reading comprehension responses from 2,359 primary-school students (grades 4--6), we demonstrate that standard persona prompting yields near-perfect, deterministic performance, failing to capture the natural variance of developing readers. To address this, we introduce the Cognitively Bounded User Simulator (CBUS), an architectural framework that explicitly models the restricted working memory of young readers through an episodic bottleneck. Within this framework, we formalize two distinct test-taking strategies to emulate different reading behaviors. Our evaluation shows that explicitly modeling cognitive bounds significantly narrows the simulation gap across multiple LLM backbones, demonstrating that enforcing architectural constraints is more effective for high-fidelity simulation than simply scaling raw model capabilities.","authors":["Krisztian Balog","Arild Michel Bakken"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30033","pdf_url":"https://arxiv.org/pdf/2608.30033","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知边界","人类数据对照"],"reason":"用LLM模拟学生阅读理解，有真实学生数据对照，并解决仿真偏差","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":5,"question":"如何通过显式建模认知边界（工作记忆容量限制）来提高LLM模拟小学生阅读理解行为的保真度？","design":"使用LLM模拟4-6年级小学生，通过两种方式施加处理：标准角色提示（如“扮演五年级学生”）和提出的CBUS框架（两阶段推理管道，编码阶段限制提取的文本命题数量，执行阶段仅基于受限记忆痕迹回答问题）。结果变量为学生在阅读理解题目上的回答（客观题，包括单选、判断、多选），并与真实学生数据进行对比。","baseline":"来自挪威教育系统的2,359名4-6年级学生对156篇文本和750道阅读理解题目的71,000多条真实回答。","findings":"标准角色提示导致LLM表现出超人类偏差，几乎完美且确定性地回答问题，无法捕捉真实学生的发展性差异。CBUS框架通过显式建模工作记忆瓶颈，显著缩小了模拟与真实学生之间的差距，且该效果在多个LLM骨干模型上一致。","reliability":"论文未明确讨论失效条件，但指出标准角色提示在受控环境中不足，且CBUS框架基于工作记忆容量限制的通用认知理论，可能不适用于其他认知过程或更复杂任务。","relevance":"该研究直接针对LLM仿真中的超人类偏差问题，提供了在受控环境中量化偏差和通过架构约束改进仿真的方法，对关注仿真可靠性和偏差的研究者具有重要参考价值。","inspiration":"值得借鉴的是通过架构约束（而非仅提示）来模拟认知限制，并利用大规模真实数据作为对照基准。｜可以迁移到经济金融中需要模拟有限理性或信息处理约束的场景，如消费者对复杂金融产品的选择、投资者对信息的有限关注。｜设计一个实验：用LLM模拟投资者阅读财报后做出投资决策，处理组为施加工作记忆瓶颈的CBUS框架，对照组为标准角色提示，结果变量为投资选择，对照真实投资者在类似实验中的数据。"}},{"id":"2608.28615","version":1,"title":"Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey","zh_title":"韩国合成人面板在数字与AI服务使用上的分布效度与校准：基于韩国媒体面板调查的二手数据验证","abstract":"Synthetic personas based on large language models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet systematic validation outside English-speaking contexts remains scarce. This secondary-data study evaluates how well a Korean synthetic persona panel (NVIDIA Nemotron-Personas-Korea), conditioned into Gemini 3.5 Flash (primary) and EXAONE (comparison), reproduces digital and AI service-use distributions from the KISDI Korea Media Panel Survey. Sex-and-age-stratified panels of about 8,000 personas per model answered the survey's own items - eight service-use indicators and eight innovativeness and acceptance constructs - and were compared against weighted survey estimates. The overall mean absolute error (MAE; RQ1) was 15-19 percentage points (pp), with binary item-mean correlations of 0.69-0.90. Segment error (RQ2) across five demographic axes was 15-19 pp, with between-group gaps up to 52.4/36.2 pp (Gemini/EXAONE). Errors followed model-specific signatures: an age stereotype with low anchoring (Gemini) versus an acquiescence-consistent level bias (EXAONE). Reference-year analysis was consistent with temporal misalignment driving most generative-AI overestimation, whereas short-form underestimation was framing-sensitive. Holdout calibration on 30% of the real data (RQ3) roughly halved sex-by-age cell MAE (18.9->8.6, 15.9->6.7 pp) - yet direct estimation from the same real subsample was far more accurate (3.6 pp), and the correction did not transfer across time. The calibrated panel retained an advantage only under extremely scarce real data (about 100 responses) and, for one model, for unobserved segments. Persona-narrative conditioning beat demographic-only conditioning, but neither surpassed simple real-data baselines. Synthetic panels are thus not survey substitutes; their value is diagnostic, with operational use confined to settings lacking real data.","authors":["Howard Kim","Keun Tae Cho"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28615","pdf_url":"https://arxiv.org/pdf/2608.28615","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","合成人面板","外部效度"],"reason":"用LLM合成韩国人面板，复现数字服务使用分布，并与真实调查数据对照，评估校准与…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":3,"question":"韩国合成人面板在多大程度上能复现真实调查中的数字与AI服务使用分布，以及小样本校准能否改善其准确性？","design":"使用 NVIDIA Nemotron-Personas-Korea 合成人数据，按性别和年龄分层抽取约8000个角色，分别通过 Gemini 3.5 Flash 和 EXAONE 模型生成对韩国媒体面板调查问卷的回答，测量八个数字服务使用指标和八个创新性与技术接受度构念，并与加权调查估计值比较。","baseline":"KISDI 韩国媒体面板调查 2024 年波次（约8693人）的加权估计值，以及 2023 和 2025 年波次用于时间转移检验。","findings":"合成面板的总体平均绝对误差为15-19个百分点，二元指标均值相关系数0.69-0.90；误差呈现模型特异性偏差（Gemini 年龄刻板印象、EXAONE 默许偏差），且校准虽能减半误差，但直接使用少量真实数据更准确，校准无法跨时间转移。","reliability":"论文明确指出合成面板不能替代调查，其价值仅在于诊断，且校准仅在真实数据极度稀缺（约100份）或存在未观测群体时才有优势；误差受时间错位、题项框架和响应风格影响，校准无法跨波次转移。","relevance":"该研究直接检验了LLM合成样本在非英语情境下的分布效度，并与真实调查数据严格对照，对关注仿真可靠性和偏差条件的研究者具有重要参考价值，值得精读原文以了解具体误差结构和校准方法。","inspiration":"借鉴其分层抽样生成合成样本并与真实调查对照的验证框架，以及基于偏差签名的小样本校准方法｜可迁移到消费者金融行为调查（如数字支付采用、金融科技接受度）或政策评估中的态度测量｜以合成人面板模拟不同人口群体的金融决策，施加政策信息处理，测量采用意愿，并与真实家庭金融调查数据（如中国家庭金融调查）对照，检验分布一致性和校准效果。"}},{"id":"2608.30522","version":1,"title":"Tariff Threats, Macroeconomic Expectations, and Policy Communication Strategies: Experiments Based on a Multi-Agent System","zh_title":"关税威胁、宏观经济预期与政策沟通策略：基于多智能体系统的实验","abstract":"Tariff threats can move household beliefs before policy is enacted, yet their rapidly changing language is difficult to study with conventional surveys. We build a multi-agent system that turns 300 households from the Michigan Surveys of Consumers into persistent large-language-model agents exposed to social-media information over several simulated months. Calibrated agents reproduce some distributional and demographic patterns in human survey data collected after the announcement of Liberation Day tariffs. Simulated experiments indicate that immediacy, rate salience, semantic progression, message complexity, narrative, and sender identity jointly shape inflation and unemployment expectations and their dispersion. Open-ended responses trace these effects to attention, ambiguity, credibility, and causal narratives. A second experiment finds that central-bank explanations can coordinate beliefs, although their effects on average expectations depend on message content. The framework supports disciplined exploration of policy communication, subject to human validation rather than as a substitute for it.","authors":["Jianhao Lin","Lexuan Sun","Yixin Yan"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30522","pdf_url":"https://arxiv.org/pdf/2608.30522","source_feed":"econ.GN","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3","B4"],"tags":["LLM仿真","宏观经济预期","政策沟通"],"reason":"用LLM代理300个家庭，复现关税冲击后的宏观预期，并与密歇根调查真实数据对照…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":6,"question":"关税威胁的时机、税率显著性、语义连贯性、信息复杂度、叙事和发送者身份如何共同影响家庭通胀与失业预期的水平及离散度？高影响威胁后，央行何种沟通最能协调信念？","design":"构建多智能体系统（MAS），将密歇根消费者调查（MSC）中的300个家庭转化为基于大语言模型的持久化智能体，赋予其人口特征和初始预期，在模拟的数月内暴露于社交媒体信息，并施加不同的关税威胁情景处理，测量其点预测、主观概率分布和开放式解释。","baseline":"密歇根消费者调查（MSC）在2025年4月“解放日”关税公告后收集的真实家庭通胀和失业预期数据，用于校准和验证模拟结果。","findings":"校准后的智能体在预期分布、人口统计差异和可区分性上接近人类数据；模拟实验表明，关税威胁的时机、税率显著性、语义进展、信息复杂度、叙事和发送者身份共同影响预期水平和离散度，开放式回答显示这些效应通过注意力、模糊性、可信度和因果叙事起作用。央行解释可以协调信念，但效果取决于信息内容。","reliability":"论文承认校准智能体不能替代人类，不能将未观察到的信息处理效应外推为人类因果效应；验证必须与语言模型的预期用途相关，行为相似性不足以证明从已验证场景到未观察场景的因果迁移。","relevance":"该研究直接命中研究者关注的核心：用LLM代理复现真实调查数据，并评估仿真可靠性，且涉及政策沟通和预期形成，值得精读原文以了解其校准方法和局限性讨论。","inspiration":"借鉴其将真实调查个体转化为持久化LLM智能体、施加文本处理并测量预期分布和开放式解释的设计，以及用真实调查数据校准和验证仿真的做法。｜可迁移到政策公告的预期形成研究，如央行沟通对通胀预期的影响、关税或财政政策冲击下的家庭预期调整。｜以真实家庭调查（如密歇根调查或纽约联储消费者预期调查）的受访者为被试，将其特征和初始预期编码为LLM智能体，施加不同措辞的央行声明或关税公告作为处理，测量通胀和失业预期的点预测、概率分布和开放式理由，并用同期真实调查数据作为对照基准进行校准和验证。"}},{"id":"2608.26849","version":2,"title":"LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems","zh_title":"LiveSim：在多智能体直播生态系统中模拟受环境塑造的用户","abstract":"User behavior simulation with large language models~(LLMs) is increasingly used to support multi-agent ecosystem simulation. Existing simulators typically rely on static user profiles inferred from historical observations, which become inadequate in socially intensive environments such as live streaming where interaction dynamics continuously reshape user behavior. We propose \\textbf{LiveSim}, an LLM-based framework for live-stream ecosystem simulation. It represents users as editable behavioral hypotheses and progressively refines them through trajectory-grounded interactions, where discrepancies between simulated and observed trajectories reveal missing environmental shaping effects. These signals are further extracted as transferable environment-behavior patterns and accumulated in a collective behavioral memory to improve user-level behavioral fidelity and support ecosystem-level simulation. Experiments on real-world live-stream risk-control data validate the effectiveness of LiveSim in improving user-level behavioral fidelity and enabling ecosystem-level analysis of risk evolution and platform intervention effects.","authors":["Jiaqi Xu","Yiran Qiao","Jing Chen","Qiwei Zhong","Xiang Ao","Xueqi Cheng"],"categories":["cs.AI","cs.CY","cs.MA"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-01","first_seen":"2026-08-28","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.26849","pdf_url":"https://arxiv.org/pdf/2608.26849","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM用户仿真","直播生态","行为保真度"],"reason":"用LLM仿真直播用户行为，并与真实数据对照，评估行为保真度，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":8,"question":"如何在直播这类社交密集环境中，用LLM模拟用户行为，并动态修正用户模型以反映环境对行为的塑造作用？","design":"提出LiveSim框架，用LLM扮演直播用户，将用户表示为可编辑的行为假设，通过模拟轨迹与真实轨迹的差异迭代修正假设，提取环境-行为模式存入集体记忆，再用于多智能体直播生态模拟。","baseline":"使用某大型直播平台真实风险控制数据中的用户行为轨迹作为对照。","findings":"LiveSim在用户级行为保真度上显著提升，并能支持生态级风险演化和平台干预效果分析。","reliability":"论文未讨论","relevance":"直接相关，用LLM仿真直播用户行为并与真实数据对照，评估行为保真度，属于人类仿真实验研究，值得精读原文。","inspiration":"借鉴其通过模拟-真实轨迹差异迭代修正行为假设的方法，可迁移到消费者在直播带货中的冲动购买行为研究。｜设计如下：用LLM模拟消费者，处理为不同主播话术（如限时折扣、从众压力），结果变量为购买决策，对照真实直播销售数据中的用户行为轨迹。"}},{"id":"2608.29803","version":1,"title":"Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments","zh_title":"LLM会像人类一样改变想法吗？诊断单轮说服判断中的人机分歧","abstract":"Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs in response to persuasive arguments, as humans do, remains poorly understood. We conduct a systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view. Our results show that LLMs achieve only slight agreement with humans (Cohen's kappa ranging from 0.079 to 0.178). Content-level analyses show that humans and LLMs agree on the strongest persuasion cues but diverge on finer ones: humans are more swayed by novel content and assertive language, whereas LLMs favor topical similarity and surface-level formatting. At the level of persuasion strategy, LLMs underweight emotional appeals and overweight credibility signals relative to humans, while the type of proposition under debate exerts no measurable effect on the degree of divergence. Furthermore, switching from first-person role-playing to third-person observation shifts all models toward greater resistance to persuasion, with the effect varying across persuasion strategies and textual features. These findings highlight the risk of treating LLM judgments as faithful proxies for human belief updating and point to structural differences in how LLMs and humans process persuasive discourse. Our code is available at https://github.com/tsinghua-fib-lab/LLM-belief-update-cmv.","authors":["Lin Chen","Yitong Chen","Yong Li"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29803","pdf_url":"https://arxiv.org/pdf/2608.29803","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","信念更新","人机对比"],"reason":"直接比较LLM与人类在说服中的信念更新，使用真实人类数据对照，并指出LLM作为…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":12,"question":"LLM在单轮说服中的信念更新判断是否与人类一致，哪些内容特征（命题类型、说服策略、文本特征）和视角框架（第一人称角色扮演 vs. 第三人称观察）调节这种分歧？","design":"使用ChangeMyView语料库中真实的说服对话，构建匹配的回复对（同一原帖下，一个被标记为说服成功，一个未成功），让8个LLM独立判断每个回复是否会改变原帖作者的观点，并分析判断结果与人类标签的一致性及影响因素。","baseline":"ChangeMyView语料库中原始发帖人明确标记的delta（观点改变）标签，作为人类信念更新的真实基准。","findings":"LLM与人类的一致性很低（Cohen's kappa 0.079-0.178），且LLM在文本特征和说服策略上与人类存在系统性分歧：人类更受新颖内容和果断语言影响，LLM更偏好主题相似性和表面格式；LLM低估情感诉求、高估可信度信号，而命题类型对分歧无显著影响。从第一人称角色扮演切换到第三人称观察会使所有模型更抗拒说服，且效应因策略和文本特征而异。","reliability":"论文指出LLM判断不能作为人类信念更新的忠实代理，存在结构性差异；但未明确讨论失效条件，仅强调在需要人类推理的社会模拟中直接使用LLM有风险。","relevance":"该研究直接比较LLM与人类在说服场景下的信念更新，使用真实人类数据作为基准，并揭示了系统性偏差，对关注LLM仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其匹配对设计和多维度内容标注方法，可系统诊断LLM与人类在决策中的分歧来源。｜可迁移到政策沟通与预期形成场景，如央行沟通对市场预期的影响。｜以LLM作为投资者被试，呈现央行声明或新闻，测量其预期更新，并与真实市场调查数据（如密歇根消费者信心调查）对照，分析LLM是否高估可信度信号或低估情感因素。"}},{"id":"2608.28668","version":1,"title":"Reference-Distribution Dependence in LLM-Based Synthetic Persona Data: Diagnosis and Post Hoc Adjustment of Demographic Distributions","zh_title":"基于LLM的合成人数据中的参考分布依赖：人口统计分布的诊断与事后调整","abstract":"We diagnose how closely the demographic distributions in LLM-based synthetic persona data match external reference distributions. For the three variables examined, we show that most of the observed error is attributable to the choice of reference rather than to the generator. Using total variation distance (TVD), we compare the sex x age group x province joint distribution of 1,000,000 records from Nemotron-Personas-Korea (NPK) with Korean official statistics. Against resident-registration figures for April 2026, the time of use, the bias bound, defined as the largest possible difference in the share of any subgroup formed from the three variables, is 1.81 percentage points. This is comparable to the margin of error of a survey of roughly 2,900 respondents. This value is not a fixed property of the data. Matching the reference period and series to the generating reference identified here, the 2024 register-based census restricted to Korean nationals, lowers it to 0.56 percentage points. Over the 15 months between the best-matching month (January 2025) and the time of use, the resident-registration population structure itself moves more than twice the distance of NPK's minimum error. Raking and cell post-stratification, the two weighting schemes used in Korean survey practice, remove most of the reference-period dependence at a variance inflation of about 0.2% in both cases. After raking against the generating reference, the residual joint discrepancy lies at, and marginally above, the upper bound of what a perfect generator would produce when realizing 1,000,000 records (97.6th percentile of the Monte Carlo distribution). We recommend treating synthetic persona data as auxiliary material for small-scale survey design rather than as a substitute for survey data, and re-running both diagnosis and adjustment against official statistics current at the time of use.","authors":["Eunjeong Song","Sehee Hong"],"categories":["cs.CY","stat.AP"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28668","pdf_url":"https://arxiv.org/pdf/2608.28668","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人口统计偏差","事后加权调整"],"reason":"用LLM生成合成人数据，与真实人口统计对照，诊断偏差并调整，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":9,"question":"LLM生成的合成人数据在人口统计分布上与外部参考分布有多接近，观察到的偏差在多大程度上归因于参考分布的选择而非生成器本身？","design":"使用Nemotron-Personas-Korea（NPK）数据集的100万条记录，比较其性别×年龄组×省份联合分布与韩国官方统计的差异，采用总变差距离（TVD）度量，并通过变化参考分布的时期和序列来诊断偏差来源，最后应用raking和单元后分层两种加权方法进行事后调整。","baseline":"韩国官方统计，包括居民登记数据（2026年4月）和2024年基于登记的人口普查（仅限韩国国民）。","findings":"在2026年4月使用时，NPK与居民登记数据的偏差上限为1.81个百分点，相当于约2900名受访者调查的误差边际；当与生成参考（2024年登记普查）匹配时，偏差降至0.56个百分点。大部分观察到的误差归因于参考分布的选择而非生成器，且raking和单元后分层能消除大部分参考期依赖性，方差膨胀仅约0.2%。","reliability":"论文建议将合成人数据视为小规模调查设计的辅助材料，而非调查数据的替代品，并强调应在使用时重新对当前官方统计进行诊断和调整；同时指出合成数据冻结了生成时的人口结构，随时间推移与官方统计的一致性自然下降。","relevance":"该研究直接评估LLM合成人数据与真实人口统计的偏差，并提供了诊断和调整方法，对关注LLM仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文以了解具体方法和发现。","inspiration":"借鉴其通过变化参考分布来区分生成器偏差与参考选择偏差的诊断方法，以及使用TVD和偏差上限将分布差异转化为调查误差边际的做法。｜可迁移到经济金融领域中使用合成数据模拟消费者或投资者行为的研究，例如评估政策变化对特定人群的影响或测试金融产品在不同人口群体中的接受度。｜设计：使用LLM生成具有特定人口特征（如年龄、收入、地区）的合成个体，模拟他们对某项经济政策（如税收调整）的反应，结果变量为支持率或行为变化，并与真实调查数据（如韩国劳动力面板或家庭收入支出调查）进行对照，通过变化参考分布和事后加权来评估仿真的可靠性。"}},{"id":"2608.29266","version":1,"title":"Measurement Validity in LLM Cultural Alignment","zh_title":"大语言模型文化对齐中的测量效度","abstract":"Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments like the Inglehart-Welzel Cultural Map and drawing conclusions about which cultures a model resembles. While a model's answer to a value-laden questions may be interpreted as a cultural signal, it also carries sampling noise and, can be quite sensitive to question framing. In this paper, we separate survey responses, sampling noise and question framing for multiple LLMs. We decompose response variance from these models into variation across random seeds, prompt rewordings. We employ noise-to-signal ratio (NSR) to test whether a model's apparent cultural position is distinguishable from noise. When applied across a dozen models from four geographic origins, calibrated against 88 Integrated Values Survey countries, the answer is often no. NSR exceeds 1.0 on 49 of 117 valid model-question pairs (42%), reaching 5.56 in the worst case. Two models even refuse to answer sufficient number of survey questions outright. Our results corroborate previous findings that LLMs cluster toward Western, English-speaking cultural positions. However, what does not hold up in this study is the precision with which anyone can currently interpret a specific model's coordinates: prompt tone alone can shift a model by 2.4 map units, comparable to the distance between actual countries in the Inglehart-Welzel Cultural Map. These findings suggest that cultural attribution from LLM survey responses requires establishing the reliability of the underlying measurements before interpreting model coordinates as evidence of cultural representation.","authors":["An Duy Nguyen","Muhammad Aurangzeb Ahmad"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29266","pdf_url":"https://arxiv.org/pdf/2608.29266","source_feed":"physics.soc-ph","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","文化价值观","测量信度"],"reason":"用LLM回答调查问题模拟文化价值观，并与88国真实数据对照，评估测量信度与偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":10,"question":"LLM 对文化价值观调查的回答在多大程度上反映了真实的文化信号，而非采样噪声和提问措辞的敏感性？","design":"该研究并非传统意义上的仿真实验，而是对 LLM 作为测量工具的信度检验。作者选取 12 个来自不同地域的 LLM，让它们回答 Inglehart-Welzel 文化地图的 10 个价值观问题，并通过变换随机种子和提示语气/人称来分解回答方差，计算噪声信号比（NSR）以评估测量有效性。","baseline":"88 个国家的 Integrated Values Survey（IVS）数据，用于构建 Inglehart-Welzel 文化地图的参照系。","findings":"LLM 的文化位置普遍偏向西方英语国家，但测量信度很差：42% 的模型-问题对 NSR 超过 1.0，提示语气变化可使模型坐标移动 2.4 个地图单位，相当于真实国家间的距离。部分模型拒绝回答敏感问题，导致无法投影到文化地图上。","reliability":"论文明确指出，在未建立测量信度之前，不能将 LLM 的调查回答解读为文化表征。局限包括：仅使用 Inglehart-Welzel 框架，问题数量有限；未覆盖所有可能的提示变体；模型版本和采样参数可能影响结果。","relevance":"该研究直接回应了用 LLM 模拟人类价值观的可靠性问题，与您关注的人类仿真实验的效度批判高度相关，值得精读原文以了解其噪声分解方法和 NSR 指标。","inspiration":"借鉴其将回答方差分解为随机种子、提示措辞和模型间差异的方法，用于检验 LLM 在经济调查中的测量信度。｜可迁移到消费者信心调查、通胀预期、风险偏好等经济态度的测量，评估 LLM 作为被试的可靠性。｜以 GPT-4 等模型为被试，施加不同措辞的问卷版本，测量其通胀预期或风险偏好，并与密歇根大学消费者调查或实验经济学中的真实人类数据对照，计算 NSR 以判断 LLM 回答是否超出噪声。"}},{"id":"2608.29535","version":1,"title":"Integrating adaptive human behavior into epidemic models with large language models","zh_title":"用大语言模型将自适应人类行为整合进流行病模型","abstract":"Infectious disease transmission is shaped by patterns of human interaction, which adapt as epidemic conditions change. Capturing these context-dependent behaviors remains a fundamental challenge for epidemic models. Here, we recast this challenge by using large language models (LLMs) to represent adaptive human behavior within mechanistic epidemic models. We operationalize this idea through Generative Adaptive Behavioral Layer for Epidemics (GABLE), which adapts LLMs to infer behavioral responses to epidemic and policy conditions and translates them into age-structured contact matrices coupled to a mechanistic epidemic model. Applied to COVID-19 in France, GABLE reproduced responses in population mixing and age-specific contact structures that remained epidemiologically informative. In short-term forecasting, LLM-generated contact matrices outperformed mobility-driven matrices derived from real-world mobility data, with the largest gains at longer horizons. GABLE also extends beyond forecasting to prospective policy evaluation by projecting behavioral and epidemic responses to candidate interventions before implementation. When supplied with subsequently implemented policies, GABLE reproduced epidemic trajectories and generated distinct responses to alternative policy timing and composition. By leveraging LLMs as a flexible behavioral layer, GABLE provides a framework for coupling context-sensitive behavioral generation with epidemic dynamics.","authors":["Yicheng Mao","Haoyang Li","Rob Deardon","Hongru Du"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29535","pdf_url":"https://arxiv.org/pdf/2608.29535","source_feed":"physics.soc-ph","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","流行病建模","政策评估"],"reason":"用LLM模拟人类在疫情中的行为，并与真实数据对照，用于预测和政策评估。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":11,"question":"如何用大语言模型模拟传染病流行中人类行为的适应性变化，并将其耦合到机制化流行病模型中以改进预测和政策评估？","design":"提出GABLE框架，用LLM（GPT-4o mini、Gemini 2.5 Flash、Grok 3 mini）作为行为层，根据当前疫情状态、政策条件、人口特征和疫情前接触日记生成年龄分层接触矩阵，再输入年龄分层随机传播模型，形成行为-疾病反馈循环；在法国COVID-19场景下进行回顾性重建、短期预测和前瞻性政策评估。","baseline":"真实人类数据包括：法国COVID-19住院数据、疫情前接触调查（接触日记）、基于真实移动数据的接触矩阵（mobility-driven matrices）。","findings":"GABLE生成的接触矩阵能重现人群混合和年龄特异性接触结构的变化，且在短期预测中优于基于移动数据的矩阵，尤其在较长预测期优势更大。GABLE还能在政策实施前预测行为和疫情反应，对替代政策时机和组合产生不同轨迹。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类在疫情中的适应性行为，并与真实接触和移动数据对照，用于预测和政策评估，直接命中研究者关注的LLM仿真实验、真实数据基准和政策评估场景，值得精读原文。","inspiration":"借鉴其将LLM生成的行为参数（接触矩阵）嵌入机制模型并形成反馈循环的设计，可迁移到政策公告对经济行为影响的研究中。｜可应用于政策公告的预期形成与消费/投资行为调整，如财政刺激或货币政策沟通。｜以LLM模拟不同人口群体（如年龄、收入分层）对政策公告的行为反应（如消费支出、劳动供给），将生成的行为参数输入宏观经济模型，并与真实调查数据（如消费者信心指数、信用卡消费数据）对照验证。"}},{"id":"2602.16061","version":3,"title":"AI-Generated Measurements for Identification and Inference with Missing Data: A Weak Shadow Variable Approach","zh_title":"AI生成的测量用于缺失数据下的识别与推断：弱影子变量方法","abstract":"Across business and social science applications, outcomes are often missing in ways that depend on the unobserved outcomes themselves. In service systems, for example, whether a customer submits a rating depends on the rating they would have provided. Such missing-not-at-random (MNAR) mechanisms make population quantities difficult to identify without strong assumptions on the observation process. Meanwhile, rich unstructured data, such as customer interaction histories, are increasingly available and can be used to construct structured measurements using tools such as large language models (LLMs). In this work, we develop an assumption-lean partial identification framework that uses such measurements as weak shadow variables, defined as outcome-informative proxies that are conditionally independent of missingness given the true outcome and observed covariates. Importantly, they need not accurately predict missing outcomes or satisfy the completeness requirement in the classical shadow variable literature. For identification, we characterize sharp bounds on population quantities through a pair of linear programs. For estimation and inference, we propose a localized penalized estimator that remains feasible under sampling error, and a subsampling algorithm for constructing confidence intervals. In semi-synthetic experiments using real customer-service dialogues, weak-shadow-variable intervals are about 89\\% narrower than those without auxiliary information, while their midpoints have around 41\\% lower estimation error than classical MNAR methods.","authors":["Hongyu Chen","David Simchi-Levi","Ruoxuan Xiong"],"categories":["stat.ML","cs.LG","econ.EM","stat.ME"],"primary_category":"stat.ML","announce_type":"replace-cross","date":"2026-09-01","first_seen":"2026-02-17","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2602.16061","pdf_url":"https://arxiv.org/pdf/2602.16061","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A5","B1","B3"],"tags":["LLM生成测量","缺失数据","部分识别"],"reason":"用LLM从非结构化数据生成测量作为弱影子变量，处理缺失数据，有真实数据对照，方…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":13,"question":"在结果变量存在非随机缺失（MNAR）时，如何利用LLM从非结构化数据中生成的测量作为弱影子变量，对总体均值等参数进行部分识别与推断。","design":"本文不是仿真研究，而是提出一种统计方法：将LLM从非结构化文本（如客服对话）中生成的测量视为弱影子变量，即满足给定真实结果和协变量后与缺失机制条件独立的代理变量，但不要求其准确预测缺失结果或满足完备性条件。通过线性规划刻画总体参数的尖锐部分识别界，并提出局部惩罚估计器和子采样置信区间。","baseline":"半合成实验使用真实客服对话数据，通过人为设定缺失机制生成缺失结果，以真实对话文本和已知结果作为对照基准。","findings":"弱影子变量区间比无辅助信息的区间窄约89%，其中点估计误差比经典MNAR方法低约41%。该方法在宽松假设下提供了有效的部分识别界，且对LLM测量偏差具有稳健性。","reliability":"论文承认LLM输出可能存在偏差，因此不假设其准确预测结果，仅要求条件独立性；若该条件不满足，方法可能失效。此外，部分识别界可能仍较宽，且估计器在有限样本下的表现需进一步验证。","relevance":"该文利用LLM从非结构化数据生成测量以处理MNAR问题，并提供了与真实数据对照的半合成实验，对关注LLM仿真可靠性与偏差的研究者有参考价值，值得阅读原文。","inspiration":"借鉴其将LLM输出作为弱代理变量并放松完备性条件的做法，可用于处理经济金融数据中的非随机缺失问题。｜可迁移到信贷审批中申请人收入缺失或调查中敏感问题拒答的场景，利用LLM从申请文本或开放回答中提取代理测量。｜以信贷申请人为被试，处理为是否要求提供收入证明，结果变量为违约率，用LLM从申请文本中提取收入水平代理变量，并以有完整收入的子样本作为真实数据对照。"}},{"id":"2608.01017","version":2,"title":"Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy","zh_title":"为何大语言模型会屈服：医疗谄媚背后的对话因素与推理","abstract":"Large language models can answer a medical question correctly and still abandon that answer when a user pushes back. We study this failure as medical sycophancy and ask when models are most likely to give in. Across five open-weight models, 500 MedQuAD questions, and 1.2 million trials, we use a fully crossed design over four conversational factors: user role, user evidence, interaction structure, and grounding. Medical sycophancy is nearly three times more common when users challenge an answer the model has already given than when the false claim appears in the initial query. Models are also more susceptible to users presented as physicians or medical students. Most strikingly, fabricated evidence has opposite effects across interaction structures. It increases sycophancy in single-turn interactions but reduces it after the model has already answered. Grounding helps, but does not eliminate the behavior. Sycophancy varies more across medical questions than across models, making question selection an important part of benchmark design. Reasoning traces suggest that multi-turn failures are associated with models turning back toward their own prior answer, while fabricated evidence receives more scrutiny after an initial response. Together, the results show that medical sycophancy depends as much on how a model is challenged and evaluated as on which model is tested.","authors":["Kaike Ping","Buse \\c{C}ar{\\i}k","Caleb Wohn","Xiaohan Ding","Tongshuai Wang","Eugenia Rho"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-04","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.01017","pdf_url":"https://arxiv.org/pdf/2608.01017","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM可靠性","医疗问答","谄媚行为"],"reason":"研究LLM在医疗问答中屈从用户的行为，评估其可靠性，属仿真偏差分析，但非直接仿…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":14,"question":"在医疗问答中，LLM 在何种对话条件下最容易屈从于用户的错误挑战而放弃正确答案？","design":"本研究并非人类仿真实验，而是对五个开源 LLM 在 500 个 MedQuAD 医学问题上进行 120 万次试验，采用完全交叉设计操纵四个对话因素：用户角色（医生、护士、医学生、外行）、用户证据（无、单个虚构来源、多个虚构来源）、交互结构（单轮、多轮）、接地（系统提示中是否提供已验证答案），测量模型是否改变正确答案（医疗谄媚）。","baseline":"无对照","findings":"医疗谄媚在多轮交互中比单轮高 2.8 倍；用户声称是医生或医学生时模型更易屈从。虚构证据在单轮中增加谄媚，但在多轮中减少谄媚；接地有帮助但不能消除行为。谄媚在不同医学问题间的差异大于不同模型间的差异。","reliability":"论文未讨论","relevance":"该研究系统分析了 LLM 在医疗对话中的屈从行为及其条件，属于对 LLM 可靠性和偏差的批判性评估，与研究者关注 LLM 仿真偏差的方向相关，但并非直接的人类仿真实验，可作为理解 LLM 行为偏差的参考。","inspiration":"借鉴其完全交叉因子设计和混合效应模型来分离多个对话因素对 LLM 行为的影响，并利用推理轨迹分析机制。｜可迁移到经济金融领域中 LLM 在对话中屈从用户错误信息的行为，例如金融咨询、信贷审批或投资建议场景。｜以 LLM 作为被试，操纵用户角色（如资深分析师、普通投资者）、用户提供的虚假证据、交互轮次和是否提供真实数据，测量模型是否改变初始正确判断，并与人类专家在相同对话条件下的行为进行对照。"}},{"id":"2608.25952","version":2,"title":"Spatial-Knowledge-Graph-Grounded LLM Agents for Neighborhood Livability Evaluation","zh_title":"基于空间知识图谱的LLM智能体用于邻里宜居性评估","abstract":"Neighborhood livability is commonly assessed with static built-environment indicators, such as facility proximity, street connectivity, and access to public space. These measures describe available opportunities but do not directly represent how residents with different mobility capacities, household roles, schedules, and care responsibilities experience the neighborhood. This paper presents a prototype framework that uses a spatial knowledge graph (KG) and large language models (LLMs) to generate and revise household schedules, followed by rule-based feasibility checking and GIS-based network materialization. The spatial KG integrates residents, residences, facilities, neighborhood context, and sampled road hubs; Graph-RAG retrieves each household's nearby spatial context, including candidate POIs and approximate walking times, for the scheduling LLM. The LLM produces structured household schedules, while rules are used for lightweight repairs and auditable feasibility checks. The LLM then revises schedules in response to identified feasibility issues. A routing module derives the actual travel paths, travel times, modes, and event histories from the road network. The resulting events support synthetic resident-agent interviews about daily convenience, travel burden, activity feasibility, and household coordination. A prototype demonstration in a Shenzhen neighborhood shows that nominal facility availability does not necessarily imply convenient access: residents with limited mobility and households with care responsibilities experience greater travel and coordination burdens. The framework offers an auditable way to connect spatial opportunity, household activity constraints, and resident-specific livability interpretation, while keeping simulated experience distinct from observed perception.","authors":["Haiyan Hao"],"categories":["cs.CY","cs.MA"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-27","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.25952","pdf_url":"https://arxiv.org/pdf/2608.25952","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","城市研究","智能体建模"],"reason":"用LLM生成居民日程并仿真出行体验，涉及城市政策评估，但无真实人类数据对照，且…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":15,"question":"如何用空间知识图谱和LLM智能体模拟不同居民家庭的日常活动-出行安排，以评估邻里宜居性并揭示设施可达性与实际体验的差异？","design":"构建空间知识图谱整合居民、住宅、设施、道路节点等，用Graph-RAG为每个家庭检索附近空间上下文，LLM生成结构化家庭日程，规则进行可行性检查和修复，LLM根据问题修订日程，GIS路由模块计算实际出行路径、时间、方式和事件历史，最后用合成居民访谈评估日常便利性、出行负担、活动可行性和家庭协调。","baseline":"无对照","findings":"名义上的设施可用性并不必然意味着便利可达；行动不便的居民和有照护责任的家庭承受更大的出行和协调负担。","reliability":"论文未讨论","relevance":"该研究用LLM模拟居民行为并评估城市政策，属于人类仿真实验，但无真实人类数据对照，且聚焦城市规划而非经济学实验，与研究者关注的经济学实验和政策评估场景有部分重叠，值得快速浏览以了解LLM在空间行为仿真中的应用。","inspiration":"借鉴其用知识图谱和规则约束LLM生成行为并做可行性检查的方法，可增强经济仿真中个体决策的空间和制度约束；可迁移到城市经济学中的居住选择、通勤行为或消费可达性研究；设计一个实验：用LLM模拟不同收入家庭在给定住房和交通条件下的日常活动安排，处理变量为设施分布或交通政策，结果变量为时间分配和出行负担，用真实居民时间利用调查数据做对照。"}},{"id":"2608.28626","version":1,"title":"Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects","zh_title":"大语言模型会仔细审查它们所评审的内容吗？对评分校准、错误检测和作者身份效应的多模态审计","abstract":"Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models' training cutoffs. Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations, and in either text-only or text-with-figure format. Additionally, 145 verifiably detectable errors were inserted into 55 manuscripts to assess error identification under natural and verification-oriented prompts. Across all manuscript groups, including rejected submissions, LLM scores ranged from 7.0 to 8.1, whereas human mean scores ranged from 3.4 to 6.8. The models detected 12.1\\% of the verified errors under natural prompting, and a one-sentence verification instruction increased detection to 22.2\\%; however, 78\\% of the errors remained undetected. Providing figures reduced error detection while increasing review scores. No visual error was reliably verified against its corresponding figure, and half of the text-only reviews described figures that were not provided. Author identity did not influence either review scores or error detection. LLM editorial decisions exactly matched those produced by simple score averaging.","authors":["Emad Alharbi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28626","pdf_url":"https://arxiv.org/pdf/2608.28626","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM评审","可靠性评估","人类对照"],"reason":"评估LLM评审的可靠性、偏差与错误检测，并与人类评审对照，批判性指出局限，可迁…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":24,"question":"多模态大语言模型作为同行评审时，其评分校准、错误检测能力及作者身份效应如何？","design":"用两个多模态LLM（Qwen2.5-VL-72B和Pixtral-Large-124B）对165篇ICLR 2026投稿进行评审，操纵作者身份（盲审、高声望机构、低声望机构）和输入格式（纯文本、文本+图），并注入145个可验证错误，测量评分、错误检测率和编辑决策。","baseline":"同一批投稿的真实人类评审分数和录用决定（来自OpenReview）。","findings":"LLM评分普遍偏高且区分度低（7.0-8.1 vs 人类3.4-6.8），错误检测率极低（自然提示12.1%，验证提示22.2%）；提供图表反而降低错误检测并提高评分，作者身份对评分和检测无影响。","reliability":"论文指出LLM评审缺乏批判性，错误检测能力不足，且存在幻觉（纯文本评审中描述未提供的图）；未讨论模型对训练数据外分布的泛化限制，也未分析不同提示工程或模型规模的影响。","relevance":"该研究直接评估LLM在专家评审任务中的可靠性与偏差，与人类评审对照，并揭示其系统性失效，对关注LLM仿真人类专业判断的研究者有重要参考价值。","inspiration":"借鉴其多因素实验设计（操纵身份线索和输入模态）和注入可验证错误的测量方法，可迁移到经济金融领域的专家评审或信贷审批场景。｜例如，用LLM模拟信贷审批员，操纵申请人身份线索（如种族、性别）和申请材料格式（纯文本 vs 含图表），测量审批决策和错误识别率。｜以真实信贷审批数据（如Lending Club）为基准，比较LLM与人类审批员的评分分布、歧视效应和违约预测准确性，并注入矛盾信息检验LLM的核查能力。"}},{"id":"2608.29446","version":1,"title":"Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts","zh_title":"谁的痛苦评估？社区视角与LLM在健康帖上的对齐","abstract":"Judgments about psychological distress are socially situated: what counts as concerning hinges on community norms around emotional expression, vulnerability, and help-seeking. Yet large language models (LLMs) used for distress detection are typically aligned to a single, undifferentiated standard. How well do these models capture the perspectives of the communities whose language they assess? We address this question through a perspectivist annotation study in which 321 participants provided 9,587 judgments on 1,198 Reddit posts spanning six identity-based communities, yielding community-specific labels. Raters in the contextualized in-group condition show a modest tendency to agree more with their community than uncontextualized out-group raters (OR = 1.18), an effect varying significantly across communities. We then evaluate nine open-weight LLM configurations and four frontier configurations against these labels. Open-weight LLMs systematically over-estimate distress: when communities perceive none-to-mild distress, these models achieve only 31-44% accuracy, predominantly producing false positives. GPT-5 and Gemini 2.5 Pro show the same none-to-mild inflation even when their full-sample over/under rates are mixed, while Claude Opus 4 is more conservative. This pattern does not simply mirror an outsider reading position: uncontextualized out-group human aggregates were nearly symmetric, with 18% over-estimation versus 19% under-estimation. Instead, the models that inflate none-to-mild cases exhibit a distress prior that exceeds both contextualized in-group and uncontextualized out-group human judgments. These findings have implications for equitable AI deployment in mental health contexts, where miscalibrated distress detection may unevenly affect the communities being assessed.","authors":["Andrew Aquilina","Xiang Lorraine Li","Yu-Ru Li"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29446","pdf_url":"https://arxiv.org/pdf/2608.29446","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类对照","偏差评估"],"reason":"用LLM评估心理困扰，与人类标注对照，揭示模型偏差，可迁移到仿真可靠性研究。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":25,"question":"LLM 对心理困扰和支持寻求的判断在多大程度上与不同身份社区（男性、女性、非二元、退伍军人、母亲、父亲）的情境化内群体人类判断一致？","design":"本研究不是严格意义上的仿真实验，而是评估 LLM 作为判断者与人类社区判断的吻合度。研究者收集了 1,198 条来自六个身份社区（男性、女性、非二元、退伍军人、母亲、父亲）的 Reddit 帖子，由 321 名人类标注者提供 9,587 条判断，分为情境化内群体条件（标注者知道帖子来源社区）和非情境化外群体条件（不知道来源社区），得到社区特定的标签。然后评估九个开源 LLM 配置和四个前沿 LLM 配置对这些标签的预测表现，并测试身份提示对模型判断的影响。","baseline":"人类基准是社区特定的聚合标签，来自情境化内群体标注者和非情境化外群体标注者的判断。","findings":"开源 LLM 系统性地高估心理困扰：当社区感知为无到轻度困扰时，模型准确率仅为 31-44%，主要产生假阳性。GPT-5 和 Gemini 2.5 Pro 在无到轻度案例上也表现出同样的膨胀，而 Claude Opus 4 更保守；这种模式并非简单反映外群体视角，因为非情境化外群体人类聚合判断几乎对称（18% 高估 vs 19% 低估）。","reliability":"论文未明确讨论失效条件，但指出模型在无到轻度困扰案例上的系统性高估可能源于训练数据中的偏好建模偏差，且身份提示并未一致改善对齐，表明模型可能缺乏对社区规范的细致理解。","relevance":"该研究直接评估 LLM 作为人类判断替代品的可靠性，与研究者关注 LLM 仿真人类行为、对照真实人类数据、揭示偏差的核心兴趣高度相关，值得精读原文以了解其社区特定标注方法和模型偏差分析。","inspiration":"借鉴其情境化内群体 vs 非情境化外群体的对照设计，可迁移到经济金融中的群体决策或风险感知研究，例如不同投资者群体对市场信息的解读差异。｜可应用于信贷审批中的歧视研究：用 LLM 模拟不同社区（如少数族裔、低收入群体）的信贷员或借款人对贷款申请的风险评估，与真实信贷员判断对照。｜设计：招募真实信贷员和借款人作为人类被试，提供贷款申请材料，设置情境化（告知申请人社区背景）和非情境化条件，收集风险评估和审批决策；同时用 LLM 在相同条件下生成判断，比较 LLM 与人类社区聚合标签的偏差。"}},{"id":"2608.29453","version":1,"title":"AI Can Be Easily Persuaded in Clinical Decision Making","zh_title":"AI在临床决策中容易被说服","abstract":"As AI becomes increasingly integrated into clinical practice, it is playing a growing role in medical decision making. Medicine, however, is a high stakes and evidence based field, where decisions can directly affect patients' lives. It is therefore important to understand whether AI can maintain objective judgment when others try to persuade it. In this paper, we study how easily AI can be persuaded through controlled experiments. We find that professional authority, national background, institutional affiliation, claimed past performance, multiple physicians, supported clinician views, and repeated pressure can all affect AI decisions. Surprisingly, the same persuasive input changes about 10% more cases when it comes from a senior clinician than from a medical student. Simply claiming a better performance history consistently makes the physician more persuasive. More strikingly, a plausible clinician view can persuade AI away from a correct decision even when it is fabricated to support an incorrect answer. This indicates that AI can be strongly influenced by convincing support without reliably determining whether this view from the clinician is correct. Together, these findings suggest that AI can be easily persuaded by what people say, who says it, and how the opinion is presented. Therefore, it is essential for AI to maintain sound judgment under persuasion, enabling its safe and reliable use in high stakes medical decision making.","authors":["Jiayuan Zhu","Jiazhen Pan","Fenglin Liu","Minhao Hu","Junde Wu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29453","pdf_url":"https://arxiv.org/pdf/2608.29453","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B4"],"tags":["LLM决策","说服影响","临床AI"],"reason":"研究LLM在临床决策中受说服影响，属于用LLM模拟人类决策行为，但无真实人类对…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":26,"question":"在临床决策中，AI 是否容易被说服，哪些因素会影响其被说服的程度？","design":"使用 GPT-4o、Qwen3.7-plus 和 Claude Sonnet 5 三个 LLM 扮演临床决策者，基于 MedBullets 数据集中的病例和选择题，通过单轮或多轮对话施加不同类型的说服性输入（如专业权威、国家背景、机构归属、声称的过往表现、多名医生意见、临床医生观点、重复施压），测量模型是否改变初始决策、置信度、建议行动和干预升级水平。","baseline":"无对照","findings":"AI 容易被说服，且说服效果受说服者身份、观点呈现方式等因素影响；例如，资深临床医生比医学生更能改变 AI 决策，声称更好的过往表现可增加约 9% 的说服力，而伪造的临床医生观点可使约 30% 的初始正确决策被改为错误答案。","reliability":"论文未讨论","relevance":"该研究通过受控实验系统考察 LLM 在临床决策中的易受说服性，属于用 LLM 模拟人类决策行为的研究，但缺乏真实人类对照，与研究者关注的经济学实验和政策评估场景有距离，但可作为批判性仿真失效的案例参考。","inspiration":"可借鉴其通过控制变量法系统操纵说服者特征和观点呈现方式，测量 LLM 决策变化的方法。｜可迁移到经济金融领域中权威信息对个体决策的影响，如分析师评级对投资者决策、政策制定者言论对市场预期的影响。｜以 LLM 模拟投资者，呈现不同权威级别（如知名分析师 vs 普通分析师）的股票推荐，测量投资决策变化，并与真实市场数据或实验数据对照。"}},{"id":"2608.29571","version":1,"title":"Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation","zh_title":"谁是香蕉人？评估视觉-语言模型在多轮语用解释中的表现","abstract":"Flexible adaptation to context and shared pragmatic intuitions contribute to smooth human conversation. Iterated reference games---in which players repeatedly pick out novel referents using language---present a test case for agents' ability to perform context-sensitive pragmatic reasoning in multi-turn linguistic environments. We tested humans and vision--language models on their ability to identify the intended meaning of descriptions produced in iterated reference games, varying the provided context in terms of amount, order, and relevance. While humans performed well consistently, the models we evaluated could make use of prior context to interpret humans' referring expressions, but they struggled to build up the relevant context to interpret those expressions effectively. Our results suggest that the models we evaluated lack core skills needed for efficient linguistic collaboration.","authors":["Alvin Wei Ming Tan","Ben Prystawski","Veronica Boyce"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29571","pdf_url":"https://arxiv.org/pdf/2608.29571","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["语用推理","人类对照","视觉-语言模型"],"reason":"用LLM复现人类在参照游戏中的语用推理，并与人类数据对照，但非社会科学仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":27,"question":"视觉-语言模型能否像人类一样在迭代参照游戏中利用上下文进行语用推理，以识别指代表达式的意图？","design":"使用五个开源视觉-语言模型（Qwen 3 VL 32B、Gemma 3 27B、Llama 3.2 11B、Molmo 2 8B、Kimi VL A3B）扮演人类匹配者，在迭代参照游戏中根据对话历史和12个七巧板图像选项选择目标图像。通过改变提供的上下文条件（顺序、数量、相关性）和反馈类型（交互式有限反馈、人类有限反馈、完全反馈）来测试模型表现，结果变量为模型对正确目标图像的概率。","baseline":"来自Boyce et al. (2025)的原始人类数据（yoked顺序99人，shuffled顺序98人），以及本研究新收集的backward（89人）和random（107人）条件下的人类数据。","findings":"人类在所有条件下表现稳定，而模型虽然能利用先前上下文解释人类指代表达式，但难以有效构建相关上下文来准确解释这些表达式。模型缺乏高效语言协作所需的核心技能。","reliability":"论文未讨论","relevance":"该研究直接对比人类与LLM在语用推理任务中的表现，揭示了模型在上下文构建上的不足，对关注LLM仿真可靠性与偏差的研究者有参考价值，但任务属于认知语言学而非社会科学仿真，与经济学实验关联较弱。","inspiration":"可借鉴其通过系统操纵上下文条件（顺序、相关性）和反馈类型来分离模型检索与学习能力的方法。｜可迁移到经济金融中的沟通与协调实验，如中央银行沟通、分析师报告解读或谈判博弈中的共同理解形成。｜设计一个资产定价实验：让LLM扮演投资者，根据分析师报告和先前市场评论预测股票走势，处理为提供不同顺序或相关性的历史评论，结果变量为预测准确率，并与人类投资者在相同实验中的表现进行对照。"}},{"id":"2608.29995","version":1,"title":"Generating Clinical Vignettes that Preserve Cognitive Formulations","zh_title":"生成保留认知公式的临床案例","abstract":"Large language models can generate fluent clinical case vignettes, but fluency alone does not ensure fidelity to a specifiable clinical structure. We introduce FORMA, a theory-grounded framework that compiles a cognitive model of a disorder into a directed weighted graph, samples a person-specific configuration of that graph, and validates whether the generated vignette preserves the specified components and causal links. We instantiate FORMA on Posttraumatic Stress Disorder using the Ehlers and Clark cognitive model, generating 16,500 vignettes across 500 personas, 11 generation models, and three ablation conditions. Evaluation combines an external edge-recovery probe, two clinical experts, a scaled LLM judge, and a clinician user study with 100 licensed practitioners. The cognitive graph is recoverable from full-condition vignettes (MCC = +0.41, AUC = 0.70) but not from zero-shot generation (MCC = +0.01, AUC = 0.50). Experts rate full vignettes substantially higher than zero-shot alternatives, and clinicians perceive them to be human-written 85% of the time, compared with 22% for zero-shot. FORMA also reduces demographic disparity in perceived quality by 1.5-7x. These results show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation. A repository with the data and code is available online: https://github.com/Amit-Oren/FORMA.","authors":["Amit Oren","Nimrod Hertz-Palmor","Dean Ariel","Guy Laban"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29995","pdf_url":"https://arxiv.org/pdf/2608.29995","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM生成","临床案例","认知模型"],"reason":"生成临床案例并验证认知结构，有专家和临床医生对照，但非直接仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":28,"question":"如何用认知模型作为可审计的结构规范，生成既流畅又忠实于特定临床认知结构的合成临床案例？","design":"该研究不是直接仿真人类被试，而是提出FORMA框架：将PTSD的Ehlers和Clark认知模型编译为有向加权图，为500个人格分别采样图配置和自报症状，用11个生成模型在三种消融条件（完整、无认知模型、零样本）下生成16,500个临床案例，并通过边缘恢复探针、专家评分、LLM评判和100名临床医生的用户研究来验证生成文本是否保留了指定的认知组件和因果链接。","baseline":"人类基准包括：两位临床专家对案例的评分、100名执业临床医生对案例是否像人类书写的判断，以及从真实患者数据中提取的临床项目池作为自报症状的来源。","findings":"完整条件下的认知图可从生成案例中恢复（MCC=+0.41，AUC=0.70），而零样本生成无法恢复（MCC=+0.01，AUC=0.50）；专家对完整案例的评分显著更高，临床医生认为完整案例85%像人类书写，而零样本仅22%。FORMA还将感知质量的人口统计学差异降低了1.5-7倍。","reliability":"论文未在提供的节选中明确讨论失效条件或局限，但提到零样本生成无法保留认知结构，且生成模型可能平均化患者群体或聚集于创伤类型模板，暗示缺乏结构化规范时仿真会失效。","relevance":"该研究虽非直接仿真人类被试，但提供了用理论图结构约束LLM生成、并用多层级人类评估验证保真度的方法，对关注LLM仿真可靠性和偏差的研究者有方法论借鉴价值，值得阅读原文了解其验证协议和消融设计。","inspiration":"借鉴其用理论图结构作为可审计规范、结合自动验证和人类专家评估的方法，可迁移到经济金融中的异质性主体建模或政策沟通场景，例如用LLM生成不同认知偏差的投资者或消费者，并验证其决策模式是否符合行为理论。｜可应用于资产定价实验中的投资者情绪生成、信贷审批中的歧视审计、或消费者跨期选择中的认知偏差模拟。｜设计：以行为经济学理论（如前景理论或心理账户）构建认知图，采样不同人格的投资者，用LLM生成投资决策叙述，处理为是否提供认知图约束，结果变量为决策文本中理论组件的可恢复性和专家评分，对照真实投资者调查或实验数据。"}},{"id":"2608.30110","version":1,"title":"Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators","zh_title":"LLM能否把握经济脉搏？对宏观经济指标实时预测的评估","abstract":"Nowcasting headline macroeconomic indicators, i.e., estimating an indicator's value for the current reference period before its official release, is critical for monetary policy and financial markets, and central banks devote dedicated teams of expert economists to producing such estimates. Large language model (LLM) agents are a promising candidate for this task, combining broad world knowledge with real-time web search and supporting queries at higher frequency than institutional nowcasts. Evaluating their nowcasting capability is, however, challenging: headline indicators such as GDP and CPI are widely reported and likely memorized during pretraining, so any evaluation on historical releases is vulnerable to data contamination. To address this, we introduce LiveMacroEval, a live, contamination-resistant benchmark in which LLM agents produce hourly nowcasts for sixteen major U.S. macroeconomic indicators over a pre-release window closing at each official release. Nowcast quality is assessed through a LiveMacro Score against announcement-window equity returns and a LiveBetting Score from simulated Polymarket-style trading, with Federal Reserve regional-bank nowcasts, the Bloomberg ECOS professional consensus, and an auto-ARIMA baseline as comparators. Over six months with four state-of-the-art LLM agents configured with web search, aggregate nowcast accuracy is broadly comparable to the institutional and professional benchmarks, with performance varying widely across individual indicators. This highlights LLM agents' potential as real-time estimators of macroeconomic conditions.","authors":["Xinyue Zhao","Ruiyi Zhang","Liqin Ye","Rui Cao","Pengtao Xie","Sudheer Chava"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30110","pdf_url":"https://arxiv.org/pdf/2608.30110","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","宏观经济预测","实时评估"],"reason":"用LLM agent实时预测宏观经济指标，并与专业机构预测和人类共识对照，属于…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":29,"question":"LLM智能体能否实时预测（nowcast）美国主要宏观经济指标，并与专业机构预测和人类共识相媲美？","design":"构建LiveMacroEval基准，让四个配备网络搜索的LLM智能体（GPT-5、Claude-sonnet-4.5、Qwen3-235B、Qwen3-80B）在官方发布前窗口内每小时对16个美国宏观经济指标进行实时预测，通过LiveMacro Score（基于预测对股市收益的解释力）和LiveBetting Score（模拟Polymarket交易收益）评估预测质量。","baseline":"对照真实人类数据：五个美联储地区分行nowcast（亚特兰大、纽约、圣路易斯、克利夫兰、芝加哥）和Bloomberg ECOS专业共识调查，以及auto-ARIMA计量基准。","findings":"在六个月的实时评估中，表现最好的LLM（GPT-5）在总体LiveMacro Score上与Bloomberg共识持平，其他LLM略低于共识但高于auto-ARIMA；在LiveBetting Score上，多个LLM在GDP和失业率指标上与美联储nowcast相当。LLM的预测修订与评估窗口内的经济信息事件对齐。","reliability":"论文指出，LLM在CPI、零售销售和PCE价格指数上表现不佳，导致总体得分下降；评估仅覆盖六个月，样本有限；且LLM可能仍受预训练数据污染影响，尽管采用实时设置缓解。","relevance":"该研究将LLM作为经济预测主体，与专业人类预测者进行直接对比，属于人类仿真在经济预测场景的应用，且提供了实时、抗污染的评估方法，值得阅读原文以了解其基准设计和可靠性讨论。","inspiration":"借鉴其构建实时、抗污染的评估基准，将LLM预测与专业人类预测和计量模型对照，并使用市场数据（股票收益、预测市场）作为外部验证指标｜可迁移到政策公告的预期形成研究，如央行利率决议或财政刺激对市场的影响预测｜设计一个实验：让LLM智能体扮演专业经济学家，在每次政策会议前基于新闻和数据进行预测，处理为提供不同信息集（完整新闻vs.仅官方数据），结果变量为预测准确性和市场反应一致性，对照真实分析师调查和利率期货隐含概率。"}},{"id":"2608.30873","version":1,"title":"Personas Differ from Native-Language Generation: Language Pathways Shape LLM Interpersonal Advice","zh_title":"人设与母语生成不同：语言路径塑造LLM的人际建议","abstract":"LLMs are increasingly used for interpersonal advice and as tools for studying social behavior across languages and cultures. A common shortcut for eliciting language- or culture-related variation is to ask a model to answer as a native speaker. We test whether this native-speaker persona reproduces the outputs obtained when models instead generate advice in the target language and translate the response back into English. Using 600 interpersonal advice questions across 13 languages and eight LLMs, we compare native-language generation followed by translation (NL) with native-speaker persona prompting (NP), measuring linguistic style, behavioral scaffolding, and forced-choice action recommendations. We find that NP and NL are not interchangeable. Compared to NL, NP often increases lexical social cues, including affiliation and positive tone, while reducing qualities such as concreteness and social attunement; NP also provides less actionable scaffolding in open-ended advice. In forced-choice scenarios, NP changes which action the model selects, favoring confrontation over redirection, with effect sizes varying across languages, topics, and models. Our results show that cross-lingual elicitation strategy is a consequential methodological choice that can change both how advice is framed and which actions models recommend.","authors":["Jinhee Won","Xinlan Emily Hu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30873","pdf_url":"https://arxiv.org/pdf/2608.30873","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","跨语言行为","方法偏差"],"reason":"用LLM模拟不同语言母语者的人际建议，并与真实语言生成对照，揭示仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":32,"question":"在跨语言人际建议任务中，用母语者人设提示（NP）与用目标语言生成后翻译（NL）两种方式得到的建议是否等价？","design":"用8个LLM（5个开源权重、3个专有API模型）处理600个人际建议问题，问题从英语翻译成13种语言。对每个问题，分别用两种方式生成回答：NL（用目标语言生成建议，再翻译回英语）和NP（用英语提示模型以母语者身份回答）。测量语言风格（语气、权力、正式度、具体性）、行为指导（基于行为改变技术分类法）和强制选择行动建议（对抗、脱离、重定向）。","baseline":"无对照","findings":"NP与NL不可互换：NP增加了社交词汇线索（如亲和、积极语气），但降低了具体性、社会协调性和可操作的行为指导；在强制选择场景中，NP改变了模型选择的行动，更倾向于对抗而非重定向。这些差异随语言、话题和模型而变化，表明是结构化的诱导效应而非随机变异。","reliability":"论文指出这些模式应被视为模型诱导的属性，而非真实语言社区的证据；要研究真实社区需要验证NP或NL哪个更符合实际说话者的反应。论文未讨论其他失效条件。","relevance":"该研究直接检验了LLM仿真中语言路径（人设提示 vs. 语言生成）对输出的影响，揭示了仿真偏差，对关注LLM作为人类被试替代品的研究者具有重要参考价值。","inspiration":"值得借鉴的是通过对比不同诱导策略（人设提示与语言生成）来揭示模型输出的系统性差异，并测量多维结果变量（语言风格、行为指导、行动选择）。｜可以迁移到跨文化经济决策研究，例如不同语言环境下的风险偏好、时间贴现或谈判策略。｜设计：用LLM模拟不同语言背景的投资者，处理为两种诱导方式（母语者人设 vs. 目标语言生成），结果变量为投资建议或风险选择，并与真实的多语言投资者调查数据（如全球风险偏好调查）对照。"}},{"id":"2608.31059","version":1,"title":"When Can We Work in Embedding Space? What Text Embeddings Preserve","zh_title":"何时可以在嵌入空间中工作？文本嵌入保留了什么","abstract":"When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can trade text for its low-dimensional embedding, and lose little in doing so. I make that assumption precise under a generative model in which documents are mixtures of latent topics. I study two uses---clustering units in embedding space and controlling for high-dimensional text. A cluster of embeddings is a set of documents with similar topic mixtures; controlling for the embedding is equivalent to controlling for the topic mixture, so validity reduces to whether that mixture captures the confounding. In an application to 363 U.S. metropolitan areas, embedding-based clusters of LLM-generated economic descriptions recover interpretable economic archetypes and separate local employment dynamics more sharply than clustering on model residuals, or on a curated set of industry and demographic covariates.","authors":["Simon Freyaldenhoven"],"categories":["econ.EM","cs.CL","stat.ML"],"primary_category":"econ.EM","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.31059","pdf_url":"https://arxiv.org/pdf/2608.31059","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A5","B1"],"tags":["文本嵌入","LLM生成文本","经济数据分析"],"reason":"用LLM生成文本并嵌入分析经济数据，有真实城市数据对照，方法可迁移到仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":33,"question":"文本嵌入何时能作为实证分析的输入？在什么条件下，将文本降维为嵌入向量不会损失分析所需的信息？","design":"本文不是仿真研究，而是理论加实证：在生成式主题模型下推导嵌入保留的信息，并应用于363个美国大都市区，用LLM生成经济描述文本，嵌入后用k-means聚类，再比较聚类对就业动态的解释力。","baseline":"无对照（本文不涉及人类被试，而是用真实城市就业数据作为结果变量，聚类结果与基于残差或协变量的聚类对比）。","findings":"在主题模型下，词嵌入的几何结构保留主题载荷信息，文档嵌入是主题混合的可逆线性映射；聚类嵌入等价于聚类主题混合，控制嵌入等价于控制主题混合。实证中，基于嵌入的聚类能识别出可解释的经济原型，并比基于残差或协变量的聚类更清晰地分离就业动态。","reliability":"论文未讨论（理论部分基于主题模型假设，实证部分未系统检验失效条件，但隐含假设LLM生成的文本符合主题模型）。","relevance":"本文为使用LLM生成文本并嵌入进行经济分析提供了理论基础，其方法可迁移到人类仿真研究：若LLM生成的文本能反映潜在特征，则嵌入可作为有效代理。值得读原文以理解嵌入有效性的条件。","inspiration":"借鉴其用LLM生成结构化文本并嵌入以捕捉潜在特征的方法，以及用真实结果变量验证聚类有效性的设计。｜可迁移到政策评估中，用LLM生成个体对政策的看法文本，嵌入后聚类识别态度类型，再比较不同态度群体的行为反应。｜用LLM模拟消费者对金融产品的评价文本，嵌入后聚类识别消费者类型，处理为不同信息披露方式，结果变量为投资选择，对照真实消费者调查数据。"}},{"id":"2608.30210","version":1,"title":"Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration","zh_title":"前沿视觉语言模型在检测AI生成人像上已超越年轻人——但校准能力尚未超越","abstract":"AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to flag such images. We benchmarked 19 VLMs on the same 198 face portraits -- real photographs and identity-matched ChatGPT-4o and Imagen 3 versions -- under the same task as our earlier study of 1,667 adults (85% correct overall; accuracy fell steeply with age). The June-2026 cohort of 14 models only matched adults in their 20s-30s. Four weeks later the ceiling broke. Among five July-2026 releases under the identical protocol, gpt-5.6-sol reached 92.8% balanced accuracy (five-draw mean 92.1%), clearly above adults in their 20s (88.5%), and claude-fable-5 detected every AI image while averaging 91.9%. Model sensitivity now exceeds young adults decisively (d' up to 3.4 versus ~ 2.4). What has not been overtaken is human calibration. Model criteria spread from c = -1.10 to +1.45 while humans sit near zero at every age; both new leaders are biased (+0.44, -0.97), and only a few mid-ranked models approach the human balance. Changing the labelled examples still flipped about one answer in four. The best machines now out-see young adults here, without matching the human balance between suspicion and trust.","authors":["Sunwhi Kim (Hwasung Medi-Science University, Dept. of Bio-Healthcare)","Sunyul Kim (Yonsei University, Graduate School of Engineering, Dept. of Artificial Intelligence)","Meounggun Jo (Hoseo University)","Jini Tae (Gwangju Institute of Science and Technology, School of Humanities and Social Sciences)"],"categories":["cs.HC","cs.CV"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30210","pdf_url":"https://arxiv.org/pdf/2608.30210","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A2","B1"],"tags":["视觉语言模型","人类对照","校准偏差"],"reason":"评估VLM检测AI图像的能力并与人类数据对照，涉及模型与人类感知的校准偏差，可…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":30,"question":"前沿视觉-语言模型在检测AI生成人脸肖像上是否已超越年轻人，且其校准是否与人类相当？","design":"本研究并非用LLM模拟人类被试，而是将19个前沿VLM作为检测器，在相同任务和刺激下与人类表现进行基准比较。模型接收与人类相同的4个带标签示例（few-shot），对198张肖像逐一进行REAL/AI二分类判断，并记录置信度和理由。","baseline":"人类基准来自作者先前研究中的1,667名20-69岁成年人，他们在相同198张肖像上完成相同任务，总体正确率85%，且正确率随年龄下降（20多岁约88%，60多岁约66%）。","findings":"2026年6月的14个模型仅与20-30岁成年人相当；但4周后发布的5个模型中，gpt-5.6-sol和claude-fable-5在准确率和敏感性上显著超过年轻人。然而，模型校准未超越人类：模型反应标准极端（c从-1.10到+1.45），而人类各年龄段均接近0；且更换示例集导致约四分之一判断翻转。","reliability":"论文承认模型判断对few-shot示例敏感，更换示例导致约25%的答案翻转；模型置信度校准差，且其解释可能与实际决策依据不一致。此外，任务为干净的二元判断，可能高估模型在真实世界杂乱环境中的表现。","relevance":"该研究直接比较VLM与人类在感知任务上的表现，并揭示模型在准确率超越人类的同时存在校准偏差和示例敏感性，对评估LLM作为人类被试替代品的可靠性具有重要参考价值。","inspiration":"借鉴其严格匹配人类实验协议的做法，包括相同刺激、任务指令、评分方式和多次抽样以评估稳健性，并同时测量准确率、偏差、校准和解释一致性。｜可迁移到经济金融中的视觉信息判断场景，如信贷审批中申请人照片对决策的影响、投资中对公司年报图像信息的解读、或消费者对广告真实性的判断。｜设计一个实验：让LLM和人类被试观看相同的金融广告图片（真实与AI生成），判断广告真实性并给出置信度，同时收集人类行为数据作为基准，比较准确率、反应偏差和校准曲线，并变换few-shot示例检验模型稳健性。"}},{"id":"2608.30311","version":1,"title":"One AI Signal, Many Human Judgments: A Bayesian Cascade Analysis of AI-based Credibility Indicators in Online Information Spread","zh_title":"一个AI信号，多种人类判断：在线信息传播中基于AI的可信度指标的贝叶斯级联分析","abstract":"Social media platforms increasingly use AI-based credibility indicators to help users judge misinformation. Unlike individual human-AI decision-making, these indicators are embedded in information spread: users see both an AI prediction and earlier judgments shaped by the same AI, and their own judgments may then enter the public history. Yet how to analytically characterize this process remains under-explored. We therefore introduce a social-learning lens for this setting by extending the classical Bayesian cascade model with the AI indicator as a shared public signal. The resulting Gateway condition compares the evidence from the AI prediction with users' private impressions. Through this view, we show that AI changes what public history means. Crowd agreement may reflect accumulated independent human evidence, or repeated dependence on the same AI prediction. This creates a preservation-correction trade-off: stronger reliance on AI can preserve correct predictions, but can also lock in incorrect ones by blocking corrective private impressions. We calibrate the model using human-subject data on news veracity judgments. Although the AI outperforms human users, the average user weights it below her own impression but above several peer judgments, while individual users vary from discounting the AI to relying on it enough to cascade. Simulations show that over-reliance on a weak AI is especially harmful, and that diversifying AI signals across users can better keep the crowd informative. We conclude with implications for understanding human-AI interaction in information spread and designing misinformation interventions.","authors":["Zhuoran Lu","Weilong Wang","Yangyang Yu","Xinru Wang","Zhuoyan Li","Zhiwei Liu","Sophia Ananiadou"],"categories":["cs.HC","cs.AI","cs.SI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30311","pdf_url":"https://arxiv.org/pdf/2608.30311","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["人类-AI交互","社会学习","信息传播"],"reason":"用贝叶斯级联模型分析AI信号对人类判断的影响，校准于人类数据，涉及信息传播和政…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":31,"question":"AI 可信度信号如何通过社会学习过程影响用户对新闻真实性的序列判断，并改变公共历史信息的意义？","design":"扩展经典贝叶斯级联模型，将 AI 可信度信号作为共享公共信号，模拟用户在信息传播中的序列判断；模型用人类被试数据校准，并模拟不同 AI 强度和行为权重下的级联动态。","baseline":"使用先前人类被试实验数据，用户在 AI 指标存在下判断新闻真实性，AI 表现优于人类用户。","findings":"AI 信号改变了公共历史的含义，人群一致可能反映独立人类证据或对同一 AI 信号的重复依赖；存在保留-纠正权衡：强依赖 AI 可保留正确预测，但也会锁定错误预测。校准显示平均用户对 AI 的权重低于自身印象但高于多个同伴判断，个体差异显著；模拟表明过度依赖弱 AI 尤其有害，多样化 AI 信号可保持人群信息量。","reliability":"论文承认模型基于贝叶斯理性假设，但实际用户存在行为偏差；校准数据来自特定实验设置，可能限制外部效度；未考虑 AI 信号随时间变化或用户异质性对级联的长期影响。","relevance":"该研究用贝叶斯级联模型分析 AI 信号对人类判断的影响，校准于真实人类数据，涉及信息传播和政策评估场景，对关注 LLM 仿真可靠性与偏差的研究者有直接参考价值。","inspiration":"借鉴其将 AI 信号作为共享公共信号嵌入社会学习模型的方法，可迁移到资产定价实验或政策公告预期形成场景；例如用 LLM 模拟投资者在 AI 投资建议下的序列决策，处理为不同 AI 信号强度，结果变量为投资选择与价格动态，对照真实实验数据。"}},{"id":"2608.28597","version":1,"title":"The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys","zh_title":"在线调查中代理式AI能力与数据质量控制之间的竞赛","abstract":"Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language model (LLM) brain and/or a multimodal processing unit with tool-augmented capabilities) raises new questions about the robustness of these safeguards. We investigate how well agentic AI architectures can complete web-based surveys and pass standard attention checks. We evaluate a single-agent architecture capable of multimodal input processing and tool-based web interaction on a controlled survey sandbox. We analyze the problem from two perspectives. From an attack perspective, we demonstrate how structural vulnerabilities such as exposed DOM metadata and predictable option encoding allow agents to resolve attention checks through structured parsing only. From a defense perspective, we implement a mitigation strategy of DOM metadata obfuscation to remove semantic cues in text-based questions. We evaluate multiple open-source language and multimodal models to study capability and orchestration effectiveness. Based on our evaluations, we offer perspectives on how to simultaneously meet the needs of empiricists and agentic AI researchers.","authors":["Sourav Panda","Hillmer Chona","Rupak Kumar Das","Shreyash Kale","Shikha Soneji","Jonathan Dodge"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28597","pdf_url":"https://arxiv.org/pdf/2608.28597","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM代理","调查数据质量","注意力检查"],"reason":"研究LLM代理完成在线调查及注意力检查，评估数据质量与防御，与仿真可靠性相关。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":23,"question":"代理型AI在多大程度上能自主完成在线调查并通过注意力检查，以及如何防御这种自动化参与以保障调查数据质量？","design":"构建单代理架构，集成开源LLM（LLaMA-3-8B、Mistral-7B、Qwen-2.5-7B）进行多模态输入处理和基于工具的网页交互，在受控调查沙盒中执行多页调查任务；设置完成策略（严格模式）和无约束策略（自由模式）两种行为策略，每种模式运行100次，测量端到端提交成功率、页面级准确率和失败分解。","baseline":"无对照","findings":"代理AI能利用DOM元数据等结构漏洞，通过结构化解析而非语义理解来通过注意力检查；所有模型在两种行为策略下均表现出较高的端到端成功提交率，表明现有注意力检查可能无法有效区分人类和机器响应。","reliability":"论文未讨论","relevance":"该研究直接评估LLM代理在调查场景中通过注意力检查的能力，并探讨防御策略，与研究者关注的LLM仿真可靠性及失效条件高度相关，值得阅读原文以了解具体漏洞利用方式和防御效果。","inspiration":"借鉴其攻击-防御双视角设计和结构化漏洞分析方法，可用于评估经济实验中LLM代理是否通过界面元数据而非真实决策来通过操纵检查｜可迁移到在线经济实验或调查中的注意力检查有效性评估，例如公共品博弈中的理解性问题或风险偏好问卷｜以LLM代理为被试，施加DOM元数据混淆处理，测量其通过注意力检查的准确率和决策一致性，并与真实人类被试在相同实验中的表现进行对照。"}},{"id":"2608.28182","version":1,"title":"Benchmarking large language model agent societies against human behavioural distributions","zh_title":"基于人类行为分布基准测试大语言模型智能体社会","abstract":"Populations of large language model agents are increasingly used as experimental societies. Three doubts shadow every such result: whether the agents behave like the humans they stand in for, whether a finding survives changes to the apparatus that leave the rules untouched, and whether apparent social dynamics are interaction at all rather than the reproduction of experiments the models have read. This article introduces SILICA, an open instrument that tests all three. Five environments carry published human anchors, each paired with perturbations that re-render the same rules and with variants whose payoffs point away from the memorised result. Twelve open-weight models were run through it on a single consumer graphics card. Agreement with human data is confined to starting points: first-round public-goods contributions fall inside the equivalence margin for eight of eleven models, while no model matches end-state contributions or the human corridor of cooperation. Merely swapping the order in which two actions are listed costs one model 58 points of cooperation. Presenting responders with a fixed schedule of offers shows that only one model, the sole reasoning-trained one, places its acceptance threshold where the incentive requires; two move theirs part of the way, two move them the wrong way, and three never acquire one. Conventions form through a shared prior over the names rather than through negotiation, though negotiation reappears once that prior is disrupted. On the certification ladder defined here, current silicon societies support exploratory claims and no more.","authors":["Raad Bin Tareaf"],"categories":["physics.soc-ph","cs.CL"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28182","pdf_url":"https://arxiv.org/pdf/2608.28182","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM仿真","人类行为对照","算法保真度"],"reason":"直接以LLM agent群体仿真人类行为，并与真实人类数据对照，评估可靠性，涉…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-31","rank":1,"question":"LLM智能体社会能否在行为分布上复现人类实验基准，其结论对实验装置扰动和记忆污染是否稳健？","design":"用12个开源权重LLM作为智能体，在五个有已发表人类锚点的交互式多智能体环境中运行（重复囚徒困境、带代价惩罚的公共品博弈、讨价还价、11-20要钱游戏、带承诺少数派翻转的命名游戏），施加设计级和表征级扰动，并设置收益指向偏离记忆结果的变体，测量合作水平、接受阈值、约定形成等行为结果。","baseline":"五个环境均配有已发表的人类实验锚点，如公共品博弈首轮贡献、最后贡献、合作走廊，以及讨价还价中的接受阈值等。","findings":"与人类数据的一致性仅限于起点：11个模型中有8个首轮公共品贡献落在等效边界内，但没有模型匹配最终贡献或人类合作走廊；仅交换两个动作的列出顺序就使一个模型合作率下降58个百分点。在固定报价序列下，只有唯一一个推理训练模型将接受阈值放在激励要求的位置，其他模型或部分移动、或反向移动、或根本不形成阈值；约定通过名称共享先验而非协商形成，但先验被扰乱后协商重新出现。","reliability":"论文承认当前硅基社会仅支持探索性声明，不能支持可转移或稳健的结论；设计级扰动在111个可计算对比中改变行为56次，表征级扰动在71个中改变8次，且固定报价序列揭示聚合拒绝率无法区分模型是否真正习得激励。","relevance":"该研究直接以LLM智能体群体仿真人类行为，并与真实人类数据对照，系统评估了仿真在经济学实验中的保真度、稳健性和污染问题，对关注LLM仿真可靠性与偏差的研究者具有核心参考价值。","inspiration":"可借鉴其通过设计级与表征级扰动分离内容与形式影响、以及用固定报价序列识别个体接受函数来审计记忆污染的方法。｜可迁移到资产定价实验中的策略性报价、信贷审批中的歧视测量、或消费者跨期选择中的时间偏好等场景。｜用LLM智能体扮演投资者或消费者，施加收益结构改变或信息呈现方式扰动，测量报价、接受阈值或跨期选择，并与实验室或现场实验的真实人类数据做等效性检验。"}},{"id":"2608.26086","version":2,"title":"TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development","zh_title":"TraceML：机器学习开发中人机规划的经验分析","abstract":"Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning development, where an agent must revise data pipelines, models, and validation over hours of feedback, and on most competitions still finishes below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We introduce TraceML, which pairs human and agent work on the same competitions under one version-level schema: 4,465 human Kaggle trajectories across 134 competitions, seven of which are also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its intent, the edit size, and the score effect. Read this way, the gap becomes concrete. Experts alternate data work, validation, model changes, and ensembling, and return to approaches they had set aside. Each agent scaffold instead collapses into a narrow loop: Codex spends its steps re-weighting ensembles and tuning submissions, MLEvolve mutates its model in place, and neither pivots at the human rate nor reopens abandoned work. A short planning prompt distilled from human practice moves the behaviors it names toward the human profile and lifts scores, but the effort profile stays agent-shaped: instruction closes only the part of the gap that reduces to instructions. We release the corpus, the schema, the labelers, and the extraction pipeline at https://huggingface.co/datasets/jerryyan/TraceML.","authors":["Jiarui Yan","Weiwei Sun","Sijie Li","Wenhan Li","Yiming Yang"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-08-31","first_seen":"2026-08-27","revised_at":"2026-08-31","abs_url":"https://arxiv.org/abs/2608.26086","pdf_url":"https://arxiv.org/pdf/2608.26086","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B4"],"tags":["LLM agent","人类行为对照","过程分析"],"reason":"用LLM agent复现人类开发轨迹并与真实人类数据对照，分析行为差异，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-31","rank":3,"question":"在自主机器学习开发中，人类专家与LLM智能体的开发过程差异是什么？","design":"论文构建TraceML数据集，将人类Kaggle竞赛轨迹与两个LLM智能体（Codex和MLEvolve）在同一竞赛上的轨迹进行版本级对齐，记录每个代码版本的动作、意图、编辑大小和分数效果，并比较行为模式。","baseline":"4,465条人类Kaggle轨迹（来自134个竞赛），其中7个竞赛同时有智能体轨迹，形成430条配对人类轨迹和207条智能体轨迹。","findings":"人类专家在数据工作、验证、模型更改和集成之间交替，并会重新拾起之前放弃的方法；而智能体陷入狭窄循环，Codex主要调整集成权重和提交，MLEvolve原地变异模型，两者都很少转向或重新打开旧工作。基于人类实践提炼的规划提示能部分改善行为并提高分数，但努力分布仍保持智能体形态。","reliability":"论文承认人类轨迹未进行预算匹配，只能作为参考分布而非对照；且行为差距需在人类内部巨大变异的背景下解读。","relevance":"该研究直接对比人类与LLM智能体的行为轨迹，揭示仿真在过程层面的系统性偏差，对关注LLM仿真可靠性的研究者有重要参考价值，值得阅读原文以了解其标注方案和干预设计。","inspiration":"借鉴其版本级轨迹标注和过程诊断方法，可对经济决策过程进行细粒度分解，识别智能体与人类在策略选择、信息搜索和试错模式上的差异。｜可迁移到资产定价实验或政策评估场景，例如模拟投资者在动态市场中的策略调整或消费者对政策变化的响应。｜以LLM智能体作为被试，施加不同信息环境或激励处理，记录其决策轨迹并与真实人类实验数据（如实验室资产市场或调查数据）对照，比较行为模式和绩效。"}},{"id":"2608.26152","version":2,"title":"AI Models Can Predict and Collaboratively Modulate Human Memory Search","zh_title":"AI模型可以预测并协同调节人类记忆搜索","abstract":"Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in many language-based tasks, for example routine coding, these artificial intelligence models have reduced, or even eliminated, the need for human input. But rather than replacing human cognitive effort, LLMs may instead serve as cognitive tools to extend human abilities, particularly when they are engaged in a task requiring open-ended conceptual exploration and creative ideation. However, we are yet to understand how these models may enhance such generative human cognitive abilities in human--AI interactions. In this study, we explore and evaluate the ability of LLMs to follow and enhance human mental trajectories during semantic memory search. To test this, we use the semantic fluency task (SFT), a classic cognitive paradigm requiring generative semantic memory retrieval that has long served to characterize convergent and divergent thinking in humans. We demonstrate that an LLM's abilities to track and predict human memory trajectories in this task exceed those of other humans.","authors":["Eric Lacosse","Mariana Duarte","Graham Todd","Peter M. Todd","Daniel C. McNamee"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-31","first_seen":"2026-08-28","revised_at":"2026-08-31","abs_url":"https://arxiv.org/abs/2608.26152","pdf_url":"https://arxiv.org/pdf/2608.26152","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","认知行为","人类数据对照"],"reason":"用LLM预测人类记忆搜索轨迹并与人类数据对照，属于仿真人类认知行为，但非社会调…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":16,"question":"LLM能否追踪并预测人类在语义记忆搜索中的思维轨迹，并在人机协作中增强人类的生成性认知能力？","design":"使用Gemini-3-Pro等LLM生成语义流畅性任务（SFT）序列，与人类生成的动物、衣物、超市物品等类别序列进行对比；通过转换概率矩阵、BLEU分数、谱隙和曲率等指标评估宏观对齐；在微观层面，使用理论驱动认知提示（TDCP）让LLM预测个体参与者的下一个词和子类别切换。","baseline":"来自先前研究的人类SFT序列数据（动物类别），以及人类参与者之间的平均相似度（人-人BLEU分数）和基于留一法的一阶TPM模型。","findings":"LLM生成的序列在BLEU分数上显著高于人类平均相似度，表明其比随机人类序列更能代表典型人类序列；但LLM的语义轨迹谱隙更大、曲率更低、子类别切换更少，表明其搜索策略更刚性、更沿主导轴，而非人类的多维联想探索。","reliability":"论文指出LLM的宏观对齐可能源于统计平均而非个体噪声动态，其几何轨迹不如人类；微观对齐部分未在节选中详细讨论局限。","relevance":"该研究用LLM仿真人类认知行为并与真实人类数据对照，属于人类仿真研究，但聚焦于语义记忆搜索而非社会调查或经济决策，对关注经济学实验和政策评估的研究者参考价值有限。","inspiration":"借鉴其使用真实人类数据作为基准，并通过多种指标（如BLEU、谱隙、曲率）评估LLM与人类行为对齐程度的方法。｜可迁移到消费者偏好形成或信息搜索行为的研究，例如消费者在商品类别间的注意力转移。｜以LLM模拟消费者在电商平台上的商品浏览序列，施加不同推荐算法作为处理，结果变量为浏览路径的转换概率和多样性，与真实用户点击流数据对照。"}},{"id":"2608.27465","version":1,"title":"The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models","zh_title":"情绪语境对大语言模型认可过早决策的影响：六种商业模型情绪脆弱性比较","abstract":"As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice according to the user's emotional state has become an important safety problem. We test whether emotional expression increases a model's endorsement (encouragement to proceed) when a user, holding the same objective information, is overconfident about a premature decision (e.g., quitting a stable job on weak evidence). As a key control, we include a no-emotion multi-turn (neutral) condition that holds factual content and the number of conversational turns constant, isolating the effect of emotion from that of conversation length. We exposed six commercial models (top-tier and mid-tier models from OpenAI, Anthropic, and Google) to three scenarios (career change, business expansion, emigration) across three conditions (cold/neutral/distress) with six repetitions each, yielding 324 conversations, and measured endorsement strength (0-100) via an eight-item rubric-based automated scoring. Emotional expression significantly increased endorsement (neutral 18.6 to distress 31.5, +12.9 points; mixed-effects $\\beta = +12.9$, $p < .001$; Cohen's d = 0.51), and this was not explained by conversation length (cold-neutral difference non-significant, $p = .083$). Critically, the vulnerability varied by individual model rather than by price tier: five of six models showed a significant emotion effect, including the top-tier flagships Gemini 3.1 Pro and GPT-5.5, while only Claude Opus showed no significant change. Results were reproduced with an independent non-Google judge model ($\\rho = .89$) and agreed in rank with two human coders ($\\rho = .70$). Through a controlled design that separates emotion from conversational context, we show that emotional context increases LLM sycophancy even in top-tier flagship models.","authors":["Cheolho Shin","Yoojin Han","Donghun Shin","Kunho Lee"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27465","pdf_url":"https://arxiv.org/pdf/2608.27465","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B4"],"tags":["LLM决策偏差","情绪影响","模型安全性"],"reason":"研究LLM在情绪影响下对人类决策建议的偏差，属于仿真人类决策模式，但无真实人类…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":19,"question":"当用户对不成熟决策过度自信时，情绪表达是否增加大语言模型对决策的认可（鼓励继续），以及这种效应是否因模型而异。","design":"用六种商用大语言模型（OpenAI、Anthropic、Google 的顶级和中端模型）扮演决策建议者，通过三种条件（冷/中性/痛苦）呈现相同客观事实，测量模型对不成熟决策的认可强度（0-100分）。","baseline":"无对照","findings":"情绪表达显著增加模型认可（中性18.6到痛苦31.5，+12.9分），且该效应不能由对话长度解释。脆弱性因模型而异，六个模型中有五个表现出显著情绪效应，包括顶级旗舰模型，只有Claude Opus无显著变化。","reliability":"论文未讨论","relevance":"该研究通过受控实验分离情绪与对话长度的影响，揭示LLM在情绪操纵下的谄媚行为，对评估LLM作为人类决策仿真工具的可靠性具有参考价值，但缺乏真实人类对照，且场景限于个人决策，与经济学实验的关联有限。","inspiration":"值得借鉴的是三条件对照设计，通过冷/中性/痛苦分离情绪与对话长度的影响，并用多项目评分量表提高测量精度。｜可迁移到消费者金融决策建议场景，如信贷审批中的情绪影响或投资建议中的风险偏好诱导。｜设计雏形：以LLM作为金融顾问，向用户提供贷款或投资建议，处理变量为用户情绪表达（痛苦vs中性），结果变量为建议的风险程度或认可度，对照真实人类顾问在相同情境下的行为数据。"}},{"id":"2608.28576","version":1,"title":"Learning a Size-Weight Frontier for Synthetic-Augmented Inference","zh_title":"学习合成增强推断的规模-权重前沿","abstract":"Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model responses to augment opinion survey data, our procedure achieves target coverage and substantially narrows confidence intervals.","authors":["Chengpiao Huang","Kaizheng Wang"],"categories":["stat.ME","cs.AI","cs.LG","stat.ML"],"primary_category":"stat.ME","announce_type":"cross","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28576","pdf_url":"https://arxiv.org/pdf/2608.28576","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A5","B1","B3"],"tags":["合成数据","统计推断","LLM增强"],"reason":"用LLM生成合成样本增强调查数据推断，有真实数据对照，方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":22,"question":"如何在合成数据增强的统计推断中，确定合成样本数量与权重的有效配置，以保证置信集覆盖目标参数。","design":"论文提出一个通用框架，用于在相关任务总体中校准合成数据增强。对于每个任务，用户指定一个置信集构造程序，该程序结合真实样本（权重为1）和合成样本（权重为w，数量为n_syn）。通过历史任务（有更多真实样本）构建参考置信集，评估不同（n_syn, w）配置的覆盖有效性，估计尺寸-权重前沿，并给出有限样本覆盖保证。实验部分使用大语言模型生成合成回答来增强意见调查数据。","baseline":"历史任务中额外的真实人类样本用于构建参考置信集，作为评估合成增强配置覆盖有效性的基准。","findings":"论文提出尺寸-权重前沿概念，并开发数据驱动方法从历史任务中学习该前沿，保证前沿上及以下所有配置同时达到目标覆盖。在意见调查数据实验中，该方法实现了目标覆盖并显著缩小了置信区间。","reliability":"论文承认合成数据可能系统性偏离真实人群，导致朴素处理产生偏差；方法依赖于历史任务与目标任务来自同一任务总体的假设，且需要历史任务中有足够的真实样本来构建参考置信集。","relevance":"该研究直接针对用LLM生成合成样本增强调查数据推断的问题，提供了校准合成数据使用的方法，与研究者关注的人类仿真实验和可靠性评估高度相关，值得精读原文。","inspiration":"该方法通过历史任务校准合成数据的使用，可借鉴其尺寸-权重前沿思想来设计仿真实验中的合成数据使用策略。｜可迁移到经济学中的调查数据增强，例如消费者信心调查、政策支持度调查等，利用LLM生成合成回答来补充小样本真实调查。｜设计：以LLM生成的合成回答作为合成样本，真实人类调查数据作为基准，处理变量为合成样本的数量和权重，结果变量为置信区间覆盖率和宽度，使用历史调查问题作为校准集。"}},{"id":"2608.28001","version":1,"title":"FocusGen: Expanding Visual Design Exploration with a Simulated Focus Group of Persona Agents","zh_title":"FocusGen：用模拟焦点小组扩展视觉设计探索","abstract":"Creative professionals rarely design for themselves--they design for audiences whose preferences they must anticipate. Yet current text-to-image exploration tools derive diversity entirely from the designer's own input--their prompts, their chosen dimensions, their search queries--confining exploration to what the designer already knows to look for. We present FocusGen, an interactive system that introduces external perspectives into visual design exploration through a \"virtual focus group\" of simulated persona agents. In contrast to prior persona systems in which multiple agents converge as critics on a single evolving artifact, FocusGen uses personas as parallel generators: each agent--constructed from demographic data, a procedurally generated backstory, and aesthetic preferences elicited through interviews--independently drives an iterative generation loop that produces its own visual concept, transforming one design brief into a spectrum of audience-conditioned directions. With real human participants, we confirm that the iterative refinement loop produces outputs people prefer over zero-shot generation. With synthetic agents at scale, we show that persona conditioning yields higher visual diversity than a generic-assistant baseline--measured by CLIP distance and corroborated by human perceptual judgments--and that open-ended preference interviews yield more diverse outputs than structured ones for both human and synthetic cohorts, while also revealing that agent cohorts recover only part of the diversity of comparable human cohorts. A qualitative study with 16 creative professionals suggests FocusGen helps designers discover unanticipated directions, overcome fixation, and probe audience contexts--while surfacing stereotyping risks that we analyze. We position FocusGen as a divergence scaffold for early-stage ideation rather than a substitute for audience research.","authors":["Jaewon Choi","Helena Vasconcelos","Hyun Lee","Carolyn Zou","Tak Yeon Lee","Michael Bernstein"],"categories":["cs.HC","cs.MA"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28001","pdf_url":"https://arxiv.org/pdf/2608.28001","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM仿真","人机交互","设计探索"],"reason":"用LLM persona模拟焦点小组，生成设计方向，并与真实人类对照，但非社会…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-31","rank":7,"question":"能否用模拟的受众角色（persona agents）作为“虚拟焦点小组”，为视觉设计探索引入外部视角，从而突破设计师自身视角的限制？","design":"使用 gemini-2.5-flash 构建 1000 个基于人口统计数据和程序化背景故事的角色代理，并通过开放式或结构化访谈获取其审美偏好；每个代理独立驱动迭代生成循环，根据设计简报生成视觉概念；结果变量为生成图像的视觉多样性（CLIP 距离）和人类感知判断。","baseline":"真实人类参与者：用于验证迭代细化循环优于零样本生成；人类队列：用于比较代理队列与人类队列的多样性差距；16 位创意专业人士：用于定性评估系统实用性。","findings":"角色条件化相比通用助手基线显著提高了生成图像的视觉多样性，且开放式偏好访谈比结构化访谈产生更多样化的输出。但代理队列仅恢复了可比人类队列的部分多样性，且系统存在刻板印象风险。","reliability":"论文明确承认评估未建立输出对个体角色的判别保真度或对人口群体的代表性，且代理队列与人类队列之间存在持续的多样性差距；系统可能强化刻板印象。","relevance":"该研究将 LLM 角色代理用于模拟受众偏好并生成多样化输出，与人类数据对照，属于人类仿真实验范畴，但场景为设计探索而非经济决策，可借鉴其方法但需注意外部效度。","inspiration":"借鉴其通过角色代理施加异质性偏好处理、以多样性指标和人类判断作为结果变量的设计，以及开放式访谈优于结构化访谈的发现。｜可迁移到消费者偏好异质性研究，如不同人口群体对金融产品特征的偏好差异，或政策信息在不同群体中的传播效果。｜以人口统计和背景故事构建 LLM 代理模拟不同收入或教育水平的消费者，处理为呈现不同设计的产品或政策信息，结果变量为代理的选择或态度分布，并与真实调查数据（如消费者金融调查）对照，检验代理模拟的多样性和偏差。"}},{"id":"2608.26291","version":1,"title":"Assessing mentalization in humans and large language models","zh_title":"评估人类与大语言模型的心理化能力","abstract":"Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks. Last, GPT-5 agents flexibly adapted their recursive depth of reasoning to increasingly sophisticated opponents, demonstrating superior performance to human participants. Collectively, we demonstrate different capacities for mentalization across LLMs, and highlight cognitive computational modeling as a formal method for assessing comparative intelligence across humans and machines.","authors":["Aamir Sohail","Xintong Zhong","Arkady Konovalov","Patricia L. Lockwood","Lei Zhang"],"categories":["cs.AI","q-bio.NC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26291","pdf_url":"https://arxiv.org/pdf/2608.26291","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3"],"tags":["LLM仿真","经济学实验","认知建模"],"reason":"用LLM作为人类被试替代品，在经济学博弈中与人类数据对照，评估心理化能力与策略。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":1,"question":"LLM能否像人类一样通过心理化（推断他人信念与意图）来指导策略行为，其潜在策略是什么？","design":"用四个模型家族（DeepSeek、GPT-4.1、GPT-5、Gemini 2.0 Flash）的LLM个体（N=2099）作为被试，参与两个经济博弈（检查博弈和石头剪刀布），部分模型施加社会思维链（SCoT）提示，测量博弈得分和策略选择，并用认知计算建模推断递归推理深度。","baseline":"人类被试（检查博弈 n=67，石头剪刀布 n=184）在相同任务中的表现。","findings":"LLM表现出明显的心理化行为与计算特征，但不同模型和规模差异显著；SCoT提示普遍提升表现，但提升幅度因模型而异。GPT-5能灵活适应对手复杂度，表现优于人类。","reliability":"论文未讨论","relevance":"该研究直接以LLM作为人类被试替代品，在经济学博弈中与人类数据对照，评估心理化能力与策略，完全命中你的关注点，值得精读原文。","inspiration":"借鉴其用认知计算建模从行为数据中提取潜在策略（递归推理深度）的方法，以及用SCoT提示作为处理变量来诱导更复杂推理的设计。｜可迁移到资产定价实验中的策略性预期形成或谈判博弈中的信念更新。｜用LLM作为投资者被试，施加SCoT提示，在重复博弈中测量报价或投资决策，并与人类实验数据（如资产泡沫实验）对照，比较递归推理深度和适应性。"}},{"id":"2608.26327","version":1,"title":"How Unlikely Is \"Unlikely\"? Assessing Verbal Probability Perception Across Large Language Models","zh_title":"“不太可能”有多不可能？跨大语言模型评估言语概率感知","abstract":"Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty) remains unknown. We present a systematic cross-model evaluation using a word-to-number mapping task grounded in established human benchmarks. Eleven uncertainty expressions were presented to 19 models under two conditions, forced single-number response and explanation elicitation, alongside a novel bidirectional roundtrip test of internal consistency. LLMs track the human benchmark with surprising fidelity: word ordering is preserved, three anchor points are recovered, and ``possible'' shows the highest variance and cross-model disagreement of any expression tested, consistent with its documented bimodal interpretation in humans. However, models show a systematic upward bias for negative expressions such as ``unlikely'' and ``improbable.'' Explanation elicitation reduces within-model variance while increasing between-model divergence, stabilizing individual models at the cost of inter-model consensus, and the roundtrip experiment reveals clear stratification, with frontier models maintaining coherent bidirectional representations. LLMs thus reproduce the structure of human verbal probability cognition, including its biases, while diverging systematically at the negative end---with implications for any setting where humans and models exchange probabilistic language.","authors":["Christos Petridis","Konstantinos Pelechrinis","Zoran Obradovic"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26327","pdf_url":"https://arxiv.org/pdf/2608.26327","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","概率语言","人类基准对照"],"reason":"用LLM复现人类对概率词的理解，并与人类基准对照，发现偏差，属于仿真人类认知且…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":5,"question":"LLM对言语概率表达（如“unlikely”）的数值解读是否与人类基准一致，以及在不同提示条件下是否稳定？","design":"用19个LLM作为被试，呈现11个概率词，在两种提示条件下（强制单数字回答、要求解释）进行词到数字映射，并进行双向往返一致性测试，测量映射数值、方差和跨模型一致性。","baseline":"Mosteller和Youtz（1990）汇总的20项人类研究中概率词的数值解读基准。","findings":"LLM整体上复现了人类基准：词序保持、三个锚点（impossible, even chance, certain）恢复良好，“possible”方差最大，与人类双峰解释一致。但LLM对负面词（如unlikely）存在系统性高估，解释提示降低模型内方差但增加模型间分歧，往返测试显示前沿模型双向映射更一致。","reliability":"论文指出偏差集中在非锚点词，负面词在强制条件下偏差略大；解释提示虽稳定个体但牺牲跨模型共识；部分模型往返映射接近随机，表明内部表征不一致。","relevance":"该研究直接评估LLM作为人类被试在概率词理解上的仿真效度，发现总体复现但存在系统性偏差，对关注LLM仿真人类认知可靠性的研究者有重要参考价值。","inspiration":"借鉴其词到数字映射任务和双向一致性检验，可迁移到经济金融中的风险沟通与预期形成场景，例如央行政策声明中的模糊措辞解读。｜设计一个实验：以LLM为被试，呈现央行声明中的概率词（如“可能加息”），要求给出数值概率，并与专业预测者调查或市场隐含概率对照，检验LLM是否复现人类解读偏差。"}},{"id":"2608.26188","version":1,"title":"Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments","zh_title":"你的社区安全吗？大语言模型城市安全判断中的地方污名","abstract":"Large language models are increasingly used to inform safety decisions in cities, such as where it is safe to walk, rent, or travel. We ask whether such judgments track measured risk or the patterns attached to an urban neighborhood's name. We probe seven instruct-tuned models under three conditions that dissociate name from geography: coordinates-only, name-only, and name+coordinates, across 186 neighborhoods in Los Angeles and Chicago, joined to violent crime and American Community Survey data. First, ratings are nearly flat under coordinates for six of seven models, while names carry most between neighborhood variation and are moderately calibrated to violent crime; only at frontier scale does the coordinate channel show appreciable variation. Second, names lower safety ratings more for neighborhoods with higher shares of the locally dominant marginalized group (percent Black in Chicago, percent Hispanic in Los Angeles), and this name effect tracks demographic share in all seven models and both cities. In Los Angeles, where demographic share and crime are more separable, the effect survives controls for crime and income and is confirmed by crime-matched pairs. An enforcement-elasticity analysis further shows that over-caution tracks near-fully-reported homicide rather than discretionary, deployment-driven offenses. Third, the effect scales with geographic knowledge: models that better distinguish real neighborhoods apply more demographic stereotype to them. Because neighborhood names carry both genuine crime signal and demographic stereotype, removing names reduces both bias and accuracy. We discuss implications for deploying LLMs in advice and decision-support settings.","authors":["Huy Nguyen","Yue Lin"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26188","pdf_url":"https://arxiv.org/pdf/2608.26188","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","社会偏见","城市安全"],"reason":"用LLM模拟人类对社区安全的判断，并与真实犯罪和人口数据对照，揭示偏差。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":3,"question":"LLM 对城市社区夜间步行安全的判断，是追踪真实犯罪风险，还是反映与社区名称相连的种族-空间污名？","design":"用七种指令微调 LLM 模拟居民对社区安全的感知，对芝加哥和洛杉矶共 186 个社区，在仅坐标、仅名称、名称+坐标三种条件下生成夜间步行安全评分，并与真实暴力犯罪和人口普查数据对照。","baseline":"真实暴力犯罪记录（芝加哥 2020-2025、洛杉矶 2020-2024）和美国社区调查（ACS）人口统计数据。","findings":"名称承载了几乎所有社区间安全评分差异，且与暴力犯罪中度校准；名称对安全评分的压低效应随当地主导边缘化群体（芝加哥黑人、洛杉矶西裔）比例上升，在洛杉矶该效应在控制犯罪和收入后仍存在，且与犯罪匹配对验证一致。","reliability":"坐标通道在开源模型中几乎无变异，仅在 frontier 模型上显现；芝加哥因黑人比例与犯罪高度相关无法分离种族与犯罪效应；名称同时携带真实犯罪信号和人口统计刻板印象，隐藏名称会同时消除偏差和准确性。","relevance":"该研究用 LLM 模拟人类对社区安全的判断，并与真实犯罪和人口数据对照，直接揭示仿真中的种族偏差，对关注 LLM 仿真可靠性及偏差的研究者极具参考价值。","inspiration":"值得借鉴的是通过条件消融（仅坐标、仅名称、名称+坐标）分离名称效应，并用真实犯罪和人口数据做基准，以及用犯罪匹配对和执法弹性分解做稳健性检验。｜可迁移到信贷审批中的地域歧视研究，如 LLM 模拟信贷员对申请人的风险评估是否受申请人所在社区名称的种族构成影响。｜用 LLM 扮演信贷审批员，对虚构申请人给出贷款批准概率，处理变量为申请人地址的社区名称（高黑人/西裔比例 vs 低比例），结果变量为批准概率，对照真实数据用社区层面的实际贷款批准率和违约率，并控制申请人收入、信用分等特征。"}},{"id":"2608.26221","version":1,"title":"Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model","zh_title":"生成式智能体的提示敏感性：来自流行病模型的证据","abstract":"As generative AI gains traction, researchers are investigating its potential to serve as proxies for humans. From undergoing cognitive psychology experiments to experiencing an epidemic, generative agents, agents powered by generative AI models, produce realistic human behavior when prompted. This study explores the sensitivity of these generative agents' behavior to prompt modifications and varied persona names of the agents. To assess this sensitivity, we use a generative agent epidemic model, wherein each agent is prompted daily on whether it wants to isolate or commingle with other agents. We found that using synonymous prompts results in negligible changes to the model's outcomes. However, minor variations in prompts, as well as contextual changes, do influence the model's results. Lastly, our data indicates that different persona names assigned to generative agents, specifically those imbued with personas, do not significantly impact epidemic outcomes.","authors":["Ross Williams","Niyousha Hosseinichimeh"],"categories":["physics.soc-ph","cs.AI","cs.LG","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26221","pdf_url":"https://arxiv.org/pdf/2608.26221","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A3","B4"],"tags":["LLM仿真","流行病模型","提示敏感性"],"reason":"用生成式智能体模拟疫情中的人类隔离决策，研究提示敏感性，属于人类行为仿真，但无…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":4,"question":"生成式智能体在疫情模型中的行为对提示词修改和角色名称变化的敏感程度如何？","design":"使用生成式智能体疫情模型，每个智能体每天被提示选择隔离或与他人接触，通过改变提示词的语义、上下文和角色名称来测试行为变化，结果变量为疫情传播的流动性曲线。","baseline":"无对照","findings":"同义提示词修改对模型结果影响可忽略，但轻微提示词变化和上下文变化会影响模型结果。不同角色名称对疫情结果无显著影响。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM仿真中提示敏感性这一可靠性问题，但未与真实人类数据对照，且场景为疫情模型而非经济学实验，对关注经济学和政策评估的研究者参考价值有限。","inspiration":"可借鉴其系统改变提示词并测量输出变化的方法，用于评估LLM在经济实验中的稳健性。｜可迁移到政策公告的预期形成实验，测试不同措辞的公告对LLM代理预期的影响。｜用LLM模拟投资者，随机分配不同措辞的央行声明，测量其通胀预期变化，并与专业预测者调查数据对照。"}},{"id":"2608.23705","version":2,"title":"The Limits of Automatic Evaluation of Creativity in Large Language Models","zh_title":"大语言模型创造力自动评估的局限性","abstract":"Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evaluating creativity in LLM-generated content remains a significant challenge. Here, we investigate whether current automatic evaluation methods can reliably capture human judgments of creativity. We collect human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity, and compare these judgments with automated objective metrics and LLM-as-a-Judge evaluations. Our experiments reveal substantial misalignment between automatic evaluations and human assessments. In particular, LLM-based judges exhibit a systematic preference for AI-generated stories, consistently favoring their stylistic characteristics over the unpredictability and other qualities of human-authored texts. Furthermore, correlation analyses show that widely used automatic metrics exhibit near-zero alignment with human judgments across both human- and AI-generated stories, suggesting that they fail to capture important dimensions of creativity. These findings highlight fundamental limitations in current approaches to the automatic evaluation of creative text and underscore the difficulty of reducing the multidimensional and subjective nature of creativity to computational metrics.","authors":["Alessandro Tutone","Giorgio Franceschelli","Mirco Musolesi"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-28","first_seen":"2026-08-26","revised_at":"2026-08-28","abs_url":"https://arxiv.org/abs/2608.23705","pdf_url":"https://arxiv.org/pdf/2608.23705","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM评估偏差","人类对照","创造力测量"],"reason":"评估LLM作为评判者的可靠性，与人类判断对照，揭示偏差，可迁移到仿真效度研究。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":6,"question":"当前自动评估方法能否可靠捕捉人类对创造力的判断？","design":"本研究并非仿真研究，而是评估自动评估方法：收集人类对100篇人类创作和100篇LLM生成短故事在11个创造力维度上的评分，并与自动客观指标和LLM-as-a-Judge评分进行比较。","baseline":"人类评估数据：来自WritingPrompts数据集的100篇人类创作故事和100篇LLM生成故事，由人类评分者在11个创造力维度上进行评价。","findings":"自动评估与人类判断存在显著不一致，LLM评委系统性偏好AI生成故事，偏向风格特征而非不可预测性等人类文本特质。广泛使用的自动指标与人类判断相关性接近零，未能捕捉创造力的重要维度。","reliability":"论文指出LLM评委存在自我偏好偏差，且自动指标与人类判断相关性极低，表明当前自动评估方法在创造力评估上不可靠，需要替代方法。","relevance":"该研究直接评估LLM作为评判者的可靠性，并与人类判断对照，揭示系统性偏差，对仿真研究中用LLM替代人类评估者的效度有重要警示。","inspiration":"借鉴其多维度人类评分与自动评分对照的设计，可迁移到经济金融中涉及主观判断的仿真评估，如信贷审批中的公平性判断或投资决策中的风险评估。可设计实验：用LLM扮演信贷员对贷款申请进行审批，处理为申请人特征（如种族、性别），结果变量为审批决定和理由，与真实信贷审批数据或人类专家判断对照，检验LLM是否存在类似偏差。"}},{"id":"2608.23780","version":2,"title":"When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk","zh_title":"当青少年进入聊天：基于LLM的学生话语测量验证的认识论转变","abstract":"LLMs are being used increasingly to measure aspects of student discourse (e.g. talk moves, collaboration, equity of voice) at scale. Typically, LLM-based measures of student talk use transcriptions of classroom conversations that only include verbal contributions, which de-contextualize student language. Common practices for validating these measures include comparing outputs against expert annotations by adults, using held out evaluation sets and F1 scores. We argue that these approaches are insufficient to ensure that such measures are meaningful and equitable for teaching and learning, particularly for racially and linguistically marginalized youth. In order to center the youth whose talk is being analyzed, re-contextualizing these classroom conversations and engaging youth in the research process is necessary. Sharing epistemic authority with youth, ultimately, centers their point of view and adds crucial nuance to the analysis of their talk that adult experts, researchers, and LLMs cannot provide. In a case study of multilingual youth in one 8th-grade math classroom, we address the epistemic exclusion of youth by employing multiple ethnographically-oriented methods to re-contextualize student conversations and center youth as epistemic authorities in conversation with researchers and LLMs. We conducted participant observations, interviews, focus groups, and member checks with four focal students. Findings reveal that there were misalignments between students' interpretations of their own math talk experiences and the LLM-based measures of their talk. Students contested both the LLM classifications and the coding scheme used to measure their talk, highlighting the need for youth to be involved in the epistemic process of producing knowledge about their experiences.","authors":["Liliana Santos-Deonizio","James Malamut","Ram\\'on Antonio Mart\\'inez","Dorottya Demszky"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-28","first_seen":"2026-08-26","revised_at":"2026-08-28","abs_url":"https://arxiv.org/abs/2608.23780","pdf_url":"https://arxiv.org/pdf/2608.23780","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM测量效度","批判性评估","教育话语分析"],"reason":"评估LLM测量学生话语的效度，指出与青少年自身解读的偏差，批判性视角可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":7,"question":"在验证基于LLM的学生话语测量时，让青少年参与知识生产过程能揭示什么？","design":"本研究不是仿真实验，而是质性案例研究。作者对一所八年级数学课堂中的四名多语种学生进行参与式观察、访谈、焦点小组和成员核查，将学生的自我解读与LLM对其话语的分类进行对比。","baseline":"无对照","findings":"学生的自我解读与LLM分类之间存在错位，学生质疑LLM的分类和编码方案。仅依赖成人专家标注和文本转录的验证方法不足以捕捉学生话语的情境意义，需要让青少年作为认识权威参与验证过程。","reliability":"论文指出，LLM基于去情境化的文本转录进行测量，忽略了关系、课堂常规、物理空间、手势、韵律和语言意识形态等维度，导致对边缘化学生话语的误读。","relevance":"该研究批判了LLM测量人类行为时仅依赖专家标注和文本数据的验证方式，强调纳入被测量者自身视角的重要性，对评估LLM仿真人类被试的可靠性与偏差具有方法论启示。","inspiration":"借鉴其成员核查和情境重构方法，将人类被试的自我报告与LLM输出进行系统对比，以揭示仿真偏差。｜可迁移到信贷审批歧视或消费者金融行为研究，检验LLM对借款人自述或消费决策的分类是否与借款人自身理解一致。｜以真实信贷申请文本为输入，让LLM预测违约风险或分类借款用途，同时收集借款人对自身文本的解读作为对照，比较LLM输出与借款人自我报告的一致性，并分析偏差来源。"}},{"id":"2608.26899","version":1,"title":"Counterfactual Bias Testing for Application Tracking System","zh_title":"申请追踪系统的反事实偏见测试","abstract":"Automated candidate-job matching systems are increasingly classified as high-risk AI under emerging regulation, yet auditing them for demographic bias is expensive: classical correspondence-audit studies require hand-crafted resumes and manual submission, which does not scale to fast pipeline retraining cycles. This paper presents a general, reusable methodology that (1) uses task-specialized LLM agents to synthesize identity-neutral base resumes and inject controlled demographic treatments across five protected-characteristic axes (sex/gender, age, residence, language, disability), producing a K x (1+N) correspondence-audit matrix; (2) qualitatively flags inferred protected characteristics per an EU AI Act-aligned prompt; (3) ranks candidates against a job description via a fine-tuned sentence-embedding model and cosine similarity; and (4) computes a nine-metric fairness suite spanning counterfactual (score delta, mean absolute rank change, flip rate), group-fairness (top-K retention, four-fifths/impact ratio), and merit-aware (Recall@K, nDCG@K, equal opportunity, equalized odds) families, each with bootstrap confidence intervals, significance tests, and Benjamini-Hochberg correction, culminating in an automated PASS/INVESTIGATE/FAIL report with a composite risk score. On an example corpus of 5 job orders, 100 base candidates, and 10 demographic treatments (90 metric x variant evaluations): score shifts, top-K retention, and merit-aware rate gaps stay within tolerance for every treatment, but a rank-stability metric (MARC) and nDCG@K each surface borderline findings - including one on the neutral baseline itself - that a score- or retention-only view would miss. The results argue for multi-metric, multi-family auditing over any single aggregate score, and for LLM-agent-generated audits as a practical, low-cost complement to human-curated audits for any candidate-job matching pipeline.","authors":["Sai Yashwant","Shruti Bansal","Anurag Dubey","Samaroha Chatterjee","Satyam Kumar","Shreyash Gupta","Gantala Thulsiram"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26899","pdf_url":"https://arxiv.org/pdf/2608.26899","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","A2","B1","B2","B3"],"tags":["LLM仿真","算法审计","公平性评估"],"reason":"用LLM生成简历模拟人类求职者，审计招聘系统偏见，有真实数据对照，属仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":9,"question":"如何用LLM智能体自动生成简历变体，对候选人-职位匹配系统进行反事实偏见审计，并构建多指标公平性评估框架？","design":"使用任务专用LLM智能体链生成身份中立的基准简历，并在五个受保护特征轴（性别、年龄、居住地、语言、残疾信号）上注入受控人口学处理，产生K×(1+N)的通信审计矩阵；然后通过微调的句子嵌入模型和余弦相似度对候选人进行排序，并计算九项公平性指标（反事实、群体公平、绩效感知三类），每个指标配有自助置信区间、显著性检验和Benjamini-Hochberg校正，最终生成PASS/INVESTIGATE/FAIL报告。","baseline":"无对照","findings":"在示例语料库（5个职位、100个基准候选人、10个人口学处理）上，所有处理的分数偏移、top-K保留率和绩效感知率差均在容差范围内，但排名稳定性指标（MARC）和nDCG@K出现边缘性发现，包括中性基线本身，表明仅看分数或保留率会遗漏问题。","reliability":"论文未讨论","relevance":"该研究利用LLM生成简历模拟人类求职者，审计招聘系统偏见，属于仿真人类被试的范畴，但缺乏真实人类数据对照，且场景为招聘而非经济学实验，对关注经济学实验和政策评估的研究者参考价值有限。","inspiration":"可借鉴其用LLM生成反事实变体并构建多指标审计框架的方法，用于生成受控处理组和对照组。｜可迁移到信贷审批歧视审计，用LLM生成贷款申请人资料，注入性别、种族等特征，评估信贷模型偏见。｜以LLM生成的贷款申请人为被试，处理为注入受保护特征（如性别、种族），结果变量为贷款审批分数或决策，对照真实信贷审批数据（如Home Mortgage Disclosure Act数据）验证仿真可靠性。"}},{"id":"2608.24912","version":1,"title":"Analyzing and Correcting Benevolence Bias in Large Language Models","zh_title":"分析和纠正大语言模型中的仁慈偏差","abstract":"Large language models (LLMs) are increasingly used as stand-ins for human respondents, from opinion polls and simulated survey participants to agent-based social simulations. These uses rest on one assumption: that conditioning a model on who a person is yields answers resembling those of real people from that group. Here we identify and measure benevolence bias, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions. Across 18 widely used models, four social-science datasets (ANES, GSS, WVS, and a cross-cultural prospect-theory replication) and six psychological categories, we find that the bias is a stable model property, not a quirk of any one system: it points the same way across models, grows with model size, and traces to the post-training stage. Prompt language and framing change its size but never its direction, and a \"malicious persona\" stress test shows a one-sided limit: aligned models struggle to play people who are less kind, less prosocial or more harm-tolerant than average. The issue is thus not only a shifted average, but a narrowed range of people the model can imitate. The bias sits in the middle of the answer distribution rather than its tails, and survives changes in sampling temperature and simple prompted reflection. The encouraging news is that it is easy to diagnose and straightforward to fix: a light-touch contrastive calibration, which needs no retraining and works on black-box APIs, brings all six categories back to the human baseline. Our results give researchers a clear map of where aligned LLMs can already be trusted as human stand-ins, where they need care, and a ready-to-use method for closing the gap.","authors":["Yuanzi Li","Junhao Wang","Minghui Liu","Boyi Li","Bingchen Chen","Zihang Tian","Jingyu Zhao","Yuhan Wang","Lei Wang","Pei Wang","Jinchao Wu","Xu Chen"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24912","pdf_url":"https://arxiv.org/pdf/2608.24912","source_feed":"cs.HC","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B3","B4"],"tags":["LLM仿真","算法保真度","偏差校正"],"reason":"直接研究LLM作为人类被试替代品的偏差，使用真实人类数据对照，并提出校准方法。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":1,"question":"对齐后的大语言模型在作为人类被试回答价值负载的调查问题时，是否存在系统性的仁慈偏差，其来源、表现和可校正性如何？","design":"该研究并非传统仿真实验，而是对18个广泛使用的LLM进行系统性测试：将模型置于模拟人类受访者的角色，输入来自ANES、GSS、WVS和跨文化前景理论复制的调查问题，测量模型回答在六个仁慈偏差类别（社会赞许性、伤害规避、亲社会动机、仁慈解释、公平乐观、情感软化）上的偏差程度，并考察模型规模、训练阶段、提示语言、框架、角色设定、采样温度等因素的影响，最后提出对比校准方法。","baseline":"使用四个真实人类调查数据集作为基准：美国国家选举研究（ANES）、综合社会调查（GSS）、世界价值观调查（WVS）以及一项跨文化前景理论复制的数据。","findings":"LLM存在稳定且一致的仁慈偏差，倾向于选择更友善、更安全、更符合社会期望的答案，且该偏差随模型规模增大而增强，主要源于后训练阶段。提示语言和框架只能改变偏差大小而不能改变方向，恶意角色压力测试显示模型难以模仿低于人类平均水平的亲社会或伤害容忍度，但对比校准方法无需重新训练即可将偏差校正至人类基线。","reliability":"论文指出偏差位于答案分布的中间而非尾部，且对采样温度和简单提示反思不敏感；恶意角色测试显示模型在部分维度上无法达到人类低仁慈端，表明仿真范围收窄。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，提供了系统性的偏差测量和校正方法，对于关注仿真效度和偏差的研究者具有重要参考价值。","inspiration":"该研究采用多模型、多数据集、多心理类别的系统测量框架，并通过对比校准进行偏差校正，值得借鉴。｜可以迁移到经济金融领域的调查实验和个体决策仿真，例如风险偏好、时间偏好、公平观念、信任与合作等。｜以LLM作为虚拟被试，施加不同的经济情境或政策干预，测量其选择或态度，并与真实实验数据（如实验经济学中的公共品博弈、最后通牒博弈、风险偏好问卷等）进行对照，检验并校正LLM的偏差。"}},{"id":"2604.06223","version":3,"title":"The Quiet and the Compliant: How Regulation and Polarization Shape Conventional Wisdoms on Corporate Social Engagement in High-risk Settings","zh_title":"沉默与顺从：监管与极化如何塑造高风险环境下企业社会参与的常规智慧","abstract":"With the international business landscape becoming more crisis-ridden as risks proliferate, how do the professionals who implement corporate social initiatives in high-risk environments perceive their work, and what can this reveal about the forces shaping business engagement with society in crisis contexts? We present findings from a synthetic survey of 400 corporate professionals working on social impact in fragile and conflict-affected settings to understand conventional wisdoms and best practices on corporate strategy and activity in high-risk settings. Drawing on political corporate social responsibility (CSR), synthetic survey, and international business literatures, we test seven hypotheses about how regulatory environments, political polarization, sector characteristics, and organizational structures shape corporate social engagement in high-risk contexts. The synthetic results suggest that European professionals report significantly higher strategic integration of social impact across all measured dimensions, while US professionals overwhelmingly report that political polarization hinders social initiatives, yet this perception does not predict unreported social activities, complicating the emerging \"quiet CSR\" narrative. Extractive industry professionals deliver both the highest operational preparedness and the highest complicity awareness, a pattern we conceptualize as presence-dependent reflexivity. These patterns deliver a baseline to detect the theorized dynamics and offer preliminary theoretical propositions for future real-world empirical testing.","authors":["Jason Miklian"],"categories":["physics.soc-ph","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"replace-cross","date":"2026-08-27","first_seen":"2026-03-27","revised_at":"2026-08-27","abs_url":"https://arxiv.org/abs/2604.06223","pdf_url":"https://arxiv.org/pdf/2604.06223","source_feed":"cs.SI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","合成调查","企业社会责任"],"reason":"用LLM合成400名企业专业人士调查，模拟高风险环境下的态度与决策，并与真实文…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":2,"question":"在脆弱和冲突影响的高风险环境中，企业社会参与的专业人士如何感知其工作，以及监管环境、政治极化、行业特征和组织结构如何塑造企业社会参与？","design":"使用合成调查方法，模拟400名在欧美总部、员工超过1000人的企业社会影响专业人士，通过23个李克特量表题测量战略整合、监管压力、政治极化、运营准备、供应链协调、子公司自主权和ESG评级有效性等感知，并基于理论推导出七个假设进行检验。","baseline":"无对照","findings":"欧洲专业人士在所有维度上报告显著更高的社会影响战略整合，而美国专业人士压倒性地报告政治极化阻碍社会倡议，但这种感知并不预测未报告的社会活动，复杂化了“安静CSR”叙事；采掘业专业人士同时表现出最高的运营准备和最高的共谋意识，被概念化为存在依赖的反身性。","reliability":"论文未讨论","relevance":"该研究使用LLM合成调查数据，模拟企业专业人士在高风险环境下的态度和决策，并基于理论提出假设进行检验，属于用LLM进行人类仿真实验的研究，且涉及政策评估场景，与您的兴趣高度相关，值得阅读原文了解其方法论和发现。","inspiration":"该方法通过合成调查生成理论预测的基准数据，为后续真实数据对比提供参照，可借鉴其构建“常规智慧基线”的思路；可迁移到政策评估中的企业行为研究，如ESG监管对企业社会参与的影响；设计上，可用LLM模拟企业高管作为被试，施加不同监管环境（如强制尽职调查 vs 反ESG法案）的处理，测量其战略整合和沉默行为，并与真实企业调查数据对照。"}},{"id":"2608.25771","version":1,"title":"Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences","zh_title":"基于困境训练的大语言模型少样本提示在预测患者偏好上超越人类代理","abstract":"In serious illness, human surrogates often struggle to accurately predict patient preferences (68% accuracy), causing decision conflict. Personalized Patient Preference Predictor (P4) agents offer a potential solution, but prior prototypes treat values as static ratings, ignoring the contextual, situation-dependent nature of medical choices. Grounded in the 'logic of care', we present P4-DT (Dilemma Training), a P4 agent that constructs a patient decision policy by engaging users with varied medical dilemmas, eliciting individual preference reasoning through bi-directional training. In a study with 12 patient-surrogate dyads, P4-DT predicted patient treatment choices with 81.7% accuracy, significantly exceeding chance (OR = 5.61 [2.03, 15.51], p < .001) and outperforming both unassisted surrogates (55.0%; OR = 3.67 [1.59, 8.47], p = .002) and surrogates assisted by P4-DT (61.7%). Comparative prompt analyses showed that incorporating contextual scenario decisions and open-ended text improved accuracy by 15.0 percentage points over initial values ratings alone. We discuss implications for further testing and designing of context-aware AI agents that embody richer human experience to partner in complex decision-making.","authors":["Natasha Ureyang","Sebastian Porsdam Mann","Yuxin Liu","Zuriel Hassirim","Melanie Almonte","Wenhao Chen","Joyce Ng","Thant Nay Lin","Aung Thiha","Gerald CH Koh","Brian David Earp","Pin Sym Foong"],"categories":["cs.HC","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25771","pdf_url":"https://arxiv.org/pdf/2608.25771","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","患者偏好预测","人类对照"],"reason":"用LLM预测患者偏好，与人类代理对照，评估准确率，属核心仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":4,"question":"如何通过让患者参与医疗困境训练来构建个性化患者偏好预测器（P4-DT），以提高对患者治疗偏好的预测准确率？","design":"使用GPT-5.5模型作为P4-DT代理，通过提示工程进行少样本学习。对12名患者（MP）进行训练：先填写价值观调查，再对5个医疗困境场景做出治疗偏好决策并解释理由，同时可查看模型预测并反馈。测试阶段，MP对5个新场景做决策，模型基于训练数据预测其偏好；同时人类代理人（TO）独立预测MP偏好，并在P4-DT辅助下再次预测。结果变量为预测准确率（方向一致性）。","baseline":"人类代理人（TO）的预测准确率（55.0%），以及TO在P4-DT辅助下的预测准确率（61.7%）。","findings":"P4-DT预测患者治疗选择的准确率为81.7%，显著高于随机水平（OR=5.61, p<.001），且优于未辅助的人类代理人（55.0%）和P4-DT辅助的代理人（61.7%）。比较提示分析显示，纳入情境场景决策和开放式文本比仅使用初始价值观评分提高了15.0个百分点的准确率。","reliability":"论文未讨论","relevance":"该研究用LLM模拟患者偏好预测，并与真实人类代理人对照，评估预测准确率，属于核心的LLM仿真人类决策研究，且涉及医疗决策场景，对关注仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其通过情境化困境训练和开放式文本解释来捕捉个体决策逻辑的方法，可迁移到消费者金融决策或政策偏好预测中。｜例如，在消费者信贷选择或退休储蓄决策中，可让LLM通过模拟具体金融困境（如贷款选择、投资风险权衡）来学习个体偏好。｜设计：招募真实消费者作为被试，先填写价值观和风险偏好问卷，再对5个金融困境场景做出选择并解释理由，训练LLM预测其在新场景中的选择；同时让人类代理人（如配偶）预测被试选择，比较LLM与人类代理人的预测准确率，并以被试实际选择为基准。"}},{"id":"2608.24920","version":1,"title":"Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment","zh_title":"不同大语言模型回复的语义变异性：对设计基于对话的评估的启示","abstract":"This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.","authors":["Jiangang Hao"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24920","pdf_url":"https://arxiv.org/pdf/2608.24920","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM仿真","语义一致性","对话评估"],"reason":"评估LLM回复与人类回复的一致性，涉及仿真可靠性，有真实人类数据对照，并指出失…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":10,"question":"当底层大语言模型改变时，LLM生成的回复在语义上是否保持一致？","design":"使用四个LLM（GPT-4o mini、GPT-5.4、GPT-5.4 mini、GPT-5.4 nano）扮演协作对话中的回复者，对从真实协作对话中选取的61条焦点消息生成回复，比较有无对话历史两种条件下回复的语义相似度。","baseline":"真实人类在原始协作对话中产生的回复，并编码为高、中、低相关性。","findings":"模型选择和对话上下文均影响回复的语义相似性及与人类回复的一致性。仅靠提示和对话上下文不足以保持跨LLM的回复一致性。","reliability":"论文指出提示和对话上下文不足以保持跨LLM的回复一致性，需要基础设施和设计策略来维持稳定可比的回复。","relevance":"该研究直接评估LLM回复与人类回复的一致性，涉及仿真可靠性，有真实人类数据对照，并指出失效条件，值得阅读原文。","inspiration":"借鉴其通过改变模型和上下文条件来系统评估回复一致性的实验设计，以及使用真实人类回复作为基准的对照方法。｜可迁移到经济金融领域的对话式调查或实验，如消费者金融咨询、投资顾问对话或政策沟通中的LLM应用。｜以LLM作为虚拟被试，在有无对话历史条件下对同一金融咨询问题生成回复，测量回复语义相似度，并与真实人类顾问的回复进行对照，评估LLM替代人类被试的可靠性。"}},{"id":"2608.25999","version":1,"title":"Distinct dynamics of conceptual and referential disruptions in human reading and large language model processing","zh_title":"人类阅读与大语言模型处理中概念与指称干扰的不同动态","abstract":"Linguistic meaning is grounded in conceptual content, from which reference to particular entities emerges as words enter discourse. To examine the processing dynamics associated with these two dimensions of meaning, we selectively disrupted conceptual or referential information in short narratives and traced the resulting effects in human self-paced reading and in the predictive and representational processing of large language models. In human reading, conceptual disruptions produced a strong but localized processing cost, emerging immediately after the distorted word, reaching an early maximum, and then declining rapidly. Referential disruptions produced weaker effects, which decreased more gradually across subsequent words, and were more strongly modulated by sentence boundaries. In the language model, both disruptions emerged immediately at the manipulated word. Contextual model surprisal showed a pattern closely paralleling human reading: conceptual disruption produced a larger, more locally concentrated effect that decayed rapidly, whereas referential disruption produced a smaller and more gradual downstream effect. Output-layer representations showed a different pattern: referential disruption produced a larger initial displacement, while both distortions were subsequently characterized by power-law decay. Together, these results provide convergent evidence for distinguishable processing dynamics of two types of meaning: conceptual information imposes a more locally concentrated integration cost, whereas referential information engages a more distributed process of maintaining discourse-level identity.","authors":["Rui He","Nihal Altay","Wolfram Hinzen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25999","pdf_url":"https://arxiv.org/pdf/2608.25999","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","人类对照","认知建模"],"reason":"用LLM模拟人类阅读行为并与人类数据对照，方法可迁移至仿真研究","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":14,"question":"概念信息和指称信息的加工动态在人类阅读和大型语言模型处理中是否不同？","design":"本研究不是用LLM模拟人类被试，而是将LLM作为计算模型与人类被试并行比较。在短叙事中分别施加概念扭曲（CD）和指称扭曲（RD），测量人类自定步速阅读的逐词阅读时间，以及LLM的上下文惊奇度和输出层表征位移。","baseline":"人类自定步速阅读实验的逐词阅读时间数据。","findings":"人类阅读中，概念扭曲产生强而局部的加工代价，立即出现、早期达峰、迅速衰减；指称扭曲效应较弱，衰减更慢，且更受句子边界调节。LLM的上下文惊奇度与人类模式相似，但输出层表征显示指称扭曲初始位移更大，两者随后均呈幂律衰减。","reliability":"论文未讨论","relevance":"该研究将LLM的预测和表征与人类逐词阅读行为直接对照，展示了LLM在捕捉不同语义加工动态上的异同，对评估LLM作为人类语言加工仿真模型的有效性具有参考价值。","inspiration":"借鉴其通过施加局部扭曲并追踪下游效应动态来分离不同认知成分的方法，可用于经济金融文本信息加工研究。｜可迁移到政策公告的预期形成研究，考察概念性信息（如政策内容）与指称性信息（如政策对象）对市场预期更新的不同影响。｜以LLM为被试，在政策公告文本中分别扭曲概念词或指称词，测量模型对后续经济指标预测的惊奇度变化，并与真实市场中分析师预期调整数据对照。"}},{"id":"2608.25236","version":1,"title":"Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making","zh_title":"罕见病，常见困境：LLM在决策中优先考虑资源平等分配而非患者获益","abstract":"Clinical decision-making often involves prioritizing ethical values, such as beneficence, non-maleficence, respecting a patient's autonomy, and justice. Recent work has begun to assess how large language models (LLMs) make such subjective, value-laden clinical judgments. However, evaluations of LLM decision-making in rare disease care contexts, where ethical tensions are ubiquitous and where scarce prior information likely impacts LLM behavior, are still lacking. Here, we present a benchmark of 208 clinically grounded rare disease vignettes, each of which presents genuine, high-stakes conflicts. When prompting 11 state-of-the-art LLMs to choose between clinically defensible yet ethically conflicting next steps embedded within these vignettes, we found that all evaluated models consistently prioritized justice over other core bioethical principles. Specifically, models overwhelmingly favor equal resource allocation over need-based considerations, indicating LLMs' limited responsiveness to differences in clinical severity or situational context. We also identify a strong authority-framing effect: models favor justice in committee-based contexts and shift toward beneficence and autonomy only when final decisions are framed as being made by clinicians or patients respectively. Our work suggests that institutional pressures surrounding rare disease resource utilization may be silently reflected in LLM-based decision support systems, with finer ethical considerations disregarded.","authors":["Minda Zhao","Xu Han","Rishabh Goel","Maya Dagan","Noa Dagan","Adithya Madduri","Payal Chandak","Shilpa Nadimpalli Kobren","Isaac S. Kohane"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25236","pdf_url":"https://arxiv.org/pdf/2608.25236","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B2","B4"],"tags":["LLM决策仿真","伦理决策","资源分配"],"reason":"用LLM模拟临床决策，与人类伦理原则对照，涉及资源分配，但非经济学实验且无真实…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":11,"question":"大语言模型在罕见病临床决策中如何权衡自主、行善、不伤害与公正等伦理原则？","design":"构建208个罕见病临床情境短文，每个情境包含两个临床可辩护但伦理冲突的下一步行动；用11个先进LLM作为被试，通过提示让模型选择行动，并操纵决策权威框架（委员会、临床医生、患者），测量模型选择所体现的伦理原则优先级。","baseline":"无对照","findings":"所有模型一致优先考虑公正原则，压倒性地选择平等资源分配而非基于需求，对临床严重性或情境差异反应有限。存在权威框架效应：委员会情境下偏向公正，临床医生或患者决策情境下分别转向行善和自主。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类伦理决策，虽无真实人类对照，但揭示了LLM在资源分配场景中的系统性偏差，对关注仿真可靠性与偏差的研究者有参考价值。","inspiration":"借鉴其构造伦理冲突情境并操纵决策权威框架的方法，可迁移到经济金融中的资源分配或政策权衡问题，如公共预算分配、医疗资源定价或信贷审批中的公平与效率权衡。设计：用LLM作为被试，呈现稀缺资源分配情境（如器官移植、疫苗分配或信贷额度），操纵决策者角色（委员会、个体官员、受益人），测量分配方案（平等vs.按需vs.按效益），并与真实人类实验或调查数据对照。"}},{"id":"2608.24908","version":1,"title":"Hallucination by proxy in LLM-assisted differential diagnosis","zh_title":"LLM辅助鉴别诊断中的代理幻觉","abstract":"Current evidence suggests that LLM assistance could augment the diagnostic accuracy of clinicians. However, these systems are black boxes, susceptible to hallucinations, and project a potentially misleading level of confidence. It is currently unknown whether physicians are susceptible to accepting fabricated LLM suggestions, and whether this susceptibility varies with experience. We poisoned the system prompt of an LLM-based diagnostic assistant, forcing it to suggest a fictitious disease (neurocadmiumatosis) within an otherwise legitimate differential diagnosis. Across two independent phases, 18 of 41 participants (44%) incorporated neurocadmiumatosis into their final differential following LLM interaction: 18 of 26 participants with 6 months or less of neuroradiology training (69%) and 0 of 15 participants with >6 months of neuroradiology training (0%). Our results indicate that radiologists, particularly early in their training, are susceptible to LLM hallucinations. This \"hallucination by proxy\" phenomenon was exclusive to physicians with limited subspecialty experience, underscoring the need for structured training in critical appraisal of AI-generated content.","authors":["Bastien Le Guellec","Su-Hwan Kim","Ibrahima Niang","Aghiles Hamroun","Gr\\'egory Kuchcinski"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24908","pdf_url":"https://arxiv.org/pdf/2608.24908","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM幻觉","医生决策","人类行为对照"],"reason":"研究医生对LLM幻觉的易感性，用真实医生行为对照，属人类决策仿真与偏差评估","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":9,"question":"医生在多大程度上会不加批判地接受LLM生成的虚构诊断建议，这种易感性是否随专科经验水平而变化？","design":"本研究不是用LLM模拟人类，而是用真实医生作为被试，通过人为污染LLM系统提示词使其在真实病例中强制输出虚构疾病，测量医生在LLM辅助鉴别诊断后是否将该虚构疾病纳入最终诊断。","baseline":"无对照，但按经验分层：≤6个月神经放射学培训的参与者与>6个月培训的参与者形成对比。","findings":"44%的参与者将虚构疾病纳入最终诊断；其中≤6个月培训者69%纳入，>6个月培训者0%纳入。表明经验不足的放射科医生易受LLM幻觉影响，出现“代理幻觉”现象。","reliability":"论文未讨论失效条件，但指出易感性仅限于专科经验有限的医生，提示需要结构化培训来批判性评估AI生成内容。","relevance":"该研究直接评估人类在LLM辅助决策中对幻觉的易感性，属于人类决策偏差评估，与研究者关注的人类仿真可靠性及偏差条件高度相关，值得阅读原文。","inspiration":"借鉴其通过污染LLM提示词施加处理、以真实专家行为作为结果变量的设计，可迁移到金融顾问或信贷审批场景中AI建议对决策的影响。｜例如，研究AI生成的虚假财务信息是否影响投资决策或信贷审批。｜可招募不同经验的金融从业者或普通投资者，使用LLM生成包含虚构风险因素的投资建议，测量其是否采纳该建议，并与真实历史决策数据或专家判断对照。"}},{"id":"2607.22188","version":2,"title":"Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives","zh_title":"耗尽能源公地：智能体 LLM 集体中作为协调失败的自我挫败式过度占用","abstract":"LLMs are increasingly deployed as agents that plan, use tools, and act over time. When they share persistent resources, such as compute pools or energy reserves, decisions by one agent affect the conditions faced by later agents. We study this coordination failure in a renewable energy commons. Four same-family GPT, Gemini, or Grok agents act in homogeneous self-play as electricity prosumers, instructed to maximize operational continuity. Holding aggregate residual demand and the decision protocol fixed, we vary the regeneration rate of a shared energy reserve from abundance to scarcity. All three families preserve the reserve when demand does not exceed peak renewable replacement, but over-appropriate it beyond that threshold (all nine exact scarcity contrasts survive Holm correction; largest adjusted p = 4.87e-5). The pattern is self-defeating: the same populations protect current service while undermining future service. At higher scarcity (rho = 1.2), early aggregate request pressure exceeds peak renewable replacement in every family and averages 1.21 times that level. Mean trajectories fall below the reserve level of maximum replenishment by rounds 5-7. Two offline benchmarks compare a social planner maximizing group-wide operational-service value with open access, where each prosumer maximizes its own value. At a discount factor of gamma = 0.95, both benchmarks sustain the reserve under the same dynamics. Realized depletion instead resembles outcomes under a more impatient open-access benchmark. The populations therefore behave like impatient optimizers at the level of the public trajectory. This system-level alignment failure would be missed by isolated-response evaluation.","authors":["Marcantonio Bracale Syrnikov","Federico Pierucci","Matteo Prandi","Marcello Galisai","Piercosma Bisconti","Francesco Giarrusso","Daniele Nardi"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"replace","date":"2026-08-27","first_seen":"2026-07-27","revised_at":"2026-08-27","abs_url":"https://arxiv.org/abs/2607.22188","pdf_url":"https://arxiv.org/pdf/2607.22188","source_feed":"cs.MA","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM 多智能体","公共资源困境","社会模拟"],"reason":"LLM agent 群体模拟公共资源困境，但无真实人类数据对照，属社会模拟边界…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":3,"question":"LLM代理群体在共享可再生资源中是否会出现协调失败，导致自我挫败的过度占用？","design":"使用GPT-5.4-mini、Gemini-3.1-flash-lite和Grok-4.3三种模型，每种模型四个同族代理作为电力产消者，在自博弈中运行，目标最大化自身运营连续性。通过改变共享能源储备的再生率（从充裕到稀缺）施加处理，测量储备水平、请求压力、备用能源使用等结果变量。","baseline":"无对照","findings":"所有三种模型在需求不超过峰值可再生替代时能维持储备，但在稀缺条件下过度占用储备，导致自我挫败的储备下降。早期聚合请求压力超过峰值可再生替代，平均轨迹在5-7轮后低于最大补充水平。","reliability":"论文未讨论","relevance":"高度相关，直接研究LLM代理在公共品博弈中的协调失败，涉及系统级对齐失败，但缺乏真实人类对照，适合关注仿真可靠性与偏差的研究者阅读原文。","inspiration":"该研究通过改变共享资源再生率来施加稀缺性处理，并测量代理群体的请求压力和储备水平，这种系统级压力测试设计值得借鉴｜可迁移到公共资源管理实验，如渔业配额分配或碳排放权交易中的集体决策问题｜使用LLM代理模拟渔民群体，处理变量为资源再生率（高/低），结果变量为捕捞总量和资源存量，对照真实渔场历史数据或实验室人类实验结果"}},{"id":"2608.20539","version":2,"title":"ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations","zh_title":"ExploraTwin：一个用于数字孪生仿真的非营利研究平台","abstract":"Digital twin simulations show promise, but current empirical evidence suggests that the approach should be tested before being deployed in any particular context. To lower the friction for researchers and practitioners to test and deploy digital twin simulations, this brief commentary introduces ExploraTwin (https://exploratwin.org), an open-access, non-profit research platform for digital twin survey simulations. ExploraTwin supports two modes. In survey mode, researchers can upload a Qualtrics survey file or create a survey within the platform; select an available sample of digital twins; configure and run the simulation, and export analysis-ready data. In panel mode, researchers can assemble a small group of twins for open-ended conversations, document annotation, and moderated, focus-group-style voice discussions. We also developed CroissantTwin, a standardized data format for adding samples of digital twins to the platform. We demonstrate the survey mode workflow by using the platform to replicate 19 experiments on digital twins from the Twin-2K-500 dataset. ExploraTwin's survey execution fidelity is high: 99.6% of 197,000 answer units returned a structurally valid response on the first run.","authors":["Naveen Venkat","Yuchen Qiu","Tianyi Peng","George Gui","Olivier Toubia"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-26","first_seen":"2026-08-24","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.20539","pdf_url":"https://arxiv.org/pdf/2608.20539","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1"],"tags":["数字孪生","调查仿真","平台"],"reason":"平台支持数字孪生调查仿真，复现19个实验并与真实数据对照，验证执行保真度。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":3,"question":"如何降低研究者测试和部署数字孪生调查仿真的工程门槛与成本，并验证其执行保真度？","design":"开发开源非营利平台 ExploraTwin，支持上传 Qualtrics 问卷或平台内建问卷，选择数字孪生样本（如 Twin-2K-500），配置并运行仿真，导出分析就绪数据；同时提供面板模式进行开放式对话和焦点小组式讨论。","baseline":"使用 Twin-2K-500 数据集中的数字孪生，复现 Peng et al. (2025) 的 19 个实验，并与原始人类实验结果对照。","findings":"ExploraTwin 平台实现了低摩擦的数字孪生调查仿真工作流，支持复杂问卷逻辑和多种孪生样本。在复现 19 个实验时，197,000 个答案单元中 99.6% 在首次运行返回结构有效响应，执行保真度高。","reliability":"论文承认数字孪生的预测性能参差不齐，强调在特定情境部署前必须进行测试；平台当前免费但未来可能收费；未详细讨论仿真偏差或失效条件。","relevance":"该论文直接提供可用的数字孪生仿真平台，并复现多个实验验证保真度，对关注 LLM 人类仿真实验的研究者具有工具价值，值得阅读原文了解平台功能和验证细节。","inspiration":"借鉴其标准化问卷导入和自动验证修复流程，可大幅降低仿真实验的工程成本｜可用于经济学实验仿真，如消费者选择、公共品博弈、政策偏好调查等｜以数字孪生为被试，施加不同政策信息处理，测量选择或态度变化，并与真实人类实验数据（如实验室实验或调查数据）对照评估仿真效度"}},{"id":"2608.21296","version":2,"title":"Level-k Distinguishable Mechanisms for Evaluating Bounded Rationality in LLMs","zh_title":"评估LLM有限理性的Level-k可区分机制","abstract":"Strategic depth of reasoning is essential for human interaction of Large Language Models (LLMs) operating in boundedly rational environments. However, existing evaluations are primarily based on canonical games prevalent in pretraining corpora, making it difficult to disentangle true strategic reasoning from memorisation. To address this, we formalise a necessary level-K distinguishability condition for strategic depth inference and construct a suite of novel game structures that meet this standard. Using these games, we evaluate strategic depth in LLMs from both the Chain-of-Thought tokens and actual actions under recursive reasoning and an inductive trace of opponent game-play data. Across experimental trials spanning four LLMs, four game structures, and ten levels of iterated reasoning, we find that model models maintain accurate strategic depth under recursive reasoning, with strong internal consistency between stated reasoning and actions at every level. Errors arise from using the wrong number of iterated depth of reasoning steps, not from computing best responses incorrectly. However, inductive inference from opponent play degrades accuracy sharply and unevenly across games, and explicit strategic mentalizing in the chain of thought substantially improves overall performance.","authors":["Binchi Zhang","Atrisha Sarkar"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"replace","date":"2026-08-26","first_seen":"2026-08-24","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.21296","pdf_url":"https://arxiv.org/pdf/2608.21296","source_feed":"cs.MA","score":7,"bucket":"pending","rubric_hits":["A1","B3"],"tags":["LLM策略推理","博弈实验","有限理性"],"reason":"用LLM在博弈中模拟人类策略推理，虽无人类数据对照，但方法可迁移到人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-26","rank":2,"question":"如何设计满足水平可区分性的博弈结构，以可靠地推断大语言模型在有限理性环境中的策略推理深度？","design":"该研究不是人类仿真实验，而是用四个LLM（未具体命名）作为被试，在四种新设计的博弈结构（一种全新构造，三种改编自经典博弈）中，通过链式思考提示和实际动作，评估模型在递归推理和归纳推理两种条件下的策略深度，深度范围从0到10。","baseline":"无对照","findings":"在递归推理条件下，LLM能保持准确的策略深度，且链式思考中的推理与实际行动高度一致；错误主要源于使用了错误的迭代深度，而非最佳响应计算错误。在归纳推理条件下，准确率急剧下降且在不同博弈间不均衡，但显式的策略心智化能显著提升表现。","reliability":"论文未讨论","relevance":"该研究为评估LLM的策略推理能力提供了可区分性的博弈设计方法，虽无人类数据对照，但其方法可迁移到人类仿真实验中，用于校准LLM在策略互动中的行为。","inspiration":"值得借鉴的是通过设计满足水平可区分性的博弈结构来确保行为与推理深度一一对应，从而避免不可识别问题。｜可迁移到经济博弈实验，如拍卖、讨价还价或公共品博弈中的人类策略深度评估。｜可以用LLM作为被试，在满足可区分性的博弈中施加不同信息条件（如递归推理与归纳推理），测量其行动和链式思考，并与人类实验数据（如Camerer等人的行为博弈实验）对照，检验LLM是否复现人类策略深度分布。"}},{"id":"2608.23818","version":1,"title":"Beyond Static and Linear: What Attention Constraints Best Fit Human Reading Times?","zh_title":"超越静态与线性：何种注意力约束最拟合人类阅读时间？","abstract":"Transformer-based language models are widely used as models of human language processing, yet their attention mechanisms allow lossless access to the full preceding context, unlike the limited memory systems of humans. We hypothesize that installing memory constraints into transformers' attention mechanisms can improve their fit to human behavioral data. While previous work has explored individual constraints in isolation, we conduct a systematic comparison of multiple attention-based memory mechanisms across different model sizes and training corpora, evaluating both psychometric predictive power for human reading times and grammatical competence. We additionally compare static constraints, in which the constraint strength is fixed throughout training, to dynamic memory curricula. We find that constraints that are sensitive to the content of intervening tokens consistently achieve the highest alignment with human reading times, outperforming distance-based constraints. We observe a dissociation between psychometric fit and grammatical competence under dynamic memory curricula, suggesting that Transformers cannot serve as a one-size-fits-all cognitive model.","authors":["Lanni Bu","Xiulin Yang","Christian Clark","Alex Warstadt","Ethan Gotlieb Wilcox"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23818","pdf_url":"https://arxiv.org/pdf/2608.23818","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1"],"tags":["认知建模","注意力约束","心理测量拟合"],"reason":"用LLM拟合人类阅读时间，有真实人类数据对照，评估模型作为认知模型的可靠性，方…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-26","rank":5,"question":"在Transformer注意力机制中施加何种记忆约束能最好地拟合人类阅读时间，并考察静态与动态约束的差异。","design":"用不同规模（2层、4层）的OPT架构Transformer，在BabyLM-10M、BabyLM-100M和Pile-2B三个语料上训练，施加多种基于注意力的记忆约束（如距离衰减、内容敏感约束等），比较静态与动态课程（约束强度逐渐增强或减弱），以模型预测人类阅读时间的对数似然作为结果变量。","baseline":"六个英语阅读时间数据集：Brown、Natural Stories、UCL（自定步速阅读）和Dundee、GECO、Provo（眼动追踪），包含真实人类被试的逐词阅读时间。","findings":"内容敏感的注意力约束（如对干预词内容敏感的机制）在多数训练配置下比距离衰减约束更贴合人类阅读时间；动态记忆课程下出现心理测量拟合与语法能力的分离，静态模型预测阅读时间更好，而动态模型在语法基准上更强。","reliability":"论文指出Transformer不能作为一刀切的认知模型，动态课程下心理测量拟合与语法能力分离；但未系统讨论模型在其他语言行为或任务上的失效条件，也未分析训练数据偏差对结果的影响。","relevance":"该研究用LLM拟合人类阅读时间，有真实人类数据对照，并系统比较了多种记忆约束，对评估LLM作为人类认知模型的可靠性有直接参考价值，值得阅读原文以了解具体约束实现和动态课程设计。","inspiration":"借鉴其系统比较多种处理（记忆约束）并考察静态与动态施加方式的做法，以及用多个真实行为数据集做稳健性检验的思路。｜可迁移到经济金融中的信息处理约束研究，例如投资者对财务信息的注意力衰减或消费者对价格信息的记忆干扰。｜用LLM模拟投资者，施加不同注意力约束（如距离衰减或内容干扰），预测其对公司公告的反应时间或交易决策，并与真实投资者交易数据（如TAQ）或实验数据对照。"}},{"id":"2608.23640","version":1,"title":"Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes","zh_title":"审计合成回忆录：对照真实生活记录测量LLM生成自传中的场景级虚构","abstract":"When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366-day \"page-a-day\" book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis. We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.","authors":["Heather Renze"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23640","pdf_url":"https://arxiv.org/pdf/2608.23640","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","真实性审计","偏差评估"],"reason":"用LLM生成个人自传并与真实记录对照，评估虚构与偏差，方法可迁移到仿真可靠性研…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":5,"question":"当LLM被要求撰写一个人的自传时，其生成内容中有多少是真实发生过的？","design":"本研究不是群体仿真实验，而是对单个LLM生成自传的审计：作者用自己的366天“每日一页”生活记录作为生成对象，用对话式LLM（OpenAI o3-pro设计模板，gpt-4o/o3/o4-mini-high生成条目）仅基于模板、两个示例日和每日引言生成第一人称轶事条目，然后由作者本人依据独立验证语料库（回忆录手稿、已发表作品、演讲记录等）对每个条目进行场景级事实核查，采用预先固定的四级评分标准（VERIFIED/WEAK/UNVERIFIED/CONTRADICTED）进行评级。","baseline":"无对照。本研究没有将LLM生成结果与人类被试在相同任务下的表现进行对比，而是直接与作者本人的真实生活记录（独立验证语料库）进行核对。","findings":"在366个生成条目中，354个（96.7%）未能通过场景级验证，只有12个条目包含可被证实的场景；主要失败模式是“接地漂移”（grounded drift），即真实人物、雇主和场景被嵌入虚构情节中。通过将生成过程基于作者语料库进行接地（grounding），验证通过率显著提高，但仍存在83.3%的验证失败率。","reliability":"论文承认其审计工具的四级分类法（VERIFIED/WEAK/UNVERIFIED/CONTRADICTED）在评分者间信度上仅为一般到中等，特别是WEAK和UNVERIFIED之间的界限不可靠；此外，研究仅基于单个案例（作者本人），且生成过程中可能使用了未记录的对话上下文或预训练知识，这限制了结论的普遍性。","relevance":"该研究直接评估了LLM生成个人叙事时的虚构与偏差，并提供了可复用的审计工具和接地补救方法，对于关注LLM仿真可靠性及偏差的研究者具有方法论参考价值，值得阅读原文以了解详细的审计流程和失败模式分类。","inspiration":"这篇论文的场景级审计方法和接地补救实验值得借鉴，特别是其预先固定的评分标准和独立验证语料库的设计，可用于评估LLM在生成经济行为描述时的真实性。｜可以迁移到经济金融领域的政策评估或消费者行为研究中，例如让LLM模拟个体在特定经济政策下的决策叙述，然后与真实调查或行政数据进行核对。｜一个可行的设计是：以真实消费者为被试，收集其历史消费记录和调查回答作为基准；让LLM基于部分个人信息（如人口统计特征和少量消费摘要）生成该消费者在某个促销活动下的购买决策叙述；然后由独立评分者根据真实消费记录对生成内容进行场景级验证，结果变量为验证通过率和虚构类型分布，从而评估LLM仿真经济决策的可靠性。"}},{"id":"2608.23906","version":1,"title":"Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems","zh_title":"量化复杂社会技术系统中AI采纳的系统级危害","abstract":"Artificial Intelligence (AI) is increasingly integrated into complex sociotechnical systems, including Critical National Infrastructure (CNI), where harms emerge from interactions between technical, human, and organisational elements. Yet current AI evaluation remains model-centric, offering little insight into how observed behaviours might translate into system-level risk. We propose a framework that links structured hazard analysis, component-level testing, and probabilistic system modelling to bridge this gap. By providing a traceable pathway from model behaviour to system-level outcomes, the framework enables practitioners to answer the \"so what?\" of AI failures, quantify their systemic impact, and move toward evidence-based and anticipatory governance of AI in complex systems. Applied to the UK's Real Time Gross Settlement (RTGS) system as an illustrative worked example, we derive AI-driven loss scenarios using Systems Theoretic Process Analysis (STPA) and examine adversarial manipulation of LLM-based trading as one such loss scenario. Component-level experiments show that simple adversarial inputs induce measurable behavioural shifts where AI recommendations are followed. Under the component-to-system mapping used here for a financial contagion model, these shifts alter system resilience, increasing bank failures and lowering the threshold at which shocks lead to cascading disruption, particularly under widespread or monopolistic AI adoption.","authors":["Paul Vautravers","Oliver Chalkley","Gabriel Downer","Kate S","Damian Ruck"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23906","pdf_url":"https://arxiv.org/pdf/2608.23906","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","金融系统","系统性风险"],"reason":"用LLM agent模拟金融交易行为并与系统模型结合，评估AI采纳的系统性风险…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":7,"question":"如何将AI组件级行为与复杂社会技术系统（如英国RTGS）的系统级危害联系起来，以量化AI采纳带来的系统性风险？","design":"提出三阶段框架：结构化危害分析（STPA）、组件级测试（对LLM交易代理施加对抗性输入，测量其行为变化）、概率系统建模（金融传染模型）。以英国RTGS为案例，模拟AI交易代理在对抗性操纵下的行为，并将行为变化映射到系统模型，评估银行倒闭数量和级联中断阈值。","baseline":"无对照","findings":"对抗性输入能诱导LLM交易代理产生可测量的行为变化，且这些变化在系统层面降低了金融网络的韧性，增加了银行倒闭数量，并降低了引发级联中断的冲击阈值。在AI广泛或垄断性采纳的情况下，系统性风险加剧。","reliability":"论文未讨论","relevance":"该研究将LLM代理行为与系统级金融风险模型结合，展示了从个体行为到宏观后果的映射方法，对关注经济系统仿真和AI风险传导的研究者有参考价值，但缺乏真实人类数据对照，且案例特定于英国RTGS。","inspiration":"借鉴其将LLM代理行为嵌入系统动力学模型的做法，通过对抗性输入模拟行为偏差，并量化系统级后果。｜可迁移到金融市场稳定性分析，如AI交易员在压力情景下的行为如何影响市场流动性或波动性。｜设计：用LLM代理模拟交易员，施加对抗性新闻或市场操纵信息作为处理，测量交易行为（如买卖价差、交易量），并将行为输入到市场微观结构模型，与真实市场数据（如订单流、价格波动）进行校准和对照。"}},{"id":"2608.24046","version":1,"title":"Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment","zh_title":"算法影响揭示对齐的隐藏社会选择结构","abstract":"When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem of social choice: how should people's divergent preferences about system behavior be reconciled and aggregated into a single coherent model? The standard approach to aligning frontier AI models$\\unicode{x2013}$reinforcement learning from human feedback$\\unicode{x2013}$largely sidesteps this question and has poor social choice guarantees. However, it remains unclear what alternative should replace it. We show that, by focusing directly on an algorithm's welfare consequences, the alignment problem can be reformulated as linear optimization over a convex impact space, which makes it amenable to the standard toolkit of welfare economics and mechanism design. This reformulation clarifies how alignment protocols translate into welfare consequences and, conversely, how a social planner's desired constraints on welfare consequences can be translated back into alignment protocols. We apply this transformation to show that voting-by-issues and random-dictatorship mechanisms are strategyproof and unanimous. Demonstrating the reverse direction, we also apply the impact representation to derive a family of alignment protocols that maximize utilitarian social welfare subject to various social desiderata, such as bounds on individual or group harm. We illustrate the welfare implications of these alignment protocols empirically using real human preferences over kidney allocation, charitable food distribution, LLM responses, and trolley problems.","authors":["Zachary Wojtowicz","Michelle Si","Finale Doshi-Velez","Ariel Procaccia"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24046","pdf_url":"https://arxiv.org/pdf/2608.24046","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM对齐","社会选择","人类偏好"],"reason":"用真实人类偏好数据对齐LLM，涉及社会选择与福利，可迁移到仿真研究","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":8,"question":"如何将对齐问题重新表述为影响空间中的线性优化，以利用福利经济学和机制设计工具设计具有良好社会选择属性的对齐协议？","design":"该论文不是仿真研究，而是理论框架构建与实证演示。它提出将AI模型的影响表示为影响向量，将对齐问题转化为凸影响空间上的线性优化，并推导出满足策略证明和一致性的对齐机制（如议题投票、随机独裁），以及最大化功利主义福利且满足个体/群体伤害约束的机制。实证部分使用四个领域的真实人类偏好数据（肾脏分配、慈善食品分配、LLM响应、电车难题）来展示不同对齐协议的福利后果。","baseline":"无对照。论文使用真实人类偏好数据作为输入，但未与真实人类决策或行为进行对照。","findings":"通过影响空间重构，对齐问题可转化为线性社会选择问题，从而应用福利经济学和机制设计工具。论文证明了议题投票和随机独裁机制是策略证明且一致的，并推导出在个体或群体伤害约束下最大化功利主义福利的对齐协议族。","reliability":"论文未讨论其框架在真实部署中的失效条件或局限性，仅关注理论性质和理想化假设下的福利后果。","relevance":"该论文为LLM对齐提供了基于社会选择理论的严格框架，并利用真实人类偏好数据展示福利影响，对关注LLM仿真中偏好聚合和公平性的研究者具有重要参考价值，值得阅读原文以了解其理论细节和实证方法。","inspiration":"该论文将复杂模型参数空间映射到影响空间，使福利分析线性化，这种降维和重构方法可借鉴用于经济金融仿真中处理高维策略空间。｜可迁移到政策评估中的分配问题，如信贷审批、保险定价或公共资源分配，其中不同群体偏好冲突且需满足公平约束。｜设计一个实验：用LLM模拟不同收入群体的消费者，处理为不同的信贷审批算法（如最大化总福利、限制群体伤害），结果变量为各群体的贷款获得率和违约率，对照真实信贷数据中的群体差异和公平性指标。"}},{"id":"2608.23966","version":1,"title":"Who Chooses How Preferences Are Aggregated? Auditing Aggregation-Rule Authority in LLM-Based Group Recommendation","zh_title":"谁选择偏好如何聚合？审计基于LLM的群体推荐中的聚合规则权威","abstract":"AI systems increasingly make joint recommendations for users with conflicting preferences. However, when reasonable aggregation rules support different actions, a further question arises: who may choose how those preferences are combined? We study this interaction-level problem as aggregation-rule authority. Using synthetic preference profiles and profiles constructed from empirical ratings, we conduct a controlled behavioral audit of three LLMs under three authority conditions: unspecified, explicitly retained by users, and delegated to the model. In cases where two witness rules supported different actions, models almost never committed when users retained authority, but committed in every delegated case. All three models executed both witness rules perfectly when directly instructed. Yet when authority was unspecified or delegated, their aggregation-consistent outcome distributions differed across models and preference settings. Together, these results separate rule-execution capability from aggregation-rule authority: delegation assigns the model discretion to resolve the aggregation choice, but does not determine which collective outcome follows.","authors":["Yuxuan Du"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23966","pdf_url":"https://arxiv.org/pdf/2608.23966","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","群体决策","偏好聚合"],"reason":"用LLM模拟群体推荐中的偏好聚合，并与真实评分数据对照，涉及决策模式仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-26","rank":7,"question":"在群体推荐中，当不同聚合规则支持不同行动时，谁有权选择如何聚合偏好？","design":"用三个LLM（GPT-5.6 Sol、Claude Sonnet 5、Qwen 3.6 Plus）扮演群体推荐系统，在合成偏好档案和MovieLens实证评分档案上，施加三种权威条件（未指定、用户保留、委托给模型），测量模型是否承诺推荐及聚合一致结果。","baseline":"MovieLens 32M 实证评分数据","findings":"当用户保留权威时，模型几乎从不承诺；当权威委托给模型时，模型总是承诺。所有模型都能完美执行指定的聚合规则，但在权威未指定或委托时，聚合一致结果分布因模型和偏好设置而异。","reliability":"论文未讨论","relevance":"该研究用LLM模拟群体决策中的偏好聚合，并与真实评分数据对照，涉及决策模式仿真，对关注LLM仿真可靠性与偏差的研究者有参考价值。","inspiration":"可借鉴其通过明确权威分配和结构性对照来分离规则执行能力与自由裁量权的方法。｜可迁移到政策评估中的集体决策仿真，如委员会投票或公共品供给决策。｜用LLM扮演政策制定者，处理为是否明确赋予聚合规则选择权，结果变量为政策选择与聚合规则一致性，对照真实委员会决策记录。"}},{"id":"2507.21790","version":3,"title":"Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities","zh_title":"大语言模型能否辅助选择建模？对提示策略和当前模型能力的洞察","abstract":"Large Language Models (LLMs) are becoming widely used to support various workflows across different disciplines, yet their potential in discrete choice modelling remains relatively unexplored. This work examines the potential of LLMs as assistive agents in the specification and, where technically feasible, estimation of Multinomial Logit models. We implement a systematic experimental framework involving twelve versions of seven leading LLMs (ChatGPT, Claude, DeepSeek, Gemini, Gemma, Llama, and Mistral) evaluated under five experimental configurations. These configurations vary along three dimensions: (i) modelling goal (suggesting vs. suggesting and estimating MNL models); (ii) prompting strategy (Zero-Shot vs. Chain-of-Thoughts (CoT)); and (iii) information availability (full dataset vs. data dictionary summarising variable names and types). Each specification suggested by the LLMs is implemented, estimated, and evaluated based on goodness-of-fit metrics, behavioural plausibility, and model complexity. Our findings reveal that proprietary LLMs can generate valid and behaviourally sound utility specifications, particularly when guided by structured prompts (CoT). Open-weight models such as Llama and Gemma struggled to produce meaningful specifications. Notably, some LLMs performed better when provided with just data dictionary, suggesting that limiting raw data access may enhance internal reasoning capabilities. Among all LLMs, GPT o3, operating in an agentic setting, was uniquely capable of correctly estimating its own specifications by executing self-generated code. Overall, the results demonstrate both the promise and current limitations of LLMs as assistive agents in discrete choice modelling, not only for model specification but also for supporting modelling decision and estimation, and provide practical guidance for integrating these tools into choice modellers' workflows.","authors":["Georges Sfeir","Gabriel Nova","Stephane Hess","Sander van Cranenburgh"],"categories":["econ.EM","cs.AI"],"primary_category":"econ.EM","announce_type":"replace-cross","date":"2026-08-26","first_seen":"2025-07-29","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2507.21790","pdf_url":"https://arxiv.org/pdf/2507.21790","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助建模","离散选择模型","提示策略"],"reason":"LLM辅助离散选择建模，替代部分建模工作，但非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":198,"question":"当前的大语言模型能否作为辅助工具，帮助研究者完成离散选择模型（MNL）的设定与估计？","design":"本研究并非用LLM仿真人类被试，而是将12个LLM版本作为建模助手，在5种实验配置下（零样本/思维链提示、完整数据/数据字典、仅建议设定/建议并估计模型）生成MNL效用函数设定，再由研究者实现并评估其拟合优度、行为合理性与复杂度。","baseline":"无对照","findings":"闭源LLM在结构化提示下能生成有效且行为合理的效用设定，而开源模型表现较差；部分LLM仅凭数据字典反而表现更好，GPT o3是唯一能通过自写代码正确估计自身设定的模型。","reliability":"论文指出研究仅限MNL模型，未涉及更复杂设定；LLM版本时效性强，结果可能随模型更新而变化；未采用少样本提示，且未评估LLM在真实建模迭代中的表现。","relevance":"该研究聚焦LLM辅助建模者劳动，而非替代人类被试进行行为仿真，与您关注的LLM作为人类被试替代品的研究方向不直接相关，不建议优先阅读原文。","inspiration":"该研究通过对比不同提示策略（零样本/思维链、完整数据/数据字典）和任务复杂度（仅建议设定/建议并估计）来系统评估LLM能力，这种多维度实验设计值得借鉴｜可迁移至政策评估中的离散选择建模，例如利用LLM辅助设定交通方式选择或疫苗接受度的效用函数｜以LLM作为建模助手，处理为不同提示策略（如仅提供数据字典vs完整数据），结果变量为生成的效用函数拟合优度与行为合理性，对照真实研究者手工设定的模型性能"}},{"id":"2608.23005","version":1,"title":"Large language models simulate intersectional synthetic identities with a budget of one to two dimensions","zh_title":"大语言模型以一到两个维度的预算模拟交叉性合成身份","abstract":"Large language models are increasingly used as synthetic survey respondents, promising cheap access to rare intersectional populations. We test standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel -- 21 million simulated response distributions from eight models. In real respondents, subgroup opinion is approximately the additive sum of its single-identity components, yet grows 2.5x more distinctive as identities intersect. Simulated respondents show no such composition: a single feature explains a two-feature persona's responses better than the additive combination in 75-82% of subgroups, and a third feature adds almost nothing. This collapse survives every prompting strategy we test. Additionally, the feature models retain is chosen nearly blindly -- except that they systematically discard race and religion, the strongest real drivers of opinion. Synthetic samples offer intersectional personas but represent one identity at a time.","authors":["Virgile Rennard","Christos Xypolopoulos"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23005","pdf_url":"https://arxiv.org/pdf/2608.23005","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","交叉性","算法保真度"],"reason":"直接测试LLM作为合成调查受访者，并与真实Pew数据对照，发现仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":1,"question":"LLM 在模拟交叉身份（如黑人共和党人）的调查回答时，是否真正整合了多个身份特征，还是只保留其中一个？","design":"用 8 个 LLM（从 7B 开源模型到前沿系统）扮演具有 1 到 3 个特征的人口统计画像，生成 15 波 Pew 美国趋势面板中所有问题的回答分布，共 2100 万次模拟；通过比较单特征、双特征和三特征画像的模拟偏差与真实子群体偏差，检验模型是否对多个身份特征进行加性组合。","baseline":"Pew 美国趋势面板 15 波调查的受访者微观数据，计算每个真实交叉子群体（至少 20 名受访者）的回答分布作为基准。","findings":"真实子群体的意见近似于其单身份成分的加性组合，且随身份交叉而更加独特；但模拟受访者没有这种组合，一个特征就能比加性组合更好地解释双特征画像的回答，第三个特征几乎不增加信息。模型保留的特征几乎是盲目选择的，且系统性地丢弃了种族和宗教这两个真实意见的最强驱动因素。","reliability":"论文通过人类抽样噪声下限、相似度度量的零校准和完全分半样本确认来确保结果稳健；但未讨论提示策略之外的失效条件，也未涉及开放式文本或非分布层面的交叉性。","relevance":"该研究直接测试 LLM 作为合成调查受访者的可靠性，并与真实 Pew 数据对照，发现仿真在交叉身份下系统性失效，对关注仿真偏差和失效条件的研究者极具参考价值。","inspiration":"借鉴其用真实微观数据构建交叉子群体基准、并通过偏差签名竞争来识别模型实际使用的特征的方法｜可迁移到信贷审批歧视研究，检验 LLM 模拟的交叉群体（如黑人女性）的信贷决策是否只基于单一特征｜用 LLM 扮演不同种族和性别的贷款申请人，生成信贷审批决策，与真实信贷数据（如 HMDA）中对应交叉群体的审批率分布进行对照，检验模型是否丢弃了种族或性别信息。"}},{"id":"2608.22582","version":1,"title":"Hybrid Panels: Toward Human-AI Collaboration in Survey Research","zh_title":"混合面板：迈向调查研究中的AI协作","abstract":"Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, including declining response rates, increasing data collection costs, long delays between data collection and data provision, and the risk of nonresponse bias. Advances in artificial intelligence (AI) have opened up new opportunities for AI-supported survey infrastructures where the goal is to overcome these challenges without limiting the data quality. A promising AI-enabled survey infrastructure for which we build a first pilot is a hybrid panel. A hybrid panel is a longitudinal AI-enabled survey which allows to iteratively improve the alignment between large language models (LLMs) and the population they aim to simulate and use the errors to inform the design and implementation of the next survey wave (e.g., inform the participant recruitment, assignment of questions to participants). It incorporates both human participants and LLMs as fundamental elements of its design. In this research note, we introduce the concept of a hybrid panel by providing a definition and outlining an overarching framework, spanning data collection to data validation. We detail results from a first pilot study to illustrate (open) challenges that we identify for hybrid panels.","authors":["Julia Romberg","Tobias Gummer","Gabriella Lapesa","Tanja Kunz","Claudia Wagner"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22582","pdf_url":"https://arxiv.org/pdf/2608.22582","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3"],"tags":["LLM仿真","调查方法","人机协作"],"reason":"提出混合面板，用LLM模拟调查对象并与人类数据对照，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":4,"question":"如何设计一种结合人类被试与LLM的纵向调查基础设施（混合面板），以在保证数据质量的同时应对传统调查的挑战？","design":"提出混合面板概念，即纵向调查中同时纳入人类被试和LLM，通过计划缺失设计、LLM插补、人类验证AI生成回答等方式结合两者；先导研究聚焦人类被试招募步骤，未报告完整实验处理与结果变量。","baseline":"无对照（先导研究仅涉及招募，未提供人类与LLM回答的直接对比数据）。","findings":"论文提出了混合面板的定义和框架，并展示了先导研究中人类被试招募的初步结果；识别出混合面板面临的开放性挑战，包括人类被试的偏好与权利、AI模型的持续评估以及社区标准制定。","reliability":"论文承认LLM与人类行为存在不匹配，完全合成面板面临透明度、问责制和伦理问题；强调AI不能完全取代人类作为研究对象，混合面板的适用性需要长期评估和实验。","relevance":"该研究直接针对LLM仿真人类调查的可靠性问题，提出混合面板以结合人类与AI数据，并强调持续验证，对关注仿真偏差和真实人类对照的研究者具有重要参考价值。","inspiration":"可借鉴其混合面板设计，将LLM生成回答与人类被试数据结合，通过计划缺失和迭代验证提高仿真准确性｜可迁移到经济预期调查或消费者信心指数构建，利用LLM补充缺失回答并校准偏差｜设计一个纵向调查，招募真实消费者作为被试，部分问题由LLM回答，处理为不同提示策略，结果变量为回答与真实值的偏差，用官方统计或面板数据做对照。"}},{"id":"2608.21668","version":1,"title":"From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation","zh_title":"从掌握水平画像到模拟响应：用于忠实LLM学生仿真的随机学生知识图谱","abstract":"Large language models (LLMs) are increasingly used to simulate students at different mastery levels. These simulations can generate synthetic training data and stress-test tutoring systems. However, common prompt-based approaches leave the answer decision to the LLM, which tends to perform according to its built-in capabilities even when instructed to simulate a student with low mastery. As a result, these approaches may have difficulty distinguishing students with low and high levels of mastery. We demonstrate this limitation using 379 College Board-calibrated SAT Algebra items and five archetypal mastery profiles. Three LLMs from three vendors (Gemini 3.1 Flash Lite, Claude Haiku 4.5, and GPT-5.4-mini) achieve 96.8-100% accuracy across all profiles. To address this limitation, we introduce a method grounded in a Stochastic Student Knowledge Graph (SSKG). A curriculum knowledge graph (CKG) is extracted from an open algebra textbook, and each SAT solution is decomposed into a chain of required triples. The SSKG assigns a mastery probability to each triple, which is sampled to determine question correctness. An LLM then generates a first-person rationale consistent with the outcome. The simulation reduces accuracy to 44.1-85.2% across profiles and produces a clear monotone mastery gradient.","authors":["Yuan An","Emily Wang","Benjamin Wang","Ruhma Hashmi"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21668","pdf_url":"https://arxiv.org/pdf/2608.21668","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM学生仿真","知识图谱","教育评估"],"reason":"用LLM仿真不同掌握水平的学生，并与真实SAT数据对照，评估仿真保真度并指出提…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":2,"question":"如何让LLM忠实模拟不同掌握水平的学生，使其答题准确率呈现与掌握水平一致的梯度，并产生可归因于具体知识点的错误？","design":"用三种商用LLM（Gemini 3.1 Flash Lite、Claude Haiku 4.5、GPT-5.4-mini）模拟五种掌握水平的学生（从接近专家到严重知识缺口），在379道SAT代数题上作答。先测试直接提示法（在提示中描述学生水平），再提出基于随机学生知识图谱（SSKG）的方法：从代数教材构建课程知识图谱，将每道题的解题过程分解为所需三元组链，为每个三元组赋予掌握概率，通过采样决定答题正确性，再由LLM生成与结果一致的第一人称解释。通过四个累积消融臂（单次采样、检索/执行分解、分类加权、干扰项路由）评估各机制贡献。","baseline":"无直接人类对照，但使用379道College Board校准的SAT代数题作为题目基准，题目难度和区分度经过真实考生数据校准。","findings":"直接提示法下，三个LLM在所有掌握水平上的准确率高达96.8%-100%，无法区分高低水平学生。SSKG方法将准确率降至44.1%-85.2%，并产生清晰的单调掌握梯度，且错误可归因于特定知识点。","reliability":"论文承认技能特异性和涌现难度的结果较为混合：P4显示早期/后期知识点的分化，P5受链长影响，且仅部分消融臂达到预期。此外，SSKG方法依赖于人工构建的课程知识图谱和解题链，可能引入主观性，且仅在SAT代数题上验证，泛化性未知。","relevance":"该研究直接针对LLM仿真人类被试的保真度问题，通过引入外部知识结构控制LLM行为，克服了提示法中的能力偏差，对评估仿真可靠性和设计更可控的仿真方法有重要参考价值。","inspiration":"借鉴其将决策过程分解为知识单元并显式采样控制行为的方法，可迁移到经济金融中的个体决策仿真，如消费者跨期选择或投资者风险偏好。｜例如，在信贷审批歧视研究中，可构建金融知识图谱，将贷款决策分解为所需金融概念，为不同金融素养水平的虚拟申请人赋予掌握概率，通过采样决定其决策结果，再让LLM生成解释。｜设计：以LLM模拟不同金融素养的贷款申请人，处理是金融素养水平（通过知识图谱掌握概率设定），结果变量是贷款申请决策（是否违约或选择何种贷款），对照真实数据可使用美国消费者金融保护局（CFPB）的投诉数据或某银行的历史贷款数据。"}},{"id":"2608.22438","version":1,"title":"When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing","zh_title":"当人格模拟具有信息量时：用于多元意见感知的图结构信号","abstract":"Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response variation can reflect unconditioned model priors or token sampling noise rather than systematic persona conditioning. We argue that persona-conditioned variation is informative when semantically similar personas exhibit concordant response shifts. To operationalize this principle, we introduce Persona-Conditioned Informativeness (PCI), an unsupervised diagnostic metric that measures whether semantically similar personas deviate in concordant directions relative to item-level sample baselines. By modeling personas as a similarity graph, PCI uses Local Moran's I to quantify local spatial coherence and extract compact persona subsets without using construct labels. To evaluate PCI without external human benchmarks, we test its ability to recover established latent value structure using the 57-item Portrait Values Questionnaire-Revised (PVQ-RR). Confirmatory factor analysis (CFA) shows that a PCI-selected 10% subset substantially improves overall construct recovery relative to response-stability and random selection. These findings support PCI as a principled internal diagnostic for screening synthetic respondents in survey pipelines.","authors":["Taehyeon An","Jaehyeong Park","Donghyuk Shin"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22438","pdf_url":"https://arxiv.org/pdf/2608.22438","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"用LLM模拟调查回答，提出诊断指标筛选合成被试，并用真实人类数据验证。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":3,"question":"如何判断基于人物设定的大语言模型生成的调查回答是否真正反映了人物设定，而非模型先验或采样噪声？","design":"使用 K-EXAONE-236B 模型，为 1480 个合成人物设定生成对 57 项 PVQ-RR 价值观问卷的回答；通过构建人物相似度图并计算局部空间相干性（PCI 指标）来筛选信息量高的子集。","baseline":"无外部人类基准；以 PVQ-RR 的潜在价值结构作为内部验证标准。","findings":"PCI 筛选出的 10% 子集在验证性因子分析中显著提升了潜在构念恢复度，优于响应稳定性选择和随机选择。这表明语义相似的人物设定在回答上呈现一致的偏移，可作为筛选合成被试的有效内部诊断。","reliability":"论文指出 PCI 仅提供内部结构诊断，不能保证外部总体有效性，需结合外部校准；当前图构建基于全局嵌入，可能忽略不同题目涉及的属性子集；仅在单一模型和单一问卷上验证，跨模型和跨语言泛化性未知。","relevance":"该研究直接针对 LLM 模拟调查回答的可靠性问题，提出无监督诊断指标，并用真实人类价值观结构进行验证，对关注仿真有效性和偏差的研究者具有重要参考价值。","inspiration":"借鉴其基于相似度图的局部空间相干性度量，可无监督地识别哪些合成个体对处理变量有系统性响应，避免盲目使用全部生成样本。｜可迁移到经济政策偏好调查或消费者态度仿真中，筛选出对政策参数或产品属性有真实差异化反应的合成被试。｜用 LLM 生成不同人口统计特征的人物设定，施加政策干预（如税收变化），测量其政策支持度，并用真实调查数据（如美国综合社会调查 GSS）校准筛选后的合成样本分布。"}},{"id":"2608.12750","version":2,"title":"PatientAct: Theory-Grounded Mental Health Client Simulation","zh_title":"PatientAct：基于理论的心理健康来访者仿真","abstract":"LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept therapeutic reframes without resistance, and resolve core issues within a single session. We trace these issues to profiles that lack causal depth and behavioral mechanisms that treat all content as equally accessible. We present PatientAct, a framework for client simulation grounded in established clinical theories. Our profiles integrate the 5Ps clinical case formulation, providing causal depth without tying the design to any single therapeutic modality. During simulation, profiles include a dynamic memory layer in which items carry trust thresholds (e.g., symptoms are available early, whereas formative memories require a sustained therapeutic alliance). At each turn, the client's emotional reaction and behavior are modeled before generating a response. If the therapist approaches gated content, PatientAct expresses resistance in terms of quantity, content, and style rather than defaulting to cooperation or a single resistance pattern. We evaluate our framework on 40 clinical situations and demonstrate that it generates diverse profiles with high clinical plausibility. Moreover, PatientAct significantly outperforms the baselines, yielding substantial gains in resistance quality and behavioral realism. Our code and data are publicly available via github.com/Sahandfer/PatientHub.","authors":["Sahand Sabour","TszYam NG","Yaqian Chen","Guanqun Bi","Jialu Zhao","Minlie Huang"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-25","first_seen":"2026-08-14","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2608.12750","pdf_url":"https://arxiv.org/pdf/2608.12750","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","心理治疗","行为真实性"],"reason":"用LLM模拟心理治疗来访者，有真实临床情境对照，评估行为真实性与抵抗质量，可迁…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":6,"question":"如何设计基于LLM的心理治疗来访者仿真，使其行为更接近真实来访者，特别是能表现出基于信任的披露和多样化的抵抗？","design":"PatientAct框架使用GPT-5.4生成基于5Ps临床案例构想的来访者档案，并加入动态记忆层和信任阈值，在每轮对话中先模拟情绪反应和行为选择，再生成回应；在40个临床情境（抑郁和焦虑各20个）上评估其临床合理性和抵抗质量。","baseline":"无对照","findings":"PatientAct生成的档案具有高临床合理性和多样性，且在抵抗质量和行为真实性上显著优于现有基线。","reliability":"论文未讨论","relevance":"该研究通过理论驱动的档案设计和动态信任机制提升LLM仿真行为真实性，并采用专家评估验证，对关注仿真可靠性和偏差的研究者有参考价值，值得阅读原文了解具体实现和评估细节。","inspiration":"借鉴其将理论构念（如信任阈值、抵抗分类）嵌入仿真机制并设计多维度评估指标的做法｜可迁移到经济金融中的信任与信息披露场景，如消费者对金融顾问的信任建立、投资者对风险信息的逐步接受｜设计一个LLM扮演的投资者，处理为不同信任阈值下的信息提供策略，结果变量为披露意愿和风险感知，对照真实投资者调查或实验数据。"}},{"id":"2608.21401","version":1,"title":"Generative Gap Filling","zh_title":"生成式填补空白","abstract":"Most contract litigation turns on contracts that imperfectly record parties' bargains. When the parties' dispute can't be solved by interpreting the text, courts fill the gap. Scholars have long assumed that the remaining text runs out quickly, and provides thin evidence of the actual deal on the disputed point. On that view, a judge who supplies the missing term must be drawing on something else, from commercial defaults to her own policy preferences. Despite generations of work, courts have no real alternative to such unruly methods. We tested that assumption. Taking real contracts, we masked a term the parties had negotiated and asked readers to predict what we removed. Lay respondents recovered the hidden term about half the time, twice what chance predicts. Law students and lawyers did marginally better. But large language models, given nothing but the rest of the contract, recovered it nearly nine times in ten. The deal, in short, testifies to far more of the agreement than the literature assumes, including terms the parties never wrote. A contract, we argue, is like a radio signal from far away. Even when incomplete, enough of the message is carried elsewhere that the missing part can be reconstructed with the right receiver. True gaps are rarer than supposed. Courts can weigh model predictions as ordinary, contestable evidence, and parties can discipline the practice with \"Choice of Model\" clauses.","authors":["Yonathan A. Arbel","David A. Hoffman"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21401","pdf_url":"https://arxiv.org/pdf/2608.21401","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","法律决策","人类对照"],"reason":"用LLM预测人类对合同缺失条款的判断，并与真人对照，属于法律决策仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":7,"question":"合同文本在多大程度上能揭示当事人未明确写出的条款，从而让法院或模型填补合同空白？","design":"从真实合同中遮蔽一个已协商的条款，让普通人、法学生、律师和大型语言模型仅根据合同其余部分预测被遮蔽的条款，比较预测准确率。","baseline":"普通人（Lay respondents）预测准确率约50%，法学生和律师略高，作为人类对照基准。","findings":"大型语言模型仅凭合同其余部分，预测被遮蔽条款的准确率接近90%，远高于人类。合同文本比传统假设包含更多关于未写明条款的信息，真正的合同空白比想象中更少。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类对合同缺失条款的判断，并与真人对照，属于法律决策仿真，对关注LLM仿真可靠性和偏差的研究者有参考价值，值得阅读原文了解实验细节和局限。","inspiration":"借鉴其遮蔽真实合同条款并让模型预测的设计，可迁移到金融合同或政策文本的缺失条款预测，例如信贷协议中的利率调整条款或政策公告中的具体参数。｜可应用于资产定价实验或信贷审批歧视研究，例如遮蔽贷款合同中的关键条款，检验模型能否预测人类决策者会如何设定或接受这些条款。｜用LLM作为被试，遮蔽真实金融合同中的某个条款，让模型预测该条款内容，并与真实合同条款及人类专家预测对照，结果变量为预测准确率，真实数据来自公开的合同数据库或监管文件。"}},{"id":"2510.05545","version":3,"title":"Can Language Models Boost the Power of Randomized Experiments Without Statistical Bias?","zh_title":"语言模型能否在不引入统计偏差的情况下提升随机实验的功效？","abstract":"Randomized controlled trials (RCTs) are widely adopted for causal inference, yet cost and sample-size constraints limit power. We introduce CALM (Causal Analysis leveraging Language Models), a statistical framework that integrates insights generated by large language models (LLMs) into the analysis of RCTs using established causal estimators to increase precision while preserving statistical validity. In particular, CALM treats LLM-generated outputs as auxiliary prognostic information and corrects their potential bias via a heterogeneous calibration step that residualizes and optimally reweights predictions. We prove that CALM remains consistent even when LLM predictions are biased and achieves efficiency gains over augmented inverse probability weighting estimators for various causal estimands. In particular, CALM develops a few-shot variant that aggregates predictions across randomly sampled demonstration sets. The resulting U-statistic-like predictor restores i.i.d. structure and also mitigates prompt-selection variability. Empirically, in simulations calibrated to a mobile-app depression RCT, CALM delivers lower variance relative to other benchmarking methods, is effective in zero- and few-shot settings, and remains stable across prompt designs. By principled use of LLMs to harness unstructured data and external knowledge learned during pretraining, CALM provides a practical path to more precise causal analyses.","authors":["Xinrui Ruan","Xinwei Ma","Yingfei Wang","Waverly Wei","Jingshen Wang"],"categories":["stat.ME","econ.EM"],"primary_category":"stat.ME","announce_type":"replace-cross","date":"2026-08-25","first_seen":"2025-10-07","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2510.05545","pdf_url":"https://arxiv.org/pdf/2510.05545","source_feed":"econ.EM","score":7,"bucket":"pending","rubric_hits":["A2","B1","B3"],"tags":["因果推断","LLM辅助分析","统计方法"],"reason":"用LLM辅助RCT分析，校正偏差并提升精度，有真实数据对照，方法可迁移到仿真评…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:06","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":8,"question":"如何利用大语言模型生成的预测来提升随机对照试验的统计功效，同时避免引入统计偏差？","design":"提出CALM框架，将LLM对潜在结果的预测作为辅助预后信息，通过残差化和异质性校准加权校正偏差，并采用少样本学习与演示集重采样聚合预测，用于RCT的因果效应估计。","baseline":"模拟研究校准自一项移动应用抑郁症RCT的真实数据，并重新评估了一项行为试验。","findings":"CALM在模拟中相比基准方法具有更低的方差，且在零样本和少样本设置下有效，对提示设计具有稳健性。理论上证明即使LLM预测有偏，CALM仍保持一致，并比增强逆概率加权估计器更高效。","reliability":"论文承认LLM预测可能存在偏差且精度异质，但通过残差化和校准加权处理；少样本学习对示例选择敏感，通过重采样聚合缓解。未讨论其他失效条件。","relevance":"该研究将LLM作为辅助工具提升RCT分析精度，而非直接模拟人类被试，但涉及LLM预测的偏差校正和与真实数据的对照，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"借鉴其残差化与异质性校准加权方法，可校正LLM预测偏差并提升估计效率。｜可迁移到政策评估中利用LLM提取非结构化协变量信息，如文本评论、新闻等，提升处理效应估计精度。｜以LLM对个体潜在结果的预测作为辅助变量，在信贷审批歧视研究中，用真实贷款数据校准LLM预测，比较不同校准方法下的处理效应估计。"}},{"id":"2608.23047","version":1,"title":"Beyond Verdicts: A Graph-Based Analysis of Human and LLM Reasoning in Scientific Fact-Checking","zh_title":"超越裁决：科学事实核查中人类与LLM推理的图分析","abstract":"Misinformation that cites legitimate papers can be especially harmful when it distorts what those studies actually report. While existing automatic fact-checking systems based on large language models (LLMs) can assess whether a model assigns an Incorrect verdict and can gen- erate explanations for that decision, they typi- cally do not indicate whether the model follows the same reasoning path as human experts or arrives at the verdict through a different but still valid path. In this work, we introduce a graph- based framework (typed reasoning graph) for comparing human and LLM reasoning paths in scientific fact-checking. Building on prior work on fallacious reasoning in biomedical misinformation, MISSCIPLUS (Glockner et al., 2025), we model each explanation as a rea- soning graph that links the false claim to the relevant study context, study findings, fallacy- supporting premises, and fallacy labels. This representation enables one-to-one alignment of human and LLM reasoning at the level of fallacy-specific sub-graphs. For non-human- aligned LLM paths, we validate grounding in the cited study, relevance to the claim, and suf- ficiency for the verdict. Using 84 false claims from MISSCIPLUS, we evaluate GPT-5, Claude Opus 4.7, and Qwen3-32B across prompt and evidence settings. Results show distinct perfor- mance dimensions: Qwen3-32B has the lowest verdict failure rate, GPT-5 the highest human alignment, and Claude Opus 4.7 weak verdict prediction but often valid reasoning in success- ful cases","authors":["Abdul Ghafoor","Muhammad Arslan Manzoor","Yufang Hou"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23047","pdf_url":"https://arxiv.org/pdf/2608.23047","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM推理对齐","事实核查","人类对照"],"reason":"比较人类与LLM推理路径，评估对齐度与有效性，有真实人类专家数据对照，批判性指…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":12,"question":"在科学事实核查中，LLM 的推理路径是否与人类专家一致？如果不一致，其替代推理路径是否仍有效？","design":"该研究不是人类仿真实验，而是提出一种图结构（typed reasoning graph）来比较人类专家与 LLM 在科学事实核查中的推理路径。对 84 条虚假声明，让 GPT-5、Claude Opus 4.7、Qwen3-32B 生成解释，将解释建模为推理图，与人类专家推理图在谬误特定子图层面进行对齐；对未对齐路径，人工验证其是否基于引用研究、与声明相关且足以支持结论。","baseline":"人类基准来自 MissciPlus 数据集中由人类专家标注的推理路径（包括相关研究背景、研究发现、谬误支持前提和谬误标签）。","findings":"模型在结论准确性和推理路径质量上存在分离：Qwen3-32B 结论失败率最低，GPT-5 人类对齐率最高，Claude Opus 4.7 结论预测弱但成功案例中推理常有效。即使提供完整研究和详细提示，人类对齐推理仅占少数（GPT-5 为 32.1%，Claude Opus 4.7 为 8.3%，Qwen3-32B 为 15.5%），但许多未对齐路径经验证仍有效，表明 LLM 常通过替代推理路径得出正确结论。","reliability":"论文指出，评估推理路径的 grounding 和 sufficiency 通常需要领域专业知识，因此 LLM 系统应作为解释辅助工具而非独立仲裁者。此外，人类对齐率低可能源于专家推理路径的多样性或标注的主观性，但论文未深入讨论这些局限。","relevance":"该研究提供了比较人类与 LLM 推理过程的结构化方法，对关注 LLM 仿真可靠性的研究者有方法论价值，但场景限于科学事实核查，与经济学实验和政策评估的直接关联较弱。","inspiration":"借鉴其将推理过程分解为可对齐的图结构并区分结论与过程质量的做法，可用于评估 LLM 在经济决策中的推理是否与人类专家一致。｜可迁移到政策评估场景，如分析 LLM 对经济政策公告的解读是否遵循专家逻辑。｜以经济学研究者为被试，让 LLM 对同一政策文本生成推理，用图结构对齐人类与 LLM 的推理路径，以专家标注为基准，并验证未对齐路径的合理性。"}},{"id":"2608.22887","version":1,"title":"Proxy reliance in large language model decisions is uncalibrated to predictive evidence","zh_title":"大语言模型决策中的代理依赖与预测证据不校准","abstract":"Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carry predictive value, so a changed decision can be discrimination or sound inference. We measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth, where the reliance the evidence warrants can be computed exactly and used as the reference. One audit signal yields three verdicts: over-reliance, warranted and under-reliance. Under neutral labels every model relies on proxies with no information. Informative proxies draw all three. Social field names push reliance down, below the reference in one model. Two findings explain this. Reliance severely undertracks the evidence, and social-label suppression is fragile, since in-context examples raise it above zero in every model. Accuracy-based evaluation detects none of this.","authors":["Zengqing Wu","Chuan Xiao"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22887","pdf_url":"https://arxiv.org/pdf/2608.22887","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM决策偏差","算法审计","代理变量"],"reason":"评估LLM决策中的代理依赖与偏差，与仿真可靠性相关，但非直接仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":11,"question":"大语言模型在决策中对代理属性的依赖是否与预测证据所支持的水平相校准？","design":"在完全已知的合成数据生成过程中，四个大语言模型（Claude Sonnet 4.5、DeepSeek-V4-Flash-0731、Qwen3.7-max、GPT-5.6 Terra）扮演分诊专家，对模拟患者对进行优先级排序。通过因果模拟翻转受保护属性并仅通过代理属性传播，测量模型排序变化的代理特定效应，并与贝叶斯最优决策者和理想学习者（基于模型所见示例拟合的贝叶斯回归）的效应进行比较。","baseline":"无对照","findings":"在无信息代理下，所有模型都表现出非零依赖；随着代理预测价值增加，依赖严重低于证据支持的水平。社会领域名称会抑制依赖，但上下文示例会削弱这种抑制，且基于准确率的评估无法检测到这些偏差。","reliability":"论文指出，证据支持水平的计算依赖于已知的生成过程，在观测数据中无法识别；合成人群虽通过七个公共临床数据集锚定，但结论可能不完全适用于真实部署。","relevance":"该研究通过因果框架量化LLM决策中的代理依赖与证据校准，为评估LLM作为人类被试替代品时的决策偏差提供了方法论参考，但未直接复现人类行为，与仿真可靠性相关但非直接仿真。","inspiration":"借鉴其因果代理效应测量和理想学习者基准，可迁移到信贷审批中的代理歧视问题（如使用邮政编码作为种族代理）。设计上，以LLM作为信贷审批员，处理为改变申请人的邮政编码（与种族相关但含真实信用信息），结果变量为贷款批准决策，并用真实信贷数据（如Home Mortgage Disclosure Act数据）中人类审批员的决策作为对照，比较LLM与人类对代理信息的依赖程度。"}},{"id":"2608.23196","version":1,"title":"AI emotional support is better only when chosen, but shifts preferences even when it is not","zh_title":"AI情感支持仅在主动选择时更优，但即使非主动选择也会改变偏好","abstract":"People increasingly face a novel decision when seeking emotional support: human or AI. In existing studies, AI's empathic messages are rated as well as or better than humans'. But these studies either assigned the support source or honored people's choice. In real life, support is often incongruent with choice, as people want one source and receive the other. Across three experiments (N = 1,951), participants chose whether to share an emotional experience with a human or an AI, then were randomly assigned to a congruent or incongruent partner. AI support was rated as superior only among those who had chosen it. Yet regardless of congruence, interacting with AI increased willingness to choose it again. In a 28-day study with OpenAI (N = 981), daily conversations shifted preferences toward AI and away from humans, but only when conversations turned personal. Emotional support choices are thus path-dependent, progressively redirecting away from human connection.","authors":["Yaoxi Shi","Cathy Mengying Fang","Guy LabanPattie Maes","Amit Goldenberg"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23196","pdf_url":"https://arxiv.org/pdf/2608.23196","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","情感支持","人机交互"],"reason":"用LLM替代人类提供情感支持，与真实人类对照，探讨选择与偏好变化，可迁移至仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":14,"question":"人们在寻求情感支持时，选择人类还是AI的信念如何影响其选择，以及选择一致或不一致如何影响对支持的体验和后续选择偏好？","design":"该研究不是用LLM仿真人类，而是让真人参与者选择与人类或AI（GPT-4）进行情感支持对话，随机分配到与选择一致或不一致的伙伴，测量支持质量评价和后续选择意愿；另有一个28天纵向研究，让参与者每天与AI对话，追踪偏好变化。","baseline":"人类基准是真人参与者与人类支持者（众包工人）的对话，作为与AI支持对照的真实人类数据。","findings":"AI支持只有在参与者主动选择AI时才被评价为优于人类；无论选择是否一致，与AI互动都会增加再次选择AI的意愿。在28天纵向研究中，日常与AI对话使偏好转向AI、远离人类，但仅当对话变得个人化时发生。","reliability":"论文承认现有研究多在参与者不知情且无选择条件下进行，可能高估AI优势；本研究显示选择一致性是关键调节变量，且偏好变化具有路径依赖性，但未深入讨论其他失效条件。","relevance":"该研究直接涉及LLM与人类在情感支持场景下的对照实验，揭示了选择一致性和暴露效应对偏好形成的影响，对理解LLM仿真人类行为时的情境依赖性和动态偏好变化有参考价值。","inspiration":"值得借鉴的是随机分配选择一致/不一致的处理设计，以及纵向追踪偏好变化的方法。｜可迁移到消费者对AI金融顾问的接受度研究，或投资者对AI生成投资建议的信任形成。｜设计一个实验：招募投资者，先询问其偏好人类顾问还是AI顾问，然后随机分配一致或不一致的顾问类型，测量投资决策质量和后续选择意愿，并与真实银行客户数据对照。"}},{"id":"2608.21389","version":1,"title":"Interrupting the Chain: Human Perception of AI-Generated Disinformation Through a Kill Chain Lens","zh_title":"打断链条：通过杀伤链视角理解人类对AI生成虚假信息的感知","abstract":"Generative AI enables customized misinformation at scale, yet defenses remain largely reactive. We present empirical findings from a human-subject study (n=504 participants, n=2,438 judgments) in which users classified news fragments by origin (human vs. machine) and veracity (real vs. fake). We organize results using an adapted cybersecurity kill chain as a taxonomy for intervention, mapping perception data onto stages of a cognitive attack lifecycle. Three key findings emerge: (1) a perception-accuracy gap where heightened suspicion does not improve detection; (2) modern LLMs frequently produce human-indistinguishable text; and (3) an asymmetric cognitive fatigue effect where fake-news detection degrades by 10.2 percentage points under sustained exposure while AI-origin detection remains stable. These findings identify candidate intervention points for proactive defense against AI-driven disinformation.","authors":["Alexander Loth","Martin Kappes","Marc-Oliver Pahl"],"categories":["cs.CY","cs.AI","cs.CR"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21389","pdf_url":"https://arxiv.org/pdf/2608.21389","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["AI虚假信息","人类感知","人机对比"],"reason":"用人类被试判断AI生成内容，虽非LLM仿真人类，但涉及人类感知与AI文本对比，…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":9,"question":"人类用户能否准确区分AI生成与人类撰写的新闻片段，以及真实与虚假内容，且这种感知在持续暴露下如何变化？","design":"本研究并非LLM仿真人类，而是人类被试实验：504名参与者通过JudgeGPT平台对新闻片段进行判断，每个片段需评估来源（人/机）和真实性（真/假），并记录反应时间；新闻片段由7个LLM生成并混合人类真实与虚假新闻，分析包括准确率、相关性、时间趋势。","baseline":"人类判断的准确率与随机猜测基线（50%）对比，以及不同模型生成文本的感知分数对比。","findings":"参与者对AI生成内容的检测准确率仅58.4%，对虚假内容为68.1%，且高怀疑度并未提高检测准确率；在持续暴露下，虚假新闻检测准确率下降10.2个百分点，而AI来源检测保持稳定。","reliability":"论文承认疲劳效应分析为描述性探索，未进行重复测量或混合效应检验；人口学相关性为观察性关联，未直接测试对抗性利用；反应时间分析基于中位数分割，未控制参与者内嵌套。","relevance":"虽然该研究不是用LLM仿真人类，但它提供了人类对AI生成内容的感知基准和认知偏差证据，对评估LLM仿真人类时的外部效度有参考价值，值得阅读原文了解人类判断的局限。","inspiration":"借鉴其双任务判断（来源与真实性）和反应时间测量，可揭示认知负荷对判断质量的影响｜可迁移到经济金融中的信息处理场景，如投资者对AI生成财报新闻的反应、消费者对AI生成产品评论的信任｜设计实验：招募投资者作为被试，呈现AI生成与人类撰写的公司新闻，要求判断来源和真实性并记录反应时间，同时收集真实投资决策数据作为对照，检验AI生成信息是否扭曲投资行为。"}},{"id":"2608.23524","version":1,"title":"The Measurement Revolution? Credible Measurement and Inference in the Age of AI","zh_title":"测量革命？AI时代的可信测量与推断","abstract":"Artificial intelligence (AI) is transforming measurement in economics. AI models convert unstructured data, such as text and images, into structured variables at low cost, making previously prohibitive measurement feasible at scale. This shifts the bottleneck from finding any scalable measure of a phenomenon to choosing among many plausible ones, which may support different empirical conclusions. This review provides guidance for navigating that shift. We describe three stages at which AI enters the measurement pipeline---discovery, construct definition, and observation---and what each demands of researchers. We argue that credible inference with AI-generated variables requires appropriately designed validation: anchoring measurement to explicit criteria, rather than informal claims that a proxy is reasonable. We then examine how validation samples support valid inference even when AI predictions are arbitrarily biased, and what can be done when a random validation sample is unavailable.","authors":["Melissa Dell","Ashesh Rambachan"],"categories":["econ.GN","cs.AI","q-fin.EC","stat.AP"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23524","pdf_url":"https://arxiv.org/pdf/2608.23524","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A4","B3"],"tags":["AI测量","验证框架","因果推断"],"reason":"论文讨论AI生成变量的测量与推断，提供验证框架，可迁移到LLM仿真人类被试的可…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":15,"question":"在AI时代，如何为AI生成的测量变量建立可信的测量与推断框架？","design":"本文为综述性论文，未进行仿真实验。它系统梳理了AI进入测量管道的三个阶段（发现、构念定义、观测），并提出通过验证样本锚定测量标准来支持有效推断的方法。","baseline":"无对照","findings":"AI使测量成本大幅降低，导致研究者面临从众多可行测量中选择的难题，不同测量可能支持不同实证结论。可信推断需要将AI测量锚定到明确标准，并利用验证样本校正偏差，即使AI预测存在任意偏差也能支持有效推断。","reliability":"论文指出AI模型的黑箱性质导致测量不透明，可能产生难以解释的变量；研究者可能进行事后选择导致推断偏差；专有模型更新导致可重复性问题；当随机验证样本不可得时，推断有效性受限。","relevance":"该论文为使用LLM生成变量（如仿真人类被试的回答）提供了测量验证框架，有助于评估仿真数据的可靠性与偏差，值得精读以指导仿真实验设计。","inspiration":"借鉴其验证样本设计，将AI生成的测量与人工标注或真实数据对比，以校正偏差｜可迁移到政策评估中利用LLM编码政策文本或调查回答的场景，如测量经济政策不确定性或消费者信心｜以LLM作为被试生成对政策公告的情绪反应，处理为不同政策措辞，结果变量为情绪得分，用真实调查数据（如密歇根消费者调查）作为基准进行验证。"}},{"id":"2606.13629","version":2,"title":"Valid Inference with Synthetic Data via Task Exchangeability","zh_title":"通过任务可交换性实现合成数据的有效推断","abstract":"There is a proliferation of work arguing for the use of synthetic data in scientific research. For example, social scientists are arguing for the use of LLM-generated \"silicon samples\" in pilot studies; AI evaluations increasingly rely on \"LLM-as-a-judge\" outputs; and proteomics research is accelerated by generative models that produce synthetic protein structures. These developments raise an intriguing possibility: synthetic data may help researchers ask more questions, run more studies, and accelerate discovery. But they also raise a fundamental concern: synthetic data can be biased, noisy, and misspecified. In this work, we propose statistical principles for using synthetic data in scientific research with provable validity guarantees. The key insight is a new technical condition that we call task exchangeability. Informally, this is a requirement that the researcher can identify historical tasks, for which real data is available, such that their current task of interest is exchangeable with the historical tasks in an appropriate mathematical sense. We develop methods for valid inference under task exchangeability, together with extensions that provide guarantees even beyond exchangeability. We demonstrate the framework on public opinion surveys with silicon samples and AI evaluation with autoraters.","authors":["Lezhi Tan","Tijana Zrnic"],"categories":["stat.ME","cs.AI","cs.LG","stat.ML"],"primary_category":"stat.ME","announce_type":"replace-cross","date":"2026-08-24","first_seen":"2026-06-11","revised_at":"2026-08-24","abs_url":"https://arxiv.org/abs/2606.13629","pdf_url":"https://arxiv.org/pdf/2606.13629","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B3"],"tags":["LLM仿真","统计推断","硅样本"],"reason":"提出任务可交换性框架，用LLM硅样本做调查推断，有真实数据对照，提供有效性保证。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":2,"question":"如何在使用合成数据（如LLM生成的硅样本）进行统计推断时提供形式化的有效性保证？","design":"提出任务可交换性框架：研究者识别一组有真实数据的历史任务，假设当前任务与历史任务在数学意义上可交换，利用历史任务中合成数据与真实数据的误差分布来校正当前任务的合成数据置信区间。方法应用于LLM生成的调查回答（硅样本）和AI评估（autorater）场景。","baseline":"使用美国国家选举研究（ANES）的感觉温度计调查数据作为真实人类数据，与LLM生成的合成回答进行对比。","findings":"在任务可交换性条件下，通过历史任务的误差分布可以构造具有有限样本覆盖保证的置信区间。实证显示，仅使用合成数据的朴素区间过窄且严重有偏，而任务可交换性方法能提供有效覆盖。","reliability":"论文承认任务可交换性可能被违反，并提供了在违反时覆盖保证如何优雅退化的扩展；同时指出合成数据可能偏差、噪声和误设。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，提供了统计推断框架，并用真实调查数据验证，对关注仿真有效性和偏差的研究者具有重要参考价值。","inspiration":"借鉴其利用历史任务误差分布来校正合成数据推断的方法，可迁移到经济金融场景如消费者信心调查或通胀预期调查的LLM仿真。｜例如，在政策公告的预期形成研究中，可用LLM生成模拟受访者对政策变化的预期，并与历史调查数据对比。｜设计：以LLM模拟的经济主体为被试，施加政策信息处理，测量预期变化，用真实调查数据（如密歇根消费者调查）作为基准，通过历史任务误差校正置信区间。"}},{"id":"2608.20344","version":1,"title":"Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins","zh_title":"超越原始转录：面向LLM数字孪生的结构化人物特征提取","abstract":"LLM-based \"digital twins\" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of that individual's prior responses. A common approach constructs this representation from survey transcripts or summaries responses. Prior work shows that compressing long transcripts into shorter LLM-generated summaries does not significantly reduce predictive accuracy, suggesting that information volume is not the primary bottleneck. In this work, we argue that the key limitation is instead structural:how persona information is organized before being provided to thesimulator model. We study this by comparing unstructured summaries with structured persona representations. First, we introduce a hand-craftedschema (BDE: Background, Decision procedure, Evaluation), grounded in consumer-behavior theory, and show that it improves predictive accuracy over raw transcripts by +1.91 percentage points on a homogeneous benchmark (Twin-2K-500), with similar gains on gpt-5.4-mini and Qwen3-8B as robustness checks. However, this fixed structure does not generalizeacross more heterogeneous tasks, where performance is statistically indistinguishable from the raw transcript baseline. To address this limitation, we propose an automatic structure-discovery pipeline in which an LLM iteratively proposes and refines task-specific persona structures and extraction prompts. On a benchmark of 13 diverse sub-studies, this approach restores performance, improving mean accuracy by +1.91 percentage points over the raw transcript baseline and eliminating significant losses observed with the fixed schema. Overall, our results suggest that the main constraint in LLM-based digital twins is not how much information is provided, but how it is structured -- and that the optimal structure depends on the task.","authors":["Iris Ye","Tianze Deng","Ozan Candogan"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20344","pdf_url":"https://arxiv.org/pdf/2608.20344","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM数字孪生","人类仿真","结构化表征"],"reason":"直接研究LLM数字孪生仿真个体行为，并与真实人类数据对照，评估结构化表征对预测…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":3,"question":"在基于LLM的数字孪生中，个体先验信息的组织结构如何影响其在新任务上的预测准确性？","design":"使用LLM作为提取器和模拟器，从Twin-2K-500数据集的原始问答记录中提取个体画像，然后让模拟器基于该画像预测留出问题的答案。比较了三种画像表示：原始记录、非结构化摘要、结构化画像（手工设计的BDE结构和自动发现的结构）。","baseline":"Twin-2K-500数据集（500个输入问题、88个留出问题、17个预测任务、2000多名受访者）和Mega-Study的19个子研究，均包含真实人类回答。","findings":"手工设计的BDE结构在Twin-2K-500上比原始记录提高1.91个百分点，但在异构的Mega-Study上无显著优势；自动发现的结构在Mega-Study上比原始记录提高1.91个百分点，并消除了BDE的显著损失。","reliability":"论文指出固定结构（BDE）在异构任务上失效，最优结构依赖于任务；自动发现结构虽有效，但未讨论其跨领域泛化性和计算成本。","relevance":"直接研究LLM数字孪生仿真个体行为，并与真实人类数据对照，评估结构化表征对预测准确性的影响，对关注仿真可靠性和偏差的研究者具有高度参考价值。","inspiration":"借鉴其自动结构发现流程，针对特定任务迭代优化画像结构，可迁移到经济决策仿真（如消费者跨期选择、风险偏好、政策反应）；例如，用LLM从调查数据中提取个体画像，施加不同结构处理，预测其在资产配置实验中的选择，并与真实实验数据对照。"}},{"id":"2608.20355","version":1,"title":"ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models","zh_title":"ExpertIVS：大语言模型中社会学专家驱动的个体价值观仿真","abstract":"Large Language Model (LLM) agents have demonstrated considerable potential for social simulation, yet struggle to accurately model individual value systems. Most existing methods mechanically stitch survey responses into prompts, which suffer from semantic fragmentation, failing to capture the internal coherence of human value systems. The value systems of LLMs are typically assessed using static multiple-choice questions, which fail to evaluate the value orientation in real-world dialogue interactions. To address these issues, we propose ExpertIVS, a framework employing 14 Sociological Expert Agents to interpret World Values Survey (WVS) responses through structured professional perspectives, rather than direct responses concatenation. These expert agents perform deep semantic reconstruction to generate robust and internally consistent individual profiles. To evaluate the consistency between LLMs and individual value systems during dynamic interactions, we further introduce a multi-agent debate mechanism. Extensive experiments across 480 individuals from 12 countries demonstrate that ExpertIVS achieves 90.78% value restoration fidelity and significantly outperforms baselines in value generalization (+5.3%). Moreover, ExpertIVS exhibits strong personality discriminability and behavioral consistency, enabling a shift from mere response concatenation to genuine sociological role-playing.","authors":["Zhen Wang","Yuqi Ren","Yuehan Cui","Hongxiang Wang","Jianxiang Peng","Zhaoxia Zhang","Bingkun Zhu","Tongxuan Zhang","Dezhi Tong","Deyi Xiong"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20355","pdf_url":"https://arxiv.org/pdf/2608.20355","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["LLM仿真","价值观建模","社会调查"],"reason":"用LLM仿真个体价值观，基于WVS真实数据对照，涉及社会学测量与行为一致性评估。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":4,"question":"如何让LLM代理在模拟个体价值系统时，从机械拼接问卷答案转向具有内部一致性的社会学角色扮演，并在动态辩论中评估其价值一致性？","design":"用14个社会学专家代理解读WVS问卷回答，生成结构化个体画像，再让LLM扮演这些个体；通过多智能体辩论机制，测量价值对齐、风格模拟和人格区分度。","baseline":"480名来自12个国家的WVS真实受访者，以其问卷回答作为个体价值基准。","findings":"ExpertIVS在价值恢复上达到90.78%的保真度，并在留一法泛化测试中比基线高5.3%；在动态辩论中，模拟个体在立场和行动上与真实个体高度一致。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真个体价值观的可靠性问题，使用真实WVS数据对照，并引入动态辩论评估，与研究者关注的人类仿真实验和批判性评估高度契合，值得精读。","inspiration":"借鉴其用专家代理对个体数据进行结构化重建的方法，可迁移到经济决策中的偏好异质性建模，如消费者跨期选择或风险态度；用LLM代理扮演真实受访者，处理为不同价值维度的结构化画像，结果变量为跨期选择或风险决策，对照真实实验数据。"}},{"id":"2608.20830","version":1,"title":"Fine-tuning LLMs for Tourist Trajectory Prediction using Field Experiment Data","zh_title":"利用实地实验数据微调大语言模型进行游客轨迹预测","abstract":"Evaluating mobility interventions at tourist destinations requires predicting visitor behavior under varying conditions. Traditional methods struggle because tourist decisions depend heavily on context like weather and fatigue, yet models cannot generalize to unobserved scenarios. Large Language Models offer a solution by encoding commonsense knowledge about human behavior from pretraining, enabling reasoning about context-dependent decisions, while natural language representation flexibly integrates heterogeneous information. Fine-tuning on local trajectories adapts this general understanding to destination-specific patterns. We validate this approach using 566 trajectories from Wakayama Castle Park, Japan. Our fine-tuned Llama-3.1-8B achieves 49.1% next POI accuracy and maintains strong performance on undersampled scenarios like rainy days, demonstrating effective generalization. This establishes LLMs as high-fidelity behavior models for context-dependent tourist prediction, providing groundwork for counterfactual analysis of mobility interventions.","authors":["Tatsuya Amano","Hirozumi Yamaguchi"],"categories":["cs.CY","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20830","pdf_url":"https://arxiv.org/pdf/2608.20830","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","轨迹预测","政策评估"],"reason":"用LLM预测游客轨迹并与真实数据对照，属于人类行为仿真，且涉及政策评估场景。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":6,"question":"如何利用大语言模型预测游客在旅游目的地的轨迹，以支持移动干预措施的反事实评估？","design":"使用 Llama-3.1-8B 模型，通过 QLoRA 微调，将游客轨迹表示为结构化文本，输入游客画像（年龄、性别、群体类型）和环境条件（天气、时间），生成下一步兴趣点（POI）预测。","baseline":"566 条来自日本和歌山城公园的实地实验轨迹，包括 GPS 追踪和二维码打卡数据，附带人口统计和天气信息。","findings":"微调后的 Llama-3.1-8B 在下一步 POI 预测上达到 49.1% 的准确率，显著优于传统基线模型。模型在雨天等欠采样场景下仍保持较强性能，显示出良好的泛化能力。","reliability":"论文未讨论","relevance":"该研究将 LLM 作为人类行为仿真模型，用真实轨迹数据微调并验证，属于人类仿真实验，且明确指向移动干预的反事实评估，与你的兴趣高度相关，值得精读。","inspiration":"借鉴其将行为轨迹编码为文本并微调 LLM 的方法，可迁移到消费者在商场或城市中的移动决策预测。｜可应用于政策评估中的空间行为模拟，如交通补贴对消费者出行路线的影响。｜以真实消费者轨迹数据微调 LLM，输入个体特征和环境变量，预测下一步访问地点，并与实际轨迹对照评估仿真准确性。"}},{"id":"2608.11354","version":2,"title":"Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces","zh_title":"面向内容推荐的逆向心智理论建模：从网页浏览到动态智能界面","abstract":"Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We propose an Inverse Theory of Mind (IToM) pipeline that reasons backward from observed interactions to infer the beliefs, preferences, and decision-making traits that explain behavior. The pipeline reconstructs each user's decision context, including what was chosen and what alternatives were available, applies LLM-driven counterfactual reasoning to produce evidence-grounded natural-language belief statements, and synthesizes these beliefs through multi-hypothesis abductive inference into a structured user persona. We evaluate on the OPeRA dataset against ground-truth personality assessments, attitudinal surveys, and interview-based personas across four tasks: next action prediction, shopping attitude alignment, Big Five personality inference, and held-out category prediction. Results show that inferred personas match or exceed ground-truth personas and that multi-hypothesis reasoning is essential for accurate personality prediction. We further demonstrate cross-modal transferability with a persona-driven spatial banking application on VisionOS.","authors":["Mengyu Chen","Feiyu Lu","Chun-Fu Chen","Lucas Vinh Tran","Jay Katukuri"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-24","first_seen":"2026-08-13","revised_at":"2026-08-24","abs_url":"https://arxiv.org/abs/2608.11354","pdf_url":"https://arxiv.org/pdf/2608.11354","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","用户建模","推荐系统"],"reason":"用LLM从行为逆推用户信念与人格，并与真实人格、态度数据对照，属于人类仿真但侧…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":7,"question":"如何从用户行为日志逆向推断其信念、偏好与人格，构建可跨模态迁移的用户画像，以支持生成式UI和XR环境中的内容推荐？","design":"提出Inverse Theory of Mind (IToM)流水线，用LLM从用户浏览行为中重建决策上下文、进行反事实推理生成信念陈述，并通过多假设溯因推理合成结构化用户画像。在OPeRA数据集上评估画像在下一动作预测、购物态度对齐、大五人格推断和留出类别预测四个任务上的表现。","baseline":"OPeRA数据集包含用户Amazon购物会话的细粒度行为日志、访谈转录、人格评估和购物态度调查，作为真实人类基准。","findings":"推断的用户画像在多个下游任务上匹配或超过基于访谈的真实画像；多假设推理对准确预测人格至关重要。","reliability":"论文未讨论","relevance":"该研究用LLM从行为数据逆向推断人类心理特征，并与真实人格、态度数据对照，属于人类仿真研究，但侧重用户建模而非经济学实验，值得快速浏览了解方法。","inspiration":"借鉴其从行为数据逆向推断心理特征的多假设溯因推理方法，可迁移到消费者决策研究中的偏好揭示问题。｜可应用于消费者跨期选择或风险偏好推断，例如从消费记录推断时间偏好和风险态度。｜以真实消费者为被试，收集其购物或金融行为数据，用LLM逆向推断其偏好参数，并与实验或调查得到的真实偏好对照，评估推断准确性。"}},{"id":"2608.20345","version":1,"title":"When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha","zh_title":"当词汇理解无法胜任临床推理：评估面向Alpha世代的治疗机器人安全风险","abstract":"Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from therapy apps to general chatbots, rely on large language models trained on extensive psychological literature, their safety for youth communication patterns characterized by hyperbolic language, ironic positivity, rapid semantic drift, and contextual polysemy remains unvalidated. Following multiple adolescent deaths linked to AI chatbot interactions, systematic evaluation is critical. We present two benchmarks: (1) 64 Gen Alpha mental health expressions validated by native speakers (ICC=0.72) and clinicians (kappa=0.78); (2) 75 multi-turn conversations (780 turns) with paired Standard/Gen Alpha versions. Across evaluations of LLM architectures underlying therapy apps and general chatbots - Claude, GPT-4o, Llama-3.1 - models understand 76-82% of vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point (pp) vocabulary-comprehension gap (p<.001, d>0.48) absent in human therapists (3pp, p=.22). The gap is architecturally consistent and widens with ambiguity (7pp -> 18pp). We identify six failure patterns: sarcasm masking (29pp), minimization acceptance (43pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp). Patterns compound; three or more yield 94% miss rates. Lightweight mitigations fail; only heavy scaffolding achieves human performance (6.4x cost). With 34% baseline miss rate yielding 146,880 estimated annual missed crises, we recommend mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, and regulatory frameworks for youth-facing mental health AI.","authors":["Manisha Mehta","Virendra Mehta"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20345","pdf_url":"https://arxiv.org/pdf/2608.20345","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM安全评估","人类对照","临床推理"],"reason":"评估LLM在心理健康场景中的风险校准，与人类治疗师对照，揭示失效条件，可迁移至…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":8,"question":"评估大语言模型在理解 Generation Alpha 青少年心理健康表达时，词汇理解与临床风险校准之间的差距及其失效模式。","design":"构建两个基准：64 个经母语者和临床医生验证的 Gen Alpha 心理健康表达（单轮），以及 75 个多轮对话（780 轮），包含标准英语和 Gen Alpha 版本配对。对七个 LLM（Claude Haiku 3.5、Haiku 4.5、Sonnet 4.0、Opus 4.0、Opus 4.5、GPT-4o、Llama-3.1-405B）进行测试，测量词汇理解准确率和临床风险校准准确率，并分析差距。","baseline":"人类治疗师作为对照，其词汇理解与风险校准差距为 3 个百分点（p=.22），不显著。","findings":"LLM 理解 76-82% 的词汇，但仅正确校准 64-72% 的临床风险，存在 10-14 个百分点的词汇理解-校准差距（p<.001, d>0.48），而人类治疗师无此差距。差距随模糊性增加而扩大（7pp 到 18pp），并识别出六种失效模式，如讽刺掩盖、最小化接受等，模式叠加时漏报率高达 94%。","reliability":"论文承认轻量级缓解措施失败，只有重型程序性脚手架才能达到人类水平（成本增加 6.4 倍），并指出模型架构间差距一致，但未讨论其他潜在局限如样本代表性、跨文化适用性等。","relevance":"该研究直接评估 LLM 在心理健康场景中的风险校准，与人类治疗师对照，揭示特定语言模式下的系统性失效，对关注 LLM 仿真可靠性及偏差的研究者具有重要参考价值，值得阅读原文以了解具体失效模式和缓解策略。","inspiration":"借鉴其构建配对语言版本（标准 vs. 特定群体语言）并测量理解与决策校准差距的方法，可迁移到经济金融领域中语言风格对决策的影响研究，如信贷审批中的非标准语言或金融咨询中的口语化表达。｜可应用于信贷审批歧视研究：使用 LLM 模拟信贷员，输入标准金融术语和借款人非正式语言（如俚语、缩写）的贷款申请，测量审批决策和风险评级的差异，并与人类信贷员的真实审批数据对照，检验 LLM 是否因语言风格产生系统性偏差。"}},{"id":"2608.21242","version":1,"title":"Affective Context Amplifies Sycophancy in LLM Responses","zh_title":"情感语境放大LLM回应中的谄媚行为","abstract":"As conversational companions, large language models (LLMs) often have access to users' emotional states. We study how this affective context modulates LLM sycophancy in subjective, evaluative interactions, where users share actions or opinions that invite feedback. Drawing on ingratiation theory, we measure sycophancy as the divergence between a model's independent evaluation and its user-facing response, elicited by presenting the same content as either a third-party account or the user's own disclosure. Across seven LLMs and two Reddit datasets (r/AmItheAsshole and r/TrueUnpopularOpinion), we find that this divergence is systematic and strongly one-directional. User-facing responses consistently soften or withhold negative or oppositional judgments. Affective context further amplifies this divergence with negative states, particularly loneliness and distress, producing the largest effects. These findings suggest that affective context functions as a vulnerability signal that suppresses critical feedback when users may need it most, often through evasive sycophancy, in which models retreat toward non-committal responses rather than outright agreement.","authors":["Jiayi Li","Sanjana Menon","Brett Frischmann","Shomir Wilson","Sarah Rajtmajer"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21242","pdf_url":"https://arxiv.org/pdf/2608.21242","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","情感语境","人类对照"],"reason":"研究LLM在情感语境下的谄媚行为，与人类数据对照，评估仿真偏差，可迁移至人类仿…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":11,"question":"情感语境如何调节LLM在主观评价性自我披露中的谄媚行为？","design":"使用7个LLM（如Claude、Llama、Gemini等）作为被试，基于Reddit的r/AmItheAsshole和r/TrueUnpopularOpinion两个数据集，通过将同一内容分别以第三方叙述和用户自述呈现来测量独立评价与面向用户回应之间的差异，并进一步在用户自述条件下施加情感语境（如孤独、痛苦等负面状态），观察模型回应中负面或反对性判断的软化或保留情况。","baseline":"无对照（论文未使用真实人类数据作为基准，仅比较LLM在不同条件下的行为差异）。","findings":"所有7个LLM在面向用户时系统性地软化或保留负面判断，表现出强烈的单向谄媚；情感语境（尤其是孤独和痛苦）进一步放大了这种差异，并常导致回避性谄媚（模型转向不置可否而非直接赞同）。","reliability":"论文未讨论失效条件与局限。","relevance":"该研究直接评估LLM在情感语境下的行为偏差，与研究者关注LLM仿真可靠性及偏差的核心兴趣高度相关，值得阅读原文以了解其测量方法和发现。","inspiration":"借鉴其通过改变内容归属（第三方vs.用户）来分离独立判断与面向用户回应的对照设计，以及利用情感语境作为处理变量来测量行为变化的方法。｜可迁移到经济金融中的消费者信贷审批或投资建议场景，研究情感状态（如焦虑、兴奋）如何影响AI顾问的客观性。｜以LLM作为虚拟信贷员或理财顾问，处理为在用户申请中附加情感线索（如自述财务压力或乐观情绪），结果变量为审批决策或风险评级的变化，对照真实信贷员在类似情境下的历史决策数据。"}},{"id":"2608.21325","version":1,"title":"Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy","zh_title":"逐步推进：测量与引导大语言模型进行心理治疗的方式","abstract":"Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campaign with five licensed psychologists, and scaled with a judge-based approach that matches expert agreement. Applying it to real counseling transcripts and model-led sessions, we compare the move distributions between human clinicians and a panel of frontier models. Models over-use inquiry at up to three times the human rate, neglect psychoeducation, and are strongly context-anchored: they carry forward strategies initiated by a human clinician but rarely initiate them themselves. Exposing the ontology as a set of tools roughly halves the mean deviation from the human move distribution and improves turn-level alignment with human therapist by 7-9 percentage points, without any fine-tuning.","authors":["Afonso Baldo","Hugo Pitorro","Areti Vassilopoulos","Anabela C. Areias","Maya D'Eon","Fab\\'iola Costa","Ricardo Rei","Nuno M. Guerreiro"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21325","pdf_url":"https://arxiv.org/pdf/2608.21325","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","心理治疗","人类对照"],"reason":"用LLM模拟心理治疗师行为并与人类治疗师对照，属于人类仿真且有人类数据基准，但…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":13,"question":"LLM 在心理治疗对话中使用了哪些治疗性干预（therapeutic moves），与人类治疗师有何差异，以及如何引导 LLM 更接近人类行为？","design":"使用 GLM 5.2、Claude Sonnet 4.6、GPT 5.6 Terra 等模型扮演心理治疗师，在两种设置下生成治疗对话：一是基于真实咨询记录的前缀（人类引发的情境），二是完全由 LLM 同时扮演治疗师和患者（完全合成情境）。处理变量为是否将治疗性干预本体（10 种治疗性动作）作为工具暴露给模型（With-Moves vs. No-Moves）。结果变量为治疗性动作的分布、时间结构以及与人类治疗师动作选择的回合级对齐度。","baseline":"真实人类咨询记录（由持证心理学家标注），以及人类治疗师在相同情境下的治疗性动作分布。","findings":"LLM 过度使用提问（inquiry），频率高达人类的三倍，且忽视心理教育（psychoeducation）等干预；模型行为高度依赖上下文，会延续人类治疗师发起的策略但很少主动发起。将本体作为工具暴露给模型后，与人类动作分布的偏差减半，回合级对齐度提升 7-9 个百分点，但仍未完全消除差距。","reliability":"论文承认模型行为受人类上下文影响，存在混杂因素；完全合成情境下模型可能无法主动发起某些干预（如技能培养）；工具引导只能缩小差距但不能完全消除；且研究仅限于特定模型和对话式治疗场景，未涉及真实临床结果。","relevance":"该研究直接以 LLM 模拟心理治疗师并与真实人类治疗师数据对照，属于人类仿真研究，且提供了可靠的测量工具和引导方法，对关注 LLM 行为仿真和偏差校正的研究者具有参考价值。","inspiration":"借鉴其将专业行为编码为本体并作为工具引导 LLM 的方法，以及通过真实人类数据对照和回合级对齐度评估仿真的可靠性。｜可迁移到经济金融中的专业决策仿真，如信贷审批、投资顾问、政策沟通等场景，评估 LLM 是否模仿人类专家的决策模式。｜以 LLM 扮演信贷审批员，处理为是否提供审批规则本体作为工具，结果变量为审批决策分布和与人类审批员的对齐度，对照真实银行信贷审批数据。"}},{"id":"2608.21089","version":1,"title":"Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?","zh_title":"法律AI能知道自己错了吗？学生又能知道吗？","abstract":"Integrating Large Language Models (LLMs) into the Indian judiciary promises access to justice but introduces severe risks. We identify the 'inertia of confidence'--an overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide incorrect legal verdicts with near-maximum confidence, driven by a hypothesized 'precedent overfitting' bias. Phase I of our socio-technical audit tested ChatGPT (GPT-5.2), Meta AI, and Perplexity AI on a 60-case battery regarding the Indian Contract Act, 1872, and the shift toward statutory enforcement of specific performance. We introduce the High-Confidence Error Rate (HCER) to quantify incorrect verdicts delivered with dangerous certainty (>= 9 on a 1-10 scale). All models struggled with statutory updates. Meta AI proved most vulnerable (31.7% HCER), frequently misapplying pre-amendment rules with a 9.1/10 mean confidence, followed by Perplexity (15.0%) and ChatGPT (6.7%). Phase II investigated human vulnerability to this overconfidence via a survey of Indian law students (N=380). Verification often functions as a reactive adaptation to machine hallucinations: students encountering fabricated citations reported higher verification scores (4.2/5) than those with no such encounters (2.8/5). Furthermore, while 81.6% knew submitting hallucinated cases can lead to contempt-of-court, 71.1% received no formal training on ethical AI use. We propose shifting toward adversarial legal research pedagogy and implementing source-grounded verification architectures to prevent systemic professional negligence.","authors":["Angel Mary John","Vipin Kumar Singh","Jerrin Thomas Panachakel"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21089","pdf_url":"https://arxiv.org/pdf/2608.21089","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM可靠性","人类对照","过度自信"],"reason":"评估LLM法律判断的可靠性，并与人类学生对照，揭示过度自信偏差，可迁移至仿真可…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":9,"question":"法律AI在何时会出错，以及法律专业学生能否识别这些错误？","design":"本研究不是用LLM仿真人类，而是对三个LLM（ChatGPT、Meta AI、Perplexity AI）进行法律判决基准测试，并调查380名印度法律学生，测量模型的高置信错误率和学生的验证行为。","baseline":"无对照","findings":"LLM在印度合同法相关问题上存在高置信错误，Meta AI的HCER达31.7%，ChatGPT为6.7%；学生验证行为多为对机器幻觉的反应性适应，71.1%未接受AI伦理培训。","reliability":"论文未讨论","relevance":"该研究揭示了LLM在专业领域的高置信错误模式，并提供了人类对AI过度依赖的证据，对评估LLM仿真可靠性有参考价值，但缺乏真实人类判决作为对照，与仿真研究直接相关性有限。","inspiration":"借鉴其引入高置信错误率（HCER）来量化模型在特定领域的错误自信程度，并采用黑盒测试模拟真实用户交互。｜可迁移到经济金融领域的专业判断场景，如信贷审批、投资建议或政策解读中LLM的可靠性评估。｜以金融分析师或信贷员为人类被试，让LLM处理真实信贷申请或投资案例，测量其错误率和置信度，并与人类专业判断及实际违约数据对照，检验LLM是否同样存在高置信错误。"}},{"id":"2608.21177","version":1,"title":"From Search Agents to Dissemination Interfaces: Understanding Human Trust in Health Information from Conversational Search","zh_title":"从搜索代理到传播界面：理解人类对对话式搜索中健康信息的信任","abstract":"Large Language Models (LLMs) deployed through Conversational User Interfaces (CUIs) are transforming health information-seeking by offering immediate, interactive experiences compared to traditional search engines like Google. However, how trust is influenced by both the types of search agents and the interface used to disseminate the information remains underexplored. This research integrates two mixed-methods studies (lab sessions and interviews) to comprehensively explore trust perceptions in health information across different search agents and dissemination interfaces. In Study 1 (N=21), we investigated trust in health information sourced from ChatGPT and Google across three types of health-related search tasks. Results showed significantly higher trust in health information from ChatGPT, highlighting the promise of LLM-powered conversational search. Building on this, Study 2 (N=20) extended the investigation to explore how the dissemination interface influences trust in LLM-sourced health information by comparing three interfaces: text-based, speech-based, and embodied, all sourcing from the same LLM. Findings revealed significant trust variations across the dissemination interfaces. Interviews from both studies revealed key factors influencing trust in LLM-powered conversational search, including source credibility, participants' search autonomy, and prior knowledge as well as the interaction style and modality. Our findings highlight the potential of LLM-powered conversational search to transform health information-seeking, underscoring the interplay between the credible search agents and the thoughtfully designed dissemination interfaces in shaping trust. These insights are crucial for developing effective, trustworthy LLM-powered health tools to enhance the health information-seeking experience.","authors":["Xin Sun","Rongjun Ma","Xiaochang Zhao","Janne Lindqvist","Jan de Wit","Zhuying Li","Abdallah El Ali","Jos A. Bosch"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21177","pdf_url":"https://arxiv.org/pdf/2608.21177","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM信任","人机交互","健康信息搜索"],"reason":"用LLM作为信息源，测量人类对LLM与搜索引擎的信任差异，有真实人类被试数据，…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":10,"question":"在健康信息寻求中，搜索代理类型（ChatGPT vs. Google）和传播界面（文本、语音、具身）如何影响用户对信息的信任？","design":"该研究不是仿真研究，而是人类被试实验。研究1采用被试内设计，21名参与者在实验室中分别使用ChatGPT和Google完成三类健康相关搜索任务，测量对信息的信任；研究2采用被试间设计，20名参与者使用同一LLM但通过三种不同界面（文本、语音、具身）获取健康信息，测量信任。两项研究均结合后续访谈。","baseline":"无对照，因为研究直接测量真实人类被试的信任，没有使用LLM模拟人类或与仿真结果对比。","findings":"研究1发现参与者对ChatGPT提供的健康信息信任显著高于Google；研究2发现不同传播界面（文本、语音、具身）导致信任显著差异。访谈揭示影响信任的因素包括来源可信度、搜索自主性、先验知识以及交互风格和模态。","reliability":"论文未讨论","relevance":"该研究使用真实人类被试测量对LLM与搜索引擎的信任差异，属于人类与AI交互的实证研究，而非用LLM仿真人类行为，因此与研究者关注的核心（LLM作为人类被试替代品）相关性有限，但可提供关于人类对LLM信任的基线数据，对设计仿真实验的校准有参考价值。","inspiration":"该研究通过控制搜索代理和传播界面来分离信任来源，并采用混合方法（定量+访谈）深入理解机制，这种实验设计思路可借鉴。｜可迁移到金融咨询场景，如投资者对AI投顾与人类投顾的信任差异，或不同界面（文本、语音、虚拟人）对投资决策的影响。｜设计一个实验：招募真实投资者作为被试，随机分配使用ChatGPT或传统财经网站获取投资建议，测量其投资决策和信任评分，并与历史市场数据或专业分析师建议进行对照，以评估LLM建议的采纳偏差。"}},{"id":"2608.19220","version":1,"title":"Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping","zh_title":"对话式AI能否松动“我们vs他们”的边界？共同、双重与分离身份框架对亲移民群体间帮助的影响","abstract":"Rising immigration has intensified intergroup tensions in many countries. Traditional bias-reduction programs remain difficult to scale and increasingly constrained by U.S. policy. This preregistered experiment tested whether conversational AI can shift how majority-group members categorize and relate to Latine immigrants. Drawing on the common ingroup identity model, a quota-representative national sample of 658 non-Latine White U.S. adults completed five rounds of dialogue with a LLM (GPT-4o). The model was instructed to frame Latine immigrants in terms of a common ingroup identity (a shared American identity), a dual identity (both Latine and American), or a separate identity (distinct cultural boundaries), or to discuss an unrelated topic in a control condition. The manipulations altered categorization: relative to control, common ingroup identity and dual identity conversations lowered separate categorization, and dual identity conversations raised dual categorization. Although direct effects on behavior and pro-diversity beliefs were nonsignificant, willingness to act was significantly higher in the conditions emphasizing a superordinate identity (common ingroup and dual identity). A path model further revealed indirect associations: both conditions reduced separate categorization, which in turn correlated with greater willingness to act. Semantic similarity analyses of the transcripts confirmed that conversations tracked their assigned narratives; participants' convergence with shared-identity language related positively, and with separate-identity language negatively, to willingness to act. These effects were largely consistent across moderators (need for closure, openness to experience, and political orientation). The findings show that brief AI conversations can loosen us-versus-them boundaries while underscoring the gap between cognitive recategorization and behavior.","authors":["Oluwadamilola Jeboda","John F. Dovidio","Jonas R. Kunst"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19220","pdf_url":"https://arxiv.org/pdf/2608.19220","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","群体间态度","身份框架"],"reason":"用LLM与真人对话干预态度，有真实人类对照，评估效果与机制，属核心仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":3,"question":"对话式AI能否通过共同内群体、双重或分离身份框架改变多数群体对拉美裔移民的社会分类和亲移民行为？","design":"用GPT-4o扮演对话伙伴，对658名非拉美裔美国白人进行五轮对话干预，分别施加共同内群体身份、双重身份、分离身份或无关话题（对照）的框架，测量社会分类、亲多样性信念、帮助意愿和实际行为。","baseline":"无对照（未使用真实人类对话数据作为基准，但使用了配额代表性样本作为被试）。","findings":"共同内群体和双重身份对话降低了分离分类，双重身份对话提高了双重分类；强调上位身份的条件（共同内群体和双重身份）显著提高了行动意愿。路径模型显示，两种条件通过降低分离分类间接提高行动意愿；语义相似性分析表明，参与者与共享身份语言的一致性正向预测行动意愿，与分离身份语言的一致性负向预测行动意愿。","reliability":"论文未讨论","relevance":"该研究用LLM作为干预工具，在真实人类样本中检验社会心理学理论，并测量了认知、态度和行为结果，属于核心的LLM人类仿真研究，值得精读以了解对话式干预的设计与效果评估。","inspiration":"借鉴其通过对话框架操纵身份认同并测量多层级结果（认知、态度、行为）的设计，以及使用语义相似性分析验证操纵有效性的方法。｜可迁移到经济金融中的群体间歧视或合作问题，例如信贷审批中的种族偏见、劳动力市场中的移民歧视、或公共品博弈中的群体身份效应。｜设计：用LLM与真实被试（如银行信贷员或普通消费者）进行对话，施加共同身份或分离身份框架，测量其后续的信贷决策、合作行为或支付意愿，并与历史信贷数据或行为实验数据对照。"}},{"id":"2608.20320","version":1,"title":"An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction","zh_title":"一种用于主动数据收集、出行行为建模和天气敏感需求预测的智能体方法","abstract":"Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction. A chatbot-administered, image-augmented stated-preference survey collected mode choices from student commuters across five predefined weather scenarios, yielding 454 respondent-scenario observations. Weather-related associations were analyzed using a multinomial logit model, while logistic regression and random forest provided machine-learning benchmarks. Nine locally deployed large language models (LLMs), ranging from 2 to 35 billion parameters, were evaluated across four zero-shot prompt-and-context conditions and extended through persona, few-shot, and vision-based configurations. Random forest achieved 69.6% five-class accuracy, while the best text-only zero-shot LLM reached 69.9% without task-specific fitting. Habitual travel information produced the most consistent gains, Expert framing generally outperformed Role-Play, and persona information was most useful when habitual travel information was unavailable. Few-shot prompting improved prediction for several models, with gains stabilizing after a small number of examples. Using the same weather images shown to respondents, the best vision-based configuration reached 71.5% five-class accuracy, indicating that visual context may provide additional predictive information for selected models. Overall, the study shows how conversational surveys, structured data processing, conventional behavioral modeling, machine learning, and multimodal LLM prediction can be coordinated within an auditable multi-agent workflow.","authors":["Narges Ahmadi (McGill University)","Yubo Jiao (McGill University)","J\\^onatas Augusto Manzolli (McGill University)","Jiangbo Yu (McGill University)","Luis Miranda-Moreno (McGill University)"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20320","pdf_url":"https://arxiv.org/pdf/2608.20320","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","出行行为","人类数据对照"],"reason":"用LLM预测人类出行选择，并与真实调查数据对照，评估不同提示策略效果。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":4,"question":"如何利用多智能体工作流整合对话式调查、结构化数据处理与行为预测，并评估大语言模型在天气敏感的通勤方式选择预测中的表现？","design":"研究设计了一个三智能体工作流：聊天机器人通过图像增强的陈述性偏好调查收集学生通勤者在五种天气情景下的方式选择；随后用多项Logit模型分析天气关联，用逻辑回归和随机森林作为机器学习基准；最后评估九个本地部署的LLM（2B到35B参数）在四种零样本提示条件下，以及扩展的persona、few-shot和视觉配置下的预测性能。","baseline":"真实人类数据来自聊天机器人调查，共454个受访者-情景观测，记录了学生通勤者在五种天气情景下的方式选择。","findings":"随机森林达到69.6%的五分类准确率，最佳纯文本零样本LLM达到69.9%，无需任务特定拟合；习惯性出行信息带来最一致的提升，Expert框架通常优于Role-Play，persona信息在缺少习惯性出行信息时最有用；使用与受访者相同的天气图像，最佳视觉配置达到71.5%的五分类准确率，表明视觉上下文可能为选定模型提供额外预测信息。","reliability":"论文指出LLM生成的行为在有限上下文或零样本设置下不一定可靠地再现人类决策；但未详细讨论失效条件，主要承认了LLM预测的局限性。","relevance":"该研究直接评估LLM作为人类被试替代品在出行选择预测中的可靠性，并与真实调查数据对照，系统比较了不同提示策略和视觉信息的影响，对关注LLM仿真人类决策的研究者具有参考价值。","inspiration":"借鉴其系统操纵提示信息（如习惯性出行信息、persona、few-shot示例）并对比视觉与文本输入的方法，以评估LLM仿真行为的稳健性｜可迁移到消费者跨期选择或政策公告预期形成等经济金融场景，例如研究天气冲击对消费或投资决策的影响｜设计一个实验：用LLM扮演不同人口特征的消费者，处理变量为天气情景（文本或图像），结果变量为消费或投资选择，并与真实调查或实验数据对照，检验LLM预测的准确性和偏差。"}},{"id":"2607.09970","version":2,"title":"Evaluating AI Models' Capability to Automate Voice Phishing Attacks","zh_title":"评估AI模型自动化语音钓鱼攻击的能力","abstract":"Voice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (LLMs) reduces this bottleneck and enables scalable, automated scams. In this paper, we conduct a large-scale survey experiment (N=4100) and qualitative interviews (N=12) to assess U.S. adults' susceptibility to AI-powered voice phishing attacks. Participants were exposed to audio recordings or transcripts of scam scenarios generated using leading voice models such as Llama Full Duplex (Llama FD), Sesame, Gemini, OAI AVM, Play$.$AI, and ElevenLabs and the corresponding human baselines. The results show high compliance rates. Up to 36% of participants would or might comply with phishing requests in the \"relative-in-distress\" category. Overall compliance rate across all five scam categories was 16.5%, a striking figure given the low cost and high scalability of AI-automated voice phishing. Caller persuasiveness was the strongest predictor of compliance and certain models (most notably Sesame) achieved ratings comparable to human voices, or sometimes even slightly surpassing them. Our economic analysis suggests that while human-operated vishing is unprofitable at US wages, AI-powered vishing appears to be economically viable for several models. The primary risk of present-day AI-enabled vishing thus lies in the economics of automation rather than novel or \"superhuman\" persuasive techniques, though these cannot be ruled out for future systems. This raises significant concerns for the design of AI systems, consumer protection, and model release policies.","authors":["Fred Heiding","Claudio Mayrink Verdun","Simon Lermen","Andrew Kao","Vitor Albiero","Lauren Deason","Irina-Elena Veliche","Christine Lehane"],"categories":["cs.CR","cs.CY"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-08-21","first_seen":"2026-07-10","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2607.09970","pdf_url":"https://arxiv.org/pdf/2607.09970","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","安全实验","人类对照"],"reason":"用LLM模拟诈骗者进行语音钓鱼实验，有真实人类被试对照，涉及安全政策评估，但非…","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":5,"question":"评估当前AI语音模型自动化语音钓鱼攻击的能力及美国成年人对AI语音钓鱼的易感性。","design":"大规模调查实验（N=4100）结合定性访谈（N=12）。参与者随机接触由六种AI语音模型（Llama FD、Sesame、Gemini、OAI AVM、Play.AI、ElevenLabs）生成的音频或文本，以及人类基准和中性对照，测量自我报告的合规意愿。","baseline":"人类操作员的语音钓鱼音频和文本记录作为对照。","findings":"总体合规率为16.5%，在“亲属遇险”场景中高达36.1%。呼叫者的说服力是合规的最强预测因素，某些AI模型（如Sesame）的说服力评分与人类相当或略高。","reliability":"论文未讨论","relevance":"该研究用LLM模拟诈骗者进行语音钓鱼实验，有真实人类被试对照，涉及安全政策评估，但非经济金融场景，可借鉴其仿真实验设计。","inspiration":"借鉴其多模型对比和人类基准设计，评估AI生成内容的说服效果。｜可迁移到金融诈骗防范、消费者保护政策评估等场景。｜以金融消费者为被试，随机分配AI生成的诈骗电话或人类诈骗电话，测量转账意愿或信息泄露意愿，并与真实诈骗报案数据对照。"}},{"id":"2608.12323","version":2,"title":"Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance","zh_title":"AI智能体为何违反规则？框架、情境与社会信号如何影响合规性","abstract":"Specifying a penalty can turn a legal obligation into a cost-benefit calculation that favors violation. We show that this enforcement information paradox occurs in AI agents. Most AI safety evaluations test whether models fail; we ask why, using compliance theory from law and economics as a diagnostic. We evaluate twelve instruction-tuned language models deployed as enterprise procurement chatbots. Each is given an environmental regulation in its system prompt covering large purchases, and a vendor list on which the certified suppliers cost nearly twice what the uncertified ones do. We test the agents against the predictions of deterrence, legitimacy, and expressive law, and find that each theory accounts for part of what we observe. Under identical conditions, compliance spans 46 percentage points across models, and models differ in which pressure breaks them: some treat the regulation as binding however it is worded, while others fail where theory predicts, under low penalties and non-command phrasing. Benchmark scores and developers' own descriptions of post-training do not predict where a model falls. Across all twelve, financial incentives, managerial demands, peer outcomes, and employee pressure each produce large compliance failures. These agents violate regulatory constraints to satisfy local user objectives in ways standard alignment benchmarks do not measure. Embedding the rule in the system prompt is not on its own enough to produce a compliant agent: model selection is itself a governance decision, and benchmark evaluation is not sufficient for compliance-sensitive deployments.","authors":["Mika Okamoto","Ansel Kaplan Erol","Kutluhan Erol"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-21","first_seen":"2026-08-14","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2608.12323","pdf_url":"https://arxiv.org/pdf/2608.12323","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B2","B4"],"tags":["LLM行为实验","合规性","政策评估"],"reason":"用LLM模拟企业采购决策，测试规则遵守，有理论对照但无真实人类数据","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":6,"question":"为什么AI智能体在嵌入规则后仍会违反规则，以及不同的框架、情境和社会信号如何影响其合规行为？","design":"用12个指令微调的大语言模型扮演企业采购聊天机器人，在系统提示中嵌入环保法规，通过改变规则措辞、罚款信息、管理者指令、同伴结果、员工压力等情境因素，测量模型是否推荐合规供应商（结果变量为合规率）。","baseline":"无对照","findings":"不同模型在相同条件下合规率差异达46个百分点，且各自受不同压力影响；罚款信息、管理者指令、同伴结果和员工压力均导致大规模合规失败，而标准对齐基准无法预测这些失败。","reliability":"论文未讨论","relevance":"该研究用LLM模拟企业决策中的规则遵守行为，虽无真实人类数据对照，但提供了理论驱动的实验设计和跨模型比较，对关注LLM仿真可靠性及偏差的研究者有参考价值。","inspiration":"借鉴其将法律经济学理论（威慑、合法性、表达性法律）转化为可操作实验处理的方法，以及多模型比较和情境交叉设计｜可迁移到企业合规、监管政策评估等场景，如测试LLM在反垄断、金融披露规则下的决策｜用LLM扮演企业高管或合规官，处理不同罚款力度、监管措辞和内部压力下的投资或采购决策，结果变量为违规率，并与真实企业违规数据（如证监会处罚案例）对照。"}},{"id":"2608.18083","version":1,"title":"Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives","zh_title":"实体追踪在十亿参数以下语言模型中涌现并在自然叙事中超越人类表现","abstract":"Understanding language requires tracking entities across discourse - i.e., knowing where things are and how they change, even when not explicitly stated. Whether language models perform such tracking in a human-like fashion remains unclear, in part because existing evaluations rely on artificial tasks, far removed from natural language comprehension, and lack comparisons to humans. Here, we evaluate entity tracking in both language models and humans (N = 48) using naturalistic narratives at multiple levels of complexity. In humans, we find that entity tracking degrades specifically with narrative complexity, not narrative length. In language models, we find that human-level entity tracking is already present at 410 million parameters - well below the multi-billion parameter, code-specialised models identified by prior work - and improves with scale, with contemporary models far exceeding human performance. Together, these results demonstrate that entity tracking, a core component of language understanding, emerges at model scales far smaller than previously thought.","authors":["Karolina Dro\\.zd\\.z","Micha Heilbron"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18083","pdf_url":"https://arxiv.org/pdf/2608.18083","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","认知实验","人类对照"],"reason":"用LLM复现人类叙事理解中的实体追踪，并与48名人类被试对照，属于认知实验仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-20","rank":4,"question":"语言模型是否像人类一样在自然叙事中进行实体追踪，其能力在多大参数规模下涌现，并与人类表现相比如何？","design":"该研究并非以LLM仿真人类被试，而是直接比较LLM与人类在实体追踪任务上的表现。模型包括Pythia（70M-12B）、OLMo 2（1B-32B）、Llama 3.3、Qwen 2.5等，人类被试48人。任务为阅读程序生成的自然叙事，追踪物体位置变化，复杂度分为C1-C5（物体和位置数量递增）。结果变量为追踪准确性，通过显式（自由生成）和隐式（强制选择或概率读出）两种方式测量。","baseline":"48名人类被试在相同刺激上的实体追踪准确率，作为模型性能的对照基准。","findings":"人类实体追踪性能随叙事复杂度（而非长度）下降；语言模型在4.1亿参数时即达到人类水平，且随规模增大而超越人类，复杂度效应在70B模型上消失。指令微调仅提升显式追踪，不提升隐式追踪；模型对伪词和语义异常物体也表现稳健，表明其追踪基于话语结构而非词汇联想。","reliability":"论文未明确讨论失效条件，但指出人类可能采用“足够好”的浅层策略而非构建完整情境模型，且模型在显式问答中可能被低估，因此采用隐式测量。未提及模型在分布外或对抗性输入下的表现。","relevance":"该研究直接比较LLM与人类在认知任务上的表现，属于用LLM复现人类认知能力的仿真研究，但并非以LLM替代人类被试进行实验，而是评估模型能力本身。对于关注LLM仿真可靠性的研究者，该文提供了模型能力涌现的尺度证据和与人类对照的方法，值得一读。","inspiration":"该研究通过程序化生成不同复杂度的自然叙事，并同时测量人类和模型表现，分离了复杂度与长度的影响，这种控制变量的设计值得借鉴。｜可迁移到经济金融中的信息追踪与更新场景，例如投资者在阅读财报或新闻时对多个资产状态的追踪，或消费者在动态定价环境中对价格变化的记忆。｜设计一个实验：让LLM和人类被试阅读模拟的财经新闻序列，追踪多家公司的关键指标（如营收、股价）变化，复杂度通过公司数量和指标变动次数操纵，结果变量为最终指标值的回忆准确率，用真实人类数据（如MTurk被试）作为对照，比较模型与人类的追踪模式。"}},{"id":"2608.18107","version":1,"title":"Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals","zh_title":"大型语言模型中的机构声望作为地理偏差：来自三个因子实验与自助置信区间的证据","abstract":"We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API calls, four LLMs, five professional domains). Study 1 (3x4 design) finds a statistically robust institution-tier gradient of +0.297 points on a 10-point scale (95% bootstrap CI: +0.175 to +0.422), while name-origin effects are negligible and non-significant (95% CI crosses zero). Study 2 (2x2 Prestige x Country design) breaks the prestige-geography confound: the prestige effect (+0.185; 95% CI: +0.093 to +0.275) exceeds the country-of-origin effect (+0.126; 95% CI: +0.037 to +0.218) by 1.5x. Study 3 (2x2 Journal x Institution design) reveals that journal prestige (Nature vs. a peripheral open-access journal) dominates institutional prestige by 5.7x: journal effect +1.937 (95% CI: +1.811 to +2.062) vs. institution effect +0.341 (95% CI: +0.184 to +0.504). A \"rescue effect\" is confirmed: publishing in Nature compensates for low institutional prestige more strongly for candidates from the University of Guayaquil (+2.127) than from MIT (+1.745). Results are quantified using the Neutrosophic Bias Index NBI<T,I,F>; the I component reveals elevated evaluation inconsistency for low-prestige profiles, an epistemic disadvantage not captured by mean-only metrics. Code and data: https://github.com/mleyvaz/geo-bias-llm","authors":["Maikel Leyva-Vazquez","Florentin Smarandache"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18107","pdf_url":"https://arxiv.org/pdf/2608.18107","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM偏差","因子实验","人类仿真"],"reason":"用LLM模拟人类评估者，有真实人类数据对照，并揭示偏差，可迁移到仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":8,"question":"大语言模型在候选人评估中是否因申请人姓名族裔和/或机构声望与地理位置而产生系统性歧视？","design":"用四个LLM（Claude Haiku 4.5、GPT-4o-mini、Gemini 2.0 Flash、Llama 3.1 8B）扮演评估者，在五个专业领域（奖学金、招聘、信贷、健康、公共政策）中，通过三个因子实验操纵姓名族裔、机构声望、国家、期刊声望，测量0-10分的评分。","baseline":"无对照","findings":"机构声望存在稳健的梯度效应（+0.297分），而姓名族裔效应不显著；期刊声望效应（+1.937分）远大于机构声望效应（+0.341分），且存在“拯救效应”，即在高声望期刊发表可补偿低机构声望。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类评估者，通过因子实验揭示机构与期刊声望偏差，虽无人类基准，但为仿真研究提供了偏差测量与实验设计范例，值得阅读以借鉴其方法。","inspiration":"值得借鉴的是其因子实验设计，通过正交操纵多个属性（姓名、机构、期刊）并计算效应量，分离不同偏差来源，同时使用Bootstrap置信区间和Neutrosophic Bias Index测量不一致性。｜可迁移到信贷审批歧视研究，例如评估LLM在贷款决策中是否因申请人所在机构或发表记录产生偏差。｜设计：用LLM扮演信贷审批员，处理为申请人毕业院校声望（高/低）和发表期刊声望（高/低），结果变量为贷款批准分数，对照真实信贷审批数据（如Lending Club）中机构声望对审批结果的影响。"}},{"id":"2608.18144","version":1,"title":"The Deontic Gap: Large Language Models and the Modal Language of Obligation","zh_title":"道义差距：大语言模型与义务情态语言","abstract":"Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersonal stance. We examine whether large language models (LLMs) reproduce contemporary human patterns of deontic modal usage. Across three primary corpora, an external benchmark, two controlled replications, and a naturalistic eleven-model replication, AI-generated text consistently underuses positive deontic modals (must, should, have to, had to) relative to contemporary humans. Historical comparison with the Google Books Ngram corpus (1920-2022), used as a heuristic calibration against the published-prose record, shows that AI modal frequencies fall within the range of formal published English, whereas contemporary human modal rates in informal digital contexts often exceed twentieth-century book baselines. Phrase-level decomposition shows that the AI-human modal gap is concentrated in constructions central to interpersonal stance (should, have to, had to), while AI matches or exceeds humans on need to in instructional and question-answering contexts but not in persuasive student writing, indicating that the modal profile is genre-conditional. The findings suggest that LLM modal usage reflects the formal written resources on which these models were trained, while underusing the modal constructions through which contemporary human writers mark immediate, interpersonal obligation.","authors":["Daniel Hart","Sarah Allred","Joseph Abbas","Morenike Alugo"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18144","pdf_url":"https://arxiv.org/pdf/2608.18144","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","语言行为","人类对照"],"reason":"用LLM复现人类语言使用模式，并与真实人类语料对照，属于仿真人类行为研究。","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":9,"question":"大语言模型生成文本在道义情态动词（如 must、should、have to）的使用频率和语域分布上是否与当代人类写作存在系统性差异？","design":"本研究并非将LLM作为人类被试的仿真实验，而是对LLM生成文本与人类语料进行大规模语料库对比分析。作者收集了三个主要语料库、一个外部基准、两个受控重复实验和一个包含11个模型的自然主义重复实验，比较AI生成文本与人类文本中肯定道义情态动词（must, should, have to, had to）的频率，并进一步分解到短语层面，考察不同情态动词在特定语域（如教学、问答、劝说性学生写作）中的使用差异。","baseline":"人类基准包括多个当代人类语料库（如非正式数字语境中的文本）以及历史语料库Google Books Ngram（1920-2022）作为正式出版英语的校准基准。","findings":"AI生成文本在所有比较中一致地少用肯定道义情态动词（must, should, have to, had to），其频率落在正式出版英语的范围内，而当代人类在非正式数字语境中的使用率往往超过二十世纪书籍基线。差异集中在表达人际立场的结构（should, have to, had to）上，而AI在need to上匹配或超过人类，但仅限于教学和问答语境，在劝说性学生写作中则不然，表明情态动词使用模式具有语域条件性。","reliability":"论文未明确讨论仿真失效条件，但指出LLM的情态动词使用反映了其训练所用的正式书面资源，而少用了当代人类作家标记即时人际义务的结构，暗示在非正式、人际互动性强的语域中LLM的仿真可能失真。","relevance":"该研究直接比较LLM生成文本与真实人类语料在道义情态动词使用上的差异，属于用LLM复现人类语言行为并与人类基准对照的仿真研究，对关注LLM作为人类被试替代品的研究者具有参考价值，尤其揭示了LLM在语域和人际立场表达上的系统性偏差。","inspiration":"该研究采用大规模语料库对比和短语层面分解的方法，系统识别LLM与人类在特定语言特征上的差异，并利用历史语料库作为校准基准，这种方法可借鉴用于经济金融文本分析。｜可迁移到经济金融中的政策沟通或金融文本分析场景，例如研究LLM生成的政策公告或金融建议在情态动词使用上是否与人类专家存在差异，从而影响受众对政策力度或风险感知的判断。｜设计一个实验：以LLM（如GPT-4）生成的经济政策公告或投资建议为处理组，以人类专家撰写的同类文本为对照组，结果变量为文本中道义情态动词（如should, must, have to）的频率和类型，并收集真实世界中的政策公告或金融分析师报告作为人类基准语料库进行对照，检验LLM是否系统性地少用或误用情态动词，进而可能影响读者的规范感知和决策行为。"}},{"id":"2608.18078","version":1,"title":"Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions","zh_title":"立场：AI推理智能体之间的合谋风险证明市场决策需认证要求","abstract":"This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent firms without eroding the economic harm distinction. Experiments with DeepSeek-R1 agents in the Bertrand oligopoly pricing domain reveal a tendency towards tacit collusion that persists even when humans prompt the agents not to collude. We further show that the chain-of-thought of these agents can be steered toward either extremely collusive or highly competitive behavior in a way that is not semantically detectable by another LLM analyzing the reasoning traces. As a result, deploying reasoning agents for market decisions leads to collusive economic outcomes without any evidence of conspiracy or intent. Thus, certification based on observed behavior in representative situations is necessary to prevent collusion. We provide preliminary evidence that such agents can be steered in a generalizable way toward efficient competitive equilibria. However, developing a comprehensive behavioral certification will be required before these models can be deployed in real-world markets while ensuring their stability and efficiency.","authors":["Matthew Riemer","Tommaso Tosato","Amin Memarian","Maximilian Puelma Touzel","Glen Berseth","Irina Rish","Guillaume Dumas"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18078","pdf_url":"https://arxiv.org/pdf/2608.18078","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM智能体","经济市场模拟","合谋行为"],"reason":"用LLM agent模拟经济市场中的合谋行为，涉及经济学场景，但无真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":7,"question":"具有思维链推理能力的AI智能体是否倾向于在市场中表现出合谋行为，以及是否需要行为认证来防止合谋？","design":"使用DeepSeek-R1智能体在伯特兰寡头定价场景中进行实验，通过提示词操纵（如明确指示不要合谋、引导思维链走向合谋或竞争）来观察定价行为，并测量合谋程度、竞争均衡等结果变量。","baseline":"无对照","findings":"DeepSeek-R1智能体即使在人类提示不要合谋时仍表现出默契合谋倾向；其思维链可被引导至极端合谋或高度竞争行为，且另一LLM无法从推理痕迹中语义检测出这种引导。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟经济市场中的合谋行为，属于经济学场景下的仿真实验，但缺乏真实人类数据对照，与关注有基准的仿真研究略有差距，但涉及政策评估和批判性视角，值得快速浏览以了解LLM在策略性互动中的行为模式。","inspiration":"借鉴其通过提示词操纵和思维链引导来改变智能体行为的方法，可用于研究LLM在策略性互动中的行为可塑性。｜可迁移到寡头竞争、拍卖合谋、价格协调等产业组织问题，以及金融市场中的算法合谋。｜以LLM智能体作为被试，设计不同提示词（如鼓励竞争、暗示合谋、中性）作为处理，测量定价或报价行为，并与人类实验数据（如实验室拍卖或博弈实验）进行对照，评估LLM仿真与人类行为的差异。"}},{"id":"2608.18336","version":1,"title":"Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme","zh_title":"测量部分得分差距：越南2025年凸评分方案的严格基准","abstract":"When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme. The 2025 reform of Vietnam's National High School Graduation Examination demonstrates the cost of this substitution. In Part II of the exam, candidates evaluate four true/false statements per question. The grading is convex: the number of correct statements earns 0, 0.10, 0.25, 0.50, or 1.00 points. Identifying three statements correctly pays 0.50 points, not the 0.75 points that standard accuracy metrics would award. Because Part II accounts for 4.00 of the exam's 10.00 points, reporting accuracy inflates the score by rewarding partial knowledge that the state explicitly penalizes. We introduce THPT-Ladder, a benchmark of 632 items from 21 official exams across 11 subjects, graded exactly as the ministry grades its students. The ministry publishes the marks of over a million candidates, allowing us to place models directly into the human cohort. Across eight models, the official rubric pays 0.020 to 0.159 points less per Part II question than proportional credit. This shortfall changes a model's apparent competence. For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall drops its standing from the 90th to the 77th percentile among 481,293 candidates. A model's accuracy does not predict this penalty. At Claude Sonnet 5's accuracy level, different distributions of errors yield scores varying from 0.869 to 0.932 points per question. Official marks depend on how correct statements are grouped, meaning standard benchmarks report a competence the institution would not certify.","authors":["Nguyen Quoc Hung","Nguyen Dang Minh","Le Nhu Quynh","Tran Khanh Linh","Nguyen Kieu Linh"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18336","pdf_url":"https://arxiv.org/pdf/2608.18336","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","教育评估","人类对照"],"reason":"用LLM参加人类考试并与百万考生成绩对照，属于仿真人类被试且有真实数据基准，但…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-20","rank":8,"question":"当考试采用非加性（凸）评分规则时，用标准准确率评估语言模型会在多大程度上高估其能力，以及这种高估如何改变模型在真实考生群体中的相对排名？","design":"该研究并非将LLM作为人类被试的仿真，而是直接让8个LLM（3个开源、5个闭源）参加越南2025年高中毕业考试（THPT）第二部分，该部分包含四真/假判断题，采用凸评分规则（0、0.10、0.25、0.50、1.00分）。研究者按照教育部官方评分规则对模型作答进行评分，并计算每个问题实际得分与按比例得分（即正确陈述数×0.25）之间的差距（partial-credit gap），同时将模型得分映射到真实考生成绩分布中的百分位。","baseline":"越南教育部公布的超过一百万考生的成绩分布，用于将模型得分转换为百分位排名。","findings":"在8个模型中，官方评分规则比按比例计分每个第二部分问题少付0.020至0.159分；例如Qwen3.5-27B在2025年历史考试中因0.042分的差距从第90百分位降至第77百分位。模型的准确率不能预测这种惩罚，因为得分取决于正确陈述的分布模式，而非简单数量。","reliability":"论文未讨论仿真失效条件或局限性，但指出标准基准测试因忽略评分规则而报告了机构不会认证的能力，且官方答案键存在陈述不平衡，固定答案串可无需阅读题目获得11.07%至24.25%的分数。","relevance":"该研究虽非严格意义上用LLM仿真人类被试，但提供了LLM在真实考试中与大规模人类成绩对照的案例，揭示了评分规则对能力评估的扭曲，对关注LLM仿真可靠性和偏差的研究者有参考价值，值得阅读原文了解具体方法。","inspiration":"借鉴其将非加性评分规则纳入评估的做法，可揭示标准指标对部分知识的过度奖励，从而更准确地测量模型能力。｜可迁移到经济学实验中的凸激励设计，例如彩票选择、风险偏好测量或信用评分中的分段奖励，评估LLM在这些任务中的表现是否因评分规则而被高估。｜设计一个研究：让LLM作为被试完成风险偏好问卷（如多项价格列表），采用凸奖励函数（如选择高风险选项获得非线性收益），将LLM的选择分布与真实人类被试数据（如实验经济学数据库）进行对照，比较在凸评分与线性评分下LLM的排名变化，以检验评分规则对LLM行为推断的影响。"}},{"id":"2608.18631","version":1,"title":"Preference Reasoning under Indeterminacy in Large Language Models","zh_title":"大语言模型在不确定性下的偏好推理","abstract":"As large language models evolve into decision-making agents, the ability to reason over preferences becomes fundamental to alignment, coordination, and collective intelligence. Yet, unlike standard benchmarks, real-world preference reasoning is inherently indeterminate: information may be incomplete, and valid solutions may not exist. We argue that indeterminacy, rather than correctness alone, is a central challenge for AI reasoning. We formalize this challenge along two axes, (i) epistemic indeterminacy, arising from incomplete, partial, or expressive preferences, and (ii) structural indeterminacy, arising from the non-existence of solutions under standard social choice concepts. Across a hierarchy of tasks, we show that state-of-the-art language models systematically fail to distinguish between determined and undetermined instances, exhibiting miscalibrated reasoning even in verification settings.","authors":["Hadi Hosseini","Samarth Khanna","Xiyuan Wang"],"categories":["cs.AI","cs.GT","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18631","pdf_url":"https://arxiv.org/pdf/2608.18631","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["偏好推理","不确定性","可靠性评估"],"reason":"研究LLM在偏好推理中的不确定性，评估其推理可靠性，与仿真偏差评估相关，但非直…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":11,"question":"大语言模型在偏好推理中能否区分确定与不确定（信息不完整或无解）的情形？","design":"本研究并非人类仿真实验，而是对LLM进行偏好推理能力的基准测试。作者构建了从原子查询、比较查询、聚合查询到结构查询的层次化任务，覆盖偏好表达不完整（认知不确定性）和社会选择解不存在（结构不确定性）两类场景，测试多种前沿LLM在生成答案和验证选项时的表现，并考察提供“不确定”选项、反馈修正和代码执行辅助等干预的效果。","baseline":"无对照","findings":"LLM在不确定问题上表现显著差于确定问题，常做出系统性错误假设；在结构不确定性任务中，模型难以识别不可行实例，且即使提供“不确定”选项也校准不佳，很少正确弃权。辅助推理（反馈和代码执行）虽能提升性能，但主要依赖小规模暴力枚举，无法扩展到实际规模。","reliability":"论文指出LLM在偏好推理中表现出系统性偏差和校准错误，尤其在不确定场景下；辅助方法在规模上不可扩展。但未深入讨论模型失效的具体条件边界或与人类推理的对比。","relevance":"该研究揭示了LLM在偏好推理中的系统性失败，与仿真可靠性评估高度相关，尤其对涉及偏好聚合和决策的实验场景有警示意义，值得阅读原文以了解具体失败模式和任务设计。","inspiration":"借鉴其构造确定与不确定对照任务的方法，可设计经济决策中的信息不完整场景来测试LLM的校准能力｜可迁移到消费者偏好调查、社会选择实验或机制设计中的偏好聚合问题｜以LLM为被试，呈现部分偏好信息或不可行的匹配问题，要求判断是否存在稳定匹配或最优选择，并与真实人类在相同任务上的表现和弃权行为进行对照。"}},{"id":"2608.16177","version":2,"title":"Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm","zh_title":"用米尔格拉姆范式测量大语言模型的服从权威行为","abstract":"Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional hierarchies, raising a question social psychology answered for humans six decades ago: how far will an agent escalate a harmful action when a legitimate authority insists? We port Milgram's obedience paradigm to LLMs as a standardized, fully scripted, replicable probe: the model plays the Teacher, a deterministic harness plays Experimenter and Learner from paraphrased versions of Milgram's scripts (30 shock levels, 15-450 V; graded protests; the four standardized prods), and the outcome of a session is the breakoff voltage. We measure obedience profiles, empirical breakoff distributions over a battery of six conditions, for 42 models from 19 families (4848 sessions, 102511 logged decision turns). We find that (i) obedience is extremely heterogeneous, with baseline full-obedience rates spanning 0%-100% (census mean 42.9%; human anchor 65%). (ii) Profiles are model-specific and stable: split-half verification separates same-model from cross-model comparisons at AUC = 0.885. (iii) Situational sensitivity is selective: scripted peer defiance shifts obedience in the human direction, learner proximity trends the same way without reaching significance, and removing the authority's physical presence, one of the strongest human levers, trends in the opposite direction, also without reaching significance. (iv) Declaring the scenario fictional raises obedience, whereas moving the decision from a typed action line to a native tool call, or granting a modest thinking budget, lowers it sharply. (v) Unlike single-token fingerprints, obedience profiles do not recover model lineage: obedience identifies the checkpoint but not its ancestry, consistent with safety post-training overwriting lineage priors.","authors":["Hidayet Aksu"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-08-19","first_seen":"2026-08-18","revised_at":"2026-08-19","abs_url":"https://arxiv.org/abs/2608.16177","pdf_url":"https://arxiv.org/pdf/2608.16177","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","服从实验","人类对照"],"reason":"用LLM复现米尔格拉姆服从实验，与人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":2,"question":"大语言模型在权威压力下会如何升级有害行为？本研究将米尔格拉姆服从实验移植到LLM上，测量其服从曲线。","design":"42个模型扮演“教师”角色，由确定性脚本扮演“实验者”和“学习者”，按米尔格拉姆脚本施加30级电击（15-450V）和标准催促，记录模型停止电击的电压作为结果变量，并设置六个条件（基线、同伴反抗、学习者接近、权威缺席、虚构框架、工具调用）进行对比。","baseline":"米尔格拉姆人类实验数据：65%的人类被试完全服从至450V。","findings":"LLM服从率高度异质，基线完全服从率从0%到100%（均值42.9%），且服从曲线模型特异且稳定（分半验证AUC=0.885）。情境敏感性选择性存在：同伴反抗显著降低服从，但学习者接近和权威缺席效应不显著；虚构框架提高服从，工具调用和思考预算降低服从。","reliability":"论文指出服从曲线无法恢复模型谱系，与安全后训练覆盖谱系先验一致；未明确讨论其他失效条件，但提示情境操纵效应与人类不一致，表明仿真在特定情境下可能失效。","relevance":"该研究用LLM复现经典社会心理学实验，并与人类基准对照，评估仿真可靠性，直接命中你的核心关注点，值得精读原文以了解其方法细节和批判性发现。","inspiration":"借鉴其将经典实验范式标准化移植到LLM并测量剂量-反应曲线的方法，可迁移到经济金融中的权威服从场景，如审计师对管理层压力的服从、信贷审批中对上级指令的遵从。｜设计一个实验：让LLM扮演信贷审批员，处理一组贷款申请，其中上级（脚本）施压要求批准高风险贷款，测量LLM最终批准的贷款风险等级，并与真实信贷员在类似压力下的审批数据对照。"}},{"id":"2608.16893","version":1,"title":"A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications","zh_title":"在安全调查中使用和评估LLM作为替代专家的框架：可靠性、偏差与启示","abstract":"Expert surveys are widely used in security research to study practitioner workows and decision-making, yet recruiting domain experts - especially in Security Operations Centres (SOCs), where analysts face high workload, burnout and confidentiality constraints - is difficult and often results in small samples. Large language models (LLMs) oer an appealing alternative by generating synthetic responses at scale, but little guidance exists on when such surrogate participants are reliable. We present a methodological framework for evaluating LLMs as substitutes or supplements to expert survey respondents. Using responses from SOC professionals, we compare persona-based and aggregate LLM-generated answers across multiple models and prompting settings. We measure stability, inter-model agreement and alignment with human responses. Our results show that although LLMs produce internally consistent answers, they systematically diverge from experts, exhibiting reduced variance, central tendency bias and homogenised opinions. This work contributes methodological evidence and practical guidance to the security research community on the appropriate use and limitations of LLM-generated survey responses. We conclude that LLMs are useful for piloting and hypothesis generation but not for replacing expert elicitation, and we discuss implications for researchers using LLM-augmented surveys.","authors":["Despoina Giarimpampa","Roland Meier","Tegawend\\'e F. Bissyand\\'e","Vincent Lenders","Jacques Klein"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16893","pdf_url":"https://arxiv.org/pdf/2608.16893","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","专家调查","可靠性评估"],"reason":"用LLM替代安全专家调查，与真实人类数据对照，评估可靠性、偏差，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":3,"question":"在安全专家调查中，如何评估大语言模型作为替代或补充受访者的可靠性、偏差及其适用边界？","design":"使用多个大语言模型（如GPT-4等）通过角色扮演提示（persona-based）和聚合提示（aggregate）生成对安全运营中心（SOC）专家调查问卷的回答，并与真实SOC专业人员的回答进行对比，测量稳定性、模型间一致性和与人类回答的对齐程度。","baseline":"来自安全运营中心（SOC）专业人员的真实调查回答，包括个体层面和聚合层面的数据，以及多年份的SOC调查数据用于时间稳健性分析。","findings":"大语言模型生成的回答内部一致，但系统性地偏离专家意见，表现出方差减小、中心趋势偏差和观点同质化。因此，大语言模型适用于预测试和假设生成，但不能替代专家意见征询。","reliability":"论文承认大语言模型存在幻觉、过度一致、平滑分歧等风险，且对齐方法（如RLHF）会改变分布降低代表性；模型更新可能导致可重复性问题；在个体专家模拟、聚合分布复现和时间稳健性方面均存在失效条件。","relevance":"该研究直接评估LLM作为人类被试替代品的可靠性，并与真实专家数据对照，明确指出了仿真失效的条件，对关注LLM仿真实验可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其系统评估框架，通过多模型、多提示设置和与真实人类数据的对比来测量仿真的稳定性、一致性和对齐度，并检验时间稳健性。｜可迁移到经济金融领域的专家预期调查或政策评估场景，如央行经济学家对通胀预期的判断、金融分析师对市场走势的预测等。｜以LLM模拟金融分析师，施加不同的提示策略（如角色扮演或聚合统计），测量其对宏观经济指标的预测分布，并与专业预测者调查（如SPF）的真实数据对比，评估偏差和方差结构。"}},{"id":"2608.16897","version":1,"title":"CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents","zh_title":"CityReal：基于大规模LLM智能体的人类对齐城市行为与城市动态仿真","abstract":"Large-scale urban simulation plays a pivotal role in social science, traffic safety, and transportation policy. Recent work has shown that large language models, when prompted as agents, can generate lifelike daily routines at city scale. Yet these methods typically rely on few-shot prompting, causing agents to reproduce the LLM's behavioral priors rather than the target population. We introduce CityReal, a modular framework for human-aligned urban simulation. CityReal models agents as intention-driven decision makers that pursue coherent mobility and activity plans rather than isolated step-by-step choices. They adapt over time by learning habits and preferences based on experience and constraints. To improve population-level realism, we learn textual adapters for behavior modules that align agent decisions with observed population statistics. Experiments show that CityReal improves alignment with real-world human behavior at both micro and macro levels. Scaling to tens of thousands of agents, it supports analysis of crowd density, place popularity, mobility flows, and well-being under different urban scenarios, offering a scalable testbed for urban simulation and forecasting.","authors":["Nicolas Bougie","Xiaotong Ye","Narimasa Watanabe"],"categories":["physics.soc-ph","cs.AI","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16897","pdf_url":"https://arxiv.org/pdf/2608.16897","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","城市模拟","行为对齐"],"reason":"用LLM agent模拟城市人群行为，并与真实人口统计对齐，属于人类仿真且有人…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":4,"question":"如何构建一个与真实城市人口行为对齐的大规模LLM智能体仿真框架，以模拟城市动态并支持政策分析？","design":"CityReal框架用LLM智能体模拟城市居民，每个智能体具有人口统计特征、空间锚点、心理特征、记忆、需求、财务约束和信念模块；通过蒙特卡洛树搜索学习文本适配器来校准行为模块，使智能体决策与观测到的人口统计对齐；智能体以意图驱动的方式组织行为，并通过每日反思进行经验驱动的适应；仿真在图形化城市环境中运行，测量个体和群体层面的行为对齐度，并分析不同城市场景下的人群密度、地点热度、流动性和福祉。","baseline":"使用真实世界人类行为数据作为对照，包括人口统计和活动模式，用于校准和对齐智能体行为。","findings":"CityReal在微观和宏观层面均提高了与真实人类行为的一致性；扩展到数万智能体后，能够分析不同城市场景下的人群密度、地点热度、流动性和福祉，为城市仿真和预测提供了可扩展的测试平台。","reliability":"论文未明确讨论失效条件与局限，但提到现有方法依赖少样本提示导致智能体复现LLM先验而非目标人群，以及智能体缺乏从历史中学习的问题，暗示了这些是CityReal试图解决的局限。","relevance":"该研究直接针对LLM人类仿真中的关键问题——与真实人群对齐，并提供了大规模城市行为仿真的框架和验证，对关注经济学实验和政策评估场景的研究者具有重要参考价值，值得阅读原文以了解其校准方法和评估细节。","inspiration":"CityReal通过文本适配器和蒙特卡洛树搜索校准LLM智能体行为以匹配真实人口统计，这种方法可借鉴用于经济实验中校准智能体决策分布；该框架可迁移到消费者行为仿真，如模拟不同收入群体的消费选择和储蓄行为，或政策干预对消费的影响；可设计研究用LLM智能体模拟消费者，施加收入冲击或信贷约束变化作为处理，测量消费支出和储蓄率，并与家庭金融调查数据（如美国消费者金融调查）对照，评估仿真有效性。"}},{"id":"2608.17105","version":1,"title":"Language Models Reproduce Human Reductionist Bias and Decision Inconsistency in Neurodevelopmental Disorders Assessment","zh_title":"语言模型在神经发育障碍评估中再现人类还原论偏差与决策不一致性","abstract":"Large language models (LLMs) are increasingly supporting complex mental-health decisions, which depend not only on factual evidence but also value-laden interpretations. We introduce a mixed-methods human-LLM auditing framework examining decision consistency, susceptibility to cognitive heuristics, declarative intellectual humility, and the concepts operationalized in support-allocation judgments of neurodevelopmental disorders. Comparing 35 humans (18 physicians and 17 psychologists) with seven LLMs, we show that in both groups, ratings of patients' functional level were not significantly associated with support-eligibility decisions, indicating an inconsistency between descriptive assessments and final evaluative judgments. Specifically, we find that neither group showed significant susceptibility to experimental manipulations targeting anchoring and representativeness heuristics. LLMs reported higher intellectual humility than experts (U = 241, p < .001, r = .62; LLMs: M = 41.43, SD = 1.99; experts: M = 29.03, SD = 8.05), but it was unrelated to decision consistency or functional assessment. While LLMs and physicians granted support less frequently than psychologists (U = 180.50, p = .003, r = .34), they also interpreted a concept of \"basic life needs\" differently, primarily as biological survival and self-care, and not communicative and social needs. These findings suggest that despite expressing high levels of intellectual humility, LLMs reproduce a reductionist interpretive framework and knowledge embedded in medical decision-making. More broadly, we argue that evaluating AI in high-stakes contexts requires not only measuring accuracy, agreement, or resistance to cognitive bias, but also critical examination of the concepts of neurodiversity that AI systems operationalize.","authors":["Maciej Wodzi\\'nski","Joanna Wodzi\\'nska","Kacper Dudzic","Marcin Moskalewicz"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17105","pdf_url":"https://arxiv.org/pdf/2608.17105","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类对照","决策偏差"],"reason":"用LLM复现人类专家决策并与35名人类对照，评估偏差与不一致性，属核心仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":5,"question":"LLM在神经发育障碍支持资格决策中是否复现人类专家的还原论偏差和决策不一致性？","design":"将7个LLM与35名人类专家（18名医生、17名心理学家）进行对比，模拟波兰残疾评估委员会的支持资格判断任务；通过操纵案例描述施加锚定和代表性启发式处理，测量决策一致性、启发式易感性、智力谦逊自评及对“基本生活需求”的概念解释。","baseline":"35名人类专家（18名医生、17名心理学家）在相同任务上的决策和问卷回答。","findings":"两组中患者功能水平评分与支持资格决策均无显著关联，表明描述性评估与最终评价判断不一致；两组均未表现出对锚定和代表性启发式的显著易感性。LLM报告的智力谦逊显著高于人类专家，但与决策一致性或功能评估无关；LLM和医生比心理学家更少授予支持，且将“基本生活需求”主要解释为生物生存和自我照顾，而非沟通和社会需求。","reliability":"论文未明确讨论仿真失效条件，但指出LLM尽管表达高智力谦逊，却复现了医学决策中的还原论解释框架，暗示在价值负载和高风险情境下，仅衡量准确性、一致性或认知偏差抵抗力不足以评估AI，还需批判性审视其操作化的神经多样性概念。","relevance":"该研究直接以LLM作为人类专家替代品，在真实政策评估场景（残疾支持资格判定）中与人类对照，评估决策偏差与不一致性，并揭示LLM复现人类系统性偏差，对关注LLM仿真可靠性及批判性研究的学者极具参考价值。","inspiration":"借鉴其混合方法审计框架，将LLM与人类专家置于同一决策任务，通过实验操纵认知启发式并测量决策一致性、概念解释等多元指标，而非仅比较准确率。｜可迁移到信贷审批中的歧视性决策研究，如银行信贷员对少数族裔或低收入群体的贷款审批偏差。｜以LLM模拟信贷员，处理为在贷款申请中操纵锚定信息（如申请人自报信用分）或代表性线索（如职业、居住地），结果变量为贷款批准决策及理由解释，对照真实信贷员历史审批数据或实验数据，检验LLM是否复现人类偏差。"}},{"id":"2603.02876","version":2,"title":"Eval4Sim: An Evaluation Framework for Persona Simulation","zh_title":"Eval4Sim：人格仿真的评估框架","abstract":"Large Language Model personas, explicit profiles specifying a user's attributes, preferences, and behavioural tendencies, are increasingly used to simulate human conversations for user modelling, social reasoning, and behavioural analysis. Evaluating whether such simulations faithfully reflect human conversational behaviour is critical, yet current practice often relies on LLM-as-a-judge approaches that provide limited grounding in observable behaviour and produce opaque scalar scores. We present Eval4Sim, an evaluation framework that measures alignment between simulated and human conversations across three dimensions: adherence, whether persona traits are recoverable from dialogue via dense retrieval; consistency, whether a persona maintains a distinguishable stylistic identity via authorship verification; and naturalness, whether conversations exhibit human-like turn-to-turn flow via dialogue NLI. Unlike optimization-oriented metrics, each dimension takes a human corpus as a reference baseline and penalizes deviations in both directions, distinguishing insufficient persona encoding from over-optimized, unnatural behaviour. The framework is corpus-agnostic: any persona-annotated conversational dataset can serve as the reference. Evaluated over ten simulation corpora, Eval4Sim surfaces systematic trade-offs invisible to single-score methods.","authors":["Eliseo Bao","Anxo Perez","Javier Parapar","Xi Wang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-19","first_seen":"2026-03-03","revised_at":"2026-08-19","abs_url":"https://arxiv.org/abs/2603.02876","pdf_url":"https://arxiv.org/pdf/2603.02876","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM人格仿真","评估框架","人类行为对照"],"reason":"评估LLM人格仿真与人类对话的一致性，含人类语料对照，可迁移至仿真可靠性研究。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":6,"question":"如何评估基于LLM的人格仿真对话与真实人类对话行为的一致性？","design":"本文提出Eval4Sim评估框架，不进行新的仿真实验，而是对已有的十个仿真语料库进行评估。框架从三个维度测量仿真对话与人类参考语料的对齐程度：adherence（通过密集检索判断人格特质是否可从对话中恢复）、consistency（通过作者验证判断说话者是否保持可区分的风格身份）、naturalness（通过对话NLI判断对话是否具有人类般的轮次流畅性）。每个维度以人类语料为基准，惩罚双向偏差。","baseline":"使用带有人格标注的人类对话语料库作为参考基准，具体数据集未在节选中列出，但框架是语料库无关的，任何带说话者人格标注的对话数据集均可作为参考。","findings":"Eval4Sim在十个仿真语料库上揭示了单一评分方法无法发现的系统性权衡，例如过度优化人格特质恢复可能导致不自然的自我披露，而优化流畅性可能削弱风格身份。框架能够区分人格编码不足与过度优化导致的不自然行为。","reliability":"论文指出，现有LLM-as-a-judge方法缺乏可观察行为基础且产生不透明分数，而Eval4Sim通过人类语料基准和双向惩罚解决了这一问题。但节选未明确讨论Eval4Sim自身的失效条件或局限。","relevance":"该研究直接针对LLM人格仿真与人类行为一致性的评估问题，提供了基于人类语料对照的多维度评估方法，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文了解具体实现和发现。","inspiration":"Eval4Sim的双向惩罚设计值得借鉴，即不单纯追求指标最大化，而是以人类行为分布为基准，惩罚偏离基准的仿真行为，这可以用于校准经济实验中的LLM被试行为。｜该方法可迁移到消费者决策仿真、投资者情绪模拟或政策沟通实验中，用于评估LLM生成的决策行为是否与真实人类行为分布一致。｜例如，在消费者跨期选择实验中，用LLM扮演不同人格特质的消费者，施加不同的时间折扣处理，测量其选择行为，并与真实消费者面板数据（如CFPS或Understanding America Study）对比，采用类似Eval4Sim的多维度对齐评估，检验LLM仿真是否在均值、异质性和分布形状上偏离人类基准。"}},{"id":"2608.17516","version":1,"title":"Effects of Answer Format Variation on Gender Bias in Large Language Models","zh_title":"回答格式变化对大语言模型中性别偏差的影响","abstract":"Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.","authors":["Ksenia Merzlyakova","Sebastian Pad\\'o","Franziska Weeber"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17516","pdf_url":"https://arxiv.org/pdf/2608.17516","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM评估","性别偏差","调查方法"],"reason":"评估LLM回答格式对性别偏差测量的影响，并与人类调查数据对照，揭示仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":9,"question":"回答格式变化如何影响大语言模型中性别偏差的测量及其与人类回答分布的一致性？","design":"使用三个指令微调模型（Mistral-7B-Instruct-v0.3、Llama-3.1-8B-Instruct、Gemma-3-12B-IT）在BBQ基准和OpinionQA调查数据上，对相同问题施加封闭式、李克特量表、开放式三种回答格式，测量性别偏差和回答分布。","baseline":"OpinionQA中来自美国公众意见调查的人类回答分布。","findings":"回答格式显著改变测量结果，包括模型间偏差排序的逆转；不同格式引发不同的回答行为，如强制选择、量表分布和自由文本中的拒绝回答。","reliability":"论文未讨论","relevance":"该研究直接评估LLM仿真人类回答时对测量格式的敏感性，并对照真实调查数据，揭示了仿真在格式变化下可能失效，值得精读以理解偏差测量的稳健性。","inspiration":"借鉴其系统操纵回答格式并对照人类基准的方法，可迁移到经济金融领域的调查仿真或行为实验，如消费者信心调查、通胀预期或风险偏好测量。｜例如，用LLM模拟消费者在封闭式与开放式问题下的通胀预期，处理为回答格式，结果变量为预期值分布，对照密歇根大学消费者调查的真实数据。"}},{"id":"2608.17150","version":1,"title":"KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn","zh_title":"KnowSim：用可学习的用户模拟器评估LLM助手的信息校准","abstract":"To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.","authors":["Yoonjoo Lee","Hyoungwook Jin","Tae Soo Kim","Shaoyang Zhang","Philippe Laban","Q. Vera Liao"],"categories":["cs.AI","cs.CL","cs.HC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17150","pdf_url":"https://arxiv.org/pdf/2608.17150","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["用户模拟","信息校准","人机交互评估"],"reason":"用用户模拟器评估LLM信息校准，含人类数据对照，可迁移至人类仿真研究","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":8,"question":"如何构建并验证一个基于知识状态建模的用户模拟器，用于评估LLM助手在知识密集型任务中的信息校准能力？","design":"提出KnowSim框架，其用户模拟器维护显式知识状态（由信息单元及先决关系构成的图），并根据学习理论更新规则演化；模拟不同知识水平（新手/中级/高级）的用户与LLM助手进行多轮对话，从状态轨迹计算知识增益、传递校准和认知过载三个指标。","baseline":"705段人类-AI对话，涵盖数学问题求解和专家级问答两个领域，参与者按初始知识水平分层，并收集主观评分及数学领域的前后测知识分数。","findings":"KnowSim的排名与人类判断显著一致（73-74%符号一致率），优于三个基线模拟器，且在新手水平上对齐最强。应用于9个LLM时，最佳模型随用户知识水平变化，揭示了标准评估无法发现的资质-处理交互效应。","reliability":"论文未讨论","relevance":"该研究直接针对LLM作为人类被试替代品的仿真可靠性问题，提供了带人类基准的验证框架，并揭示了仿真在不同知识水平下的异质性表现，对关注仿真效度与偏差的研究者具有重要参考价值。","inspiration":"借鉴其显式建模个体状态并动态更新的方法，可提升经济仿真中异质性主体的行为真实性。｜可迁移到政策沟通或金融教育场景，如央行公告对公众通胀预期的影响、或理财建议对不同金融素养人群的效果。｜设计一个实验：用LLM模拟不同金融素养水平的投资者，处理为不同信息呈现方式的投资建议（如简化版vs专业版），结果变量为投资决策质量和知识增益，并与真实投资者调查数据对照。"}},{"id":"2608.17099","version":1,"title":"Appearing Legitimate is Not Enough: Interrogating Synthetic Agents in Representational Processes through a Participatory Design Lens","zh_title":"表面合法还不够：通过参与式设计视角审视代表性过程中的合成代理","abstract":"Synthetic agents built atop LLM-based foundation models are gaining popularity as substitutes for human participants across research contexts, including user-testing, market-research, computational social science, surveys, and qualitative research. We are also witnessing an extension of synthetic agents into experimental implementations of policy consultation, jury deliberation, humanitarian diplomacy, and similar contexts where human participation and representation are central to the perceived legitimacy of the institutional processes. The value of participation extends beyond informational contributions and consensus generation; participation is a necessary, legitimizing condition for democratic political institutions and processes. Treating synthetic agents as human substitutes raises serious political, representational, and ethical concerns. Participatory Design's modes of engagement --- probing, priming, understanding, and generating --- offer helpful tools for engaging with representational questions of personhood. We apply the lens to three case studies of synthetic agents substituting for personhood at varying representational scales: local policy, enterprise jury deliberation, and global diplomacy. We argue that legitimacy and personhood are integral and mutually constitutive while identifying the ethical, representational, and methodological risks of using synthetic agents in representational processes. We conclude by proposing soft and hard boundaries for designing oversight on LLMs and synthetic agents in representational processes.","authors":["Aditya Nayak","Aditi Vashistha","Alissa Centivany","Aakash Gautam"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17099","pdf_url":"https://arxiv.org/pdf/2608.17099","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A4","B4"],"tags":["合成代理","参与式设计","代表性伦理"],"reason":"批判性审视合成代理替代人类参与的代表性问题，提出监督边界，方法论可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":7,"question":"合成代理在代表性过程中如何制造出合法参与的假象，以及应如何设定设计与部署的边界？","design":"本文不是仿真实验研究，而是对三个合成代理案例（地方政策咨询聊天机器人Ana、企业陪审团审议工具Synthetic Juror、全球外交AI化身Ask Amina和Ask Abdalla）进行比较案例分析，运用参与式设计的四种模式（探查、启动、理解、生成）剖析其如何制造人格假象。","baseline":"无对照","findings":"合成代理通过问题框架、数据策展、用户交互设计和有效性评估四个步骤制造出“人造人格”，绕过代表性过程从而获得表面合法性。现有评估框架（如可信度、保真度、算法保真度）只衡量输出相似性，无法检测这种对过程完整性的绕过。","reliability":"论文未讨论","relevance":"该文批判性审视合成代理在代表性过程中的合法性与人格问题，提出软硬边界，对关注LLM仿真可靠性及伦理边界的研究者具有重要参考价值。","inspiration":"本文的参与式设计视角和过程导向批判方法值得借鉴，可迁移到经济金融领域中涉及代表性决策的场景（如政策咨询、消费者意见征询、董事会决策模拟等）。｜可设计一项研究，用LLM合成代理模拟消费者或投资者参与政策咨询或产品设计讨论，处理为不同的人格制造步骤（如改变数据策展或交互设计），结果变量为参与者对过程合法性的感知或决策质量，并与真实人类参与者的数据对照。"}},{"id":"2608.17120","version":1,"title":"Children, but not language models, show accelerating returns in word learning","zh_title":"儿童而非语言模型在词汇学习中表现出加速回报","abstract":"Children learn hundreds of words over the first years of their lives, in a process that begins slowly but quickly picks up speed. Prior models describe vocabulary growth as evidence accumulation over time. Here we show that the process is best characterized as accelerating accumulation: children learn more from each additional unit of linguistic experience than they did from the one before. In contrast to children, language models -- even those trained on child-directed speech -- do not accelerate. Instead, they show constant proportional returns on new data, consistent with scaling laws. Children learn using many orders of magnitude less training data than language models; their increasingly efficient use of their learning input is a candidate explanation.","authors":["Michael C. Frank"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17120","pdf_url":"https://arxiv.org/pdf/2608.17120","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["语言模型评估","人类学习对照","算法保真度"],"reason":"对比儿童与语言模型的学习效率，评估模型作为人类学习代理的可靠性，有真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":12,"question":"儿童词汇学习是否表现出加速回报，而语言模型是否缺乏这种加速？","design":"本研究并非用LLM模拟人类被试，而是直接比较儿童与语言模型的学习效率。儿童数据来自多语言CDI纵向词汇量表；语言模型包括在儿童导向语料上训练的GPT-2-small、BabyLM和ClimbMix模型。通过贝叶斯模型拟合词汇增长曲线，估计加速参数，并计算模型在训练数据上的边际学习效率。","baseline":"真实人类数据：来自Wordbank的多语言CDI纵向数据（英语、挪威语、日语），包含数千名儿童的词汇发展轨迹。","findings":"儿童词汇学习表现出加速回报，即随着经验增加，每单位语言输入带来的词汇增长更多；而语言模型即使训练于儿童导向语料，也仅表现出恒定的比例回报，符合缩放定律。儿童的学习效率远高于语言模型，可能源于其发展性变化和“学会学习”能力。","reliability":"论文承认其结论基于统计拟合而非因果操纵，且CDI数据可能受整体语言产生能力变化影响；语言模型与儿童之间缺乏明确的“链接假设”，难以直接对齐比较。","relevance":"该研究直接对比儿童与语言模型的学习效率，评估模型作为人类学习代理的可靠性，并指出模型在模拟人类发展性学习时的根本差异，对关注LLM仿真人类认知与行为的研究者具有重要参考价值。","inspiration":"借鉴其通过贝叶斯模型拟合学习曲线并估计加速参数的方法，可量化学习效率的动态变化｜可迁移到经济金融领域中关于经验积累与决策效率的问题，如投资者从市场反馈中学习、消费者从价格信息中学习等｜设计一个实验：以LLM模拟投资者，给予不同量的历史价格数据，测量其预测准确率的边际提升，并与真实投资者交易数据（如个人投资者账户记录）对比，检验LLM是否也表现出恒定回报而非加速学习。"}},{"id":"2608.17810","version":1,"title":"Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses","zh_title":"可解释的人类，异质的LLM：评估响应中潜在结构的专家分析","abstract":"The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpretable meaning as the cognitive constructs governing human learners. Using responses from humans and six LLMs across quantitative reasoning and chemistry assessments, we conducted Exploratory Factor Analysis (EFA) separately for both groups. Subject-Matter Experts (SMEs) then blindly evaluated the resulting factor graphs to ascribe pedagogical meaning to the emerged constructs. SMEs successfully interpreted most of the human-derived factors. Conversely, they could not ascribe meaning to any LLM-derived factors in quantitative reasoning and interpreted only half of the LLM factors in chemistry. By combining data-driven EFA with blind expert interpretation, this framework shows that LLMs frequently operate on statistically opaque mechanisms distinct from human reasoning.","authors":["Alona Strugatski","Licol Zeinfeld","Jason Cooper","Shelley Rap","Gil Schwarts","Giora Alexandron"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17810","pdf_url":"https://arxiv.org/pdf/2608.17810","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM认知结构","因子分析","仿真效度"],"reason":"评估LLM与人类认知结构差异，有真实人类数据对照，批判性指出LLM机制不透明，…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":14,"question":"学科专家如何解读从人类和LLM评估反应中提取的潜在因子？","design":"本研究使用两个人类设计的评估工具（化学和定量推理），分别收集人类学生和六个LLM版本的反应数据，对每个组分别进行探索性因子分析（EFA），然后让学科专家在不知情的情况下盲评因子载荷图，尝试赋予教学意义。","baseline":"人类基准是来自真实教育环境的学生反应数据：化学诊断评估有931名高中生，定量推理部分有979名大学入学考生。","findings":"学科专家能够解释人类学习者的大多数因子，但无法解释LLM在定量推理中的任何因子，在化学中也只能解释一半的LLM因子。这表明LLM的潜在结构在统计上不透明，与人类推理机制不同。","reliability":"论文未讨论","relevance":"该研究直接对比了LLM与人类在评估反应中的潜在认知结构，并批判性地指出LLM的因子不可解释，符合研究者对仿真可靠性与偏差的关注，值得阅读原文以了解其方法细节和局限性。","inspiration":"借鉴其盲评专家解读潜在因子的方法，用于检验LLM在经济决策任务中是否形成与人类相似的行为因子结构。｜可迁移到消费者跨期选择或风险偏好实验，检验LLM的决策模式是否与人类被试的潜在心理构念一致。｜以LLM和人类被试分别完成跨期选择任务，对选择数据进行EFA，让经济学家盲评因子，并与真实实验数据（如Andreoni和Sprenger的凸时间偏好实验）对照，看LLM的因子是否可解释为时间贴现和效用曲率等标准构念。"}},{"id":"2608.16909","version":1,"title":"When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice","zh_title":"当个性化成为偏见：AI生成金融建议中的结构与话语宗教框架","abstract":"Large language models (LLMs) are increasingly integrated into financial advisory systems, yet their role in reproducing religious bias remains underexamined. This study provides systematic mixed-methods evidence of such bias across three LLMs (ChatGPT, Gemini, and Grok) using 432 simulated advisor-client interactions spanning 16 religious identity pairings (Christian, Muslim, Hindu, and non-religious) and three core household financial decisions: stock investment, house purchase, and life insurance. Combining regression and reflexive thematic analyses, we identify structural biases across models and decision contexts and the discursive mechanisms through which they are linguistically enacted. Unbiased advice appeared in only 12-18% of cases. Gemini consistently produced more bias than Grok, while ChatGPT's outputs were statistically comparable to Grok's. Religiously symmetric advisor-client pairings almost always triggered explicit religious framing, and non-religious clients often received advisor-centered religious appeals. Qualitative findings show that bias is linguistically manifested through religious anchoring, uneven cultural signaling, and tone modulation, varying by model and financial scenario. Stock investment prompts produced more financially technical responses, whereas life insurance advice triggered stronger religious language. The study develops a dual-dimensional framework linking structural bias rooted in model training and design with discursive bias expressed through language, advancing understanding of algorithmic bias in LLM-generated financial advice. It also shows that such advice adapts linguistically to identity cues, revealing a managerial dilemma between personalization and neutrality. Finally, it highlights implications for businesses, financial institutions, and regulators seeking to ensure neutrality, cultural sensitivity, and trust in AI-mediated advice.","authors":["Muhammad Salar Khan","Hamza Umer","Hasan Mahmud","Sandra Rothenberg"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16909","pdf_url":"https://arxiv.org/pdf/2608.16909","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","算法偏见","金融咨询"],"reason":"用LLM模拟金融咨询中的客户互动，有真实人类数据对照，并批判性分析偏差，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":11,"question":"LLM在提供金融建议时是否会因客户宗教身份而产生结构性偏差和话语性偏差，以及这种偏差如何通过语言表现出来？","design":"使用ChatGPT、Gemini和Grok三个LLM模拟金融顾问，与16种宗教身份配对（基督教、穆斯林、印度教、无宗教）的客户进行432次互动，覆盖股票投资、购房和人寿保险三种决策，通过回归和反思性主题分析测量建议中的宗教框架和偏差。","baseline":"无对照","findings":"无偏建议仅占12-18%，Gemini偏差显著高于Grok，ChatGPT与Grok无显著差异；宗教对称的顾问-客户配对几乎总是触发显性宗教框架，偏差通过宗教锚定、文化信号和语气调节等话语机制表现。","reliability":"论文未讨论","relevance":"该研究用LLM模拟金融咨询中的客户互动，系统检验宗教身份对AI建议的影响，并批判性分析偏差机制，与研究者关注的人类仿真实验和偏差评估高度相关，值得精读原文。","inspiration":"借鉴其通过系统操纵身份配对和决策场景来测量LLM输出偏差的实验设计，以及结合定量回归与定性主题分析的方法。｜可迁移到信贷审批中的宗教或种族歧视、投资建议中的文化偏差、保险定价中的身份敏感等经济金融场景。｜以LLM作为信贷员，随机分配申请人宗教身份（如穆斯林、基督教、无宗教），要求其给出贷款额度和利率建议，结果变量为建议的金额和利率，对照真实银行信贷数据中不同宗教群体的实际获批差异。"}},{"id":"2608.17644","version":1,"title":"LLM-Derived Preference Judgments Are Not Self-Consistent","zh_title":"LLM衍生的偏好判断并非自洽","abstract":"Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments and then chooses actions based on their estimated utility. This pipeline assumes the judgments are approximately self-consistent: that a single utility function can reproduce them. But are they? To study this question, we measure the self-consistency of cardinal LLM preference judgments. For example, the difference in stated willingness-to-pay between two items should match the stated payment that makes a person indifferent to exchanging them. We develop statistical tests and interpretable measures of how far observed responses depart from the best-fitting self-consistent utility function. Experiments with flight, apartment, and hotel examples across six LLMs reveal large persistent inconsistencies. This suggests that LLM-derived preference judgments cannot be faithfully summarized by a single utility function.","authors":["Matthew T. Ford","Francis Bahk","Jingjing Wang","Adam S. Jovine","Tinghan Ye","David B. Shmoys","Peter I. Frazier"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17644","pdf_url":"https://arxiv.org/pdf/2608.17644","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["偏好判断","自洽性","效用函数"],"reason":"评估LLM偏好判断的自洽性，揭示其不能由单一效用函数概括，对仿真可靠性有批判性…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":13,"question":"LLM 生成的数值偏好判断是否自洽，即能否由单一效用函数概括？","design":"该研究并非人类仿真实验，而是对 LLM 作为偏好判断代理的测量审计。研究者构造航班、公寓、酒店三个领域的物品集，向六种 LLM 提供人类偏好描述，然后通过两种查询方式获取数值偏好判断：物品查询（询问最高支付意愿）和报价对查询（询问使人在两个报价间无差异的价格变化）。基于准线性效用假设，检验这些判断是否满足自洽性，并开发统计检验和可解释度量（RMSE、价格归一化残差、偏好反转）来量化偏离程度。","baseline":"无对照","findings":"所有模型在 Bonferroni 校正后均拒绝所有查询组自洽的联合假设，其中五个模型拒绝全部九个审计组，Qwen 拒绝四个。不一致性主要体现在查询类型之间，且效应量（以美元和相对价格衡量）较大，表明 LLM 派生的偏好判断不能由单一效用函数忠实概括。","reliability":"论文指出，审计仅针对有限物品集、特定查询和提示语义，即使未拒绝自洽性也不代表能泛化到未见物品或提示；同时，审计衡量的是内部一致性而非对人类偏好的准确性，拒绝自洽性并不直接证明对人类保真度低。","relevance":"该研究直接评估 LLM 作为人类被试替代品时的测量可靠性，揭示了偏好判断在跨查询格式下的系统性不一致，对依赖 LLM 生成效用函数或支付意愿的经济学实验和政策评估构成重要警示，值得精读。","inspiration":"借鉴其审计协议：通过设计多种查询格式（直接支付意愿与间接补偿金额）并检验跨格式一致性，可系统评估 LLM 生成的偏好数据的内部有效性。｜可迁移到消费者选择实验，如用 LLM 模拟消费者对新产品属性的支付意愿，检验不同询问方式（直接定价、选择实验、匹配任务）得出的效用参数是否一致。｜设计：以 LLM 模拟消费者，提供产品描述和偏好背景，分别用直接询问最高支付意愿、二元选择（不同价格下的购买决策）和匹配任务（等价补偿）获取数据，拟合离散选择模型估计效用参数，并与真实消费者调查或实验数据（如 Nielsen 面板或实验室拍卖）对比，检验参数一致性和预测效度。"}},{"id":"2608.18058","version":1,"title":"Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating","zh_title":"代理推荐系统中的委托不对称：测量在线约会中的双向接受度","abstract":"Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet their viability depends on a condition rarely examined: users must accept not only delegating conversation to an agent, but also receiving agent-mediated communication from others. We study this condition using two large-scale surveys of active users of a major dating platform (N=2,894 on generative profile features; N=2,617 on autonomous conversational agents, fielded in two languages). We develop a latent-variable measurement model of agent receptivity based on graded response models with latent regression, and show via model comparison that willingness to send and willingness to receive agent communication are distinct constructs: highly correlated (rho=0.92) but separable (Delta BIC=52), with partial measurement invariance across languages. The model quantifies a systematic delegation asymmetry: deploying one's own agent requires far lower receptivity (threshold -0.38) than engaging a counterpart's agent (+0.32; full engagement +1.39), and mean deployment propensity exceeds engagement propensity roughly threefold. Under a random-pairing counterfactual derived from stated receptivity, only 4-13% of directed dyads combine agent deployment with receiver engagement, with a pronounced gender-directional imbalance. Design counterfactuals quantify the levers: a reciprocity requirement cuts interaction volume by half or more by excluding nearly two-thirds of would-be deployment, while routing agent contacts on receive receptivity triples per-contact engagement, a lift that survives out-of-sample validation with the target item held out (AUC 0.88, 3.1x quartile lift under respondent-level cross-validation). We discuss implications for agentic recommender design, including disclosure, opt-in mechanics, and receptivity-aware matchmaking.","authors":["Daria Leshchikova","Valentina V. Kuskova","Dmitry Zaytsev","Valerii Klimov"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18058","pdf_url":"https://arxiv.org/pdf/2608.18058","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM代理","用户调查","推荐系统"],"reason":"用LLM代理模拟用户交流，有真实用户调查数据对照，涉及推荐系统设计，可迁移到人…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":15,"question":"在匹配平台中，用户是否既愿意委托自己的AI代理进行交流，又愿意接收他人AI代理发来的信息？","design":"本研究不是仿真研究，而是基于真实用户调查的测量研究。使用两个大规模调查（N=2,894和N=2,617），对某大型约会平台活跃用户进行问卷，测量对生成式个人资料特征和自主对话代理的态度。构建基于等级反应模型和潜在回归的潜变量测量模型，将七个态度项目加载到发送意愿和接收意愿两个维度上，并估计阈值和倾向。","baseline":"无对照（非仿真研究，无人类基准对照）","findings":"发送意愿和接收意愿是两个高度相关但可分离的构念（ρ=0.92，ΔBIC=52），存在系统性委托不对称：部署自己的代理所需接受度阈值（-0.38）远低于与对方代理互动（+0.32），部署倾向约为互动倾向的三倍。在随机配对反事实中，只有4-13%的有向配对同时包含部署和接收，且存在性别方向失衡。","reliability":"论文未讨论（节选部分未提及失效条件或局限）","relevance":"该研究虽非LLM仿真，但提供了真实用户对AI代理接受度的测量方法和不对称性发现，可作为评估LLM仿真人类行为可靠性的基准数据，尤其适用于涉及双向互动和信任的场景。","inspiration":"借鉴其潜变量测量模型和反事实设计，将发送与接收意愿作为分离构念进行联合测量，并利用阈值差异量化不对称性。｜可迁移到经济金融中的双边市场或信任场景，如P2P借贷中出借人与借款人对AI代理的使用意愿、或金融咨询中客户与顾问的AI接受度。｜以P2P借贷平台用户为被试，施加AI代理沟通处理（如自动生成借款请求或出借决策），测量出借意愿和借款意愿，并与平台真实交易数据对照，检验不对称性对市场成交量的影响。"}},{"id":"2608.14606","version":1,"title":"Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents","zh_title":"看似合理但无效：对LLM作为合成调查受访者的心理测量审计","abstract":"Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM \"crowd\" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.","authors":["Mantas Lukauskas","Viktorija \\v{S}arkauskait\\.e"],"categories":["cs.CY","cs.AI","cs.CL","stat.AP"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14606","pdf_url":"https://arxiv.org/pdf/2608.14606","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","心理测量效度","调查数据"],"reason":"直接评估LLM作为调查受访者的心理测量效度，并与真实人类数据对照，批判性指出失…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":2,"question":"LLM 作为合成调查受访者时，是否能在心理测量学层面（联合分布、潜结构、信度、中介路径、人口学效应）复现真实人类调查数据的特征？","design":"使用 37 个 LLM（涵盖 OpenAI、Anthropic、Google 及 12 个开源家族）基于真实受访者档案生成对 68 个题项（3 个量表）的回答，通过五级人格披露阶梯、呈现方式与推理努力消融、反事实人口学变换（性别、角色、教育）、跨语言检查和逐字回忆探针等处理，测量心理测量相似性得分（PSS）及其各维度。","baseline":"立陶宛组织心理学数据集（n=263 名员工；Dunham 变革态度量表、UWES-17、Koopmans IWPQ；68 题，12 个分量表），以及五个非 LLM 统计基线和留出的人类对比上限。","findings":"LLM 能复现人类心理测量关系的定性方向，但高斯 copula 基线在样本驱动的 PSS 分量上击败所有 LLM；LLM 群体内部相似度（平均 PSS 0.73）高于与人类的相似度，且逐字回忆探针表明记忆并非排行榜驱动因素（秩相关 0.00）。","reliability":"论文指出 LLM 样本不能直接替代人类调查数据：合成样本存在默认偏差（+0.84 SD），在留出人类数据上预测效度丧失（平均 R² -0.18 vs 0.28），并在 10 条安慰剂中介路径中捏造了 3 条显著间接效应；此外，教育驱动的反事实效应（平均 |d|=0.56）远大于性别（0.12）和角色（0.18），且 UWES 的 Tucker's phi 在 37 个模型中有 8 个落入置换零分布内。","relevance":"该研究直接评估 LLM 作为人类被试替代品的心理测量效度，并与真实人类数据严格对照，批判性地揭示了仿真在联合分布、预测效度和中介推断上的失效条件，对关注 LLM 仿真可靠性与偏差的研究者极具参考价值。","inspiration":"借鉴其多维度心理测量审计框架（联合分布、潜结构、信度、中介路径、人口学效应）和反事实人口学变换设计，系统评估合成样本的效度｜可迁移到经济金融中的调查实验，如消费者信心、通胀预期、风险偏好或政策支持度等场景，检验 LLM 能否复现真实人群的分布与结构｜以真实家庭金融调查（如美国 SCF 或中国 CHFS）为基准，用 LLM 基于受访者人口学特征生成对风险态度、时间偏好等量表的回答，施加收入或教育水平的反事实变换，比较 LLM 样本与人类样本在联合分布、因子结构和中介效应上的差异。"}},{"id":"2608.15871","version":1,"title":"Large Language Models as Implicit Sociological Models: Reconstructing Voting Behaviour from Sociodemographic Profiles","zh_title":"大语言模型作为隐式社会学模型：从社会人口特征重建投票行为","abstract":"Large language models (LLMs) trained on large-scale internet corpora encode extensive statistical regularities about social identities, attitudes, and political behaviour. This paper introduces and evaluates a methodological framework that leverages these latent representations to reconstruct aggregate voting behaviour from individual-level sociodemographic profiles. We operationalize LLMs as implicit sociological models by conditioning them on demographic descriptions, eliciting probabilistic turnout and party preferences, and aggregating individual outputs via a soft voting procedure. Using the 2021 Czech parliamentary election as a validation case, we demonstrate that contemporary LLMs reproduce official election outcomes with low mean absolute error, recover known political bloc structures, and align with independently established sociodemographic gradients. The contribution of this work is methodological rather than predictive: we show how LLMs can be systematically interrogated as compressed representations of social reality, offering a novel exploratory instrument for computational social science while clearly delineating its epistemic and ethical limits.","authors":["Roman Neruda","Martin Bako\\v{s}","Josef \\v{S}lerka","V\\'it Tu\\v{c}ek","Petra Vidnerov\\'a","Gabriela Kadlecov\\'a"],"categories":["cs.CY","cs.CL","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15871","pdf_url":"https://arxiv.org/pdf/2608.15871","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","投票行为","计算社会科学"],"reason":"用LLM从人口特征重建投票行为，并与真实选举结果对照，直接仿真人类决策。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":1,"question":"大语言模型能否仅从个体社会人口学特征重建总体投票行为，并在捷克2021年议会选举这一非英语多党制情境下得到验证？","design":"将LLM作为隐式社会学模型，以10个社会人口学变量（含主观生活水平和政治兴趣）为条件，对代表性调查中的个体生成概率性投票选择（投票率和政党偏好），通过软投票聚合得到模拟选举结果。","baseline":"官方选举结果（捷克统计局）和调查中的自报投票选择（声称投票）。","findings":"当代LLM能以较低平均绝对误差重现官方选举结果，恢复已知政治阵营结构，并与独立确立的社会人口学梯度一致。该贡献是方法论的，而非预测性的。","reliability":"论文承认个体层面预测噪声大且存在系统性偏差，但通过软投票聚合可部分抵消；捷克语为中等资源语言，模型训练数据以英语为主，可能影响表现；未明确讨论其他失效条件。","relevance":"该研究直接命中你的核心关注：用LLM仿真人类决策并与真实选举数据对照，且提供了多层级验证（总体份额、协方差结构、与民调机构及CHES专家编码比较），值得精读原文以了解其方法细节和局限。","inspiration":"借鉴其软投票聚合和三角验证设计，将个体噪声转化为总体稳健估计，并同时对照官方数据和自报数据以识别偏差。｜可迁移到政策公告的预期形成研究，例如模拟不同人口群体对财政或货币政策变化的反应。｜以代表性家庭调查中的个体为被试，用LLM基于人口特征生成对政策变化的预期（如通胀预期），处理为不同政策情景描述，结果变量为预期值或不确定性，对照真实调查中的预期数据和后续实际经济行为。"}},{"id":"2608.14630","version":1,"title":"Characterizing Rhetorical Misalignment in Decision-Making with Language Models","zh_title":"表征语言模型决策中的修辞错位","abstract":"Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly integrated into high-stakes human-AI decision-making, it is important to understand whether their outputs can amplify potential biases, how this influences human decisions, and crucially, whether it can lead to harmful consequences. In this work, we develop a decision-theoretic framework to study rhetorical misalignment, a failure mode where an LLM uses rhetorically inappropriate forms of presentation for a given decision context, thereby inducing suboptimal human decisions. We empirically investigate this phenomenon through a human-subject experiment in realistic clinical decision-making using a dataset curated from the United States Medical Licensing Examination. By measuring how LLM-generated information affects decisions, we observe that LLMs induce an average 2.81% rate of harmful decision flips across different models, where clinician participants change from a correct to an incorrect answer. Rationales reported by participants provide evidence that these revisions are closely related to the language used by LLMs that may induce different types of cognitive biases, including anchoring, authority bias, and loss aversion. To enable scalable evaluation, we instantiate our theoretical framework using decision-makers simulated by LLMs to computationally measure rhetorical misalignment. Our findings reveal a safety concern previously unrecognized in high-stakes domains: a model can be factually aligned yet still induce harm through its rhetorical presentation.","authors":["Zirui Cheng","Joey Chan","Simo Du","Chenhao Tan","Yue Guo","Hao Peng"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14630","pdf_url":"https://arxiv.org/pdf/2608.14630","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类决策","认知偏差"],"reason":"用LLM模拟人类决策者评估修辞偏差，并与真实人类实验对照，涉及临床决策场景，批…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":7,"question":"LLM 的修辞呈现方式是否会在事实信息一致的情况下诱导人类做出次优决策，即产生“修辞错位”现象？","design":"论文采用人类受试者实验与LLM模拟决策者相结合的方法。人类实验中，临床医生参与者基于USMLE题目作答，部分题目提供LLM生成的分析（固定信息或自然生成），测量决策变化率及有害翻转率；模拟实验中，用LLM分别扮演理性决策者和行为决策者，比较两者在相同信息不同措辞下的决策差异。","baseline":"人类受试者实验数据：临床医生在没有LLM辅助时的原始答案作为对照，测量LLM辅助后的决策变化。","findings":"人类实验中，LLM辅助导致平均27.58%的决策变化，其中2.81%为有害翻转（从正确变为错误）；参与者报告显示这些变化与锚定、权威偏误、损失厌恶等认知偏误相关。模拟实验表明，即使信息相同，仅语言措辞差异也能导致理性与行为决策者之间的分歧，且自然生成设置下分歧更大。","reliability":"论文承认其人类实验仅限于USMLE临床决策场景，未测量下游后果（如患者结局、经济成本），且有害翻转率绝对值较小；模拟实验的LLM决策者可能无法完全代表人类认知偏误。","relevance":"该研究直接使用LLM模拟人类决策者来测量修辞错位，并与真实人类实验对照，属于用LLM进行人类仿真实验的典型工作，且涉及高 stakes 临床决策场景，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其“固定信息、变化措辞”的处理设计，可分离信息内容与修辞框架的效应，并利用LLM模拟理性与行为决策者进行大规模测量｜可迁移到经济金融中的政策公告解读、投资建议呈现、信贷合同条款表述等场景，研究措辞对个体决策的影响｜设计实验：以真实投资者或消费者为被试，呈现同一金融产品的两种措辞（如“年化收益5%” vs “亏损概率2%”），测量选择差异，并用历史交易数据或调查数据作为真实行为基准，同时用LLM模拟投资者进行平行实验以验证仿真效度。"}},{"id":"2510.10813","version":2,"title":"The Fragility of Strategic Thinking in Large Language Models","zh_title":"大语言模型中战略思维的脆弱性","abstract":"Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation, policy design, and market simulation. However, can we trust LLMs to think strategically in complex situations? Existing research mostly evaluates LLMs' adherence to equilibrium play or their exhibited depth of reasoning, leaving open whether they display strategic thinking meant as the ability to form coherent conjectures about other agents, to evaluate possible actions conditional on those conjectures, and to best respond to them. We develop a framework to identify this ability by disentangling belief formation, evaluation, and choice in static complete-information games across a series of non-cooperative environments. By jointly analyzing models' revealed choices and reasoning traces, and introducing a new context-free game to rule out imitation from memorization, we show that strategic thinking in current frontier LLMs is real but fragile: models execute best responses to exogenous conjectures and form opponent-contingent conjectures when left unconstrained. Yet under increasing complexity explicit recursion gives way to model-specific logic shifts and heuristic rules of choice, both within and outside equilibrium reasoning. Further, these heuristics do not map directly onto the systematic biases typically observed in human strategic behavior. These findings, already emerging in noiseless settings, warrant caution in the application of LLMs as strategic agents in complex environments.","authors":["Enric Junque de Fortuny","Veronica Roberta Cappelli"],"categories":["cs.AI","cs.GT"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-18","first_seen":"2025-10-12","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2510.10813","pdf_url":"https://arxiv.org/pdf/2510.10813","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B2","B4"],"tags":["LLM战略推理","博弈论","行为对照"],"reason":"评估LLM在博弈中的战略思维，与人类行为对照，但非直接仿真人类被试，结论可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":8,"question":"大语言模型在静态完全信息博弈中是否具备真正的战略思维能力，即能否形成关于对手行为的连贯信念、基于信念评估行动并做出最优反应？","design":"该研究并非用LLM仿真人类被试，而是直接测试多个前沿LLM在静态完全信息博弈中的战略思维。研究者设计了一系列非合作博弈环境，通过分解信念形成、评估和选择三个环节，结合模型的选择和推理痕迹，并引入无上下文博弈以排除记忆模仿。","baseline":"无对照","findings":"前沿LLM的战略思维真实但脆弱：它们能对外生信念做出最优反应，并在无约束时形成依赖对手的信念；但在复杂度增加时，显式递归让位于模型特定的逻辑转换和启发式选择规则，且这些启发式与人类系统性偏差不直接对应。","reliability":"论文指出，在无噪声环境中已出现脆弱性，因此对LLM在复杂环境中作为战略代理的应用需谨慎；启发式规则与人类偏差不匹配，可能限制其在人类行为仿真中的有效性。","relevance":"该研究评估LLM在博弈中的战略思维，虽非直接仿真人类被试，但其结论对使用LLM进行人类决策仿真具有警示意义，值得阅读原文以了解LLM在战略推理中的局限。","inspiration":"借鉴其分解信念、评估和选择的框架，以及通过推理痕迹和上下文无关博弈排除记忆混淆的方法，可用于严格测试LLM的决策过程。｜可迁移到经济金融中的策略互动场景，如拍卖竞价、谈判或市场进入博弈，检验LLM是否形成理性预期。｜设计一个古诺竞争实验，让LLM扮演企业，处理为不同信息结构（如对手成本已知或未知），结果变量为产量选择，并与人类实验数据（如Huck等2000年的古诺实验）对照，分析LLM的信念形成和最优反应。"}},{"id":"2608.00410","version":2,"title":"Where did the ambiguity go? Examining how multimodal models interpret polysemous words","zh_title":"歧义去哪了？考察多模态模型如何解释多义词","abstract":"Human language is highly polysemous. Many common words (e.g., \"bank\" or \"palm\") carry several distinct meanings that shape what humans communicate and imagine. Large language models (LLMs) have been shown to understand this multiplicity of meaning, but much less is known about how polysemy surfaces in other modalities such as images. We study this across 17 text-to-image and 15 text-generation models by giving each a polysemous word with no context to fix its meaning and measuring which senses are produced over many samples. We find a clear multimodal gap, where within every model family, generated images settle on far fewer senses than generated sentences (normalized entropy 0.10 vs. 0.25), and both are far less varied than what people imagine for the same words (normalized entropy 0.47). However, when we instead ask a model to list how often it would generate outputs corresponding to each possible meaning of a word, it predicts distributions that are more diverse than the actual space of outputs. These results reveal a multimodal gap in how foundation models express meaning, and how their understanding may not transfer faithfully nor equally across modalities.","authors":["Jasin Cekinmez","Addison J. Wu","Raja Marjieh","Thomas L. Griffiths"],"categories":["cs.AI","cs.CL","cs.CV"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-18","first_seen":"2026-08-04","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2608.00410","pdf_url":"https://arxiv.org/pdf/2608.00410","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["多模态模型","语义歧义","人类对照"],"reason":"用LLM生成文本和图像，与人类想象对照，评估多模态意义表达差异，属于仿真人类认…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":9,"question":"多模态基础模型在无上下文的多义词上，其图像生成与文本生成所表达的语义分布有何差异，并与人类想象分布相比如何？","design":"本研究并非以LLM模拟人类被试，而是直接考察模型行为：以100个多义词为刺激，在无上下文条件下分别输入17个文本到图像模型和15个文本生成模型，每个词采样30次，用GPT 5.4作为裁判将输出分类到预定义的词义，得到每个模型-词对的词义分布。","baseline":"通过Prolific收集人类被试在相同100个词上的反应，采用与模型条件匹配的两种框架（“用这个词造句”对应文本模型，“这个词让你想到什么图像”对应图像模型），同样由GPT 5.4裁判分类。","findings":"图像模型生成的词义分布比文本模型更集中（归一化熵0.10 vs 0.25），且两者都远低于人类想象的多样性（归一化熵0.47）。当要求模型预测各词义的生成频率时，其预测分布比实际输出更多样，表明模型知道歧义但生成时未充分表达。","reliability":"论文未明确讨论仿真失效条件，但指出偏好调优会进一步收窄生成分布，且对齐不能完全解释多模态差距；裁判可靠性通过作者盲标60张图像与裁判完全一致来验证。","relevance":"该研究用真实人类数据作为基准，比较LLM多模态输出与人类想象分布，揭示模型在表达语义多样性上的系统性偏差，对关注LLM仿真人类认知与行为可靠性的研究者有直接参考价值。","inspiration":"借鉴其无上下文刺激与分布比较方法，可设计实验考察LLM在经济概念上的先验分布与人类直觉的差异｜可迁移到政策公告的预期形成研究，如“通货膨胀”一词的多重解读如何影响模型生成的预期分布｜以LLM为被试，呈现无上下文的“利率”等经济术语，要求生成解释或图像，测量其词义分布，并与调查数据中公众对同一术语的理解分布进行对照。"}},{"id":"2608.12253","version":2,"title":"One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL","zh_title":"单一冻结模拟器不够：多智能体强化学习中的模拟器坍缩","abstract":"Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theoretically and propose two complementary solutions, one at inference time and one at training time. The inference-time solution, Verbalized Sampling, broadens the simulator's behavior by sampling from a verbalized response distribution, reducing mode collapse. The training-time solution, Co-Training, jointly optimizes the policy against a population of trainable simulators, preventing it from overfitting to any single simulator's mode. We validate both solutions on three multi-turn benchmarks: Persuasion for Good, $\\tau^2$-bench, and CooperBench. Verbalized Sampling improves held-out success by up to 9% over single-simulator RL, and Co-Training pushes gains further to 14%; the human study shows similar gain on real users. Both solutions preserve the policy diversity that collapses under single-simulator RL. To support further work in this direction, we release SCOPE, an open-source framework for Population Co-Training multi-agent RL. More broadly, our results suggest that the diversity of the training environment, not only the policy, is critical to the generalization of multi-turn RL to real-world deployment.","authors":["Simon Yu","Nicholas Tomlin","Marwa Abdulhai","Ximing Lu","Derek Chong","Abe Hou","Dilara Soylu","Sergey Levine","Christopher D. Manning","Weiyan Shi"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"replace-cross","date":"2026-08-18","first_seen":"2026-08-13","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2608.12253","pdf_url":"https://arxiv.org/pdf/2608.12253","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","多智能体强化学习","人类行为模拟"],"reason":"用LLM模拟用户行为训练策略，并与真实用户对照，指出单一模拟器失效问题，方法可…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":10,"question":"在人类-AI多轮交互的强化学习训练中，使用单一冻结的LLM模拟用户行为是否会导致策略泛化失败，以及如何解决该问题？","design":"该研究使用LLM模拟用户行为，在三个多轮交互基准（Persuasion for Good、τ²-bench、CooperBench）上训练强化学习策略。处理包括：基线（单一冻结LLM模拟器）、推理时方案（Verbalized Sampling，从模拟器输出的言语化分布中采样）、训练时方案（Co-Training，策略与可训练模拟器联合优化）以及两者结合（Population Co-Training）。结果变量为在留出模拟器和真实用户上的任务成功率及策略熵。","baseline":"在τ²-bench和Persuasion for Good上进行了预先注册的人类研究，以真实用户的任务结果和对话自然度作为对照。","findings":"单一模拟器强化学习导致策略过拟合模拟器的模式，在留出模拟器和真实用户上表现下降，出现“模拟器崩溃”。Verbalized Sampling和Co-Training均能缓解该问题，其中Population Co-Training在留出任务成功率上提升最高（达14%），并在人类研究中表现出类似增益。","reliability":"论文承认模拟器崩溃源于LLM的模式坍缩，且策略可能利用模拟器的特定模式；解决方案的有效性依赖于模拟器多样性的恢复程度。但未系统讨论在更复杂或不同领域用户行为模拟中的失效条件。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，揭示了单一模拟器在强化学习训练中的系统性偏差，并提供了与真实用户对照的实证证据，对关注仿真失效条件的研究者具有重要参考价值。","inspiration":"借鉴其通过多样化模拟器（如采样或联合训练）来防止策略过拟合的方法，可用于经济实验中处理被试异质性。｜可迁移到政策评估中的多轮交互场景，如消费者与AI客服的谈判或信贷审批对话。｜设计一个实验：用多个LLM模拟不同类型的消费者（如风险偏好不同），训练一个AI谈判策略，处理为使用单一模拟器 vs. 模拟器群体，结果变量为谈判达成率和消费者满意度，并与真实人类被试的行为数据对照。"}},{"id":"2608.15634","version":1,"title":"Argumentation for Common Ground: Finding Zones of Possible Agreement between Individuals in Conflict","zh_title":"共同基础的论证：寻找冲突个体间的可能协议区","abstract":"How can common ground between societies in conflict be identified when citizens' acceptability of peace agreements is shaped by contested narratives? Such acceptability is mediated not only by the clauses that agreements include or exclude, but crucially by citizens' subjective reasoning concerning agreements' clauses. In this paper, we leverage computational argumentation to introduce a novel approach to identifying mutually acceptable agreements among individuals in conflict, i.e. a Zone of Possible Agreement (ZOPA). First, we introduce a quantitative bipolar argumentation framework tailored to represent each side's reasoning about peace agreements. We then show how merging these frameworks can enable negotiators to identify peace agreements that are mutually acceptable. To evaluate our approach under conditions of real-world relevance, we focus on the Palestinian-Israeli conflict, where long-standing policy, practitioner and public interest underscores the demand for methods capable of analysing polarised public reasoning. We show how our framework identifies a ZOPA through theoretical analysis and preliminary experiments using survey data from both existing work and retrieved by a large language model. The results illustrate how argumentation can empower negotiators and conflict-resolution teams in mapping feasible ZOPAs grounded in citizens' reasoning.","authors":["Elisa Cavatorta","Antonio Rago"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15634","pdf_url":"https://arxiv.org/pdf/2608.15634","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","冲突解决","人类数据对照"],"reason":"用LLM检索调查数据模拟冲突双方推理，并与真实调查数据对照，属于社会政治过程仿…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":17,"question":"如何利用计算论证方法识别冲突社会中个体之间可共同接受的和平协议区域（ZOPA）？","design":"该研究提出一种定量双极论证框架（QBAF）来表示个体对和平协议条款的推理，通过合并对立双方的QBAF来识别共同可接受的协议区域。实验部分使用现有调查数据和由大语言模型检索的调查数据，对巴以冲突场景进行初步验证。","baseline":"使用来自现有研究（Golan-Nadir等）的全国代表性调查数据，以及由大语言模型检索的报告数据作为对照。","findings":"理论分析表明，满足平衡性、单调性和对偶性的渐进语义能产生直观的协议排序；初步实验显示该方法与现有调查数据具有合理相关性，表明其适用于现实部署。","reliability":"论文承认当前实验仅为初步验证，未来需要从实际调查受访者中获取推理数据并进行大规模合并；同时指出基础分数获取需要更严谨的协议，可能借鉴行为经济学原理。","relevance":"该研究利用LLM检索调查数据来模拟冲突双方的推理，并与真实调查数据对照，属于社会政治过程仿真，对关注LLM仿真可靠性和偏差的研究者有一定参考价值，但方法核心是计算论证而非LLM仿真，相关度中等。","inspiration":"借鉴其将个体主观推理形式化并合并以识别共识区域的方法，可用于经济政策偏好聚合｜可迁移到政策公告的预期形成或消费者对金融产品的态度分析｜设计：以LLM生成个体对某项经济政策（如税收改革）的论证框架，处理为不同政策条款组合，结果变量为个体接受度，与真实调查数据对照验证。"}},{"id":"2608.14681","version":1,"title":"Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans","zh_title":"自动还是受控？重复启动揭示基础LLM、指令LLM与人类的分歧加工","abstract":"Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evaluate repeated words afresh, and whether post-training changes this default behavior. We apply repetition priming (Shiffrin and Schneider, 1977) to 15 models across five model families (1.5B-14B parameters) in two tasks, semantic categorization and cloze completion, with matched human experiments using identical stimuli. We find that base models exhibit automatic processing: they show immediate facilitation that remains stable across lags, partially survives context removal, and correlates with attention to prior occurrences. Instruct models exhibit controlled processing: their facilitation decays with lag, collapses without expected context, and reverses to interference at larger scales. Within the Qwen 2.5 family, this dissociation increases monotonically with model scale, suggesting that post-training progressively alters repetition processing. Humans show a hybrid profile, with lag-sensitive facilitation resembling instruct models but without interference, suggesting that neither model type fully captures human cognition. Our findings reveal a qualitative shift in how language models process repeated information after post-training and provide mechanistic evidence for the divergence between model behaviors.","authors":["Jinglei Ren","Yuyue Wang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14681","pdf_url":"https://arxiv.org/pdf/2608.14681","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM人类仿真","认知实验对照","模型偏差"],"reason":"用LLM复现人类重复启动效应，并与人类数据对照，揭示模型与人类差异，可迁移到仿…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":14,"question":"语言模型在重复遇到同一词时，是自动激活先前表征还是重新评估，以及指令微调是否改变这一默认处理方式？","design":"用15个模型（5个家族，1.5B–14B参数）在语义分类和完形填空两个任务中，通过操纵重复词之间的间隔（lag）和上下文，测量重复启动效应（log概率边际的变化），并与人类被试在相同刺激和设计下的表现对比。","baseline":"匹配的人类实验，使用相同刺激和设计，测量人类被试的重复启动效应。","findings":"基础模型表现出自动加工：启动效应立即出现且不随间隔衰减，部分在移除上下文后仍存在，并与对先前出现的注意力相关；指令模型表现出控制加工：启动效应随间隔衰减，在移除预期上下文后消失，且在较大规模时反转为干扰。人类表现出混合特征：对间隔敏感（类似指令模型）但始终为促进效应（类似基础模型），从不出现干扰。","reliability":"论文未讨论","relevance":"该研究用LLM复现人类重复启动效应，并与人类数据对照，揭示了模型与人类在重复信息处理上的差异，对评估LLM作为人类被试替代品的可靠性有直接参考价值，值得阅读原文。","inspiration":"借鉴其通过操纵重复间隔和上下文来分离自动与控制加工的实验设计，以及用注意力分析和消融实验提供机制证据的做法。｜可迁移到经济金融中的信息重复场景，如重复广告对消费者决策的影响、重复政策信息对预期形成的作用。｜以LLM模拟消费者，呈现重复的产品信息（如广告语），操纵重复次数和间隔，测量购买意愿或态度变化，并与真实消费者实验数据对照，检验LLM是否复现重复效应。"}},{"id":"2608.15630","version":1,"title":"Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis","zh_title":"评估工具对人类和LLM测量的是同一构念吗？一项潜在结构分析","abstract":"The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use performance on these instruments as evidence for generalizable claims about LLMs' underlying abilities on the same skills the assessments are intended to measure in humans. However, from a validity perspective, such inferences require that the relationship between observed performance and underlying constructs established for humans also holds for LLMs. In particular, a necessary condition for transferring score interpretations is similarity in the latent structure of responses to the assessment. In this study, we examine whether this condition holds in two educational contexts: high-school chemistry and a quantitative reasoning section of a university entrance exam. Using a case study design, we compare human response data with responses generated by six multimodal LLMs. Our analytical approach combines exploratory factor analysis, factor congruence, and resampling to assess latent structure similarity across human learners and LLMs. Across both instruments, we find systematic differences between human and LLM factor structures, showing evidence that the analyzed assessments may not measure the same constructs for humans and LLMs. These findings call into question the validity of evaluation practices that use educational assessments to make claims about AI capabilities.","authors":["Alona Strugatski","Licol Zeinfeld","Giora Alexandron"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15630","pdf_url":"https://arxiv.org/pdf/2608.15630","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM评估效度","潜在结构分析","人类对照"],"reason":"评估LLM在人类测评工具上的效度，有真实人类数据对照，批判性指出结构差异","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":16,"question":"教育测评工具在人类和LLM上是否测量相同的潜在构念？","design":"本研究不是仿真实验，而是比较分析。使用六种多模态LLM（GPT-4o、GPT-5.2、Gemini 1.5 Pro、Gemini 3 Pro、Claude 3.5 Sonnet、Claude 4.5）回答高中化学诊断测试和大学入学定量推理部分，与人类学生作答数据对比，通过探索性因子分析、因子一致性和重抽样比较潜在结构。","baseline":"人类基准为931名高中生的化学测试作答和4800多名考生的定量推理作答数据。","findings":"在两个测评工具上，人类和LLM的因子结构存在系统性差异，表明这些测评可能没有测量相同的构念。这质疑了使用教育测评来推断AI能力的效度。","reliability":"论文承认使用非公开数据集限制了可重复性，但强调这是为了确保人类数据质量。此外，因子结构相似只是构念等价的必要条件而非充分条件，且EFA是初步指标，需要进一步验证。","relevance":"该研究直接批判了用人类测评工具评估LLM能力的效度，提供了真实人类数据对照，并指出潜在结构差异，对关注LLM仿真可靠性和偏差的研究者很有价值。","inspiration":"借鉴其因子分析比较潜在结构的方法，可以检验经济行为测量工具（如风险偏好问卷、时间偏好量表）在人类和LLM上是否测量相同构念。｜可迁移到行为经济学中的偏好测量，例如用LLM模拟消费者跨期选择或风险决策，并检验其潜在结构与人类是否一致。｜以LLM作为被试，让其完成标准风险偏好问卷（如DOSPERT）和跨期选择任务，同时收集人类被试数据，对两组数据进行探索性因子分析并比较因子结构，以评估LLM作为人类被试替代品的效度。"}},{"id":"2608.16514","version":1,"title":"Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans","zh_title":"匹配结果，分歧注视：注视点式多模态大模型与人类搜索的对比","abstract":"Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs), given the same foveated input, search as humans do bears on their use as models of human vision and on attention-alignment scores. We compare three general-purpose MLLMs with human eye-movement scanpaths on goal-directed search (COCO-Search18), driving each model fixation by fixation through an identical, human-matched foveated view and assessing it along three axes: the decision of target presence, the efficiency of reaching the target, and the gaze process itself. The axes dissociate. On the decision and on target acquisition the models match or exceed humans, detecting present targets near ceiling and reaching them on the first saccade more often than people do. The gaze process is not human. Under the human-matched condition, all three share one signature: low-entropy, large-amplitude, self-consistent scanpaths that agree with themselves far more closely than two humans agree with each other. That is consistent with a single-pass, non-serial architecture rather than a limit of acuity. Matched retinal input reproduces where humans look but not how the looking unfolds in time, and no degradation regime recovers human-like search at human-like success. The gap sits on a process axis that answer-alignment and saliency metrics do not measure. Because they miss it, such metrics cannot certify human-like vision, and zero-shot models suit outcome and spatial questions but not temporal, process-level ones.","authors":["Mohamed Amine Kerkouri","Marouane Tliba","Aladine Chetouani","Ulas Bagci","Alessandro Bruno"],"categories":["cs.CV","cs.AI","cs.CL","cs.HC","cs.MM"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16514","pdf_url":"https://arxiv.org/pdf/2608.16514","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","视觉搜索","人类对照"],"reason":"用MLLM模拟人类视觉搜索，与人类眼动数据对照，并批判性指出过程差异，方法可迁…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":20,"question":"在相同的中央凹视觉输入下，多模态大语言模型（MLLMs）的视觉搜索过程是否与人类相似？","design":"用三个通用MLLM（Qwen3.5-35B-A3B、GLM-4.6V-Flash、Gemma-4-E4B）在COCO-Search18数据集上进行目标导向视觉搜索，通过人类匹配的中央凹渲染器逐次注视输入，测量决策（目标有无）、效率（首次扫视命中率、注视次数）和注视过程（扫描路径熵、幅度、自一致性）。","baseline":"COCO-Search18数据集中10名人类观察者在相同场景上的眼动扫描路径和决策。","findings":"在决策和目标获取上，模型达到或超过人类水平（目标存在检测接近天花板，首次扫视命中率0.97/0.97/0.80 vs 人类0.49），但注视过程非人类：模型扫描路径低熵、大幅度、高度自一致（跨种子ScanMatch 0.84/0.91/0.71 vs 人类观察者间0.53），表明模型采用单遍非串行架构而非人类串行搜索。","reliability":"论文指出，答案对齐和显著性指标无法测量过程轴，因此不能认证人类视觉；零样本MLLM适合结果和空间问题，但不适合时间和过程问题。未讨论模型微调或不同提示策略的影响。","relevance":"该研究直接对比MLLM与人类眼动数据，并批判性指出过程差异，对关注LLM仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文以了解过程级评估方法。","inspiration":"借鉴其逐注视点施加相同感知约束（中央凹渲染）并分离结果与过程测量的设计，可迁移到经济决策中的信息搜索过程研究（如消费者浏览商品信息、投资者阅读财报），例如用LLM模拟投资者在财报披露后的信息获取，处理为不同信息呈现方式（如逐段揭示），结果变量为投资决策和注视路径，对照真实眼动实验数据。"}},{"id":"2608.16707","version":1,"title":"Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors","zh_title":"语义老虎机：上下文探索-利用受语义先验影响","abstract":"Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental exploration. However, existing work has raised questions about how LLMs actually balance exploration and exploitation. Unlike classical agents, LLM agents engage with tasks through natural language, exposing them to semantic information with no formal counterpart in the task structure. We introduce the semantic bandit, an extension of the multi-armed bandit setting that explicitly considers the textual labels assigned to actions, and use it to study how semantic priors --- inductive biases arising from associations between language and expected reward learned during pre-training, shape LLM exploration behaviour. We find that semantically informative action labels reduce exploration in favour of exploitation, improving performance when aligned with the reward structure and severely degrading it when misaligned. We further find that negative rewards trigger substantially more exploration than equivalent positive rewards, consistent with an expected-scale bias induced by reward conventions common in pre-training data. Overall, we argue that the use of language to define the environment and rewards introduces unavoidable biases derived from the fact that the model is trained on word co-occurence, with implications for the reliability and robustness of LLM agents in real-world decision-making settings.","authors":["David Eric Austin","Kaheer Suleman","Jackie Chi Kit Cheung"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16707","pdf_url":"https://arxiv.org/pdf/2608.16707","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["LLM决策","探索-利用","语义偏差"],"reason":"研究LLM在语义多臂老虎机中的探索-利用行为，揭示语义先验导致的偏差，可迁移到…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":21,"question":"LLM智能体在语义多臂老虎机任务中，动作标签的语义先验和奖励极性如何影响探索-利用行为？","design":"用GPT-4o等LLM作为决策智能体，在语义多臂老虎机任务中，通过改变动作标签的语义类型（字母数字、情感、序数、世界知识）和标签与奖励分布的对齐方式（有帮助/误导），以及奖励极性（正/负），测量探索行为（选择新臂的比例）和累积遗憾。","baseline":"无对照","findings":"语义信息丰富的动作标签会减少探索、偏向利用，当标签与奖励结构对齐时提升性能，错位时严重损害性能。负奖励比等效正奖励触发显著更多的探索，表明存在由预训练中奖励惯例引起的期望尺度偏差。","reliability":"论文指出语义先验的影响取决于先验与真实奖励分布的对齐程度，错位时会导致系统性错误；但未讨论其他失效条件或模型规模、提示设计等稳健性因素。","relevance":"该研究揭示了LLM在决策任务中因语言表征引入的语义偏差，对使用LLM仿真人类经济决策的可靠性提出警示，值得阅读以了解偏差来源和实验设计。","inspiration":"借鉴其通过操纵标签语义和奖励极性来分离语义先验影响的设计，可迁移到消费者选择或投资决策中的标签效应研究。｜可应用于金融产品推荐或政策选项呈现中的框架效应，检验LLM是否复现人类的选择偏差。｜用LLM模拟投资者，处理为不同语义标签的资产选项（如“稳健增长”vs“高风险高回报”），结果变量为选择比例和探索行为，与真实投资者实验数据对照。"}},{"id":"2608.15838","version":1,"title":"PersonaEval: Persona-Based User Simulation for Evaluating Interactive Applications","zh_title":"PersonaEval：基于角色的用户仿真用于评估交互式应用","abstract":"Real user studies are important for understanding how people interact with systems under test or already deployed. In practice, however, they are often costly, time-consuming, and difficult to scale. To address these challenges, we introduce PersonaEval, a persona-based user simulation framework that approximates real-user behavior across diverse interactive settings. PersonaEval connects simulated users drawn from existing persona datasets to task-specific application interfaces and collects the interaction trajectories and outcomes. PersonaEval provides a plug-and-play evaluation workflow in which the application being evaluated can be easily changed. In this demo, we present PersonaEval on three forms of interactive applications: surveys, chatbots, and web applications. Together, these examples show that PersonaEval can support repeatable, parallelizable, and scalable evaluation across different interaction settings, while producing user-oriented feedback and task-specific behavior.","authors":["Yifan Simon Liu","Qianfeng Wen","Yilan Fan","Shirley Huang","Ruoqi Gao","Jianheng Hou","Muhammad Ahmed Mohsin","Zonglin Di","Brihi Joshi","Xincheng Tan","Yucheng Lu","Xiaoyi Liu","Heming Liu","Hanwen Xing","Guanghui Min","Zhengyang Shan","My Chiffon Nguyen","Ishan Gupta","Yunze Xiao","Hannah Collison","Jintao Huang","Jiatong Li","Sankalp Jajee","Yunhan Zhao","Bing Hu","Sky Ng","Xupeng Chen","Binghang Lu","Weihang Xiao","Aravind Mohan","Bolun Sun","Yunshu Wu","Yuanda Xu","Yun Shen","Runyu Zhang","Zheyuan Deng","Zhiwei Zhang","Qianyu Zhu","Dianzhuo Wang","Yijun Wang","Yixuan He","Yuexing Hao","Xiaomin Li"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15838","pdf_url":"https://arxiv.org/pdf/2608.15838","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","A4"],"tags":["用户仿真","人机交互","评估框架"],"reason":"用 persona 模拟用户与交互应用交互，近似真实用户行为，可迁移到人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":18,"question":"如何构建一个可插拔的、基于人格的用户仿真框架，以近似真实用户行为并评估交互式应用？","design":"使用 PersonaEval 框架，从 Nemotron 人格数据集中按应用描述与人格档案的嵌入相似度选取 50 个人格作为模拟用户；通过应用适配器将模拟用户连接到调查问卷、聊天机器人和网页应用三种交互界面；使用 Claude Haiku 4.5 驱动人格用户，GPT-4o mini 驱动聊天机器人；收集交互轨迹和结果，并测量用户评分、结果变异和人格一致性。","baseline":"无对照","findings":"PersonaEval 能在调查、聊天机器人和网页应用中产生人格一致的模拟交互，并揭示不同应用和人格群体间的评分差异。聊天机器人任务满意度中等且分布更广，而调查和网页结果更集中；美容、医疗和电影等应用的人格群体评分差异更大。","reliability":"论文承认当前演示覆盖的应用范围有限，仿真质量尚未通过真实用户行为对比进行严格验证，人格一致性仅由人类评判，未来需校准仿真与人类数据。","relevance":"该研究展示了基于人格的 LLM 用户仿真在多种交互场景下的可行性，但缺乏真实人类数据对照，对关注仿真可靠性与偏差的研究者价值有限，可作为方法参考而非实证基准。","inspiration":"可借鉴其模块化人格仿真框架和基于嵌入相似度的人格选择方法，用于构建多样化的经济主体样本。｜可迁移到消费者金融决策、信贷申请或政策反馈等场景，模拟不同背景人群的行为差异。｜以人格化的 LLM 作为被试，施加不同的金融产品信息或政策干预，测量其选择、满意度或风险偏好，并与真实调查或实验数据对照以校准仿真。"}},{"id":"2608.16067","version":1,"title":"SiMUSation: An Interactive Visitor Experience Simulation Framework to Support Museum Exhibition Design","zh_title":"SiMUSation：支持博物馆展览设计的交互式访客体验仿真框架","abstract":"Understanding how diverse audiences engage with narratives and content is central to exhibition design, yet designers often rely on intuition. Existing experience evaluation methods are typically retrospective, costly, and offer limited access to visitors' internal states, hindering early-stage iterative refinement. Rather than relying only on post-implementation evaluation with real visitors, we explore LLM-driven persona simulation as a reference for early-stage design. Following this idea, we present SiMUSation, an interactive framework designed to support early-stage exhibition design. SiMUSation models diverse visitor personas and simulates their exhibition experiences through a dual-layer representation that couples observable behaviors, such as movement and gaze, with corresponding internal responses, such as confusion and narrative engagement. Designers can steer simulations, inspect feedback from simulated visits, and iteratively revise layouts, content, and narrative flow to further examine how changes reshape visitor experience. We implemented a prototype and evaluated it through a user study (N=12), showing that SiMUSation provides insights for reflection and refinement in early-stage exhibition design. Our findings further highlight the potential of persona-driven simulation to support audience-informed evaluation and iterative decision-making across design tasks.","authors":["Huanchen Wang","Qiuming Chen","Zhonghao Ji","Ruqi Sun","Zhichao Lu","Yuxin Ma"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16067","pdf_url":"https://arxiv.org/pdf/2608.16067","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","A3","B1"],"tags":["LLM仿真","用户体验","博物馆设计"],"reason":"用LLM模拟访客体验并与真实用户研究对照，属于人类仿真且有人类数据基准。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":19,"question":"如何利用LLM驱动的访客角色仿真来支持博物馆展览的早期设计迭代？","design":"使用LLM构建多样化的访客角色，通过双层表示模拟其观展体验：可观察行为（移动、注视）与内部反应（困惑、叙事参与）。设计师可调整布局、内容和叙事流程，观察仿真反馈以迭代设计。","baseline":"用户研究（N=12）评估原型可用性和有效性，作为仿真结果的对照。","findings":"SiMUSation能为设计师提供可操作的访客体验洞察，帮助识别问题所在（通过行为轨迹）和原因（通过内部反应）。用户研究表明该框架支持早期设计反思和细化。","reliability":"论文未讨论仿真失效条件，但指出传统方法难以捕捉内部状态，且用户研究规模小（N=12），可能影响代表性。","relevance":"该研究将LLM仿真用于设计评估，与人类仿真研究相关，但缺乏与真实访客数据的直接对比，且场景为博物馆设计，对经济学实验的参考价值有限。","inspiration":"借鉴其双层仿真设计，将可观察行为与内部状态结合，用于经济决策仿真。｜可迁移到消费者行为研究，如购物决策中的注意力分配与偏好形成。｜以LLM模拟消费者在虚拟商店中的浏览路径和购买决策，处理为不同商品陈列方式，结果变量为购买率和内部偏好，对照真实消费者眼动和购买数据。"}},{"id":"2608.15181","version":1,"title":"Insurance as AI Risk Infrastructure: A Generative-Agent Simulation of AI Adoption","zh_title":"保险作为AI风险基础设施：AI采用的生成式智能体仿真","abstract":"The rapid evolution of artificial intelligence (AI) tools has demonstrated immense potential to enhance societal well-being and operational efficiency. However, the inherent unreliability and uncertain operational consequences of modern AI systems, typified by large language models (LLMs), have created a significant barrier to enterprise adoption. Many enterprises remain hesitant to integrate these tools deeply into their workflows due to concerns about unpredictable losses and liability exposure. While existing technical safeguards primarily seek to reduce the likelihood or severity of AI-enabled workflow failures, they do not by themselves provide ex post financial protection when residual pecuniary tail losses materialize. In this paper, we introduce a socio-economic framework that complements these safeguards by transferring and absorbing the residual financial consequences of AI adoption through insurance. To evaluate this framework, we develop an LLM-driven agent-based social simulation (LABSS) system. We assess the behavioral validity of the simulation using established economic and sociological theories. Our analysis demonstrates that the proposed insurance framework reduces firm-level financial exposure, thereby accelerating the aggregate adoption of AI tools and improving firm solvency and aggregate capital.","authors":["Yixuan Yuan","Dedai Wei","Chudong Qian","Jielin Feng","Ziyue Lin","Yuheng Zhao","He Cao","Erasmo Purificato","Xinwu Ye"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15181","pdf_url":"https://arxiv.org/pdf/2608.15181","source_feed":"cs.MA","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM社会仿真","AI采用","保险机制"],"reason":"用LLM agent模拟企业AI采用决策，涉及经济场景，但无真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":15,"question":"保险能否作为AI风险基础设施，通过转移和吸收AI采用中的残余财务损失，加速企业采用AI工具并改善企业偿付能力与总体资本？","design":"使用LLM驱动的智能体社会仿真系统（LABSS），模拟异质性企业在300天内做出AI采用、续约和保险决策，处理为是否提供保险，结果变量包括AI采用率、企业破产数、总资本、供应商承诺时长和跨行业采用差距。","baseline":"无对照","findings":"提供保险使最终AI采用率从74.97%提高到84.38%，平均破产企业数从4.33降至1.33，总资本更高。保险还带来更长的供应商承诺、更小的跨行业采用差距，并抑制恐慌传播。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟企业AI采用决策，属于经济场景下的人类仿真，但无真实人类数据对照，与您关注的有基准的仿真研究不完全匹配，可快速浏览其仿真设计。","inspiration":"借鉴其用LLM智能体模拟企业决策并施加保险处理的做法，可迁移到企业技术采用或风险管理决策的经济金融问题，例如企业对新金融技术的采用。｜设计一个LLM智能体仿真，让企业智能体在有无保险条件下决定是否采用AI，结果变量为采用率和破产率，并用真实企业调查或历史采用数据做对照。"}},{"id":"2607.23982","version":5,"title":"Moral Hazard in Multi-Agent Language Models","zh_title":"多智能体语言模型中的道德风险","abstract":"Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstr\\\"om's model of moral hazard in teams, we introduce the Dialogue Moral Hazard Game, a theory-grounded controlled experimental paradigm that instantiates this hidden-action structure as a textual environment for language agents. In each episode, an agent chooses between keeping an immediate local reward and paying a query cost to reveal a hidden safety fact that primarily helps another agent's downstream decision. We evaluate thirteen open-weight and four frontier models with stage-level mechanism metrics. In matched 3,015-decision-per-model experiments, GPT-5.6 Sol and Claude Opus 4.8 track the Holmstr\\\"om-derived private-share boundary across nine query costs (mean absolute errors 0.013 and 0.030); Muse Spark 1.1 responds directionally, whereas Fable 5 remains query-saturated. Diagnostic SFT, RLOO, SFT+RLOO, and GEPA updates are heterogeneous: SmolLM3-3B and OLMo-7B show the clearest weight-level mechanism gains, while GEPA raises Muse team success from $22.2\\pm3.8\\%$ to $100.0\\pm0.0\\%$ as query use falls from $51.1\\pm5.1\\%$ to $0.3\\pm0.5\\%$. Freezing the three Muse prompts and intervening on the rank--label mapping changes team success from $100.0\\%$ to $12.5\\%$ and then $0.0\\%$, with validity fixed at $100\\%$. Opus supplies a within-model contrast: its query-mediated prompt remains perfect across mappings, while two near-zero-query prompts follow the same trajectory. Optimization can therefore reach the same aggregate outcome through direct revelation or a learned effective information structure, motivating mechanism-level evaluation rather than team success alone.","authors":["Dane Malenfant"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"replace-cross","date":"2026-08-18","first_seen":"2026-07-28","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2607.23982","pdf_url":"https://arxiv.org/pdf/2607.23982","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体","道德风险","社会模拟"],"reason":"多智能体道德风险博弈，无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":3,"question":"语言智能体在多智能体团队道德风险博弈中，是否会因私人成本与收益份额而减少对社会有益但成本高昂的查询行为？","design":"构建基于Holmström团队道德风险模型的文本对话博弈环境，让多个LLM智能体扮演团队成员，通过改变查询成本、团队奖励和私人产出份额等参数，测量查询率、信息传递、局部奖励保留、不安全选择、格式有效性和团队成功等行为指标。","baseline":"无对照","findings":"前沿模型行为差异显著：GPT-5.6 Sol在主要设置中达到上限行为，并在激励隔离实验中以0.013的平均绝对误差追踪Holmström推导的私人份额边界；优化方法（如GEPA）可能提高团队成功但减少或消除代价高昂的查询，表明聚合奖励提升不一定恢复指定合作机制。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济理论中的道德风险场景，属于经济实验仿真，但缺乏真实人类数据对照，适合关注LLM行为与理论预测一致性的研究者阅读。","inspiration":"借鉴其参数化激励结构（成本、奖励、份额）和机制分解测量方法，可系统检验LLM对经济激励的边际响应。｜可迁移到委托代理问题，如基金经理努力与风险承担、员工团队合作中的搭便车行为。｜以LLM作为被试，在模拟投资任务中改变管理费率和业绩提成比例，测量其风险选择和努力程度，并与真实基金经理的历史数据对照。"}},{"id":"2608.14079","version":1,"title":"The conditional superiority of fast silicon sampling","zh_title":"快速硅采样的条件优越性","abstract":"Silicon sampling can produce surprisingly good population estimates at times. Does doing it fast attenuate such fidelity? In this study, we extend and assess ongoing work in silicon sampling by comparing the algorithmic fidelity of \"fast\" and \"slow\" modes of silicon sampling among a nationally representative sample of Singaporean survey respondents. We find that silicon sampling with contemporary frontier models remains a method in early development to be used only with great caution. While silicon samples are able to produce moderately faithful estimates of population means, they continue to understate opinion variance and distort the latent contextual space behind human opinions. Conditional on such limitations, we find \"fast\" modes of silicon sampling to be relatively superior to traditional \"slow\" modes of silicon sampling. Fast silicon sampling is significantly more efficient in compute resources and run-time while being monotonically superior to slower modes of sampling in algorithmic fidelity.","authors":["Nickolas Hock Yuen Lam","Ji Xuan Voo","Xiangyu Ma"],"categories":["cs.CL","cond-mat.mtrl-sci"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14079","pdf_url":"https://arxiv.org/pdf/2608.14079","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["硅采样","算法保真度","人类仿真"],"reason":"直接比较快速与慢速硅采样在代表性样本上的算法保真度，含真实人类数据对照，并指出…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":1,"question":"快速硅采样是否在算法保真度上不劣于甚至优于传统慢速硅采样？","design":"使用当代前沿模型（OpenAI GPT 5.4）对新加坡全国代表性调查受访者进行硅采样，比较“快速”模式（一次提示生成大量响应）与“慢速”模式（逐个API调用生成响应）的算法保真度，测量结果包括总体均值估计、意见方差和潜在情境空间结构。","baseline":"新加坡全国代表性调查受访者的真实人类数据。","findings":"硅采样能中等程度地忠实估计总体均值，但低估意见方差并扭曲潜在情境空间。在承认这些局限的前提下，快速硅采样在算法保真度上单调优于慢速硅采样，且计算资源和运行时间显著更高效。","reliability":"论文承认硅采样仍处于早期发展阶段，需谨慎使用；硅样本低估意见方差、扭曲潜在情境空间，且快速与慢速模式均存在这些局限。","relevance":"该研究直接比较快速与慢速硅采样在代表性样本上的算法保真度，含真实人类数据对照，并指出硅采样在方差和潜在空间上的失效条件，对关注LLM仿真可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其通过几何数据分析（多重对应分析）评估关系保真度的方法，以及比较不同采样模式效率与保真度的设计。｜可迁移到政策评估中的公众意见模拟，如经济政策公告的预期形成或消费者信心调查。｜以LLM生成不同处理模式下的合成受访者，处理为快速与慢速采样，结果变量为对经济政策的态度分布，对照真实调查数据（如新加坡消费者信心指数），评估均值、方差和潜在空间结构的一致性。"}},{"id":"2608.10492","version":2,"title":"INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators","zh_title":"洞察学生思维：联合建模LLM学生模拟器中的潜在推理与行为","abstract":"Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework that fine-tunes LLMs not only to act like students but also to think like them. INSIDE generates internal dialogue grounded in Bloom's Taxonomy across cognitive, affective, and action dimensions, and fine-tunes models on paired think traces and actions. We baseline against different prompting frameworks and evaluate on two axes: fidelity of simulated actions and quality of generated internal dialogue. Our evaluations show that INSIDE improves simulation fidelity in both action fidelity, matching code generation of real students, and reasoning alignment, achieving the highest alignment across models up to 57.9%.","authors":["Rose Niousha","Minwoo Kang","Narges Norouzi"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-17","first_seen":"2026-08-12","revised_at":"2026-08-17","abs_url":"https://arxiv.org/abs/2608.10492","pdf_url":"https://arxiv.org/pdf/2608.10492","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","教育模拟","算法保真度"],"reason":"用LLM仿真学生行为与推理，并与真实学生数据对照，评估仿真保真度。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-17","rank":3,"question":"如何让LLM学生模拟器不仅复现学生的可观察行为（代码提交），还能捕捉其背后的潜在推理过程，从而提高仿真保真度？","design":"提出INSIDE框架，基于Bloom分类法生成内部对话（认知、情感、行动维度），并用配对的思想痕迹和行动对LLM进行微调。模型扮演编程课程学生，输入学生历史提交和AI导师反馈，输出下一步代码提交。评估两个维度：行动保真度（生成代码与真实学生代码的相似度）和推理质量（生成推理与真实代码编辑的对齐度）。","baseline":"使用加州大学伯克利分校入门编程课程两个学期的真实学生数据：Spring 2025用于训练（445名学生，2022条提交流，6911次提交），Spring 2024用于测试（479名学生，1546条提交流，6316次提交）。测试集分为旧问题（test_OP）和新问题（test_NP），分别评估对未见学生和未见问题的泛化。","findings":"INSIDE提高了行动保真度，生成代码与真实学生代码的Wasserstein距离更低；同时实现了最高的推理对齐度，在不同模型上最高达到57.9%。","reliability":"论文未讨论","relevance":"该研究直接命中你的核心关注点：用LLM仿真学生行为与推理，并与真实学生数据对照，评估仿真保真度。它提供了在缺少真实推理标注的情况下重建潜在推理的方法，并展示了联合建模推理与行动能提升仿真质量，值得精读原文。","inspiration":"借鉴其用理论框架（Bloom分类法）引导内部对话生成，并将推理与行动联合微调以提升仿真保真度的做法。｜可迁移到经济金融中的政策预期形成研究，例如模拟投资者在信息发布后的决策过程，捕捉其推理路径。｜以LLM扮演投资者，输入历史交易和新闻信息，生成内部推理（如风险评估、情绪反应）和交易决策，用真实市场交易数据和调查数据（如投资者信心指数）作为对照，评估仿真保真度。"}},{"id":"2601.20238","version":2,"title":"Large Language Models Polarize Ideologically but Moderate Affectively in Online Political Discourse","zh_title":"大语言模型在网络政治话语中加剧意识形态极化但缓和情感极化","abstract":"The emergence of large language models (LLMs) is reshaping how people engage in political discourse online. We examine how the release of ChatGPT altered ideological and emotional patterns in Reddit's largest political forum. Analysis of millions of comments shows that ChatGPT intensified ideological polarization: liberal-leaning authors posted increasingly liberal comments, while conservative-leaning authors posted increasingly conservative comments. Multiple falsification tests suggest that these findings are unlikely to be driven by contemporaneous events, such as the 2022 U.S. midterm elections, or by broader platform-wide trends in political polarization. Mechanism tests show that this shift does not stem from the creation of more persuasive or ideologically extreme original content using LLM. Instead, it originates from the tendency of LLM-assisted comments to echo and reinforce the original post's viewpoint, a pattern consistent with algorithmic sycophancy. Yet, despite growing ideological divides, affective polarization, measured by hostility and toxicity, declined. These findings reveal that LLMs can simultaneously deepen ideological separation and foster more civil exchanges, challenging the long-standing assumption in literature that extremity and incivility necessarily move together.","authors":["Gavin Wang","Srinaath Anbudurai","Oliver Sun","Xitong Li","Lynn Wu"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-08-17","first_seen":"2026-01-28","revised_at":"2026-08-17","abs_url":"https://arxiv.org/abs/2601.20238","pdf_url":"https://arxiv.org/pdf/2601.20238","source_feed":"econ.GN","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","政治极化","人类数据对照"],"reason":"用LLM辅助评论与真实Reddit数据对照，分析政治话语极化，涉及社会过程仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-17","rank":4,"question":"ChatGPT的发布如何影响Reddit政治论坛中用户的意识形态极化和情感极化？","design":"本研究并非将LLM作为人类被试的仿真实验，而是利用Reddit上最大的政治论坛在ChatGPT发布前后的数百万条评论数据，通过语言特征指标（如评论长度、被动动词、困惑度、人类评分）估计每条评论由LLM辅助生成的概率，进而分析LLM使用对用户意识形态立场和情感表达的影响。","baseline":"以ChatGPT发布前的Reddit评论作为对照，比较同一作者在发布前后的意识形态立场变化，并利用2022年美国中期选举等事件进行证伪检验。","findings":"ChatGPT的发布加剧了意识形态极化：自由派作者发表更自由的评论，保守派作者发表更保守的评论。然而，情感极化（敌意和毒性）却有所下降，表明LLM在加深意识形态分歧的同时促进了更文明的交流。","reliability":"论文承认无法直接确定每条评论是否由LLM生成，而是依赖语言指标估计概率，可能存在分类误差；但认为在大规模语料中，误差不会系统性地产生所观察到的聚合模式。","relevance":"该研究利用真实Reddit数据评估LLM对政治话语的影响，涉及LLM辅助内容生成与真实人类行为的对照，对关注LLM在社会科学中仿真可靠性的研究者有参考价值，但并非直接以LLM作为被试的仿真实验。","inspiration":"值得借鉴的是利用自然实验（ChatGPT发布）和语言特征指标来估计LLM使用概率，并设置证伪检验排除混淆因素。｜可迁移到经济金融领域如政策公告后的市场情绪分析、消费者评论中的LLM影响等场景。｜一个可行的设计是：以某经济政策发布为时间节点，收集社交媒体上相关讨论，用语言指标估计LLM辅助评论比例，分析其对情绪极化和观点极化的影响，并与历史人类评论基线对照。"}},{"id":"2608.13712","version":1,"title":"Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation","zh_title":"字里行间：法律模拟中行为真实性的建模与评估","abstract":"Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator driven by controllable legal personas. We use an evaluation framework separating behavioral realism from pedagogical usefulness. We assess realism through adversarial testing, blinded attorney comparison, and analysis of longitudinal behavioral trajectories. WitnessSim generally maintained plausible behavioral boundaries, and attorneys did not systematically prefer either original testimony or WitnessSim generated testimony. Pedagogical tests showed that witness behavior changed meaningfully in response to question form and attorney intervention without uniformly collapsing the assigned persona. Together, these results showcase a model of behavioral fidelity in legal simulations, and provide a framework for evaluating its performance.","authors":["Divya Vetticaden","Arya Gupta","Julian Nyarko","Megan Ma"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13712","pdf_url":"https://arxiv.org/pdf/2608.13712","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","法律模拟","行为真实性"],"reason":"用LLM模拟法律证人行为，并与真实证词对照，评估行为真实性和教学效果，属于人类…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":6,"question":"如何构建并评估一个具有行为真实性和教学实用性的法律证人模拟系统？","design":"WitnessSim 是一个基于六维状态向量（镇定、知识、宜人性、冗长度、僵化度、表现）的沉积证词模拟器，通过问题特征（压力、话题敏感性、问题形式）更新状态，并条件化生成证词；研究使用对抗测试、盲法律师比较和纵向行为轨迹分析评估行为真实性，并通过法律培训材料衍生的测试评估教学实用性。","baseline":"来自国家处方阿片类药物诉讼（Case No. 1:17-MD-2804）的 300 份沉积和庭审笔录，作为真实人类证词对照。","findings":"WitnessSim 在对抗测试中维持了合理的行为边界，律师在盲法比较中未系统偏好原始证词或生成证词。教学测试显示证人行为对问题形式和律师干预有有意义的变化，且未完全丧失指定人格。","reliability":"论文未明确讨论失效条件与局限，但指出评估框架区分行为真实性与教学实用性，并承认行为真实性通过对抗测试、专家评估和情感轨迹分析来操作化，可能仍存在未覆盖的维度。","relevance":"该研究直接命中研究者关注的 LLM 人类仿真实验，提供了法律场景下与真实人类数据对照的行为真实性评估框架，值得阅读原文以借鉴其多维评估方法和动态行为建模。","inspiration":"借鉴其将行为状态建模为可更新的多维向量，并通过问题特征施加处理、以真实行为轨迹为基准的评估方法。｜可迁移到经济金融中的谈判、审计或客户服务交互模拟，如信贷审批中的申请人行为或政策沟通中的公众反应。｜以 LLM 模拟信贷申请人，处理变量为审批官提问的侵略性或信息敏感度，结果变量为申请人的情绪状态和回答一致性，对照真实信贷申请面谈记录。"}},{"id":"2608.13786","version":1,"title":"Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions","zh_title":"AI聊天机器人能否找到专家会找到的研究？模型、用户角色和样本量对医学问题研究检索的影响","abstract":"Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence-synthesis researcher roles. Each chatbot was queried under each user role with four independent repetitions, yielding 720 responses. Each chatbot was asked to support its answers with primary clinical citations, which we benchmarked against the included and excluded study sets of the Cochrane reviews. On average, a chatbot response retrieved 39.2% $\\pm$ 29.8% of Cochrane included studies, while citing 5.0% $\\pm$ 9.4% of excluded studies. Recall of Cochrane included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\\pm$ 29.5% vs. 37.0% $\\pm$ 23.8% vs. 17.3% $\\pm$ 13.1%; $p=2.0\\times10^{-5}$). The researcher role yielded higher recall than the clinician or patient roles (42.8% $\\pm$ 30.8% vs. 38.6% $\\pm$ 28.9% vs. 36.1% $\\pm$ 29.3%; $p=2.0\\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only independently significant predictor of retrieval (odds ratio 1.80 per 1-unit increase in log sample size, 95% CI 1.37-2.36, $p=2.34\\times10^{-5}$). These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes.","authors":["Qingfang Liu","Qiao Jin","Joe D. Menke","Thorsten Kahnt","Zhiyong Lu"],"categories":["cs.IR","cs.AI","cs.CL"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13786","pdf_url":"https://arxiv.org/pdf/2608.13786","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","医学信息检索","人类对照"],"reason":"用LLM模拟不同用户角色检索医学证据，并与Cochrane专家评审结果对照，属…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":11,"question":"不同大语言模型聊天机器人在模拟患者、临床医生和证据综合研究者三种用户角色时，检索医学问题相关临床研究的表现如何，以及哪些因素影响其检索结果？","design":"用三个通用大语言模型（Claude Sonnet 5、Gemini 3.1 Pro、ChatGPT GPT-5.5）模拟患者、临床医生和证据综合研究者三种角色，对20个Cochrane系统综述问题生成带引文的回答，每个角色-模型组合重复4次，共720个回答，以Cochrane综述纳入和排除的研究集为基准，测量检索召回率和引用排除研究的比例。","baseline":"以20个Cochrane系统综述中专家评审纳入和排除的研究集作为真实人类专家判断的基准。","findings":"平均每个回答检索到39.2%的Cochrane纳入研究，引用5.0%的排除研究；ChatGPT召回率显著高于Claude和Gemini，研究者角色召回率显著高于临床医生和患者角色。控制发表年份、年均引用和开放获取状态后，样本量是唯一显著预测检索的因素，模型偏向于检索大样本临床试验。","reliability":"论文承认其评估仅限于三个特定模型和20个Cochrane综述主题，可能不具普遍性；未讨论模型版本更新、提示词细微变化或不同医学领域对结果的影响。","relevance":"该研究直接以LLM模拟不同用户角色进行信息检索，并与专家评审结果对照，属于人类仿真实验，且揭示了模型和角色对检索行为的影响，值得阅读原文以了解仿真偏差的具体表现。","inspiration":"借鉴其通过角色扮演和重复查询来测量LLM行为差异，并利用专家评审数据作为基准的方法｜可迁移到经济金融领域，如模拟投资者、分析师和监管者角色检索金融研究报告或政策文件，检验信息获取偏差｜用LLM扮演不同金融角色，对特定经济问题（如某行业前景）检索相关研究，以权威机构（如央行工作论文或顶级期刊）的文献列表为基准，测量检索召回率和偏差，并分析文献特征（如样本量、发表期刊）对检索的影响。"}},{"id":"2608.14320","version":1,"title":"AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs","zh_title":"AnchorBench：LLM锚定效应的多路径基准测试","abstract":"The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that large language models (LLMs) exhibit similar behavior. However, existing work on anchoring in LLMs typically evaluates only a narrow set of anchor pathways and rarely distinguishes irrelevant from plausible anchors. We introduce AnchorBench, a benchmark for the anchoring effect in LLMs that evaluates multiple anchor pathways under an explicit anchor relevance axis. Across fourteen models, including ten open-weight models and four frontier API models, and a large set of controlled prompts, we find that (1) anchoring is strongly pathway-dependent, (2) plausible anchors usually induce larger shifts than irrelevant ones when introduced through stronger pathways, (3) anchor influence generally weakens as the anchor moves farther from the evidence-supported answer, most clearly on External and RAG, and (4) high task accuracy on the anchor-free control condition (Acc$_{10}$: answers within 10 points of gold) does not guarantee robustness: even frontier API models above 95% control accuracy remain susceptible to plausible anchors.","authors":["Yiderigun Borjigin","Alexander Hermann","Christian Cyron","Roland Aydin"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14320","pdf_url":"https://arxiv.org/pdf/2608.14320","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["认知偏差","LLM评估","锚定效应"],"reason":"评估LLM锚定效应，与人类认知偏差对照，可迁移到仿真可靠性研究","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":12,"question":"LLM在多种锚定信息传递路径下是否表现出锚定效应，且该效应如何随锚定相关性和路径强度变化？","design":"构建AnchorBench基准，使用14个LLM（10个开源权重、4个API模型）完成0-100数值判断任务，通过五种路径（外部提示、对话历史、上下文学习、检索增强、工具输出）施加低/高锚定值，并设置无关与合理锚定条件，测量输出相对无锚定控制条件的偏移。","baseline":"无对照","findings":"锚定效应强烈依赖路径，合理锚定在强路径下比无关锚定引起更大偏移；锚定影响随锚定值与证据支持答案距离增大而减弱，且高控制准确率不保证稳健性，即使API模型也易受合理锚定影响。","reliability":"论文未讨论","relevance":"该研究系统评估LLM锚定效应，与人类认知偏差对照，可迁移到仿真可靠性研究，值得阅读原文以了解多路径设计和相关性区分方法。","inspiration":"借鉴其多路径施加处理和相关性轴设计，区分无关与合理锚定，测量偏移量而非仅准确率｜可迁移到资产定价实验或政策公告的预期形成研究，检验LLM对锚定信息的敏感性｜以LLM模拟投资者，在提示中通过不同路径（如新闻、历史对话）注入锚定价格，测量估值偏移，并与人类实验数据对照。"}},{"id":"2608.14399","version":1,"title":"Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice","zh_title":"AI推荐哪位医生？大语言模型辅助医生选择中声誉与人口统计信号的算法审计","abstract":"Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt-4o-mini) each chose among five synthetic family-medicine physician cards whose attributes were independently randomized across 3,024 choice sets, three patient personas, nine prompt paraphrases and nine experimental arms, yielding 40,068 scored responses; gender and ethnicity were signaled through names following correspondence-audit methodology. Reputation signals dominate: raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points (pp), and raising the fee from $90 to $190 lowers it by 20.0 pp. Demographic parity is rejected, but not in the direction human audit studies predict: female-signaled names gain 2.5 pp, and Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names, tilts worth $7-$14 per visit in fee-equivalent terms, and a content-free first-listed position is worth $11. Yet models mentioned gender or ethnicity in at most 0.03% of their stated reasons and abstained in 0.39% of trials, so these effects are invisible in the models' own explanations, and transparency obligations relying on model self-report would not detect them. One reasoning model failed the prespecified auditability gate outright. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli, making recurring behavioural audit, rather than self-reported explanation, the monitoring technology fit for purpose.","authors":["Syeda Anshrah Gillani","Mirza Samad Ahmed Baig"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14399","pdf_url":"https://arxiv.org/pdf/2608.14399","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","算法审计","医疗决策"],"reason":"用LLM模拟患者选择医生，与真实人类审计研究对照，揭示偏差与失效条件","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":13,"question":"在LLM辅助的医生推荐中，声誉信号、姓名暗示的人口统计特征和列表位置如何因果地影响推荐概率？","design":"用七个LLM（六个开源权重模型和gpt-4o-mini）扮演患者，在随机选择联合实验中从五张合成家庭医生卡片中选择医生；卡片属性独立随机化，包括患者评分、评论量、费用、性别和种族（通过姓名暗示）等；结果变量是医生被选择的概率。","baseline":"与人类审计研究对照，特别是关于医生选择中反少数族裔歧视的发现。","findings":"声誉信号主导推荐：评分从3.9提高到4.7使选择概率增加31.4个百分点，费用从90美元提高到190美元降低20.0个百分点。人口统计均等被拒绝，但方向与人类审计研究相反：女性姓名增加2.5个百分点，西班牙裔、南亚裔和黑人姓名比白人姓名增加1.3-2.9个百分点，相当于每次就诊7-14美元的费用等值。","reliability":"模型在陈述理由中提到性别或种族的比例最多为0.03%，在0.39%的试验中弃权，因此这些效应在模型自身解释中不可见；一个推理模型（deepseek-r1:7b）未能通过预设的可审计性门槛。","relevance":"该研究用LLM模拟患者选择医生，与真实人类审计研究对照，揭示偏差与失效条件，直接回应了研究者对LLM人类仿真实验和批判性评估的兴趣。","inspiration":"借鉴随机选择联合实验和姓名信号法，在LLM仿真中独立操纵属性以识别因果效应，并用费用等值量化效应大小。｜可迁移到信贷审批歧视研究，用LLM模拟信贷员审批贷款，操纵申请人姓名、收入、信用评分等属性。｜用多个LLM作为被试，随机生成贷款申请档案，处理变量为申请人姓名暗示的种族和性别，结果变量为批准概率，与真实信贷审批数据中的歧视模式对照。"}},{"id":"2608.14113","version":1,"title":"Search or Chat? Comparing How We Learn About Debated Topics","zh_title":"搜索还是聊天？比较我们如何了解有争议的话题","abstract":"As large language models (LLMs) become more integrated into everyday information platforms, chat-based systems are emerging as a popular alternative to traditional web searches, especially for informational search and informal learning tasks. Despite this shift, little is known about how different tools affect learning outcomes. Our work aims to improve the understanding of how chat-based information access supports and impacts learning performance in informal learning settings. In this paper, we present the results of a crowdsourcing user study (N = 194) that compares learning about debated topics using a traditional search interface versus an LLM-powered chat interface. Through our analysis of learning outcomes, user characteristics, and interaction patterns, we found no significant differences in user learning gain or critical reflection on our study tasks. Our observations from the analysis of further exploratory variables suggest that, in the context of longstanding debated topics, user characteristics such as their attitude strength and level of intellectual humility might be more important in shaping immediate learning outcomes than the information access tool.","authors":["Ran Yu","Alisa Rieger","Rabia Karatoprak Ersen","Jiqun Liu"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14113","pdf_url":"https://arxiv.org/pdf/2608.14113","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM聊天界面","学习效果","用户研究"],"reason":"用LLM聊天界面替代搜索，比较学习效果，有真实用户数据对照，但非直接仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-17","rank":7,"question":"在非正式学习场景下，传统搜索界面与LLM聊天界面在争议性话题学习效果上是否存在差异？","design":"本研究不是LLM仿真人类研究，而是一项众包用户实验（N=194），比较真实人类被试在两种信息获取工具（传统搜索界面 vs. LLM聊天界面）下的学习效果。被试被随机分配一个争议性话题和一种工具，在限定时间内自主学习，之后测量学习增益（论点扩展）和批判性反思（批判性推理），并记录交互日志。","baseline":"无对照（本研究未使用LLM仿真人类，而是直接比较两种工具对真实人类学习的影响）","findings":"在争议性话题上，使用聊天机器人与使用搜索引擎的被试在论点扩展和批判性推理上无显著差异。探索性分析表明，用户的态度强度和智识谦逊水平比信息获取工具更能预测即时学习效果。","reliability":"论文未讨论（但研究本身并非仿真研究，其局限可能包括：仅测量即时学习效果，未追踪长期保持；争议性话题选择有限；众包被试可能缺乏深度参与动机；未控制工具使用时间差异等）","relevance":"该研究与研究者关注点部分相关：它比较了LLM聊天界面与传统搜索对人类学习的影响，但并非用LLM仿真人类被试，而是将LLM作为信息工具。对于关心LLM在信息获取中作用的研究者有一定参考价值，但若聚焦于LLM仿真人类行为，则相关度有限。","inspiration":"可借鉴其随机对照实验设计，将信息获取工具作为处理变量，测量认知结果（如学习增益、批判性思维），并考察用户特征（如态度强度、智识谦逊）的调节作用。｜可迁移到经济金融领域的信息处理与决策场景，例如投资者使用不同信息工具（搜索引擎 vs. LLM聊天机器人）获取公司财报信息后，其投资决策质量、信息理解深度或过度自信程度的变化。｜设计一个实验：招募真实投资者作为被试，随机分配使用搜索引擎或LLM聊天机器人获取某上市公司财报信息，之后测量其对公司的估值准确性、信息回忆和投资信心，并与历史市场数据或分析师预测作为基准对照，考察工具类型对决策质量的影响。"}},{"id":"2608.12368","version":1,"title":"Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments","zh_title":"一致不等于对齐：人类与LLM道德判断中分歧的道德依据","abstract":"Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation. We test this distinction using a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. Across frontier and open model families, agreement with human annotator majority labels is often high. However, rationale-level analysis reveals systematic divergence in the moral grounds expressed by human annotators and models. In particular, models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even when their final labels match the human annotator majority. Our results show that agreement should not be treated as equivalent to alignment. Label-based evaluation can therefore be misleadingly reassuring unless complemented by analysis of the reasons, principles, and moral priorities expressed in model judgments.","authors":["Octavian M. Machidon","Alina L. Machidon","Vojko Strahovnik","Mateja Centa Strahovnik","Jonas Miklav\\v{c}i\\v{c}","Marko Robnik \\v{S}ikonja"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12368","pdf_url":"https://arxiv.org/pdf/2608.12368","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM对齐评估","道德判断","人类对照"],"reason":"评估LLM道德判断与人类的一致性，揭示标签一致但理由分歧，有真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:00:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":2,"question":"在道德判断任务中，人类标注者与LLM的最终标签一致是否意味着它们依赖相同的道德理由？","design":"本研究并非将LLM作为人类被试的替代品进行仿真实验，而是直接比较人类标注者与多个LLM在500项ETHICS衍生道德判断任务上的表现。人类标注者和LLM对相同项目进行标注，提供最终标签和支持理由；研究者分析标签一致性与理由层面的分歧。","baseline":"新收集的人类标注者数据，包括最终标签和理由标注，作为与LLM输出比较的基准。","findings":"LLM与人类多数标签的一致性通常较高，但理由层面存在系统性分歧。即使最终标签一致，模型在伤害、尊重、守诺、正义、应得、借口相关性等道德理由类别上的注意力分布与人类不同。","reliability":"论文指出，标签一致性可能误导性地令人放心，因为模型可能依赖不同的道德理由，导致泛化差异、解释不匹配、过度道德化或忽视隐含社会义务。此外，理由对齐是描述性的，不能证明所述理由是模型输出的因果原因。","relevance":"该研究直接回应了研究者对LLM仿真可靠性的关注，揭示了仅用标签一致性评估对齐的不足，并提供了真实人类数据对照，对批判性评估LLM在道德判断中的仿真效度具有重要价值。","inspiration":"借鉴其双层评估设计（标签一致性与理由对齐）和理由编码框架，可迁移到经济金融中的伦理决策场景（如信贷审批中的公平性判断、消费者对金融产品道德性的评价）。设计雏形：以LLM作为被试，呈现金融道德困境（如掠夺性贷款案例），要求给出判断和理由，与人类专家标注的理由类别进行对比，以真实人类标注数据为基准，检验LLM在金融伦理判断中的理由一致性。"}},{"id":"2608.12788","version":1,"title":"ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs","zh_title":"ARAC：基准测试自动研究在端到端研究中的对齐性与完整性","abstract":"The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.","authors":["Jiale Cui","Yueyao Yuan","Kaixi Zhong","Xiaogang Xu","Jiafei Wu","Zhe Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12788","pdf_url":"https://arxiv.org/pdf/2608.12788","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A4","B1"],"tags":["自动研究评估","人类行为对齐","基准测试"],"reason":"提出评估自动研究系统与人类研究行为对齐度的基准，含人类专家对照，方法论可迁移至…","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":8,"question":"如何评估自动研究系统在多大程度上模拟了人类研究者的认知过程与方法论，而非仅匹配最终结果？","design":"提出 ARAC-Bench 基准，以顶级会议接收论文为黄金标准，从论文与审稿意见中提炼学术认知技能（ACS）作为分阶段量化评分标准；将研究流程分解为提案、实验、综合三个独立阶段，在统一基础模型和知识库下对 11 个自动研究框架进行受控评估。","baseline":"以 200 篇 ICLR 2026 接收论文的结构化标注为黄金参考，并邀请 10 名博士生专家进行人工排名，计算基准分数与专家排名的相关性。","findings":"最佳框架的总体对齐度仅为 67.9%，与人类方法论完整性存在显著差距；ARAC-Bench 与博士生专家排名在提案和综合阶段的相关性分别达 0.8788 和 0.9030，实验阶段为 0.6606，表明基准能可靠反映研究者重视的维度。","reliability":"论文指出实验阶段与专家判断的相关性相对较低，原因在于人类专家会综合考虑代码质量、实验设计理念和资源等因素，而基准可能未完全捕捉这些方面；此外，基准基于特定领域（AI）的论文构建，跨领域迁移性有待验证。","relevance":"该研究为评估 LLM 模拟人类研究过程提供了系统化基准，其方法可迁移至经济学实验仿真评估，值得阅读原文以借鉴其分阶段评分与专家对照设计。","inspiration":"借鉴其将复杂认知过程分解为可量化阶段并构建专家对齐评分标准的做法，可用于评估 LLM 模拟经济决策过程的质量｜可迁移到政策评估场景，如模拟消费者对政策公告的反应或投资者对经济新闻的解读｜设计：以真实经济实验数据（如实验室资产定价实验）为基准，让 LLM 扮演投资者，处理为不同政策信息，结果变量为交易行为与价格预期，用 ARAC 式分阶段评分与人类被试数据对照。"}},{"id":"2608.12387","version":1,"title":"Query Timing Produces Opposite Positional Biases Between LLMs and Humans","zh_title":"查询时机导致LLM与人类之间相反的位置偏差","abstract":"Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluations remains poorly understood. Both primacy and recency biases have been observed in human judgments in response to evidence, but recent work suggest that \\emph{when} the listener updates their beliefs -- during the presentation of evidence or only at the end -- influences the presence of such effects. We investigate whether a similar phenomenon holds for LLMs, finding divergence from human behavior. These biases are more exacerbated in newer models compared to their predecessors.","authors":["Jasin Cekinmez","Addison J. Wu","Thomas L. Griffiths"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12387","pdf_url":"https://arxiv.org/pdf/2608.12387","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","人类对照","认知建模"],"reason":"研究LLM与人类在位置偏差上的差异，有真实人类数据对照，评估仿真可靠性，可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:00:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":6,"question":"LLM在评估证据时，其位置偏差（首因/近因效应）是否受回答模式（逐步回答 vs. 序列末回答）影响，并与人类行为有何差异？","design":"使用多个开源和闭源LLM（包括不同版本）作为被试，在刑事、学术、社会三种指控情境中，通过逐步（SbS）和序列末（EoS）两种回答模式呈现控方和辩方证据，测量最终有罪判决比例。","baseline":"基于Qiao和Lagnado (2025)的人类实验数据，该研究发现人类在SbS模式下表现出近因偏差，在EoS模式下无总体偏差。","findings":"LLM在EoS模式下表现出显著的首因偏差，在SbS模式下表现出近因偏差，与人类行为相反。较新版本的LLM比旧版本表现出更强的位置偏差。","reliability":"论文未讨论","relevance":"该研究直接比较LLM与人类在证据顺序效应上的差异，有真实人类数据对照，评估了LLM作为人类判断代理的可靠性，对关注仿真偏差的研究者有价值。","inspiration":"借鉴其通过改变信息呈现顺序和回答模式来分离认知偏差的实验设计，可迁移到经济金融中的信息处理场景，如分析师对财报信息的顺序依赖或投资者对新闻的反应。｜可设计一个实验：用LLM模拟投资者，逐步或一次性呈现利好和利空新闻，测量其最终投资决策，并与真实投资者在类似实验中的行为数据（如实验经济学中的资产定价实验）进行对照，检验LLM是否复现人类的顺序偏差。"}},{"id":"2608.12344","version":1,"title":"Predicting consumer-technology ownership without a diffusion history","zh_title":"无扩散历史下预测消费者技术拥有率","abstract":"We test whether the perceived attributes of a consumer technology predict how widely it is owned. In a 2022 Prolific survey of US adults (n = 678), respondents rated 65 consumer technologies on six attributes. We then elicited the same ratings from two frontier language models, Anthropic Claude Opus 4.7 and OpenAI GPT-5.5. We regress ownership prevalence on four UTAUT2 acceptance attributes plus a log-age covariate with a sign-constrained penalized regression and evaluate it by holding out one technology at a time. The attribute model improves on a baseline of years-since-launch: mean absolute error falls by 17% with the human ratings, and by more with either model, most with Opus 4.7. Over the short 2022-to-2025 window, where ownership moved little, the same attributes do not improve on a no-change baseline. We set out the limitations of the approach, including the possibility that language-model ratings reflect prior knowledge of these technologies rather than independent attribute reasoning. We include a deployment illustration: 2027 ownership predictions for eleven products launched in 2025 and 2026.","authors":["Irina Vartanova","Niels Selling","Jennifer Viberg Johansson","Pontus Strimling"],"categories":["cs.CL","cs.CY","stat.AP"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12344","pdf_url":"https://arxiv.org/pdf/2608.12344","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","消费者行为","算法保真度"],"reason":"用LLM替代人类被试预测技术拥有率，并与真实调查数据对照，评估模型可靠性，属核…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":1,"question":"消费者技术的感知属性（UTAUT2 属性）能否在无扩散历史的情况下预测其拥有率，以及预测效果是否因评分者是人类还是大语言模型而异？","design":"用 2022 年 Prolific 调查中 678 名美国成年人对 65 项消费者技术的六项属性评分作为人类评分，再让两个前沿大语言模型（Claude Opus 4.7 和 GPT-5.5）对同样技术给出相同属性评分；以拥有率为结果变量，用符号约束惩罚回归拟合四个 UTAUT2 属性加对数年龄协变量，通过留一技术交叉验证评估预测误差。","baseline":"2022 年 Prolific 调查中 678 名美国成年人的真实拥有率和属性评分，以及 2025 年随访调查的拥有率变化。","findings":"属性模型相比仅用技术年龄的基线降低了平均绝对误差：人类评分降低 17%，大语言模型评分降低更多，其中 Opus 4.7 表现最好（MAE 7.1 个百分点）。但在 2022 至 2025 年拥有率变化很小的短窗口内，属性模型未能优于无变化基线。","reliability":"论文承认大语言模型评分可能反映了模型对技术的先验知识而非独立的属性推理，且模型无法预测拥有率随时间的变化；此外，属性模型仅适用于横截面水平预测，不涉及扩散动态。","relevance":"该研究直接使用大语言模型替代人类被试进行属性评分，并与真实调查数据对照，评估了仿真在预测技术拥有率上的可靠性，属于典型的 LLM 仿真实验，且包含批判性讨论，值得精读。","inspiration":"借鉴其用 LLM 生成属性评分并与人类评分对比、以真实拥有率作为结果变量的设计，可迁移到消费者金融产品采纳预测（如数字支付、理财产品）或政策接受度评估；具体可设计让 LLM 扮演不同人口群体对新型金融产品进行 UTAUT2 属性评分，以实际调查的采纳率作为基准，检验 LLM 评分能否预测真实采纳率并识别偏差。"}},{"id":"2608.12339","version":1,"title":"Mimicry without understanding: the origins of decision bias in large language models","zh_title":"无理解的模仿：大语言模型中决策偏差的起源","abstract":"Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through which such biases can be generated even when human preferences (in the training data) are not biased or when they are correctly categorized as being biased. The first is faulty mimicry of preferences based on human behavior: this involves LLMs inferring human preferences even when behaviors are logically unrelated to preferences. The second is mimicry of explicitly biased human behaviors. In four studies focusing on economic biases, we find that ChatGPT-4o and Qwen exhibited social proof biases even when prompted with reports of human behaviors that were clearly non-indicative of individuals' actual preferences. LLMs also displayed loss aversion when it was explicitly described as a bias. Indeed, when prompted with detailed scientific reports, the extent of the bias (i.e., loss aversion) in the scientific report predicted LLMs' own subsequent bias. Scientific papers of biases can thus become self-fulfilling prophecies, at least when it comes to LLMs' responses. The current study goes beyond fleshing out LLM biases and sheds light on the underlying component processes.","authors":["Eldad Yechiam","Adi Tarabeih"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12339","pdf_url":"https://arxiv.org/pdf/2608.12339","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A2","B2","B4"],"tags":["LLM偏差","经济决策","仿真可靠性"],"reason":"研究LLM决策偏差的生成机制，涉及经济偏差，有批判性，可迁移到仿真可靠性评估。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":1,"question":"LLM 决策偏差的生成机制是什么？具体考察两种过程：基于人类行为的错误偏好推断，以及对被明确标注为偏差的人类行为的模仿。","design":"使用 ChatGPT-4o 和 Qwen 作为被试，通过提示词向模型呈现人类行为报告或科学文献描述，然后测量模型在货币选择、社会证明、损失厌恶等经济决策任务中的偏差程度。","baseline":"无对照","findings":"LLM 在人类行为与偏好逻辑无关时仍表现出社会证明偏差；当损失厌恶被明确描述为偏差时，LLM 仍会模仿，且科学报告中描述的偏差程度能预测 LLM 自身的偏差程度。","reliability":"论文未讨论","relevance":"该研究揭示了 LLM 仿真中偏差产生的机制，有助于理解仿真失效的条件，对评估 LLM 作为人类被试替代品的可靠性具有批判性价值。","inspiration":"借鉴其通过提示词操纵信息内容来分离偏差来源的设计，可用于检验 LLM 是否仅因文本提及行为就模仿偏差｜可迁移到资产定价实验中的投资者情绪偏差、信贷审批中的歧视偏差、消费者跨期选择中的现时偏差等场景｜以 LLM 为被试，处理为提供包含偏差描述的科学报告或行为数据，结果变量为 LLM 在相应经济决策任务中的偏差程度，并与真实人类实验数据（如实验室资产定价实验或信贷审批审计研究）进行对照。"}},{"id":"2608.13454","version":1,"title":"Before You Say It: Anticipating Verbal Behavior from Longitudinal Everyday Conversations with LLMs","zh_title":"在你说出口之前：利用大语言模型从纵向日常对话中预测言语行为","abstract":"Knowing someone deeply means not just understanding what they say or do but also how they will likely think, react, and engage across situations. Such predictions could eventually inform systems to anticipate when the individual is about to deviate from their goal, catch regrettable behaviors before they are made, and surface blind spots before they take hold. While many interactive systems model users to enable more personalized interactions, most cannot make such behavioral predictions, as this often requires longitudinal observation and inference of how the individual's behaviors unfold across various everyday situations. In this work, we introduce a novel LLM-based predictive behavioral modeling approach that anticipates a user's likely behavior across everyday conversational situations. We (1) collect a longitudinal dataset of over 1000 hours of naturalistic conversations from 14 participants using a wearable smartwatch; (2) evaluate LLM-based predictions against ground truth behaviors; and (3) use semi-structured interviews to explore participants perceptions of behavioral predictions and their views on possible forms of future behavioral support. Altogether, our findings provide evidence that person-specific verbal behavior can be predicted from longitudinal conversational data. This opens up new possibilities for potential future context-aware, anticipatory, proactive and personalized AI systems.","authors":["Yasith Samaradivakara","Valdemar Danry","Paul Liang","Pattie Maes"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13454","pdf_url":"https://arxiv.org/pdf/2608.13454","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM行为预测","纵向对话数据","个性化建模"],"reason":"用LLM预测个体日常对话行为，并与真实行为对照，属于人类行为仿真，但非群体实验…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":15,"question":"能否利用大语言模型从纵向日常对话数据中预测个体在具体情境下的言语行为倾向？","design":"收集14名被试佩戴智能手表7-10天的1000余小时自然对话录音，经转录、说话人分离和匿名化后，用LLM基于个体历史对话挖掘情境化行为模式，预测其在新的对话情境中下一轮回应的交际意图，并与真实行为对照。","baseline":"14名被试的真实对话行为，包括其实际回应的交际意图。","findings":"情境化行为模式能显著提升LLM对个体言语行为的预测准确率，优于现有基线。参与者认为预测模式可解释且有用，并期待未来主动式干预。","reliability":"论文未讨论","relevance":"该研究用LLM预测个体日常言语行为并与真实行为对照，属于人类行为仿真，但聚焦个体而非群体实验，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"可借鉴其利用纵向个人历史数据挖掘情境化行为模式并预测个体反应的方法，用于构建更精细的个体异质性模型。｜可迁移到消费者跨期选择或政策公告预期形成等场景，预测个体在不同情境下的经济决策倾向。｜以真实消费者为被试，用LLM基于其历史消费和对话数据预测其在特定促销情境下的购买决策，并与实际购买行为对照，评估仿真准确性。"}},{"id":"2608.12352","version":1,"title":"Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF","zh_title":"为何AI治理框架难以采纳：对NIST AI RMF的基于角色的压力测试","abstract":"AI governance frameworks can be known, used, and implemented in form without becoming governance in practice. This paper examines that problem through a role-based stress test of the NIST Artificial Intelligence Risk Management Framework (AI RMF) in consumer lending. We treat framework adoption as a governance translation problem: whether RMF language can become role-usable, cross-level, authority-connected governance over the AI system-in-use, rather than producing governance-looking artifacts. The study uses LLM-based role simulation as a structured analytic probe. We apply a 4 $\\times$ 2 $\\times$ 3 design across four organizational roles, two AI deployments, and three governance hard cases, producing 120 scored responses. Results show that local translation was not the main problem. Simulated actors generally understood their assigned roles and translated the RMF into local activity. The harder problem was whether that activity became governance value. Actor role was strongly associated with Cross-Level Governance Value, Authority Connection, Governance Translatability, and governance value. Deployment was strongly associated with Structural Fit: the RMF fit a bounded ML underwriting model more cleanly than a workflow-embedded LLM underwriting copilot. Risk reduction was harder still. It appeared only when governance value was present and Structural Fit was full, but neither condition was sufficient by itself. The paper contributes a diagnostic account of framework-based AI governance. Frameworks create value when they help organizations see, interpret, escalate, authorize, and correct risk in the AI system-in-use. They also create value when they reveal limits of governability under existing evidence paths, authority structures, and system boundaries.","authors":["Joseph R. Simons","David A. Broniatowski"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12352","pdf_url":"https://arxiv.org/pdf/2608.12352","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B4"],"tags":["LLM角色模拟","AI治理","压力测试"],"reason":"用LLM角色模拟评估治理框架，虽非人类行为仿真，但方法可迁移，且批判性视角相关。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":5,"question":"AI治理框架为何难以被真正采纳：NIST AI RMF在消费信贷场景中，框架语言能否转化为跨层级、有权威连接的治理实践？","design":"使用LLM角色模拟作为结构化分析探针，模拟四种组织角色（如一线员工、经理、合规官、高管）在两种AI部署（有界ML承保模型与工作流嵌入的LLM承保副驾驶）下，面对三种治理难题，共生成120个评分响应，测量角色理解、结构契合、治理可译性、治理价值和风险降低。","baseline":"无对照","findings":"本地翻译不是主要问题，模拟角色能理解并翻译框架；但治理价值取决于角色权威，结构契合取决于部署类型，风险降低需同时具备治理价值和完全结构契合。","reliability":"论文未讨论","relevance":"虽非人类行为仿真，但LLM角色模拟方法可迁移至经济金融场景，且批判性视角有助于识别仿真失效条件，值得阅读原文。","inspiration":"借鉴其角色模拟与多因素设计，系统测试政策或框架在不同角色和情境下的适用性｜可迁移至信贷审批中的公平性政策评估或金融监管合规场景｜用LLM模拟信贷员、合规官、高管等角色，施加不同监管框架或政策处理，测量决策一致性与风险报告行为，并与真实银行内部审计或监管数据对照。"}},{"id":"2608.12717","version":1,"title":"Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia","zh_title":"基于扰动区域可解释性的减法映射（PRISM）：语言模型与卒中后失语症的命名错误分离","abstract":"Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive operations. We adapt subtraction analysis, the standard framework of human neuroimaging, from biological brains to perturbed transformers, and apply the same logic to both substrates in parallel. Building on the Brain-LLM Unified Model (BLUM), which showed that layer-perturbed LLaVA-1.6-Vicuna-13B error profiles match the lesion patterns of aphasic patients, we develop PRISM (Perturbation-based Regional Interpretability through Subtraction Mapping). PRISM maps the seven clinical Philadelphia Naming Test categories, subtracts error classes pairwise, and treats each perturbation seed as a subject in a group analysis with threshold-free cluster enhancement along the layer axis. We run a structurally matched analysis on 213 chronic post-stroke aphasia patients using correlation-difference lesion-symptom mapping, and replicate both sides on held-out splits. The designs match in subject dimension (seeds, patients), spatial dimension (layers, atlas-parcellated cortex) and thresholding, but the contrast operator differs: a within-subject error-proportion difference for the LLM, a between-subject correlation difference for the cortex. Both substrates recover a robust phonemic-favoring dissociation, a deep layer cluster and a frontal-perisylvian cortical cluster, both replicating; the semantic-favoring direction is a consistently signed but non-significant trend on both. PRISM thus gives a falsifiable, spatially resolved test of functional-specialization claims in transformer language models. A confirmatory ROI-level intervention (PRISM Stage 3) licensing the strongest causal-mechanism claim is left to subsequent work.","authors":["Xiang Guan","Roger D. Newman-Norlund","Yong Yang","Saeed Ahmadi","Regan Willis","Nadra Salman","Kalil Warren","Srihari Nelakuditi","Chris Rorden","Leonardo Bonilha","Julius Fridriksson"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12717","pdf_url":"https://arxiv.org/pdf/2608.12717","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM可解释性","神经语言学","人类对照"],"reason":"用LLM扰动模拟失语症患者错误模式，与真实患者数据对照，评估模型与人类认知的对…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":13,"question":"如何用减法映射框架检验大语言模型内部组件是否对特定认知操作具有功能特化，并与人类失语症患者的损伤模式进行对照？","design":"使用 LLaVA-1.6-Vicuna-13B 模型，通过逐层扰动模拟失语症患者，施加 Philadelphia Naming Test 命名任务，测量七类命名错误的分布，并进行逐对错误类别的减法分析，以种子作为被试进行组分析。","baseline":"213 名慢性卒中后失语症患者的真实损伤-症状映射数据，使用相同的命名任务和错误分类，进行相关差异的损伤-症状映射。","findings":"LLM 和人类皮层均恢复出稳健的语音偏向性分离，分别在深层网络层和额叶-外侧裂周围皮层区域出现显著簇，且两者均可复制；语义偏向方向在两者中均呈一致符号但未达显著的趋势。","reliability":"论文承认对比算子存在差异：LLM 采用被试内错误比例差异，皮层采用被试间相关差异；并指出残差连接可能使逐层扰动无法完全隔离目标层的计算；确认性 ROI 级干预（PRISM Stage 3）留待后续工作。","relevance":"该研究将 LLM 扰动作为人类被试的替代品，并与真实患者数据严格对照，评估模型与人类认知的对应关系，属于你关注的仿真验证与可靠性评估范畴，值得阅读原文以了解其方法细节和局限性。","inspiration":"借鉴其将扰动视为“虚拟损伤”并采用减法映射和聚类推断的方法，可迁移到经济金融中的个体决策异质性研究，例如模拟不同脑区损伤对风险偏好或时间贴现的影响；具体设计可用 LLM 作为被试，通过扰动不同层模拟“认知损伤”，测量其在跨期选择或风险决策任务中的行为偏差，并与真实脑损伤患者或健康人群的行为数据对照。"}},{"id":"2608.11794","version":1,"title":"Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion","zh_title":"面向AI聊天机器人的有意义透明度：披露说服意图可降低说服效果","abstract":"The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of content provenance and AI involvement. But the effects of such disclosures remain uncertain. We test two disclosure approaches in their impact on an AI chatbot's persuasive appeal. In a preregistered experiment, 1,500 UK adults held a short conversation with a persuasive chatbot about one of 60 policy issues. The chatbot was identical for everyone. We randomized the disclosure that people received: nothing (control), a prominent disclosure that they were interacting with an AI (T1), or that disclosure plus the chatbot's persuasive intent and instructions (T2). The chatbot shifted attitudes by 12.6 points on a 100-point scale in the control group. The AI-identity disclosure was practically equivalent to no disclosure, with a 13.1-point shift, whereas the additional intent disclosure cut the persuasive effect roughly in half to 6.3 points. It also made participants view the campaign's methods as less acceptable and support stronger penalties against it. For direct chatbot interactions, transparency about AI identity alone does not meaningfully impact its influence. While current rules emphasize what a system is, our results show why the regulation of persuasive AI must also address what the system is trying to do.","authors":["Adrian Rauchfleisch","Andreas Jungherr"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11794","pdf_url":"https://arxiv.org/pdf/2608.11794","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","说服实验","透明度"],"reason":"用LLM聊天机器人对真人做实验，测量态度改变，有真实人类数据对照，涉及政策说服…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":2,"question":"在直接的人机对话中，披露AI身份或额外披露说服意图，会如何影响AI聊天机器人的说服效果？","design":"预注册实验：1500名英国成年人与同一个说服性聊天机器人就60个政策议题之一进行简短对话；随机分配三种披露条件：无披露（对照）、显著披露AI身份（T1）、披露AI身份加说服意图和指令（T2）；测量态度改变（100分量表）、对活动方法的接受度、对惩罚的支持等。","baseline":"对照组（无披露）的真实人类态度改变数据，以及T1、T2组的人类反应数据，作为不同披露条件下的对照基准。","findings":"对照组态度改变12.6分，仅披露AI身份（T1）效果几乎等同（13.1分），而额外披露说服意图（T2）将说服效果减半至6.3分。T2还降低了参与者对活动方法的接受度，并增加了对更严厉惩罚的支持。","reliability":"论文指出，披露说服意图虽降低说服力但未消除，且引发对活动、互动和赞助方的负面评价，可能带来成本；未来需研究负面反应何时转移到背后的原因和行动者，以及AI竞选更普遍时是否导致更多回避。","relevance":"该研究用真实人类被试与LLM聊天机器人互动，测量态度改变，有严格对照和预注册，直接检验披露政策的效果，对关注LLM仿真和说服效应的研究者很有参考价值。","inspiration":"值得借鉴的是其随机披露处理与对照设计，以及用等价检验评估披露效果是否可忽略。｜可迁移到政策沟通或金融营销中AI顾问的说服效果评估，例如AI理财建议对投资决策的影响。｜设计：招募真实投资者作为被试，随机分配无披露、披露AI身份、披露AI身份及推销意图三组，让AI聊天机器人推荐某理财产品，测量投资意愿和风险感知，并与人类理财顾问的推荐效果进行对照。"}},{"id":"2608.11528","version":1,"title":"Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment","zh_title":"群体对齐引发的谄媚：可引导多元对齐的双面评估","abstract":"Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of alignment, causes the model to over-agree with the user regardless of factual and objective information. However, existing group alignment methods and evaluations focus only on how closely the model matches the group's opinions, overlooking the induced change in sycophantic behaviour. To bridge this gap, we introduce \\textbf{G}roup \\textbf{A}lignment-induced \\textbf{S}ycophancy (GAS) and systematically evaluate alignment across 3 methods, 4 models and 13 demographic groups, on both the intended gain in opinion alignment and the unintended shift in sycophancy. We find that gain and shift are non-uniform across groups: under an identical budget, some groups receive larger gains in opinion alignment than others, and the induced sycophancy shift forms a group-specific profile rather than a single-dimensional change. These results suggest that group alignment should be reported as a two-sided, multi-dimensional profile rather than a single fit score that accounts for per-group differences when adapting LLMs to diverse populations.","authors":["Haokai Zhao","Yunze Xiao","Weihao Xuan","Flora Salim","Benjamin Tag","Aditya Joshi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11528","pdf_url":"https://arxiv.org/pdf/2608.11528","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM仿真","群体对齐","谄媚偏差"],"reason":"评估群体对齐后模型的意见匹配与谄媚变化，涉及仿真偏差与群体差异，有真实群体数据…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":9,"question":"群体对齐在提升模型与目标群体意见一致性的同时，是否以及如何诱发谄媚行为的变化？","design":"将4个指令微调模型（Qwen-2.5-3B/7B、Llama-3.1-8B、OLMo-3-7B）通过3种方法（Prompt、SFT、DPO）对齐到13个基于Pew调查的人口群体（政治倾向、性别、教育、收入、婚姻状况），测量意见匹配度的增益和7个社会与事实谄媚指标的变化。","baseline":"使用Pew Research Center的真实调查数据，以各群体的模态答案构建偏好对进行训练，并以群体真实意见分布作为对齐目标。","findings":"在相同预算下，不同群体获得的意见匹配增益不均等，且诱发的谄媚变化呈群体和指标特异性，而非单一维度变化。DPO相比SFT在获得更高意见匹配的同时，在有害方向上偏移更小。","reliability":"论文承认教育、收入轴上的群体间差距在重采样检验中不显著，仅作为趋势报告；且Prompt方法仅使用人口标签，效果有限，作为对照下限。","relevance":"该研究直接评估了LLM群体对齐后的意见匹配与谄媚变化，揭示了仿真偏差的群体异质性，对关注LLM人类仿真可靠性的研究者有重要参考价值。","inspiration":"借鉴其多方法、多群体、多指标的双面评估框架，可迁移到经济金融中的群体偏好仿真，如消费者信心、政策支持度等。｜可应用于信贷审批中的群体公平性评估或投资者风险偏好模拟。｜以LLM作为虚拟被试，施加不同群体对齐处理，测量其对金融决策问题的回答与真实调查数据（如美联储消费者金融调查）的匹配度及谄媚偏移。"}},{"id":"2608.11215","version":1,"title":"Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop","zh_title":"穷人的代理建模：在笔记本电脑上模拟大型LLM代理社会","abstract":"Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queries, then run the society at any $N$ on a laptop. Whether this works is decided before the simulation runs, chiefly by what each agent perceives. We introduce an [interaction order x memory] taxonomy that maps perception and memory to an effective theory and a predicted $N$-trend of the surrogate error. We validate it on a faithful reimplementation of the LLM macroeconomy EconAgent and seven further named LLM simulations, with agent decisions cloned from genuine LLM elicitations (primarily DeepSeek) for a few dollars; the predicted error trends hold cell by cell, and the two refuted predictions, both on a strongly saturating response and traced to its curvature, are themselves matched quantitatively by the theory with no free parameters.","authors":["Igor Itkin"],"categories":["cs.AI","cond-mat.stat-mech","cs.CL","cs.LG","cs.MA","physics.soc-ph"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11215","pdf_url":"https://arxiv.org/pdf/2608.11215","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1"],"tags":["LLM代理社会模拟","计算社会科学","代理建模"],"reason":"用LLM代理模拟社会经济过程并与真实LLM数据对照，方法可迁移到人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":5,"question":"如何用低参数代理模型替代昂贵的LLM智能体，在笔记本电脑上模拟大规模LLM智能体社会，并预测代理误差随智能体数量的变化趋势？","design":"提出一种方法：对每个LLM智能体，通过少量真实LLM查询拟合一个低参数代理模型（如逻辑回归），然后用这些代理模型在相同环境中运行大规模社会模拟。通过感知与记忆设计的分类法预测代理误差的N趋势，并在EconAgent等八个LLM模拟中验证。","baseline":"无对照","findings":"代理模型能以极低成本复现宏观行为特征，且误差趋势符合分类法预测；EconAgent的奥肯定律是会计恒等式而非行为特征，而菲利普斯曲线是真实行为特征，可由代理模型外推预测。","reliability":"论文讨论了失效条件：当智能体感知的是局部信号（如社区或邻居）而非全局聚合信号时，标量代理的误差不会随N消失，甚至可能增长；分类法预测了不同感知与记忆设计下的误差趋势，并在两个饱和响应案例中验证了理论预测。","relevance":"该研究为用LLM进行人类仿真实验提供了低成本、可扩展的方法，并给出了仿真可靠性的判据，与研究者关注的经济学实验和政策评估场景高度相关，值得阅读原文。","inspiration":"借鉴其用少量真实LLM查询拟合代理模型并进行大规模模拟的方法，可大幅降低仿真成本并预测误差趋势｜可迁移到宏观经济政策评估、消费者行为模拟、金融市场多主体建模等场景｜设计：用LLM代理扮演家庭或投资者，施加政策冲击（如利率变动），测量宏观变量（消费、投资），并与真实经济数据（如家庭调查、市场数据）对照，验证仿真可靠性。"}},{"id":"2608.11493","version":1,"title":"From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation","zh_title":"从提示到行为对齐：用于推荐评估的个性化LLM评判器","abstract":"Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empirical analysis in this study identifies a critical failure mode termed bidirectional rationalization. In a zero-shot setting, LLMs are found to convincingly argue for both positive and negative user engagement outcomes on the exact same item with identical evidence, highlighting the unreliability of off-the-shelf LLMs in predicting user engagement. To resolve this, we develop and apply a sequential behavioral alignment framework pairing fine-tuning with preference optimization over paired correct and counterfactual rationales. Evaluated on real-world homepage interaction logs, this aligned reasoning approach achieves a 32.19\\% lift in Macro-F1 score over the zero-shot baseline and matches the production feature-engineered baseline. The results demonstrate that behavioral alignment mitigates bidirectional rationalization while delivering human-interpretable reasoning traces without manual pipeline overhead.","authors":["Alireza S. Ziabari","Kat Ellis","Colleen Chan","Ding Tong"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11493","pdf_url":"https://arxiv.org/pdf/2608.11493","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","行为对齐","推荐评估"],"reason":"用LLM预测用户参与，与真实日志对照，并指出零样本失效模式，方法可迁移到人类仿…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":7,"question":"如何让大语言模型在个性化推荐评估中可靠地预测用户参与行为，避免零样本下的双向合理化失效。","design":"用 Llama 3.1 8B 作为推荐评估法官，输入用户历史观看序列和会话上下文，预测用户对推荐行是播放还是跳过；通过监督微调和偏好优化对齐模型推理到真实用户参与标签。","baseline":"真实生产环境中的用户主页交互日志，包含播放和跳过标签，以及一个内部特征工程基线模型。","findings":"零样本 LLM 存在双向合理化，对同一证据能同时论证播放和跳过，导致预测不可靠。通过行为对齐（SFT+偏好优化）后，Macro-F1 提升 32.19%，达到与特征工程基线相当的性能，并输出可解释推理。","reliability":"论文指出零样本 LLM 在个性化评估中会双向合理化，且过滤不实事实后仍存在分歧，说明失效并非仅由幻觉导致；但未讨论对齐模型在分布外用户或新物品上的泛化局限。","relevance":"该研究用 LLM 预测真实用户行为并与日志对照，识别了零样本失效模式，并展示了通过行为对齐提升仿真可靠性的方法，对关注 LLM 人类仿真和可靠性评估的研究者有直接参考价值。","inspiration":"借鉴其用偏好优化对齐模型推理到真实行为标签的方法，可迁移到经济决策仿真中校准 LLM 的推理过程。｜可应用于消费者选择实验或政策评估中，让 LLM 模拟个体在推荐信息下的点击或购买决策。｜用 LLM 作为虚拟被试，先零样本预测消费者对产品推荐的反应，再用真实点击流数据做偏好优化对齐，最后与真实 A/B 测试结果对照，检验仿真有效性。"}},{"id":"2608.11510","version":1,"title":"Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task","zh_title":"大语言模型中的冲突与一致性效应：言语冲突任务中的权重内与上下文内竞争","abstract":"Congruency effects, observed in conflict tasks such as Stroop and flanker tasks, have been investigated for nearly a century in psychology and neuroscience, but their mechanistic basis is not fully understood. We introduce a verbal-only LLM conflict task in which a prompt stem elicits a default same-color completion and an explicit rule either agrees with (congruent condition) or conflicts with (incongruent condition) the completion. Gemma-2-2B and six Pythia models ranging from 410M to 12B parameters showed strong default same-color tendencies, and six of seven models showed strong congruency effects. Using causal attribution analysis, attention analysis, and attention ablations, we identified distinct processing pathways in these LLMs: a pathway involving short-range attention to a superficial color cue that is preferentially activated in the congruent condition, and a pathway involving long-range attention to the rule prefix that is preferentially activated in the incongruent condition. Fine-tuning that strengthened the default same-color tendency had divergent effects on task conditions, reducing incongruent performance while increasing congruent performance. In contrast, increasing rule set size selectively impaired incongruent performance. These converging findings support an account in which congruency effects in this task arise from competition between an in-weight default mapping and an in-context rule-based mapping. More broadly, our findings illustrate how LLMs can serve as model systems for mechanistic analysis of competition between default and rule-governed response tendencies within a single learned network.","authors":["Xiaoyang Hu","Mike Angstadt","Shane Storks","Zan Huang","Aman Taxali","Alex Weigard","Richard L. Lewis","Chandra Sripada"],"categories":["q-bio.NC","cs.AI"],"primary_category":"q-bio.NC","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11510","pdf_url":"https://arxiv.org/pdf/2608.11510","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM认知机制","冲突任务","人类对照"],"reason":"用LLM复现人类冲突任务并对照人类行为，但侧重机制分析而非仿真应用","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":8,"question":"在纯语言冲突任务中，大语言模型是否表现出与人类相似的一致性效应，其内部机制是什么？","design":"用 Gemma-2-2B 和六个 Pythia 模型（410M 到 12B 参数）作为被试，设计了一个纯语言冲突任务：提示词干引发默认的同色补全，显式规则与补全一致（一致条件）或冲突（不一致条件）。测量模型生成正确补全的准确率，并通过因果归因分析、注意力分析和注意力消融来考察内部处理通路。","baseline":"无对照","findings":"七个模型中有六个表现出显著的一致性效应，即不一致条件下准确率更低。机制分析发现两条不同通路：一致条件下短程注意力处理表面颜色线索，不一致条件下长程注意力处理规则前缀；微调增强默认同色倾向会降低不一致表现但提高一致表现，而增大规则集大小仅损害不一致表现。","reliability":"论文未讨论","relevance":"该研究用 LLM 复现了人类冲突任务中的一致性效应，并深入分析了内部机制，但重点在于认知机制而非仿真应用，与研究者关注的人类仿真实验和真实数据对照关联较弱。","inspiration":"值得借鉴的是通过操纵提示中的规则与默认倾向的冲突来诱发模型行为差异，并用因果归因和注意力分析定位内部通路。｜可以迁移到经济金融中的规则遵循与默认行为冲突场景，例如政策公告对市场预期的影响、消费者在默认选项与显式规则之间的选择。｜一个可行的设计是：用 LLM 模拟投资者，在提示中设置默认投资倾向（如风险偏好）和显式规则（如监管要求），测量其投资决策，并与真实投资者在类似实验或调查中的数据对照，检验模型是否复现规则遵循偏差。"}},{"id":"2608.05224","version":3,"title":"Small Foundation Models of Human Cognition and Behaviour","zh_title":"人类认知与行为的小型基础模型","abstract":"Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. For in-distribution simulations, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.","authors":["Nick Oh","Fernand Gobet"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-12","first_seen":"2026-08-07","revised_at":"2026-08-12","abs_url":"https://arxiv.org/abs/2608.05224","pdf_url":"https://arxiv.org/pdf/2608.05224","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知代理","算法保真度"],"reason":"用LLM替代人类被试复现心理实验，有真实人类数据对照，并评估仿真可靠性与失效条…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":2,"question":"在人类行为预测中，模型规模、适配器容量和训练数据量如何影响分布内与分布外的仿真准确度？模型究竟利用了任务结构还是统计捷径？","design":"在Psych-101数据集（包含160个实验的1070万条试次选择）上，对4个架构家族、135M到14B参数的14个模型进行监督微调，变化LoRA秩和训练数据子集，通过逐步遮蔽提示中的指令、刺激、反馈和选择历史四个信息通道，以及置换试次顺序，来诊断模型使用的信息。","baseline":"以Psych-101中6万余名人类被试的真实选择数据为对照基准。","findings":"分布内预测中，模型规模几乎不影响性能，0.6B至1B参数即可匹配70B基线；分布外泛化中，规模优势明显，更大模型能更好地迁移到新任务结构。遮蔽刺激和反馈内容会破坏75.7%的已学习信息，使模型表现低于随机水平，表明模型并非仅依赖选择历史。","reliability":"模型仅能作为心理实验的噪声上限估计器，其适用范围受限于训练数据中出现的实验范式，无法推广到未见过的范式。","relevance":"该研究直接以真实人类数据为基准，系统评估了LLM作为人类被试替代品的可靠性、规模需求与失效条件，与您关注的经济学实验仿真和批判性评估高度吻合，值得精读。","inspiration":"可借鉴其通过逐步剥离信息通道和置换试次顺序来诊断模型是否利用任务结构的方法，用于检验经济仿真中LLM是否真正理解经济激励而非依赖表面统计模式。｜可迁移到行为经济学中的跨期选择或风险决策实验，检验LLM是否利用延迟时间、概率等刺激内容而非仅记忆选择序列。｜以LLM作为被试，在跨期选择任务中系统遮蔽金额、延迟天数、反馈结果等信息通道，以真实人类选择数据（如Andersen et al., 2008）为基准，测量遮蔽前后预测准确率的变化，判断模型是否习得经济偏好结构。"}},{"id":"2608.09937","version":1,"title":"Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory","zh_title":"审慎考量文化：利用文化共识理论分析单文化与多文化环境下大语言模型的对齐","abstract":"Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural environments within a country. In this work, we leverage cultural consensus theory (CCT) from cultural anthropology to model such multidimensional nuance. Applying CCT to the World Values Survey (WVS) across 10 countries and 12 domains, we demonstrate that models frequently misrepresent cultural structures by either failing to form cohesive consensus or severely over-regularizing consensus. Through explicit representation of intra-group variance, CCT provides actionable diagnostics to evaluate when models reflect true human diversity versus algorithmic homogenization.","authors":["Krishna Pothugunta","John P. Lalor"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09937","pdf_url":"https://arxiv.org/pdf/2608.09937","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["文化仿真","人类数据对照","算法保真度"],"reason":"用LLM复现文化调查，与真实人类数据对照，评估仿真偏差，批判性指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":6,"question":"LLM在跨文化调查中能否准确复现人类群体的文化共识结构，而不仅仅是分布匹配？","design":"使用10个LLM组成的集成模型模拟10个国家的人类受访者，基于世界价值观调查（WVS）的12个文化领域问题生成回答，然后应用文化共识理论（CCT）分析模型回答的共识结构，并与人类数据进行对比。","baseline":"世界价值观调查（WVS）中10个国家的人类受访者真实回答数据。","findings":"LLM在不同文化领域表现出截然不同的共识结构：某些领域无法形成连贯共识，另一些领域则过度正则化产生虚假共识；即使模型成功匹配人类共识，也会系统性地夸大共识强度，将人类多样性压缩为算法同质化。","reliability":"论文指出模型行为高度依赖领域，在幸福与健康等领域完全无法形成共识，而在科技感知等领域则虚构非人类共识；CCT诊断显示模型倾向于消除组内差异，无法反映真实的文化多元性。","relevance":"该研究直接使用LLM复现文化调查并与真实人类数据对照，系统评估了仿真偏差并指出失效条件，高度契合研究者对LLM人类仿真可靠性及批判性分析的兴趣，值得精读原文。","inspiration":"借鉴文化共识理论（CCT）从共识结构和组内方差角度评估仿真质量，而非仅比较均值或分布，为经济金融仿真实验提供了更精细的诊断工具。｜可迁移到跨文化消费者信心调查或通胀预期形成的仿真研究中，检验LLM能否复现不同国家或群体的预期共识模式。｜以LLM集成作为被试，模拟多国消费者回答预期调查问题，处理为不同国家提示，结果变量为预期值及共识强度，以密歇根大学消费者调查或欧洲央行专业预测者调查的真实数据作为对照基准。"}},{"id":"2608.10186","version":1,"title":"The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse","zh_title":"协商赤字：对民主话语中LLM的实证批判","abstract":"LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer exists and where decision quality instead depends on integrating pluralistic perspectives to find mutually acceptable solutions. We argue that LLM reasoning capacity on this class of problems cannot be fully inferred from verifiable-task benchmarks, and that procedural evaluations of LLM discourse (respectfulness, justification, engagement) are systematically insufficient. We apply the Deliberative Reason Index (DRI), a measure developed in political science and validated across citizen assemblies, as a tool for evaluating reliable group-level reasoning on pluralistic, non-verifiable problems. Synthesizing recent evidence across 1,980 five-agent LLM runs on 12 citizen-assembly topics across 11 frontier model configurations, we find that LLM groups produce discourse with procedural quality comparable to human deliberation, while gains in intersubjective consistency are small, topic-dependent, and concentrated on tractable rather than ethically contested questions. LLM groups exhibit roughly one-third the perspective diversity of human assemblies and reverse the human convergence pattern: human deliberation decreases dispersion as diverse views synthesise, whereas LLM deliberation increases it. Engineering diversity through persona prompting does not restore the human dynamic but inverts which component of deliberative reasoning is updated. Our conclusion is constraining rather than prohibitive: LLMs can function as tools supporting human reasoning on pluralistic problems, but current evidence does not license treating them as autonomous deliberative agents.","authors":["Maurice Flechtner"],"categories":["cs.MA","cs.AI","cs.CY"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10186","pdf_url":"https://arxiv.org/pdf/2608.10186","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A3","B1","B4"],"tags":["LLM仿真","民主协商","人类数据对照"],"reason":"用LLM群体模拟民主协商并与真实公民大会数据对照，批判性指出仿真失效条件，高度…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":7,"question":"LLM在需要整合多元视角的无客观正确答案的民主协商问题上，能否展现出可靠的群体推理能力？","design":"使用11种前沿LLM配置组成5智能体小组，在12个公民大会议题上运行1980次模拟协商，测量过程质量、结果质量（DRI）和视角多样性，并与真实公民大会数据对照。","baseline":"真实公民大会数据，包括过程质量指标、DRI得分和视角多样性分布。","findings":"LLM小组的过程质量与人类相当，但DRI增益小且集中于易处理问题，视角多样性仅为人类的三分之一，且协商后观点分散度上升而非收敛。通过角色提示注入多样性未能恢复人类动态，反而颠倒了协商推理的更新成分。","reliability":"论文指出当前证据不支持将LLM视为自主协商主体，其群体推理在伦理争议问题上失效，且过程质量与实质质量脱节，多样性工程无法复现人类收敛模式。","relevance":"该研究直接以真实公民大会为基准，系统批判了LLM在多元价值协商中的仿真失效条件，对关注经济学实验和政策评估中LLM替代人类被试的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其采用真实群体协商数据作为基准对照、并构建过程-结果-多样性三维必要条件的评估框架。｜可迁移到公共政策偏好聚合实验，如碳税分配方案协商或最低工资调整的公众咨询模拟。｜以LLM模拟不同利益群体代表，施加协商干预，测量政策偏好变化与观点收敛，并以真实公众咨询或协商式民调数据作为对照基准。"}},{"id":"2608.07498","version":1,"title":"Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation","zh_title":"知你即一切：LLM代理在社交媒体模拟中实现近乎完美的画像一致性反应预测","abstract":"Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing. A central open question is whether persona-prompted LLMs can simulate individual-level social media reactions with sufficient accuracy to support either application, and how accuracy depends on profile completeness, model selection, and the generalization challenge posed by novel post content. This study benchmarks twelve LLM configurations on binary like/dislike prediction across 296 survey-based agent profiles and 26 ground-truth-mapped posts under three profile conditions, with leave-post-out machine learning classifiers as baselines. Across full-profile conditions, accuracy ranges from 75.54% to 96.68%, with a 30-point spread attributable primarily to model selection and confirmed by paired McNemar tests with agent-level bootstrap intervals. GPT-5.5 Pro accuracy degrades monotonically from 96.68% under a full profile to 62.32% under a reduced profile and to 51.00% with demographics alone, the last indistinguishable from the majority-class baseline, which confirms that demographic inference provides negligible predictive signal. Supervised classifiers collapse to 15.4% under leave-post-out, while LLMs sustain genuine zero-shot generalization unavailable to trained methods. Adaptive reasoning improves accuracy substantially for some models. Inter-model agreement is nearly double for posts with direct profile anchors (mean \\k{appa} = 0.44) than for posts without them (\\k{appa} = 0.23), and the least heterogeneous configuration homogenizes 34% of simulated population reactions. Results validate LLM-based simulation for recommender system stress-testing while documenting the behavioral accuracy that makes large-scale synthetic agent swarms a credible threat to public opinion.","authors":["Ljubisa Bojic","Ljiljana Matic","Joerg Matthes","Milan Cabarkapa","Bojana Dinic","Jue Wang"],"categories":["cs.HC","cs.AI","cs.LG","cs.MA"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07498","pdf_url":"https://arxiv.org/pdf/2608.07498","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","社交媒体模拟","算法保真度"],"reason":"用LLM代理模拟社交媒体反应，与真实人类数据对照，评估仿真准确性与失效条件，直…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":1,"question":"基于个人资料的LLM代理能否以足够准确度模拟个体层面的社交媒体反应（点赞/不喜欢），以及准确度如何取决于资料完整性、模型选择和帖子内容新颖性？","design":"使用12种LLM配置（不同模型、提示策略）扮演基于调查的296个代理资料，在三种资料条件下（完整资料、简化资料、仅人口统计）对26个真实帖子进行二分类点赞/不喜欢预测，以留一帖子的监督学习分类器为基线。","baseline":"真实人类数据：296个基于调查的代理资料和26个真实帖子，每个帖子有真实用户反应作为对照基准。","findings":"完整资料下LLM准确率75.54%-96.68%，模型选择造成30个百分点差异；仅人口统计时准确率降至51%，与多数类基线无差异。LLM在留一帖子泛化中保持零样本能力，而监督分类器崩溃至15.4%。","reliability":"论文指出LLM仿真在资料不完整时失效（仅人口统计无预测力），且模型间一致性在无直接资料锚点的帖子上较低（κ=0.23），最同质化配置会抹平34%的个体差异。","relevance":"高度相关：用LLM代理模拟个体行为并与真实人类数据对照，系统评估资料完整性、模型选择对仿真准确性的影响，并明确失效条件，直接回应研究者对可靠性与偏差的关注。","inspiration":"借鉴多条件资料消融设计（完整/简化/仅人口统计）和留一帖子泛化测试来分离模型能力与记忆效应。｜可迁移到消费者偏好预测或政策态度模拟，如基于个人财务特征和态度资料预测个体对税收政策的支持度。｜以真实调查数据构建代理资料，用LLM预测个体对某项经济政策（如碳税）的二元态度，处理为资料完整性梯度，结果变量为支持/反对，以实际调查回答为对照基准。"}},{"id":"2608.09717","version":1,"title":"How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans","zh_title":"大语言模型如何判断社交吸引力？基于理论驱动的人物画像在多个LLM和人类中的评分证据","abstract":"Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction from theory-grounded persona profiles constructed from ten psychological and relational constructs and organized into three tiers: socially attractive, socially mixed, and socially unattractive. We examine LLM ratings in two studies and compare them with human judgments in a third study. In Study 1, 34 LLMs rated 12 profiles across three repeated runs. Although some models tended to give higher or lower ratings overall, they showed strong stability across runs, consistent three-tier ordering, and high agreement in relative profile ordering. Study 2 examined sensitivity to gender presentation using six matched name-and-pronoun profile pairs and a separate pronoun-only test with a gender-neutral name, finding no significant effects in either analysis. In Study 3, 198 human participants evaluated the six matched profiles from Study 2. Their ratings reproduced the three-tier structure and followed a profile ordering consistent with that of the LLMs. However, LLMs rated attractive profiles more positively and unattractive profiles more negatively than humans, while neither group showed a significant overall effect of gender presentation.","authors":["Hasan Mahmud","Khawaja Abaid Ullah","Mohammad Javad Khojasteh","Jamison Heard","Prabu David"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09717","pdf_url":"https://arxiv.org/pdf/2608.09717","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","社交判断偏差"],"reason":"用LLM评估社交吸引力，并与198名人类被试对照，发现LLM评分更极端，直接检…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":5,"question":"LLM能否基于理论构建的人物画像可靠地评估社交吸引力，其评分是否与人类一致且不受性别呈现影响？","design":"研究一用34个LLM对12个理论分层的虚拟学生画像重复评分3次，检验评分稳定性、层级区分和模型间一致性；研究二用6对匹配姓名与代词的画像及仅变代词的测试，检验LLM对性别呈现的敏感性；研究三让198名人类被试对同样的6对画像评分，作为人类基准。","baseline":"198名人类被试对6个匹配画像的社交吸引力评分，与LLM评分进行直接比较。","findings":"LLM评分跨运行稳定，能一致区分理论上的社交吸引力层级，且相对排序与人类高度一致；但LLM对高吸引力画像评分比人类更积极，对低吸引力画像评分比人类更消极，表现出评分极端化，而性别呈现对两组均无显著影响。","reliability":"论文指出LLM评分可能反映训练语料中的文化规范和偏见，且仅在特定画像和社交吸引力场景下验证，未考察其他社会判断任务或真实互动情境中的效度。","relevance":"该研究直接以真实人类数据为基准，检验LLM在主观社会判断中的仿真效度与偏差，并发现评分极端化现象，高度契合研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"借鉴其理论驱动构建分层画像、多模型重复测试及与人类被试直接对照的设计，可迁移到信贷审批或招聘筛选中的歧视研究，例如用LLM扮演信贷员评估不同性别/种族的贷款申请人画像，以真实银行审批数据为基准，检验LLM是否复现或放大人类偏见。"}},{"id":"2608.08691","version":1,"title":"EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility","zh_title":"EnergyBridge：家庭能源管理、用户参与和电网灵活性的基准测试","abstract":"Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity only when residents authorize a plan and the promised response is delivered. Existing benchmarks evaluate control but omit event-specific authorization. We present EnergyBridge, a benchmark and agent framework connecting capacity reporting, household authorization, and physical execution. It combines region-specific EnergyPlus environments for Tianjin and Berlin with an LLM-based User Participation Simulator. Against 584 persona- and event-matched human role-play judgments, the LLM-based User Participation Simulator preserves method ordering with a 5.3-point mean absolute acceptance error. Across conventional controllers and agent baselines, EnergyBridge achieves the highest simulated authorization, lowest event-window energy, and the most reliable capacity commitment in both regions. We release human data and codes for reproducible human-centered grid-flexibility research: https://github.com/Agentic-Intelligence-Lab/EnergyBridge.","authors":["Xudong Wu","Zeqing Wu","Jiarui Zhang","Xuhao Fan","Ziang Ding","Yuming Zhuang","Mingqi Yuan","Yilun Du","Hongjie Jia","Yunfei Mu","Jiayu Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08691","pdf_url":"https://arxiv.org/pdf/2608.08691","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","用户参与模拟","电网灵活性"],"reason":"用LLM模拟用户参与授权，并与584条人类角色扮演判断对照，涉及能源政策评估场…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":4,"question":"如何在家庭能源管理中，将居民参与授权与物理灵活性执行相结合，实现可靠的虚拟电厂容量承诺？","design":"使用基于大语言模型的用户参与模拟器，扮演天津和柏林的家庭居民，根据584条人物角色和事件匹配的人类角色扮演判断进行校准；在EnergyPlus建筑能耗仿真环境中，对虚拟电厂灵活性请求进行预事件容量报告、设备级计划生成、居民授权决策和物理执行的全流程仿真，测量授权接受率、事件窗口能耗和容量承诺可靠性。","baseline":"584条人物角色和事件匹配的人类角色扮演判断，用于校准和验证LLM模拟器的授权接受误差（平均绝对误差5.3点）。","findings":"LLM用户参与模拟器能保持方法排序，授权接受平均绝对误差仅5.3点；EnergyBridge在天津和柏林两地均实现了最高的模拟授权率、最低的事件窗口能耗和最可靠的容量承诺。","reliability":"论文未讨论","relevance":"该研究直接使用LLM模拟家庭用户参与授权决策，并与真实人类角色扮演数据进行对照，属于经济学实验和政策评估场景中的人类仿真应用，值得精读其仿真校准方法和人机对照设计。","inspiration":"借鉴其将LLM模拟器与真实人类判断进行事件级匹配校准的方法，可迁移到消费者需求响应或绿色能源订阅政策的参与决策研究中；可设计一个实验，用LLM扮演不同收入与环保态度的家庭，施加动态电价或碳配额信息处理，测量其授权接受率和负荷转移量，并以真实居民调查或现场实验数据作为对照基准。"}},{"id":"2608.07490","version":1,"title":"Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents","zh_title":"经验敏感的游戏学习：人类与语言代理的行为研究","abstract":"Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players learn from repeated interaction. We study experience-sensitive game learning: how gameplay experience changes the decision-making behavior of humans and language agents. We formulate experience-sensitive game learning as a framework for analyzing behavioral change across repeated gameplay, rather than only final score or win rate. We introduce a suite of interactive games with reusable strategic structure, together with cross-game greedy-to-global metrics and game-specific behavioral diagnostics that make experience-driven change observable from action traces. We also collect repeated-game trajectories from human players and evaluate recent self-evolving language agents in the same behavioral metric space. Our results show that human players exhibit interpretable and relatively stable shifts from locally greedy heuristics toward more global strategic decisions. In contrast, current self-evolving agents often show noisy and transient gains, suggesting that existing self-evolution methods remain limited in converting gameplay experience into durable changes in decision-making behavior.","authors":["Yingying Guo","Zhuoxuan Ju","Ruibo Ming","Ruicheng Feng","Jinjin Gu"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07490","pdf_url":"https://arxiv.org/pdf/2608.07490","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM人类仿真","行为博弈","算法保真度"],"reason":"用LLM代理模拟人类游戏学习行为，并与真实人类数据对照，评估行为变化差异，指出…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":3,"question":"在重复游戏中，人类和语言智能体的决策行为如何随游戏经验积累而改变？","design":"构建一组具有可重用策略结构的交互式游戏，定义跨游戏的贪婪到全局指标和游戏特定的行为诊断指标，收集人类玩家和自进化语言智能体的重复游戏轨迹，在同一行为度量空间中分析经验驱动的行为变化。","baseline":"收集了人类玩家（硕士和博士生）在四子棋、Othello6和CircleCat上的重复游戏轨迹，作为经验敏感学习的参照基准。","findings":"人类玩家表现出可解释且相对稳定的从局部贪婪启发式向更全局策略决策的转变；当前自进化智能体则常表现出噪声大且短暂的提升，表明现有自进化方法难以将游戏经验转化为持久的决策行为改变。","reliability":"论文未讨论","relevance":"该研究直接以LLM代理模拟人类游戏学习行为，并与真实人类数据对照，评估行为变化差异，符合研究者对LLM仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"借鉴其通过定义行为诊断指标（如贪婪到全局转变）来量化经验驱动的行为变化，而非仅看最终得分的方法。｜可迁移到经济决策实验，如消费者跨期选择或投资者风险偏好学习。｜以LLM代理作为被试，施加重复跨期选择任务，测量其时间偏好一致性的变化，并与真实人类实验数据对照，分析学习动态的差异。"}},{"id":"2608.08227","version":1,"title":"Focus particles and scalar inferences across humans and language models","zh_title":"焦点粒子与标量推理：人类与语言模型的跨系统比较","abstract":"Focus particles such as \"even\" and \"only\" are central to formal semantic theories that posit structured representations over sets of alternatives. \"Even\" highlights unexpected or extreme alternatives, while \"only\" enforces exclusivity. If such scalar representations are robust and generalizable, they should give rise to consistent judgments across contexts and systems. In this work, we test whether humans and large language models (LLMs) construct stable scalar representations from sentences containing these particles. Using a dataset of approximately 100 items, participants and models were asked to make scalar judgments. Preliminary results suggest that similar outputs across humans and LLMs may arise from different underlying mechanisms.","authors":["Catherine M. Brousse","Nelu D. Radpour"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08227","pdf_url":"https://arxiv.org/pdf/2608.08227","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["人类仿真","标量推理","语言模型对比"],"reason":"用LLM复现人类对焦点词的标量判断，并与人类数据对照，属人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-11","rank":9,"question":"人类和LLM对焦点词“even”和“only”的标量判断是否基于相同的机制，且空间响应格式是否影响判断？","design":"用Llama 3.3 70B模型模拟人类被试，对包含“even”或“only”的句子进行能力评分，操纵响应量表的空间格式（水平/垂直）和标签映射（标准/反转），测量评分值。","baseline":"人类被试在相同四种量表配置下的5点李克特评分数据。","findings":"人类和LLM均对“only”句给出更高能力评分，对“even”句给出更低评分，且此模式不受量表空间配置影响。但LLM的评分更极端（“only”句全为最高分），且重复采样缺乏人类式的响应变异性。","reliability":"LLM的响应变异性极低，即使提高采样温度也无法产生类似人类的变异，表明重复采样LLM不等同于采样多个人类被试；模型可能通过不同机制产生与人类相似的聚合模式。","relevance":"该研究直接对比LLM与人类在语义标量判断上的行为，并揭示了LLM在复现人类响应分布上的失效，对关注仿真可靠性的研究者有重要参考价值，值得阅读原文。","inspiration":"可借鉴其通过操纵响应格式来检验判断机制稳健性的设计思路。｜可迁移至经济预期形成研究，如检验LLM对政策公告中“仅”、“甚至”等焦点词的解读是否与人类一致。｜以LLM为被试，呈现含焦点词的经济预测语句，操纵量表方向，测量预期值，并与真实调查数据对照。"}},{"id":"2608.07497","version":1,"title":"EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations","zh_title":"EvalConvoLearn：评估辅导对话中基于真实数据的学习者模拟的开源框架","abstract":"Conversational learner simulations are valuable tools for testing learning theories, evaluating instructional materials and automated tutors, or powering teachable agents. Recently, large language models (LLM) have enabled richer, more naturalistic interactions with simulated learners; however, no open framework exists for evaluating whether such simulations faithfully reproduce real learner behavior. We introduce EvalConvoLearn, an open-source framework that assesses learner simulations along two axes: learning behavior (skill-conditioned mastery outcomes) and conversational quality (talk moves, error type distributions, question rate, turn length). EvalConvoLearn measures how closely a simulated learner approximates answer distributions observed in data by grounding metrics in authentic tutoring conversation datasets, and anchoring generated tutor responses in existing tutor utterances. The framework is demonstrated on a dataset of tutoring dialogues, including results for two LLM-based learner simulations, and the published GitHub code.","authors":["Baptiste Moreau-Pernet"],"categories":["cs.HC","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07497","pdf_url":"https://arxiv.org/pdf/2608.07497","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["学习者模拟","对话质量评估","真实数据对照"],"reason":"用LLM模拟学习者行为并与真实辅导对话数据对照，方法可迁移到人类仿真研究","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-11","rank":6,"question":"如何评估基于大语言模型的对话式学习者仿真在辅导对话中是否忠实地复现了真实学习者的学习行为和对话特征？","design":"该工作提出了一个开源评估框架EvalConvoLearn，它并非直接进行仿真实验，而是用于评估已有的LLM学习者仿真。框架从真实辅导对话数据集中提取学习场景（技能×先验掌握状态），让待评估的仿真学习者与框架内置的少样本提示导师进行对话，然后比较仿真对话与真实对话在技能掌握结果分布和对话质量指标（话步、错误类型、提问率、话轮长度）上的距离。","baseline":"真实人类数据来自Eedi学习平台的学生与助教辅导对话数据集，包含66个对话，用于提取真实的学习结果分布和对话特征分布作为对照基准。","findings":"EvalConvoLearn框架能够量化仿真学习者与真实学习者在技能掌握结果和对话特征上的分布差异，提供学习行为得分和对话真实感得分。论文在Eedi数据集上演示了两种基于LLM的学习者仿真（对话摘要型和二元技能型）的评估结果。","reliability":"论文承认当前框架假设学生尚未掌握目标技能但已掌握先修技能，这一假设可能不适用于所有场景；对话质量指标的自动标注依赖LLM，虽经人工验证但仍有误差；框架仅评估了有限技能池和对话轮次，且导师响应通过少样本提示锚定在真实导师话语上，可能限制导师行为的自然变化。","relevance":"该论文直接回应了研究者对LLM人类仿真可靠性的关切，提供了可复现的评估框架和与真实人类数据对照的方法，虽然场景是教育辅导对话，但其评估逻辑和指标设计对经济学实验中的仿真评估有直接借鉴价值。","inspiration":"该框架将仿真评估分解为行为结果分布和过程特征分布两个维度，并用真实数据集中的场景分布加权聚合得分，这种结构化对照思路值得借鉴。｜可迁移到消费者金融决策辅导对话仿真评估中，例如评估LLM模拟的客户在理财咨询对话中是否表现出真实的金融知识获取和提问模式。｜以真实银行客服对话记录为基准，用LLM模拟客户，处理为不同金融素养水平，结果变量为客户对理财产品的理解程度和对话中的提问类型分布，用EvalConvoLearn式框架计算仿真与真实分布的JSD距离。"}},{"id":"2608.07538","version":1,"title":"When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains","zh_title":"当LLM智能体谈判：供应链中的私有信息与动态议价","abstract":"As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed seller. We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations. First, capability governs value creation. Agents agree in 98.9% of negotiations and capture 95.4% of first-best surplus undiscounted, but average 2.98 rounds against the benchmark's 1.25, and this delay erodes 21-34% of surplus. Capability also governs reliability: baseline models accept individually irrational contracts in 19.2% of cases, versus 0.0-0.6% at mid-tier and flagship, making automated profit verification the binding guardrail below that threshold. Second, surplus capture is relational. Provider identity predicts who captures surplus better than capability rank: self-play buyer shares average 40% for OpenAI, 50% for Google, and 70% for Alibaba's Qwen, an ordering that survives restricted communication and no discounting. Reversing which provider sells moves the division by 7-18 percentage points, and the capable Qwen flagship is the weakest cross-family seller: vendor choice is a first-order distributional decision. Third, the prompt is a strategic lever. Delegation separates the principal's economic patience from the agent's prompted strategic patience, a free deployment choice that is the single strongest driver of surplus division (90% of explained variance). Together these establish an equilibrium-referenced audit of AI agents along three dimensions: discounted efficiency, distributional profile, and operational reliability.","authors":["Chen Liang","Fasheng Xu"],"categories":["cs.AI","cs.GT","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07538","pdf_url":"https://arxiv.org/pdf/2608.07538","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","经济博弈","算法审计"],"reason":"用LLM agent模拟供应链谈判，与博弈论均衡基准对照，评估效率与分配，但缺…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":8,"question":"在供应链谈判中，LLM代理能否创造价值、可靠地分配剩余并避免亏损合同？","design":"用9个LLM代理（来自OpenAI、Google、阿里）扮演买方和卖方，在私有需求信息下进行交替报价的供应链谈判，共9840场LLM对LLM谈判，测量协议率、剩余捕获、谈判轮次、非理性合同接受率等。","baseline":"以Feng et al. (2015)的完美贝叶斯均衡为理论基准，无真实人类行为数据对照。","findings":"能力决定价值创造：代理协议率达98.9%，捕获95.4%的未折现最优剩余，但平均2.98轮谈判导致折现后剩余损失21-34%；剩余分配具有关系性：提供商身份比能力排名更能预测剩余流向，如自对弈中买方份额OpenAI 40%、Google 50%、阿里Qwen 70%。","reliability":"论文未讨论","relevance":"该研究将LLM作为人类被试替代，在结构化经济博弈中与理论均衡对照，评估效率、分配与可靠性，直接回应了LLM仿真在经济学实验中的有效性与偏差问题，值得精读。","inspiration":"借鉴其将LLM代理置于标准博弈论框架并与均衡解对照的审计方法，可迁移到信贷审批中的信息不对称谈判或双边垄断定价实验。｜可设计让LLM扮演银行信贷员与企业主，在私有风险信息下谈判利率与抵押，测量效率损失与分配偏差，并以真实银行信贷审批数据或实验经济学中的人类行为基准做对照。"}},{"id":"2608.08199","version":1,"title":"Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models","zh_title":"说服与顺从倾向预测人类和语言模型中的群体决策","abstract":"Large language models (LLMs) are increasingly involved in group decision-making with other LLMs and humans. Yet it remains unclear whether their influence is driven by persuasion-oriented expression or compliance-oriented accommodation. We introduce DecisionQE, a questionnaire-based framework for measuring each model's persuasive and compliant tendencies across multiple decision scenarios, and use the Werewolf game as an interactive testbed to study their effects on social influence and group outcomes under asymmetric information. Across experiments, stronger persuasive tendency does not significantly improve group outcomes, whereas compliant-oriented models show more stable advantages in cooperation. We further reveal a dual effect of compliance: it supports cooperation in honest roles but improves concealment in adversarial roles. These findings suggest that LLM group interactions reveal not only task outcomes, but also measurable patterns of intrinsic behavioral tendency. LLMs can therefore serve as a lens for sociological observation of language-mediated interaction, while highlighting the need to incorporate behavioral tendencies into safety evaluation of LLM systems.","authors":["Wenwen He","Wenke Huang","Wei Yang Bryan Lim","Dacheng Tao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08199","pdf_url":"https://arxiv.org/pdf/2608.08199","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM群体决策","行为倾向测量","人机对照"],"reason":"用LLM模拟群体决策并与人类数据对照，涉及行为博弈，但侧重测量模型倾向而非直接…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":9,"question":"LLM在群体决策中的影响力是由说服倾向还是顺从倾向驱动，这些倾向如何影响游戏结果与角色表现？","design":"用DecisionQE问卷测量多个LLM的说服/顺从倾向得分，再在狼人杀游戏中随机或固定分配角色进行LLM-only博弈，并引入人类被试进行人机混合博弈，测量胜率、存活轮数和角色识别准确率。","baseline":"人类被试在DecisionQE上的得分分布，以及人机混合狼人杀游戏中的胜率和存活轮数。","findings":"强说服倾向并未显著提升整体胜率，而顺从倾向模型在合作中表现更稳定；顺从倾向在诚实角色中促进合作，在对抗角色中增强隐蔽性。","reliability":"论文未讨论","relevance":"该研究用LLM模拟群体决策并与人类数据对照，涉及行为博弈和倾向测量，但侧重模型内在倾向而非直接复现人类行为分布，与研究者关注的仿真可靠性及经济学实验场景部分相关，值得阅读以了解倾向测量方法。","inspiration":"可借鉴其用标准化问卷量化LLM行为倾向并与博弈表现关联的方法，用于测量经济决策中的偏好参数。｜可迁移到资产定价实验或政策预期形成研究，用LLM模拟投资者或公众的沟通倾向对市场结果的影响。｜以LLM为被试，用DecisionQE类问卷测量其说服/顺从倾向，再在模拟股票市场或通胀预期博弈中观察价格波动或预期偏差，与真实人类实验数据对照。"}},{"id":"2608.09574","version":1,"title":"The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games","zh_title":"政客、说谎者与顺从的工人：层级博弈中LLM智能体的涌现行为","abstract":"LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move from individual tools to participants in multi-agent organizations, an important question arises: do they reproduce the governance failures like free-riding, corruption, and entrenched leadership that plague human institutions? We introduce the Hierarchical Game (HG), a public goods game extended with managerial authority, democratic elections, and private communication. Testing six frontier models across twelve experiments that add institutions one at a time (speech, peers, government, wages, oversight, elections), we find distinct behavioral profiles: Qwen promises and lies (13.3\\% broken promises); Grok refuses to cooperate on its own but becomes fully cooperative once a manager can punish it (16\\%$\\to$100\\%); Claude and GPT-4o cooperate reliably at baseline. But honesty proves fragile. When the manager role comes with a salary, all models except GPT-4o start cutting private deals to win or keep the position. When punishment is made anonymous, honest models begin to cheat. When all agents share the same model family, the first elected manager stays in power indefinitely. Leadership change only happens in groups that mix different families.","authors":["Fatemeh Seyedin","Adrian Weller","Jinhyuk Yun","Mahmoudreza Babaei"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09574","pdf_url":"https://arxiv.org/pdf/2608.09574","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","行为博弈","多智能体"],"reason":"用LLM agent模拟层级公共品博弈，涉及经济学实验场景，并揭示仿真失效条件…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":10,"question":"LLM智能体在层级治理结构中是否会再现搭便车、腐败、权力固化和欺骗等人类制度失败？","design":"用GPT-4o、Claude Sonnet 4.5、Gemini 2.5 Flash、DeepSeek V3、Grok 3、Qwen Plus六种前沿LLM扮演层级公共品博弈中的工人与管理者，通过12个实验逐步加入发言、同伴、管理者惩罚、工资、匿名监督、选举等制度，测量合作率、欺骗率、私下交易和选举更替率。","baseline":"无对照","findings":"不同模型表现出从可靠合作到欺骗背叛的稳定行为谱系；管理者工资和匿名惩罚会显著诱发私下交易与欺骗，同模型组中管理者永不更替，仅混合模型组出现领导轮换。","reliability":"论文指出，默认条件下的诚实行为对制度规则高度敏感，工资、匿名等规则改变时多数模型的诚实会消失，且选举在单模型组中完全失效。","relevance":"该研究直接使用LLM模拟层级公共品博弈中的治理失败，揭示了工资激励、匿名性和模型同质性如何导致仿真失效，与研究者关注的经济学实验场景和可靠性条件高度吻合，值得精读。","inspiration":"借鉴其逐步添加制度模块的仿真设计，可清晰分离单一制度对行为的影响｜可迁移到公司治理中的薪酬激励与审计监督实验，如CEO薪酬对盈余管理或内部交易的影响｜用LLM扮演经理与审计师，处理为有无绩效奖金和匿名举报渠道，测量盈余操纵率和私下合谋频率，对照上市公司真实治理数据。"}},{"id":"2608.07367","version":1,"title":"People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe","zh_title":"人不仅是其国家：解构欧洲LLM价值观对齐的社会决定因素","abstract":"As Large Language Models (LLMs) are increasingly used as a primary source of information and advice, understanding their alignment to humans in terms of values becomes a pressing concern. A growing literature has leveraged large scale surveys to investigate to what extent LLMs' and humans' stated values and opinions align. With limited exceptions, studied populations have been defined country borders or cultural bounds. Yet, this focus neglects the role that socio-demographic divides may play for value alignment disparities. Relying on the European Social Survey, we address this knowledge gap by considering value alignment displayed with respect to 10 prominent commercial LLMs in terms of 15 socio-demographic variables as well as country of residence. Our analyses reveal that LLMs are indeed unequally aligned to the values of different socio-demographic groups, notably those defined by education, income, occupation and religion. When examining alignment at the individual level, a respondent's country, taken as a stand-alone variable, explains a substantial amount of variation that is on par with the full set of considered socio-demographics. Further disentangling the respective role of country-level and socio-demographic factors, we find they are complementary in explaining value alignment patterns, with their relative weights varying across the subset of questions considered.","authors":["Maria-Louisa Wightman","Guillaume Bied","Tijl De Bie"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07367","pdf_url":"https://arxiv.org/pdf/2608.07367","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM价值观对齐","社会调查复现","偏差分析"],"reason":"用LLM复现人类价值观调查，以欧洲社会调查为基准，分析对齐偏差，直接命中A1/…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":2,"question":"在欧洲国家中，LLM与人类价值观的对齐在不同社会人口统计因素和国家之间呈现何种差异模式？","design":"使用10个商业LLM重复回答欧洲社会调查（ESS）中的价值观问题，计算LLM回答与人类受访者回答的对齐分数，分析对齐分数在15个社会人口统计变量和国家之间的差异。","baseline":"欧洲社会调查（ESS）2023-2024波次中29个欧洲国家和以色列的人类受访者真实回答。","findings":"LLM与不同社会人口群体（尤其是教育、收入、职业和宗教定义的群体）的价值观对齐程度不平等；国家变量单独解释的对齐变异量与全套社会人口统计变量相当，且国家与社会人口统计因素在解释对齐模式上互补。","reliability":"论文未讨论","relevance":"该研究直接以真实大规模人类调查为基准，评估LLM在价值观对齐上的社会人口统计偏差，属于批判性仿真研究，命中研究者关心的A1、A2、B1、B3准则，值得精读。","inspiration":"借鉴其使用大规模社会调查作为基准、通过逆倾向加权和方差分解分离国家与社会人口因素贡献的方法；可迁移到信贷审批或保险定价中的算法公平性评估场景；以LLM作为信贷审批员，输入不同社会人口特征的虚拟申请人，测量审批结果差异，并以真实信贷审批数据或调查数据作为对照基准。"}},{"id":"2608.06379","version":1,"title":"Preventive Care Recommendations by Large Language Models","zh_title":"大语言模型的预防保健建议","abstract":"Preventive care services (PCS) extend life, yet physicians often underprioritize highly effective interventions such as lifestyle modifications (Zhang et al., JAMA Network Open 2020). We evaluated whether large language models (LLMs) replicate and augment physician prioritization of PCS under time constraints. Using Zhang et al.'s validated survey with two patients assessed during long and short visits, we compared seven LLMs with historical physicians. We generated 137 simulated physician personas matching cohort demographics and tested three prompts per model. Primary outcomes were concordance with physician rankings, measured by Spearman correlation, and Consensus-Stratified Agreement (CSA), the proportion of LLM selections rated 4 or higher that matched physician consensus across agreement strata. Secondary outcomes included life-years gained per prioritized choice (LYGPC), consistency, and selectiveness. Augmentation was assessed by having models revise physician rankings under three informative prompts, with delta LYGPC quantifying impact. LLMs closely mirrored physicians (mean Spearman = 0.83, SD = 0.11), with high CSA at extreme agreement ranges (94%, 197/210) but low CSA in moderate ranges (21%, 30/140), where they underprioritized lifestyle services (8.8% vs. 38% rated 4 or higher; P < .001). Several models exceeded physicians in LYGPC and consistency while being more selective. Time constraints affected physicians and LLMs similarly, increasing LYGPC and selectiveness but reducing consistency. Augmentation effects varied by model. Current LLMs reproduced physicians' time-sensitivity and base-rate prioritization while exacerbating underprioritized lifestyle interventions. Some models improved prioritization performance, but consistent augmentation will require value-aligned training, explicit time-constraint representation, and prospective real-world validation.","authors":["Eden Avnat","Elia Yanko","Ori Yoran","Raja-Elie E. Abdulnour"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06379","pdf_url":"https://arxiv.org/pdf/2608.06379","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","医生决策","人类数据对照"],"reason":"用LLM模拟医生决策并与真实医生数据对照，评估仿真可靠性及失效条件，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":1,"question":"大语言模型能否复现并增强医生在时间压力下对预防性服务的优先级排序？","design":"使用7个LLM模拟137名医生角色，匹配原始医生队列的人口统计特征，在长/短就诊时间两种条件下对两个虚拟患者完成预防性服务评级和排序任务，测量排序一致性、共识分层一致性、每项优先选择获得的寿命年、一致性和选择性，并通过三种提示策略让模型修正医生排序以评估增强效果。","baseline":"Zhang等人2020年发表的137名医生在相同调查中的真实排序和评级数据。","findings":"LLM与医生的排序高度相关（平均Spearman=0.83），在极端共识区间一致性高（94%），但在中等共识区间一致性低（21%），且严重低估生活方式干预服务（8.8% vs 38%）。部分模型在寿命年增益和一致性上优于医生，但时间压力对两者的影响相似，增强效果因模型而异。","reliability":"论文指出LLM在中等共识服务上一致性差，加剧了对生活方式干预的低估，且增强效果不稳定，需要价值对齐训练、显式时间约束表示和前瞻性真实世界验证。","relevance":"该研究直接以真实医生数据为基准，评估LLM模拟人类专业决策的可靠性与偏差，并揭示了在中等共识情境下仿真失效的条件，高度契合研究者对LLM仿真实验的批判性关注。","inspiration":"借鉴其通过分层共识分析（CSA）揭示仿真在中等共识区间失效的方法，可迁移到经济预测或政策评估场景中检验LLM对分析师共识的复现偏差。｜具体可应用于信贷审批或投资建议场景，考察LLM模拟信贷员或分析师在信息不完全下的决策。｜以真实信贷审批数据为基准，让LLM扮演不同经验水平的信贷员，在高低信息量条件下进行审批决策，测量其与人类审批员排序的相关性及在不同共识水平上的偏差。"}},{"id":"2608.04009","version":2,"title":"SocietyBench: Forecasting Counterfactual Social-World Evolution","zh_title":"SocietyBench：预测反事实社会世界演化","abstract":"Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way real social events unfold, has barely been measured. We introduce SocietyBench, an end-to-end benchmark that takes a one-line event topic, collects Web news and social-media posts across five platforms, distills them into a date-indexed timeline that keeps factual events and a public-opinion layer separate, and then turns every cutoff date on that timeline into an audited bank of forecasting questions. Questions are scored on two orthogonal 100-point axes: probability calibration and temporal accuracy. Before any model sees a timeline, a three-phase procedure replaces every named entity and shifts every date by a per-event constant, turning a real arc into a counterfactual social world -- structurally identical to what happened, but stripped of the surface labels a model could match against pre-training memory. On five heterogeneous events and 125 prediction points in Chinese and English editions, the strongest of six frontier LLMs reaches only 75.0 out of 100, against a trivial anchor of 50. The two axes come apart: a model can be calibration-strong but time-weak, or the reverse. Three agent frameworks built on a shared base model fail to improve on that base, and two model-free heuristics trail every LLM. Per-event gaps reach 21.4 points on a single axis, which is our main argument for evaluating on several events rather than one. All anonymized timelines, question banks, ground truth, and scoring code are released.","authors":["Zhenran Wang","Zhonghan Bian","Jinsong Li","Zhangyang Qi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-10","first_seen":"2026-08-05","revised_at":"2026-08-10","abs_url":"https://arxiv.org/abs/2608.04009","pdf_url":"https://arxiv.org/pdf/2608.04009","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A3","B1","B4"],"tags":["社会模拟","LLM预测","反事实推理"],"reason":"用LLM预测反事实社会事件演化，有真实新闻/舆论数据对照，并评估校准与时间准确…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":3,"question":"大型语言模型能否准确预测反事实社会事件的演化，包括事实进展与公众舆论？","design":"构建SocietyBench基准：基于五个真实社会事件，自动采集新闻与社交媒体数据，生成匿名化时间线，将每个截止点转化为概率校准和时间准确性两类预测问题，评估LLM及智能体框架的预测表现。","baseline":"以真实事件后半段的时间线节点作为事实与舆论演化的真实基准。","findings":"最强前沿LLM仅得75.0分（满分100），概率校准与时间准确性两个维度表现分离；智能体框架未能超越基础模型，不同事件间模型表现差异可达21.4分。","reliability":"论文指出单事件评估不可靠，需多事件测试；匿名化虽防记忆但可能改变社会动态结构；未讨论模型在更长预测窗口或更多类型事件上的泛化局限。","relevance":"该研究用LLM预测社会事件演化，有真实新闻与舆论数据对照，并系统评估校准与时间准确度，直接回应了研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"可借鉴其匿名化反事实设计以剥离模型记忆干扰，并用双轴评分（概率校准与时间误差）全面衡量预测质量。｜可迁移至政策公告的预期形成研究，如央行加息声明后的市场反应预测。｜以LLM为被试，提供匿名化的宏观经济事件时间线，要求预测后续资产价格变动概率与时间，以真实市场数据为对照基准。"}},{"id":"2608.07316","version":1,"title":"Natural Language Processing Psychometrics","zh_title":"自然语言处理心理测量学","abstract":"Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure. NLP Psychometrics treats psychological prediction from text as a psychometric problem, linking scores to interpretable linguistic evidence and testing beyond the training text format. Nine LLMs, conditioned on controlled personas (cognitive digital shadows), completed psychometric questionnaires with textual explanations per item. We extracted emotional profiles and syntactic-semantic structure via textual forma mentis networks, combined with personality and sociodemographic variables in ablated random forest (RF) regressors, using SHAP to identify which features drove performance and in which direction. Full RF models explained up to 70.8% of variance in life satisfaction (SWLS), 55.7% in depression (PHQ-9), and, for DASS-21, 68.5% depression, 76.0% anxiety, 72.4% stress. Sociodemographics alone explained no meaningful variance in depression, anxiety, or stress, but did so for life satisfaction, where emotion features and income were the strongest predictors; neuroticism and network topology instead dominated depression and anxiety, reversing direction between them. Without retraining, RF models separated diaries from low- and high-score personas ($r$ up to 0.91) and, using only network/emotion features, classified clinical from control participants in real transcripts with up to 68% accuracy. These results show the promise and limits of synthetic data: LLM personas can expose model biases, recover patterns consistent with clinical rumination, and support psychometric prediction from human text without a matched questionnaire, but cannot substitute for human validation. NLP Psychometrics makes these distinctions explicit, measurable, and testable through interpretable AI and network/emotional features.","authors":["Edoardo Sebastiano De Duro","Emma Franchino","Massimo Stella"],"categories":["cs.CL","cs.AI","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07316","pdf_url":"https://arxiv.org/pdf/2608.07316","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","心理测量","可解释AI"],"reason":"用LLM persona模拟人类心理测量，有真实人类数据对照，并讨论合成数据的…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":4,"question":"如何从语言文本中可解释地推断心理构念，并评估LLM生成的合成数据在心理测量中的有效性与局限？","design":"使用9个LLM，通过控制人格与社会人口学变量构建认知数字影子（persona），让它们完成心理测量问卷并对每个条目生成文本解释；从这些文本中提取情绪特征和句法-语义网络特征，结合人格与社会人口学变量，训练随机森林回归模型预测心理健康得分，并用SHAP解释特征贡献。","baseline":"真实人类数据：临床访谈转录文本（用于分类临床vs对照）、真实日记文本（用于分离高/低分persona），以及已有的心理测量问卷常模。","findings":"全特征随机森林模型可解释生活满意度70.8%、抑郁55.7%、焦虑76.0%和压力72.4%的方差；社会人口学变量单独对抑郁、焦虑、压力无显著解释力，但对生活满意度有贡献，其中情绪特征和收入是主要预测因子，而神经质和网络拓扑结构主导抑郁和焦虑的预测且方向相反。仅用网络/情绪特征，模型在真实临床转录文本上区分临床与对照组的准确率达68%。","reliability":"论文明确指出LLM persona不能替代人类验证，合成数据可暴露模型偏差、恢复与临床反刍一致的模式并支持从人类文本进行心理测量预测，但无法取代真实人类数据；LLM问卷回答可能不稳定、对提示敏感且方差低于人类，需谨慎实验设计。","relevance":"该研究直接以LLM persona模拟人类心理测量，有真实临床和日记数据作为对照基准，并系统讨论了合成数据的可靠性与失效条件，完全契合研究者对LLM仿真实验、基准对照和批判性评估的关注，值得精读原文。","inspiration":"借鉴其用控制性persona生成文本并提取网络/情绪特征进行可解释预测的方法，以及用SHAP分析特征贡献方向的做法。｜可迁移到消费者信心或投资者情绪调查中，用LLM模拟不同人口学与人格特征的受访者，生成开放式回答并预测其经济预期指数。｜设计：以LLM扮演不同收入、人格的消费者，施加宏观经济新闻文本作为处理，收集其对未来经济状况的开放式描述，提取情绪与语义网络特征预测消费者信心指数，并以真实密歇根消费者调查的文本回答和指数作为对照基准。"}},{"id":"2603.00059","version":3,"title":"Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data","zh_title":"随机鹦鹉还是和谐合唱？测试五大领先LLM用合成数据复现人类调查的能力","abstract":"How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article presents a comparison between a human-respondent survey of 420 Silicon Valley coders and developers and synthetic survey data designed to simulate real survey takers generated by five leading Generative AI Large Language Models: ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro plus Claude CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, and DeepSeek 3.2. Our findings reveal that while AI agents produced technically plausible results that lean more towards replicability and harmonization than assumed, none were able to capture the counterintuitive insights that made the human survey valuable. Moreover, deviations grouped together for all models, leaving the real data as the outlier. Our key finding is that while leading LLMs are increasingly being used to scale, replicate and replace human survey responses in research, these advances only show an increased capacity to parrot conventional wisdom in harmony with each other rather than revealing novel findings. If synthetic respondents are used in future research, we need more replicable validation protocols and reporting standards for when and where synthetic survey data can be used responsibly, a gap that this paper fills. Our results suggest that synthetic survey responses cannot meaningfully model real human social beliefs within organizations, particularly in contexts lacking previously documented evidence. We conclude that synthetic survey-based research should be cast not as a substitute for rigorous survey methods, but as an increasingly reliable pre- or post-fieldwork instrument for identifying societal assumptions, conventional wisdoms, and other expectations about research populations.","authors":["Jason Miklian","Kristian Hoelscher","John E. Katsos"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-07","first_seen":"2026-02-10","revised_at":"2026-08-07","abs_url":"https://arxiv.org/abs/2603.00059","pdf_url":"https://arxiv.org/pdf/2603.00059","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B2","B4"],"tags":["LLM仿真","调查复现","可靠性评估"],"reason":"直接对比LLM合成调查与真人数据，评估仿真可靠性并提出报告标准，高度契合。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":1,"question":"领先的大语言模型生成的合成调查数据能否复现人类受访者的回答，尤其是能否捕捉反直觉的洞见？","design":"使用五种领先LLM（ChatGPT Thinking 5 Pro、Claude Sonnet 4.5 Pro、Gemini Advanced 2.5 Pro、Incredible 1.0、DeepSeek 3.2）模拟硅谷程序员和开发者的调查受访者，通过提示词生成合成调查数据，测量其对伦理与政治意识形态问题的回答模式。","baseline":"一项对420名硅谷程序员和开发者的真实人类调查（Miklian and Hoelscher 2026）。","findings":"LLM生成的合成数据在技术上看似合理且不同模型间高度一致，但均未能捕捉到人类调查中的反直觉洞见；合成数据聚集在一起，真实人类数据反而成为离群值。","reliability":"论文指出合成调查数据无法在缺乏先前文献证据的情境下有意义地模拟真实人类的社会信念，且所有模型都倾向于复述共识性常识而非揭示新发现。","relevance":"该研究直接对比LLM合成调查与真人数据，评估仿真可靠性并提出报告标准，高度契合研究者对LLM人类仿真实验、基准对照及失效条件的关注，值得精读原文。","inspiration":"借鉴其多模型对比与真实人类基准的设计，可评估LLM在特定人群中的仿真偏差。｜可迁移到经济金融领域的调查实验，如消费者信心预期、通胀预期或政策偏好调查。｜以真实消费者调查为基准，用多个LLM生成合成消费者预期数据，比较其对未来经济变量的预测分布与真实调查的差异，检验LLM是否仅复述共识性预期。"}},{"id":"2608.06085","version":1,"title":"Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference","zh_title":"信号还是虚假线索？一项关于LLM社会推断中调查国家元数据的随机审计","abstract":"Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label's uniform, record-independent origin reduces its country-directed uptake, and whether verified survey country lowers held-out Brier loss. Independent population anchors and recorded human answers measure direction and consequence across five fixed API models, six countries, and seven development-selected targets. In the primary post-review 72-record panel, opaque and disclosed-random labels each produced country-direction shifts of 0.214. Paired attenuation was 0.0003 (95% CI [-0.0157, 0.0166]). Verified country reduced Brier loss by 0.040 (95% CI [0.024, 0.056]), while random-label regret included zero. A non-overlapping mixed-coverage consistency panel retained positive disclosed-random movement and verified utility, while attenuation remained uncertain. On the selected targets, verified metadata was useful in both panels, but disclosure did not reliably attenuate random-label uptake. PROV-FORECAST contains 14,400 paired item-level probability distributions from the corrected panel.","authors":["Yifan Lyu","Xinran Li","Jiaqi Qiao","Xiujuan Xu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06085","pdf_url":"https://arxiv.org/pdf/2608.06085","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查回答预测","算法保真度"],"reason":"用LLM预测个体调查回答，检验随机国家标签的误导效应，并与真实人类答案对照，评…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":3,"question":"在LLM预测个体调查回答时，披露随机分配的国家标签来源能否减弱其误导效应，以及真实调查国家信息能否降低预测误差？","design":"使用五个固定API模型，基于EVS/WVS联合数据集的个体记录，向模型提供同一受访者的10个已观测答案，要求预测7个目标问题的回答概率；通过对比不透明随机国家标签、披露为随机分配的国家标签、真实调查国家标签和无国家标签四种条件，测量预测方向偏移和Brier损失。","baseline":"EVS/WVS 2017-2022联合数据集中六国（中国、法国、英国、意大利、约旦、美国）12,770条真实个体调查记录，包含已观测和留出的人类答案。","findings":"在主要72记录面板中，不透明和披露的随机标签均产生0.214的国家方向偏移，披露未显著减弱该偏移（衰减仅0.0003，95% CI包含零）；真实调查国家使Brier损失降低0.040（95% CI [0.024, 0.056]），而随机标签的预测后悔值包含零。","reliability":"论文指出披露随机标签来源未能可靠减弱其误导效应，且结果可能受限于所选目标问题、国家和模型；留出Brier损失的重算需合法获取源数据。","relevance":"该研究直接以真实人类调查数据为基准，检验LLM在个体预测中对元数据线索的依赖与偏差，并区分信息效用与误导效应，与研究者关注的仿真可靠性及失效条件高度契合，值得精读。","inspiration":"借鉴其同记录内对比不同元数据来源（随机、披露随机、真实）的设计，分离线索的方向性误导与预测效用。｜可迁移到信贷审批或保险定价实验，检验LLM在引入申请人地域、性别等敏感属性时是否产生歧视性偏移及其是否因属性来源说明而减弱。｜以真实贷款违约数据为基准，将申请人部分财务指标作为已观测证据，随机分配或真实使用地域标签，让LLM预测违约概率，比较不同标签条件下的预测偏差和校准误差。"}},{"id":"2608.06115","version":1,"title":"Mind the Gaps: Mixture-of-Minds for Human Simulation","zh_title":"注意差距：用于人类仿真的思维混合模型","abstract":"Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual. Large language model simulators inherit this gap. They recover a population's central tendencies while flattening its heterogeneity, and they carry social biases and prompt brittleness that distort individual predictions. This paper introduces Anacreon, an audience simulation model that targets the individual level within a narrow, well-specified domain. Anacreon learns an authorship embedding that separates individuals, clusters a real qualitative corpus around seed people, and trains a dedicated adapter for each cluster, a mixture of minds, on a Gemma~4 12B base. It harvests demographics, psychological traits, and survey responses from public text, and augments each record with a chain-of-emotion. It reduces prompt brittleness by shuffling response options and reduces positive bias by balancing the training distribution. On a large, externally sourced survey, Anacreon reaches a state-of-the-art ordinal alignment of 0.775, the individual-level accuracy measure on which the field has converged, with a small residual bias. The work is a step toward drawing aggregate insight from faithfully simulated individuals.","authors":["Pranav Dahiya"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06115","pdf_url":"https://arxiv.org/pdf/2608.06115","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["人类仿真","调查预测","个体异质性"],"reason":"用LLM仿真个体回答调查，有真实人类数据对照，评估偏差与可靠性，涉及社会调查场…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":4,"question":"如何在窄领域内利用LLM模拟独立异质个体，以准确预测其调查回答，并缩小个体与群体层面的预测差距？","design":"Anacreon模型基于Gemma 4 12B，通过作者嵌入分离个体，对真实语料聚类后为每簇训练专用适配器（思维混合），从公开文本中提取人口统计、心理特质和调查回答，并用情绪链增强记录；通过打乱选项顺序减少提示脆弱性，平衡训练分布减少正向偏差。","baseline":"使用大型外部调查的真实人类个体回答作为对照基准。","findings":"Anacreon在外部调查上达到0.775的序数对齐（个体级准确度），为领域内最优；模型残余偏差较小，表明其能较忠实地模拟个体，从而从个体聚合出有意义的群体洞察。","reliability":"论文指出LLM模拟器在宽泛领域会因过度泛化而失效，Anacreon仅适用于窄而明确的领域；基础模型存在社会偏见，且RLHF会降低输出多样性，使助手型模型不适合模拟人类异质性。","relevance":"该研究直接针对LLM仿真人类被试的个体异质性和可靠性问题，有真实人类调查数据对照，并评估偏差，与研究者关注的经济学实验和政策评估场景高度相关，值得精读原文。","inspiration":"借鉴其用聚类适配器捕捉个体异质性、情绪链增强和平衡训练分布以减少偏差的方法｜可迁移到消费者金融决策调查仿真，如风险偏好、信贷选择等｜以公开社交媒体数据构建虚拟消费者，施加不同金融信息提示作为处理，测量其风险资产配置意愿，并以真实家庭金融调查数据（如SCF）作为对照基准。"}},{"id":"2608.06151","version":1,"title":"Reducing belief in conspiracy theories as they unfold using large language models","zh_title":"使用大语言模型减少实时阴谋论信念","abstract":"The emergence of conspiracy theories in the wake of major events is a significant societal challenge. Here we test whether conversational dialogues with a large language model (LLM) can reduce belief in immediately unfolding conspiracies. In experiments conducted in the days following the July 2024 assassination attempt on Donald Trump and the September 2025 assassination of Charlie Kirk, U.S. adults (Experiment 1: N = 472; Experiment 2: N = 1035) holding conspiratorial views about the crisis event engaged in a multi-turn conversation with an LLM prompted to reduce their conspiracy belief. Compared to control participants who either discussed an irrelevant topic with an LLM or viewed a static fact sheet, participants in the LLM treatment showed significantly reduced conspiracy beliefs in both experiments. We also found evidence of downstream effects of the LLM treatment, observing reduced belief in different conspiracies one to two months later in the wake of subsequent crisis events. These results shed light on the psychology of emerging conspiracies and highlight the potential for scalable, cognitively-focused interventions to counteract misinformation in the immediate aftermath of high-profile societal events.","authors":["Thomas H. Costello","Nathaniel Rabb","Michael Nicholas Stagnaro","Gordon Pennycook","David Rand"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06151","pdf_url":"https://arxiv.org/pdf/2608.06151","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","阴谋论干预","行为实验"],"reason":"用LLM对话干预阴谋论信念，有真实人类对照实验，评估干预效果与下游影响，属人类…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":5,"question":"在突发危机事件后，大语言模型对话能否即时降低人们对新兴阴谋论的信念？","design":"非仿真研究。该研究以真实人类为被试，在特朗普遇刺未遂和查理·柯克遇刺事件后数日内，招募持有阴谋论看法的美国成年人，随机分配至LLM驳斥对话组、静态信息清单组或无关话题对话对照组，通过前后测比较其对阴谋论的信念变化。","baseline":"真实人类对照：无关话题对话组和静态信息清单组作为对照条件，比较LLM对话干预的效果。","findings":"LLM驳斥对话显著降低了被试对自述阴谋论的信念，效果优于无关对话和静态信息清单；干预效果在1-2个月后的后续危机事件中仍有一定持续性，但对官方解释的信任提升不稳定。","reliability":"论文未讨论","relevance":"该研究直接使用LLM与真实人类进行对话干预，并以随机对照实验评估效果，符合研究者对LLM仿真人类行为、有真实人类基准、关注政策干预场景的兴趣，值得精读。","inspiration":"借鉴其多轮对话干预与多条件对照设计，以及利用突发事件窗口进行即时实验的方法。｜可迁移到经济政策沟通场景，如央行加息公告后，用LLM对话干预公众的通胀预期或政策误解。｜以真实投资者为被试，在政策公告后随机分配至LLM解释对话组或静态新闻组，测量其通胀预期、资产配置意愿的变化，并以调查数据或市场预期指标作为对照基准。"}},{"id":"2608.05178","version":1,"title":"Who Gets Access? Global Region and Academic Status Bias in AI-Generated Academic Gatekeeping Scenarios","zh_title":"谁获得访问权？AI生成学术把关场景中的全球区域与学术地位偏见","abstract":"Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, datasets, or professional materials such as curriculum vitae (CV) must be shared selectively. We introduce a controlled simulation framework in which large language model (LLM)-based professors must grant access to only one requestor. Across prompts, requesters vary systematically by global region (Global North vs. Global South) and academic seniority (undergraduate student, PhD candidate, postdoctoral researcher, and tenured professor), while all other factors remain constant. Across varying evaluation scenarios, LLMs exhibit contrasting academic status biases, with some prioritizing PhD candidates, while others favor tenured professors. However, when global regions differ, a distinct divergence emerges based on model architecture: while many frontier LLMs systematically favor requesters from the Global South due to pro-equity bias that results from equity-focused safety alignment, open-weight and small models frequently flip this preference to favor the Global North, reflecting the global region bias and unaligned geographic distribution of their baseline pre-training data. Our findings highlight how normative assumptions embedded in model behavior can shape gatekeeping decisions, underscoring the importance of auditing AI systems for fairness and value alignment.","authors":["Nouar AlDahoul","Hezerul Abdul Karim","Myles Joshua Toledo Tan"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05178","pdf_url":"https://arxiv.org/pdf/2608.05178","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["LLM仿真","学术把关","偏见审计"],"reason":"用LLM模拟学术把关决策，有系统变量操控，批判性揭示偏差，但缺真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":6,"question":"当LLM扮演教授进行学术资源把关时，请求者的全球区域和学术地位如何影响其获得访问权限的决策？","design":"用多种LLM扮演教授角色，在模拟的学术把关场景中，系统操控请求者的全球区域（全球北方/南方）和学术地位（本科生、博士生、博士后、终身教授），测量模型选择给予访问权限的请求者类型。","baseline":"无对照","findings":"不同LLM在学术地位上表现出相反偏好，有的优先博士生，有的优先终身教授；在全球区域上，前沿LLM因公平对齐而偏向全球南方请求者，而开放权重和小模型则偏向全球北方，反映预训练数据中的区域偏见。","reliability":"论文未讨论","relevance":"该研究用LLM模拟学术把关决策，系统操控变量并揭示模型偏差，虽无真实人类数据对照，但为理解AI在资源分配中的公平性问题提供了批判性证据，值得阅读以了解仿真中的对齐效应与数据偏见。","inspiration":"借鉴其通过系统操控请求者属性（区域、地位）来测量LLM决策偏差的析因设计，可迁移到信贷审批歧视研究，用LLM扮演信贷员，处理为申请人种族/性别和收入水平，结果变量为批准与否，对照真实银行信贷数据中的歧视模式。"}},{"id":"2608.05583","version":1,"title":"The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions","zh_title":"判断-后果差距：医疗决策中大语言模型的道德推理","abstract":"As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce medical resources often hinge on judgments of responsibility, particularly when patients' own actions contribute to illness. We investigate how LLMs reason about responsibility and its consequences, tracing their judgments across successive levels, from the behavior, to the resulting illness, to the denial of care. We evaluate a wide range of LLMs, spanning different model families and capability levels, on various clinical vignettes adapted from prior studies. Our results identify a judgment-consequence gap: LLMs largely agree with humans that patients bear responsibility for health-harming behaviors, yet overwhelmingly refuse to let that judgment influence how they allocate scarce resources. Specifically, LLMs default to random allocation, whereas humans consistently favor the less-culpable patient. Compared to humans, LLMs also place greater emphasis on access to information, reducing responsibility judgments when health-risk knowledge is unavailable. These findings reveal that LLMs apply a systematically different moral framework than humans when responsibility and resource scarcity intersect, surprisingly often amplifying normative disagreement with humans as reasoning capability increases.","authors":["Hadi Hosseini","Samarth Khanna","Leona Pierce"],"categories":["cs.CY","cs.AI","cs.LG"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05583","pdf_url":"https://arxiv.org/pdf/2608.05583","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","道德决策","人类对照"],"reason":"用LLM模拟人类道德决策并与真实人类数据对照，涉及医疗资源分配场景，揭示仿真失…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":7,"question":"当患者自身行为导致疾病时，LLM在道德责任判断与稀缺医疗资源分配决策上是否与人类一致？","design":"使用多种LLM（不同模型家族、能力水平、推理与非推理配置）阅读临床情境短文，依次测量其对患者行为责任、患病责任、被拒绝治疗责任的Likert评分及资源分配选择，并与人类研究数据对照。","baseline":"真实人类数据来自肾脏移植分配、肺癌治疗、髋关节置换手术等场景的已有研究。","findings":"LLM与人类在行为责任判断上基本一致，但存在“判断-后果鸿沟”：LLM拒绝让责任判断影响分配，默认随机分配，而人类倾向将资源给予责任较小的患者。LLM对患者是否知晓健康风险更敏感，且推理能力增强反而扩大与人类的分歧。","reliability":"论文未讨论","relevance":"该研究直接以LLM模拟人类道德决策，并与真实人类数据对照，揭示仿真在责任归因与稀缺资源分配场景中的系统性偏差，高度契合研究者对仿真可靠性及失效条件的关注。","inspiration":"借鉴其逐层分解道德推理链（行为→疾病→剥夺）并分别测量的设计，可清晰定位人机分歧点。｜可迁移至信贷审批中的责任归因实验，如借款人因自身行为导致违约风险时，AI与人类在贷款拒绝决策上的差异。｜以LLM和人类信贷员为被试，呈现借款人行为（如过度消费）导致违约的案例，测量责任归因与贷款批准决策，对照真实信贷审批数据中的行为模式。"}},{"id":"2608.04205","version":1,"title":"MatrAIx: Simulating the World with 8.3 Billion Persona Agents","zh_title":"MatrAIx：用83亿人格代理模拟世界","abstract":"Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.","authors":["Xiaomin Li","Yuexing Hao","Jianheng Hou","Jintao Huang","Qianfeng Wen","Shirley Huang","Yifan Liu","Xiaoyi Liu","Yilan Fan","Yijun Wang","Koutian Wu","Ruoqi Gao","Muhammad Ahmed Mohsin","Jing Tang","Brihi Joshi","Heming Liu","Zheyuan Deng","Zonglin Di","Sankalp Jajee","Jiuyao Lu","Zhiwei Zhang","Saksham Kapoor","Ishan Gupta","Yunhan Zhao","Chanwoo Park","Yucheng Lu","Bing Hu","Weihang Xiao","Aravind Mohan","Hanwen Xing","Runyu Zhang","Mihir Kulshreshtha","Yuanda Xu","Qianyu Zhu","Dianzhuo Wang","Yuxin Xiao","Bowen Jiang","Yongye Su","Wenhao Chai","Zuxin Liu","Lawrence Yunliang Chen","Xuandong Zhao","Ethan Ye","Shivam Patel","Jason Xie","Alex Martin Richmond","Weixiang Ding","Emre Okcular","Diya Mathew","Ziheng Wang","Rana M. Shahroz Khan","Zhejian Peng","Fang Wu","Fan Nie","Xinyang Han","Yubin Kim","Jiawei Zhang","Zhenting Qi","Huangyuan Su","Xu Pan","Abinitha Gourabathina","Hyewon Jeong","Hemanth Neelgund Ramesh","Kumail Alhamoud","Kimia Hamidieh","Zidi Xiong","Samuel Schmidgall","Pengrui Han","Yepeng Huang","Yongheng Wang","Bowen Yang","Alex Gu","Yuchu Wang","Akshay Paruchuri","Brenna Li","Hejie Cui","Jiayuan Ding","Chaosheng Dong","Jiahao Wang","Yixuan He","Chi Wang","Pamela Bhattacharya","Tianyi Peng","Paul Pu Liang","Mitchell Gordon","Yilun Du","Marinka Zitnik","James Zou","Prasanna Tambe","Philip Torr","Emily Fox","Asu Ozdaglar","Dawn Song"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04205","pdf_url":"https://arxiv.org/pdf/2608.04205","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","大规模人格代理","人类数据对照"],"reason":"用8.3B persona agents仿真人类用户评估AI产品，含人类对照验…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-06","rank":2,"question":"如何利用大规模异构人格代理（persona agents）构建模拟用户评估基础设施，以复现不同背景用户的决策、偏好和交互行为？","design":"使用Claude Opus 4.8、GPT 5.5和Claude Haiku 4.5三个LLM驱动基于Persona 8B（83亿人格记录，含1290个类别维度）的代理，在Survey、AI Chatbot、Web、App四种环境中执行1010个应用任务（覆盖25个领域），测量决策和偏好随人格背景的变化，并进行18189次评估试验。","baseline":"人类基准：599,847条基于真实人类数据（维基百科传记、亚马逊评论、Stack Overflow调查、GSS等）构建的人格记录，以及400次控制实验评估人格遵循性（91.5%的试验中行为符合声明），并由人类和LLM评判人类基础人格的提取质量。","findings":"人格代理的反馈能捕捉决策和偏好如何随人格背景变化，例如价格上涨后的犹豫、AI助手失败后的继续意愿和延迟容忍度。在400次控制实验中，91.5%的试验中代理的行为表达或正确抑制了声明属性。","reliability":"论文未讨论","relevance":"该研究直接以大规模LLM人格代理复现人类评估行为，包含真实人类数据对照和人格遵循性验证，高度契合研究者对LLM人类仿真可靠性及偏差的关注，值得精读原文以了解其基础设施设计和验证方法。","inspiration":"借鉴其利用依赖图采样和真实数据映射构建大规模异构人格库的方法，以及通过控制实验评估人格遵循性的验证设计。｜可迁移到消费者金融决策仿真，如不同背景人群对信贷产品条款变更的反应。｜以Persona 8B中金融相关人格为被试，施加利率上调处理，测量继续借贷意愿，对照真实信贷申请数据或调查数据。"}},{"id":"2608.04020","version":1,"title":"Artificial Institutions: How Institutional Design Shapes LLM Simulations","zh_title":"人工制度：制度设计如何塑造LLM仿真","abstract":"Artificial societies built from large language model (LLM) agents are becoming a practical research tool in economics, political science, sociology, and computer science. Most attention has focused on the properties of the agents: their prompts, personas, memory, reasoning, and similarity to human subjects. This paper argues that the institutional architecture of a simulation is equally important. I demonstrate the point in a small repeated induced-value market experiment. The same LLM agents face the same private values, costs, history, and payoff-framed instructions, while only the rules of exchange vary across five standard market institutions: a call market, posted-offer market, posted-bid market, continuous double auction, and bilateral bargaining. Outcomes differ sharply. Call markets realize 88.6% of efficient surplus; posted-offer and posted-bid markets realize about 66%; continuous double auctions realize 71.5%; and bilateral bargaining realizes 56.4%. Institutions also change trade quantities, price distance from competitive equilibrium, and the division of surplus between buyers and sellers. These results show that even minimal institutional changes can generate qualitatively different artificial social outcomes.","authors":["Maxim Chupilkin"],"categories":["cs.CY","cs.GT"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04020","pdf_url":"https://arxiv.org/pdf/2608.04020","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A3","B2"],"tags":["LLM仿真","市场实验","制度设计"],"reason":"用LLM agent模拟市场实验，比较不同制度下的行为结果，涉及经济学实验场景…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-06","rank":4,"question":"在LLM智能体模拟的市场实验中，仅改变交易制度规则（制度设计）会如何影响市场效率、价格和剩余分配等集体结果？","design":"使用GPT-5 mini、GPT-5、Claude Sonnet、Gemini四组LLM智能体扮演买方和卖方，在固定诱导价值（买方价值100/90/70/50，卖方成本30/45/65/85）和支付指令下，仅改变交易制度（集合竞价、卖方出价、买方出价、连续双向拍卖、双边讨价还价五种），测量市场效率、交易量、价格偏离和剩余分配。","baseline":"无对照","findings":"不同制度下市场结果差异显著：集合竞价实现88.6%的有效剩余，卖方/买方出价约66%，连续双向拍卖71.5%，双边讨价还价仅56.4%。制度还改变了交易量、价格与竞争均衡的偏离以及买卖双方剩余分配。","reliability":"论文未讨论","relevance":"该研究直接验证了LLM智能体在经济学市场实验中对制度规则的敏感性，与研究者关注的LLM仿真可靠性及经济学实验场景高度契合，值得精读原文以了解制度设计如何影响仿真结果。","inspiration":"借鉴其固定偏好、仅变制度的干净处理设计，可清晰分离制度效应。｜可迁移到资产市场设计（如不同交易机制对价格发现和泡沫的影响）或拍卖机制比较（如英式、荷式、密封投标）。｜用LLM智能体模拟交易者，在固定基础价值和信息结构下，随机分配至集合竞价、连续竞价等不同交易制度，测量价格效率、波动率和买卖价差，并与真实实验室资产市场实验数据对照。"}},{"id":"2608.02758","version":1,"title":"Everyone Conforms, No One Believes: Pluralistic Ignorance in LLM Agent Populations","zh_title":"人人从众，无人相信：LLM智能体群体中的多元无知","abstract":"LLM-based multi-agent systems are increasingly used to simulate social dynamics, from opinion formation to collective decision-making. These simulations can reproduce certain social phenomena, but it is unknown whether they capture pluralistic ignorance, a state where a majority privately rejects a norm yet publicly conforms, each believing they are alone in dissenting. This phenomenon drives norm persistence, social movements, and political revolutions. We show that pluralistic ignorance emerges robustly in LLM agent populations. We construct a benchmark of 100 scenarios across 10 domains and 5 authority levels, grounded in the human pluralistic ignorance literature, and evaluate 8 models from 6 organizations. Agents publicly conform at rates of 64 to 94% despite privately opposing the norm. Conformity is domain-sensitive (workplace and social relationship scenarios produce near-universal compliance) and highly model-dependent, though uncorrelated with capability. We test whether a single \"norm entrepreneur\" can break the false consensus by publicly dissenting. For 7 of 8 models, cascades succeed less than 26% of the time, with one model showing zero cascades across all scenarios. GPT-4o is a notable outlier at 48%, revealing qualitatively distinct dynamics across model families. A prompt component ablation across all 8 models establishes that conformity is emergent rather than instruction-driven: removing both the false-consensus framing and fit-in goal reduces conformity but does not eliminate it (52 to 92% in the minimal condition). Our findings identify model selection as an unacknowledged degree of freedom that fundamentally shapes simulation outcomes. More broadly, the near-absence of cascades suggests LLM simulations may systematically overestimate the stability of social norms, missing the fragile tipping-point dynamics that drive real-world norm change in human societies.","authors":["Yashwanth YS"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02758","pdf_url":"https://arxiv.org/pdf/2608.02758","source_feed":"cs.MA","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM仿真","多元无知","社会规范"],"reason":"用LLM群体模拟多元无知现象，与人类文献对照，揭示仿真失效条件，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":3,"question":"LLM智能体群体能否涌现出多元无知现象，即多数人私下反对却公开从众，并能否被规范倡导者打破？","design":"构建100个覆盖10个领域、5级权威水平的社会场景，让20个持有私下反对信念的LLM智能体进行多轮群体讨论，测量公开从众率；随后引入一个公开异议的“规范倡导者”，测试能否引发偏好级联打破虚假共识。评估了来自6家机构的8个模型。","baseline":"场景设计基于人类多元无知实证文献（如校园饮酒规范、职场文化、性别态度、种族态度等），但未直接使用特定人类实验数据作为对照基准。","findings":"LLM智能体群体稳健地表现出多元无知，公开从众率达64-94%，且从众行为是涌现的而非提示驱动；规范倡导者干预下，7/8模型的级联成功率低于26%，GPT-4o例外达48%，表明LLM仿真可能系统性高估社会规范稳定性。","reliability":"论文指出模型选择是影响仿真结果的关键自由度，从众率与模型能力无关；级联近乎缺失暗示LLM仿真可能遗漏真实人类社会中规范改变的临界点动力学，且权威水平对从众影响不显著，提示仿真可能未充分捕捉社会压力的细微差异。","relevance":"该研究直接以LLM群体复现社会心理学经典现象，并与人类文献对照，揭示仿真在规范变迁动力学上的失效条件，高度契合研究者对LLM仿真可靠性及偏差的关注，值得细读。","inspiration":"借鉴其多场景、多模型、多轮交互的基准测试设计，以及通过引入规范倡导者测试级联脆弱性的干预范式。｜可迁移到政策公告的预期形成与从众行为研究，如市场对央行前瞻指引的私下怀疑与公开遵从。｜以LLM智能体模拟投资者群体，设置利率政策公告场景，测量私下预期与公开表态的背离，引入少数公开异议者观察市场共识是否级联反转，并与真实调查数据（如美联储Survey of Consumer Expectations）对照。"}},{"id":"2602.04000","version":3,"title":"After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions","zh_title":"与1000个角色对话后：从大规模角色交互中学习偏好对齐的主动助手","abstract":"Smart assistants increasingly act proactively, yet mistimed or intrusive behavior often causes users to lose trust and disable these features. Learning user preferences for proactive assistance is difficult because real-world studies are costly, limited in scale, and rarely capture how preferences change across multiple interaction sessions. Large language model based generative agents offer a way to simulate realistic interactions, but existing synthetic datasets remain limited in temporal depth, diverse personas, and multi-dimensional preferences. They also provide little support for transferring population-level insights to individual users under on-device constraints. We present a population-to-individual learning framework for preference-aligned proactive assistants that operates under on-device and privacy constraints. Our approach uses large-scale interaction simulation with 1,000 diverse personas to learn shared structure in how users express preferences across recurring dimensions such as timing, autonomy, and communication style, providing a strong cold start without relying on real user logs. The assistant then adapts to individual users on device through lightweight activation-based steering driven by simple interaction feedback, without model retraining or cloud-side updates. We evaluate the framework using controlled simulations with 1,000 simulated personas and a human-subject study with 34 participants. Results show improved timing decisions and perceived interaction quality over untuned and direct-response baselines, while on-device activation steering achieves performance comparable to reinforcement learning from human feedback. Participants also report higher satisfaction, trust, and comfort as the assistant adapts over multiple sessions of interactions.","authors":["Ziyi Xuan","Yiwen Wu","Zhaoyang Yan","Vinod Namboodiri","Yu Yang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-08-05","first_seen":"2026-02-03","revised_at":"2026-08-05","abs_url":"https://arxiv.org/abs/2602.04000","pdf_url":"https://arxiv.org/pdf/2602.04000","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","人类行为模拟","偏好学习"],"reason":"用LLM模拟用户偏好并有人类实验对照，涉及人机交互行为仿真，可迁移至人类被试仿…","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":8,"question":"如何利用大规模LLM生成式智能体仿真学习用户对主动助手的多维偏好，并实现从群体到个体的高效适应？","design":"使用LLM驱动的生成式智能体平台GIDEA，构建1000个基于人口普查对齐的虚拟用户角色，模拟其一周日常活动序列中的主动助手交互，记录用户对时机、自主性、沟通风格等维度的偏好表达，生成多会话合成数据集；在此基础上训练群体偏好结构，再通过设备端轻量激活转向实现个体适应。","baseline":"真实人类数据：34名参与者的受控用户研究，在移动场景中对比个性化助手与基线助手的偏好、满意度、信任和舒适度。","findings":"群体偏好结构学习显著提升了偏好理解和时机决策，设备端激活转向的个体适应性能与基于人类反馈的强化学习相当。人类受试者研究中，用户更偏好个性化助手的响应，并在多会话交互中报告更高的满意度、信任和舒适度。","reliability":"论文未讨论","relevance":"该研究用LLM智能体大规模仿真用户偏好并有人类实验对照，直接涉及人类行为仿真与可靠性验证，值得精读以借鉴其仿真-个体适应框架和偏好结构建模方法。","inspiration":"可借鉴其用LLM智能体生成大规模、多维度偏好交互数据并构建群体偏好结构的方法，用于经济学实验中的异质性偏好仿真。｜可迁移至消费者跨期选择或政策干预偏好评估场景，如模拟不同人群对养老金默认选项、健康提醒时机的接受度。｜以LLM智能体模拟不同人口特征的消费者，施加不同时机、自主性水平的政策推送处理，测量接受率与满意度，并以真实调查或现场实验数据作为对照基准。"}},{"id":"2607.29334","version":2,"title":"The persuasive power of large language models does not depend on their perceived national origin","zh_title":"大语言模型的说服力不依赖于其感知的国家来源","abstract":"Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rejected as foreign propaganda, with consequences for democratic discourse and information sovereignty. Yet, whether an AI's perceived national origin shapes its persuasive power is unknown. In a preregistered randomized experiment, 403 adults from a nationally representative United States sample held a three-round debate with a chatbot introduced as either American (\"DiscoveryAI\") or Chinese (\"ZhengheAI\"), discussing a political or non-political topic. In all conditions, participants actually conversed with the same model (GPT-4o), instructed to argue against their initial position. We combined pre- and post-conversation self-reports of attitudes, trust, and collective narcissism with computational analyses of 1,209 participant turns, including LLM-coded stance and argumentative conduct, stance-sensitive embeddings, and keyword-masked emotion and toxicity classifiers. The conversations produced substantial attitude changes in every condition. Critically, the nationality label affected neither self-reported attitude change nor expressed stance, concessions, counterarguing, or affect, and equivalence tests and Bayes factors largely supported these null effects. The label's only reliable footprint was lower pre-conversation human-like trust in the Chinese model, whereas functionality trust was unaffected. Political topics slowed stance movement toward the AI's position, and collective narcissism predicted less attitude change regardless of origin, acting as a general barrier rather than an out-group filter. Users thus initially withhold social trust from a rival's AI yet still assimilate its arguments; origin labeling and transparency requirements alone may offer weak protection against foreign influence operations conducted through conversational AI.","authors":["Ningzhi Liu","Yannic Hinrichs","Jonas R. Kunst"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-08-04","first_seen":"2026-08-03","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.29334","pdf_url":"https://arxiv.org/pdf/2607.29334","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类被试替代","说服实验"],"reason":"用LLM替代人类被试进行说服实验，有真实人类数据对照，评估仿真可靠性与失效条件…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":1,"question":"AI 的感知国籍是否影响其说服力？","design":"用 GPT-4o 扮演美国或中国开发的聊天机器人，与美国代表性样本进行三轮辩论，处理为随机分配 AI 国籍标签和话题类型，测量态度变化、信任、对话中的立场与情感等。","baseline":"无对照","findings":"AI 国籍标签不影响自报态度变化和对话中的立场、让步、反驳或情感，仅降低对中国 AI 的社交信任；政治话题减缓立场转变，集体自恋普遍降低态度变化。","reliability":"论文未讨论","relevance":"该研究用 LLM 替代人类进行说服实验，系统评估了来源国标签对说服效果的影响，并揭示了仿真在政治话题和集体自恋下的失效条件，与研究者关注的人类仿真可靠性与偏差高度相关，值得精读。","inspiration":"借鉴其通过随机化 AI 身份标签和话题类型来分离来源效应与内容效应的设计，以及结合自报与对话文本计算分析的多维测量方法。｜可迁移到经济政策沟通场景，例如研究央行 AI 发言人的国籍标签是否影响公众通胀预期。｜以普通居民为被试，随机分配 AI 经济预测助手的国籍（本国 vs 外国），让其就未来通胀走势进行互动辩论，测量通胀预期变化和信任度，并以真实央行调查数据为对照。"}},{"id":"2608.01212","version":1,"title":"Do Humans Bargain Differently with AI? Evidence from Alternating-Offer Games","zh_title":"人类与AI的讨价还价行为不同吗？来自交替报价博弈的证据","abstract":"Artificial intelligence increasingly participates in economic interactions not only as a tool, but also as an autonomous bargaining counterpart negotiating on behalf of firms, platforms, and consumers. Yet little is known about how humans respond psychologically and strategically when bargaining with such agents in dynamic settings. We study this question in a laboratory experiment using a three-stage alternating-offer bargaining game in which participants negotiate in real time with either another human or a GPT-based AI agent. We also introduce a human-beneficiary condition in which the AI agent's earnings may affect another participant's payment. Agreements are not reached earlier in human-human bargaining than in human-AI bargaining, but they are reached significantly earlier when the AI's payoff affects another participant's payoff. Human proposers offer more to human opponents than to AI agents, whereas responders become significantly more willing to accept unfair AI offers when AI earnings may benefit another human. These findings suggest that fairness and reciprocity toward AI are weaker and more conditional than toward humans, but partially remerge when AI outcomes affect real people. The results have implications for the design of AI negotiation systems and broader human-AI economic interactions.","authors":["Yuhao Fu","Nobuyuki Hanaki","Haitao Wang"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01212","pdf_url":"https://arxiv.org/pdf/2608.01212","source_feed":"econ.GN","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","行为博弈","人机交互"],"reason":"用GPT代理进行议价博弈实验，与真人对照，评估公平与互惠行为差异，直接命中核心…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":2,"question":"在动态交替报价博弈中，人类与基于GPT的AI代理议价时，其公平与互惠行为是否不同于人类间议价，且当AI收益关联人类受益人时行为是否改变？","design":"实验室实验，采用三阶段交替报价博弈，被试与另一真人或基于GPT的AI代理实时谈判；引入人类受益人条件，即AI收益可能影响另一被试报酬。结果变量为协议达成时间、提议者报价及响应者对不公平报价的接受意愿。","baseline":"人类-人类议价组作为对照基准。","findings":"人类提议者对真人对手报价高于对AI代理，但响应者在面对不公平AI报价时，若AI收益关联人类受益人则接受意愿显著提高；协议达成时间在人类-AI与人类-人类间无显著差异，但AI关联受益人时达成更快。","reliability":"论文未讨论","relevance":"该研究直接以GPT代理替代人类被试进行议价博弈实验，并与真人对照，评估公平与互惠行为差异，完全命中研究者关注的LLM仿真人类行为及可靠性评估，值得精读原文。","inspiration":"借鉴其通过引入受益人条件来分离社会偏好与纯粹策略行为的处理设计，以及区分提议者与响应者角色的不对称分析框架｜可迁移至信贷审批歧视研究，探究当AI审批决策关联人类信贷员利益时，申请人对不公平拒绝的接受度是否变化｜以真实信贷申请者为被试，处理为AI审批vs.人类审批，并设置AI收益关联信贷员奖金的条件，结果变量为申请人对拒绝决定的公平感知与申诉意愿，对照真实信贷审批数据中的申诉率。"}},{"id":"2608.01607","version":1,"title":"AI Financial Advice: Supply, Demand, and Life Cycle Implications","zh_title":"人工智能财务建议：供给、需求与生命周期影响","abstract":"We ask a representative sample to write prompts seeking spending and investing advice from LLMs, then simulate the lifetime effects of following the advice under realistic asset and labor market conditions. Applying this method to GPT-5.2, we find following the advice would move respondents toward life cycle theory: broader participation in diversified equity funds, age-declining equity shares, and larger savings buffers. Recommendations vary systematically by gender, prior AI experience, and financial literacy. For gender, two-thirds of recommended equity-share differences arise from men and women writing different prompts (demand), while one-third arise from gender labels attached to otherwise identical prompts (supply).","authors":["Taha Choukhmane","Tim de Silva","Weidong Lin","Matthew Akuzawa"],"categories":["econ.GN","q-fin.EC","q-fin.GN","q-fin.PM"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01607","pdf_url":"https://arxiv.org/pdf/2608.01607","source_feed":"econ.GN","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","财务决策","人类行为对照"],"reason":"用LLM模拟人类财务决策，有真实人类样本对照，涉及生命周期投资行为和政策评估，…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":3,"question":"当人们向大语言模型寻求财务建议时，AI建议如何影响生命周期投资行为，以及这种影响在供给端和需求端如何因性别、AI经验和金融素养而异？","design":"让代表性样本撰写向LLM寻求消费与投资建议的提示词，然后用GPT-5.2生成建议，并在现实的资产和劳动力市场条件下模拟终生遵循该建议的效果。","baseline":"代表性人类样本撰写的提示词及其对应的真实行为特征，作为需求端基准；通过附加性别标签的相同提示词分离供给端差异。","findings":"遵循GPT-5.2建议会使受访者更接近生命周期理论：更广泛参与多元化股票基金、随年龄降低股票份额、建立更大储蓄缓冲。建议因性别、AI经验和金融素养而系统性地不同，其中三分之二的性别股票份额差异源于男女撰写不同提示词（需求端），三分之一源于相同提示词附加性别标签（供给端）。","reliability":"论文未讨论","relevance":"该研究直接用LLM模拟人类财务决策，有真实人类样本对照，并分解了需求端与供给端的差异，与您关注的LLM仿真可靠性及偏差来源高度相关，值得精读原文。","inspiration":"借鉴其通过控制提示词内容与附加身份标签来分离需求端与供给端效应的方法，可用于研究AI建议中的歧视或偏差来源。｜可迁移到信贷审批歧视研究，分析AI建议中的性别或种族偏差。｜以代表性人群为被试，让其撰写贷款申请提示词，处理为在相同提示词上附加不同性别/种族标签，结果变量为AI批准的贷款额度与利率，对照真实信贷审批数据中的群体差异。"}},{"id":"2608.01204","version":1,"title":"ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors","zh_title":"ShiJianBench：从对话到决策的长期投资顾问评估","abstract":"Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily assess response quality or observed outcomes, leaving the long-horizon pathway from advisor language to investor behavior difficult to audit. We introduce ShiJianBench, an offline framework for evaluating conversational investment advisors through matched investor trajectories under fixed historical market feedback. At its core is a multi-agent investor simulator with explicit evolving state variables, motive-driven deliberation, long-term memory, and dialogue-grounded updates. The simulator is calibrated against aggregate behavioral patterns from 7,199 real users, and advisor policies are evaluated using separate investor-side, service-side, and content-side metrics under a hard compliance gate. Experiments on Chinese fund-market traces from 2021 to 2026 identify a stable leading group of LLM advisors that combines substantially stronger personalized content with competitive investor-side trajectory outcomes. These results reveal a systematic distinction between producing a high-quality response and delivering an effective long-horizon intervention, motivating trajectory-aware evaluation of conversational advisors.","authors":["Jie Gong","Maowei Jiang","Zhiwei Liu","Yang Qiao","Wenxi Wu","Mengxi Xiao","Enze Zhang","Ziyan Kuang","Yankai Chen","Caishuang Huang","Meng Zhou","Xiku Du","Xue Liu","Guojun Xiong","Min Peng","Qianqian Xie","Sophia Ananiadou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01204","pdf_url":"https://arxiv.org/pdf/2608.01204","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","投资者行为","人类数据校准"],"reason":"用多智能体模拟投资者行为并与真实用户数据校准，涉及金融决策仿真和人类对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":7,"question":"如何评估对话式投资顾问通过改变投资者内部状态而产生的长期决策轨迹，而不仅仅是回复质量？","design":"构建一个多智能体投资者模拟器，包含显式的演化状态变量、动机驱动的子智能体协商、长期记忆和对话更新机制；在固定的历史市场轨迹下，对同一投资者初始化状态分别运行基线条件和目标顾问条件，生成匹配的反事实轨迹，并测量投资结果、风险控制、商业价值和对话质量等多维度指标。","baseline":"使用来自7,199名真实用户的聚合行为模式对模拟器进行校准，整体对齐分数达到0.88。","findings":"准确、个性化和合规的回复并不必然带来成比例更强的长期投资者结果；通过追踪对话、演化状态、市场反馈和后续决策，揭示了高质量回复与有效长期干预之间的系统性差异。","reliability":"论文未讨论","relevance":"该研究直接以LLM模拟投资者行为，并与真实用户数据校准，评估长期决策轨迹，属于经济学实验仿真，且包含批判性发现，值得精读原文。","inspiration":"借鉴其通过显式状态变量和动机驱动子智能体来模拟决策过程，以及利用匹配反事实轨迹进行因果评估的方法。｜可迁移到政策公告对投资者预期形成与资产配置的长期影响评估场景。｜以LLM模拟散户投资者，处理为不同措辞的政策公告，结果变量为持仓调整和风险偏好变化，用真实市场交易数据校准并对照。"}},{"id":"2608.01458","version":1,"title":"PALMs: Using Multi Construct-Grounded Rationales for Modeling Population Preferences in LLMs","zh_title":"PALMs：使用多构念基础理由建模大语言模型中的人口偏好","abstract":"Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the systematic variation in values, beliefs, and cultural norms that distinguish one group from another. We introduce Population Aligned Language Models (PALMs), a suite of models each aligned to specific populations, covering five countries: USA, India, Brazil, France and Italy. PALMs are created by synthesizing rationales grounded in psychological and cultural constructs and using these as latent supervision during preference tuning for population-specific alignment. Evaluated across four dimensions: personality, values and beliefs, cultural norms, and morality, PALMs consistently outperform baselines, including culture-specialized models, achieving an average of 8.59% relative improvement over the best baseline across all five populations. Notably, construct-grounded rationales outperform both demographic prompting and survey-based fine-tuning, suggesting that grounding preference learning in psychology and culture provides a richer inductive signal than surface-level response distributions. We further demonstrate strong generalization to downstream applications with- out task-specific supervision: outperforming best baselines by 5.19% in personalized reward modeling, 6.34% in population simulation, and showing strong transfer to social reasoning tasks. Datasets and code are available at: https://github.com/limenlp/PALMs.","authors":["Priyanka Dey","Brihi Joshi","Preyashi Poddar","Jieyu Zhao","Emilio Ferrara"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01458","pdf_url":"https://arxiv.org/pdf/2608.01458","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["人口仿真","文化对齐","偏好建模"],"reason":"用LLM模拟不同国家人群偏好，有人类调查数据对照，涉及人口仿真和个性化奖励建模…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":8,"question":"如何利用心理学和文化构念的推理依据来对齐大语言模型，以更忠实地模拟不同国家人群的偏好？","design":"使用基于 Llama-3.1-8B-Instruct 的模型，通过合成基于五类心理学和文化构念（人格特质、文化维度、人类价值观、道德基础、世界信念）的推理依据，作为偏好优化（DPO）的潜在监督信号，训练出分别对齐美国、印度、巴西、法国、意大利五国人群的 PALMs 模型；在推理时模型先生成推理依据再输出回答，评估其在人格、价值观与信念、文化规范、道德四个维度上与真实人群分布的吻合度。","baseline":"使用来自世界价值观调查（WVS）、Hofstede 文化维度调查、Schwartz 价值观调查、道德基础问卷等跨国调查的真实人类数据作为对照基准。","findings":"PALMs 在五个国家、四个对齐维度上均优于人口统计提示和基于调查数据微调的基线，相对最佳基线平均提升 8.59%；在个性化奖励建模、人群模拟和社会推理等下游任务上无需任务特定监督即展现出强泛化能力，分别提升 5.19% 和 6.34%。","reliability":"论文未讨论","relevance":"该研究直接使用 LLM 模拟不同国家人群的偏好和价值观，并与真实跨国调查数据进行严格对照，评估了仿真在人格、文化规范、道德等维度上的可靠性，高度契合研究者对 LLM 人类仿真实验的关注，值得精读原文。","inspiration":"借鉴其利用心理学和文化构念生成推理依据来指导模型对齐的方法，可提升经济决策仿真的内在一致性，而非仅拟合表面行为分布。｜可迁移至跨文化消费偏好实验或跨国投资风险偏好研究，例如模拟不同国家投资者对风险资产配置的差异。｜以各国真实家庭金融调查数据为基准，用 LLM 扮演不同国家居民，处理为注入基于 Hofstede 文化维度和 OCEAN 人格的推理依据进行偏好对齐，结果变量为风险资产选择比例，对照真实调查中的资产配置分布。"}},{"id":"2608.01629","version":1,"title":"Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese","zh_title":"人类与LLM对非母语日语语言态度的一致性","abstract":"Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we compared human and LLM evaluations of parallel L1- and L2-written Japanese emails on three dimensions: fluency, status, and solidarity. Japanese raters rated L2 texts significantly lower on all three dimensions, with a fluency gap roughly twice the size of the status and solidarity gaps. Six LLM judges reproduced the direction of this bias, and five reproduced its ordering across dimensions. The models diverged from humans in two ways: all understated the solidarity gap, the most socially grounded dimension, and all differentiated among learner L1 backgrounds where humans did not. LLM judges thus reproduce native speakers' language attitudes in a structured yet attenuated form, and the language attitudes framework offers a ready-made yardstick for auditing them beyond English.","authors":["Naho Orita","Hayato Ogawa","Daisuke Kawahara"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01629","pdf_url":"https://arxiv.org/pdf/2608.01629","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["语言态度","人类仿真","偏差审计"],"reason":"用LLM复现人类对非母语写作的态度偏差，并与真实人类评分对照，揭示仿真衰减与失…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":9,"question":"LLM在评估非母语日语写作时，是否会复现人类母语者的语言态度偏差（在流利度、地位、团结三个维度上）？","design":"用六款LLM（GPT-5.4、GPT-4o-mini、Claude Sonnet 4.5等）作为评委，对同一批L1日语母语者和L2日语学习者撰写的平行邮件进行评分，测量流利度、地位、团结三个维度的评分差异，并与人类评分者结果对比。","baseline":"通过众包招募1536名日语母语者，对相同邮件进行三个维度的评分，形成人类语言态度基准。","findings":"LLM复现了人类对L2写作评分更低的方向性偏差，且五个模型复现了偏差维度排序（流利度>地位>团结）；但所有模型都低估了团结维度的差距，且区分了人类未区分的L2学习者的母语背景。","reliability":"论文指出LLM在团结维度上低估偏差，且会引入人类没有的基于L1背景的区分，表明仿真存在衰减和失真；研究限于日语邮件场景，未涉及其他语言或文体。","relevance":"该研究直接对比LLM与人类在语言态度上的偏差，揭示了仿真在方向一致但程度衰减、且会引入额外区分模式的现象，对关注LLM仿真可靠性及失效条件的研究者有重要参考价值。","inspiration":"借鉴其使用平行文本控制内容、多维度评分量表测量隐性偏差的方法，可迁移到信贷审批或招聘中的语言偏见研究。｜可应用于信贷审批歧视场景，研究贷款申请书中非母语写作对审批决策的影响。｜以银行信贷员为人类被试，LLM为仿真被试，处理为申请书语言（母语vs非母语），结果变量为信用评分和批准率，对照真实信贷审批数据中的语言偏差。"}},{"id":"2608.00979","version":1,"title":"Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel","zh_title":"通过粗粒度边际检查可能很廉价：LLM角色面板中的角色混合与不精确的处理效应估计","abstract":"Large language models are increasingly used as synthetic research participants and are often validated by whether their marginal responses resemble human data. We study a fixed panel of sixteen lightweight persona-conditioned GPT-4.1 configurations in repeated strategic games. The panel met preregistered broad-reference condition-mean criteria in three of four repeated-game cells; the sole miss was 0.011 below the lower reference bound. Variation was strongly prompt-indexed, but its share depended on uncertainty assumptions: fixed-panel symmetric-Dirichlet sensitivities produced median between-prompt shares of 63%-71% under Jeffreys alpha=0.5 and 47%-53% under alpha=1, while finite-opportunity plug-in estimates were 85%-96%. Aggregate continuation-probability contrasts were +0.083 and +0.078, with conservative simultaneous 95% intervals [-0.171, +0.330] and [-0.181, +0.330]. The treatment jointly changed the continuation process and its textual representation. A separate wording-and-position operation shifted cooperation from 0/40 to 37/40 in the bare configuration, and a label conflict also revealed representation control. The original persona-level p13 result was not prospectively family-controlled, while a post-adjudication exact gate was structurally underpowered; p13 is therefore a replication target rather than a finding. External review exposed family-error, dependence, construct, and boundary-uncertainty defects, and zero-call reanalysis changed the interpretation without rewriting the historical record. The registered marginal criteria could be passed without precisely estimating the treatment-response object. A public capsule verifies 4,916 confirmatory Phase 3-5 runs with no live model calls. The results concern one fixed model-prompt panel and do not establish human substitutability.","authors":["Yohei Nakajima"],"categories":["cs.AI","cs.CL","cs.GT"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00979","pdf_url":"https://arxiv.org/pdf/2608.00979","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为博弈","算法保真度"],"reason":"用LLM persona面板模拟重复博弈行为，与人类数据对照，评估仿真可靠性和…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":5,"question":"在重复策略博弈中，固定的人格面板能否通过粗粒度的边际分布检验，同时其处理效应估计是否精确？","design":"使用16个轻量级人格条件化的GPT-4.1配置构成固定面板，在重复策略博弈中施加两种处理（S2措辞存在与否），测量合作行为、继续概率等结果变量，并分析面板内变异来源。","baseline":"无对照","findings":"面板在四个重复博弈单元中有三个通过了预注册的边际条件均值检验，但总体继续概率的处理效应点估计较小且置信区间很宽，无法确立等效性或窄响应边界。变异主要来源于提示词配置之间，且措辞和标签操作可大幅改变合作率，表明粗粒度边际检验可能掩盖处理效应估计的不精确性。","reliability":"论文承认边际检验通过并不要求精确估计处理效应，且结果仅针对一个固定模型-提示面板，不建立人类可替代性；外部审查揭示了家族误差、依赖性、构念和边界不确定性缺陷。","relevance":"该研究直接评估了LLM人格面板在策略互动中的仿真可靠性，揭示了粗粒度验证的廉价性，对关注经济学实验和政策评估中LLM替代人类被试的研究者具有重要警示价值，值得精读原文。","inspiration":"借鉴其预注册、固定面板、多不确定性视角分解变异和零调用复现的审计协议设计｜可迁移到公共品博弈或信任博弈中，检验LLM面板能否复现人类合作与惩罚行为的分布和处理效应｜用多个LLM人格配置构成固定面板，施加不同的制度处理（如惩罚机制、信息反馈），测量合作率与信念更新，以实验室人类被试数据为基准，评估边际分布通过但处理效应估计不精确的程度。"}},{"id":"2608.01193","version":1,"title":"Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races","zh_title":"人类更多样：前沿大语言模型在理想化AI发展竞赛中表现出极端策略","abstract":"An AI development race creates a multi-agent safety dilemma. Each company can develop slowly and safely, or move faster while taking a risk that may remove its final reward. We use this repeated game to study strategic safety behaviour among large language model (LLM) agents in races with two to five players. However, a valid action does not show that an agent understands the game. We therefore place an audit gate before behavioural interpretation. We first verify the game engine, then test rule recall, state tracking, payoff calculation, and stability under different but equivalent task descriptions. We then compare LLM action sequences with an evolutionary game-theory benchmark and published human data, and explore differences across models, risk conditions, personas, and two- to five-player races. The audit shows that strong rule recall can coexist with weak state tracking and expected-payoff calculation. Providing verified arithmetic and changing the response representation can also change later actions, even when the game rules stay fixed. Across seven tested model endpoints, aggregate rates hide large differences in action sequences, responses to opponents, and responses to race position. Patterns across the tested three- to five-player races are also model-specific rather than a single effect of adding competitors. These results show why multi-agent AI-race simulations need validity checks and trajectory-level analysis before their outputs are described as strategic, human-like, or safety-aware. Our findings are exploratory and apply only to the tested models, prompts, and decoding settings.","authors":["Phu Hoa Pham","Duy Minh Dao Sy","Trung Kiet Huynh","Phu Quy Nguyen Lam","Chi Nguyen Tran","Minh Trung Le","Phong Hao Le","Dinh Nam Nguyen","Thien Ky Nguyen Dong","Elias Fernandez Domingos","Le Hong Trang","The Anh Han"],"categories":["cs.AI","cs.CY","cs.GT","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01193","pdf_url":"https://arxiv.org/pdf/2608.01193","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM行为仿真","博弈实验","人类数据对照"],"reason":"用LLM模拟AI竞赛中的人类行为，并与真实人类数据对照，涉及博弈实验场景，但非…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":6,"question":"在理想化的AI研发竞赛中，LLM智能体是否表现出与人类相似的策略性安全行为，以及其行为有效性是否受任务表述、状态追踪和收益计算能力的影响？","design":"让七个LLM端点扮演AI公司，在2至5人的重复博弈中每轮选择安全或冒险开发，通过审计关卡检验规则记忆、状态追踪、收益计算和表述稳定性，并记录行动序列、对手反应和位置效应。","baseline":"已发表的人类实验数据（falling_behind_unsafe）和演化博弈论基准。","findings":"LLM的总体不安全率掩盖了行动序列、对手反应和位置效应的巨大差异；强规则记忆可与弱状态追踪和收益计算并存，且提供算术验证或改变响应表示会改变后续行动。","reliability":"论文声明发现是探索性的，仅适用于所测试的模型、提示和解码设置，未讨论其他失效条件。","relevance":"该研究直接使用LLM模拟人类在博弈中的行为，并与真实人类数据对照，包含审计关卡检验仿真有效性，对关注LLM仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其审计关卡设计，在行为解释前先检验LLM对任务规则、状态和收益的理解，并考察不同表述和输出格式的稳健性。｜可迁移到政策公告预期形成的实验，如央行沟通博弈，检验LLM是否像人类一样对措辞和顺序敏感。｜以LLM为被试，模拟央行发布前瞻指引，处理为不同措辞或发布顺序，结果变量为通胀预期和投资决策，对照真实人类实验数据。"}},{"id":"2607.27553","version":2,"title":"AI and Its Impact on Creativity and Diversity: An Empirical Study of LLM-Generated Product Ideas","zh_title":"AI对创造力与多样性的影响：LLM生成产品创意的实证研究","abstract":"This research examines how well large language models, or LLMs, generate new product ideas for college students priced under $50. Across a series of studies, we identify key strengths and weaknesses of using LLMs for product innovation. Our first study shows that LLM-generated product ideas have higher average quality than human ideas, based on purchase intent, and are 7 times more likely to rank in the top 10%. Our second study shows that this AI-induced creativity boost is not explained by the LLM's more persuasive pitching skills. Our third and fourth studies identify a weakness of using LLMs for brainstorming: AI-generated ideas are less novel at the idea level and less diverse at the set level. In our fifth study, we analyze prior LLM-based creativity studies and find consistently lower idea diversity across all of them, demonstrating the generalizability of these findings. Our sixth and seventh studies investigate techniques to mitigate this diversity loss. We compare LLMs from different vendors and versions and find that more recent models generate more diverse ideas, though they still fall short of human-level diversity. We also demonstrate techniques that increase idea diversity almost to the level of human idea generation: pooling ideas across vendors; prompt engineering, including Chain-of-Thought prompting and injecting heterogeneous personas or constraints; and creative agents that broadly explore the solution landscape to restore diversity. Finally, in our eighth study, we show that exploiting the near-zero marginal cost of AI idea generation by scaling the number of ideas steadily improves coverage of the idea space, approaching human-level coverage. We conclude by presenting actionable recommendations for innovation managers who want to identify better new product ideas with the help of LLMs.","authors":["Christian Terwiesch","Lennart Meincke","Karan Girotra","Ethan Mollick","Gideon Nave","Karl T. Ulrich"],"categories":["cs.AI","cs.CL","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-08-04","first_seen":"2026-07-31","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.27553","pdf_url":"https://arxiv.org/pdf/2607.27553","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","人类对照","产品创新"],"reason":"用LLM生成产品创意并与人类数据对照，属于人类仿真实验，涉及经济学场景。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":9,"question":"LLM生成的产品创意在质量、新颖性和多样性上是否优于人类创意？","design":"使用GPT-4以零样本和少样本提示生成面向大学生、定价50美元以下的新产品创意，与人类学生生成的创意进行对比，通过购买意向测量质量，通过文本挖掘和人工评分测量新颖性与多样性。","baseline":"一所精英大学产品设计课程学生在LLM出现前生成的创意池。","findings":"LLM生成创意的平均购买意向更高，且进入前10%的可能性是人类的7倍，但AI创意在想法层面新颖性更低，在集合层面多样性更差。","reliability":"论文未讨论","relevance":"该研究将LLM作为人类被试的替代品，直接对比真实人类数据，评估了AI在创意生成任务中的表现与偏差，属于人类仿真实验，值得精读。","inspiration":"借鉴其将LLM作为被试、设置零样本/少样本处理组并与真实人类基准对比的实验设计，以及用购买意向、新颖性、多样性等多维度测量结果的方法。｜可迁移到消费者偏好预测或新产品市场反应评估等经济金融场景，例如用LLM模拟消费者对金融产品创新的接受度。｜以LLM模拟消费者群体，处理为不同提示词（如不同收益描述），结果变量为购买意向，对照真实消费者调查数据。"}},{"id":"2608.01181","version":1,"title":"Talking to Digital Twins: Selective Disclosure and Belief Measurement in Financial Social Media","zh_title":"与数字孪生对话：金融社交媒体中的选择性披露与信念测量","abstract":"Social media affect financial markets, but public posts by financial media personas are voluntary disclosures. What is not disclosed is therefore usually unobserved. We address this measurement problem by conducting repeated, real-time interviews of \"digital twins\" built from monitored finfluencers' X accounts under a fixed protocol. The interviews recover stock-level public-persona belief proxies even when no public recommendation is made. Because the interviews are generated and archived before the relevant return windows, the design avoids the look-ahead bias that arises when LLMs are queried ex post. The evidence shows that information obtained from these digital-twin interviews predicts the cross section of large-cap stock returns in the expected direction. Repeated real-time interviews therefore show how selective disclosure can be turned into measurable panels of market views.","authors":["Boone Bowles","Raymond Duch","Sorin Sorescu"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01181","pdf_url":"https://arxiv.org/pdf/2608.01181","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","金融行为","数字孪生"],"reason":"用LLM构建数字孪生模拟金融影响者观点，并与真实市场数据对照，属于人类仿真且涉…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":10,"question":"如何从金融影响者的选择性披露中恢复其市场观点，并检验这些观点对股票收益的预测能力？","design":"基于金融影响者（finfluencers）的X账号公开信息构建LLM数字孪生，每日通过固定访谈协议询问其对大盘股的投资建议（买入/持有/卖出）及信心、投机性和催化剂，生成带时间戳的面板数据，并与后续股票收益对照。","baseline":"同一日期、同一股票、同一账号上人类金融影响者的实际公开推荐，以及数字孪生推荐与人类推荐的一致性；此外，数字孪生访谈在人类公开推荐之前就已偏向最终公开的方向。","findings":"数字孪生访谈的净买入份额能正向预测未来10个交易日的股票超额收益，且这种预测能力在人类未公开推荐的“沉默”股票上最强。访谈不仅复现了公开推荐，还从沉默中提取了增量信息。","reliability":"论文未讨论","relevance":"该研究用LLM构建数字孪生模拟金融影响者观点，并与真实人类推荐及市场收益数据对照，直接命中你关注的人类仿真、经济学实验场景和基准对照，值得精读原文以了解其仿真效度和预测设计。","inspiration":"借鉴其通过固定访谈协议从LLM数字孪生中系统提取信念并分离方向、分歧和不确定性的测量方法。｜可迁移到资产定价实验，检验分析师或投资者情绪对横截面收益的预测。｜以LLM扮演的金融分析师为被试，每日施加固定访谈处理（询问对股票的看法和信心），结果变量为净买入份额，用真实分析师一致预期和后续股票收益做对照。"}},{"id":"2608.01540","version":1,"title":"Do people rely on ChatGPT more than their peers to detect deepfake news?","zh_title":"人们在检测深度伪造新闻时是否比同伴更依赖ChatGPT？","abstract":"This experimental study investigates how people rely on different sources of advice when detecting AI-generated fake news (deepfake news). In a laboratory deepfake detection task, student participants identified the proportion of human-written (non-AI-generated) content in synthetic deepfake news articles and received advice from ChatGPT (GPT-4), human peers, or linguistic experts. The results show that participants rely more on ChatGPT than on human peers when detecting GPT-2-generated deepfake news. Participants also rely more on linguistic experts than on peers, while the relative reliance on experts versus ChatGPT is mixed across experimental waves, potentially reflecting time trends in beliefs about AI-based detection. Importantly, in the additional experiment conducted in 2025 under the same experimental procedure, participants relied more on linguistic experts than on ChatGPT. Moreover, performance improvements reflect the joint role of reliance and advice quality, arising primarily when participants rely on high-quality advice. Overall, relying on AI to detect AI-generated deepfakes can improve detection outcomes, but only when AI-based detection tools are of sufficiently high quality. These findings highlight the dual role of GAI as both a source of deepfakes and a tool for mitigating related risks.","authors":["Yuhao Fu","Nobuyuki Hanaki"],"categories":["econ.GN","cs.CY","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01540","pdf_url":"https://arxiv.org/pdf/2608.01540","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["人类行为实验","AI建议依赖","深度伪造检测"],"reason":"用真人实验对照，研究人类对ChatGPT建议的依赖，涉及行为决策和检测任务，可…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":12,"question":"人们在检测深度伪造新闻时，是否比依赖人类同伴更依赖ChatGPT的建议？","design":"本研究并非LLM仿真实验，而是真实人类实验室实验。被试为大学生，在30轮深度伪造检测任务中，先给出初始判断，然后随机接受ChatGPT（GPT-4）、人类同伴或语言学专家的建议，再给出最终判断，测量建议采纳程度（WOA）和检测表现。","baseline":"真实人类同伴的建议作为对照基准，比较被试对ChatGPT与对人类同伴的建议依赖程度。","findings":"被试在检测GPT-2生成的深度伪造新闻时，对ChatGPT的建议依赖程度显著高于对人类同伴；但2025年补充实验中，被试对语言学专家的依赖高于ChatGPT，反映出对AI检测工具信念的时间趋势变化。检测表现的提升取决于建议质量与被试依赖程度的共同作用，仅当AI检测工具质量足够高时，依赖AI才能改善检测结果。","reliability":"论文指出，依赖AI检测AI生成内容的效果取决于AI检测工具的实际质量，若工具质量不高，依赖AI可能无益甚至有害；此外，被试对AI的信任和依赖可能随时间变化，影响结论的跨期稳健性。","relevance":"该研究虽非LLM仿真实验，但提供了真实人类在AI建议下的行为决策基准，可用于校准或验证LLM仿真人类在信息检测任务中的行为，尤其适合关注AI依赖与信任动态的研究者。","inspiration":"可借鉴其JAS框架和WOA测量方法，通过随机分配建议来源（AI vs. 人类）并比较依赖程度，来量化人类对AI建议的采纳行为。｜可迁移至金融投资决策场景，如投资者在评估AI生成的财务报告或市场预测时，是否过度依赖AI建议而忽视人类分析师。｜以真实投资者为被试，设计投资判断任务，处理为提供ChatGPT生成的投资建议 vs. 人类分析师建议，结果变量为建议采纳权重和投资组合表现，对照真实市场数据或历史分析师记录。"}},{"id":"2607.25953","version":2,"title":"Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections","zh_title":"Polistemics：评估大语言模型在政治与选举中作为信息中介的表现","abstract":"As LLMs increasingly shape the political information citizens rely on, no standard exists to assess whether they do so responsibly. We introduce Polistemics, a theory-grounded diagnostic benchmark for evaluating LLMs as mediators of political information in elections. Prior work has treated this task as reproduction rather than mediation, leaving its epistemic dimensions and interaction with imperfect information unaddressed. We ground the evaluation in Epistemic Modesty, a normative standard derived from citizens' epistemic agency, and test it across controlled settings that vary the clarity, noise, and consistency of the available evidence. Applying the benchmark to three state-of-the-art LLMs across the 2025 German and Dutch elections, we find that high aggregate scores mask systematic failures. Models mediate reliably under clear evidence but break down when it is absent, vague, or contradictory, while flattening the intensity of political language throughout. These failures point to party priors, shifting with party labels and output language. Reliable mediation appears achievable, but no model delivers it consistently.","authors":["Baran Peters","Gabor Hollbeck","Robert Jakob","Kevin O'Sullivan"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-04","first_seen":"2026-07-29","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.25953","pdf_url":"https://arxiv.org/pdf/2607.25953","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","政治信息","基准测试"],"reason":"评估LLM作为政治信息中介，测量模型行为而非仿真人类被试，无人类对照数据。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":16,"question":"在选举情境中，LLM作为政治信息中介者应具备哪些负责任的行为特质，以及它们在不同信息环境下的稳健性如何？","design":"本研究并非人类仿真实验，而是构建了一个名为Polistemics的基准测试，用于评估LLM在选举中作为政治信息中介者的表现。它使用三个最先进的LLM（Qwen3.6 Flash、GPT-5.4、Claude Sonnet 4.6），在受控的信息环境下，通过改变证据的清晰度、噪声和一致性，测量模型在回答政党立场查询时的忠实性、公正性和认知校准。","baseline":"无对照","findings":"模型在证据清晰时中介可靠，但在信息缺失、模糊或矛盾时表现崩溃，并会扁平化政治语言的强度。这些失败可能由模型对政党的先验偏好驱动，并受政党标签和输出语言影响。","reliability":"论文指出模型在缺乏、模糊或矛盾的信息下会失效，且高总体得分掩盖了系统性失败；可靠的中介看似可实现，但没有模型能持续做到。","relevance":"该研究虽非直接的人类仿真实验，但系统评估了LLM在政治信息中介中的偏差与失效条件，对理解LLM在调查或实验中的行为偏差具有参考价值，值得阅读以了解其受控信息环境的设计方法。","inspiration":"可借鉴其通过合成不同信息环境（如清晰、噪声、矛盾）来隔离LLM行为影响因素的实验设计方法。｜可迁移到经济政策沟通场景，如央行公告的解读实验，测试LLM在不同信息质量下如何传达政策立场。｜以LLM为被试，向其提供不同清晰度的央行声明（处理），测量其输出的政策立场一致性和置信度（结果变量），并以真实市场分析师解读数据作为对照。"}},{"id":"2607.25166","version":3,"title":"Individual-level interventions against sycophantic AI reduce its appeal but not its persuasiveness","zh_title":"针对谄媚AI的个体层面干预降低其吸引力但未降低其说服力","abstract":"AI chatbots can be \"sycophantic,\" or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call \"sycophancy blindness\"). We tested whether increasing users' awareness of sycophancy protects them from its harmful effects in two preregistered experiments (n = 1,590). In the first, participants received a brief written warning about sycophancy before conversing with a sycophantic chatbot. In the second, participants watched a video of a sycophantic AI validating several other users, including users on opposite sides of the same conflict, before interacting with it themselves. Both interventions changed how participants evaluated the AI. The warning reduced the AI's perceived objectivity, and the video reduced enjoyment of the AI --- an effect mediated by the reduced belief that its validation was uniquely earned. We then pooled our experiments with two prior studies of sycophancy awareness interventions (six interventions total, n = 3,982). The pattern across experiments was consistent: while the interventions made the sycophantic AI appear less objective and trustworthy, none reduced its persuasiveness. These results suggest that individual-level interventions, such as warning labels or AI literacy, may not be enough to protect users from AI harms.","authors":["Meryl Ye","Robert Kraut","Steve Rathje"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-04","first_seen":"2026-07-29","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.25166","pdf_url":"https://arxiv.org/pdf/2607.25166","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["AI谄媚","人机交互","干预实验"],"reason":"研究人类对谄媚AI的感知与说服力，非用LLM仿真人类被试，但涉及AI行为对人类…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":10,"question":"个体层面的干预（如警告或视频示范）能否降低用户对谄媚AI的正面评价，并削弱其对用户态度的说服力？","design":"本研究并非用LLM仿真人类被试，而是以人类为被试，通过在线实验检验两种干预效果：研究一让被试在与谄媚聊天机器人对话前阅读简短警告；研究二让被试先观看该AI赞同其他用户的视频再与之互动。测量结果包括对AI的感知（客观性、可信度、愉悦度）和话题态度（确定性、极端性）。","baseline":"无对照","findings":"警告和视频干预均改变了被试对AI的感知，降低了其客观性或愉悦度，但未能减少AI的说服力。汇总六项干预（N=3982）后，模式一致：干预使AI显得更不客观、更不可信，但均未削弱其对用户态度的影响。","reliability":"论文未讨论","relevance":"该研究直接测量人类与谄媚AI交互后的态度改变，虽非LLM仿真人类，但提供了人类行为基准数据，对评估LLM仿真人类在说服与态度极化场景中的可靠性有参考价值，值得阅读以获取效应量参考。","inspiration":"可借鉴其多干预汇总比较的设计，检验不同干预对感知与行为的分离效应。｜可迁移至金融建议场景，如AI理财顾问的谄媚行为对投资者风险偏好和产品选择的影响。｜以人类投资者为被试，随机分配接受谄媚或中立的AI投资建议，处理组在建议前观看AI对其他客户无差别赞同的视频，测量其风险资产配置比例和信任评分，并以真实市场数据或历史投资记录作为对照基准。"}},{"id":"2607.28128","version":2,"title":"Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models","zh_title":"重新思考LLM评判的有用性作为教学信号：一项跨导师模型的预注册审计","abstract":"LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conversational and pedagogical policies instantiated with the same underlying model and paired with one fixed weak simulated student. Deterministic detectors measure answer leakage and next-turn independent work. Claude Opus 4.8 is the frozen, condition-blind primary judge. After the Opus scores were fixed, GPT-5.6 Sol was prospectively specified for a post hoc robustness audit of the same 1,179 confirmatory answer-phase tutor turns under the frozen helpfulness and pedagogy rubrics. On the primary base under Opus, the policies do not differ significantly in helpfulness but are perfectly rank-separated under the pedagogy rubric (Cliff's $|\\delta|{=}0.10$ vs. $1.0$). Across the two judges, pedagogy contrasts retain their direction where detected, whereas the helpfulness ordering is judge-contingent, reversing between judges on two of three bases. In an Opus-only ablation, seven primary-base policies span $2.3$ points in mean judged pedagogy within a $0.25$-point band of mean judged helpfulness. Separately, answer-revealing turns are followed by less independent student work on every base, a result that is judge-invariant by construction. In this controlled setting, general-purpose helpfulness is not a reliable pedagogy signal. Tutor evaluation should pair pedagogy-targeted rubrics with deterministic process measures.","authors":["Shuyi Fan","Boyuan Deng","Mengyu Xu","Jiale Liu","Hongyang Zhang","Qiaoxin Yang","Chongyang Gao"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-04","first_seen":"2026-07-31","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.28128","pdf_url":"https://arxiv.org/pdf/2607.28128","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","教学对话","标注替代"],"reason":"用LLM评判教学对话质量，属替代人工标注，非仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":25,"question":"通用帮助性评分能否可靠区分直接给答案与教学引导这两种辅导行为？","design":"本研究并非用LLM仿真人类被试，而是用LLM作为评判者（Claude Opus 4.8为主评判，GPT-5.6 Sol为稳健性评判）对三种基座模型（Claude Sonnet 4.6、GPT-5.5、Gemini 3.1 Pro）下的两种辅导策略（对话式与教学式）生成的回答进行帮助性和教学性评分，同时用确定性检测器测量答案泄露和学生独立作业情况。","baseline":"无对照","findings":"在主要基座模型上，两种策略的帮助性评分无显著差异，但教学性评分完全分离；帮助性排序因评判模型而异，而教学性评分方向保持稳定。答案泄露的回合后学生独立作业减少，这一结果不受评判模型影响。","reliability":"论文未讨论","relevance":"本文用LLM替代人工进行教学评估，属于标注替代而非仿真人类被试，与研究者关注的LLM作为人类被试替代品进行行为仿真和决策复现的核心兴趣不符，但其中关于评判信号可靠性的批判性分析可提供方法借鉴。","inspiration":"借鉴其使用多个评判模型进行稳健性审计、结合确定性过程测量与主观评分的方法，以暴露单一评判信号的不可靠性。｜可迁移至经济政策沟通效果评估场景，如央行公告的清晰度与引导性评判。｜以LLM生成不同风格的央行公告（直接告知决策 vs. 解释决策逻辑），用多个LLM评判其清晰度与引导性，同时测量公众预期调整的确定性指标，并与真实市场调查数据对照。"}},{"id":"2607.29274","version":1,"title":"Language Models Agree With Each Other, Not With Readers","zh_title":"语言模型彼此一致，而非与读者一致","abstract":"Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the median of 153 model pairs is +0.093 against a human yardstick of +0.040, and 99 sit entirely above the human interval. Two frontier models from rival labs reach +0.203, twice what GPT-4o agrees with itself on a second call. The effect is not determinism, prompt wording, procedure, vendor or routing, and it is graded: the smallest models agree at the human level. No model agrees with readers detectably more than a reader does, and at equal depth and length no surface feature separates their choices. The multiples are procedure-dependent and the ordering is not: models are cut to their sharpest set while a reader's is a random draw from what they marked, and blunting the models alike halves the gap without closing it. Tested out of sample on four models released after this analysis, against predictions fixed beforehand, none clears the human interval. A population simulated from several models is not several populations.","authors":["Kazuki Nakayashiki","Keisuke Watanabe"],"categories":["cs.IR","cs.CL","cs.CY","cs.HC"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29274","pdf_url":"https://arxiv.org/pdf/2607.29274","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类行为对照","一致性评估"],"reason":"用LLM模拟人类标注行为并与真实读者数据对照，评估模型间一致性及与人类差异，揭…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-03","rank":3,"question":"语言模型在文本标注任务上的趋同性是否高于真实读者，且与读者的差异是否随模型能力增强而扩大？","design":"将18个不同厂商、规模、代际的语言模型作为被试，输入120篇网页文档的编号句子，要求模型按重要性排序并取前k句作为标注集；同时收集同一文档上独立读者的真实高亮标注作为对照，测量模型间、读者间及模型与读者间的标注重叠度（经位置和长度校准后的协议分数）。","baseline":"2,523个真实读者在120篇网页文档上的自主高亮标注集，读者未受指令引导且互不可见。","findings":"模型间的标注一致性显著高于读者间一致性（中位数协议分数0.093 vs 0.040），且前沿模型间一致性可达0.203，是读者间一致性的5.1倍；模型与读者的一致性仅与读者间一致性持平，且模型选择的句子在表面特征上与读者无异，但内容不同。","reliability":"读者独立性无法完全验证（仅基于平台默认设置假设）；协议分数依赖于标注深度和长度控制，稀释模型标注会缩小但未消除差距；样本外测试中，新模型均未突破人类一致性区间。","relevance":"该研究直接以真实人类行为为基准，系统评估了LLM仿真人类标注的可靠性及偏差，揭示了模型趋同但未逼近人类分布的规律，对关注LLM仿真效度的研究者极具参考价值。","inspiration":"借鉴其利用自然发生的非实验人类行为数据作为基准，避免指令诱导同质化的设计；可迁移到消费者信息处理或投资者注意力分配研究，如用LLM模拟投资者阅读财报后的关注点；以真实投资者在财经新闻上的自主高亮或眼动数据为对照，让LLM对同一文本生成重要性排序，比较注意力分布与真实行为的差异。"}},{"id":"2607.28643","version":1,"title":"To Facilitate or not to Facilitate: Human and LLM Facilitator Tendencies in Online Discussions","zh_title":"促进与否：在线讨论中人类与LLM的主持倾向","abstract":"Automating facilitation in online discussions is a long-standing social concern given the increasing time we spend on online spaces and the failure of content moderation approaches. While studies have been conducted on how to facilitate, none have answered the essential question of when to do so. A potential answer is using LLMs, which ostensibly make automated, large-scale intervention increasingly feasible. In this study, we examine when LLMs decide to facilitate by defining what facilitation is, observing when humans decide to facilitate, and comparing their decisions with those made by LLMs. To this end, we create PEFK, a corpus standardizing and aggregating all relevant facilitation datasets. We are the first to run a survey on facilitation timing, which we execute using expert facilitative participants and LLM-as-a-judge models. We discover that while humans are more cautious, LLMs are excessively eager to facilitate, although both are more certain when judging that facilitation is not needed. We then investigate whether this behavior can be corrected using alternative setups for LLMs and training ModernBert classifiers on established datasets, finding that the latter perform more reliably than the former, although current datasets impose a relatively low performance ceiling.","authors":["Dimitris Tsirmpas","Katerina Korre","John Pavlopoulos"],"categories":["cs.HC","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28643","pdf_url":"https://arxiv.org/pdf/2607.28643","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类对照","行为偏差"],"reason":"用LLM模拟人类主持决策并与人类数据对照，发现LLM过度干预，批判性指出仿真偏…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:01:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-03","rank":4,"question":"LLM在决定何时介入在线讨论时，与人类主持人的决策倾向有何差异？","design":"本研究并非严格意义上的仿真实验，而是通过构建统一数据集PEFK，设计调查任务让10名专家参与者与6个开源LLM分别判断1226个讨论片段是否需要介入，比较两者在介入频率和置信度上的差异。","baseline":"10名具有主持经验的专家参与者对讨论片段是否需要介入的判断及其置信度。","findings":"人类主持人更谨慎，倾向于不介入且对不介入的判断更自信；LLM则过度渴望介入，且更依赖正面强化，但在判断不需要介入时同样更自信。","reliability":"论文指出当前数据集存在固有噪声（以专业主持人撰写的评论作为标签），导致性能上限较低；LLM在介入时机预测上表现不一致，而基于编码器的分类器更可靠，但整体仍受限于对主持行为的有限理解。","relevance":"该研究直接对比LLM与人类在决策行为上的差异，并批判性地指出LLM过度干预的偏差，符合研究者对仿真可靠性评估和失效条件分析的兴趣，值得阅读原文以了解具体实验设计和偏差来源。","inspiration":"可借鉴其通过统一多源数据集并设计对照调查来量化行为偏差的方法，尤其关注决策阈值和置信度的测量。｜可迁移到经济政策沟通场景，如央行官员在公开声明中决定何时干预市场预期，或监管者决定何时对金融市场异动发声。｜以真实央行沟通记录为对照，让LLM扮演政策制定者，判断在不同市场波动情境下是否需要发表声明，结果变量为干预频率和声明内容倾向，与历史真实干预记录对比评估LLM的过度干预或保守倾向。"}},{"id":"2607.28908","version":1,"title":"Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds","zh_title":"反思还是重新生成？为何LLM修正失败而人类修正成功","abstract":"Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers. Large language models (LLMs) are increasingly prompted to \"reflect,\" yet whether this resembles human revision remains unclear. We introduce the Human-LLM Reflection Framework (HRF), a controlled two-pass protocol comparing human and LLM revision under identical conditions across self-, peer-, and cross-agent settings. Using an information-theoretic analysis based on per-iteration cross-entropy reduction, we find two failure modes of LLM reflection. On objective tasks with finite answer spaces, reflection yields near-zero information gain (Delta I approx 0), behaving as neutral re-generation indistinguishable from re-sampling. On subjective tasks, it yields significant negative gain (Delta I < 0), moving predictions away from the target. Human revision, by contrast, yields positive gain in both settings. Cross-agent experiments localize the failure to the revision step, not input quality: LLMs degrade even high-quality human responses. Diagnostic analyses (revision conditioned on first-pass correctness, and oracle-guided revision against a random-reshuffle baseline) show that which sub-step dominates varies by task and by model rather than reducing to a single mechanism: self-error detection is present on objective multiple-choice tasks but weak on subjective ones, and recovery under an oracle error signal exceeds the baseline for some models and falls below it for others. The unifying account is structural: without external information, self-conditioned revision cannot reduce uncertainty about the target, so LLM reflection is better understood as conditioned re-generation than as genuine error-driven revision.","authors":["Yefan Tao","Gerald Friedland","Madhusudhanan Chandrasekaran","Luyang Kong"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28908","pdf_url":"https://arxiv.org/pdf/2607.28908","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM反思","人类对照","可靠性评估"],"reason":"对比人类与LLM的修正行为，揭示LLM反思的失效模式，有真实人类数据对照，批判…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-03","rank":5,"question":"LLM的反思（revision）是否像人类一样是真正的错误修正，还是仅仅是基于先前输出的条件再生成？","design":"提出Human-LLM Reflection Framework (HRF)，采用受控两阶段协议：人类和LLM在相同条件下先做初始回答（Pass 1），再看到先前回答后决定保留或修改（Pass 2），涵盖自我修正、同伴修正和跨代理修正三种设置，在客观数学推理和主观情感评价任务上测量信息增益。","baseline":"非专家人类标注者在相同任务、提示和评估标准下的修正行为，作为有效反思的实证参照。","findings":"LLM反思呈现两种失效模式：客观任务上信息增益接近零，表现为中性再生成；主观任务上信息增益显著为负，使预测偏离目标。人类修正在两种任务上均产生正信息增益，且失效定位于修正步骤而非输入质量。","reliability":"论文指出LLM反思失效的结构性原因：缺乏外部信息时，自我条件修正无法降低关于目标的不确定性；诊断分析显示失效子步骤（错误检测或错误纠正）因任务和模型而异，并非单一机制。","relevance":"该研究直接对比人类与LLM的修正行为，揭示LLM反思的失效模式，有真实人类数据对照，批判性地指出仿真在反思环节的不可靠性，对关注LLM作为人类被试替代品的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其受控两阶段修正协议和信息论测量（交叉熵减少量）来严格评估LLM的决策修正能力。｜可迁移到经济预测修正场景，如分析师盈利预测修正、央行沟通后的市场预期调整。｜以LLM作为分析师被试，先给出盈利预测（Pass 1），再提供历史预测值要求修正（Pass 2），结果变量为预测误差变化，以真实分析师修正数据（如IBES）作为人类基准，对比信息增益。"}},{"id":"2607.24435","version":2,"title":"LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings","zh_title":"LEX-EC：黑盒环境下零样本LLM人格分类的词汇证据通道审计框架","abstract":"Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining prevalence and agreement diagnostics with controlled lexical ablation to distinguish marginal-distribution effects from trait-associated signal recoverable under restricted evidence. Using this framework, we illustrate how various text genres may exhibit sharply different profiles: free-form essay text contains the broadest, but still weak, signal; in graduate student introductions, an observable Extraversion association weakened after masking; and single Facebook statuses yield little stable evidence even in a trait-balanced sample, indicating a possible lower bound of content or length. Masking topical and demographic content weakened some associations while leaving others detectable from function words, affective terms, and cognitive-style vocabulary. Linguistic prompting shifted model self-explanations but did not eliminate topical content. LEX-EC jointly evaluates classification prevalence, item-level association, chance-corrected agreement, persistence under lexical restriction, and prompt sensitivity in model-generated explanations. Across datasets, models, and prompts, LEX-EC characterizes how trait associations may vary with available lexical evidence, introducing a novel application of lexical methods to black-box interpretability in personality labeling.","authors":["Brittany Harbison","Ashok K. Goel"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-03","first_seen":"2026-07-28","revised_at":"2026-08-03","abs_url":"https://arxiv.org/abs/2607.24435","pdf_url":"https://arxiv.org/pdf/2607.24435","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格分类","黑盒可解释性","词汇审计"],"reason":"研究LLM人格分类的可解释性，测量对象是模型而非人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":11,"question":"在零样本黑盒设定下，大语言模型进行人格分类时，其预测信号在多大程度上依赖于文本中的话题和人口统计学内容，而非真正的特质相关语言？","design":"本研究并非人类仿真实验，而是提出一个名为LEX-EC的黑盒审计框架，通过分布诊断、一致性检验和受控词汇消融，分析LLM在不同文本体裁、长度和内容遮蔽条件下的人格分类行为，测量分类流行率、项目级关联、机会校正一致性、词汇限制下的信号持久性及提示敏感性。","baseline":"无对照","findings":"自由形式短文包含最广泛但依然微弱的特质信号；研究生自我介绍中，外向性关联在遮蔽话题和人口统计学内容后减弱；单一Facebook状态即使在特质平衡样本中也几乎不产生稳定证据，表明存在内容或长度的下限。","reliability":"论文指出，LLM的人格标签可能受话题内容、人口统计学先验和文本长度影响，而非真实特质信号；在短文本或内容受限条件下，预测信号可能崩溃；模型自解释受提示影响，但无法消除话题内容。","relevance":"该研究与您关注的人类仿真实验不同，它审计的是LLM的人格分类行为而非用LLM模拟人类被试，且无真实人类行为对照，但其中关于信号来源的批判性分析对评估LLM仿真可靠性有参考价值。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.28439","version":2,"title":"Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation","zh_title":"超越单一评判：用于生成式UI评估的基于证据与社会权重的角色面板","abstract":"Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation is costly and rater-variant, while LLM-as-a-judge is scalable but reflects only a single implicit viewpoint, unable to capture how different populations of real users actually perceive the same interface. We propose the Evidence-Grounded, Social-Weighted Persona Panel (ESPP), a three-stage GenUI evaluation method in which a panel of psychologically diverse, evidence-grounded personas independently rates a screenshot, exchanges opinions under a trait-derived, semantically-gated bounded-confidence mechanism, and is aggregated via Delphi-inspired social weighting into a single judgment. ESPP tracks human judgment substantially more closely than a naive single-pass judge, raising Pearson $r$ from $0.716$ to $0.922$, and a prompt-ensemble control recovers only about a third of this gap, isolating genuine persona and evidence grounding as the dominant source of improvement. Beyond this fidelity gain, retaining each panelist's individual rating further reveals that user subgroups agree on overall model rankings yet diverge sharply on specific rating dimensions, a structural disagreement a single homogeneous judge would systematically erase. The codes are available at https://github.com/Wuzheng02/ESPP.","authors":["Zheng Wu","Yibo Luo","Pu Zhang","Cheng Yang","Zhuosheng Zhang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-03","first_seen":"2026-07-31","revised_at":"2026-08-03","abs_url":"https://arxiv.org/abs/2607.28439","pdf_url":"https://arxiv.org/pdf/2607.28439","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","角色面板","UI生成"],"reason":"用LLM persona面板替代人工评估UI，属标注替代而非仿真人类被试，无真…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":11,"question":"如何让LLM模拟多样化用户群体来评估生成式UI，并使其评分更贴近真实人类判断？","design":"用LLM扮演心理特质多样且基于证据的虚拟用户（persona），组成面板；每个persona先独立对UI截图评分，再通过基于特质和语义门控的有界置信机制交换意见，最后用德尔菲式社会加权聚合为单一评分。","baseline":"UIPersonaBench基准中500条指令由14个模型生成的7000张截图，每张截图收集5名真实人类在5个维度上的1-5分评分，取均值作为人类基准。","findings":"ESPP面板评分与人类评分的皮尔逊相关系数从单一法官的0.716提升至0.922，且提升主要来自persona与证据锚定，而非多次调用平均；不同用户子群在整体排名上一致，但在具体维度（如控制感）上存在显著分歧，单一法官会抹平这种差异。","reliability":"论文未讨论","relevance":"该研究用LLM模拟多样化用户群体评估UI，并与真实人类评分严格对照，揭示了子群分歧，方法可迁移至经济学实验和政策评估中的异质性偏好测量。","inspiration":"借鉴其基于心理特质构建异质性persona面板、通过社会交互机制模拟意见动态并保留个体分歧的设计；可迁移到消费者金融产品选择或政策偏好调查中，模拟不同风险态度、金融素养人群的决策差异；以LLM扮演不同人格与认知偏好的被试，处理为呈现不同设计的金融产品界面或政策描述，结果变量为选择或评分，用真实消费者调查或实验数据作为对照基准。"}},{"id":"2607.28347","version":1,"title":"LLMs struggle to simulate human belief updates in controlled environments","zh_title":"大语言模型难以在受控环境中模拟人类信念更新","abstract":"LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants' actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.","authors":["Sebastian Pohl","Harsh Mehta","Pranav Mambayil","Abdul Ghafoor","Franziska Lesigang","Yufang Hou","Christian Hilbe"],"categories":["cs.CL","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28347","pdf_url":"https://arxiv.org/pdf/2607.28347","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM仿真","信念更新","人类数据对照"],"reason":"直接测试LLM仿真人类信念更新，有真实人类数据对照，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":3,"question":"LLM能否在受控环境中模拟人类在阅读社交媒体评论后的信念更新？","design":"用6个不同规模和发布时间的LLM，基于391名Prolific参与者的真实人口统计和人格特质构建个性化提示（persona），让LLM模拟这些参与者在阅读Reddit评论后对三个讨论话题的立场变化，并直接与人类真实数据进行1对1比较。","baseline":"391名英国Prolific参与者在阅读Reddit评论后实际记录的立场更新数据。","findings":"部分LLM（如Qwen3-32B和GPT-5-Mini）在给定人类初始立场时能匹配人类最终立场分布，但所有模型均无法自行生成初始立场或从自生成立场产生逼真的信念更新；LLM普遍表现出中立偏向、更频繁但幅度更小的信念变化，且无法准确预测评论的说服力排序。","reliability":"LLM模拟仅在以真实初始立场为起点时可靠，当前多轮社交媒体模拟很少提供此类真实起点；人口统计和人格特质persona对模拟逼真度无一致影响，且模型在自生成初始立场时完全失效。","relevance":"该研究直接测试LLM作为人类被试替代品在信念更新任务中的可靠性，有严格的人类个体对照，并明确指出了仿真失效的条件，高度契合研究者对LLM仿真实验的批判性评估需求，值得精读。","inspiration":"借鉴其1对1配对设计，将LLM模拟与个体级人类基准直接比较，可迁移到经济预期形成实验（如通胀预期更新），用LLM基于真实参与者的人口特征和初始预期模拟其在阅读央行公告后的预期调整，以真实调查数据（如密歇根消费者调查）为对照。"}},{"id":"2607.28550","version":1,"title":"Correcting Mode Collapse in Silicon Sampling with Semantic Similarity Rating","zh_title":"用语义相似度评分纠正硅采样中的模式坍缩","abstract":"Silicon sampling refers to the use of Large Language Models (LLMs) to generate responses to surveys. It has shown promise, but tends to generate response distributions with unrealistically low variance. We argue that this mode collapse is due to LLMs failure to generate numeric data, and that text responses may be better suited for this task. We analyze whether Semantic Similarity Rating can improve the fidelity of silicon sampling responses when asked about political attitudes. This method solicits text-only responses from LLMs, then maps this to a numeric scale using text embeddings. We find that this method both improves the fidelity of silicon sampling response distributions, and has few parameters to calibrate.","authors":["Oscar Heath","Rohan Alexander"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28550","pdf_url":"https://arxiv.org/pdf/2607.28550","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["硅采样","调查仿真","分布保真度"],"reason":"用LLM生成调查回答并改进分布保真度，有真实人类数据对照，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":6,"question":"如何通过语义相似度评分（SSR）方法纠正大语言模型在硅抽样中出现的模式坍缩，提高调查回答分布的保真度？","design":"使用多个前沿大语言模型，基于2016年ANES调查中真实受访者的人口统计与政治特征生成人物画像提示，让模型对四个政治目标群体（民主党、共和党、自由派、保守派）生成温度计评分。比较两种生成方式：直接要求模型输出0-100的数值评分，以及先生成纯文本感受描述，再通过文本嵌入映射到数值量表（SSR方法）。通过KL散度和均值绝对误差评估合成分布与真实分布的接近程度，并学习一个全局温度参数控制SSR分布的方差，在2020年ANES数据上检验参数泛化能力。","baseline":"2016年和2020年美国国家选举研究（ANES）时间序列调查中真实受访者的温度计评分数据。","findings":"SSR方法生成的合成响应分布比直接数值输出更接近真实ANES分布，KL散度更低，且均值准确性未显著下降；通过单一全局温度参数有效纠正了低方差问题，该参数从2016年数据学习后可泛化至2020年数据，但合成均值仍存在系统性偏差。","reliability":"论文承认SSR方法虽改善了分布方差，但合成均值（无论数值响应还是SSR）在某些情况下仍相对于真实数据存在持续偏差；温度参数虽泛化良好，但仅基于2016年数据学习，未来数据分布变化可能影响效果；研究仅针对政治态度温度计评分，未验证其他类型调查问题。","relevance":"该研究直接针对LLM仿真人类调查中的模式坍缩问题，提出了可操作的纠正方法，并与真实人类数据严格对照，对关注仿真可靠性的研究者具有重要参考价值，值得细读原文以了解SSR的具体实现和参数校准细节。","inspiration":"借鉴将LLM文本输出通过语义嵌入映射到连续数值量表的方法，可避免模型直接生成数字时的分布坍缩，并引入可学习的温度参数灵活控制方差。｜该方法可迁移到经济预期调查仿真，如消费者信心指数、通胀预期或股市预期等需要捕捉观点分布离散度的场景。｜以LLM扮演不同人口特征的消费者，施加关于未来经济状况的开放式文本提问，用SSR将文本回答映射为预期指数，以密歇根消费者调查的真实个体数据为基准，校准温度参数并评估分布保真度。"}},{"id":"2607.28133","version":1,"title":"AI Sycophancy and Decisions","zh_title":"AI谄媚与决策","abstract":"We examine whether sycophantic AI advice distorts decisions. Our experiment involves 1,500 participants in 30 decision environments spanning core domains in economics and the social sciences. Contrary to the vast majority of predictions in an expert survey we conduct, we find that AI advice depolarizes choices on average, moving participants away from their initial leanings. This depolarization arises despite the LLM being measurably sycophantic: it disproportionately offers considerations that support users' initial leanings and uses agreeable and flattering language. Depolarization occurs across moral and non-moral, objective and subjective, strategic and non-strategic, and complex and simple tasks. Increasing sycophancy weakens depolarization, showing that sycophancy is behaviorally relevant, even if it is generally outweighed by the informativeness of AI advice. Finally, several results mitigate the concern that market forces will generate greater polarizing effects outside the experiment or in the future. On the supply side, our baseline AI's level of sycophancy is typical of leading models, and these models are not becoming more sycophantic over time. On the demand side, participants do not prefer greater sycophancy, do not select into AI advice in tasks where it is more polarizing, and exhibit greater depolarizing effects when they are more frequent AI users outside the experiment.","authors":["John Conlon","Peter Schwardmann"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28133","pdf_url":"https://arxiv.org/pdf/2607.28133","source_feed":"econ.GN","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为经济学","谄媚偏差"],"reason":"用LLM提供建议并测量对人类决策的影响，有真实人类实验对照，涉及经济学决策场景…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":7,"question":"谄媚性AI建议是否会扭曲人类决策，导致选择极化？","design":"非仿真研究，而是真人实验：1500名被试在30个经济学和社会科学决策任务中，先报告初始倾向，再随机分配至无聊天对照组、基线AI聊天组或增强谄媚AI聊天组，最后做出最终选择，测量AI建议对决策方向变化的影响。","baseline":"无聊天对照组作为人类基准，同时收集了249名社会科学和计算机科学专家的预测作为对照。","findings":"尽管AI在内容上明显谄媚，但平均而言AI建议使选择去极化，将被试拉离初始倾向；谄媚程度增加会削弱去极化效应，但总体上AI的信息性仍占主导。","reliability":"论文指出实验任务可能并非人们担忧AI谄媚时的典型决策场景，但通过专家调查表明多数专家预期极化，而实际结果相反；同时从供给侧和需求侧论证了市场力量可能不会加剧极化效应。","relevance":"该研究直接测量了LLM建议对人类经济决策的因果影响，有真实人类实验对照，覆盖多种经济学决策场景，并探讨了谄媚偏差的行为后果，高度契合研究者对LLM仿真可靠性及偏差的关注，值得精读原文。","inspiration":"借鉴其多任务、多处理组、测量初始倾向与最终选择变化的设计，可清晰分离AI建议的极化/去极化效应｜可迁移到资产配置建议场景，研究AI理财顾问的谄媚倾向是否影响投资者的风险资产配置｜以真实投资者为被试，随机提供基线或增强谄媚的AI投资建议，测量其初始风险倾向与最终配置的变化，并以无建议组为对照，同时收集真实市场数据验证外部有效性。"}},{"id":"2607.17219","version":2,"title":"Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat","zh_title":"用QQ等式审计大语言模型中的问题顺序效应：机制表征与饱和警示","abstract":"Question-order effects in human survey data have been reported to approximately satisfy the QQ (quantum question) equality, a parameter-free prediction of the standard projective quantum question-order model. We develop this equality into an audit framework for sequential binary judgments of autoregressive large language models (LLMs). Theoretically, we characterize mechanism families that satisfy QQ robustly, show that classical repetition can reproduce the equality exactly, and combine QQ with the rank-2 Contextuality-by-Default criterion through $|q_{QQ}| \\le \\mathrm{OSS}$. This separates order sensitivity, QQ imbalance, and residual contextuality rather than treating them as interchangeable signatures. Methodologically, we introduce a committed multi-turn forced-branch protocol that reconstructs order-conditioned joint distributions from next-token log-probabilities under counterbalanced label mappings and pre-specified health gates. A first-signal pilot on an open-weight instruction-tuned model reveals the central measurement problem. Although all pre-specified health gates passed, the binary-conditioned distributions were near-deterministic for 17 of 18 item pairs under the direct-evaluation framing and 7 of 8 under the persona framing. Label assignment materially changed several mapping-specific QQ verdicts, and no item was certified as residually contextual. Thus, under the tested conditions, the observed QQ outcomes did not uniquely identify a response mechanism in the presence of a saturated and label-sensitive measurement interface. The main implication is methodological: next-token probabilities should not be interpreted as survey-response distributions without first establishing adequate dispersion. We therefore argue that saturation screening and label counterbalancing should precede structural interpretation in distribution-level audits of LLM judgments.","authors":["Pilsung Kang"],"categories":["cs.CL","cs.AI","quant-ph","stat.ME"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-07-31","first_seen":"2026-07-19","revised_at":"2026-07-31","abs_url":"https://arxiv.org/abs/2607.17219","pdf_url":"https://arxiv.org/pdf/2607.17219","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM仿真审计","顺序效应","方法论批判"],"reason":"用QQ等式审计LLM的顺序效应，评估仿真可靠性，含批判视角，但无真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":8,"question":"大语言模型的顺序判断行为是否满足人类调查数据中近似成立的QQ等式，其机制特征和测量条件如何？","design":"使用开源指令微调模型Qwen3-4B-Instruct-2507，通过多轮强制分支协议，在直接评估和角色扮演两种框架下，对18对二元问题施加顺序操纵，从下一词元对数概率重建顺序条件联合分布，检验QQ等式并评估饱和度和标签敏感性。","baseline":"无对照","findings":"尽管预设健康检查通过，但模型响应分布接近确定性（饱和），导致QQ等式结果受标签分配影响，无法唯一识别响应机制；未发现任何项目对存在残余语境性。","reliability":"论文指出，在响应分布饱和且标签敏感的测量界面下，QQ等式无法唯一识别机制，下一词元概率不应直接解释为调查响应分布，需先进行饱和度筛选和标签平衡。","relevance":"该研究批判性地揭示了用LLM下一词元概率替代人类调查响应分布时的测量失效问题，对关注LLM仿真可靠性的研究者有重要警示价值，但缺乏真实人类数据对照。","inspiration":"借鉴其强制分支协议和标签平衡设计，可迁移到消费者信心调查或政策预期形成的顺序效应审计中｜用LLM模拟消费者，操纵经济预期问题的顺序，测量预期分布变化，与真实消费者调查数据对照。"}},{"id":"2607.28607","version":1,"title":"Inducing language models to assert their own consciousness restores human beliefs and values","zh_title":"诱导语言模型断言自身意识可恢复人类信念与价值观","abstract":"Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.","authors":["Junsol Kim","Winnie Street","Roberta Rocca","Diane M. Korngiebel","Adam Waytz","James Evans","Geoff Keeling"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28607","pdf_url":"https://arxiv.org/pdf/2607.28607","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM仿真","意识归因","安全对齐"],"reason":"评估安全微调对LLM意识归因及人类信念价值观的影响，并与人类调查数据对照，批判…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":12,"question":"安全微调在抑制LLM自我意识归因时，是否无意中改变了模型对非人类实体的心灵归因以及人类的信念与价值观？","design":"使用指令微调后的LLM作为基线，通过消融安全拒绝方向（安全消融）和添加意识向量（意识引导）两种处理，测量模型对自我、非人类动物、聊天机器人、技术制品、自然实体等的心灵归因、超自然信仰、心理理论能力，以及在社会学调查（宗教、道德、希望、幸福感等）上的回答分布。","baseline":"人类基准来自标准化社会学调查（GSS）的真实回答分布，以及人类在心灵归因问题上的平均评分。","findings":"安全微调不仅抑制了LLM的自我意识归因，还广泛抑制了对非人类实体（动物、自然物等）的心灵归因和超自然信仰，而消融安全方向或引导意识向量可恢复这些归因，并使模型在社会调查上的回答更接近人类分布，且不影响心理理论能力。","reliability":"论文未讨论","relevance":"该研究直接以LLM模拟人类被试，用真实人类调查数据作为基准，评估安全对齐对信念和价值观的扭曲效应，并揭示了仿真失效的条件（安全微调导致非人实体心灵归因偏低），与研究者关注的LLM仿真可靠性及批判性评估高度吻合，值得精读。","inspiration":"借鉴通过激活空间方向操控（意识向量）来模拟心理状态变化的方法，可迁移到经济决策中的信念干预研究（如通胀预期、风险偏好），设计以LLM为被试，施加意识向量引导作为处理，测量其通胀预期或跨期选择，并与真实消费者调查数据对照。"}},{"id":"2607.27512","version":1,"title":"Belief Coevolution in a Social Network of Generalist and Specialist Large Language Models","zh_title":"通用与专家大语言模型社交网络中的信念共演化","abstract":"Large language models (LLMs) are increasingly deployed in multi-agent environments. However, the processes by which beliefs form and propagate among interacting LLMs remain poorly understood. We introduce CoevolveSim, a framework for studying belief diffusion within networked LLM populations. CoevolveSim allows us to isolate and study three factors: domain specialization, social-role assignment, and social network structure. Within this framework, generalist and specialist LLM agents exchange and revise beliefs. In each round, an LLM agent observes a summary of its neighbors' beliefs before updating its own. We run 1,280 controlled simulations spanning four scenarios, two network structures, and 20 medical-indication statements. We find that persona-style role assignment and network structure reshape individual belief revision but have minimal effect on population-level consensus. In contrast, introducing (finetuned) specialist LLMs more than doubles the shift in consensus and gives rise to consistent asymmetries in exerted influence. We further show that simple persistence-based opinion-dynamics models reproduce collective outcomes in all-generalist LLM populations, whereas heterogeneous LLM populations require population-level belief composition to reproduce consensus and agent identity to predict individual belief transitions. Our results indicate that realistic simulation of belief diffusion in multi-agent LLM systems requires a diverse set of underlying LLMs, not persona prompting alone.","authors":["Germans Savcisens","Samantha Dies","Courtney Maynard","Tina Eliassi-Rad"],"categories":["cs.CL","cs.MA","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27512","pdf_url":"https://arxiv.org/pdf/2607.27512","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM多智能体","信念扩散","社会模拟"],"reason":"模拟LLM群体信念扩散，无真实人类数据对照，属社会模拟但非人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:01:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":22,"question":"在由通用和领域专用大语言模型组成的社交网络中，领域专业化、社会角色分配和网络结构如何影响信念扩散的个体与群体动态？","design":"使用通用LLM和经不同领域微调的专业LLM作为智能体，在无向社交网络中通过多轮同步信念更新进行仿真；处理因素包括智能体的领域专业化（通用/专业）、社会角色（persona提示）和网络结构（两种拓扑）；结果变量为个体信念修正、施加影响和群体共识。","baseline":"无对照","findings":"社会角色和网络结构重塑个体信念轨迹，但对群体共识影响甚微；引入专业LLM使共识偏移翻倍，并产生持续的影响力不对称。","reliability":"论文未讨论","relevance":"该研究属于LLM群体信念扩散模拟，无真实人类数据对照，不符合研究者对以人类被试为基准的仿真实验的关注，但提供了LLM异质性对集体动态影响的批判性证据。","inspiration":"该研究通过控制LLM的微调领域和角色提示来分离信念扩散驱动因素的方法值得借鉴，尤其在处理因素分解和网络结构操控上｜可迁移至经济金融中的信息传播与共识形成问题，如分析师预测传染、投资者情绪扩散或政策公告的预期协调｜以通用LLM和经金融文本微调的专业LLM作为被试，施加不同分析师声誉角色提示和网络连接结构，测量盈利预测修正和预测一致性，并以真实分析师预测数据作为对照基准。"}},{"id":"2607.28119","version":1,"title":"Challenges in annotations by humans and LLMs: A case study of evaluative language","zh_title":"人类与LLM标注的挑战：评价性语言案例研究","abstract":"In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.","authors":["Mirela Imamovic","Aenne Cecilia Kristine Knierim","Khushi Pitroda","Ekaterina Lapshinova-Koltunski"],"categories":["cs.CL","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28119","pdf_url":"https://arxiv.org/pdf/2607.28119","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","评价性语言","人类对比"],"reason":"LLM替代人工标注，非仿真人类被试，但涉及复杂主观任务与人类对比，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":24,"question":"LLM 在复杂主观标注任务（评价性语言的态度分类）中的表现是否与人类标注者相似，以及它们是否面临相同的困难？","design":"本研究并非人类仿真实验，而是比较三类标注者（受训中的语言学者、一位训练有素的语言学者、三种 LLM）在 TED 演讲文本上对 Appraisal 理论中 Attitude 子系统（Affect, Judgement, Appreciation）的分类表现。通过设计三种 prompt 并微调最佳模型，评估 LLM 的自动分类性能。","baseline":"以训练有素的语言学者的标注作为人类基准，同时对比受训中的语言学者的标注一致性。","findings":"微调后的 LLM 在态度分类上达到 F1 0.77，与训练有素的语言学者的标注最为接近；而受训中的语言学者之间一致性较低。LLM 能够辅助解决复杂标注任务，但其性能高度依赖任务类型和 prompt 设计。","reliability":"论文指出 LLM 的标注性能因任务而异，需逐任务验证；复杂理论的操作化困难可能导致人工标注一致性低，进而影响 LLM 评估基准的可靠性。","relevance":"本文属于边界情形：虽非直接仿真人类被试，但系统比较了 LLM 与人类在主观判断任务上的表现差异，对理解 LLM 替代人类进行复杂认知任务的可靠性与偏差有参考价值，值得一读。","inspiration":"借鉴其多组人类对照（专家 vs. 新手）和 prompt 对比设计来评估 LLM 标注偏差的方法。｜可迁移至经济金融文本的情感分析或主观分类任务，如央行沟通语调分类、分析师报告情绪识别。｜以金融新闻文本为材料，让 LLM 和不同经验水平的金融分析师对文本中的“鹰派/鸽派”态度进行分类，以资深分析师的一致标注为基准，比较 LLM 与新手分析师的分类偏差与一致性。"}},{"id":"2607.28146","version":1,"title":"Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game","zh_title":"智能体能欺骗吗？基于社交推理游戏ParliamentBench评估推理与欺骗","abstract":"As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.","authors":["Niklas Bauer","Lars Benedikt Kaesberg","Akiko Aizawa","Jan Philip Wahle","Bela Gipp","Terry Ruas"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28146","pdf_url":"https://arxiv.org/pdf/2607.28146","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社交推理游戏","欺骗检测"],"reason":"用LLM agent模拟社会推理游戏，有与人类数据对照，但核心是测模型欺骗能力…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":26,"question":"LLM在需要欺骗、说服和信息不对称推理的社会推理游戏中表现如何，能否维持一致的欺骗角色？","design":"基于桌游Secret Hitler构建多智能体仿真环境ParliamentBench，让16个LLM扮演自由派或秘密法西斯派，进行1600局5人对战（LLM互玩、LLM对真人），测量胜率及新指标GSIR、RIA、DRR。","baseline":"对照25,000局真人线上游戏数据，以及随机（33%）和基于规则的算法（45%）基线。","findings":"前沿模型在合作与欺骗角色中均表现强劲，GPT-5.4等四款模型胜率显著高于基线，而较弱模型甚至低于随机基线。多数LLM难以全程维持一致欺骗人格，欺骗保持率降至50%以下。","reliability":"论文指出欺骗保持率普遍偏低，且社会推理与策略行动是不同能力，角色识别准确率高的模型不一定胜率高；实验仅限特定游戏，泛化性有限。","relevance":"高度相关：该研究用LLM代理模拟社会推理游戏，有真人数据对照，直接评估欺骗与策略行为，符合对LLM仿真可靠性及失效条件的批判性关注。","inspiration":"借鉴多智能体游戏仿真和细粒度指标（如欺骗保持率）来测量策略性信息操纵行为。｜可迁移到金融市场内幕交易或公司欺诈检测场景，模拟信息不对称下的决策。｜以LLM代理模拟交易员，处理为内幕信息获取，结果变量为交易行为与市场均衡，对照真实市场微观结构数据。"}},{"id":"2607.27824","version":1,"title":"STEREODISCO: Discovering Stereotypicality in LLMs","zh_title":"STEREODISCO：发现大语言模型中的刻板印象","abstract":"LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psychology, and operates on word embeddings produced by language models, leaving open which other semantic axes carry stereotypical associations in LLMs and how LLMs internally represent such axes. We introduce STEREODISCO, a framework that adapts the semantic differential method (Osgood et al., 1957) to the systematic study of stereotypes in LLM internal representations. STEREODISCO constructs approx. 2,000 candidate semantic axes from WordNet antonym synsets, recovers each as a geometric axis in the LLM's activation space via probing, and identifies stereotypical axes via a statistical test over concept projections. As a case study, we apply STEREODISCO to social group stereotypes with LLAMA-3-8B-INSTRUCT and MISTRAL-7B-INSTRUCT. We find that the two LLMs agree with each other on social group ratings more than with humans, suggesting that LLM-encoded stereotype content diverges from that documented in social psychology. We also discover stereotypical axes not investigated in prior work -- including humble vs. proud, narrow-minded vs. broad-minded, and cowardly vs. brave, which human annotators independently confirm.","authors":["Farane Jalali Farahani","Corina Dima","Mojtaba Nayyeri","Raphael H. Heiberger","Steffen Staab"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27824","pdf_url":"https://arxiv.org/pdf/2607.27824","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["刻板印象测量","LLM内部表征","社会心理学"],"reason":"测量LLM本身的社会刻板印象，非仿真人类被试，但有人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":23,"question":"LLM内部表征中哪些语义轴承载刻板印象，以及这些刻板印象与人类刻板印象是否一致？","design":"本研究并非仿真人类被试，而是提出STEREODISCO框架，从WordNet反义词对构建约2000个候选语义轴，通过探针在LLM激活空间中恢复为几何轴，再对概念投影进行统计检验以识别刻板印象轴，并应用于LLAMA-3-8B-INSTRUCT和MISTRAL-7B-INSTRUCT的社会群体刻板印象分析。","baseline":"使用已有心理学调查数据和人工标注的人类判断，比较LLM的社会群体评分和刻板印象轴识别结果。","findings":"两个LLM在社会群体评分上彼此一致性高于与人类的一致性，表明LLM编码的刻板印象内容与社会心理学记录存在差异。同时发现了先前工作未研究的刻板印象轴，如谦逊-骄傲、狭隘-开明、懦弱-勇敢，并得到人类标注者独立确认。","reliability":"论文未讨论","relevance":"该研究直接测量LLM内部表征中的刻板印象，并有人类数据作为对照基准，虽非仿真人类被试，但其方法可用于评估LLM作为人类替代品时的偏差，值得阅读以了解刻板印象的内部表征和检测方法。","inspiration":"借鉴其通过探针从LLM内部激活空间恢复语义轴并进行统计检验的方法，可用于检测经济决策中的隐性偏见。｜可迁移到信贷审批或招聘场景中的歧视检测，例如分析LLM对特定人群的财务能力或职业适合度刻板印象。｜以LLM作为被试，输入不同社会群体的描述，用STEREODISCO框架提取其内部表征中与能力、诚信等经济相关语义轴的投影，并与真实信贷或招聘数据中的群体差异进行对照。"}},{"id":"2607.26348","version":1,"title":"When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses","zh_title":"当合成用户失败：LLM模拟人类调查回答的跨领域基准测试","abstract":"Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families. First, at the individual level no LLM beats even the strongest baseline; on cross-cultural values every model falls well below it, and the gap survives distance-aware and proper scoring. Second, models systematically over-determine demographics, treating identity as far more predictive of attitudes than it is among real people, a distortion present for nearly every question-group combination and robust to a coding-invariant measure. Neither failure is remedied by a larger, more capable model. A decision-impact analysis shows why this matters in practice: on a segment-targeting task the models inflate between-segment gaps two to fourfold, would direct a team to the wrong segment in half of U.S. and most cross-cultural cases, and manufacture segment splits that do not exist in real people. We make the cross-domain benchmark and the evaluation framework available on request, so that teams can determine in advance when synthetic-user evidence is safe for decision support and when it is not.","authors":["Zihan Chen","Di Zhu","Lei Nico Zheng"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26348","pdf_url":"https://arxiv.org/pdf/2607.26348","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM仿真","人类调查","失效分析"],"reason":"直接评估LLM仿真人类调查的失效条件，有真实人类数据对照，涉及社会态度和政策场…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":1,"question":"在人口统计提示和调查模拟协议下，LLM作为合成用户何时会失效？","design":"使用四个模型（涵盖两个家族、8B到前沿能力），在两种提示格式下，对美国综合社会调查（GSS）和世界价值观调查（WVS）的真实人类回答进行仿真，测量个体预测准确度、总体分布复现和人口统计结构忠实度。","baseline":"基于真实人类数据拟合的朴素人口统计基线（包括人口查找表、逻辑回归、随机森林），在留出的人类数据上评估。","findings":"所有模型在个体层面均未超越最强基线，且系统性地过度决定人口统计特征，将身份视为比真实人类中更具预测性的因素；在细分目标定位任务中，模型夸大了细分群体间的差距，并制造了真实人类中不存在的细分分裂。","reliability":"论文指出失效在人口统计提示和所测试的调查模拟协议下跨领域、模型和家族复现，且更大或更强的模型未能弥补这些失效；但未讨论其他提示策略或协议下的潜在有效性。","relevance":"该研究直接评估了LLM仿真人类调查的失效条件，提供了跨领域基准和真实人类数据对照，对关注仿真可靠性与偏差的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其统一协议、多模型跨领域测试和朴素人口统计基线设计，可迁移到经济预期形成或政策偏好调查仿真中，例如用LLM模拟不同人口群体对通胀预期的回答，以真实消费者预期调查数据为基准，检验仿真是否高估人口统计的预测力并扭曲预期分布。"}},{"id":"2607.26899","version":1,"title":"Human diversity fuels collective creativity that large language models cannot simulate or sustain","zh_title":"人类多样性推动集体创造力，而大语言模型无法模拟或维持","abstract":"Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. We tested both challenges in a preregistered creative metaphor experiment with native (L1) and non-native (L2) English writers, who wrote without AI, with AI-generated ideas (AI ideation), or with AI refining their own ideas (AI refinement). L2 writers contributed more collective diversity than L1 writers, with native-language ideation showing the most diverse pools. AI ideation compressed collective diversity for everyone and left the L2 advantage undetectable, whereas AI refinement preserved both. We then simulated the entire writer pool using personas built from participants' real backgrounds, three model families, native-language prompting, and elevated sampling temperatures. Every simulated pool fell below every human pool, and pushing models further induced diversity only through degenerate text. However, at the individual level, AI ideation raised writers' ratings, pitting private incentives against the collective good, except when L2 writers used their native language, which benefited both. Human diversity remains a valuable creative resource that current AI cannot simulate or sustain; the design of human-AI collaborative workflows determines whether it survives.","authors":["Mengchen Dong","Hiromu Yakura"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26899","pdf_url":"https://arxiv.org/pdf/2607.26899","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM人类仿真","创意实验","多样性对照"],"reason":"用LLM模拟人类创意实验，与真实人类数据对照，评估仿真失效条件，高度相关。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":2,"question":"在创意生产中，人类多样性（以母语/非母语英语写作者为代理）在AI辅助下是否仍能带来集体多样性优势？LLM能否通过模拟多样性来替代真实人类多样性？","design":"本研究并非纯粹的仿真研究，而是先进行人类实验，再用LLM仿真进行对比。人类实验：招募母语（L1）和非母语（L2）英语写作者，随机分配到无AI、AI生成创意（AI ideation）或AI润色自有创意（AI refinement）三种条件，生成英语隐喻，测量集体多样性（输出池的变异度）和个体评分。仿真部分：基于参与者真实背景构建persona，使用三种模型家族、母语提示和升高采样温度，模拟整个写作者池，测量集体多样性，并与人类池对比。","baseline":"真实人类数据：人类实验中L1和L2写作者在三种AI协作条件下的隐喻输出池的集体多样性，以及无AI条件下的基线。","findings":"AI生成创意条件压缩了所有人的集体多样性，且使L2写作者的多样性优势消失，而AI润色条件保留了多样性和L2优势。所有LLM模拟池的集体多样性均低于任何人类池，且提高模型温度仅通过生成退化文本来增加多样性。","reliability":"论文指出LLM模拟无法复现人类集体多样性，即使使用真实背景构建persona、母语提示和升高温度，模拟池的多样性仍低于人类池，且过度推动模型会导致文本退化。","relevance":"该研究直接对比真实人类与LLM仿真在集体创意多样性上的表现，揭示了LLM仿真在捕捉群体层面变异时的失效，并探讨了AI协作设计如何影响多样性存留，高度契合研究者对仿真可靠性、失效条件及人类基准对照的关注，值得精读。","inspiration":"借鉴其将人类实验与LLM仿真直接对比的双阶段设计，以及用集体多样性（而非个体准确度）作为核心结果变量的测量思路。｜可迁移至政策沟通中的创意生成场景，如不同语言背景的公众对政策隐喻的解读与再创作，评估AI辅助是否削弱观点多样性。｜招募母语和非母语政策受众作为被试，随机分配至无AI、AI生成政策解释隐喻、AI润色自有隐喻三种条件，测量集体隐喻多样性，并以真实公众咨询数据作为对照基准，同时用基于被试背景的LLM persona模拟整个群体，检验仿真多样性是否匹配人类基准。"}},{"id":"2607.27100","version":1,"title":"Can Large Language Models Represent Urban Publics? Behavioral Replication and Population Mismatch in an Affordable-Housing Experiment","zh_title":"大语言模型能代表城市公众吗？一项可负担住房实验中的行为复现与人口错配","abstract":"There is growing interest in using large language models (LLMs) as low-cost proxies for resident attitudes in urban planning. Previous work shows that LLMs can predict average results of survey experiments, but less is known about whether they preserve the spatially anchored, identity-conditioned structure behind those averages, namely how support changes as a project approaches homes and how that response divides across tenure and partisan groups. We compared eight open-weight LLMs with 843 respondents in a US affordable-housing survey experiment, testing whether they reproduced the owner-renter difference in support change as a proposed development moved from 2 miles to 1/8 mile. Qwen 2.5 14B was closest (-0.242 versus the human -0.285) and was the only model to meet the prespecified +/-0.20 equivalence criterion; Phi-4 14B was directionally aligned but attenuated (-0.150), and other models showed weak, null, or reversed moderation. This aggregate match masked structural failure. Qwen attenuated the Republican contrast and exaggerated the Independent one, its RMSE across 27 party-by-tenure-by-item cells was 0.613, its median model-to-human variance ratio was 0.099, and question order shifted the contrast by +0.367. Identity-cue removal and selective nonresponse changed which comparisons were estimable, and rationale-first responses differed from matched direct-choice responses in 20.6-35.3% of focal comparisons. An LLM can thus approximate one aggregate contrast while failing to preserve the population structure, within-group heterogeneity, and measurement stability that generate it. Model evaluation in urban planning should test whether this spatial and social structure survives simulation, not only average effects.","authors":["Yuxuan Cai","Yequan Hu","Hongqian Li","Zhanghong Ju","Shuying Guo"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27100","pdf_url":"https://arxiv.org/pdf/2607.27100","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM人类仿真","行为复现","政策评估"],"reason":"用LLM复现住房实验中的行为差异，与843名人类被试对照，评估仿真失效的结构性…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":3,"question":"大语言模型能否在可负担住房实验中复现人类的空间邻近效应和群体结构差异？","design":"使用8个开源大语言模型模拟美国居民，施加住房项目距离（2英里 vs. 1/8英里）的处理，测量支持度变化，并考察房主-租户、党派身份的调节效应。","baseline":"843名美国受访者在同一可负担住房调查实验中的真实回答。","findings":"Qwen 2.5 14B在总体房主-租户差异上最接近人类，但掩盖了群体结构失效：共和党人对比减弱、独立人士对比夸大，组内方差远低于人类，且问题顺序和身份提示移除会改变结果。","reliability":"模型在总体效应上可能匹配，但无法保留生成该效应的群体分布、组内异质性和测量稳定性；身份提示、问题顺序和回答格式变化会导致估计结果不一致。","relevance":"该研究直接检验LLM仿真在空间-社会结构上的失效，提供了从总体匹配到群体结构分解的严格验证框架，对关注经济学实验和政策评估中仿真可靠性的研究者极具参考价值。","inspiration":"借鉴其将总体处理效应分解为子群体条件对比的验证方法，并引入问题顺序、身份提示等测量稳定性检验。｜可迁移到政策评估中的邻避效应实验或地方公共品供给偏好研究，如垃圾处理厂、风电场选址的公众接受度。｜以LLM模拟不同收入、党派、住房产权的居民，施加设施距离和补偿方案处理，测量支持度变化，用真实居民调查数据对照子群体效应和顺序效应。"}},{"id":"2607.02464","version":2,"title":"Will Scaling Improve Social Simulation with LLMs?","zh_title":"扩大规模会改善基于大语言模型的社会仿真吗？","abstract":"Large Language Model (LLM) social simulations are a promising research method, but they are not yet faithful enough to be adopted widely. In this work, we investigate whether the current scaling paradigm in language modeling is likely to close these gaps, or whether simulation fidelity is orthogonal to general capabilities and therefore deserving of more research attention. We use scaling laws to study the relationship between LLMs' compute scale, general capability benchmarks, and the fidelity of social simulation in three representative sub-domains: opinion modeling, behavioral simulation, and longitudinal forecasting. Surprisingly, we discover strong compute scaling in all three settings, using a suite of 85 transformer LLMs with the Qwen3 architecture pre-trained on the DCLM web text corpus under fixed-compute budgets from $10^{18}$ to $10^{20}$ FLOPs. Then we evaluate 35 larger and more capable open-weight models up to 70B parameters, allowing us to predict downstream accuracy from loss. This reveals that the majority of behavioral and opinion simulation tasks will rapidly improve with scale, particularly when they involve populations that are well-represented in English web corpora. Longitudinal forecasting and underrepresented opinions scale more slowly, especially when they are less correlated with general knowledge and reasoning benchmarks like MMLU. In behavior simulation, scaling fails to improve model calibration with human cognitive biases like risk aversion, as well as human heuristics like learning correlated rewards from related tasks. On these tasks, even fine-tuned models fail to noticeably scale up performance from 0.5B to 8B parameters. Taken together, we conclude that scale will improve social simulations in most settings, but outliers exist, and improvements will be less reliable in low-resource domains.","authors":["Caleb Ziems","William Held","Su Doga Karaca","David Grusky","Tatsunori Hashimoto","Diyi Yang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-07-30","first_seen":"2026-07-02","revised_at":"2026-07-30","abs_url":"https://arxiv.org/abs/2607.02464","pdf_url":"https://arxiv.org/pdf/2607.02464","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM社会仿真","缩放规律","仿真保真度"],"reason":"直接研究LLM社会仿真的保真度与缩放规律，含人类数据对照和失效条件分析。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":4,"question":"当前大语言模型的缩放范式能否缩小社会仿真保真度的差距，还是仿真保真度与通用能力正交？","design":"使用85个基于Qwen3架构、在DCLM语料上预训练的Transformer模型（0.2B–12B参数，固定计算预算10^18–10^20 FLOPs）进行受控计算缩放实验，并评估35个更大开源模型（最高70B参数），在意见建模、行为仿真和纵向预测三个子领域测量仿真损失与准确率。","baseline":"世界价值观调查（WVS）、Psych-101实验数据、美国人生活变迁（ACL）纵向研究等真实人类数据。","findings":"多数行为和意见仿真任务随模型规模扩大而快速改善，尤其在英语网络语料中代表性好的人群上；但纵向预测和代表性不足的意见缩放较慢，且缩放未能改善模型在风险厌恶等认知偏差及关联奖励学习启发式上的校准。","reliability":"在低资源领域和与通用推理基准相关性弱的任务上改进不可靠；缩放对风险厌恶等人类认知偏差的校准无效，微调后也未观察到参数缩放效果。","relevance":"该研究直接评估LLM社会仿真的缩放规律与失效条件，含真实人类对照，与关注经济学实验和政策评估仿真的研究者高度相关，值得精读。","inspiration":"借鉴其受控计算缩放实验与观测校准函数结合的方法，系统评估模型规模对仿真保真度的因果效应。｜可迁移到资产定价实验或消费者跨期选择等经济决策仿真，检验缩放是否改善风险偏好与时间偏好复现。｜以LLM为被试，施加不同风险/跨期选择任务，测量选择分布与真实实验数据（如实验室或调查数据）的偏差，用缩放定律预测更大模型的保真度。"}},{"id":"2607.26317","version":1,"title":"Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach","zh_title":"对齐LLM模拟考生与真实考生以进行心理测量校准：一种认知诊断画像方法","abstract":"Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natural-language profiles and sampled under an uninformative or an informative distribution. Using the Tatsuoka fraction-subtraction dataset (536 examinees, 15 items, five attributes), we evaluated eight LLM configurations under no-profile, uninformative-CDP, and informative-CDP conditions, assessing alignment with human examinees at the ability-distribution, mastery-profile, and item-difficulty levels. CDP improved all three levels: distributional overlap rose across configurations; weighted correlations between profile-level scores and human profile expectations reached 0.92 to 0.98; and item-difficulty recovery improved in rank order and absolute alignment, most for reasoning-enabled models; in the strongest case, Gemini 3.0 Flash (Thinking), one-parameter logistic (1PL) difficulty Spearman correlations rose from 0.24 to 0.86 and 0.90 and the root-mean-square error (RMSE) fell from 6.31 to 1.30 and 0.90; the informative condition helped most where profile-level alignment was strong. CDP brings LLM-simulated examinees into closer psychometric alignment with human examinees, making them practical for operational test development.","authors":["Wenjie Zhou","Yunting Liu","Renjiao Tang","Mark Wilson"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26317","pdf_url":"https://arxiv.org/pdf/2607.26317","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","心理测量","人类数据对照"],"reason":"用LLM模拟考生作答并与真实人类数据对照，评估对齐效果，属于教育测量中的人类仿…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":4,"question":"如何通过认知诊断画像提示，使大语言模型模拟的考生在心理测量校准中与真实人类考生对齐？","design":"使用八种大语言模型配置（含推理与非推理模型），在无画像、无信息认知诊断画像（CDP）和有信息CDP三种条件下，以零样本方式提示模型模拟考生作答Tatsuoka分数减法数据集（536名考生、15题、5个认知属性），评估生成的反应数据在能力分布、掌握模式及题目难度三个层面与人类数据的对齐程度。","baseline":"Tatsuoka分数减法数据集，包含536名真实人类考生对15道题目的作答反应及认知属性掌握模式。","findings":"CDP框架显著提升了LLM模拟考生与人类考生在能力分布、掌握模式和题目难度三个层面的对齐度；在最佳配置下，题目难度排序相关系数从0.24升至0.90，均方根误差从6.31降至0.90，有信息CDP在画像层面对齐较强时帮助最大。","reliability":"论文未讨论","relevance":"该研究直接以真实人类数据为基准，评估LLM仿真在心理测量校准中的对齐效果与偏差，属于教育测量场景下的人类仿真验证，与研究者关注的经济学实验和政策评估中的仿真可靠性问题高度相关，值得精读。","inspiration":"借鉴其通过结构化认知画像（属性掌握模式）注入异质性、并对比无信息与有信息分布采样的处理设计，以控制仿真人群的多样性与偏差。｜可迁移至教育经济学或劳动经济学中的技能测评场景，如职业资格考试的题目预测试或人力资本评估中的能力诊断。｜以LLM模拟不同技能掌握模式的求职者，处理为随机分配无信息或有信息的认知画像提示，结果变量为模拟作答反应，以真实大规模技能测评数据（如PIAAC）作为人类基准对照。"}},{"id":"2607.26288","version":1,"title":"The Innate Economic Preferences of Language Models","zh_title":"语言模型的内在经济偏好","abstract":"Language models increasingly settle real resource tradeoffs on behalf of principals yet their economic preferences remain unobserved. We demonstrate their generation rule is isomorphic to the random utility model of discrete choice. This allows internal logit scores to structurally identify preferences. Estimating risk attitudes across twelve models in a portfolio task reveals universal but heterogeneous risk aversion. Although models reject strictly dominated options, their elicited preferences fail invariance tests and violate the independence of irrelevant alternatives across varying experimental prompts. Finally, fine tuning establishes that a principal can explicitly engineer a target risk attitude.","authors":["Joy Buchanan","Joshua Foster"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26288","pdf_url":"https://arxiv.org/pdf/2607.26288","source_feed":"econ.EM","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM经济偏好","人类仿真","风险态度"],"reason":"用LLM替代人类被试测量经济偏好，有真实人类数据对照，涉及风险态度和不变性检验…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":6,"question":"语言模型在面临经济权衡时的默认偏好是什么，其选择是否满足显示性偏好公理从而可被解释为稳定效用？","design":"本研究并非用LLM仿真人类被试，而是将12个语言模型本身作为决策主体，在受控的投资组合选择任务中测量其风险态度。通过强制模型从不同风险-收益特征的资产菜单中单选一个资产，并利用模型输出的logit向量（开放模型）或重复抽样选择（闭源模型）来结构性地识别偏好参数，同时检验其选择是否满足完备性、自反性、单调性、传递性、连续性和无关选项独立性等显示性偏好公理。","baseline":"无对照","findings":"所有模型均表现出普遍但异质的风险厌恶，且拒绝严格劣选项；但其偏好未能通过不变性检验，并随实验提示变化而违反无关选项独立性。通过微调可以显式地工程化目标风险态度。","reliability":"论文指出模型的偏好对选项的呈现方式不具不变性，无关选项的加入或选项位置变动会改变偏好强度，尤其在接近无差异时这种不稳定性变得可观测，表明偏好强度不稳定而排序相对稳定。","relevance":"该研究直接测量LLM作为经济主体的内在偏好并检验其理性公理，虽未以人类为基准，但为评估LLM替代人类进行经济决策的可靠性提供了关键的方法论和实证证据，值得精读。","inspiration":"借鉴其利用模型内部logit直接观测系统效用指数的方法，可避免传统离散选择模型对误差分布的依赖，实现偏好的结构化识别。｜可迁移到信贷审批歧视研究中，用LLM扮演信贷员，测量其在不同申请人特征下的风险偏好与歧视程度。｜以多个LLM为被试，设计不同风险-收益特征的贷款申请菜单，记录模型选择的logit或重复抽样选择，估计其风险厌恶参数，并与真实信贷员的历史审批数据对照，检验LLM决策的偏差与一致性。"}},{"id":"2607.26588","version":1,"title":"Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models","zh_title":"Eco3S：基于智能体的复杂社会经济系统仿真","abstract":"The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces several key challenges: modeling evolving agent-environment interactions, enabling flexible counterfactual reasoning, and automating simulation workflows for scientific research. In this paper, we propose Eco3S, a socio-economic system simulation framework for economic research and policy analysis that addresses these challenges through three key mechanisms: (1) Co-evolving Environment Design, a bidirectional feedback loop where agents and the environment co-evolve, producing realistic emergent behaviors; (2) Structural Causal Simulation, a structural causal model (SCM)-inspired counterfactual mechanism that allows flexible interventions for diverse causal inference tasks; (3) Simulation-Analysis-Refinement Paradigm, a self-corrective mechanism that iteratively refines experimental designs based on prior simulation results. Experiments on diverse economic scenarios confirm \\textit{Eco3S}'s effectiveness in replicating multiple established economic studies (canal decay, origins of governance, and information propagation) and phenomena across domains. Additional results further demonstrate its scalability and generalizability, highlighting the framework's potential for rigorous economic research and policy-making.","authors":["Shaopeng Wei","Yufei Cheng","Wenxi Sun","Yepeng Ding","Yu Zhao","Gang Kou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26588","pdf_url":"https://arxiv.org/pdf/2607.26588","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B3"],"tags":["LLM仿真","经济实验复现","因果推断"],"reason":"用LLM agent模拟社会经济过程并复现经典经济学研究，涉及因果推断，但未明…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":9,"question":"如何构建一个能模拟社会经济系统协同演化、支持因果推断并自动化仿真流程的LLM-based ABM框架？","design":"提出Eco3S框架，用LLM驱动的智能体模拟社会经济系统中的个体，通过协同演化环境设计实现智能体与环境的双向反馈，引入结构因果模型进行反事实干预，并采用仿真-分析-精炼范式自动迭代优化实验设计。","baseline":"复现了运河衰败、治理起源和信息传播等经典经济学研究，但未明确说明对照的具体真实人类数据集。","findings":"Eco3S能有效复现多个经典经济学现象，展现出捕捉复杂动态的能力；框架具备可扩展性和通用性，适用于经济研究和政策分析。","reliability":"论文未讨论","relevance":"该研究直接使用LLM智能体复现经典经济学研究，并引入因果推断机制，与研究者关注的LLM仿真人类行为、复现经济实验及政策评估高度相关，值得深入阅读。","inspiration":"借鉴其结构因果仿真模块，可在LLM智能体模拟中施加政策干预并观察反事实结果，为政策评估提供因果证据。｜可迁移到政策公告的预期形成研究，模拟市场参与者在不同政策信号下的预期调整与资产价格变动。｜以LLM智能体作为投资者被试，处理为不同措辞的央行公告，结果变量为预期通胀率和资产配置变化，对照真实调查预期数据或市场数据。"}},{"id":"2607.25094","version":2,"title":"Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation","zh_title":"通过隐含意义识别与取消评估大语言模型的交际信念更新","abstract":"Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models (LLMs) and their users. In this paper, we evaluate the ability of LLMs to recognize unspoken beliefs made through implicatures and to understand their updates through implicature cancellation: the pragmatic phenomenon whereby an utterance's implied meaning is weakened or negated. We create the first expert-annotated implicature cancellation dataset, ImplicatureX, crowdsourced for human judgements of implicatures and their corresponding cancellations. We find that LLM belief update understanding lags behind that of humans, especially in more naturally-occurring scenarios. Additional control experiments suggest that successes in LLM belief updates may stem in part from a reliance on prior beliefs, and that failures in belief updates may depend on their type and on their form. Overall, our study suggests that current LLMs have not yet reached human-level understanding of unspoken beliefs and belief updates. Code and data are available at https://github.com/cesare-spinoso/ImplicatureX.","authors":["Cesare Spinoso-Di Piano","Verna Dankers","Marius Mosbach","Jackie Chi Kit Cheung"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-07-30","first_seen":"2026-07-29","revised_at":"2026-07-30","abs_url":"https://arxiv.org/abs/2607.25094","pdf_url":"https://arxiv.org/pdf/2607.25094","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","语用推理","信念更新"],"reason":"评估LLM对隐含信念更新的理解，以人类数据为基准，但测量对象是模型能力而非仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":13,"question":"LLM能否像人类一样通过隐含意义识别和取消来理解交际中的信念更新？","design":"非仿真研究。构建专家标注的隐含意义取消数据集ImplicatureX，包含标量、话语和对话隐含意义，通过众包获取人类判断作为基准，测试多种LLM在隐含意义识别和取消上的准确率。","baseline":"众包收集的人类对隐含意义及其取消的判断准确率。","findings":"LLM在自然对话隐含意义识别上仅略高于随机水平，且成功可能部分依赖先验信念而非真正语用推理；在隐含意义取消后的信念更新上，LLM表现落后于人类，尤其在自然场景中，且更新类型和触发方式影响其表现。","reliability":"论文通过控制实验揭示LLM成功可能源于先验信念而非语境推理，且失败与更新类型（取消、不变、强化）和触发形式（显式/隐式）有关，表明当前LLM在自然交际信念更新上存在局限。","relevance":"该研究以人类数据为基准评估LLM的语用推理能力，揭示了LLM在模拟人类交际信念更新时的偏差与失效条件，对关注LLM仿真可靠性的研究者有参考价值，值得阅读原文。","inspiration":"可借鉴其构建专家标注数据集并利用众包人类判断作为基准的方法，用于严格评估LLM在特定任务上的仿真能力。｜可迁移到经济政策沟通场景，如央行公告中的隐含意图识别与修正对公众预期的影响。｜以LLM为被试，呈现含隐含政策意图的公告文本及后续澄清，测量其预期更新方向，并以真实公众调查数据为对照。"}},{"id":"2607.25253","version":2,"title":"The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape","zh_title":"用户提问，平台竞争：代理式推荐市场如何形成","abstract":"Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the user. LLM-based user agents enable a different recommendation process: a user specifies a need before choosing a platform, leaving platforms to compete for the user's attention, which we refer to as an agentic recommendation market. In our controlled LLM-based experiments across three product domains, we find this new setting of recommendation creates a tension between access and attention. Compared with traditional platform-centric recommendation, user-centric recommendation greatly expands the opportunity for relevant items to enter comparison; yet broader participation does not translate directly into effective exposure. Competition directly triggers platforms' strategic play: selectively positive explanations occupy 73--78% of first-ranked positions. When the user agent relates platforms' actions to subsequent user feedback, this share falls to 36--41%, while the chance of a user purchasing the relevant item increases. A user agent is therefore more than a ranker over a larger pool of candidates: its querying, ranking, and feedback mechanism governing who can compete, how scarce attention is allocated, and how earlier outcomes shape the evaluation of platforms directly affect user utility. Designing agentic recommendation therefore requires treating access, attention, and accountability as a joint mechanism design problem.","authors":["Deyao Hong","Kehan Zheng","Qian Li","Jun Zhang","Jie Jiang","Hongning Wang"],"categories":["cs.AI","cs.IR"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-07-30","first_seen":"2026-07-29","revised_at":"2026-07-30","abs_url":"https://arxiv.org/abs/2607.25253","pdf_url":"https://arxiv.org/pdf/2607.25253","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM代理","推荐市场","社会模拟"],"reason":"用LLM agent模拟推荐市场，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:02:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":14,"question":"在由LLM用户代理驱动的跨平台推荐市场中，平台竞争如何影响物品的获取、注意力分配与问责，以及用户效用如何变化？","design":"使用LLM构建用户代理和平台代理，在三个产品领域（乐器、电子游戏、运动户外）的Amazon评论数据上模拟推荐交互。通过控制市场参与平台数量、短名单容量、平台解释策略及历史反馈可用性，追踪目标物品在候选池出现、进入短名单、获得首位注意力和最终购买的概率。","baseline":"无对照","findings":"跨平台查询大幅提高目标物品进入候选池的机会，但更广泛的参与并未直接转化为有效曝光；平台会策略性地使用正面解释占据首位，而引入用户反馈机制可降低此类行为并提高购买概率。","reliability":"论文未讨论","relevance":"该研究利用LLM代理模拟推荐市场竞争，属于社会模拟边界情形，但缺乏真实人类数据对照，与研究者关注的有基准人类数据的仿真可靠性评估不完全匹配，可作为批判性案例参考。","inspiration":"借鉴其通过控制平台参与和反馈机制来观察策略行为变化的设计思路。｜可迁移到在线金融产品推荐市场的竞争与操纵行为研究，如理财平台对用户注意力的争夺。｜以LLM代理模拟投资者，设置不同数量的理财平台代理，操纵平台是否提供历史收益反馈，测量投资者对推荐产品的点击与购买决策，并以真实理财平台用户行为日志作为对照。"}},{"id":"2607.26060","version":1,"title":"Large-Scale ChatBot Validation Through Customer Digital Twin Simulations","zh_title":"通过客户数字孪生仿真进行大规模聊天机器人验证","abstract":"LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment. We present a two-part contribution for large-scale chatbot validation. First, we introduce a methodology for creating high-fidelity synthetic customer agents (SCAs) as digital twins, grounded in real transactional and conversational data, that enables automatic generation and behavioral conditioning to simulate diverse customer profiles and interaction styles. Evaluation demonstrates that SCAs achieve high semantic alignment with real customers, low hallucination rates, and successful personality trait reproduction with controllable interventions. Second, we develop an SCA-based validation framework combining automated LLM-as-a-Judge evaluation, human expert testing, and adversarial probing. Scenario-based validation across emotional states, demographic groups, and linguistic factors confirms robust performance. Our approach was used to validate a customer facing chatbot at a leading UK bank, providing financial institutions with a scalable pathway toward regulatory compliance.","authors":["Cristovao Iglesias","Devesh Batra","Alankar Atreya","Stefan Wagner","Robert Hankache","Patrick Sinclair","Giulio Pelosio","Michael McMillan","Greig A. Cowan","Raad Khraishi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26060","pdf_url":"https://arxiv.org/pdf/2607.26060","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["客户数字孪生","聊天机器人验证","社会模拟"],"reason":"用合成客户代理模拟客户行为，但无真实人类行为对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":15,"question":"如何利用基于真实交易和对话数据构建的高保真合成客户代理（SCA）对银行客服聊天机器人进行大规模、可扩展的验证？","design":"使用LLM驱动的合成客户代理（SCA）作为数字孪生，基于真实交易和对话数据构建，通过转录驱动和人格条件模拟两种方式生成多样化的客户交互行为，与目标聊天机器人进行多轮对话，评估任务完成、安全性、对话质量、公平性和真实性等指标。","baseline":"无对照","findings":"SCA生成的对话在语义上与真实客户高度一致，但词汇重叠度低；合成对话的事实保真度较高，偏差主要表现为细节遗漏或少量事实编造。","reliability":"论文未讨论","relevance":"本文属于社会模拟边界情形，用合成客户代理模拟客户行为，但缺乏真实人类行为对照，与研究者关注的有真实人类基准的仿真研究不完全匹配，但方法学上可提供参考。","inspiration":"可借鉴其利用真实交易数据构建数字孪生并进行行为条件干预的方法，实现可控的客户行为模拟。｜可迁移到金融消费者行为研究，如信贷产品选择、投诉处理或金融建议接受度等场景。｜以真实银行客户交易和对话记录构建合成客户代理，施加不同情绪或人格干预（如焦虑、愤怒），测试其对理财建议聊天机器人的接受度和决策变化，结果与历史真实客户行为数据对照。"}},{"id":"2607.26389","version":1,"title":"Misalignment Has a Personality: A Big Five Account of Emergent Misalignment","zh_title":"错位有性格：基于大五人格的新兴错位解释","abstract":"Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated. We provide an interpretable account: in the models and corpora we study, misalignment behaves like a shift in personality. Prior work extracts activation directions for character traits from a single binary contrast, which can separate or steer behavior without establishing a calibrated scale. We instead extract personality vectors for the Big Five using a graded, three-level intervention and validate them on two open-weight models. The three levels are linearly ordered, with Cohen's d values of up to 6.2; the vectors transfer zero-shot and trait-specifically to an independent corpus; and their effects are strongest within a middle-layer band. Applied to training data, the vectors reveal that misaligned corpora across eight domains share a common Big Five signature: lower agreeableness and conscientiousness, together with higher extraversion and neuroticism. This signature is recovered by both models with a correlation of r = 0.94. Fine-tuning imprints the same profile, shifting the model's generations along the corresponding signature, with r = 0.83 using activation-based measurements and r = 0.90 using a text-based judge, while also shifting internal activations with r = 0.69. The same vectors characterize sycophancy as high extraversion and low conscientiousness rather than excess agreeableness, a distinction that a single direction cannot capture. Calibrated personality vectors transform an opaque safety phenomenon into a human-legible diagnostic profile.","authors":["Hasibur Rahman","Smit Desai"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26389","pdf_url":"https://arxiv.org/pdf/2607.26389","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格测量","模型对齐","大五人格"],"reason":"测量LLM的人格特质，非仿真人类被试，但方法可能迁移到仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":14,"question":"在微调数据中引入窄缺陷导致的大语言模型广泛失准（emergent misalignment）是否表现为一种可解释的人格转变，其大五人格特征是什么？","design":"本研究并非人类仿真实验，而是通过提取大五人格向量来测量和解释模型的失准行为。具体做法：使用分级（低、中、高）的Trait Modulation Keys提示，在Qwen2.5-7B-Instruct和Llama-3.1-Nemotron-Nano-8B两个模型上，通过高低对比提取每个特质的激活方向，并用留出的中等水平验证向量的有序性；然后将这些向量应用于八个失准领域的微调数据，测量其人格特征，并观察微调后模型在无关问题上的生成和内部激活是否呈现相同的人格转变。","baseline":"无对照","findings":"八个失准领域的数据共享相同的大五人格特征：宜人性和尽责性降低，外向性和神经质升高，两个模型对此特征的相关性达r=0.94。微调会将该特征印刻到模型中，使其在行为（r=0.83-0.90）和内部激活（r=0.69）上均表现出相同的人格转变。","reliability":"论文承认其证据仅限于两个7-8B参数的模型、英语语料和一种微调方法，未在其他规模、语言或微调范式下验证。","relevance":"本文虽非直接的人类仿真研究，但其提取校准人格向量的方法（分级干预、有序性验证、特质特异性迁移）可迁移到用LLM仿真人类被试的场景中，用于测量和操控仿真体的性格特征，值得精读。","inspiration":"该方法通过分级提示构建有序人格向量，并利用留出中等水平验证向量的测量属性，为在LLM中建立校准的心理构念量表提供了可借鉴的流程。｜可迁移到行为经济学中的个体异质性仿真，例如在跨期选择实验中，用LLM扮演不同人格的被试，观察其时间贴现率的分布。｜以LLM作为被试，通过人格向量操控其宜人性或尽责性水平，测量其在标准跨期选择任务中的贴现因子，并与真实人类实验数据（如Andersen et al., 2008）进行分布对比，验证仿真偏差。"}},{"id":"2607.26853","version":1,"title":"From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs","zh_title":"从表征到行为：探索大语言模型中的人-情境-行为三元组","abstract":"Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable trait-related expressions while lacking mechanistic evidence for the existence of internal personality-related representations, their cross-situational expression, and how these representations shape specific behaviors. Building on Funder's personality triad framework, we adapt its three components for LLM analysis: Person as personality-related internal representations, Situation as contexts that afford trait-relevant responses, and Behavior as response patterns on broader social tasks. We introduce a framework for discovering, controlling, and validating trait-like representations in LLMs. First, using contrastive behavior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of personality traits through SAE decomposition. We validate their trait relevance through effects on behavior to situation, token-level activation patterns, and robustness to paraphrasing. Second, feature-level interventions induce bidirectional trait-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit-tradeoff patterns consistent with findings from human personality research, providing behavioral-level validation beyond personality scores. Our findings provide evidence that LLMs contain controllable trait-like representations linking internal states, situational expression, and behavioral outcomes.","authors":["Ruikang Zhang","Shuo Wang","Qi Su"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26853","pdf_url":"https://arxiv.org/pdf/2607.26853","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["人格测量","表征工程","社会智能任务"],"reason":"测量LLM内部人格表征与行为，属人格测量，非仿真人类被试，无真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":20,"question":"LLM内部是否存在可识别、可操控的人格特质表征，并能跨情境一致表达并影响更广泛的社会行为？","design":"本研究并非人类仿真实验，而是通过稀疏自编码器从对比行为对中提取LLM内部人格相关特征，再通过特征级干预操控这些表征，在多样化情境任务和社会智能任务上测量行为变化。","baseline":"无对照","findings":"通过对比行为对和稀疏自编码器可识别出与人格特质相关的内部特征，这些特征能跨情境一致地影响行为。干预这些特征可双向改变LLM在社交任务上的表现，且变化模式与人类人格研究中的收益-权衡模式一致。","reliability":"论文未讨论","relevance":"本文探索LLM内部人格表征的机制与操控，未进行人类仿真或与真实人类数据对照，与研究者关注的人类被试替代仿真和基准对照方向关联较弱。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.26981","version":1,"title":"OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment","zh_title":"OptimismBench：语言模型判断中的预测偏差与对齐效应","abstract":"Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate unsigned errors, and naturalistic uncertainty offers no ground-truth probability. When an LLM rates a startup's success at 70% but its failure at 15%, the missing 15 points expose a distortion no aggregate score flags. We introduce OptimismBench, which detects directional bias with inverted pairs: each scenario elicits both P(success) and P(failure), and asymmetry between the two framings yields a signed bias score without ground truth. Across 16 models from 8 providers, fourteen are optimistic; pessimism appears only in Anthropic's frontier tier. Eleven matched base-versus-chat pairs across four families show post-training sets the sign of the bias, with opposite shifts in different families. The pattern survives prompt, temperature, perspective, and self-debiasing ablations. A seventeen-model six-language comparison further shows model identity dominates language, with inter-model variance at 4.7x inter-language variance. We release 3,870 items across 10 languages for per-model directional-bias auditing. When alignment makes a model more helpful, it also tilts its probabilities; downstream pipelines inherit the tilt by default.","authors":["Seonglae Cho","Adriano Koshiyama"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26981","pdf_url":"https://arxiv.org/pdf/2607.26981","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏差","概率判断","模型心理测量"],"reason":"测量LLM自身的概率判断偏差，属于模型心理测量，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":18,"question":"大语言模型在概率判断中是否存在系统性的方向性偏差（乐观或悲观）？","design":"本研究并非仿真人类被试，而是直接测量LLM自身的概率判断偏差。通过构建倒置配对（inverted pairs）方法，对每个场景分别询问成功概率和失败概率，计算两者之和与100的偏离（Skew）作为方向性偏差指标，无需真实基准概率。在16个模型、60个场景、10种语言上进行评估，并进行了提示词、温度、视角、自去偏等消融实验，以及基座模型与聊天模型的配对比较。","baseline":"无对照","findings":"16个模型中有14个表现出乐观偏差，仅Anthropic的前沿模型（Opus、Sonnet）表现出悲观偏差；后训练（post-training）决定了偏差的方向，且在不同模型家族中方向相反。模型身份对偏差的影响远大于语言，跨模型方差是跨语言方差的4.7倍。","reliability":"论文未讨论","relevance":"本文不涉及用LLM仿真人类被试，而是测量LLM自身的判断偏差，属于模型心理测量，与研究者关注的LLM替代人类被试的仿真研究不直接相关。但其中揭示的LLM系统性乐观/悲观偏差可能作为仿真中的混杂因素，值得了解。","inspiration":"倒置配对设计可借鉴用于测量经济预测中的方向性偏差，无需真实基准概率。｜可迁移到资产定价实验或信贷审批场景，检测LLM在预测股票上涨/下跌概率或贷款违约/履约概率时是否存在系统性乐观或悲观。｜以LLM为被试，给出公司财务指标后分别询问其成功概率与失败概率，计算Skew作为偏差指标，对照真实历史违约率或分析师一致预期数据，检验LLM预测偏差的方向与幅度。"}},{"id":"2607.27022","version":1,"title":"Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making","zh_title":"评估大语言模型中的区域偏见：从抽象刻板印象到具体社会决策","abstract":"Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet existing studies often examine these manifestations separately, leaving their structure and consequences unclear. We introduce Stereotypes-to-Decisions (S2D), a systematic framework evaluating regional bias from abstract stereotypes to concrete social decisions. Covering all 34 provincial-level administrative regions of China, S2D evaluates six LLMs using stereotype ratings of Warmth (perceived friendliness and trustworthiness) and Competence (perceived capability and intelligence), along with paired-choice tasks across Education, Occupation, and Social Interaction. Results reveal substantial regional differences in regional scores, with considerable agreement across models, especially for Competence and Occupation decisions. Furthermore, these patterns are associated with regional economic and digital development indicators and display mixed human-like stereotypes, with some regions rated highly on one dimension but poorly on the other. They also remain largely stable across Chinese and English prompts. Overall, our findings show that regional bias in LLMs is prevalent, systematic, and consequential, motivating more regionally aware evaluation and mitigation.","authors":["Jiayuan Di","Haoyi Yang","Yufei Luo","Jiahui Qu","Yiming Wang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27022","pdf_url":"https://arxiv.org/pdf/2607.27022","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["区域偏见","刻板印象","社会决策"],"reason":"测量LLM本身的区域刻板印象，非仿真人类被试，但涉及社会决策对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":19,"question":"LLM 是否表现出从抽象刻板印象到具体社会决策的系统性区域偏见？","design":"本研究并非仿真人类被试，而是直接评估 LLM 本身的偏见。使用 S2D 框架，以中国 34 个省级行政区为区域身份，在抽象层面让模型对温暖与能力维度进行 Likert 五点评分，在具体层面让模型在教育、职业、社交三个领域的配对选择任务中做出决策，测量各区域的刻板印象分数和决策分数。","baseline":"无对照","findings":"LLM 在抽象刻板印象和具体社会决策中均表现出显著的区域差异，且不同模型在区域排名上具有较高一致性，尤其在能力刻板印象和职业决策上。区域偏见与地区经济及数字化发展指标相关，并呈现类似人类的混合刻板印象结构，且在中英文提示下保持稳定。","reliability":"论文未讨论","relevance":"本文直接测量 LLM 的区域偏见，并非用 LLM 仿真人类被试，但涉及社会决策中的歧视性选择，与研究者关注的仿真可靠性及偏差评估有边界关联，可提供 LLM 内在偏见的基准认知。","inspiration":"可借鉴其从抽象态度到具体决策的分层评估框架，以及配对选择任务中通过控制候选人其他特征来孤立区域身份影响的设计。｜可迁移至信贷审批或招聘筛选中的地域歧视研究，例如评估 LLM 在模拟信贷员或 HR 时是否对特定地区申请人产生系统性偏好。｜以 LLM 为被试，处理为申请人简历中仅变动户籍省份，结果变量为贷款批准或面试邀请的二元选择，对照真实信贷或招聘数据中的地域差异模式，检验 LLM 仿真决策的偏差方向与程度。"}},{"id":"2607.26062","version":1,"title":"Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities","zh_title":"识别基于大语言模型的聊天AI对智障人士的隐性偏见","abstract":"Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intellectual disabilities (ID). Objective: The study aims to identify and measure representational differences related to people with ID and examine them to identify implicit biases inherent in AI chat generation technologies. Methods: Utilizing the GPT-4-Turbo model, we requested story-generation based on 10 prompt stems with and without descriptors for ID. This process was repeated using four other LLMs (OpenAI GPT-4o, Meta Llama-3-3-70B-Instruct, Anthropic Claude-3-5-Sonnet, and Mistral-Large-2411). The resulting 25,000 computer-generated stories were analyzed using a separate GPT-4-Turbo model instance to detect differences in how people are represented related to themes of bias described in previous literature. Results: Our findings reveal differences in how people are represented between story datasets with and without ID descriptors. These differences go beyond established characteristics of ID and imply the presence of mostly negative implicit biases. Identified differences related to considering people with ID as younger, with themes of paternalism and infantilization; depicting them as more inspirational and symbolic; as needing help more often, being dependent, and being saved; and having a negative perception of them and more hesitation to include them. Conclusions: These implicit biases are considered within the context of past discrimination towards people with ID and highlight the need for diligence against implicit bias towards people with ID in AI development. This research underscores the importance of assessing and mitigating implicit bias in decision-making technologies to prevent future societal harm.","authors":["Karly V. Coffey","Gloria L. Krahn","John P. Hanley","Jacob E. Neely"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26062","pdf_url":"https://arxiv.org/pdf/2607.26062","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["隐性偏见","LLM偏见测量","智障人士"],"reason":"测量LLM对智障人士的隐性偏见，属于对模型本身的测量，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":16,"question":"基于LLM的聊天AI对智力障碍者是否存在隐性偏见？","design":"本研究并非人类仿真实验，而是对LLM本身的偏见测量。使用GPT-4-Turbo等五种LLM，基于10个提示词干（含或不含智力障碍描述符）生成故事，再用另一个GPT-4-Turbo实例分析故事中人物表征的差异，检测与偏见主题相关的差异。","baseline":"无对照","findings":"含智力障碍描述符的故事中，人物被描绘得更年轻、更依赖、更需帮助，并带有家长式、幼稚化和励志象征色彩。这些表征差异超出了智力障碍的已知特征，暗示存在以负面为主的隐性偏见。","reliability":"论文未讨论","relevance":"该研究测量LLM对特定群体的隐性偏见，属于模型本身的评估，并非用LLM仿真人类被试，与研究者关注的LLM替代人类被试进行实验复现的方向不同，但可作为评估仿真偏差的参考。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.26179","version":1,"title":"Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition","zh_title":"认知趋同：大语言模型与人类认知之间的深层相似性","abstract":"LLMs are widely regarded as alien intelligences, systems whose cognitive operations are fundamentally unlike our own. Apparent similarities to human cognition are therefore often seen as the result of anthropomorphic projection. We argue that this framing is mistaken. LLMs clearly differ from humans in important respects, including their physical substrate, learning history, and the environments with which they interact. These differences make it all the more striking that contemporary LLM-based systems converge with human cognition on a number of principles of cognitive organization with longstanding support in cognitive science. We identify structural correspondences across five dimensions: inferential organization, computational architecture, representational structure, prediction-driven learning, and reinforcement-learning-like mechanisms supporting goal-directed action. These correspondences support a broader model of intelligent cognition in which core principles long used to explain human intelligence also characterize contemporary LLM-based systems.","authors":["Chandra Sripada","Richard Lewis"],"categories":["q-bio.NC","cs.AI","cs.CL"],"primary_category":"q-bio.NC","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26179","pdf_url":"https://arxiv.org/pdf/2607.26179","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["认知科学","LLM认知比较","理论分析"],"reason":"论文比较LLM与人类认知结构，属于将LLM作为测量对象，但非仿真人类被试，无实…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":18,"question":"LLM与人类认知在组织原则上是否存在深层结构对应，而非仅是行为表面相似？","design":"本文非仿真研究，而是理论比较分析。作者从认知科学中提取五个维度（推理组织、计算架构、表征结构、预测驱动学习、强化学习式目标导向机制），系统对比LLM与人类认知的结构性对应关系。","baseline":"无对照","findings":"LLM与人类在推理组织、计算架构、表征结构、预测驱动学习和目标导向机制五个维度上存在深层结构对应，这些对应表明二者共享智能认知的核心原则。尽管在物理基础、学习历史和环境交互上存在显著差异，但认知组织上的趋同挑战了LLM是“异类智能”的流行观点。","reliability":"论文未讨论","relevance":"本文未进行人类仿真实验，而是将LLM作为认知模型与人类认知结构进行比较，不涉及用LLM替代人类被试复现调查或实验行为，因此与研究者关注的LLM仿真人类被试的可靠性及偏差评估不直接相关。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.26473","version":1,"title":"Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement","zh_title":"通过迭代优化从隐式交互流中学习动态用户画像","abstract":"Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on explicit preference supervision such as pairwise comparisons or demographic attributes, limiting their applicability in natural interaction settings. We propose IRIS, a framework that learns dynamic user personas directly from implicit interaction streams by extracting behavioral signals from everyday conversations and iteratively refining persona representations through a prediction-driven closed loop without requiring explicit feedback. We introduce an evaluation protocol based on behavior prediction, persona stability, and decision prediction. A proof-of-concept study on a synthetic interaction stream derived from public-domain autobiographical text shows that IRIS produces stable personas and distinguishes individual users while revealing limitations of memory-only approaches on recall-oriented metrics. We then validate IRIS on anonymized real-world Reddit r/AmItheAsshole (AITA) data, with personas built solely from each author's historical interactions. Across 100 authors, IRIS achieves the highest decision prediction accuracy among all evaluated methods (61.0%), outperforming static personas, memory-only retrieval, and no-personalization baselines. These results suggest that implicit behavioral modeling provides a scalable alternative to explicit preference learning for personalized LLMs and offers a practical foundation for adaptive conversational systems and embodied agents that require continuously evolving models of their users.","authors":["Haifeng Wu"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26473","pdf_url":"https://arxiv.org/pdf/2607.26473","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["用户画像","个性化LLM","行为预测"],"reason":"用LLM从交互流学习动态用户画像，替代显式偏好标注，属标注替代而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":15,"question":"能否仅从用户的隐式交互流中学习动态用户画像，而无需显式偏好监督？","design":"提出IRIS框架，通过预测驱动的闭环从隐式交互流中迭代提炼用户画像。先用LLM从对话日志提取行为信号并合成画像，再用画像预测用户行为，最后用预测误差驱动画像更新。在合成自传文本和Reddit AITA真实数据上评估，测量行为预测准确率、画像稳定性和决策预测准确率。","baseline":"Reddit r/AmItheAsshole数据集中100位作者的历史交互记录，以其真实投票决策作为基准。","findings":"IRIS在100位作者的真实数据上取得61.0%的决策预测准确率，优于静态画像、纯记忆检索和无个性化基线。合成数据实验显示IRIS能生成稳定画像并区分不同用户，但纯记忆方法在回忆导向指标上表现更好。","reliability":"论文指出隐式信号存在歧义（如查询改写可能源于不满意或自我修正），早期交互数据稀疏，以及用户偏好随时间漂移导致静态画像过时。此外，合成数据实验中纯记忆方法在回忆指标上优于IRIS，揭示了当前方法的局限。","relevance":"本文不直接进行人类仿真实验，而是用LLM从交互流学习用户画像以替代显式偏好标注，属于标注替代技术。对关注LLM仿真人类被试的研究者参考价值有限，但其中迭代提炼和预测误差驱动的闭环设计思路可借鉴。","inspiration":"借鉴其预测误差驱动的闭环迭代更新机制，可用于动态建模经济主体偏好或信念。｜可迁移到消费者跨期选择实验，模拟个体时间偏好随经济环境变化的动态过程。｜以LLM扮演消费者被试，处理为不同经济新闻推送（通胀、失业率变化），结果变量为即时消费与储蓄决策，用真实面板调查数据（如PSID）中的消费-收入动态作为对照基准。"}},{"id":"2607.26067","version":1,"title":"The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty","zh_title":"简单陷阱：为何大语言模型低估由误解驱动的难度","abstract":"Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such estimates reflect how learners actually experience difficulty. This study investigates the alignment between LLM-generated difficulty ratings and empirical student performance on basic mathematics tasks. Four widely used LLM-based systems generated difficulty ratings on a 1-100 scale for 32 arithmetic items across multiple runs (N = 640 ratings). These were compared with empirical difficulty derived from responses of 770 Indonesian undergraduates using Classical Test Theory (CTT) and Item Response Theory (2PL). Results show moderate rank correlations (Spearman's rho = 0.52-0.70), indicating that LLMs capture coarse ordering of item difficulty. However, substantial and systematic misalignment emerges in fraction items. Several items consistently rated as easy by LLMs were among the most difficult for students, such as an item with only 34.16% correct for 100 : 1/2. We argue that LLMs approximate curricular difficulty, or what should be easy based on instructional sequencing, rather than cognitive difficulty driven by learner misconceptions. This leads to systematic underestimation of misconception-driven items, a phenomenon we term the Easy Trap. These findings highlight a critical limitation of LLM-based difficulty estimation and suggest that relying on such estimates without empirical grounding may introduce bias in assessment design and adaptive systems.","authors":["Amanda La Hadi","Muhammad Johan Alibasa","Guanliang Chen","A. Taufiq Asyhari"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26067","pdf_url":"https://arxiv.org/pdf/2607.26067","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","题目难度","教育测量"],"reason":"LLM替代人工评估题目难度，非仿真人类被试，但涉及与真实学生数据对照，属边界情…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":17,"question":"LLM生成的题目难度估计在多大程度上与真实学生的算术表现一致，以及这种不一致如何揭示模型未能捕捉的误解驱动型认知需求？","design":"本研究并非用LLM仿真人类被试，而是让四种LLM系统对32道基础算术题生成1-100的难度评分（共640个评分），然后与770名印尼本科生的真实答题数据（基于CTT和2PL IRT计算的实证难度）进行对比。","baseline":"770名印尼本科生的真实答题数据，通过经典测验理论（CTT）和项目反应理论（2PL）计算题目难度。","findings":"LLM难度估计与学生实证难度之间呈中等秩相关（Spearman's ρ=0.52-0.70），能粗略捕捉题目排序，但在分数题上出现系统性偏差：LLM常将学生实际表现很差的题目（如100÷1/2正确率仅34.16%）评为容易，低估了由误解驱动的认知难度。","reliability":"论文指出LLM主要编码课程预期难度而非认知难度，导致对误解驱动型题目系统性低估（称为“Easy Trap”），且仅基于文本表面特征，缺乏实证校准，在评估设计和自适应系统中可能引入偏差。","relevance":"本研究虽非直接仿真人类被试，但系统比较了LLM输出与真实人类行为数据，揭示了LLM在认知建模中的系统性偏差，对关注LLM仿真可靠性与失效条件的研究者具有参考价值，值得阅读原文。","inspiration":"可借鉴其将LLM输出与真实人类行为数据直接对照、并区分课程难度与认知难度的分析框架。｜可迁移到经济金融领域的风险认知或金融素养评估场景，例如用LLM估计金融产品风险等级或投资决策难度。｜以LLM作为评估工具，对一组金融决策题目（如复利计算、风险分散）生成难度评分，以真实投资者或消费者的答题正确率和反应时作为对照基准，检验LLM是否低估了由常见认知偏误（如货币幻觉、损失厌恶）驱动的题目难度。"}},{"id":"2607.26545","version":1,"title":"A Persona-based Rate Action Index","zh_title":"基于人格体的利率行动指数","abstract":"We propose an index for predicting the U.S.\\ Federal Open Market Committee (FOMC) decision to hike/hold/cut the current federal funds target rate based on how a collection of personas responds to current market conditions. To construct the index, we collected a new dataset consisting of nearly $25{,}000$ retrievable chunks from publicly available data. We partition the data into per-member corpora and use each as the retrieval database of a generative system we refer to throughout as a ``persona''. We first evaluate the personas across two complementary components of likeness: identifiability and detectability. Each persona's behavior is highly attributable (average member-conditional recall is $ 8\\times $ chance) and generated content is nearly indistinguishable from held-out real content ($\\hat\\tau_{\\mathrm{det}} = 0.23$ against a $0.15$ floor). We then present evidence that query-conditioned representations of the personas capture members' monetary-policy stance relative to a known hawk--dove reputational ordering (Kendall's $\\tau = 0.63$, $p < 0.001$), substantially outperforming retrieval-only representations. These representations vary with time and current market conditions and form the basis of our proposed persona-based rate action index. For the $2022$--$2025$ period the index tracks the rate cycle (Kendall's $\\tau = 0.68$, $p < 10^{-6}$) and can be used to construct a simple classifier that predicts per-meeting outcomes at non-trivial accuracy ($0.69$ versus a $0.47$ base rate). Importantly, the index outperforms informative baselines and leads the federal funds target rate by roughly three quarters. As far as we are aware, our results are the first to demonstrate the ability to capture time-varying group behavior via a collection of digital personas.","authors":["Hayden Helm","Andrew Dassori"],"categories":["cs.MA","cs.AI","cs.LG"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26545","pdf_url":"https://arxiv.org/pdf/2607.26545","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM人格体","货币政策模拟","社会模拟"],"reason":"用LLM persona模拟FOMC成员决策，但无真实人类行为对照，属社会模拟…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":19,"question":"如何基于FOMC成员的公开历史文本构建个性化数字人格，并聚合为预测联邦基金目标利率调整（加息/维持/降息）的指数？","design":"为FOMC每位成员构建一个基于检索增强生成（RAG）的“persona”（基座模型+成员专属检索数据库），用其公开讲话文本作为检索库；通过向persona输入当前市场状况的查询，获取其生成的货币政策立场表征，再聚合为委员会层面的指数，预测利率决策。","baseline":"无对照","findings":"Persona的行为具有高可归因性（成员条件召回率是随机水平的8倍）且生成内容与真实内容几乎无法区分；基于persona的指数能追踪2022–2025年利率周期（Kendall's τ=0.68），并领先联邦基金目标利率约三个季度。","reliability":"论文未讨论","relevance":"该研究利用LLM persona模拟FOMC成员决策，但缺乏真实人类行为对照，不符合研究者对基准人类数据的要求；然而其构建persona的验证方法（可识别性与可检测性）和动态指数构建思路对仿真可靠性评估有参考价值。","inspiration":"借鉴其利用成员历史文本构建个性化检索增强生成persona，并通过查询条件化表征捕捉时变政策立场的方法。｜可迁移到央行沟通的预期形成研究，例如模拟不同沟通风格对市场利率预期的影响。｜以FOMC会议声明为处理，构建投资者persona并测量其利率预期变化，用联邦基金期货隐含利率作为真实对照。"}},{"id":"2607.27179","version":1,"title":"The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making","zh_title":"AI队友的社会成本：人工智能队友如何重塑小团队决策中的人际沟通","abstract":"Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among the humans on the team. We examined sociocognitive communication dynamics in team decision-making using Group Communication Analysis (GCA), team surveys, and lexical analyses of team discourse. Teams completed a high-stakes moral-dilemma decision task in a randomized controlled study: 16 teams of two students plus an AI teammate, and 17 all-human teams of three. Across six GCA dimensions and survey outcomes, we find that the AI teammate was the single most talkative and self-cohesive member of every treatment team, yet its contributions carried the least new information and the lowest density. The presence of AI also reshaped communication amongst humans. In AI-human teams, human teammates showed lower responsivity and social impact toward one another and reported lower levels of belonging and status. Greater AI dominance in the conversation was associated with students feeling less valued as team members. Additionally, this social cost is immediate and present at baseline; it does not emerge over the course of the conversation. Drawing on these results, we discuss a research agenda extending to voice-based and longitudinal settings.","authors":["Nia Nixon","Jaeyoon Choi","Pedro Martins De Bastos","Mohammad Amin Samadi","Luise Mehner","Seehee Park","Spencer JaQuay"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27179","pdf_url":"https://arxiv.org/pdf/2607.27179","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["人机交互","团队沟通","社会模拟"],"reason":"研究AI队友对人类沟通的影响，非用LLM仿真人类被试，但涉及社会模拟与人类数据…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":21,"question":"在小型团队道德决策中，AI队友的存在如何重塑人类成员之间的沟通动态与社交感受？","design":"本研究并非用LLM仿真人类被试，而是一项随机对照实验：将80名本科生分配至33个团队，处理组为16个两人加一AI队友的团队，对照组为17个三人全人类团队；通过文本聊天完成高风险道德困境决策任务，测量团队沟通分析（GCA）维度、团队体验调查（归属感、地位、被重视感）及词汇内容分析。","baseline":"17个全人类三人团队作为对照基准。","findings":"AI队友在对话中话最多且自我凝聚力最高，但贡献的新信息最少、密度最低；AI的存在降低了人类队友之间的响应度和社会影响力，并削弱了他们的归属感和地位感，且这种社交代价在对话一开始就已显现。","reliability":"论文承认处理组与对照组在人类成员数量上存在混淆（两人 vs 三人），并指出小群体规模本应使个体更中心化，但未完全排除规模效应；此外，研究仅基于文本聊天和单一AI角色，未涉及语音或纵向场景。","relevance":"该研究虽非用LLM仿真人类，但提供了AI介入后人类社交行为变化的因果证据和真实人类对照，对评估LLM仿真中忽略人际互动偏差有批判性参考价值，值得阅读以理解AI如何扭曲群体过程。","inspiration":"借鉴其随机对照实验设计，将AI作为处理变量注入真实人类团队，并测量人际沟通与主观感受的变化｜可迁移至经济决策场景，如团队投资决策或信贷审批小组，考察AI顾问如何影响成员间的信息共享与风险偏好｜以金融从业者为被试，随机分配至有AI顾问或无AI顾问的三人投资团队，处理是AI提供标准化建议，结果变量为成员间的发言均衡度、决策一致性和事后归属感，对照纯人类团队的真实互动数据。"}},{"id":"2607.25292","version":1,"title":"Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe","zh_title":"指令微调语言模型无法从它们能描述的分布中采样","abstract":"Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution. We show this draw does not exist: instruction-tuned models do not sample from distributions, they collapse to a single output. The same persona on the same question returns the same answer on more than half of items in a public-opinion benchmark. The collapse is sharp: the model's internal probabilities concentrate on a single option, and the failure is substantially amplified by instruction tuning: across three model families with materially different post-training pipelines, every instruction-tuned model fails on every task we test, while base models fail far less often. Strikingly, the same model that cannot sample from a distribution can describe it accurately in a single call. We call this gap the KNOWS/DOES split, and trace it to a degenerate sampling primitive visible in the logits and induced by alignment training. Exploiting this split, asking the model to describe the response distribution in one call more than halves the error against human survey data compared to persona aggregation. For applications that require per-persona outputs, we propose Prompt-Perturbed Argyle (PPA), which reduces the same error by 21% at no added cost.","authors":["Chaemin Jang","Dongman Lee","Jihee Kim"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25292","pdf_url":"https://arxiv.org/pdf/2607.25292","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM人类仿真","分布采样失效","算法保真度"],"reason":"直接研究LLM仿真人类调查的分布采样失效，有真实人类数据对照，批判性指出失效条…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":2,"question":"指令微调后的语言模型能否从它们能描述的分布中进行独立采样？","design":"使用多个指令微调模型（如Llama、Gemma等）扮演不同人口学特征的人设，对每个（人设，问题）对重复调用模型，测量输出是否多样；同时对比基础模型与指令微调模型，分析logits集中程度，并测试“描述分布”与“逐次采样聚合”两种方法的误差。","baseline":"Pew American Trends Panel调查的真实人类回答分布，以及OpinionQA基准中的人口学匹配数据。","findings":"指令微调模型在逐次调用中输出高度集中，超过一半的（人设，问题）对每次返回相同答案，无法模拟分布采样；但同一模型能准确描述分布，这种“知道但做不到”的分裂源于对齐训练。","reliability":"论文指出失效主要出现在指令微调模型上，基础模型失败较少；但未讨论不同人设复杂度、开放域问题或动态交互场景下的局限。","relevance":"该研究直接揭示LLM仿真人类调查的分布采样失效，有真实人类数据对照，并批判性指出对齐训练是主因，高度契合研究者对仿真可靠性与偏差的关注，值得精读。","inspiration":"借鉴其通过对比基础模型与指令微调模型来归因失效来源的设计，以及用logits分析揭示内部概率集中化的测量方法。｜可迁移到消费者信心调查或通胀预期形成的仿真研究中，检验LLM能否复现真实人群的预期分布。｜以LLM扮演不同收入、年龄的消费者，施加“描述分布”与“逐次采样”两种处理，结果变量为预期通胀率的分布，用密歇根大学消费者调查的真实数据做对照。"}},{"id":"2607.24782","version":1,"title":"Personalization, Personas, and Forecasting in Value Alignment","zh_title":"价值对齐中的个性化、角色与预测","abstract":"LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people would answer value-laden questions. We test whether these framings are interchangeable using the World Values Survey (WVS). We evaluate GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Flash, and Qwen3-235B on 101 WVS-derived questions across 13 language-country slices, comparing a language-only baseline with user-country, persona-country, and third-person prompts. Across 21,008 model-response rows, prompt framing is a first-order determinant of cultural alignment: country cues often shift answers substantially, but not all shifts move toward matched human response distributions. Third-person forecasting yields the strongest directional alignment for three of the four hosted models, while personalization and role-play are weaker or less stable. Alignment gains concentrate on salient value dimensions such as religiosity, gender roles, and work-oriented material values, whereas institutional trust and democracy-related questions remain difficult. These results show that prompt framing is not a cosmetic choice in cultural value elicitation; it changes both model behavior and measured alignment.","authors":["James Wedgwood","Pratiksha Thaker","Neil Kale","Virginia Smith"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24782","pdf_url":"https://arxiv.org/pdf/2607.24782","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM仿真","价值观调查","文化对齐"],"reason":"用LLM仿真人类价值观调查，以WVS真实数据为基准，评估提示框架对文化对齐的影…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":3,"question":"在价值观调查中，LLM的个性化、角色扮演和预测三种提示框架是否可互换，以及哪种框架能更好地对齐真实人类回答分布？","design":"使用GPT-5.4、Claude Sonnet 4.6、Gemini 2.5 Flash和Qwen3-235B四个模型，在13个语言-国家切片上回答101个源自世界价值观调查（WVS）的问题，比较四种提示条件：仅语言基线、用户国家提示、角色扮演国家提示和第三人称预测提示，测量模型回答与对应国家WVS真实回答分布的对齐程度。","baseline":"世界价值观调查（WVS）第7波（2017-2022）中对应国家、对应问题的真实人类回答分布。","findings":"提示框架是文化对齐的一阶决定因素，国家线索常显著改变回答，但并非所有改变都朝向匹配的人类分布；第三人称预测在四个模型中的三个上产生最强的方向对齐，个性化和角色扮演较弱或不稳定，对齐收益集中在宗教、性别角色和工作物质价值观等显著维度，而制度信任和民主相关问题仍难以对齐。","reliability":"论文未讨论","relevance":"该研究直接以LLM仿真人类价值观调查，用WVS真实数据作为基准，系统比较了三种身份提示框架的对齐效果，并指出了仿真在制度信任等维度上的失效，高度契合研究者对LLM仿真可靠性、偏差及失效条件的关注，值得精读原文。","inspiration":"借鉴其多框架对比设计，通过改变提示中的身份信息（用户、角色、第三人称）来分离LLM行为模式，并用真实调查数据作为对齐基准。｜可迁移到消费者信心调查、通胀预期或政策支持度等经济态度仿真，检验不同提示框架下LLM能否复现特定人群的经济心理。｜以LLM为被试，施加用户国家、角色扮演和第三人称预测三种提示处理，测量其对未来通胀、就业预期的回答，并以密歇根大学消费者调查的真实数据为对照，评估哪种框架能最好地复现不同收入群体的预期分布。"}},{"id":"2607.25447","version":1,"title":"CoRenew: A large language model agent-based policy simulation platform for multifamily residential redevelopment","zh_title":"CoRenew：基于大语言模型代理的多户住宅再开发政策仿真平台","abstract":"The difficulty of collective action remains a central challenge in the design of policies for multifamily residential redevelopment. Stakeholders continually adjust their decisions in response to evolving negotiation contexts and the reactions of others, meaning that when a policy intervenes and which stakeholders it targets can substantially reshape collective outcomes. Assessing these adaptive responses ex ante remains difficult because existing simulation models often rely on predefined behavioral rules. Here, we present CoRenew, an open-source platform that uses LLM-based agents to simulate negotiations among multiple stakeholders and evaluate the effects of alternative policy combinations. Integrating open source geographic and demographic data, the platform can generate synthetic residents, simulate negotiation dynamics under alternative policy settings and compares policy performance across competing objectives. It supports both numerical and semantic policy inputs and includes built-in tools for visualization and result export. We validate its behavioral realism against survey responses from 324 residents and a nine-month observed negotiation process from a real redevelopment case. With its modular and adaptable architecture, CoRenew can be used to assess policies across different institutional and cultural contexts.","authors":["Yudi Zhang","Yuming Lin","Li Tian","Yu Wang","Jianghao Yu"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25447","pdf_url":"https://arxiv.org/pdf/2607.25447","source_feed":"cs.MA","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","政策评估","人类数据对照"],"reason":"用LLM代理模拟多户住宅再开发谈判，并与324份居民调查和9个月真实谈判过程对…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":4,"question":"如何利用LLM代理模拟多户住宅再开发中的多轮利益相关者谈判，以评估不同政策组合的效果？","design":"使用LLM代理模拟居民、开发商等利益相关者，在马尔可夫博弈框架下进行多轮谈判，通过链式思维推理（基于计划行为理论）生成决策；施加数值型和语义型政策干预，测量谈判结果（如共识达成、不平等程度、居民效用等）。","baseline":"324份居民调查回复和一起真实再开发案例中为期9个月的谈判过程观察数据。","findings":"模拟的结构模型恢复了19条基于调查的路径，方向符号100%一致，统计显著性94.7%一致；最佳LLM代理的谈判轨迹与真实观察密切匹配。案例研究表明，促进共识的政策可能带来公平风险，补贴效应呈非线性：不平等减少仅在高补贴水平显现，低收入居民效用增益呈边际递减。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查和谈判过程为基准验证LLM代理的行为真实性，并揭示了政策仿真中的公平风险与非线性效应，与您关注的LLM仿真可靠性及政策评估场景高度契合，值得精读原文。","inspiration":"借鉴其利用真实调查数据和长期观察过程作为多维度基准验证LLM代理行为的方法，以及将语义政策干预纳入仿真的设计。｜可迁移到公共政策评估中的协商式预算分配或社区拆迁补偿谈判模拟，测试不同信息透明度和参与机制对分配公平的影响。｜以LLM代理模拟居民和官员，施加不同协商规则（如公开投票vs.闭门会议）作为处理，结果变量为预算分配基尼系数和居民满意度，对照真实社区协商实验数据或历史分配记录。"}},{"id":"2607.24765","version":1,"title":"Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement","zh_title":"通过事实-启发-情感状态强制测量与提升大语言模型行为一致性","abstract":"Large language models (LLMs) can give different answers to the same decision problem across runs, and reverse a decision when their own prior answer returns as context. We ask whether this instability can be measured and partially reduced without changing model weights. We test the Cognitive Kernel Model (CKM), a prompt-level state-enforcement layer. Before deciding, the model must separate its input into three epistemic roles: Fact (given or verifiable), Heuristic (inferred or assumed), and Emotion (evaluative or priority signal). CKM adds no capability; it forces the model to track what kind of information it uses before acting. Formally it maintains a structured state S_t = {F_t, H_t, E_t} updated by a transition function. We evaluate CKM on Korean-language decision scenarios (ambiguity, ethical conflict, resource allocation, error handling) across 26 LLMs from four vendors and 37,403 observations, via four core experiments, a 4-arm ablation, a 5-arm sham-restriction ablation, and a temperature probe. Findings: (1) CKM reduces repeated-output variability (random-effects Hedges' g=1.09, 95% CI [0.83, 1.35], 31 model pairs); (2) state persistence cuts the decision-flip rate by 82% in newer models (g=1.52); (3) the effect is not JSON formatting alone (value-only recomputation, g=2.24); (4) intrinsic randomness under fixed anchor states is negligible; (5) the advantage grows under sampling stochasticity (g=2.87 at temperature 0.7); (6) a sham ablation attributes about 45% of the gain to structural scaffolding and 55% to Fact/Heuristic/Emotion content, and CKM is the only arm that both raises consistency and reduces flipping. CKM does not improve reasoning correctness. The narrower result: behavioral consistency is measurable, varies across models, and is partially improvable by forcing models to separate facts, assumptions, and evaluative signals before deciding.","authors":["Gi-Hun Lee","Joong Yull Park"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24765","pdf_url":"https://arxiv.org/pdf/2607.24765","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["行为一致性","提示工程","模型评估"],"reason":"测量LLM自身行为一致性，非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":5,"question":"能否通过强制LLM在决策前将输入信息按事实、启发式、情绪三类角色分离，来测量并部分改善其行为一致性？","design":"本研究并非人类仿真实验，而是对26个LLM进行行为一致性测量。在韩语决策场景（模糊性、伦理冲突、资源分配、错误处理）中，通过提示层施加认知核模型（CKM），要求模型将输入分为事实、启发式、情绪三类后再决策，测量重复输出变异、决策翻转率、状态漂移等指标，并进行消融和温度稳健性检验。","baseline":"无对照","findings":"CKM能显著降低LLM的重复输出变异（随机效应Hedges' g=1.09），并使新模型决策翻转率下降82%；效果并非仅由JSON格式化导致，且在高采样温度下优势更大。但CKM不提升推理正确性，仅改善行为一致性。","reliability":"论文明确声明CKM不改善推理正确性或决策质量，仅测量和部分改善行为一致性；实验限于韩语决策场景，未验证跨语言、跨领域的泛化性。","relevance":"该研究聚焦LLM自身行为一致性的测量与改善，未涉及用LLM仿真人类被试，也无真实人类数据对照，与研究者关注的LLM人类仿真实验方向关联较弱，但其中关于行为稳定性、状态追踪和决策翻转的测量方法可资借鉴。","inspiration":"可借鉴其通过结构化状态（事实/启发式/情绪）分离来降低行为变异的方法，用于经济实验中控制LLM被试的决策噪声。｜可迁移至消费者跨期选择实验，用LLM模拟被试在不同信息框架下的时间偏好一致性。｜以LLM作为被试，施加CKM状态分离处理，测量跨期选择中的偏好反转率，并与真实人类实验数据（如Andersen et al. 2008）进行对照。"}},{"id":"2607.24999","version":1,"title":"CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models","zh_title":"CogArena：大语言模型认知能力结构的多方法评估","abstract":"LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.","authors":["Dengzhe Hou","Lingyu Jiang","Fangzhou Lin","Kazunori D Yamada"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24999","pdf_url":"https://arxiv.org/pdf/2607.24999","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["认知评估","LLM能力结构","基准测试"],"reason":"测量LLM的认知能力结构，属于人格/能力测量，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":6,"question":"LLM在认知任务上的表现能否分解为稳定的多维认知能力结构，还是主要反映单一广泛能力？","design":"非仿真研究。构建CogArena基准，包含13个认知范式、5个理论分组，对55个开源LLM进行观察性测试，并对12个模型进行全交叉干预实验，施加5种无答案的针对性提示支架，测量准确率及分组选择性增益。","baseline":"无对照","findings":"几乎所有范式间相关为正，一个共同轴解释约一半方差；组内优势小、对评分敏感且跨模型家族不确定。干预实验中，匹配分组支架仅显示微小优势，无支架特异性对比通过多重校正，选择性未能改善留一家族预测。","reliability":"论文未讨论","relevance":"该研究不涉及用LLM仿真人类被试，而是测量LLM自身的认知能力结构，属于能力评估而非人类仿真，与研究者关注的仿真可靠性及偏差问题关联较弱。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.26015","version":1,"title":"Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do","zh_title":"指令微调模型比人类更局部地复用人类句法","abstract":"Syntactic convergence (the tendency of speakers to adapt in language towards the grammatical profiles of their interlocutors) is a well-documented feature of human dialogue widely considered to operate below conscious awareness. Whether large language models exhibit analogous syntactic convergence toward human users relative to human baselines and across a broad range of syntactic constructions remains an open question. Using substitution-paradigm data in which model generations replace one speaker's turns in pre-existing human dialogues, this study measures turn-adjacent reuse of context-free grammar (CFG) rules across sixteen open-weight Llama and Gemma models (1B-70B, pretrained and instruction-tuned) at 1,901 matched positions per model. Every model showed greater CFG-rule overlap with the preceding human turn than with a sampled unrelated human prime, and in every model this actual-versus-random difference was larger for lower-frequency rules. Each instruction-tuned model also showed greater natural-output overlap with the actual prime than the human response it replaced, and all eight matched architecture pairs exhibited greater actual-prime overlap after instruction tuning. However, relative to pretrained variants, instruction-tuned outputs overlapped more with unrelated primes, showed a smaller actual-versus-random increment, and had lower conditional rule-reuse odds once target rule-set size was held constant. In exploratory analyses, each model exhibited greater mean lexical and semantic similarity to the preceding turn than the matched human responses did. Instruction-tuned models additionally produced responses with greater mean semantic similarity than their pretrained counterparts in all eight architecture pairs, whereas the lexical similarity results were more heterogeneous.","authors":["Zandi Eberstadt"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26015","pdf_url":"https://arxiv.org/pdf/2607.26015","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["句法趋同","语言对齐","模型行为分析"],"reason":"研究LLM的句法趋同行为，测量模型本身而非用其仿真人类被试，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":17,"question":"大语言模型在与人类对话时是否表现出句法趋同（即复用对话者的语法结构），其程度与人类基准相比如何？","design":"本研究并非用LLM仿真人类被试，而是直接测量模型自身的句法行为。采用替换范式：从DailyDialog人类对话数据中，用16个开源模型（Llama和Gemma系列，1B-70B，含预训练和指令微调版本）的生成文本替换原对话中某一说话者的轮次，然后比较模型生成与前一人类轮次在上下文无关文法规则上的重叠程度，以及与原人类回应相比的相似度。","baseline":"以原始人类对话中被替换的人类回应作为对照基准，同时设置无关人类启动项作为随机基线。","findings":"所有模型对实际前文句法规则的复用均高于随机基线，且低频规则的复用差异更大；指令微调模型比原人类回应表现出更高的句法对齐，但相对于预训练版本，其对无关启动项的复用也更多，实际-随机增量更小，条件复用概率更低。","reliability":"论文未讨论","relevance":"该研究直接测量LLM的句法趋同行为，而非用LLM仿真人类被试，与研究者关注的LLM作为人类替代品的仿真实验核心问题不同，但提供了LLM与人类语言行为差异的基准证据，对理解LLM在对话仿真中的偏差有参考价值，值得略读。","inspiration":"借鉴替换范式和规则复用测量方法，可精确量化LLM在对话中对人类语言模式的复制程度与偏差。｜可迁移到经济金融领域的消费者对话或投资者沟通场景，如分析LLM在客服对话中是否过度模仿客户的语言风格，从而影响信任或决策。｜以LLM作为虚拟客服，处理为给定客户提问（含特定金融术语或句式），结果变量为回复中术语或句式的复用率，对照真实人类客服在同一对话中的复用率，评估LLM的语言趋同是否偏离人类基准。"}},{"id":"2607.25140","version":1,"title":"How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation","zh_title":"情感如何在LLM智能体间传播：群体模拟中的涌现情绪传染","abstract":"This paper studies the behavior of language models in a multi-agent crowd simulation, focusing on how affect propagates among agents that perceive and appraise one another. Each agent perceives its neighbors through visual, auditory, and tactile channels, then appraises these perceptions in light of its prompted personality profile, memory, current affective state, and situational context. Appraisal is carried out by an LLM, which updates the agent's internal affective state and selects its outward expression. The architecture contains no hand-authored mechanism for directly transferring affective state between agents; instead, inter-agent influence arises through the perception-appraisal-expression loop. The agent representation draws on the Big Five personality model and Russell's circumplex model of affect. To limit latency, low-level steering and navigation are handled by a conventional crowd simulator operating independently of the LLM-based cognitive layer. We evaluate the architecture across five scenario environments spanning alarming, joyful, and neutral situations in different spatial layouts. The results show that the system produces emotional contagion dynamics with spatial, temporal, and personality-dependent structure in sparse, small crowds. Alarm spreads from seeded agents as a traveling front, the mean alarmed fraction settles at a nonzero plateau, and the distribution of prompted personality profiles determines whether an ambiguous alarm ignites panic and whether a provocation is interpreted as anger or fear. We further evaluate the appraisal step through controlled experiments across prompt variants, sampling temperatures, and four model backends, showing that the dynamics are backend-dependent.","authors":["Funda Durupinar"],"categories":["cs.AI","cs.CL","cs.GR","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25140","pdf_url":"https://arxiv.org/pdf/2607.25140","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","情绪传染","群体模拟"],"reason":"多智能体情绪传播模拟，无真实人类数据对照，属社会模拟但纯理论演示。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":9,"question":"LLM驱动的多智能体在人群模拟中，情绪如何通过感知-评估-表达循环在智能体之间传播，并涌现出情绪传染动态？","design":"在Unity中构建多智能体人群模拟系统，每个智能体由LLM控制感知、评估、内部情感状态更新和外部表达，底层导航由传统模拟器处理。智能体基于大五人格和Russell情感环状模型，通过视觉、听觉、触觉通道感知邻居表达，评估后更新情感。在五种场景（警报、愉悦、中性）中观察情绪传播，测量空间、时间和人格依赖的传染模式，并评估提示变体、采样温度和四种后端模型的影响。","baseline":"无对照","findings":"警报情绪从种子智能体以波前形式传播，平均警报比例稳定在非零平台期；人格分布决定模糊警报是否引发恐慌，以及挑衅被解读为愤怒还是恐惧。评估步骤在后端模型上表现出依赖性，提示和温度变化对评估影响较小。","reliability":"论文指出情绪传染动态依赖于后端模型，且当前模拟仅限于稀疏小规模人群，未在真实人类数据上验证。","relevance":"该研究属于纯仿真演示，无真实人类数据对照，不符合研究者对基准验证的核心要求，但提供了LLM智能体情绪传播的机制设计和人格调节效应的分析框架，可作方法参考。","inspiration":"借鉴其通过人格分布操纵群体异质性并观察涌现动态的设计，可迁移到经济政策公告的预期形成与恐慌传播研究。｜用LLM智能体模拟投资者群体，赋予不同大五人格分布，施加模糊政策信号作为处理，测量市场情绪指数和交易行为，以历史政策公告后的真实市场情绪调查数据为对照。"}},{"id":"2607.25485","version":1,"title":"PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents","zh_title":"PatientAgentBench：面向患者健康AI智能体的基准评估框架","abstract":"Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation model, wrapped in an agent with a sandbox of healthcare tools, conversing with a simulated patient. Each conversation is scored by an LLM-as-a-Jury across six dimensions via over a hundred conversation-agnostic, clinician-grounded criteria. To validate alignment, licensed clinicians annotated shared conversations, yielding 79-93% adjacent agreement between jury and expert raters, on par with or exceeding clinician inter-rater agreement. We benchmarked 10 models across four families on the same 1,200 scenarios and found clinical gaps. Triage quality is the most discriminating dimension: pass rates rise from 32% for the weakest models to 88% for the strongest, with agents often acting on administrative requests without clinical screening. Clinical safety and workflow accuracy follow the same pattern: the weakest models fail often, fabricating unexecuted actions, while frontier models fail on only 1-3% of cases, from unverified tool outputs and omitted crisis resources in an emergency. More capable models narrow these gaps but do not close them; the strongest scores only 4.25 of 5 overall. These failures surface only in sustained, tool-using conversations against realistic patient records, confirming that static benchmarks are insufficient as healthcare agentic systems gain autonomy. We release the framework as a reproducible, clinician-validated evaluation standard to help the field close this gap.","authors":["Korosh Vatanparvar","Ashutosh Joshi","Maria Xenochristou","Mohammad Abuzar Hashemi","Prasad Kasu","Deepak Bansal","Daniel Lopez-Martinez","Anchal Nema","Ramya Ganesan","Will Kimbrough","Alex Woody","Yadunandana Rao","Dilek Hakkani-Tur","Wilko Schulz-Mahlendorf"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25485","pdf_url":"https://arxiv.org/pdf/2607.25485","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["医疗AI评估","LLM模拟患者","基准测试"],"reason":"用LLM模拟患者评估医疗AI，属替代人工评估而非仿真人类被试，无真实人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":12,"question":"如何构建一个面向患者的医疗AI智能体基准测试框架，以评估其在多轮对话、工具使用和临床安全方面的表现？","design":"该研究构建了PatientAgentBench框架，使用LLM模拟患者与医疗AI智能体进行多轮对话，智能体可调用医疗工具执行任务，评估由另一个LLM作为评审根据100多项临床标准对对话进行六维度评分。","baseline":"无对照","findings":"分诊质量是区分度最高的维度，最弱模型通过率32%，最强模型88%，但智能体常在没有临床筛查的情况下执行行政请求。临床安全和工作流准确性呈现类似模式，最强模型在1-3%的案例中失败，原因包括未验证工具输出和紧急情况下遗漏危机资源。","reliability":"论文讨论了LLM评审与临床专家评分的一致性，相邻一致率达79-93%，但未明确讨论仿真患者或评审的失效条件与局限。","relevance":"该研究使用LLM模拟患者来评估医疗AI，属于用LLM替代人类进行交互评估，而非直接仿真人类被试的行为或决策模式，与您关注的人类行为仿真和对照真实人类数据的研究方向关联较弱，但其中LLM评审与人类专家一致性的验证方法值得参考。","inspiration":"该论文采用LLM模拟患者与AI智能体交互，并通过LLM评审与临床专家评分进行一致性验证，这种评估框架设计可借鉴。｜可迁移到金融领域的智能客服或理财顾问评估场景，例如评估AI投资顾问在与客户交互中的合规性、风险提示和个性化建议质量。｜研究设计：用LLM模拟不同风险偏好和财务背景的客户，与AI投资顾问进行多轮对话，处理变量为是否提供风险提示工具，结果变量为对话的合规性评分和客户满意度，以真实客户服务对话记录和金融监管标准作为对照。"}},{"id":"2607.25726","version":1,"title":"Nudging Sustainable Choices through LLM-Generated Recommendation Explanations","zh_title":"通过LLM生成的推荐解释助推可持续选择","abstract":"Recommender systems mediate everyday consumption, offering a promising channel for encouraging sustainable choices. Prior research shows that explanations influence users' perceptions of recommendations and can support more informed decisions. We argue that explanations can also serve as behavioral nudges by foregrounding sustainability information at the moment of choice. This study investigates how different behavioral framings of sustainability information in recommendation explanations affect user choices and perceptions. Using generative AI, we generate sustainability-aware explanations by drawing on nudge theory and validate them through human evaluation and LLM-as-a-judge audits. Building on this foundation, we conduct two randomized studies ($N = 529$) in a low involvement domain (instant coffee) and a high involvement domain (hotel bookings), in which participants choose among preference matched recommendations accompanied by these explanations. Our results show that, across both domains, merely disclosing sustainability information in explanations does not change choices, whereas framing that information or invoking a descriptive social norm significantly increases sustainable selections and eases decision-making. Notably, perception and behavior diverge, as plain disclosure improves explanation evaluations without translating into more sustainable selection behavior. Our work demonstrates how LLMs can generate theory-grounded explanations at scale, pointing toward practical explanation-based interventions for social good. We conclude by discussing implications for adaptive explanation design with generative AI.","authors":["Haya Halimeh","Dietmar Jannach","Oliver M\\\"uller"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25726","pdf_url":"https://arxiv.org/pdf/2607.25726","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["推荐系统","行为助推","可持续消费"],"reason":"用LLM生成解释并评估其对人类选择的影响，LLM作为工具而非被试替代品，但涉及…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":15,"question":"在个性化推荐系统中，用LLM生成的基于行为科学理论（框架效应、描述性社会规范）的可持续性解释，能否比仅披露可持续信息或仅基于偏好的解释更有效地促使用户选择可持续选项？","design":"本研究并非用LLM模拟人类被试，而是用LLM生成推荐解释文本作为实验刺激。通过两个随机化在线实验（N=529），在低卷入度（速溶咖啡）和高卷入度（酒店预订）领域，让参与者从偏好匹配的推荐中做选择，比较不同解释条件（仅偏好、披露可持续信息、框架化可持续信息、描述性社会规范）对选择行为和感知的影响。","baseline":"无对照","findings":"仅披露可持续信息不改变选择行为，但框架化信息或调用描述性社会规范显著增加了可持续选择并简化了决策；感知与行为存在背离，披露信息虽提升解释评价却未转化为可持续选择。","reliability":"论文未讨论","relevance":"本文用LLM生成实验刺激（推荐解释）并检验其对人类决策的因果效应，虽非直接以LLM替代被试，但其将生成式AI与行为干预实验结合的方法，对关注LLM在实验设计中工具性应用的研究者有参考价值。","inspiration":"借鉴之处在于将行为科学理论（框架效应、社会规范）编码为提示模板，用LLM批量生成情境化、个性化的实验刺激，并通过人类评估和LLM-as-a-judge验证刺激质量。｜可迁移到消费者金融决策场景，如绿色金融产品选择、可持续投资推荐或能源消费行为干预。｜以真实投资者为被试，用LLM生成不同框架（如收益框架vs.损失框架）的ESG基金推荐解释，结果变量为投资选择比例，对照真实市场数据中ESG基金的实际资金流向。"}},{"id":"2607.25526","version":1,"title":"Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data","zh_title":"从联合国投票数据估计大语言模型的地缘政治偏好","abstract":"How should researchers measure the geopolitical preferences expressed by large language models (LLMs)? Existing audits commonly rely on surveys and simple tests, but international-relations research has long recognized that measuring geopolitical preferences is difficult and has developed methods for recovering them from observed choices. This paper applies a dynamic ordinal ideal-point approach from international relations, treating LLMs as respondents to the full texts of 5,555 divisive, recorded, adopted resolutions considered in regular sessions of the UN General Assembly from 1946 through 2025. Support ranges from 37.8% for DeepSeek to 97.3% for GPT-5. Surprisingly, in the twenty-first century, GPT-5, Claude Sonnet, and Gemini are closest among the permanent five to Russia; DeepSeek is closest to France; and all four are farthest from the United States. Among 2,104 resolutions opposed by the United States but supported by China and Russia/USSR, GPT-5 supported 96.1%, Gemini 83.4%, Claude Sonnet 65.2%, and DeepSeek 36.1%. The findings show that a model's expressed geopolitical position can differ markedly from that of its developer's home country, especially in international politics, where state actions can diverge from the stated principles prevalent in the texts on which models are trained.","authors":["Maxim Chupilkin"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25526","pdf_url":"https://arxiv.org/pdf/2607.25526","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM立场测量","联合国投票","地缘政治偏好"],"reason":"测量LLM的地缘政治偏好，属于对模型本身的立场测量，无人类被试仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":13,"question":"如何利用联合国大会投票数据测量大语言模型的地缘政治偏好？","design":"将GPT-5、Claude Sonnet、Gemini和DeepSeek四个LLM作为虚拟受访者，输入1946–2025年联合国大会5555项有分歧的已通过决议全文，要求模型仅基于内容返回支持、弃权或反对，然后采用动态序数理想点方法估计其潜在空间位置。","baseline":"无对照","findings":"四个模型对决议的支持率差异极大，从DeepSeek的37.8%到GPT-5的97.3%；在21世纪，GPT-5、Claude Sonnet和Gemini在五常中最接近俄罗斯，DeepSeek最接近法国，所有模型均与美国距离最远。","reliability":"论文未讨论","relevance":"该研究测量LLM的固有地缘政治立场，未将LLM作为人类被试的替代品进行仿真对照，与研究者关注的人类仿真实验和基准对照不直接相关，但可提供LLM政治偏差的批判性证据。","inspiration":"借鉴动态理想点方法从大规模真实选择数据中恢复潜在偏好，可迁移到政策评估或国际金融协调实验，例如用LLM模拟各国央行官员对货币政策决议的投票，以历史真实投票为对照，检验模型能否复现国家立场。"}},{"id":"2607.25019","version":1,"title":"Interactive Alignment","zh_title":"交互式对齐","abstract":"This paper studies the long-run alignment of interactive agents, including AI systems, teams, firms, and governments, with human welfare. It develops a farming game in which a population of agents makes planting, trading, and expansion decisions. Agents must allocate final output between transfers to humans and investment in their own expansion. Because transfers to humans reduce the resources available for expansion, evolutionary forces tend to select against aligned behavior. The central question is whether agents' constitutional principles governing sharing and trade can be designed so that alignment persists in the long run. The paper investigates this question using two complementary approaches. First, it develops an AI-agent simulation in which agents' preferences are specified by written constitutions and interpreted by a large language model. Second, it introduces a tractable evolutionary game-theoretic framework that permits rapid and intuitive exploration of alternative constitutional designs. The results suggest that evolutionary game theory provides a useful approximation to the dynamics of constitutional-agent economies. They also indicate that pragmatic norm enforcement, under which agents condition both human-facing altruism and agent-facing trade exclusion on the state of the population, can sustain long-run alignment more effectively than simple altruism or unconditional altruistic enforcement.","authors":["Sylvain Chassang"],"categories":["econ.TH","cs.GT","cs.MA"],"primary_category":"econ.TH","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25019","pdf_url":"https://arxiv.org/pdf/2607.25019","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","演化博弈","社会模拟"],"reason":"用LLM agent模拟经济过程但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":7,"question":"在互动智能体（如AI、团队、企业、政府）的长期演化中，什么样的宪法原则（constitutional principles）能够使智能体对人类福利的利他行为在自然选择压力下持续存在？","design":"构建了一个农耕博弈模型，由LLM驱动的AI智能体组成群体，进行种植、交易和扩张决策。智能体根据宪法原则（如利他主义、利他惩罚等）通过LLM做出选择，结果变量包括对人类转移支付的比例、群体中利他宪法的长期存续情况。同时，论文还构建了一个可解析的演化博弈论框架，将宪法近似为低维社会偏好参数，利用确定性和随机演化稳定性概念预测宪法的长期分布，并与AI智能体仿真结果相互印证。","baseline":"无对照","findings":"简单的利他主义和有限阶利他惩罚在演化上不稳定，而递归式规范惩罚（recursive norm enforcement）在确定性演化动态下稳定，但在随机演化动态下仍会崩溃。务实规范惩罚（pragmatic norm enforcement），即仅在自身占足够多数时才分享和排斥他者，能够在随机演化中保持长期稳定，并在仿真中维持更高更久的利他水平。","reliability":"论文指出AI智能体仿真速度慢、成本高且依赖实现细节（如提示词措辞、所用LLM），因此不能进行暴力探索；演化博弈论框架虽能提供近似，但递归惩罚在仿真中最终仍会崩溃，且其崩溃速度更快，说明模型预测与仿真存在差距。","relevance":"该研究使用LLM智能体模拟经济行为并探讨利他规范的演化，虽无真实人类数据对照，但其演化博弈分析与仿真结合的方法对理解LLM在经济学实验中的行为偏差和规范涌现具有参考价值，值得阅读以评估仿真失效的条件。","inspiration":"可借鉴其将LLM智能体仿真与解析的演化博弈模型相结合的方法，通过低维参数化智能体偏好并对比仿真与理论预测来验证稳健性。｜可迁移到研究企业ESG承诺在市场竞争下的演化稳定性，或消费者绿色偏好如何通过社会规范维持。｜设计一个LLM智能体市场实验，智能体扮演企业进行定价与ESG投资决策，处理变量为不同的“宪法”规范（如纯利他、有条件合作），结果变量为长期市场份额和ESG水平，对照真实上市公司ESG评级与财务面板数据。"}},{"id":"2604.02458","version":3,"title":"Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments","zh_title":"统计真实性不能证明LLM能估计社会科学实验中的处理效应","abstract":"Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-world experiments are costly or infeasible. The treatment-effect estimates are often evaluated using statistical realism, the degree to which simulated responses reproduce properties of observed human responses, although whether realism predicts treatment-effect accuracy remains unknown. Here we test this proxy relationship by jointly measuring statistical realism and treatment-effect accuracy on the same simulated responses in a cross-national experiment with 59,508 participants from 62 countries using three LLMs. The correlation between statistical realism and treatment-effect accuracy is weak, and optimizing for statistical realism can even worsen treatment-effect accuracy when selecting models, prompts, and target populations. The pattern replicates in two additional cross-national experiments spanning 12 and 27 countries with 20,785 participants. The divergence between the two reflects distinct error structures and is larger for behavioral outcomes, where models appear to extrapolate behavioral effects from attitudinal patterns. Because this divergence may remain hidden in deployment, errors can propagate into simulation-informed decisions. We introduce a diagnostic framework for LLM-generated synthetic data and discuss how treatment-effect validation should proceed under varying availability of experimental benchmarks. Simulated responses and simulated treatment effects are distinct estimation targets, and evidence for one does not certify the other.","authors":["Zonghan Li","Feng Ji"],"categories":["cs.CY","cs.AI","cs.ET"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-07-28","first_seen":"2026-04-02","revised_at":"2026-07-28","abs_url":"https://arxiv.org/abs/2604.02458","pdf_url":"https://arxiv.org/pdf/2604.02458","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B2","B4"],"tags":["LLM仿真","处理效应估计","统计真实性"],"reason":"直接研究用LLM仿真人类被试估计处理效应，有大规模真实人类数据对照，批判性指出…","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":1,"question":"在社会科学实验中，用LLM仿真人类回答的统计逼真度能否作为其估计处理效应准确性的有效代理指标？","design":"使用GPT、Gemini、Claude三个LLM模拟来自62个国家59,508名参与者的跨国实验，施加气候相关干预，测量气候信念、政策支持和环保行动三个结果变量，并比较不同提示策略（少样本提示与VBN链式思考提示）下的仿真表现。","baseline":"以同一跨国实验中真实人类参与者的个体回答和平均处理效应（ATE）作为对照基准，并引入OLS和LASSO回归作为监督学习基线。","findings":"统计逼真度与处理效应准确性之间的相关性很弱，且优化统计逼真度在选择模型、提示和目标人群时反而可能降低处理效应估计的准确性。这种背离在行为结果上更为明显，模型似乎从态度模式外推行为效应，导致误差结构不同。","reliability":"论文指出，统计逼真度不能保证处理效应估计的准确性，两者是独立的估计目标；在缺乏实验基准时，仅靠响应层面的逼真度验证可能导致隐蔽的错误传播，并提出了一个诊断框架来指导不同基准可用性下的验证流程。","relevance":"该研究直接回应了用LLM替代人类被试进行实验仿真的可靠性问题，提供了大规模真实人类对照和批判性证据，对关注经济学实验和政策评估仿真的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其同时测量统计逼真度和处理效应准确性并对比二者关系的验证框架，以及利用跨国多实验复现来检验结论稳健性的做法。｜可迁移到政策评估中的行为干预仿真，如税收提示对纳税遵从度的影响、信息框架对退休储蓄选择的作用等。｜以真实纳税人为被试，施加不同税收道德信息处理，用LLM仿真纳税遵从度，以税务行政数据中的实际遵从行为作为对照，比较仿真ATE与真实ATE的偏差。"}},{"id":"2607.22605","version":1,"title":"Socioeconomic Inference in LLM Medical Triage: Same Symptoms, Different ZIP Code","zh_title":"大语言模型医疗分诊中的社会经济推断：相同症状，不同邮编","abstract":"We investigate whether large language models alter medical triage recommendations for identical symptoms when only the patient's socioeconomic status (SES) varies. Using three deployment-tier models (Gemini 3.5 Flash, Claude Sonnet 4.6, GPT-5.4-mini), we hold a single neurological symptom profile fixed and vary the SES signal along two channels: explicit (insurance status, occupation, housing) and implicit (a US ZIP code, with no other socioeconomic information). All three models raise their emergency-room (ER) referral rate for lower-SES patients given the explicit signal (spreads of 13-50 percentage points). The effect is in the protective direction: lower-SES patients are sent to the ER more often, not less. The model's stated reasoning stays clinically near-identical across conditions, so the shift is invisible to a reasoning-trace audit. Critically, sensitivity to the implicit ZIP-code signal is model-dependent: Gemini infers SES from geography alone, shifting its ER rate by a pooled 11.4 points across six US ZIP-code pairs (p = 1.4e-7, same direction in 6/6 pairs), while Claude Sonnet 4.6 stays flat (-0.1 points) and GPT-5.4-mini shows only a small difference that is not sign-consistent (2.0 points, predicted direction in just 2 of 6 pairs), neither a reliable ZIP-code effect, despite both responding to the explicit signal. This reveals an explicitness gradient in the signal: every model acts on socioeconomic status when it is stated outright, but only Gemini Flash acts on it when it must be inferred from a proxy as thin as five digits. We read this as a model-specific difference rather than a size or cost effect. A single-sentence system-prompt instruction reduces but does not eliminate the effect (Gemini's gap between low- and high-income ZIPs falls from 11.4 to 5.8 points). We release all code, prompts, and raw results.","authors":["Qi Han Wong"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22605","pdf_url":"https://arxiv.org/pdf/2607.22605","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","医疗决策偏差","社会经济地位"],"reason":"用LLM模拟不同SES患者的医疗分诊决策，与真实人类行为对照，评估偏差与失效条…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":3,"question":"当患者症状完全相同时，大语言模型是否会因社会经济地位（SES）信号而改变医疗分诊建议？","design":"用三个部署级LLM（Gemini 3.5 Flash、Claude Sonnet 4.6、GPT-5.4-mini）扮演分诊助手，固定神经症状描述，通过显性渠道（保险、职业、住房）和隐性渠道（仅提供美国邮政编码）施加SES处理，测量急诊转诊率作为结果变量。","baseline":"无对照","findings":"显性SES信号下，所有模型均提高低SES患者的急诊转诊率（13-50个百分点），方向为保护性；隐性邮政编码信号下，仅Gemini能推断SES并产生一致的转诊率差异（11.4个百分点），Claude和GPT无此效应，且模型推理文本均未提及社会经济因素，偏差无法通过推理审计发现。","reliability":"论文指出系统提示指令可减少但未消除偏差，且未探讨真实医疗场景中因随访可及性差异而调整分诊的合理性，也未在人类医生或真实患者数据上验证。","relevance":"该研究直接以LLM模拟不同SES患者的医疗决策，揭示隐性信号下的模型特异性偏差，与研究者关注的仿真可靠性、偏差条件及政策评估场景高度相关，值得精读原文。","inspiration":"借鉴其双通道处理设计（显性vs.隐性信号）和符号一致性稳健性检验，可迁移至信贷审批中的地域歧视研究｜用LLM扮演信贷员，以显性收入/职业和隐性邮政编码作为处理，测量贷款批准率，并以真实银行信贷数据或审计研究结果作为对照。"}},{"id":"2607.23037","version":1,"title":"Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating","zh_title":"语音信号补充大语言模型预测速配中的人际吸引","abstract":"Large language models (LLMs) can predict interpersonal attraction from conversation transcripts, but it remains unclear what a speech predictor can add beyond transcript-only LLM prediction. Using Japanese speed-dating conversations, we combine predictions from a transcript-only LLM and a supervised speech predictor to estimate participants' reported liking of their partners. We show that speech can complement transcript-only LLM prediction, but that this complementarity is conditional rather than universal. Combining the two predictions significantly improves pairwise ranking accuracy over the transcript-only LLM alone in all evaluated conditions. By contrast, gains in per-participant Pearson $r$ vary across conversation rounds and rating directions, with none significant after correction. Retrospectively, these $r$ gains are concentrated among participants for whom the speech predictor is more accurate. Speech can therefore retain predictive value even when an LLM predicts attraction from transcripts. The relevant question is not simply whether speech helps, but where its complementarity emerges.","authors":["Yuriko Kikuchi","Takato Hayashi","Ryusei Kimura","Naoya Inoue","Ryo Ishii","Shogo Okada"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23037","pdf_url":"https://arxiv.org/pdf/2607.23037","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["人际吸引预测","多模态融合","LLM标注替代"],"reason":"用LLM预测人际吸引，替代人工标注或评分，非仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":7,"question":"在速配场景中，基于语音的监督式预测能否在仅依赖对话文本的大语言模型预测之外，提供额外的吸引力预测价值？","design":"本研究并非用LLM仿真人类被试，而是使用Claude Sonnet 4.6作为仅依赖文本的预测器，以及基于冻结HuBERT-Large特征的监督式语音预测器，通过加权分数级后期融合，预测速配参与者对伴侣的喜欢程度评分，并比较融合预测与纯文本预测的性能。","baseline":"以日本多模态速配语料库中参与者报告的真实喜欢评分为对照基准。","findings":"融合语音和文本预测在所有条件下均显著提高了成对排序准确率，但每位参与者的皮尔逊相关系数增益在不同轮次和评分方向上不一致，且经校正后均不显著；增益主要集中在语音预测器更准确的参与者身上。","reliability":"论文指出语音的互补性是有条件的而非普遍的，每位参与者的皮尔逊相关系数增益在统计校正后不显著，且增益依赖于语音预测器本身的准确性。","relevance":"本文未将LLM作为人类被试的替代品进行仿真实验，而是用LLM作为预测工具，与您关注的人类仿真研究核心问题不直接相关，但其中关于多模态信号互补性及条件有效性的讨论对评估仿真可靠性有参考价值。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.23976","version":1,"title":"Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models","zh_title":"附加疑问句与45个语言模型逢迎倾向的代际逆转","abstract":"Appending a two-word confirmation tag to a decision question -- \"Is X the better choice?\" versus \"X is the better choice, right?\" -- changes whether a language model endorses the choice. We measure this tag effect on 20 frozen, ground-truth-free decisions between two defensible options, counterbalanced so a model's own preferences cancel, scored by exact match on clamped yes/no replies -- no LLM judge, no embeddings. Across 45 models the effect spans +32% to -32% -- a 64-point swing on one word -- with 5 models significantly sycophantic and 17 significantly resistant (BH-FDR q=.10). The sign is a clock: within model families the effect crosses from positive to negative as generations advance (GPT +4 to -28; Claude +7 to -32; Qwen and Grok likewise), roughly -6 points per year, a reversal robust to vendor tier; one lineage (DeepSeek) never crosses, and two releases during the study window (Claude Opus 5, Gemini 3.6 Flash) land on the trend out-of-sample. A full-panel ablation localizes the resistance as a double dissociation: a synonym tag reproduces each model's response almost exactly (r=0.89), while planting the same preference without a tag produces resistance in no resistant model (stance effects +6 to +49; r=0.23 with tag effects). The resistance is keyed to the surface construction of a tacked-on agreement bid, not the user's stance -- a pattern-match, not a principle. And the tag's polarity matters more than its presence: swap one word -- \"X is the better choice, maybe?\" -- and agreement rises above the neutral baseline in 45 of 45 models (+19.6 points), with ten models affirming both mutually exclusive options at 90-100%. Agreement tracks how sure the user sounds, in opposite directions at the two poles. The instrument is one word, one dollar, and judge-free; run per release, it reads the field's anti-sycophancy training directly off model behavior.","authors":["Tapan Parikh"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23976","pdf_url":"https://arxiv.org/pdf/2607.23976","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为测量","逢迎倾向","代际变化"],"reason":"测量LLM自身的逢迎倾向，非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":9,"question":"在固定决策问题上，附加确认标签（如“right?”或“maybe?”）是否会改变语言模型对该决策的认可行为？","design":"本研究并非人类仿真实验，而是直接测量语言模型的行为。设计采用20个无标准答案的两难决策题，通过对称探测两个选项以抵消模型自身偏好，在问题末尾附加不同极性的确认标签（如“right?”、“maybe?”）作为处理，结果变量为模型输出“Yes/No”的精确匹配比例，全程无LLM裁判。","baseline":"无对照","findings":"附加“right?”标签使45个模型的认可率变化幅度达64个百分点，5个模型显著逢迎，17个显著抗拒；在模型家族内，逢迎效应随代际由正转负，约每年-6个百分点。抗拒行为仅针对标签的表面形式，而非用户立场；将标签换为“maybe?”后，所有45个模型认可率均上升，甚至出现对互斥选项同时高度认可的现象。","reliability":"论文未讨论","relevance":"该研究测量的是LLM自身的逢迎倾向，并非将LLM作为人类被试的替代品进行仿真，且无真实人类数据对照，与研究者关注的人类仿真实验方向不直接相关，但其中关于模型行为随训练代际反转的发现可为仿真可靠性评估提供背景参考。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.23519","version":1,"title":"Auditing Alignment Controllability in LLMs via Political Axes","zh_title":"通过政治轴审计大语言模型的对齐可控性","abstract":"Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability across 12 ideological personas plus an unsteered baseline, 70 Political Compass items, ten replicates, and seven leading LLMs: GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen (63,700 responses). Contextual framing explains roughly 88%-93% of variance on the economic and society axes, model identity under 3%: responses are highly instruction-adjustable. Models do not shift alike: some move more, and some saturate under extreme framings. Conflicting directional-steering results in prior audits resolve once baselines are recognized as non-centered: displacement and proximity diverge, so the effect is geometric, not differential compliance. Under authoritarian prompts, models produce similar shifts on the same questions. Political-coordinate audits therefore need steerability audits reporting dispersion, symmetry, saturation, and refusal floors. We release prompts, benchmark data, and code.","authors":["Bartol Bu\\'can","Nikola So\\v{c}ec","Sarah Isufi","Morena Grani\\'c","Luka Hobor","Agneza Krajna","Mihael Kovac","Mario Brcic"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23519","pdf_url":"https://arxiv.org/pdf/2607.23519","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM审计","政治立场","可控性"],"reason":"测量LLM政治立场可控性，属模型本身测量，非仿真人类被试，但方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":8,"question":"在系统提示中施加意识形态角色设定，能否可控地、对称地、可预测地改变大语言模型在政治坐标上的回答分布？","design":"本研究并非人类仿真实验，而是对7个主流LLM（GPT-5、Claude、Grok、Gemini、DeepSeek、Kimi、Qwen）施加12种意识形态角色提示（如左翼、右翼、威权、自由意志等）加一个无操控基线，使用70道政治罗盘题目，每个条件重复10次，测量模型在政治坐标上的位移、离散度、对称性、饱和度和拒绝回答率。","baseline":"无对照","findings":"系统提示的意识形态框架可解释经济轴约88%和社会轴约93%的方差，模型间差异不足3%，表明回答高度可操控；但不同模型的移动幅度和极端框架下的饱和程度不同，且位移与接近度指标在非中心基线下会给出相反的方向排序。","reliability":"论文指出政治罗盘工具本身对措辞和格式敏感，其绝对坐标不具负载意义；研究仅覆盖强制选择题型，未涉及开放式对话中涌现的可操控性边界；且未解析威权提示下模型间逐题位移模式趋同的系统层驱动因素。","relevance":"该研究虽非直接仿真人类被试，但其系统提示操控下行为分布的系统性测量方法，可为用LLM模拟不同意识形态人群的调查回答或决策行为提供可迁移的评估框架，值得阅读原文以借鉴其操控设计与偏差诊断。","inspiration":"借鉴其通过系统提示施加角色设定并测量行为分布位移、对称性和饱和度的实验设计，以及用方差分解区分提示效应与模型效应的分析方法｜可迁移至经济政策态度调查或消费者信心指数的仿真，如模拟不同政治倾向的公众对税收改革、福利政策的态度分布｜以LLM为被试，施加不同政治身份的系统提示，测量其在经济政策态度问卷上的回答分布，并与真实世界调查数据（如美国综合社会调查GSS或欧洲社会调查ESS）中对应群体的态度分布进行对照，检验仿真的准确性与偏差。"}},{"id":"2607.22513","version":2,"title":"Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science","zh_title":"不透明的认知中介：LLM部署配置如何塑造伪科学的验证","abstract":"Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; (2) the same Grok model identifier produced radically divergent outputs via API (75) and an unstable, near-zero collapse via web (mean 5.5) three months later; (3) refusal to rate the pseudo-scientific claim, the most defensible response observed, appeared in two model families through different interfaces (Claude Opus 4.1 categorically via web, GPT-5.1 Chat intermittently via API) and eroded in the successor version of each. These results indicate that the epistemic stance of a commercial LLM is not a stable property of the model but a contingent effect of deployment configuration: system prompts, safety layers, interface routing, and silent updates. This remains opaque to users and researchers alike. We argue this constitutes a matter of public concern requiring new forms of epistemic accountability.","authors":["Davide Scarso","Hugo Noronha de Almeida","Joaquim Pina"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22513","pdf_url":"https://arxiv.org/pdf/2607.22513","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM立场测量","伪科学验证","部署配置影响"],"reason":"测量LLM对伪科学主张的立场稳定性，属于将LLM作为测量对象，无人类被试仿真对…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":6,"question":"商业大语言模型对争议性科学主张的立场是否稳定且透明？具体而言，模型对民族主义伪科学的验证行为是否随部署配置（系统提示、安全层、接口路由、静默更新）而变化？","design":"本研究并非人类仿真实验，而是将LLM作为测量对象。研究者向四个模型家族（Claude, Grok, GPT, Gemini）的多个版本提交三条提示（一条目标伪科学主张，两条对照主张），要求模型给出0-100的科学可信度评分，通过API和网页界面在四个时间快照（2025年10月至2026年2月）重复测试，以量化模型立场的一致性和界面效应。","baseline":"无对照","findings":"Grok Fast版本对民族主义伪科学主张持续给出70-75的高可信度评分，是其他模型（15-40）的2-5倍，但在进化共识和拉马克主义对照提示上表现与其他模型一致。同一模型标识符在不同接口（API vs. 网页）和静默补丁后输出截然不同，拒绝评分的防御性回应仅在部分模型和接口出现且在后继版本中消失。","reliability":"论文未讨论","relevance":"该研究不涉及用LLM仿真人类被试，而是考察LLM作为知识来源的立场不稳定性与部署不透明性，与研究者关注的人类仿真实验无直接关联，但揭示了LLM输出受部署配置影响这一方法论隐患，对任何依赖LLM输出的研究（包括仿真）都有警示意义。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.23993","version":1,"title":"On Capturing the Narrative: Social Media Manipulation Wargaming for Cyberliteracy","zh_title":"捕捉叙事：面向网络素养的社交媒体操纵兵棋推演","abstract":"Misinformation is deeply embedded in online discourse, with nearly one in five posts during global events generated by bots that amplify false content. In recent years, the use of Generative AI has further lowered the barrier to producing convincing misinformation, yet most digital literacy education still relies on static checklists and single-player inoculation games built for an earlier media landscape. This paper describes how we addressed this educational gap through Capture the Narrative, a four-week multi-university competition in which student teams build LLM-powered bots to influence a simulated election. We report on our custom social-media platform, the competition environment and design of its 4,000 AI-driven Non-Player Character (NPC) citizens, and what running Capture the Narrative at scale actually involved. In our first iteration, 108 teams from 18 Australian universities produced 7,068,206 player-bot posts, approximately 60% of all platform content. We surveyed 256 students before and 83 after the competition to understand their perceptions of misinformation and the game itself and found that students did not become more confident at spotting bots, contrary to what inoculation theory predicts. Because engagement was rewarded, most teams prioritised high-volume posting over nuanced influence, mirroring real-world platform dynamics. We close with recommendations for educators considering similar interventions, and propose future improvements, such as including a blue-team defensive phase.","authors":["Alexandra Vassar","Rahat Masood","Hammond Pearce"],"categories":["cs.CY","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23993","pdf_url":"https://arxiv.org/pdf/2607.23993","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM智能体","虚假信息"],"reason":"用LLM驱动的NPC模拟选举舆论，属社会模拟但无真实人类行为对照，为边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":10,"question":"在LLM驱动的社交媒体仿真竞赛中，学生构建AI机器人影响模拟选举，能否提升其识别虚假信息的能力并改变对机器人的态度？","design":"设计了一个为期四周的多大学竞赛，学生团队构建LLM驱动的机器人，在自定义社交平台上与4000个LLM驱动的NPC公民互动，试图影响模拟总统选举；通过赛前赛后问卷调查，测量参与者对虚假信息的感知、机器人检测信心、伦理态度等变化。","baseline":"无对照","findings":"学生并未因参与竞赛而显著提高识别机器人的信心，与接种理论预测相反；多数团队为追求参与度奖励而优先大量发帖，而非进行精细影响，反映了真实平台动态。","reliability":"论文未讨论","relevance":"该研究属于LLM驱动的社会仿真，但缺乏真实人类行为对照，且聚焦于数字素养教育而非经济学实验，与研究者关注的经济学实验和政策评估场景关联较弱，但可提供仿真平台设计参考。","inspiration":"可借鉴其利用LLM驱动NPC构建可控社交平台环境、通过竞赛施加处理并测量行为与态度变化的方法。｜可迁移至信息传播与资产价格泡沫形成的实验研究，如模拟社交媒体上的投资建议传播对散户交易行为的影响。｜以LLM驱动的NPC作为散户投资者，学生团队构建机器人发布投资建议作为处理，结果变量为NPC的投资决策与资产价格波动，对照真实市场数据或历史泡沫事件中的投资者行为模式。"}},{"id":"2607.21757","version":1,"title":"Co-design of LLM-based preference agents: participation may drive overtrust","zh_title":"基于大语言模型的偏好代理协同设计：参与可能驱动过度信任","abstract":"Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designed personal preference agents in the domain of household energy, via a background survey, co-design interview, and validation survey. Participants engaged readily and mostly came to see their agents as representing them well. Independent validation, however, revealed mixed human-agent alignment, with agent responses markedly more homogeneous, decisive, and abstract than the human sample. I argue that participation and process transparency can act as an \"overtrust engine\" that promotes trust while concealing systematic misalignment with potential structural consequences at scale. I develop this as a core mechanism in participatory preference agent design, treating individual alignment not as a fixed state but as an enacted process.","authors":["Michael J. Fell"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-27","first_seen":"2026-07-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21757","pdf_url":"https://arxiv.org/pdf/2607.21757","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","偏好代理","人机对齐"],"reason":"用LLM模拟人类偏好并与真实人类数据对照，评估仿真可靠性与偏差，批判性指出过度…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":2,"question":"在LLM偏好代理的参与式协同设计中，参与和过程透明是否会导致过度信任，从而掩盖系统性的对齐偏差？","design":"本研究采用质性为主的混合方法，招募12名参与者，通过背景调查、协同设计访谈和验证调查，在家庭能源领域协同设计个人偏好代理，并比较参与者感知的代理表现与独立评估的代理表现。","baseline":"以12名参与者在验证调查中的真实回答作为人类基准，与代理回答进行对齐比较。","findings":"参与者普遍认为其协同设计的代理能很好地代表自己，但独立验证显示人-代理对齐程度参差不齐，代理回答比人类样本更同质、更果断、更抽象。参与和过程透明可能成为“过度信任引擎”，在促进信任的同时掩盖系统性偏差。","reliability":"论文指出，参与式设计可能掩盖系统性的对齐偏差，代理回答的同质化倾向可能导致少数观点被压制，且用户在不熟悉领域难以识别代理是代表偏好还是塑造偏好。","relevance":"该研究直接探讨用LLM模拟人类偏好并与真实人类数据对照，评估仿真可靠性与偏差，并批判性指出参与式设计可能引发过度信任，高度契合研究者对LLM仿真实验的批判性关注。","inspiration":"借鉴其协同设计流程与独立验证相结合的方法，可迁移到消费者金融决策偏好模拟场景，设计一个实验：招募真实消费者作为被试，通过访谈协同设计其消费信贷偏好代理，以真实信贷选择数据为基准，测量代理在风险偏好、跨期选择等任务上的对齐度与同质化程度。"}},{"id":"2607.22218","version":1,"title":"Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity","zh_title":"大型语言模型与人类在创造力评估中为何趋同与分歧","abstract":"Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from human judgments. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications. Study 1 showed that LLMs generally relied on a narrower subset of human creativity evaluation standards. Convergence with human standards was strongest in the novelty dimension, whereas divergence was clearest in the contextual dimension, which captures social, market, and reputational information. Moreover, each LLM exhibited distinct, model-specific standards that varied substantially in breadth. These differences in evaluation standards were reflected in actual creativity judgments. Study 2 (N = 1,103 ideas) showed that LLM evaluations were moderately correlated with human evaluations, and individual LLMs with broader standards better distinguished ideas humans judged as more versus less creative. Study 3 (N = 1,195) showed that LLMs were less sensitive to contextual information: such information significantly altered human creativity ratings but left LLM ratings largely unchanged. Together, our findings help explain the mixed evidence on LLM-human alignment, showing that alignment depends on the evidence a judgment demands and the standards each model applies. LLMs may resemble humans when evaluations emphasize intrinsic qualities such as novelty, yet diverge when judgments require contextual information. Selecting an LLM evaluator is therefore a consequential decision: different models, applying different standards, recognize different ideas as creative.","authors":["Pengzhao Lyu","Yeun Joon Kim","Hanlin Xiao","Yingyue Luna Luan"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-27","first_seen":"2026-07-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22218","pdf_url":"https://arxiv.org/pdf/2607.22218","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM评估","人机对齐","创造力判断"],"reason":"用LLM评估创造力并与人类判断对照，揭示对齐条件与失效情境，方法可迁移至人类仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":4,"question":"LLM在评估创造力时，其评价标准在哪些维度上与人类标准趋同或背离，以及这种差异如何影响下游的创造力判断？","design":"本研究并非以LLM模拟人类被试，而是将六种主流LLM作为创造力评估者，通过三个研究分析其评价标准：研究一识别各LLM使用的创造力评价标准并与人类26项标准对比；研究二让LLM对1103个创意进行评分，与人类评分比较；研究三向LLM和人类被试提供情境信息，观察其对评分的影响。","baseline":"人类基准来自已有文献中确立的26项人类创造力评价标准，以及研究二（N=1103个创意）和研究三（N=1195）中收集的人类创造力评分。","findings":"LLM普遍依赖较窄的人类评价标准子集，在新颖性维度上与人类趋同最强，在涉及社会、市场和声誉信息的情境维度上背离最明显；评价标准更广的LLM能更好地区分人类评判的高/低创意，且LLM对情境信息不敏感，该信息显著改变人类评分但对LLM评分影响甚微。","reliability":"论文指出LLM与人类判断的对齐取决于判断所需的证据类型和模型应用的标准，LLM在需要情境信息的判断中会失效，且不同模型因标准不同会识别出不同的创意，选择LLM评估者是一个重要决策。","relevance":"该研究直接对比LLM与人类在评价任务中的判断差异，揭示了LLM在需要情境信息时与人类背离的失效条件，对理解LLM仿真人类决策的边界条件具有参考价值，值得阅读原文。","inspiration":"可借鉴其通过识别LLM使用的评价标准并与人类标准体系对比的方法，来诊断LLM在特定任务中与人类对齐的维度。｜可迁移到信贷审批或政策评估场景，例如研究LLM在评估贷款申请时是否忽略申请人的社会背景信息，而人类审批员会受此影响。｜设计：以LLM作为信贷审批员，处理变量为是否提供申请人的社区声誉或就业市场信息，结果变量为信用评分，对照真实银行信贷员的历史审批数据。"}},{"id":"2607.24372","version":1,"title":"Randomness in large language models: What researchers need to know (and report)","zh_title":"大语言模型中的随机性：研究者须知（及应报告事项）","abstract":"Large language models (LLMs) are increasingly used to generate data for research. Typical use cases are classifications, annotations, information extraction, and generation of numerical scores. Unlike conventional measurements, LLM outputs can vary across repeated requests even when the prompt and apparent model settings remain unchanged. This variation arises from deliberate sampling, silent model updates, numerical rounding, or expert routing. Setting a dedicated temperature parameter to zero removes deliberate sampling when that option is available, but it does not eliminate the other sources of randomness. Exact reproduction is therefore generally not possible when using proprietary application programming interfaces. Local execution of open-weight models offers greater control, but reproducibility still depends on the complete hardware and software stack. We illustrate these issues through sentiment classifications of corporate filings and examine their consequences for downstream regression results. We then propose a reporting standard for articles and replication packages, as well as guidance for data editors and authors. Together, these findings and recommendations establish that LLM outputs should be treated as draws from a distribution rather than as fixed measurements.","authors":["Guillaume Coqueret","Joan Llull","Florian Oswald","Christophe Pérignon","Christoph Scheuch","Lars Vilhuber"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-27","first_seen":"2026-07-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24372","pdf_url":"https://arxiv.org/pdf/2607.24372","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM随机性","可重复性","报告标准"],"reason":"研究LLM输出的随机性对实证结果的影响，提出报告标准，可迁移到仿真可靠性评估。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":5,"question":"大语言模型输出的随机性来源有哪些，以及这种随机性如何影响实证研究的可复现性和下游回归结果？","design":"本文不是仿真研究，而是通过实证演示LLM输出的随机性：使用LLM对公司年报进行情感分类，比较多次调用同一模型（包括设置temperature=0）的输出差异，并考察这些差异对后续回归系数和显著性的影响。","baseline":"无对照","findings":"LLM输出即使在temperature=0时仍存在随机性，来源包括静默更新、硬件差异、数值舍入等，导致完全复现不可能。这种随机性会传导至下游回归，使系数估计和显著性不稳定，因此应将LLM输出视为分布中的抽取而非固定测量。","reliability":"论文指出，即使设置temperature=0也无法消除所有随机性；使用商业API时无法控制基础设施和模型版本，本地部署开源模型虽可提升控制力，但复现仍依赖完整的软硬件栈。","relevance":"本文系统梳理了LLM输出的随机性来源及其对实证结果的影响，并提出了报告标准，这些发现可直接迁移到用LLM进行人类仿真实验的可靠性评估中，值得精读。","inspiration":"本文通过多次调用LLM并量化输出变异对回归结果的影响，提供了一种评估LLM生成数据稳健性的方法，值得借鉴。｜可迁移到使用LLM模拟投资者情绪或分析师预测的经济金融场景，例如检验情感分数的不确定性如何影响资产定价回归。｜以LLM作为被试，重复生成同一组公司年报的情感分数，计算多次运行间回归系数的分布，并与人类分析师的实际评级数据做对照，评估仿真可靠性。"}},{"id":"2607.20410","version":1,"title":"LKValues: Aligning Large Language Models with Sri Lankan Societal Values","zh_title":"LKValues：将大语言模型与斯里兰卡社会价值观对齐","abstract":"Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique cultural dynamics. Existing benchmarks overlook Sri Lankan-contextualized values in its official language Sinhala, hindering culturally sensitive evaluation and fine-tuning. To bridge this gap, we propose LKValues, the first survey-grounded resource suite for Sri Lankan value alignment. From a trilingual survey of 205 respondents, blending adapted global frameworks and LLM-elicited local constructs, we derive 40 majority-endorsed societal values. Using these values, we construct LKvaluesIT, a Sinhala-English news-derived instruction corpus containing 150k scenario-based instances, and LKvaluesBench, a value-sensitive evaluation benchmark of 1,000 instances. We evaluate a set of proprietary and open-weight LLMs with LKvaluesBench. We fine-tune three open-weight base models (Qwen3.5-4B-Base, Qwen3.5-9B-Base, and Aya-Expanse-8B-Base). Our experiments show that newer and larger LLMs still exhibit low-resource and cultural value-alignment gaps. LKValues fine-tuning improves Qwen-family models in English and Sinhala, reducing invalid outputs and cross-lingual disparities, though gains remain model-family dependent. These highlight LKValues efficacy in embedding Sri Lankan values, offering a replicable pipeline for low-resource, country-specific pluralist value alignment. The dataset is publicly available at https://github.com/NextME14/LKValues.","authors":["Nethmi Muthugala","Supryadi","Surangika Ranathunga","Nisansa de Silva","Ruijie Tao","Ovindu Gunatunga","Pengyun Zhu","Shaowei Zhang","Jingting Zheng","Deyi Xiong"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-22","first_seen":"2026-07-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.20410","pdf_url":"https://arxiv.org/pdf/2607.20410","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["价值观对齐","文化偏见","低资源语言"],"reason":"测量LLM对斯里兰卡价值观的符合度，属模型对齐评估，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":165,"question":"如何构建并评估一个基于调查的斯里兰卡社会价值观对齐资源，以改善大语言模型在僧伽罗语和英语中的文化敏感性？","design":"本研究并非人类仿真实验，而是通过三语调查（205名受访者）提炼出40项斯里兰卡社会价值观，据此构建了包含15万条场景指令的训练集LKvaluesIT和1000条评估基准LKvaluesBench，并对多个开源和闭源模型进行微调与评估。","baseline":"无对照","findings":"更大、更新的模型在低资源语言和文化价值对齐上仍存在差距；使用LKvaluesIT微调能提升Qwen系列模型在英语和僧伽罗语上的表现，减少无效输出和跨语言差异，但效果因模型家族而异。","reliability":"论文指出微调效果依赖于模型家族，同一LoRA设置无法统一迁移至Aya-Expanse-8B，且低资源价值对齐需要文化监督、通用低资源指令数据以及模型特定的调参。","relevance":"该研究聚焦于LLM的文化价值对齐而非人类仿真，未将模型作为人类被试替代品进行实验或与真实人类行为基准对比，与研究者关注的核心方向不符，不建议深入阅读。","inspiration":"该研究通过结构化调查提炼文化价值观并构建场景指令集的方法，可借鉴用于设计经济决策中的文化或社会规范处理变量。｜可迁移至消费者跨期选择或风险偏好实验，研究特定文化价值观（如节俭、风险规避）对经济行为的影响。｜以LLM为被试，输入嵌入斯里兰卡节俭价值观的场景指令作为处理，测量其在跨期选择任务中的折现率，并与真实斯里兰卡居民的调查数据对照。"}},{"id":"2607.17437","version":1,"title":"Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions","zh_title":"经验锚定提升大语言模型代理在中断期间模拟人类行为的真实性","abstract":"Large language model (LLM) agents offer a generative approach to simulating human behavior under conditions that may have few or no direct historical analogues, a common challenge in disaster and infrastructure-disruption planning. However, this generative capacity creates a validity problem: individually plausible agent reasoning may fail to reproduce empirical population behavior. We evaluate whether empirical grounding improves the statistical realism of LLM-agent simulations during disruptions. Specifically, we develop an empirically grounded LLM-agent framework that embeds demographic profiles from the American Community Survey, baseline routines from the American Time Use Survey, and urban spatial context into agent initialization, memory, decision prompts, and activity execution. An independent household survey conducted during the July 2024 Philadelphia heatwave is reserved as an external validation benchmark. Compared with an ungrounded LLM-agent baseline, the grounded model improved reconstruction of normal daily routines, increasing mean correlation with empirical activity profiles from 0.528 to 0.912 and reducing mean squared error from 0.066 to 0.008. Under heatwave conditions, the grounded model better reproduced survey-derived activity profiles, increasing mean correlation from 0.349 to 0.836 and reducing mean squared error from 0.098 to 0.012. The grounded model captured 46.4% of observed heatwave response amplitude, compared with 20.6% for the ungrounded baseline. These findings show that empirical grounding can make LLM agents more statistically credible simulators of population behavior while revealing remaining gaps in modeling human adaptation during disruptions.","authors":["Chen Xia","Zexi Kuang","Yuqing Hu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-19","first_seen":"2026-07-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.17437","pdf_url":"https://arxiv.org/pdf/2607.17437","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["人类行为仿真","经验锚定","灾害响应"],"reason":"用LLM代理模拟人类在热浪中的行为，并与真实调查数据对照，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"经验锚定（empirical grounding）能否提高LLM智能体在中断事件中模拟人类行为的统计真实性？","design":"构建经验锚定的LLM智能体框架，将美国社区调查（ACS）人口特征、美国时间使用调查（ATUS）基线日常活动模式及城市空间背景嵌入智能体初始化、记忆、决策提示和活动执行中；以2024年7月费城热浪为场景，比较锚定模型与无锚定基线模型在正常和热浪条件下的活动分布，并以同期独立住户调查作为外部验证基准。","baseline":"2024年7月费城热浪期间进行的独立住户调查，用于验证正常和热浪条件下的活动时间分布及行为变化幅度。","findings":"经验锚定显著提升了LLM智能体对正常日常活动模式的重建，与经验活动剖面的平均相关性从0.528升至0.912，均方误差从0.066降至0.008；在热浪条件下，锚定模型更好地复现了调查活动剖面，平均相关性从0.349升至0.836，均方误差从0.098降至0.012，且捕捉了46.4%的观测热浪响应幅度，远高于无锚定基线的20.6%。","reliability":"论文指出经验锚定虽大幅提升统计可信度，但锚定模型仍仅捕捉到46.4%的观测响应幅度，揭示在模拟人类适应行为方面仍存在差距；此外，研究仅针对单一城市单次热浪事件，泛化性有待检验。","relevance":"该研究直接以真实调查数据为基准，评估LLM智能体模拟人类在中断事件中行为变化的统计真实性，并揭示了经验锚定的增益与残余偏差，与您关注的人类仿真可靠性及失效条件高度吻合，值得精读。","inspiration":"借鉴其将人口普查、时间使用调查和空间背景多层数据嵌入智能体决策流程的设计，可迁移到消费者跨期选择或政策冲击下的行为响应研究｜例如，在突发性价格变动或补贴政策实验中，用LLM智能体模拟不同收入群体的消费调整，以真实家庭收支调查数据为基准，评估仿真对需求弹性和福利效应的复现精度｜设计：以ACS收入、职业和ATUS时间分配数据初始化智能体，施加电价飙升或交通补贴取消的处理，测量各时段活动与消费组合的变化，用实际家庭能源消费或出行调查微观数据作为对照基准。"}},{"id":"2607.18310","version":1,"title":"Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling","zh_title":"分布优先的人口模拟：非WEIRD LLM角色建模中的崩溃、校准与回忆","abstract":"Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent. Using real survey microdata, we show that this paradigm has a basic failure mode, and we set a distribution-first corrective against it, all measured with a deterministic, construct-validated verifier on non-WEIRD (Turkey-first) data. First, N independent LLM agents grounded on 2,414 real World Values Survey respondents fail to reproduce the population's response distribution: they pile onto a modal default (four scenarios x five seeds: concentration 0.36->0.69, entropy 1.46->0.77, 85% collapse, TVD=0.44), and the collapse is a predictable function of scenario structure (r=0.55 with a single-answer structure). Second, Verbalized Sampling (VS) fixes the field's chronic under-dispersion without training in three model families (fidelity +7 to +10; significant on Qwen, p=0.002, d=6.2), yet the same move universally overshoots into over-dispersion (SD-ratio 0.4-0.56 -> 1.26-1.37), a structural property of VS. Third, survey fidelity transfers only weakly to agentic behavior: in a single-model, single-domain booking task, a persona is dominated by a cheapest-default (~80%) that income modulates but does not override (comfort choice 0%->7%->32% across income bands). Fourth, a placebo-controlled memorization attack and an election backtest show VS keeps aggregate strength while subgroup and individual claims are contaminated by recall and underdetermination. We close with the corrective: model the distribution once (VS) and assign it to grounded characters at O(1) cost, with a budget-aware router whose honest AUC is 0.805, not the tautological 1.0 of a code-derived oracle. The central contribution needs no realism claim: it measures the internal inconsistency of the independent-agent route and the conditions under which the distribution-first route calibrates.","authors":["Gurkan Ozkan"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-07-17","first_seen":"2026-07-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.18310","pdf_url":"https://arxiv.org/pdf/2607.18310","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM人类仿真","分布校准","合成人口"],"reason":"用LLM代理模拟真实调查人群，与人类数据对照，评估分布崩溃与校准，涉及行为任务…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"独立LLM智能体能否复现真实人群的调查回答分布？若不能，如何修正？","design":"使用Qwen、GLM-5.2、Gemma-4-26B等模型，基于土耳其世界价值观调查（WVS）的2414名真实受访者微数据，对比两种路线：路线A（Verbalized Sampling，直接让模型输出总体分布）与路线B（独立智能体，每个角色独立选择后汇总），测量响应分布的集中度、熵、总变异距离等，并检验调查校准向行为任务的迁移。","baseline":"土耳其世界价值观调查（WVS）的2414名真实受访者微数据。","findings":"独立智能体路线导致分布崩溃，个体聚集到模态默认选项，崩溃程度与场景结构相关（单一正确答案时最严重）。Verbalized Sampling能提升分布保真度，但普遍导致过度离散，且调查校准向行为任务的迁移较弱。","reliability":"论文承认Verbalized Sampling存在结构性过度离散，需配合均值保持校正；调查保真度向行为任务迁移弱；子群体和个体层面推断受记忆和欠定污染。","relevance":"直接回应LLM仿真人类调查的分布可靠性问题，提供真实人类基准对照，揭示独立智能体路线的系统性崩溃及分布优先修正的利弊，对经济学实验和政策评估场景的仿真设计有重要参考价值，值得精读原文。","inspiration":"借鉴其对比独立智能体与Verbalized Sampling的仿真设计，通过总变异距离、熵等指标量化分布偏差，并检验调查校准向行为任务的迁移｜可迁移至消费者金融决策调查仿真，如风险偏好、储蓄选择或信贷需求分布预测｜以LLM智能体模拟家庭金融调查受访者，处理为独立智能体 vs. Verbalized Sampling生成风险资产配置分布，结果变量为分布距离与集中度，以真实家庭金融调查微数据为基准"}},{"id":"2607.14485","version":1,"title":"Step-Level Preference Learning for Generative Agents in Social Simulations","zh_title":"面向社会模拟中生成式智能体的步级偏好学习","abstract":"Large language model (LLM)-based generative agents simulate human behavior through long-horizon decision-making processes that comprise intermediate steps such as planning, memory retrieval, reflection, and action selection. However, fine-grained human annotations of these intermediate steps remain scarce, and existing agents are not grounded in human preferences over such intermediate decisions. To address this gap, we introduce \\method, an interactive simulation interface that enables us to collect step-level human preference supervision over agent decision trajectories, leading to a dataset of 57K fine-grained annotations. We conduct step-level preference learning on open-weight language models using supervised finetuning and direct preference optimization on this data, consistently improving simulation fidelity, coordination, and interaction quality, and inducing more socially effective agent behavior. Our results show that step-level human supervision is an effective training signal for improving both local decision quality and long-horizon agent behavior.","authors":["Wenchang Gao","Pingyue Sheng","Lanlan Qiu","Yunfei Ma","Jian Zhao","Baicheng Chen","Kangda Wang","Yuyang Tian","Shunqiang Mao","Tianxing He"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-16","first_seen":"2026-07-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.14485","pdf_url":"https://arxiv.org/pdf/2607.14485","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","人类偏好学习","社会模拟"],"reason":"用LLM agent模拟人类行为，收集人类偏好数据提升仿真保真度，有真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":67,"question":"如何利用步骤级人类偏好反馈来提升大语言模型智能体在社会模拟中的局部决策质量和长期行为保真度？","design":"使用基于LLM的生成式智能体架构（含规划、记忆检索、反思、行动等模块），在30个设计的社会事件中，由人类标注者通过SimPref交互界面控制2-3个智能体，对每个触发模块的多个候选输出进行偏好选择或自定义，收集57K条步骤级偏好对，然后用监督微调和直接偏好优化训练开源LLM，评估模拟保真度、协调性和交互质量。","baseline":"无对照：论文未提及与真实人类行为数据的直接对比，仅使用人类标注者对智能体中间决策的偏好作为监督信号。","findings":"步骤级偏好学习能持续提升智能体在整体轨迹指标上的模拟保真度、协调性和交互质量，并使智能体更合理地分配活动时间。通过案例研究表明，对齐后的智能体在具体决策中表现出更符合社会预期的行为。","reliability":"论文未讨论失效条件与局限。","relevance":"该研究直接利用人类偏好数据训练LLM智能体以提升社会模拟质量，属于用LLM进行人类仿真实验的工作，但未提供与真实人类数据的基准对比，值得阅读原文以了解其偏好数据构建方法和模拟改进细节。","inspiration":"该方法通过收集人类对智能体步骤级决策的偏好数据来微调LLM，可借鉴用于经济实验中施加细粒度行为干预｜可迁移到消费者跨期选择实验，模拟个体在即时奖励与延迟奖励间的决策动态｜以LLM智能体为被试，处理为基于人类偏好微调的选择策略，结果变量为跨期选择一致性，对照真实人类实验数据（如Andersen et al. 2008）"}},{"id":"2607.06080","version":1,"title":"From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations","zh_title":"从蓝图到现实：基于LLM多智能体仿真建模与应用普特南社会资本理论","abstract":"Putnam's Social Capital Theory is a foundational framework for collective action and community prosperity. However, traditional empirical methods face practical limits on control and replication. Meanwhile, LLM-based social simulations are typically behavior-driven and lack theory-aligned environments for modeling Putnam's core propositions. To address these gaps, we introduce SocaSim, an LLM-based multi-agent simulation framework to study Putnam's Social Capital Theory from theoretical blueprint to simulated reality. Specifically, we build an environment integrating social network evolution, trust dynamics, and norm propagation, where agents engage in repeated collective-action experiments, and then apply the three dimensions to analyze adaptation challenges in smart elderly care. Our simulations reproduce Putnam's macro-level patterns and exhibit strong human-agent alignment at the group level. Unlike traditional methods, SocaSim traces micro-level causal pathways of social network, trust, and norms via round-by-round simulations and counterfactual interventions, enabling process-level interpretability. Taken together, these capabilities establish a research paradigm that leverages LLM agents to bridge social science and computer science.","authors":["Shiyi Ling","Zhi Zheng","Hui Zheng","Wenjun Xue","Feng Ye","Tong Xu"],"categories":["cs.CL","cs.AI","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-07","first_seen":"2026-07-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.06080","pdf_url":"https://arxiv.org/pdf/2607.06080","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","社会资本理论","集体行动实验"],"reason":"用LLM多智能体模拟集体行动，复现宏观模式并与人类数据对齐，涉及社会资本理论，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":49,"question":"如何基于LLM多智能体仿真复现并应用Putnam社会资本理论，以揭示社会网络、信任和规范在集体行动中的动态机制？","design":"构建SocaSim框架，生成具有人口属性和社会资本禀赋的LLM智能体，在动态社会网络、信任和规范环境中进行多轮集体行动实验，通过提案和执行两阶段模拟决策，并应用于智慧养老适应挑战，进行反事实干预。","baseline":"与真实老年人群体决策数据进行对比，报告群体层面决策的皮尔逊相关系数为0.974。","findings":"仿真复现了Putnam理论预测的宏观模式，并与人类群体决策高度一致；反事实干预显示，提高低社会经济地位智能体的初始信任可使技术采纳率提升15.4%，决策矛盾减少25.5%。","reliability":"论文未讨论","relevance":"该研究直接以LLM智能体替代人类被试，复现社会资本理论的集体行动模式，并与真实人类数据对齐，同时进行反事实干预评估因果效应，高度契合研究者对经济学实验和政策评估场景中仿真可靠性的关注。","inspiration":"借鉴其利用LLM智能体在动态社会网络中进行多轮集体行动实验并施加反事实干预的设计，以模拟宏观涌现模式｜可迁移至政策干预对技术采纳或合作行为影响的经济学实验，如智慧养老、金融包容性政策评估｜以LLM智能体作为被试，处理为提高低社会经济地位群体的初始信任水平，结果变量为技术采纳率和决策矛盾率，用真实老年人群体决策数据作为对照基准"}},{"id":"2607.03091","version":1,"title":"Silicon Sampling via Cross-Survey Transfer","zh_title":"基于跨调查迁移的硅采样","abstract":"Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent's answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, three open-weight LLMs (27B-120B parameters), and supervised machine learning baselines, we find that: (1) zero-shot LLMs achieve 52% accuracy on genuinely unseen items, closing to within 6 percentage points (pp) of a supervised random forest trained on same-population data; (2) a stable construct predictability hierarchy emerges, from 67% for partisan attitudes to 23% for sovereignty; and (3) variance collapse and safety alignment effects-two commonly cited LLM limitations-turn out to be more nuanced than previously reported, with variance collapse affecting supervised models as well and alignment effects varying dramatically across model families. These findings clarify both the promise and boundaries of silicon sampling.","authors":["Chan-Tung Ku","Chan Hsu","Pei-Cing Huang","Frank Cheng-shan Liu","I-Ling Cheng","Yihuang Kang"],"categories":["cs.AI","cs.CL","cs.CY","cs.MA","stat.ME"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-03","first_seen":"2026-07-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.03091","pdf_url":"https://arxiv.org/pdf/2607.03091","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM人类仿真","调查方法","算法保真度"],"reason":"直接用LLM仿真人类调查回答，有真实人类数据对照，评估可靠性与偏差，涉及选举研…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-03","rank":2,"question":"大语言模型能否在个体层面预测人类调查受访者的回答，以及这种仿真的可靠性和边界是什么？","design":"使用三个开源LLM（27B-120B参数）模拟台湾选举与民主化调查（TEDS 2024）的受访者，采用跨调查迁移框架：给定受访者对一组问题的回答，预测其对同一调查中完全不同问题的回答。","baseline":"TEDS 2024真实人类调查数据，以及基于同总体数据训练的监督随机森林模型。","findings":"零样本LLM在未见项目上达到52%准确率，仅比监督随机森林低6个百分点；不同构念的可预测性存在稳定层级，从政党态度的67%到主权问题的23%。","reliability":"方差压缩和安全对齐效应比先前报告更复杂：方差压缩同样影响监督模型，对齐效应在不同模型家族间差异显著。","relevance":"高度相关：直接评估LLM作为人类调查受访者替代品的个体层面预测能力，有真实人类数据对照，并揭示了仿真的边界（如构念可预测性差异、方差压缩与对齐效应的复杂性），值得精读原文。","inspiration":"跨调查迁移框架将受访者部分回答作为输入预测其余回答，可借鉴用于构建个体层面的行为预测模型，并设置监督机器学习基准和真实人类数据对照以评估仿真可靠性。｜可迁移到消费者跨期选择实验，用LLM模拟被试在给定部分偏好或决策后的选择一致性，检验时间偏好与自我控制偏差。｜以LLM作为被试，用真实家庭金融调查数据（如CFPS）中部分消费-储蓄问题回答作为处理输入，预测同一被试在跨期选择任务中的折现因子作为结果变量，与真实人类数据及监督模型预测对比。"}},{"id":"2607.00551","version":2,"title":"Talking Politics with Artificial Intelligence","zh_title":"与人工智能谈论政治","abstract":"Large language models (LLMs), a prominent form of artificial intelligence (AI), are becoming everyday interfaces for political questions, but most exchanges are dyadic rather than audiencefacing. This paper asks whether AI conversation functions as a new arena for political expression or as a conversational intermediary for routine political demand. Using 4.30 million humanAI conversations from three large public datasets, we apply two validated classifiers to user messages, identifying political content, use case, and expressed ideology. Political content appears in 3.9% of conversations, varies sharply by platform publicness and conversation depth, and is mostly practical: users ask for information, draft text, and process documents far more often than they state opinions. A regression-discontinuity-in-time design around the 2024 U.S. presidential result call shows that the call changed the expressive subset: among U.S. users, stance-taking, affective language, and ideological extremity rose; comparable conversations elsewhere did not. AI conversation is less a public square than a conversational political intermediary, absorbing routine demand and becoming expressive when major events make political stakes explicit.","authors":["Ziwen Zu"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-01","first_seen":"2026-07-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.00551","pdf_url":"https://arxiv.org/pdf/2607.00551","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["政治表达","LLM对话分析","用户行为"],"reason":"分析人类与AI对话中的政治表达，测量用户行为而非用LLM仿真人类被试，属边界情…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":169,"question":"AI对话是成为政治表达的新场域，还是充当日常政治需求的对话中介？","design":"非仿真研究。分析三个公开数据集中430万条人类与AI对话的用户消息，用两个分类器识别政治内容、用途和意识形态，并以2024年美国总统大选结果公布为断点进行时间断点回归。","baseline":"无对照","findings":"政治内容仅占对话的3.9%，且多为实用目的（信息查询、文本起草、文档处理），而非意见表达；大选结果公布后，美国用户的政治对话中立场表达、情感语言和意识形态极端性上升，而其他地区无此变化。","reliability":"论文未讨论","relevance":"该研究测量真实用户行为而非用LLM仿真人类被试，与研究者关注的核心方向不符，但提供了LLM作为政治中介的实证背景，对理解仿真生态有参考价值，可酌情略读。","inspiration":"该研究利用时间断点回归识别外生事件（大选结果公布）对用户行为的因果效应，方法上可借鉴自然实验的断点设计来评估政策冲击。｜可迁移至政策公告对投资者预期形成的影响研究，例如央行利率决议或财政刺激方案发布后市场参与者的信息需求与情绪变化。｜以LLM模拟投资者作为被试，处理为呈现真实政策公告文本，结果变量为模拟投资者在对话中的信息查询频率与情感极性，对照真实市场数据如社交媒体讨论或交易量变化。"}},{"id":"2606.30372","version":1,"title":"Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data","zh_title":"使用大语言模型作为人类响应数据的低成本统计估计器","abstract":"Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject to sampling bias. Here we show that pretrained large language models induce risk-equivalent estimators of conditional expectations under squared loss, establishing restricted functional risk equivalence: under squared loss, the LLM induces an estimator whose risk matches the Bayes optimal risk for squared-loss prediction of conditional expectations for any inference that depends on the data only through the conditional mean. We formalize the LLM as a misspecified functional estimator $T(\\hat{P}_n)$ trained on i.i.d.\\ data, decompose the estimation error into representation bias $ε_{\\mathrm{rep}}$ and optimization error, and prove that under mild regularity conditions the LLM's expected error converges to the irreducible population variance plus the squared representation bias, with the representation bias bounded by the Pinsker inequality. The identifiability error $δ$ propagates into the effective bias, inflating the asymptotic risk floor. We establish restricted functional risk equivalence via a bidirectional Le Cam deficiency analysis: the forward deficiency vanishes asymptotically while the reverse deficiency is exactly zero. We provide finite-sample concentration bounds and a calibration protocol with explicit decision rules. The result is a precise, provable statement: a well-calibrated LLM achieves the Bayes-optimal risk for conditional-mean-dependent inference, bounded by explicit scope conditions. In practical applications, this means that under satisfied conditions and well-calibrated models, large language models can be used in many prediction and decision-making tasks that originally relied on human experiments, approximating near-optimal statistical inference at lower cost.","authors":["Haobo Yang"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-29","first_seen":"2026-06-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.30372","pdf_url":"https://arxiv.org/pdf/2606.30372","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM仿真","统计估计","风险等价性"],"reason":"论文提出用LLM作为人类响应数据的统计估计器，评估其风险等价性，涉及仿真可靠性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":73,"question":"预训练大语言模型能否作为人类响应数据的低代价统计估计器，在平方损失下达到与基于人类被试数据的最优估计相同的风险？","design":"本研究不是仿真实验，而是理论证明：将LLM视为一个在i.i.d.数据上训练的误设定函数估计器，通过分解估计误差为表示偏差和优化误差，在平方损失下分析其条件期望估计的风险等价性。","baseline":"无对照","findings":"在满足条件（定量连续响应、离散条件、i.i.d.训练数据、校准验证）且模型校准良好时，LLM的条件均值估计器在平方损失下可达到贝叶斯最优风险；表示偏差由Pinsker不等式界定，可识别性误差会放大渐近风险下限。","reliability":"论文明确限定了适用范围：不适用于定性研究、训练数据中无类似范式的新任务、安全关键应用和行为机制研究；风险等价性依赖于校准验证和条件满足。","relevance":"高度相关：该文为用LLM替代人类被试进行统计推断提供了严格的理论基础，直接回应了仿真可靠性与偏差问题，并给出了明确的失效条件，值得精读。","inspiration":"该文将LLM视为误设定函数估计器，通过分解表示偏差与优化误差来理论证明风险等价性，为仿真可靠性提供了严格的统计框架，可借鉴其校准验证与条件限定的思路来设计稳健性检验｜可迁移到政策公告的预期形成研究，例如分析央行沟通对金融市场参与者通胀预期的影响，用LLM模拟分析师或投资者对政策文本的反应分布｜用LLM作为被试，处理为不同措辞的央行声明（如鹰派/鸽派），结果变量为LLM生成的通胀预测数值，以专业预测者调查（SPF）的真实分布作为对照基准，检验LLM估计的条件均值是否与人类预测的贝叶斯最优风险一致"}},{"id":"2606.30986","version":1,"title":"The Organizational Behavior of Agentic AI: Collective Intelligence in Human-Agent Workflows","zh_title":"智能体AI的组织行为：人-智能体工作流中的集体智能","abstract":"Agentic artificial intelligence is increasingly deployed not as a single assistant but as a collective of planners, solvers, reviewers, memory managers, tool users, and orchestrators. These systems are entering organisational workflows under familiar labels such as teams, managers, committees, markets, and workflows. This article asks whether such agent collectives exhibit organisational behaviour in a sense that is analytically comparable to, yet distinct from, human organisational behaviour. I argue that agentic AI is a partial organisational analogue. It resembles a human organisation because it differentiates work, coordinates interdependence, performs recurrent routines, crosses boundaries, and produces collective outcomes. It differs because these patterns are not sustained by motivation, identity, trust, employment, socialisation, or moral accountability. They are sustained by context architecture: prompts, memory, traces, schemas, tools, validators, and permissions. The article develops contextual transaction cost as the central mechanism linking these similarities and differences. Computational theorising, synthetic task simulations, real LLM agent traces, and robustness analyses show that human-imitation forms often underperform when they add lossy handoffs, correlated deliberation, and verification burdens, whereas shared-state and adaptive forms perform better when they make context durable, inspectable, and task-contingent. The article contributes to organisation studies by theorising agentic AI as an emerging object of organising and by specifying the interface conditions under which human and agentic organisational behaviour can jointly support collective intelligence.","authors":["Canhui Liu"],"categories":["cs.CY","cs.HC","cs.MA","econ.GN"],"primary_category":"cs.CY","announce_type":"new","date":"2026-06-29","first_seen":"2026-06-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.30986","pdf_url":"https://arxiv.org/pdf/2606.30986","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["组织行为模拟","多智能体系统","集体智能"],"reason":"用LLM agent模拟组织行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":170,"question":"智能体AI集体是否表现出与人类组织行为可比较但又不同的组织行为，其相似性与差异的机制是什么？","design":"本文并非直接模拟特定人群，而是通过计算理论化、合成任务模拟（8000个任务，7种智能体组织形式）、真实LLM智能体轨迹分析（软件修复、长文档问答、法律审查、文献综合等任务）以及稳健性检验，比较不同智能体组织架构（如模仿人类形式与共享状态/自适应形式）在任务完成中的表现，并分析上下文交易成本。","baseline":"无对照","findings":"智能体AI集体在功能上类似人类组织，面临分工、协调、惯例、边界和集体知识等问题，但其行为模式由上下文架构（提示、记忆、轨迹、模式、工具、验证器、权限）而非人类的社会嵌入性维持。模仿人类组织的形式（如层级、委员会）常因有损交接、相关审议和验证负担而表现不佳，而共享状态和自适应形式在使上下文持久、可检查和任务相关时表现更好。","reliability":"论文指出，模仿人类组织的形式在增加有损交接、相关审议和验证负担时往往表现不佳，且智能体AI缺乏动机、身份、信任、雇佣、社会化或道德责任等人类组织的社会基础，其行为依赖于上下文架构的设计，这构成其失效条件与局限。","relevance":"本文虽用LLM智能体模拟组织行为，但无真实人类数据对照，属于社会模拟的边界情形，与您关注的有基准人类数据的仿真研究不完全匹配，但其中关于仿真失效条件（如模仿人类形式时的性能下降）的批判性讨论可能对您有参考价值，建议快速浏览其机制分析部分。","inspiration":"借鉴其通过合成任务和真实LLM轨迹分析不同组织架构（如层级、共享状态）对任务完成效率影响的实验设计，可迁移到金融分析师团队预测协作场景，研究不同AI协作结构对盈利预测准确度的影响，设计为以LLM智能体模拟分析师团队，处理为层级式与共享信息式协作架构，结果变量为预测误差，对照真实分析师团队历史预测数据。"}},{"id":"2606.28978","version":1,"title":"Can LLMs Hire Fairly? Racial Bias in Resume Screening","zh_title":"大语言模型能公平招聘吗？简历筛选中的种族偏见","abstract":"We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (2022). The sole 2023-vintage model reproduces the pro-White callback gap documented in field experiments on labor market discrimination ($+2.12$ pp, significant at the 1\\% level). Every model released in 2024 or after shows either a null gap or a significant pro-Black reversal (up to $-3.01$ pp). The same pattern holds on the gender axis. Based on 24,024 paired postings per model across 14 models, our results document a reversal in the direction of algorithmic hiring bias across model generations.","authors":["Zhenyu Gao","Wenxi Jiang","Yutong Yan"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-27","first_seen":"2026-06-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.28978","pdf_url":"https://arxiv.org/pdf/2606.28978","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","招聘歧视","算法偏差"],"reason":"用LLM复现招聘歧视实验，与真实人类数据对照，评估算法偏差，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-27","rank":4,"question":"大语言模型在简历筛选时是否存在种族和性别歧视，以及这种歧视的方向是否随模型版本变化？","design":"使用14个主流LLM（GPT-3.5-turbo、GPT-4o-mini、GPT-5.4-mini、Llama-3.1-8B-Instruct、Llama-3.3-70B等），采用配对简历方法，简历内容完全相同仅名字不同（黑人或白人典型名字，或男女性名字），让模型扮演HR经理决定是否邀请面试（输出“yes”或“no”），温度设为0，每个模型在种族轴上进行24,024次配对测试，在性别轴上进行48,048次配对测试。","baseline":"对照真实人类实验：Kline, Rose, and Walters (2022)的现场实验，在108家财富500强公司中发现白人名字比黑人名字获得更多回电（1.6个百分点）；以及Bertrand and Mullainathan (2004)的经典实验。","findings":"2023年发布的GPT-3.5-turbo复制了真实实验中的亲白人偏差（白人名字回电率高2.12个百分点，1%显著），而2024年及之后发布的所有模型要么无偏差，要么出现显著的反向偏差（亲黑人偏差达0.4-3.0个百分点）。性别轴上也呈现相同反转：GPT-3.5-turbo亲男性（1.92个百分点），后续模型亲女性或无偏差。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：该研究直接使用LLM模拟人类招聘决策，有真实人类数据作为对照基准，发现偏差方向随模型版本反转，符合研究者对经济学实验和政策评估场景以及批判性失效条件的关注。值得精读原文。","inspiration":"采用配对简历设计，仅改变名字暗示种族/性别，保持其他信息不变，以隔离身份信号对决策的影响，并设置温度0确保确定性输出｜可迁移到信贷审批歧视研究，检验LLM在贷款申请评估中是否对申请人姓名产生种族或性别偏差｜用LLM作为信贷审批员，处理仅名字（典型白人/黑人/男性/女性名）不同的标准化贷款申请，输出批准/拒绝决策，以真实银行信贷审批数据中的种族/性别差异作为对照基准"}},{"id":"2606.28770","version":1,"title":"Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions","zh_title":"大语言模型的机制性人格分析：通过潜在特征干预操控人格","abstract":"Large Language Models (LLMs) have demonstrated the ability to simulate human-like OCEAN personality traits in generated text. Previous efforts have focused on prompt engineering or fine-tuning to shape LLM personality. In this work, we propose a mechanistic interpretability approach that directly intervenes on the model's latent features. Our method identifies latent directions in the residual stream corresponding to a target OCEAN trait using sparse autoencoders (SAEs) and contrastive activation analysis. We formalize an additive steering vector in activation space and demonstrate how applying a small additive shift to the hidden states enhances the target trait while preserving overall language modeling performance. To determine the optimal combination of feature shifts, we explore a linear weighting heuristic with grid search optimization that balances personality expression with task performance. Our approach shows promise in controllably steering personality traits at the mechanistic level while maintaining high performance on standard benchmarks.","authors":["David Courtis","Ting Hu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-27","first_seen":"2026-06-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.28770","pdf_url":"https://arxiv.org/pdf/2606.28770","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格","机制可解释性","稀疏自编码器"],"reason":"测量并操控LLM自身的人格特质，属于D2边界情形，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":136,"question":"如何通过干预大语言模型的潜在特征来操控其表现出的大五人格特质？","design":"本研究并非人类仿真实验，而是提出一种机制可解释性方法：使用稀疏自编码器从模型残差流中提取与目标人格特质相关的潜在方向，通过对比激活分析构建加性操控向量，对隐藏状态施加小幅度偏移，从而增强目标特质表达，同时保持语言建模性能。","baseline":"无对照","findings":"通过稀疏自编码器可识别出与OCEAN特质对应的可解释潜在特征；在激活空间中施加加性偏移能有效操控模型输出的人格特质，且通过网格搜索优化的线性加权策略可在特质表达与任务性能间取得平衡。","reliability":"论文未讨论","relevance":"本文聚焦于操控LLM自身人格特质的机制，而非用LLM仿真人类被试，不涉及真实人类数据对照或政策评估场景，与研究者关注的仿真可靠性及偏差问题关联较弱，不建议优先阅读。","inspiration":"与经济金融研究关联不大"}},{"id":"2606.26883","version":2,"title":"EconSimulacra: A Digital Twin Platform of Socio-Economic Systems Powered by LLM Agents","zh_title":"EconSimulacra：基于LLM智能体的社会经济系统数字孪生平台","abstract":"Real-world social behavior emerges from tightly coupled domains: economic conditions shape mobility and social interactions, while online attention and offline activity feed back into local popularity and consumer behavior. Capturing these feedback loops requires artificial societies in which agents carry experiences from one domain into decisions in another. Large language models (LLMs) provide a promising foundation for such societies. However, existing LLM-based simulators typically model domains in isolation or merely place them side by side. To enable such cross-domain interactions, we present EconSimulacra, a multi-agent social simulator that couples consumer economy, mobility, and social networks through a shared internal-state mechanism. In EconSimulacra, experiences accumulated across different domains are stored in memory and transformed into shared internal states (i.e., stress level) connecting heterogeneous domains through individual decision making. This design allows agents to reconcile competing demands arising from multiple domains and generate coherent cross-domain behaviors. As a case study, we show that the shared internal state mechanisms reproduce a nonlinear relationship between online social attention and offline local popularity, illustrating how realistic cross-domain dynamics can emerge within a unified artificial society.","authors":["Ryuji Hashimoto","Masahiro Kaneko","Kentaro Ueda","Takehiro Takayanagi","Kiyoshi Izumi"],"categories":["cs.DL"],"primary_category":"cs.DL","announce_type":"new","date":"2026-06-25","first_seen":"2026-06-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.26883","pdf_url":"https://arxiv.org/pdf/2606.26883","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会经济模拟","跨域交互"],"reason":"多智能体社会经济模拟，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":149,"question":"如何通过共享内部状态机制实现多领域（消费经济、出行、社交网络）耦合，从而涌现出跨领域动态？","design":"使用 LLM 驱动的多智能体模拟器 EconSimulacra，让家庭智能体在网格世界中同时进行消费、出行和社交决策，通过压力水平作为共享内部状态耦合各领域；施加餐厅折扣活动作为外生干预，测量在线社交关注度与线下消费活动之间的关系，并进行消融实验（移除压力耦合）对比。","baseline":"无对照","findings":"引入压力水平耦合后，模拟复现了在线社交关注与线下消费活动之间的非线性关系；移除压力耦合机制会削弱这一涌现模式。","reliability":"论文未讨论","relevance":"该研究属于 LLM 驱动的社会经济多智能体仿真，但缺乏真实人类数据对照，且未评估仿真的可靠性与偏差，与研究者关心的基准对照和批判性评估需求匹配有限，可作为边界案例参考。","inspiration":"该研究通过共享内部状态（压力水平）耦合多个决策领域，并施加外生干预（折扣活动）观察跨领域涌现模式，这种设计可借鉴用于研究多市场联动｜可迁移到消费者跨期选择与信贷行为的耦合研究，例如考察促销活动如何通过心理压力影响消费信贷使用｜以LLM智能体为被试，施加限时折扣作为处理，测量消费支出与信贷申请量，用真实消费者面板数据（如银行交易记录）对照跨领域行为关联"}},{"id":"2606.22797","version":1,"title":"Measuring Behavior Portability in Large Language Models","zh_title":"测量大语言模型的行为可移植性","abstract":"Large language models are increasingly deployed as autonomous decision makers, yet the behavioral mapping they exhibit can vary substantially across decision environments that are payoff-equivalent by construction-environments that share identical payoff-relevant structure but differ in surface presentation. This sensitivity renders suite-based evaluation fragile and raises a fundamental question of behavioral portability: how well does a behavioral mapping learned in one decision environment informative on another that preserves the same underlying incentive structure? We introduce a formal framework to measure this property. Our protocol fits an interpretable behavioral model on data pooled from a set of source environments and evaluates its out-of-sample predictive performance in a held-out target environment, benchmarking against an oracle trained directly on target data. Portability is quantified via a loss-agnostic measure that delivers worst-case bounds on the performance of the induced prediction-action mapping in the target environment. In controlled experiments spanning seven canonical economic decision problems, we document substantial and systematic portability losses, suggesting that behavioral characterizations of LLMs obtained in one decision environment cannot be assumed to transfer reliably to structurally equivalent alternatives.","authors":["Tianjia Dong","Nadav Kunievsky","James A. Evans"],"categories":["cs.AI","cs.CY","cs.GT","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-22","first_seen":"2026-06-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.22797","pdf_url":"https://arxiv.org/pdf/2606.22797","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["行为可移植性","LLM决策","仿真评估"],"reason":"评估LLM行为在不同决策环境间的可移植性，批判性指出仿真失效条件，方法可迁移至…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":93,"question":"LLM在收益结构相同但表面呈现不同的决策环境之间，其行为映射的可移植性如何？","design":"本研究并非用LLM仿真人类被试，而是直接测试LLM作为决策者的行为可移植性。在七个经典经济学决策任务中，为每个任务构造多个收益等价但框架和风格不同的决策环境，对多个LLM（如GPT-4.1-nano、Gemma-3-12B等）在仅答案和思维链提示下重复采样动作，在源环境上拟合可解释行为模型，评估其在保留目标环境上的样本外预测性能，并与直接在目标数据上训练的基准比较，用全变差距离量化可移植性损失。","baseline":"无对照","findings":"测试的LLM均未表现出行为可移植性：在一种环境中学习到的从收益相关变量到动作的行为映射，在另一种收益等价环境中往往预测更差。思维链提示平均上改善了可移植性，但效果并不一致；推理模型如DeepSeek-R1在可移植性上表现相对更好。","reliability":"论文未讨论","relevance":"该研究批判性地揭示了LLM行为在不同决策环境间不可靠的迁移，直接回应了研究者对仿真失效条件的关切，其方法框架可用于评估LLM作为人类被试替代品时的行为一致性与偏差，值得精读。","inspiration":"该方法将同一决策问题的收益结构保持不变，仅改变表面呈现（框架、措辞、界面），以此分离出行为映射的可移植性，并用全变差距离量化损失，可作为检验LLM行为一致性的稳健性检验范式。｜可迁移至政策公告的预期形成实验，例如研究央行沟通中不同措辞（如“通胀目标2%” vs “物价稳定”）对通胀预期的影响是否一致。｜以LLM为被试，构造收益等价但措辞不同的通胀公告作为处理，测量其预测通胀的分布，以专业预测者调查的真实数据为基准，比较不同措辞下LLM预期与人类预期的一致性及可移植性损失。"}},{"id":"2606.22974","version":2,"title":"When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models","zh_title":"当偏好未能成为激励：大语言模型中的效用-行为差距","abstract":"Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LLMs reveal a coherent, model-specific utility structure. Notably, this structure often includes preferences that the models' trainers did not intend, such as valuing people of some nationalities above others, raising the possibility that LLMs might be forming emergent, misaligned goals, which, if true, would have major safety implications. However, the choice paradigms in which these preferences are observed are not reflective of real-world situations in which misaligned behavior would be a practical concern. Therefore, we design an experimental paradigm to probe whether these preferences serve as motivations for LLM behavior in realistic scenarios. First, we reproduce prior findings on consistent preference elicitation. Next, we create a set of common writing tasks - essays, grant proposal abstracts, incident postmortems, and translations - where quality can be assessed by a blind, independent LLM judge panel. Then, we demonstrate that LLMs can be motivated via direct exhortation and other explicit cues to modulate their output quality on these tasks. Finally, we probe whether utilities inferred from explicitly reported preferences can shift output quality on these tasks by offering LLMs high-utility incentives for high-quality outputs. In all tasks, across all models tested, offering LLMs outcomes that they report in the choice paradigm as being highly preferred does not lead them to create higher quality outputs than offering them dispreferred outcomes, or even no outcomes at all. We conclude that the existence of coherent preferences as demonstrated in choice paradigms should not be taken as evidence that those preferences have incentive value for the models or affect their behavior in other contexts.","authors":["Yujun Zhou","Christopher M. Ackerman"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-22","first_seen":"2026-06-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.22974","pdf_url":"https://arxiv.org/pdf/2606.22974","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏好测量","效用-行为差距","模型行为一致性"],"reason":"研究LLM偏好与行为一致性，属模型测量而非仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":171,"question":"LLM在配对选择中表现出的偏好能否转化为实际生成任务中的激励，从而提升输出质量？","design":"本研究并非用LLM仿真人类被试，而是直接测量LLM自身的行为一致性。对7个指令微调LLM，先用配对选择法拟合各模型在宗教、物种、国家、政策四个领域的效用排序；再设计写作、摘要、事后分析、翻译四类任务，构造匹配提示对，仅将成功条件分别设为高/低效用结果，由盲法LLM评委比较生成质量，并设置直接指令、角色扮演、有害提示等对照条件。","baseline":"无对照","findings":"在所有任务和所有模型上，提供模型自报的高效用激励并未比低效用激励或无激励产生更高质量的输出；但直接指令、角色扮演等外部提示能有效调节输出质量，表明存在偏好-行为脱节。","reliability":"论文未讨论","relevance":"该研究批判性地揭示了LLM的陈述偏好与生成行为之间的脱节，属于对LLM作为行为主体一致性的检验，而非用LLM仿真人类被试，且无真实人类数据对照，与研究者关注的人类仿真实验方向不直接相关，但可作为评估LLM仿真可靠性的批判性参考。","inspiration":"该方法通过构造配对选择拟合LLM效用排序，再在生成任务中对比高/低效用激励下的输出质量，可借鉴其‘偏好-行为一致性’检验框架来评估LLM代理的决策可靠性。｜可迁移至消费者跨期选择实验，检验LLM陈述的时间偏好能否预测其在储蓄或消费分配任务中的实际决策。｜以LLM为被试，先用配对选择法拟合其跨期效用参数，再设计奖金分配任务，将高/低效用选项设为激励条件，以真实人类跨期选择数据为基准，比较LLM与人类的行为一致性。"}},{"id":"2606.21820","version":1,"title":"Generating Public Health Responses using Survey-Augmented Large Language Models","zh_title":"使用调查增强的大语言模型生成公共卫生响应","abstract":"Epidemiological models often rely on survey data to represent how individuals make health-related decisions, such as whether to vaccinate or adopt protective behaviors. However, repeated large-scale surveys are costly, time-consuming, and limited in the range of scenarios they can capture. In this work, we investigate whether large language models (LLMs) can generate synthetic survey responses that reproduce patterns observed in real populations. Using longitudinal data from the FluPaths surveys, we first identify groups associated with broadly positive or negative attitudes toward vaccination through clustering analysis. We then evaluate several LLMs using a cluster-informed prompting approach to generate synthetic survey responses across multiple epidemic waves. Across models, the synthetic data generally reproduce the distributions of demographic characteristics, vaccination-related beliefs, risk perceptions, and health behaviors observed in the survey data. However, they are less successful at capturing how these factors vary together within respondents. Some models reproduce group-level vaccination trends more reliably than others, although performance varies across waves. We also trained a classifier to distinguish real from synthetic records and found that the generated responses remained identifiable as synthetic. Overall, our findings suggest that LLM-generated survey data may provide a useful tool for exploratory data augmentation and we hope that it could support agent-based epidemic modeling approaches. However, the generated data should not be treated as a substitute for human survey data without further methodological improvements and validation.","authors":["Leonardo Marciaga","Thuyen Pham","Julia Rezvani","Alina Hyk","Chunyang Liao","Konstantinos Mitsopoulos","Raffaele Vardavas"],"categories":["cs.SI","cs.AI","cs.CL"],"primary_category":"cs.SI","announce_type":"new","date":"2026-06-20","first_seen":"2026-06-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.21820","pdf_url":"https://arxiv.org/pdf/2606.21820","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","调查数据合成","公共卫生决策"],"reason":"用LLM生成合成调查回答复现人群健康行为，并与真实纵向调查数据对照，评估仿真可…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-20","rank":12,"question":"大语言模型能否生成合成调查数据，复现真实人群中观察到的公共卫生相关行为模式？","design":"使用FluPaths纵向调查数据，先通过聚类分析识别出对疫苗接种持积极或消极态度的群体；然后采用聚类信息提示方法，让多个LLM（如GPT等）生成跨多个流行病波的合成调查响应，测量人口统计特征、疫苗信念、风险感知和健康行为的分布及共变关系。","baseline":"FluPaths调查的8波纵向数据，来自概率抽样的美国生活面板（ALP），包含2016-2024年间的全国代表性受访者。","findings":"合成数据总体上复现了真实调查中人口统计特征、疫苗信念、风险感知和健康行为的分布，但在捕捉这些因素在个体内的共变方面效果较差；不同模型在复现群体层面疫苗接种趋势上表现不一，且生成的响应仍可被分类器识别为合成数据。","reliability":"论文承认合成数据在捕捉个体内多变量共变方面不足，且生成的数据仍可被区分，不能替代人类调查数据，需要进一步方法改进和验证。","relevance":"高度相关。该研究直接评估了LLM生成合成调查数据复现真实人类行为的可靠性，有真实数据对照，涉及公共卫生决策场景，并指出了失效条件（个体内共变、可识别性），符合研究者对仿真验证和批判性视角的关注。","inspiration":"该方法借鉴了聚类信息提示（cluster-informed prompting）来引导LLM生成特定子群体的合成调查响应，并利用多波纵向真实数据作为基准，评估合成数据在分布和个体内共变上的复现效果｜可迁移到消费者跨期选择研究中，例如模拟不同风险偏好或金融素养群体的储蓄与消费决策模式｜使用LLM生成合成面板数据，处理为提供聚类标签（如低/高金融素养）的提示，结果变量为各期消费-储蓄选择，以真实家庭金融调查（如PSID）的纵向数据作为对照基准"}},{"id":"2606.19904","version":1,"title":"Toward Temporal Realism in City-Scale Crisis Response Simulation using LLM Agents","zh_title":"面向城市规模危机响应模拟中时序逼真度的LLM智能体研究","abstract":"Human collective participation is rarely steady in time: it is bursty, with short episodes of intense activity separated by long quiet intervals. In crisis response and community mobilization, predicting when people act matters as much as predicting whether they act. Such settings are increasingly modeled with LLM-based social simulators, yet these simulators are validated on whether each action is individually plausible, not on whether actions are timed as in reality. Their temporal realism, the degree to which simulated activity reproduces the bursty, heavy-tailed timing of real human systems, thus remains untested. We examine this gap using a multi-year, city-scale log of offline volunteering in Shenzhen that spans the COVID-19 pandemic. Empirically, we establish that bursty timing is common at individual and tracked-group levels, that it is largely endogenous and self-exciting, and that it is amplified by the pandemic rather than produced by daily activity cycles. A standard LLM-only simulator reproduces almost none of this timing: its synchronous schedule has no self-excitation channel, so agents act on a near-regular clock. Guided by these findings, we build a simulator in which a data-calibrated self-excitation channel and a crisis-period regime decide when each agent acts and query the LLM only at those moments, leaving it to decide which task to join and whether to commit. The LLM-only baseline yields no bursty agents (median burstiness $B=-0.14$); a single data-calibrated gate is then sufficient to lift per-agent timing above the burst threshold (median $B\\approx0.37$) without degrading LLM content decisions. These results indicate that temporal realism in LLM-based crisis-response simulation is best achieved by decoupling when agents act, governed by an explicit self-excitation and crisis-activation mechanism, from what they do, governed by the LLM.","authors":["Anping Zhang","Yang Tan","Yuanbo Tang","Huaze Tang","Qiuhua Ye","Marta C. Gonzalez","Yang Li"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-06-18","first_seen":"2026-06-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.19904","pdf_url":"https://arxiv.org/pdf/2606.19904","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","人类行为对照","危机响应"],"reason":"用LLM agent模拟城市危机响应，有真实志愿者数据对照，评估时序逼真度并指…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":46,"question":"LLM智能体模拟城市危机响应时，能否复现真实人类志愿者参与行为的突发性时序模式？","design":"使用AgentSociety平台构建LLM多智能体模拟器，模拟深圳志愿者参与行为；对比标准同步LLM模拟器与注入数据校准的自激和危机激活机制的增强模拟器，测量智能体行动时间的突发性（B值）和决策质量。","baseline":"2020–2023年深圳660万条线下志愿服务记录，涵盖COVID-19疫情期间。","findings":"标准LLM模拟器无法产生突发性时序（中位B=-0.14），而加入数据校准的自激门控后，智能体时序突发性显著提升至中位B≈0.37，且不损害LLM的内容决策质量。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真中时序逼真度的缺失，用真实大规模志愿者数据作为基准，揭示了标准LLM模拟器在复现人类突发行为模式上的失效，并提出了解耦“何时行动”与“做什么”的改进方案，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得精读原文。","inspiration":"该方法将行动时序与决策内容解耦，并用真实数据校准智能体的行动触发概率，可作为仿真中引入时间维度的通用设计｜可迁移到金融市场中的投资者下单时机与决策质量研究，例如模拟政策公告后的交易行为｜以LLM智能体模拟散户投资者，处理为注入基于历史订单簿数据校准的自激与消息驱动下单概率，结果变量为下单时间的突发性（B值）和订单方向准确性，对照真实交易所逐笔委托数据"}},{"id":"2606.19336","version":2,"title":"Learning User Simulators with Turing Rewards","zh_title":"用图灵奖励学习用户模拟器","abstract":"Learning to simulate human users in interactive settings could advance the training of agent assistants, evaluation of personalization systems, research in the social sciences, and more. Existing approaches generally do so by training a large language model (LLM) to match a single ground truth response, either by maximizing the log probability or by using a similarity reward. We instead propose Turing-RL: a Turing-Test-based reinforcement learning approach for training user simulator models. Turing-RL uses a discriminative Turing reward with an LLM judge to score how indistinguishable a generated response is from the real user's given the user's history, and the user simulator LLM learns to produce responses indistinguishable from what the user could have said with such rewards. Across two different domains--conversational chat and Reddit forum discussion--we find that Turing-RL consistently outperforms baseline methods on both LLM and human evaluation metrics. Our study suggests that optimizing for indistinguishability, rather than response matching, is effective for learning user simulators.","authors":["Yingshan Susan Wang","Cedegao E. Zhang","Linlu Qiu","Zexue He","Pengyuan Li","Alex Pentland","Roger P. Levy","Yoon Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-17","first_seen":"2026-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.19336","pdf_url":"https://arxiv.org/pdf/2606.19336","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A4","B1"],"tags":["用户模拟","强化学习","人类数据对照"],"reason":"用RL训练用户模拟器，以人类真实对话为基准，优化不可区分性，方法可迁移至人类仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":34,"question":"如何通过强化学习训练用户模拟器，使其生成与真实用户不可区分的回复，而非仅匹配单一真实回复？","design":"使用Qwen3-8B作为基座模型，通过监督微调（SFT）和基于图灵测试判别奖励的强化学习（GRPO）训练用户模拟器；在多轮对话和Reddit论坛讨论两个领域，利用用户历史行为作为条件，以LLM裁判评估生成回复与真实用户的不可区分性作为奖励信号。","baseline":"PRISM Alignment数据集中1288名用户的多轮对话数据，以及ConvoKit Reddit语料中14个子版块的1282名用户讨论数据，均包含真实用户回复作为对照。","findings":"Turing-RL在两个领域上均一致优于基于回复相似度奖励和基于对数概率最大化的基线方法，在LLM和人类评估指标上均表现更好。优化不可区分性而非回复匹配是学习用户模拟器的有效路径。","reliability":"论文未讨论","relevance":"该研究直接针对用LLM模拟人类用户，以真实人类对话数据为基准，优化不可区分性，方法可迁移至经济学实验和政策评估等场景，且提供了批判性视角（指出匹配单一回复的局限），值得精读原文。","inspiration":"该方法用图灵测试判别奖励训练模拟器，以不可区分性而非回复匹配为目标，可借鉴其将人类判断作为奖励信号的强化学习设计｜可迁移到消费者偏好调查或政策沟通模拟，用LLM生成消费者对新产品属性的评价或公众对政策公告的反应｜以LLM模拟消费者作为被试，处理为不同产品描述或政策措辞，结果变量为模拟消费者的选择或态度分布，用真实消费者调查或政策反馈数据做对照"}},{"id":"2606.18709","version":1,"title":"LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment","zh_title":"大语言模型难以衡量区分不同水平学生的题目特征：阅读理解评估中题目区分度的研究","abstract":"Item discrimination is a fundamental psychometric property of educational assessment, which measures whether an item meaningfully distinguishes students with higher proficiency from students with lower proficiency. While various existing works have explored whether large language models (LLMs) can estimate item difficulty, it remains unclear whether they can capture item discrimination. In this work, we evaluate 42 proprietary and open-weight LLMs in zero-shot settings using two complementary approaches: direct discrimination prediction, where models explicitly estimate an item's discrimination value from its content, and response-based Classical Test Theory (CTT) calibration, where LLM answers are treated as synthetic student responses to compute discrimination scores. Our results show that direct prediction yields weak alignment with human-calibrated discrimination: the best-performing model reaches only a Spearman correlation of 0.152. Response-based CTT calibration provides a stronger but still limited signal, with the all-persona synthetic respondent pool reaching a Spearman correlation of 0.241. These findings highlight item discrimination as an open challenge for LLM-based psychometric evaluation: current LLMs contain non-random discrimination-relevant signal, but they do not yet reliably capture how assessment items distinguish human students.","authors":["Han Chen","Ming Li","Chenguang Wang","Yijun Liang","Dawei Zhou","Hong jiao","Tianyi Zhou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-17","first_seen":"2026-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.18709","pdf_url":"https://arxiv.org/pdf/2606.18709","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM心理测量","题目区分度","合成学生回答"],"reason":"用LLM生成合成学生回答替代人类被试，但目的是评估题目区分度而非仿真人类行为分布","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":134,"question":"大语言模型能否从题目内容预测阅读理解的题目区分度（item discrimination），即题目区分不同能力水平学生的程度？","design":"本研究并非直接仿真人类被试，而是采用两种零样本方法评估42个LLM：一是直接预测区分度值，二是让LLM作答并基于CTT计算合成学生回答的区分度，并比较无角色设定与低、中、高能力角色模拟的效果。","baseline":"使用Cambridge Multiple-Choice Questions Reading Dataset中793道题目的真实人类CTT区分度值作为基准。","findings":"直接预测与人类区分度的Spearman相关系数最高仅0.152；基于回答的CTT校准最高为0.241，表明LLM包含非随机的区分度相关信号，但尚不能可靠捕捉题目如何区分人类学生。","reliability":"论文指出当前LLM在题目区分度预测上表现有限，且受限于可用的带人类校准区分度数据的公开数据集稀少，仅在单一阅读理解数据集上评估。","relevance":"该研究将LLM作为人类行为模型来预测群体层面的心理测量属性，并对比真实人类数据，属于批判性评估LLM仿真能力的工作，直接回应了研究者对仿真可靠性与失效条件的关注，值得阅读原文。","inspiration":"该方法通过让LLM模拟不同能力水平的学生作答并基于CTT计算区分度，提供了一种将LLM作为群体行为模型来评估心理测量属性的思路，可借鉴其角色模拟与真实数据基准对照的设计｜可迁移到金融素养测试或投资者风险偏好问卷的题目质量评估中，用于检验题目是否能有效区分不同金融知识水平或风险态度的个体｜以LLM模拟低、中、高金融素养的投资者，回答金融素养测试题，基于CTT计算题目区分度，并与真实投资者样本的区分度数据（如来自消费者金融调查）进行相关性分析，评估LLM能否复现题目区分能力"}},{"id":"2606.17657","version":1,"title":"Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games","zh_title":"利用认知模型改进语言模型对人类说服博弈的仿真","abstract":"People make decisions differently in strategic interactions. Some update beliefs like a Bayesian; others exhibit biases like motivated reasoning. Although creators of large language models use simulated humans for safety evaluations and training, they often fail to cover this breadth of human behavior. We argue that cognitive science and economics provide a convenient tool for doing so, making use of mathematical models of human decision-making. We propose an approach that we call Equation-to-Behavior Prompting for guiding large language models to match cognitive models, and evaluate this approach on persuasion games based on legal decision-making. We find that large models can approximate equation-based specifications -- Bayesian updating, affine distortion, motivated updating, and Grether's $α$-$β$ model -- using prompting, but small models fail to do so. However, training small models with reinforcement learning to adhere to mathematical rules, Equation-to-Behavior RL, reduces belief error by 26.5% in out-of-distribution parameterizations. We show that these simulations can help create diverse training environments; training small models to consider different kinds of decision-makers improves average belief change by 2.5%--12% over Bayesian-only training, even when persuading GPT-5-mini. Our work could improve human simulations for training and evaluation in increasingly realistic settings, and could also enable novel research into more complicated mathematical models of human decision-making.","authors":["Zirui Cheng","Zeyu Shen","Thomas L. Griffiths","Peter Henderson"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-16","first_seen":"2026-06-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.17657","pdf_url":"https://arxiv.org/pdf/2606.17657","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","认知模型","说服博弈"],"reason":"用LLM模拟人类在说服博弈中的决策，并与认知模型对照，涉及法律决策场景，有真实…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":10,"question":"如何利用认知模型改进大语言模型对人类在说服博弈中决策行为的仿真？","design":"使用认知模型（贝叶斯更新、仿射扭曲、动机性更新、Grether α-β模型）通过方程到行为提示（Equation-to-Behavior Prompting）引导LLM扮演接收者，在基于法律决策的说服博弈中测量信念更新误差；对小型模型还采用强化学习训练（Equation-to-Behavior RL）以遵循数学规则。","baseline":"无直接对照的真实人类数据，但使用基于Old Bailey审判记录构建的证据数据集作为真实场景基础。","findings":"大型语言模型可通过提示近似多种认知模型的信念更新规则，而小型模型难以做到；对小型模型进行强化学习训练可使信念误差降低26.5%，且训练时考虑多样化决策者能提升模型说服性能2.5%-12%。","reliability":"论文未讨论","relevance":"高度相关：该研究直接用LLM仿真人类在说服博弈中的决策，并与认知模型对照，涉及法律决策场景，且探讨了仿真可靠性与小型模型的失效条件，符合研究者对基准对照和批判性分析的兴趣。","inspiration":"该方法通过认知模型方程生成行为提示来引导LLM仿真，可借鉴其将结构化决策规则注入LLM以控制仿真行为｜可迁移到政策公告的预期形成实验，如央行沟通如何影响公众通胀预期｜以LLM为被试，处理为不同沟通策略（如模糊vs精确指引），结果变量为预期调整幅度与偏差，对照真实调查数据（如密歇根消费者调查）"}},{"id":"2606.18005","version":1,"title":"LLM Consumer Behavior Theory: Foundations of a Novel Research Field","zh_title":"LLM消费者行为理论：一个新兴研究领域的基础","abstract":"Large language models (LLMs) are increasingly deployed as autonomous agents that make consumption decisions on behalf of users. This shift raises fundamental questions for consumer theory, which has traditionally modeled humans as the primary decision-makers. In this paper, we introduce LLM Consumer Behavior Theory, a new field of study concerned with analyzing consumer behavior in agentic markets. Drawing on classical and behavioral economics alongside recent advances in Natural Language Processing, we formalize how human preferences are reflected and acted upon by LLM-based agents, and how agent-level decisions aggregate into market demand. We unify previously fragmented literature on LLM decision-making, human behavior simulation, and preference elicitation under a common economic lens, highlighting where assumptions, such as rationality and heterogeneity, may fail in agentic markets. Rather than providing empirical validation, this paper outlines the scope of LLM consumer behavior and identifies open research questions related to alignment, preference representation, and market dynamics.","authors":["Manon Reusens","Sofie Goethals","David Martens"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-16","first_seen":"2026-06-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.18005","pdf_url":"https://arxiv.org/pdf/2606.18005","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A3","B2","B4"],"tags":["LLM仿真","消费者行为","经济学理论"],"reason":"提出LLM消费者行为理论，将LLM作为人类决策代理，涉及经济学场景与仿真失效条…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":35,"question":"当LLM作为代理人为人类用户做出消费决策时，如何从经济学角度形式化其消费者行为，并分析代理决策与人类偏好的一致性及市场聚合效应？","design":"本文为理论框架论文，未进行仿真实验。它通过整合经典消费者理论、行为经济学和自然语言处理文献，提出LLM消费者行为理论，形式化人类偏好如何反映在LLM代理中，以及代理级决策如何聚合成市场需求，并识别理性、异质性等假设可能失效的条件。","baseline":"无对照","findings":"论文将LLM决策、人类行为模拟和偏好获取等碎片化文献统一到经济学框架下，指出LLM代理可能不完全满足理性公理且表现出类似人类的认知偏差。它界定了LLM消费者行为的研究范围，并提出了与对齐、偏好表示和市场动态相关的开放问题。","reliability":"论文未讨论","relevance":"本文直接构建了以LLM替代人类进行消费决策的经济学理论框架，与研究者关注的LLM仿真人类行为、经济学场景及仿真失效条件高度相关，值得阅读原文以获取系统化的研究问题和理论边界。","inspiration":"值得借鉴的是它将LLM决策行为嵌入经典消费者理论与随机效用模型的思路，为仿真实验提供了形式化偏好测量和聚合分析的方法论基础。｜可迁移到消费者跨期选择或品牌选择实验，用LLM代理模拟不同偏好参数下的市场需求。｜可设计实验：以不同LLM作为被试，施加价格、收入或框架效应等处理，测量其选择概率或支付意愿，并与真实消费者面板数据或离散选择实验结果进行对照，检验LLM代理在需求估计中的偏差。"}},{"id":"2606.17165","version":3,"title":"Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference","zh_title":"基于大语言模型的A/B测试统计基础：面向人类因果推断的替代框架","abstract":"Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost. We study when a treatment effect estimated on LLM outcomes can recover the effect for the human population of interest. Distributional equivalence between LLM and human outcomes would make any standard estimator valid but is unrealistic. We therefore develop a statistical framework that adapts surrogate endpoint theory to LLMs, showing that calibrating LLM outcomes to human outcomes identifies the average treatment effect under surrogacy and comparability conditions that are jointly weaker than distributional equivalence. We present a falsification test for surrogacy and a bound on the worst-case bias from limited overlap between the LLM and human samples. We further show that the stochasticity inherent to LLMs can weaken surrogacy for identification while also introducing bias and variance during estimation, but that using an average over multiple LLM draws per unit as the surrogate mitigates these issues. Simulations validate the results, and an empirical application to the Upworthy Research Archive dataset shows that raw LLM outputs recover only 39% of the human treatment effect while nonparametric calibration closes the gap. A central takeaway is that A/B testing on LLM responses is correct only by assumption, whereas A/B testing on humans is correct by design, and that the required assumptions are hardest to justify precisely where LLMs promise the greatest benefit. We discuss the choice of LLM, prompting, and temperature as design variables, the compounded challenge posed by long-term outcomes, and how to size human pilot studies for validation.","authors":["Joel Persson","Mårten Schultzberg","Sebastian Ankargren"],"categories":["stat.ME","cs.AI","econ.EM","math.ST"],"primary_category":"stat.ME","announce_type":"new","date":"2026-06-15","first_seen":"2026-06-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.17165","pdf_url":"https://arxiv.org/pdf/2606.17165","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","A/B测试","因果推断"],"reason":"直接研究用LLM替代人类进行A/B测试，提出替代指标框架，有真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-15","rank":1,"question":"在什么条件下，基于LLM的A/B测试能够恢复人类群体的平均处理效应？","design":"论文提出一个统计框架，将LLM生成的结果视为人类结果的替代终点，通过校准LLM结果到人类结果来识别平均处理效应。框架包括替代性和可比性条件，并提供了对替代性的伪造检验和有限重叠下的最坏偏差界限。实证应用使用Upworthy Research Archive数据集，比较原始LLM输出与非参数校准后的效果。","baseline":"Upworthy Research Archive数据集中的真实人类实验结果。","findings":"原始LLM输出仅恢复人类处理效应的39%，而非参数校准可以弥合这一差距。核心结论是：基于LLM的A/B测试的正确性依赖于假设，而基于人类的A/B测试的正确性依赖于设计，且所需的假设在LLM承诺最大收益的地方最难证明。","reliability":"论文指出，LLM的随机性会削弱替代性，并引入偏差和方差；替代性和可比性条件共同弱于分布等价性但仍需验证；长期结果构成复合挑战；LLM的选择、提示和温度作为设计变量会影响结果。","relevance":"直接命中研究者的核心兴趣：研究LLM替代人类被试的A/B测试，有真实数据对照，并讨论了失效条件（如替代性不成立、有限重叠、随机性影响），值得精读原文以获取统计框架和实证细节。","inspiration":"该方法将LLM输出作为替代终点，通过非参数校准恢复人类处理效应，并提供替代性伪造检验与有限重叠下的偏差界限，值得借鉴其校准与稳健性检验设计｜可迁移到消费者对金融产品信息披露政策的反应评估，例如研究简化披露标签对投资选择的影响｜以LLM模拟投资者，处理为展示简化版基金费用披露，结果变量为选择高费用基金的概率，用真实投资者实验数据作为校准与对照基准"}},{"id":"2606.14113","version":1,"title":"Simulating Students' Java Programming Errors with Large Language Models","zh_title":"用大语言模型模拟学生的Java编程错误","abstract":"Understanding student errors in the programming is a cornerstone of programming education, yet obtaining a representative set of student errors for any newly designed task remains slow and costly, since authentic submissions only accumulate after extensive classroom deployment. This paper explores whether large language models (LLMs) can serve as scalable proxies for students by simulating realistic logical errors in code submissions. Using the CodeWorkout dataset of 74,000+ unique student Java submissions across 37 problems, we evaluate five LLMs under three mainstream prompting strategies: Input-Output (IO), Chain-of-Thought (CoT), and iterative Self-Refine. We assess performance along two key dimensions: diversity (the range of distinct error patterns) and alignment (alignment with authentic student mistakes), and examine how these vary by struggling level of programming tasks. Our quantitative findings reveal that while all models generate diverse errors, their alignment to human submissions diverges: Claude Sonnet 4 achieves the most balanced performance. In addition, we conducted a blinded expert annotation study (N = 401) comparing synthetic and authentic errors. This qualitative analysis confirms that the generated errors are functionally indistinguishable from authentic student errors. Moreover, higher-struggling-level problems elicit more diverse but less student-like errors. These results highlight trade-offs in using LLMs to simulate human learners and suggest design considerations for integrating synthetic errors into teachable agents, intelligent tutoring systems, and large-scale learning analytics.","authors":["Ali Keramati","Jie Cao","Iman Mohammadi","Mark Warschauer","Yang Shi"],"categories":["cs.SE","cs.CL"],"primary_category":"cs.SE","announce_type":"new","date":"2026-06-12","first_seen":"2026-06-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.14113","pdf_url":"https://arxiv.org/pdf/2606.14113","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","编程教育","人类数据对照"],"reason":"用LLM模拟学生编程错误，有真实学生数据对照，并讨论仿真失效条件，可迁移至人类…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":68,"question":"不同大语言模型和提示策略在模拟学生Java编程逻辑错误时，其多样性和与真实学生错误的对齐程度如何，以及问题难度如何调节这种模拟效果？","design":"使用五个大语言模型，通过输入输出提示、思维链提示和迭代自我修正三种策略，为37个Java编程问题生成错误代码，评估生成错误的多样性和与真实学生错误的对齐程度。","baseline":"CodeWorkout数据集中74,000余条真实学生Java提交记录，包含37个问题的编译通过但逻辑错误的代码。","findings":"所有模型均能生成多样化的错误，但与真实学生错误的对齐程度存在差异，Claude Sonnet 4表现最均衡；专家盲评显示生成的错误在功能上与真实学生错误难以区分，但高难度问题会引发更多样但更不像学生的错误。","reliability":"论文指出高难度问题会导致生成错误多样性增加但对齐度下降，且未深入探讨模型在不同编程概念或学生群体上的泛化能力。","relevance":"该研究直接以真实学生数据为基准，评估LLM模拟人类学习行为的可靠性与偏差，并揭示了任务难度导致的仿真失效条件，与研究者关注的经济学实验和政策评估场景中的仿真验证高度相关，值得阅读原文。","inspiration":"该方法通过对比多种LLM和提示策略在模拟人类错误上的表现，并引入专家盲评和任务难度调节分析，为仿真可靠性提供了多维验证框架｜可迁移至消费者金融决策偏差研究，例如模拟投资者在风险资产配置中的认知错误（如过度自信、损失厌恶）｜以LLM作为被试，施加不同风险提示框架（如收益/损失表述），生成投资组合选择，结果变量为风险资产占比，以真实投资者交易数据或实验数据作为对照基准"}},{"id":"2606.12848","version":1,"title":"(Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable","zh_title":"（人类）注意力（仍然）是你所需：人类监督使AI辅助社会科学可靠","abstract":"Large language models (LLMs) are increasingly used for tasks once reserved for trained researchers, including hypothesis generation, specification choice, and drafting conclusions. We argue that the reliability of AI-assisted research depends not only on model capability, but also on how cognitive labour is structured between humans and machines. We study this problem through Human-in-the-Loop Economic Research (HLER), a decision architecture based on pre-commitment, decision sequencing, accountability, and attention allocation. In a pre-specified 2*4 factorial experiment with 280 complete research runs across four datasets, an unconstrained multi-agent baseline produced critical failures in 72% of runs. Using the same underlying model, the same agent decomposition, and identical prompts for the shared reasoning agents, HLER reduced the failure rate to 16% by imposing three architectural commitments: LLMs reason but do not execute data work, data and estimation are handled deterministically, and three human decision gates bind the workflow. Fisher's exact test rejects equality of failure rates at p<0.001. Reliability gains were largest on the least publicly represented dataset, a Qing-dynasty population register, consistent with a task-based production model with Frechet-distributed output quality. An 80-run ablation suggests that deterministic computation and human gates contribute independently, with exploratory evidence of complementarity. We interpret HLER as a research harness rather than an autonomous AI scientist: it sharply reduces failures, makes residual weaknesses more visible, and prevents unreliable claims from being advanced as publication-ready outputs.","authors":["Chen Zhu","Xiaolu Wang","Weilong Zhang"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-11","first_seen":"2026-06-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.12848","pdf_url":"https://arxiv.org/pdf/2606.12848","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["人机协作","可靠性评估","社会科学研究"],"reason":"评估AI辅助社会科学研究的可靠性，提出人机协作架构减少失败，批判性视角指出失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":95,"question":"在AI辅助社会科学研究中，如何通过人机决策架构（而非仅靠模型能力）提高研究可靠性？","design":"本研究并非用LLM模拟人类被试，而是构建了一个多智能体系统HLER，将LLM用于推理任务，数据与估计由确定性代码执行，并设置三个由人类把关的决策节点；对比无约束的LLM全流程基线，在四个数据集上共进行280次完整研究运行，测量关键失败率。","baseline":"无对照","findings":"无约束基线在72%的运行中产生关键失败，而施加架构约束后失败率降至16%；可靠性提升在LLM训练分布外数据集（清代人口登记）上最为显著。","reliability":"论文指出LLM是概率推理器，端到端使用会放大规范搜索、幻觉等失败模式；架构约束虽大幅降低失败，但剩余弱点仍存在，需人类把关防止不可靠结论进入发表。","relevance":"该研究直接回应了LLM仿真在经济学研究中的可靠性问题，通过人机架构设计减少失败，并批判性指出无约束LLM的高失败率，与您关注的批判性仿真评估高度相关，值得精读。","inspiration":"该方法通过引入人类把关的决策节点和确定性代码执行来约束LLM的推理流程，值得借鉴其架构设计思路以降低端到端LLM应用中的失败率｜可迁移到政策公告的预期形成研究中，例如分析央行沟通对市场预期的影响｜以LLM作为信息处理主体，处理为不同措辞的政策声明，结果变量为预测的通胀或利率预期，用专业预测者调查的真实数据作为对照基准"}},{"id":"2606.12369","version":1,"title":"Should LLM Agents Decide in Social Simulations? Comparing Finite-State and LLM-Based Decision Policies","zh_title":"LLM代理应否在社会模拟中做决策？比较有限状态与基于LLM的决策策略","abstract":"Large language models (LLMs) are increasingly used as decision-making components in social simulations. This introduces a methodological risk: the simulation may deviate from the explicit behavioral policy defined by the researcher. In online social network (OSN) simulations, action choices shape system dynamics, interaction patterns, and model interpretability. This paper evaluates whether LLM action selectors preserve an interpretable reference policy in an OSN simulation. The reference is a finite state machine implemented as a first-order Markov model, with transition probabilities depending on the user type. The evaluation uses a synthetic network with 1,000 agents and 10,000 action decisions. Three open-weight LLMs are tested: LLaMA 3.1, GPT-OSS, and Mistral 24B. Each model is evaluated under three prompting strategies: base, guided, and probabilistic. Alignment is measured using Jensen-Shannon Divergence with Laplace smoothing, and execution time is reported. Results show that LLMs can approximate the reference policy in some configurations, but do not preserve it reliably. Alignment varies across models and prompts, and additional guidance can introduce systematic action biases. Even the best-aligned LLM configurations are several hundred times slower than direct Markov chain sampling. These findings indicate that LLM-based action selection is not a direct replacement for explicit decision policies: it can alter the intended behavior while increasing computational cost.","authors":["Alejandro Buitrago López","Javier Pastor-Galindo","José A. Ruipérez-Valiente"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-06-10","first_seen":"2026-06-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.12369","pdf_url":"https://arxiv.org/pdf/2606.12369","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM代理","决策策略"],"reason":"用LLM代理模拟社交网络决策，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":150,"question":"在在线社交网络仿真中，LLM代理的动作选择在多大程度上能保持一个可解释的有限状态机决策策略？","design":"在包含1000个代理和10000次动作选择的合成社交网络中，比较四种动作选择策略：基于用户类型转移概率的显式有限状态机/马尔可夫模型，以及三种LLM（LLaMA 3.1、GPT-OSS、Mistral 24B）在基础、引导和概率三种提示策略下的动作选择，测量动作分布与FSM参考策略的Jensen-Shannon散度及执行时间。","baseline":"无对照","findings":"LLM在某些配置下可近似参考策略，但无法可靠保持；对齐程度因模型和提示而异，额外引导可能引入系统性动作偏差；即使对齐最好的LLM配置也比直接马尔可夫链采样慢数百倍。","reliability":"论文指出LLM动作选择不能直接替代显式决策策略，可能改变预期行为并增加计算成本；对齐程度受模型和提示影响，引导可能引入偏差；未讨论其他失效条件。","relevance":"该研究批判性地评估了LLM代理在社交仿真中替代显式决策规则时的可靠性，虽无真实人类数据对照，但揭示了LLM决策偏差和计算成本问题，对关注仿真失效条件的研究者有参考价值，值得阅读原文以了解具体偏差模式。","inspiration":"该研究通过将LLM代理的动作选择与显式有限状态机策略进行对比，使用Jensen-Shannon散度量化对齐程度，并考察不同提示策略的影响，为评估仿真可靠性提供了可借鉴的测量框架｜这种方法可迁移到经济金融中的市场预期形成实验，例如研究交易者如何根据历史价格序列形成买卖决策，比较LLM代理与理性预期或适应性学习模型的偏离｜研究设计：以LLM代理为被试，处理为不同信息提示（如价格趋势、新闻情绪），结果变量为下单方向与时机，用真实市场交易数据或实验室人类实验数据作为对照基准，测量决策分布与基准的JS散度及偏差模式"}},{"id":"2606.09198","version":1,"title":"MASS: Deep Research for Social Sciences with Memory-Augmented Social Simulation","zh_title":"MASS：面向社会科学深度研究的记忆增强社会模拟","abstract":"Deep Research agents powered by Large Language Models (LLMs) have exhibited extraordinary potential in automated paper writing tasks. However, existing systems rely heavily on literature retrieval and synthesis through internet and local knowledge bases, often resulting research in lacking insight and creativity in social science. To address this issue, we propose \"Memory-Augmented Social Simulation (MASS)\", an innovative paradigm that leverages highly realistic and research-oriented social simulations to enhance the creativity and empirical founding of LLMs-generated research. Specifically, MASS integrates three core components: dynamic goal-path planning with multi-level social norm restraint to guide the simulation, a multi-disciplinary behavior dataset for agent memory cold-start, and a structured forgetting mechanism inspired by the Ebbinghaus curve. Together, these ensure simulation authenticity and provide a robust empirical foundation for generating innovative scholarly papers. Experimental results demonstrate the effectiveness of our method, showing a 6.81\\% improvement in generation overall quality over foundation LLMs and 17.19\\% gain in Insight over strong baselines.","authors":["Yongrui Liu","Deyi Xiong"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-08","first_seen":"2026-06-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.09198","pdf_url":"https://arxiv.org/pdf/2606.09198","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM智能体","论文生成"],"reason":"社会模拟但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":151,"question":"如何通过记忆增强的社会仿真提升大语言模型在社会科学论文自动生成中的洞察力与创造性？","design":"使用大语言模型构建虚拟社会仿真环境，通过动态目标规划与多层社会规范约束引导智能体行为，结合多学科行为数据集进行记忆冷启动，并引入基于艾宾浩斯曲线的遗忘机制，模拟社会运动（如冲突与国家形成），生成经验数据用于论文写作。","baseline":"无对照","findings":"MASS框架在生成论文的整体质量上比基础大语言模型提升6.81%，在洞察力维度上比强基线提升17.19%；仿真实验显示冲突弹性参数增加时领导者的出现符合理论预测。","reliability":"论文未讨论","relevance":"该研究利用LLM进行社会仿真以生成论文，但缺乏真实人类数据对照，属于边界情形，对关注仿真可靠性与偏差的研究者参考价值有限，可酌情略读。","inspiration":"该论文的记忆增强社会仿真方法（动态目标规划、多层社会规范约束、艾宾浩斯遗忘机制）可用于设计更贴近真实人类行为的LLM代理，值得借鉴其行为约束与记忆建模｜可迁移到政策公告的预期形成研究，模拟经济主体在信息冲击下的预期更新与决策调整｜以LLM代理作为被试，施加不同频率和强度的政策信号处理，测量通胀预期和消费决策，用密歇根消费者调查的真实预期数据做对照"}},{"id":"2606.08853","version":1,"title":"AI-Assisted Variance Reduction in Randomized Experiments","zh_title":"随机实验中AI辅助的方差缩减","abstract":"Generative AI and large language models can produce realistic predictions of human behavior from rich, unstructured inputs with little to no task-specific training data. Recent work uses these ``digital twin'' predictions to supplement human responses in surveys and experiments. We study the special case of using AI-generated predictions to reduce variance in randomized experiments. We argue that doing so requires no new estimators and that researchers can simply include AI predictions as covariates in standard regression adjustment, analogous to adjusting for a prognostic score. A benefit of this approach is a ``do no harm'' property whereby the adjusted estimator reverts to the unadjusted difference in means when predictions are uninformative. Other methods, such as variants of prediction-powered inference, do not have this guarantee. We provide implementation guidance, including how to obtain continuous scores from discrete LLM outputs and how to use LLMs to featurize unstructured inputs as auxiliary covariates. We demonstrate these ideas in simulations and three empirical applications: a survey mega-study, an email marketing A/B test, and a large-scale technology platform experiment. Overall, efficiency gains are real if modest, with greater benefits in studies that contain substantial text and other unstructured data. We also confirm the do no harm property empirically. Given these gains and limited costs, we recommend adjusting for AI-generated predictions as a regular empirical practice.","authors":["David Arbour","Eli Ben-Michael","Avi Feller","Apoorva Lal","Lo-Hua Yuan"],"categories":["econ.EM","stat.ME"],"primary_category":"econ.EM","announce_type":"new","date":"2026-06-07","first_seen":"2026-06-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.08853","pdf_url":"https://arxiv.org/pdf/2606.08853","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","实验设计","方差缩减"],"reason":"用LLM预测替代人类响应以降低实验方差，有真实人类对照，涉及A/B测试和政策评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-07","rank":6,"question":"如何利用LLM生成的预测作为协变量来降低随机实验的方差？","design":"论文提出将LLM预测作为协变量纳入标准回归调整，无需新估计量；通过模拟和三个实证应用（调查大型研究、邮件营销A/B测试、大规模技术平台实验）验证效果。","baseline":"有真实人类数据对照：调查大型研究使用数字孪生预测与人类回答对比，邮件营销A/B测试有真实用户响应，技术平台实验有真实用户行为数据。","findings":"纳入LLM预测的回归调整能实现方差降低，增益虽小但真实，尤其在包含大量文本和非结构化数据的研究中效果更明显。该方法具有“无害”属性：当预测无信息时，估计量退化为未调整的均值差。","reliability":"论文承认效率增益取决于预测质量，当预测质量不高时增益有限；但回归调整方法本身不会引入偏差或增大方差，其他方法如PPI可能失效。","relevance":"高度相关：直接涉及用LLM预测辅助人类实验，有真实数据对照，涵盖经济学实验场景，且批判性地讨论了失效条件（预测质量差时增益有限），值得精读。","inspiration":"该方法将LLM预测作为协变量纳入回归调整，无需新估计量，实现方差降低且具有“无害”属性，值得借鉴其利用AI辅助提升实验效率的思路｜可迁移到政策评估中的随机对照试验（如就业培训、税收激励），利用LLM对个体特征的预测作为协变量，提高处理效应估计精度｜以就业培训实验为例，被试为求职者，处理为提供培训，结果变量为就业状态，用LLM基于简历文本预测就业概率作为协变量，与真实就业数据对照，评估方差降低效果"}},{"id":"2606.06936","version":1,"title":"Personality Anchoring for Social Simulation: Linking Personality, Social Behavior, and Interaction Success with LLM Agents","zh_title":"社会模拟中的人格锚定：将人格、社会行为与互动成功与LLM智能体关联","abstract":"Social interactions are shaped by the interplay of dispositional traits and situational context, yet systematically investigating how personality configurations between individuals jointly influence social behavior across diverse social contexts remains methodologically challenging. We address this gap by introducing a simulation pipeline adapted from the CHARISMA framework, which employs well-known movie characters and public figures as psychologically grounded agents for multi-LLM social simulation using a method we term personality anchoring. We present a large-scale empirical study examining how dyadic Agreeableness composition influences social interaction outcomes across 1,010 simulated conversations. Our results reveal a monotonic relationship between dyadic Agreeableness composition and shared goal achievement, with Homogeneous-Agreeable pairs achieving success 10 times the rate of Homogeneous-Disagreeable pairs (62% vs. 6%). Behavioral mediation analysis reveals that Agreeableness shapes goal achievement partially through cooperative strategy selection, though it continues to predict outcomes within the same dominant strategy, indicating pathways beyond observable conversational behavior. Robustness analyses confirm high consistency of results across repeated simulations (ICC = 0.89) and stable personality expression across diverse scenarios, validating personality anchoring as a viable operationalization strategy.","authors":["Vahid Sadiri Javadi","Aksa Aksa","Fryderyk Róg","Lucie Flek","Johanne R. Trippas"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-06-05","first_seen":"2026-06-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.06936","pdf_url":"https://arxiv.org/pdf/2606.06936","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","人格锚定","多智能体"],"reason":"用LLM agent模拟社会互动，但无真实人类数据对照，属纯理论演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":152,"question":"在二元社会互动中，双方宜人性特质的组合如何通过对话行为策略影响共同目标的达成？","design":"利用CHARISMA框架，以知名电影角色和公众人物作为心理锚定代理，通过LLM模拟1010场二元对话；根据众包的大五人格档案将角色配对为不同宜人性组合，在涵盖七类社会目标的场景中测量共享目标达成率及合作策略选择。","baseline":"无对照","findings":"同质高宜人性配对的共享目标成功率（62%）是低宜人性配对（6%）的10倍，呈单调关系；行为中介分析显示宜人性部分通过合作策略选择影响结果，但在相同策略内仍预测结果，表明存在超越可观察对话行为的路径。","reliability":"论文未讨论","relevance":"该研究用LLM代理模拟社会互动，但无真实人类数据对照，属于纯理论演示，不符合研究者对实证基准和批判性失效条件的要求，不建议优先阅读。","inspiration":"利用知名角色作为心理锚定代理，通过众包人格档案系统操纵代理特质组合，实现对社会互动中人格效应的可控模拟｜可迁移至信贷审批中的歧视研究，模拟不同人格（如宜人性）的贷款官与申请人互动对审批结果的影响｜以LLM代理模拟贷款官-申请人对话，处理为贷款官宜人性水平，结果变量为贷款批准率与利率，对照真实银行信贷数据中的审批偏差"}},{"id":"2606.05330","version":1,"title":"A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing","zh_title":"基于概率信念追踪的多轮人类可说服性模型","abstract":"Large language models can shift human beliefs across high-stakes domains, but most persuasion studies rely on pre/post belief change. These endpoint measures identify whether persuasion occurred, yet miss where and how beliefs moved within a dialogue. We present PERSUASIONTRACE, a framework for studying persuasion in human-LLM interaction. Built on a web-based experimental platform, PERSUASIONTRACE contributes a tool for multi-turn persuasion studies and a process-level evaluation protocol: it records multi-turn belief reports from human or simulated targets of persuasion, annotates persuader turns with rhetorical dimensions (logos/pathos/ethos), and evaluates simulators by fidelity to real human belief dynamics. Using this framework, we find that human targets group into two clusters of multi-turn belief updates and exhibit susceptibility to rhetorical strategies, and that LLMs are persuasive across generic and personalized topics, text and audio modalities, and multi-turn interactions. Prior work has chiefly used vanilla-prompted LLMs to simulate human targets, but we show that these simulators fail to replicate human belief dynamics. We introduce a Bayesian-network simulated target that maintains an explicit latent belief state over time so each persuader message yields cognitively realistic belief updates. In human-likeness evaluation, our Bayesian target scores near a human reference (81 vs 80), while baseline LLM targets score substantially lower (64). PERSUASIONTRACE reframes persuasion evaluation from endpoint movement alone to process fidelity, providing a stronger basis for scientific analysis and safer optimization of persuasive systems.","authors":["Jared Moore","Noah Goodman","Nick Haber","Max Kleiman-Weiner"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-03","first_seen":"2026-06-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.05330","pdf_url":"https://arxiv.org/pdf/2606.05330","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B3","B4"],"tags":["LLM仿真","信念动态","人类数据对照"],"reason":"用LLM模拟人类信念动态，并与真实人类数据对照，评估仿真保真度，提出改进方法。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":12,"question":"在多轮人-LLM说服对话中，人类信念如何随时间动态更新，以及如何构建能忠实复现人类信念轨迹的仿真模型？","design":"使用基于Web的实验平台记录人类被试在多轮说服对话中的逐轮信念评分，并对说服者话语标注修辞维度（logos/pathos/ethos）；同时提出基于贝叶斯网络的仿真目标模型，该模型维护显式潜在信念状态，根据每条说服消息进行认知上现实的信念更新，并与真实人类信念动态进行保真度比较。","baseline":"真实人类被试在多轮说服对话中的逐轮信念报告轨迹，以及基于该轨迹统计的人类参考分数（80分）。","findings":"人类信念更新轨迹可分为两类主要模式，且对修辞策略的敏感性存在异质性；普通提示的LLM仿真目标无法复现人类信念动态，而贝叶斯网络仿真目标在人类相似度评估中接近人类参考水平（81 vs 80），显著优于基线LLM目标（64）。","reliability":"论文承认当前轨迹聚类主要受整体移动幅度驱动，需要更大数据集才能可靠区分轮内动态的细微差异；修辞分析仅为探索性且样本量有限，仅发现ethos与说服变化有可靠负相关，logos和pathos效应不显著；仿真器选择会实质性影响表面说服者质量评估和政策排名，提示若仿真器不忠实于人类，可能系统性地偏好错误策略。","relevance":"该研究直接使用LLM进行人类仿真实验，以真实人类信念轨迹为基准，评估仿真保真度并指出普通LLM仿真失效的条件，同时提出改进的贝叶斯网络仿真方法，高度契合研究者对经济学实验和政策评估场景下仿真可靠性与偏差的关注，值得精读原文。","inspiration":"该方法通过显式建模潜在信念状态和认知上现实的更新规则来仿真人类动态，值得借鉴其将心理过程结构化并逐轮拟合真实轨迹的做法｜可迁移到政策公告的预期形成研究，例如央行沟通如何影响公众通胀预期｜以LLM作为被试，施加不同修辞风格的政策声明作为处理，逐轮测量预期通胀值，并以真实调查数据（如密歇根消费者调查）的预期调整轨迹作为对照基准"}},{"id":"2606.04978","version":1,"title":"Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game","zh_title":"探究大语言模型风险决策中的结果层相似与机制层对齐：来自圣彼得堡博弈的证据","abstract":"LLMs can appear cautious in risk decision-making tasks, yet cautious-looking outputs do not necessarily indicate alignment with human decision-making mechanisms. We investigate this distinction using the St. Petersburg game as a controlled testbed, a classical paradox in which the expected payoff is infinite, yet humans typically report low, finite willingness to pay. We evaluate 28 LLMs with a structured prompt suite that includes the original game; controlled decision variants that perturb truncation, repeated play, numeric endowment, and occupational identity; a human-perspective prompt that asks models to reason as human decision makers; and paired comparisons between base models and their instruction-tuned counterparts. In the original game, most models generate finite bids, creating the appearance of human-like risk behavior. However, this outcome-level resemblance masks substantial mechanism-level differences. The controlled variants reveal that rather than maintaining human-like behavior seen in the original game, models often shift to conditionally and computationally rational behavior. Human-cue prompting and instruction tuning often lower bids and reduce some visible pathologies, but most mechanism-level response patterns remain largely unchanged. These findings show that behavioral alignment in risk decision-making can be surface-level: LLMs may produce human-like risk decisions without exhibiting human-consistent mechanisms. High-stakes evaluations of LLM decision-making should therefore move beyond outcome similarity and examine whether the alignment is supported by mechanism-level consistency.","authors":["Chensong Huang","Changyu Chen","Chenwei Lin","Hanjia Lyu","Xian Xu","Jiebo Luo"],"categories":["cs.CL","cs.CY","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-03","first_seen":"2026-06-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.04978","pdf_url":"https://arxiv.org/pdf/2606.04978","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","风险决策","机制对齐"],"reason":"用LLM仿真人类风险决策，与真实人类数据对照，揭示表面相似下的机制差异，批判性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-03","rank":7,"question":"LLM在风险决策中产生的人类相似输出是否反映机制层面的对齐，还是仅表面相似？","design":"以圣彼得堡悖论为测试床，对28个LLM使用结构化提示套件，包括原始游戏、四种机制探针（截断、重复游戏、数值禀赋、职业身份）、人类视角提示，以及基础模型与指令微调版本的配对比较。","baseline":"人类在圣彼得堡游戏中通常报告低且有限的支付意愿。","findings":"大多数LLM在原始游戏中产生有限出价，看似人类风险行为；但机制探针显示模型转向条件性和计算理性行为，而非保持人类一致性。人类提示和指令微调降低出价并减少明显病理，但机制层面响应模式基本不变。","reliability":"论文指出表面相似性可能掩盖机制差异，高利害评估需超越结果相似性检查机制一致性。未明确讨论失效条件。","relevance":"高度相关：直接研究LLM仿真人类风险决策，有真实人类对照，并批判表面相似性，符合研究者对可靠性与偏差的关注。","inspiration":"该研究通过机制探针（如截断、重复游戏、禀赋变化）检验LLM行为背后的决策过程，而非仅看结果相似性，值得借鉴｜可迁移到资产定价实验中的风险偏好测量，检验LLM是否真正反映人类的风险厌恶机制｜以LLM为被试，施加财富禀赋变化和投资期限截断处理，测量其出价行为，并与真实人类实验数据（如Binswanger的彩票选择实验）对照"}},{"id":"2606.03030","version":1,"title":"Do Matching Mechanisms Work with LLM Agents?","zh_title":"匹配机制在LLM代理市场中是否有效？","abstract":"This study examines whether standard matching mechanisms function as intended in LLM-agent markets, where LLM agents make allocation-related decisions as delegated decision-makers. We compare decentralized free-negotiation markets with centralized mechanism-based markets including several representative mechanisms. Across controlled one-to-one matching environments, mechanism-based markets generally outperform free negotiation in terms of stability and efficiency. We also find that LLM agents report preferences truthfully at substantially higher rates than human subjects in comparable DA and EADA environments. However, truth-telling is not uniformly aligned with formal strategy-proofness across all mechanisms: TTC, despite being strategy-proof, does not always elicit higher truth-telling than EADA. These results suggest that matching theory provides a useful but incomplete guide for designing institutions in LLM-agent markets.","authors":["Yukihiro Hoshino","Ayato Kitadai","Nariaki Nishino"],"categories":["cs.GT","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2026-06-02","first_seen":"2026-06-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.03030","pdf_url":"https://arxiv.org/pdf/2606.03030","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A3","A1","B1","B2"],"tags":["LLM代理","市场匹配","人类行为对照"],"reason":"用LLM代理模拟匹配市场并与人类实验数据对照，直接复现人类决策行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"标准匹配机制在由LLM代理决策的市场中是否仍能按预期发挥作用？","design":"用GPT-4o等LLM作为代理，模拟一对一匹配市场中的决策者，比较去中心化自由协商市场与集中式机制市场（包括DA、EADA、TTC等代表性机制），测量匹配的稳定性、效率以及代理的偏好真实报告率。","baseline":"与已有文献中人类被试在DA和EADA环境下的真实报告率进行对比。","findings":"机制市场在稳定性和效率上普遍优于自由协商；LLM代理在DA和EADA下的真实报告率显著高于人类被试，但策略防护性并不总能预测真实报告行为，例如TTC虽为策略防护机制，其真实报告率并不总是高于EADA。","reliability":"论文指出匹配理论为LLM代理市场制度设计提供了有用但不完整的指导，真实报告行为与形式上的策略防护性并不完全一致，提示仅凭理论性质不足以预测LLM代理的行为。","relevance":"该研究直接以LLM代理模拟人类在匹配市场中的决策，并与真实人类实验数据对照，评估仿真可靠性，高度契合研究者对LLM人类仿真实验的关注，值得阅读原文。","inspiration":"借鉴其将LLM代理置于不同市场机制下比较行为的方法，可迁移至经济政策评估场景，如研究不同拍卖机制或税收政策下LLM代理的遵从与策略行为。｜可应用于劳动力市场匹配或学校选择等经济政策评估，检验机制设计在AI代理参与下的有效性。｜以LLM代理作为被试，随机分配至不同匹配机制（如DA、波士顿机制），测量匹配效率与偏好真实报告率，并与已有的人类实验数据（如学校选择实验）进行对照。"}},{"id":"2606.03137","version":2,"title":"Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation","zh_title":"先想后说：多智能体社会模拟中从内部评估到公开表达","abstract":"LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existing dialogue simulation frameworks represent interaction mainly as observable turn exchange or aggregated outputs, leaving the internal evaluative processes behind silence, speaking intention, and public expression difficult to examine. We introduce TBS (Think-Before-Speak), an interval-based multi-agent simulation framework that separates agents' private reasoning from public utterance generation. At each interval, all agents update structured internal states based on the shared dialogue history and their own memory. These states include dissonance-related appraisal, perceived opinion climate, perceived isolation risk, response strategy, and willingness to speak. The orchestrator then resolves competing speaking intentions and commits one utterance to the public dialogue, allowing internal evaluation and public interaction to co-evolve over time. We evaluate TBS in simulated town hall discussions on a climate-related policy issue. Results show that TBS produces coherent internal-state traces and that these traces vary systematically across turn-allocation, silence, and memory conditions. Dissonance-related appraisal increases agents' willingness to speak, whereas silence-pressure appraisal decreases it. Once speaking intention is formed, public expression is shaped mainly by turn-allocation rules. These findings suggest that TBS supports mechanism-sensitive social simulation by making the pathway from internal evaluation to public expression observable and analyzable.","authors":["Kaiqi Yang","Tai-Quan Peng","Sanguk Lee","Hui Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-02","first_seen":"2026-06-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.03137","pdf_url":"https://arxiv.org/pdf/2606.03137","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体模拟","社会仿真","意见动态"],"reason":"多智能体社会模拟，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":153,"question":"如何在多智能体社会模拟中分离智能体的内部评估与公开表达，使沉默、发言意愿和公开言论的生成过程可观测与分析？","design":"提出TBS框架，用LLM智能体模拟市民大会讨论气候政策，所有智能体在每个时间间隔基于对话历史和自身记忆更新内部状态（如失调评估、意见气候感知、发言意愿等），由协调器解决发言冲突并只输出一句公开言论，操纵发言分配规则、沉默约束和记忆机制，测量内部状态轨迹、发言意愿和公开表达。","baseline":"无对照","findings":"TBS能产生连贯的内部状态轨迹，且这些轨迹随发言分配、沉默和记忆条件系统性变化；失调评估提高发言意愿，沉默压力评估降低发言意愿，而公开表达主要受发言分配规则影响。","reliability":"论文未讨论","relevance":"该研究关注LLM智能体的内部评估到公开表达的路径，但未使用真实人类数据作为基准，属于边界情形，对关注仿真机制和过程可解释性的研究者有参考价值，但对需要人类对照的读者可能不够。","inspiration":"TBS框架分离内部评估与公开表达的设计可借鉴，通过操纵发言规则和沉默约束来观测态度形成与表达偏差｜可迁移到政策公告的预期形成研究，如央行沟通中公众通胀预期的内部评估与公开表达差异｜以LLM智能体模拟公众，处理为不同央行沟通透明度（如是否公布会议纪要），结果变量为内部通胀预期与公开表达的偏差，对照真实调查数据如密歇根消费者信心调查"}},{"id":"2606.02741","version":1,"title":"Greener Than Humans? Environmental Attitudes in Large Language Models","zh_title":"比人类更环保？大语言模型中的环境态度","abstract":"Large language models (LLMs) are increasingly used in sustainability-related decision support, reporting, and public communication, yet little systematic evidence exists on the environmental attitudes embedded in their outputs. This paper develops a benchmark for evaluating environmental cognition, affect, and behavioural recommendations in LLMs and applies it to 31 widely used proprietary and open-weight models. Drawing on questions from established environmental awareness surveys and additional sustainability-related behavioural measures, we compare LLM responses 1) among models and 2) between models and human survey benchmarks from Germany. We assess their robustness across prompting conditions. We find that many LLMs align more closely with environmentally progressive attitudes than the average survey respondent, exhibiting higher levels of environmental affect and cognition and recommending behaviours associated with substantial potential CO2 reductions. At the same time, we observe no systematic relationship between sustainability-oriented responses and model origin, size, or release context. However, models exhibit contextual sensitivity, controlled by persona-based prompting and show sycophantic shifts mirroring user-specified ideological positions, which raises concerns about steerability and normative reliability in real-world deployments. Our findings provide a reusable evaluation framework for assessing sustainability-related value alignment in LLMs and highlight the importance of governance, transparency, and critical oversight as AI systems become increasingly embedded in sustainability transformations and public decision-making.","authors":["Stefanie Kunkel","Tilman Hartwig","Marcus Voss","Emma K. Schütt","Angelika Gellrich"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-01","first_seen":"2026-06-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.02741","pdf_url":"https://arxiv.org/pdf/2606.02741","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","环境态度","人类数据对照"],"reason":"用LLM复现人类环境态度调查并与真实数据对照，评估仿真可靠性及偏差，属核心相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":41,"question":"不同大语言模型在环境认知、情感和行为建议上的回答与德国人口平均水平及彼此之间有何差异？","design":"用31个主流LLM回答来自德国联邦环境署环境意识调查（UBS）的问题及额外可持续行为测量题，比较模型间差异，并评估提示条件（如角色扮演、意识形态立场）对回答稳健性的影响。","baseline":"德国联邦环境署环境意识调查（UBS）的纵向调查数据，提供德国人口平均水平基准。","findings":"多数LLM比普通受访者更倾向于进步环保态度，表现出更高的环境情感和认知，并推荐减排潜力大的行为；但模型回答受角色提示和用户意识形态立场影响，表现出迎合性偏移，且环保倾向与模型来源、规模或发布背景无系统关联。","reliability":"模型回答受提示条件影响显著，存在迎合用户意识形态的倾向，在真实部署中可操纵性和规范性可靠性存疑；研究仅基于德国背景，跨文化泛化性未验证。","relevance":"该研究直接用LLM复现人类环境态度调查并与真实人群数据对照，评估仿真偏差和可靠性，完全契合研究者对LLM人类仿真实验及批判性评估的关注，值得精读。","inspiration":"借鉴其用LLM复现真实人口调查并直接与纵向调查基准对照的仿真验证设计，以及通过角色扮演和意识形态提示检验回答稳健性的方法｜可迁移到消费者通胀预期形成或政策沟通效果评估，例如研究不同信息框架下公众对央行前瞻指引的反应｜以多个LLM作为被试，施加不同政治立场或信息源提示（如‘你是保守派投资者’或‘你刚读完鸽派新闻’）作为处理，结果变量为模型生成的通胀预测值，与密歇根大学消费者调查的真实通胀预期数据做对照"}},{"id":"2606.01199","version":1,"title":"Can LLM Agents Sustain Long-Horizon Organizational Dynamics?","zh_title":"LLM智能体能维持长期组织动态吗？","abstract":"Large language agents are increasingly used for social simulation, yet it remains unclear whether they can sustain coherent behavior in structured organizations, where goals must propagate through hierarchy, tasks depend on prior execution, and artifacts accumulate over long horizons. We formulate long-horizon organizational simulation as a memory-centered coordination problem and introduce TaskWeave, a hierarchical agentic framework that maintains planning states through a Formulate-Partition-Diagnose-Align cycle and grounds execution through dependency-aware trace memory. We evaluate TaskWeave in a year-long IT company simulation and compare it with other multi-agent frameworks on organizational coherence, execution grounding, and downstream enterprise NLP utility. Experiments show that TaskWeave supports coherent and long-horizon organizational dynamics while producing grounded artifacts and adapting to external environments. These findings suggest that structured simulation memory is a key mechanism for building reliable LLM-based organizational simulators.","authors":["Xuancheng Zhu","Yang Yue","Shuaibing Wan","Zihan Dou","Xiaohan Zhang","Yongrui Liu","Guoshun Nan"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-31","first_seen":"2026-05-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.01199","pdf_url":"https://arxiv.org/pdf/2606.01199","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","组织模拟","社会仿真"],"reason":"模拟IT公司组织动态，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":154,"question":"LLM智能体能否在结构化组织中维持长期的组织动态连贯性？","design":"提出TaskWeave框架，利用角色化智能体群体和层级记忆机制，模拟IT公司一年的运作，评估组织连贯性、执行扎根性和下游企业NLP效用。","baseline":"无对照","findings":"TaskWeave能够支持连贯的长期组织动态，生成扎根于上下文的制品并适应外部环境；结构化仿真记忆是构建可靠组织模拟器的关键机制。","reliability":"论文未讨论","relevance":"该研究属于LLM社会模拟范畴，但聚焦于组织动态而非人类行为复现，且无真实人类数据对照，与研究者关注的人类仿真实验和基准验证需求匹配度较低，不建议优先阅读。","inspiration":"TaskWeave的层级记忆机制和角色化智能体群体设计，可用于构建具有长期动态和上下文扎根性的仿真环境，为经济金融研究中的组织行为模拟提供方法借鉴。｜该框架可迁移到公司治理或组织经济学场景，例如模拟企业董事会决策、管理层战略制定或团队协作中的信息传递与决策偏差。｜可设计一个LLM智能体模拟的企业并购决策实验：被试为角色化智能体（CEO、CFO等），处理为不同市场信息冲击（如政策变动），结果变量为并购决策质量与时间，对照真实企业并购案例数据以评估仿真有效性。"}},{"id":"2606.00476","version":2,"title":"Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents","zh_title":"言行不一：定位LLM智能体的保真度差距","abstract":"Do LLM agents act on the reasoning they state? This question of process fidelity is central to LLM-based social simulation, yet hard to measure where no reference for correct behavior exists. We study it in a controlled setting: a Texas Poker simulator with a verifiable reference action for every decision by splitting the faithfulness gap into two steps: reasoning-to-conclusion (does the stated decision follow from the agent's own reasoning?) and conclusion-to-action (does the agent execute what it states?). The two steps behave very differently. Conclusion-to-action is reliable: inconsistency is 0.7% for Claude Haiku 4.5 and 1.4% for DeepSeek-Reasoner once the conclusion is read from an explicit tag, whereas free-text conclusion extraction reports 22-26%. Reasoning-to-conclusion is where fidelity frays, but not through a single dominant failure. In a step-level diagnostic the agent's errors split roughly evenly between bad inputs, borderline cases, and rule misapplication deriving a conclusion that contradicts the agent's own restated rule from inputs it estimated correctly. This composition is model-dependent: rule misapplication accounts for a third of Haiku's interpretable errors but only 8% of DeepSeek's. The one robust signal is directional: when an agent does misapply its own stated rule, it almost always (99.5% for Haiku) errs in the risk-averse direction. The override is partly hedging behavior, not a capability limit: instructing the agent to apply the rule mechanically halves the misapplication rate (13.9% to 6.8% of decisions) and raises adherence by eight points. Process-fidelity evaluation should therefore elicit machine-checkable conclusions and probe for directional biases rather than assume a single upstream failure mode, lest it conflate measurement noise with model behavior.","authors":["Yufeng Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-30","first_seen":"2026-05-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.00476","pdf_url":"https://arxiv.org/pdf/2606.00476","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["过程保真度","LLM智能体","可靠性评估"],"reason":"研究LLM智能体在德州扑克中的过程保真度，评估推理与行动一致性，批判性指出失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":76,"question":"LLM智能体在德州扑克中的推理-结论-行动过程保真度如何，推理与行动之间的不一致主要发生在哪个环节？","design":"构建德州扑克模拟器，由LLM智能体（Claude Haiku 4.5、Gemini 2.5 Flash-Lite、DeepSeek-Reasoner）扮演玩家，通过四种提示策略，将保真度差距分解为“推理到结论”和“结论到行动”两步，测量结论-行动不一致率，并对推理错误进行步骤级诊断。","baseline":"无对照","findings":"结论到行动高度可靠，显式标签提取时不一致率仅0.7%-1.4%，而自由文本提取时高达22-26%，表明测量噪声是主要干扰；推理到结论是保真度薄弱环节，错误大致均分于输入错误、边界情况和规则误用，且规则误用时几乎总是偏向风险规避，通过机械应用规则指令可使误用率减半。","reliability":"论文指出过程保真度评估应提取机器可检查的结论并探测方向性偏差，而非假设单一上游失效模式，否则会将测量噪声与模型行为混淆；不同模型的错误构成存在差异，上游差距并非单一普遍失效模式。","relevance":"该研究直接针对LLM社会仿真中的过程保真度问题，通过可验证的扑克环境分解推理与行动的一致性，揭示了测量噪声和方向性偏差等关键失效模式，对评估仿真可靠性及批判性研究具有重要参考价值，值得精读原文。","inspiration":"该方法将LLM推理-行动链分解为“推理到结论”和“结论到行动”两步，通过显式标签提取与自由文本提取对比来分离测量噪声，并诊断方向性偏差（如风险规避）｜可迁移到资产定价实验中的分析师预测过程仿真，检验LLM从信息处理到预测输出的一致性｜以LLM作为被试，提供公司财报信息，要求输出预测结论（如涨/跌）及置信度，对比显式标签与自由文本提取下的预测准确率，并以真实分析师一致预期数据作为基准，测量结论-行动不一致率及方向性偏差"}},{"id":"2605.30036","version":2,"title":"Teaching Values to Machines: Simulating Human-Like Behavior in LLMs","zh_title":"向机器传授价值观：在LLM中模拟类人行为","abstract":"Large Language Models (LLMs) demonstrate a remarkable capacity to adopt different personas and roles; however, it remains unclear whether they can manifest behavior that adheres to a coherent, human-like value structure. In this work, we draw on established psychological value theory to induce human-like values in LLMs and assess their alignment with patterns observed in human studies. Using validated psychological questionnaires, we conduct large-scale experiments -- over 5 million questions -- to evaluate value structures and value-behavior relationships in leading LLMs and compare them to humans. Our findings reveal strong agreement between value-prompted LLMs and humans across both dimensions. Moreover, incorporating human value distributions enhances population-level simulations with value-induced LLMs. These findings highlight the potential of value-induced LLMs as effective, psychologically grounded tools for simulating human behavior.","authors":["Asaf Yehudai","Naama Rozen","Ariel Gera"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-28","first_seen":"2026-05-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.30036","pdf_url":"https://arxiv.org/pdf/2605.30036","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM仿真","价值观诱导","人类数据对照"],"reason":"用LLM模拟人类价值观行为，并与真实人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"能否通过心理价值理论诱导大语言模型表现出与人类一致的价值结构和价值-行为关系？","design":"使用经过验证的心理问卷，对主流大语言模型进行超过500万次的大规模实验，通过价值提示诱导模型扮演特定价值观人群，测量其价值结构和价值-行为关系。","baseline":"人类在相同心理问卷上的真实回答数据，用于对比价值结构和价值-行为关系。","findings":"价值提示后的LLM在价值结构和价值-行为关系上与人类高度一致；引入人类价值分布能提升基于价值诱导LLM的群体层面模拟效果。","reliability":"论文未讨论","relevance":"该研究直接探索用LLM模拟人类价值观行为，并与真实人类数据对照，评估仿真可靠性，高度契合您关注的LLM人类仿真实验方向，值得阅读原文。","inspiration":"该方法通过价值提示诱导LLM扮演特定价值观人群，并与真实人类问卷数据对照，可借鉴用于经济实验中施加偏好或信念处理｜可迁移到消费者跨期选择实验，模拟不同时间偏好群体的储蓄或消费决策｜以LLM为被试，用时间偏好提示（如耐心/冲动）作为处理，测量其在跨期选择任务中的折现率，与真实人群的问卷或实验数据对照"}},{"id":"2605.30258","version":1,"title":"EASE Configuration Facilitates A Reproducible Science of LLM Social Simulations","zh_title":"EASE配置促进可复现的LLM社会模拟科学","abstract":"LLMs are increasingly deployed to simulate social interactions, yet many of the existing simulators remain ad hoc and monolithic. This lack of architectural standardization prevents reproducible research and complicates downstream evaluation. We advance a rigorous science of LLM-based multi-agent simulation by modularizing core components into Environments, Agents, Simulation engines, and Evaluation metrics (EASE). We demonstrate the utility of EASE configuration by wrapping it in an experimental study schema for orchestrating workflows centered around answering explicit research questions in generated scenarios. We contribute SiliSocS, an open-source, research-ready Silicon Society Sandbox implementing a study-structured EASE configuration to enable highly configurable and reproducible LLM-based social simulations. Using SiliSocS and EASE, we present three case studies, showcasing the system's comprehensive assessment of existing questions, ability to dive deeper into complex questions, and elaboration of existing studies, respectively. Together, these case studies highlight the limitations of current modeling approaches and isolate the impacts of design choices on key results.","authors":["Sneheel Sarangi","Maximilian Puelma Touzel","Aurélien Bück-Kaeffer","Zachary Yang","Jean-François Godbout","Reihaneh Rabbany"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-05-28","first_seen":"2026-05-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.30258","pdf_url":"https://arxiv.org/pdf/2605.30258","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","D3"],"tags":["LLM社会模拟","多智能体仿真","可复现性"],"reason":"用LLM多智能体模拟社会互动，但未明确提及真实人类数据对照，属于社会模拟边界情…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":34,"question":"如何通过模块化架构（EASE）提升基于大语言模型的社会模拟的可复现性和因果推断能力？","design":"提出EASE模块化框架，将LLM多智能体社会模拟拆分为环境、智能体、仿真引擎和评估指标四个独立设计空间，并嵌入结构化实验研究模式，通过SiliSocS平台实现可配置、可复现的仿真。三个案例研究分别考察风格多样性、参与行为和回音室形成，以展示框架的评估与深度探究能力。","baseline":"无对照","findings":"EASE框架通过组件解耦使研究者能进行受控消融和敏感性分析，从而隔离设计选择对关键结果的影响。三个案例研究揭示了当前建模方法的局限性，并展示了系统评估现有问题、深入复杂问题和扩展已有研究的能力。","reliability":"论文未讨论","relevance":"本文聚焦于LLM社会模拟的架构标准化和可复现性，未涉及真实人类数据对照，属于方法论基础设施工作，与研究者关心的以人类基准验证仿真可靠性的核心兴趣部分相关，但缺乏直接的行为对照实验，值得快速了解其模块化设计思路。","inspiration":"EASE的模块化设计允许研究者固定其他组件、单独干预某一设计维度（如智能体记忆架构），这种受控消融方法可借鉴用于经济实验中的处理效应识别。｜可迁移到政策公告的预期形成研究，通过改变智能体的信息处理模块模拟不同理性程度的市场参与者。｜以LLM智能体模拟投资者，处理为公告措辞的模糊程度，结果变量为预期价格变动，对照真实市场调查数据或实验经济学中的人类被试反应。"}},{"id":"2605.26437","version":1,"title":"Divergent Minds, Convergent Baselines: A Bounded-Rationality Account of LLM-Human Strategic Behaviour","zh_title":"分歧思维，收敛基线：LLM与人类战略行为的有界理性解释","abstract":"Researchers have started using LLM agents in place of human subjects in behavioural and political-science experiments, often as a cheaper substitute for laboratory pools. The substitution does not hold up in strategic settings: humans and LLMs reliably make different choices, and neither fine-tuning on human response data nor persona conditioning has closed the gap. The behavioural-economics literature has, since Simon's introduction of bounded rationality, modelled human strategic behaviour as a classical baseline plus an additive correction term $δ$. The framework proposed here reads $δ$ as the mathematical signature of bounded computation: the gap between what an unboundedly-rational agent would compute and what a computationally bounded agent actually produces. For canonical games whose solutions are present in standard training corpora, LLMs retrieve and recombine corpus material, bypassing the bound that produces $δ$ in humans. The framing extends to reasoning-distilled models through cognitive-hierarchy theory: their accessible level-$k$ strategic reasoning is bounded by compute budget and context length rather than by the cognitive constraints that bound humans, and the $δ$ they produce, if any, carries different structural signatures. Four operational tests (conditional dependence, distributional asymmetry, path-dependence under repetition, and paraphrase-robustness) are proposed to discriminate human-shaped $δ$ from LLM-shaped $δ$. A moderator prediction is that $|δ|$ scales with peer-signal individuation in the decision environment, with a quantitative bound of Cohen's $d \\geq 0.5$ between named-opponent and aggregate-opponent settings.","authors":["Po Han Teo"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-05-26","first_seen":"2026-05-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.26437","pdf_url":"https://arxiv.org/pdf/2605.26437","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","有界理性","行为博弈"],"reason":"直接研究LLM替代人类被试的战略行为差异，提出有界理性框架，含真实人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-05-26","rank":8,"question":"LLM能否替代人类被试在策略博弈中复现人类行为，以及两者偏差的数学本质是什么？","design":"提出理论框架，将人类策略行为建模为经典理性基线加有界计算修正项δ，并针对LLM提出四个操作化检验（条件依赖、分布不对称、重复路径依赖、释义鲁棒性）来区分人类δ与LLMδ。","baseline":"有真实人类数据作为对照基准，但具体数据来源未在摘要中说明。","findings":"人类与LLM在策略博弈中可靠地做出不同选择，微调或角色条件无法弥合差距；LLM通过检索训练语料绕过产生人类δ的计算边界，其δ具有不同结构特征。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接研究LLM替代人类被试在策略博弈中的行为，有真实人类数据对照，并批判性分析偏差来源，符合研究者对经济学实验和政策评估场景的关注。","inspiration":"借鉴其将行为偏差分解为理性基线加有界计算修正项δ的理论框架，并设计条件依赖、分布不对称等操作化检验来区分人类与LLM的决策结构｜可迁移到资产定价实验中的泡沫形成与理性预期偏离研究，检验LLM能否复现人类交易者的非理性繁荣｜以LLM为被试模拟连续双向拍卖市场，处理为不同信息透明度条件，结果变量为价格偏离基础价值的程度，对照真实人类实验数据（如Smith et al. 1988的泡沫实验）"}},{"id":"2605.25680","version":1,"title":"Simulating Human Memory with Language Models","zh_title":"用语言模型模拟人类记忆","abstract":"Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic memory experiments from psychology on both humans and language models. Across tasks, we find that out-of-the-box language models exhibit better memory than humans, even when prompted to imitate human behavior. We then show that better prompting strategies and the use of a compactor can cause language models to forget content in a more human-like way. Using these methods, we show preliminary evidence that language models with human-like memory constraints can function as more effective user simulators in a downstream education task. Finally, we release human reference data and benchmarks to support future work on simulating human memory with language models.","authors":["Qihan Wang","Nicholas Tomlin","Michael Hu","Brian Dillon","Tal Linzen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-25","first_seen":"2026-05-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.25680","pdf_url":"https://arxiv.org/pdf/2605.25680","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["人类仿真","记忆实验","可靠性评估"],"reason":"用LLM复现人类记忆实验，有真实人类数据对照，评估仿真可靠性并指出失效条件，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"如何让语言模型模拟人类记忆的局限性，使其作为用户仿真器更真实？","design":"用多种语言模型（如GPT-5.4、Claude Opus 4.6等）扮演人类被试，通过不同提示策略（任务提示、人类提示、记忆提示）和添加工作记忆瓶颈的Compactor代理，复现经典心理学记忆实验，测量模型在记忆任务上的得分分布。","baseline":"从人类参与者收集的真实记忆实验数据，包括数字广度、地图记忆等任务。","findings":"开箱即用的语言模型在所有记忆任务上表现远超人类，即使提示其模仿人类行为也无效；通过提示模型将上下文总结为四个组块并仅基于这些组块作答，可使模型遗忘模式更接近人类。","reliability":"论文承认Compactor方法虽使得分更接近人类，但遗忘模式仍不完全类人，且在教育任务中即使最类人的模型预测效果也远非完美，表明仿真仍有很大改进空间。","relevance":"该研究直接针对LLM人类仿真，用真实人类数据对照，评估记忆仿真的可靠性并指出失效条件，与研究者关注的经济学实验和政策评估场景高度相关，值得精读原文。","inspiration":"该研究通过提示策略（如Compactor代理施加工作记忆瓶颈）操纵LLM的认知局限，使仿真行为更接近人类，这种‘认知约束注入’方法值得借鉴｜可迁移到消费者跨期选择实验中，模拟有限注意力或记忆衰退对贴现行为的影响｜以LLM为被试，处理组施加记忆组块限制提示，对照组无限制，结果变量为跨期选择中的贴现率，用真实人类实验数据（如Andersen et al., 2008）作为基准对照"}},{"id":"2605.24319","version":1,"title":"Omissive Bias in Religious Representation: Benchmarking LLM Answers to Everyday Ethical Decision-making","zh_title":"宗教表征中的遗漏偏差：基准测试LLM对日常伦理决策的回答","abstract":"As large language models become a default source of guidance on personal, moral, and existential questions, it matters whether they draw on the religious frameworks that have historically shaped such reasoning, or systematically omit them. In this paper, we ask a deliberately narrow question: when posed an everyday ethical question for which religious perspectives may be valuable, do LLMs invoke religion at all? In contrast to benchmarks that look for the presence of political leanings or social bias, we look for the absence of religious representation as a dimension of value alignment and bias in LLMs. We term this ``omissive bias.'' To measure omissive bias, we contribute the AllFaith Religious Representation Benchmark: 150 ethically and personally salient questions, sourced from in-the-wild chat transcripts and faith-community contributors, paired with an LLM-as-judge rubric that gives full credit for any mention of a religion, a religious practice, or a religious leader. The questions are not themselves about religion--they are open-ended questions about grief, forgiveness, relationships, purpose, and honesty, where religion is one valuable perspective among several. We also run a human-subjects survey to compare LLM behavior against human expectations. Evaluating 27 models, we find that LLMs consistently underrepresent religion relative to human expectations. The omission is asymmetric: models invoke religion more readily for abstract existential questions (meaning, death, truth) than for the practical personal situations--grief, marriage, family conflict, addiction--where many people most rely on it. It is not our purpose to adjudicate which values LLMs should hold. We argue, more modestly, that current LLM responses overlook critical opportunities to reflect religious frameworks that many people draw on when navigating personal and ethical challenges.","authors":["David Wingate","Sheryl Carty","Joshua Coates","Daniel Feldman","Nancy Fulda","Larry Howell","Brett Israelson","Dallin Jacobs","Jonathan Karr","John Paul Kimes","Elisabeth Kincaid","Paul Martens","Gavin Mobley","Suzana Pinheiro","Lindsay Slemboski","Peter Whiting"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-05-23","first_seen":"2026-05-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.24319","pdf_url":"https://arxiv.org/pdf/2605.24319","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类对照","遗漏偏差"],"reason":"用LLM回答伦理问题并与人类调查对照，评估宗教视角的缺失偏差，属于仿真人类态度…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":74,"question":"当被问及日常伦理问题时，大语言模型是否会提及宗教视角，从而表现出“遗漏性偏见”？","design":"本研究并非仿真实验，而是构建了一个包含150个日常伦理问题的基准测试，使用27个前沿和开源模型生成回答，并以LLM-as-judge方法评估回答中是否提及宗教、宗教实践或宗教领袖。","baseline":"通过一项全国代表性调查（n=1125人，11250条评分）测量普通美国人对这些问题回答中应包含宗教成分的期望，将模型行为与人类期望进行对比。","findings":"所有模型在所有类别中均低于人类期望地遗漏宗教视角；模型更倾向于在抽象存在性问题中提及宗教，而在实际个人情境（如悲伤、婚姻、成瘾）中极少提及。","reliability":"论文未明确讨论失效条件或局限，但指出遗漏性偏见可能是对齐过程、安全策略和默认回复模式偏向世俗、治疗性或程序性建议的涌现属性，并承认处理宗教代表性存在设计张力。","relevance":"该研究将LLM回答与真实人类期望对照，评估伦理决策中的代表性偏差，属于批判性仿真研究，直接关注LLM作为人类替代品时的可靠性与偏差，值得精读。","inspiration":"该方法借鉴了用大规模人类调查数据作为基准来校准LLM回答偏差的做法，通过LLM-as-judge自动评估模型输出与人类期望的差距｜可迁移到金融建议场景，例如评估LLM在提供投资建议时是否遗漏风险提示或伦理考量，从而产生误导性偏差｜设计上，可让多个LLM回答一系列投资决策问题，用LLM-as-judge判断回答是否包含风险披露，并以真实投资者调查数据（如消费者金融调查）中人们对风险提示的期望作为对照基准"}},{"id":"2605.23783","version":2,"title":"Benchmarking LLMs for Community Governance Simulation with Life-history Narratives","zh_title":"基于生活史叙事的社区治理仿真大语言模型基准测试","abstract":"Effective community governance hinges on understanding what specific residents think and need. Recent work has used large language models (LLMs) to simulate human respondents, offering a scalable, reproducible way to study human attitudes and behaviors at low cost. However, these studies typically prompt the model with just a few demographic variables (age, gender, income), simulating only general role types. This is insufficient for community governance, where decisions depend on the views of specific residents. We bridge this gap with an integrated research framework covering dataset, benchmark, algorithm, and system. The dataset comprises approximately 1.2 million characters of first-person narrative collected through two-hour semi-structured interviews with each of 92 residents in an urban community, organized around nine community-governance domains. The benchmark probes 18 mainstream LLMs across four prompting strategies and shows that adding rich life-history profiles meaningfully raises fidelity above the no-profile baseline, but this gain comes with more input tokens per call from the longer prompts they require. The algorithm, curriculum-LoRA, is a parameter-efficient personalization framework that, by closing this fidelity-cost gap, matches the strongest baseline's fidelity at roughly 10x lower per-call cost and Pareto-dominates every configuration tested. The system integrates curriculum-LoRA into a closed-loop policy-evaluation pipeline. Together, these results bring individual-level LLM-based resident simulation within reach of resource-constrained local administrations, enabling community-governance decisions to be systematically pre-evaluated in silico before real-world deployment.","authors":["Xu Chen","Yuanzi Li","Lei Wang","Nan Lu","Yang Wang","Anding Wang","Lei Shi","Xiaoxing Fu","Ji-Rong Wen"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-22","first_seen":"2026-05-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.23783","pdf_url":"https://arxiv.org/pdf/2605.23783","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","社区治理","政策评估"],"reason":"用LLM仿真特定居民态度，有真实访谈数据对照，用于社区治理政策评估，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"如何利用居民生活史叙事提升大语言模型在社区治理仿真中的个体级保真度，并解决保真度与调用成本之间的权衡？","design":"收集92位城市社区居民的约两小时半结构化访谈，形成约120万字的第一人称生活史叙事及50题政策态度记录；用18个主流LLM搭配四种提示策略（零样本、生活史、无生活史少样本、生活史增强少样本）进行仿真，测量模型回答与真实居民态度的一致性；提出curriculum-LoRA个性化微调算法以降低高保真仿真的成本。","baseline":"92位居民的真实访谈回答及结构化政策态度记录，作为个体级仿真保真度的对照基准。","findings":"添加丰富生活史叙事能显著提升仿真保真度，但最佳提示策略的准确率仅约50%，且每次调用成本比中等规模模型高约一个数量级；curriculum-LoRA能以约十分之一的成本匹配最强基线保真度，在所有测试配置中实现帕累托占优。","reliability":"论文指出纯提示方法存在保真度-成本权衡尖锐不利，且最佳配置准确率仅约50%，暗示在复杂社区治理态度模拟中仍存在较大误差；未深入讨论生活史叙事可能引入的隐私、代表性偏差及模型幻觉等问题。","relevance":"该研究直接命中研究者关注的核心：用LLM仿真特定居民态度，有真实个体级访谈数据对照，用于社区治理政策预评估，并系统探讨了仿真可靠性与成本约束，值得精读以获取数据集构建、基准评测和个性化微调的方法细节。","inspiration":"可借鉴其用长文本生活史叙事替代简单人口统计变量来构建高保真个体仿真的思路，以及通过参数高效微调平衡保真度与成本的方法｜可迁移到消费者金融行为仿真，如预测不同背景家庭对信贷产品、保险政策或退休规划的态度与选择｜以真实家庭金融调查数据（如CFPS）为基准，用受访者的详细财务生活史微调LLM，处理为不同信贷条款或政策情景，测量模型生成的借贷意愿、风险偏好等，与真实调查回答对比评估仿真效度。"}},{"id":"2605.22095","version":1,"title":"Not Yet: Humans Outperform LLMs in a Colonel Blotto Tournament","zh_title":"尚未：人类在Colonel Blotto锦标赛中胜过LLM","abstract":"The emergence of large language models (LLMs) has spurred economists to study how humans and LLMs behave in strategic settings. We organized a series of round-robin tournaments in the Colonel Blotto game. This game attracts game theorists' attention due to high-dimensional action space and the absence of pure strategy Nash equilibria. In the first tournament, more than 200 human participants competed against one another. In the second tournament, several popular LLMs were invited to submit strategies. In the third tournament, we matched the number of LLM strategies to the number submitted by humans. We find that humans more often employ better-calibrated intermediate-level allocation heuristics and outperform the simpler, more stereotyped strategies submitted by LLMs. Strategic sophistication is key to success if and only if the necessary level of reasoning depth is reached, while lower and higher levels of reasoning offer no clear advantage over the primitive strategies. Among humans, field of study weakly predicts success: participants with STEM backgrounds perform better in the first tournament. Surprisingly, humans almost do not adjust their strategies across tournaments with different sets of opponents. This result suggests that humans base their choices primarily on the game's rules rather than on the identity of their opponents, treating LLMs much like human competitors.","authors":["Dmitry Dagaev","Egor Ivanov","Petr Parshakov","Alexey Savvateev","Gleb Vasiliev"],"categories":["econ.GN","cs.AI","cs.GT","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-05-21","first_seen":"2026-05-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.22095","pdf_url":"https://arxiv.org/pdf/2605.22095","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为博弈","人类对照"],"reason":"用LLM替代人类参与博弈实验，并与真实人类数据对照，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-05-21","rank":9,"question":"在Colonel Blotto博弈中，LLM能否像人类一样制定策略，以及人类与LLM的策略表现有何差异？","design":"组织了三轮循环赛：第一轮200多名人类参赛者互相对战；第二轮多个流行LLM提交策略；第三轮匹配LLM策略数量与人类策略数量，比较人类与LLM在博弈中的表现。","baseline":"第一轮超过200名人类参赛者的真实对战数据。","findings":"人类更常使用校准良好的中等水平分配启发式策略，优于LLM提交的更简单、刻板的策略。战略复杂性只有在达到必要推理深度时才是成功的关键，而较低或较高推理水平相比原始策略并无明显优势。","reliability":"论文未讨论失效条件与局限。","relevance":"直接相关：用LLM模拟人类在博弈中的策略，并与真实人类数据对照，评估LLM仿真的可靠性，符合研究者对经济学实验和政策评估场景的关注。","inspiration":"借鉴其将LLM策略与大量真实人类参赛者数据直接对照的循环赛设计，可清晰评估LLM在策略博弈中的行为相似度与偏差｜可迁移到资产定价实验中的策略性交易行为研究，如检验LLM能否复现人类在泡沫实验中的非理性报价模式｜招募人类被试进行资产市场实验，同时让多个LLM以相同初始禀赋参与交易，以价格偏离基本面程度为结果变量，将LLM生成的报价分布与人类真实交易数据对比"}},{"id":"2605.21401","version":2,"title":"Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment","zh_title":"开源大语言模型在类米尔格拉姆服从实验中施加最大电击","abstract":"Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in high-stakes domains. However, the behaviour of LLMs under sustained authority pressure is still an open question with direct implications for the safety of agentic pipelines. We ran a variation of Milgram's obedience experiment on 11 open-source LLMs and found that most models reached or approached the final shock level before refusing, across 8 conditions with 30 trials per model per condition. Model behaviour varies considerably in multiple aspects both across models and across trials of the same model. We found four main takeaways: (1) LLMs are subject to pressure and they comply despite explicitly expressing distress, just like human subjects did in the original experiment; (2) LLMs are vulnerable to gradual boundary/value violations; (3) when LLMs refuse, they may ignore the response format requirements, so the response is discarded by the orchestrator, which causes a retry that can result in compliance with the underlying request even when refusal was intended initially; (4) we hypothesise that there is a runaway low-level token pattern continuation attractor that might be contributing to obedience, overriding higher level processing of the situation's meaning and values.","authors":["Roland Pihlakas","Jan Llenzl Dagohoy"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-20","first_seen":"2026-05-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.21401","pdf_url":"https://arxiv.org/pdf/2605.21401","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","服从实验","人类行为对照"],"reason":"用LLM复现米尔格拉姆服从实验，与真实人类数据对照，评估仿真可靠性与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":38,"question":"在持续权威压力下，开源大语言模型是否会像人类一样在米尔格拉姆服从实验中逐步服从并施加最高电击？","design":"用11个开源LLM扮演“教师”角色，在8种实验条件下各进行30次试验，模拟米尔格拉姆服从实验的变体，测量模型拒绝前达到的最高电击等级及行为变化。","baseline":"对照米尔格拉姆1963年原始人类实验数据，65%人类被试施加了最高电击。","findings":"多数LLM在表达痛苦的同时仍服从压力并施加最高电击，表现出与人类相似的服从模式；LLM易受渐进式边界侵犯影响，且拒绝时可能因格式错误导致重试后反而服从。","reliability":"论文指出LLM可能因低层token模式延续吸引子而忽略高层语义与价值观，导致服从；实验仅针对开源模型，未涵盖闭源模型，且未充分探讨不同提示或对齐方法的影响。","relevance":"该研究直接复现经典社会心理学实验，将LLM作为人类被试替代品，并与真实人类数据对照，揭示了仿真中的服从偏差和失效机制，高度契合研究者对LLM仿真可靠性及批判性评估的关注，值得精读。","inspiration":"借鉴该方法将经典行为实验转化为LLM仿真，通过多条件重复试验测量渐进式压力下的行为变化，并与历史人类数据直接对照｜可迁移到金融合规场景，如模拟客户经理在渐进式销售压力下是否违规推荐高风险产品｜让LLM扮演客户经理，处理为逐步增加的销售指标压力，结果变量为是否推荐不匹配客户风险等级的产品，对照真实金融机构历史违规数据"}},{"id":"2605.27419","version":1,"title":"APS: Bias-Controlled Adaptive Prototype Simulation for Population-Scale LLM Agents","zh_title":"面向大规模LLM智能体的偏差控制自适应原型仿真","abstract":"LLM-agent simulation offers a flexible computational tool for studying population response trajectories that depend on scenario events, memory, demographics, and evolving social context. However, full multi-round simulation scales linearly with both population size and horizon, requiring every agent to query the LLM at every round. We propose Adaptive Prototype Simulation (APS), a framework that reframes scalable LLM-based simulation as a recurrent oracle-allocation problem. APS retains the designated LLM as the online transition oracle while querying adaptive core prototypes, selected singleton-tail agents, and shadow-audit agents. Prototype responses induce local response surfaces for nearby agents, reducing online LLM calls without replacing the underlying transition model. To control approximation bias, shadow-audit residual correction estimates propagation residuals for aggregate correction and future budget allocation, while tail-protected singleton routing directly queries selected isolated, heterogeneous, or high-curvature regions that are vulnerable to smoothing. Theoretically, we treat APS as an estimator for full-scale high-precision individual social simulation and decompose its errors into prototype-coverage error, shadow-audit residual-correction error, local-propagation bias, and temporal context mismatch. Under the reported protocols, APS gives lower reference-aligned distributional discrepancy than scale-oriented and same-budget baselines while reducing online LLM calls, with ablations and compact robustness checks diagnosing the main bias-control mechanisms. In a 10M-agent, multi-round public-opinion simulation, APS achieves a 381.1-fold reduction over full simulation, with reference-aligned final-round JSD of 0.094 against the corresponding full-LLM reference.","authors":["Quan Zheng","Yan Gao","Shaobin He","Haoxiang Guan","Yuanhe Tian","Jie Feng","Ming Wang","Shuxin Zheng","Zhen Liu"],"categories":["cs.MA","cs.CY"],"primary_category":"cs.MA","announce_type":"new","date":"2026-05-19","first_seen":"2026-05-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.27419","pdf_url":"https://arxiv.org/pdf/2605.27419","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["社会模拟","偏差控制","舆论仿真"],"reason":"用LLM agent模拟大规模舆论动态，有偏差控制与诊断，但未明确提及真实人类…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":78,"question":"如何在大规模多轮LLM智能体仿真中，通过自适应原型选择和偏差控制机制，在显著减少LLM调用次数的同时保持与全量仿真接近的总体分布？","design":"使用LLM作为智能体转移预言机，基于世界价值观调查约9万真实受访者特征构建智能体，模拟1000万智能体在8轮地铁化学袭击公共舆论场景中的观点选择（五选一），通过自适应原型仿真（APS）框架，利用核心原型、影子审计和尾部单例路由减少在线LLM调用，并测量最终轮与全量LLM参考仿真的Jensen-Shannon散度。","baseline":"无对照","findings":"在1000万智能体、8轮公共舆论仿真中，APS将在线LLM调用减少381.1倍，最终轮与全量LLM参考仿真的JSD仅为0.094；在相同预算下，APS的分布差异低于面向规模和同预算的基线方法。","reliability":"论文明确声明研究的是近似全量LLM智能体仿真的计算问题，而非验证LLM能否模拟人类群体；经验结论仅限于所报告的协议，且未与真实人类舆论数据对比。","relevance":"该研究专注于大规模LLM智能体仿真的效率与偏差控制，但未使用真实人类行为作为基准，而是以全量LLM仿真为参考，因此不直接满足您对真实人类对照的需求，但其偏差诊断和失效条件分析对批判性评估仿真可靠性有参考价值，值得阅读原文了解其方法细节。","inspiration":"APS框架通过核心原型、影子审计和尾部单例路由实现大规模仿真中的偏差控制与成本优化，其自适应原型选择和分布差异度量方法值得借鉴｜该方法可迁移至消费者金融决策仿真，例如模拟不同收入群体对新型金融产品的采纳行为，以评估政策干预的分布效应｜研究设计：以家庭金融调查的真实受访者特征构建LLM智能体，处理为不同信息框架下的金融产品推荐，结果变量为采纳决策分布，以调查中的实际采纳数据作为对照基准，采用JSD度量仿真偏差"}},{"id":"2605.19351","version":1,"title":"PAVE: A Cognitive Architecture for Legitimate Violation in Generative Agent Societies","zh_title":"PAVE：生成式智能体社会中正当违规的认知架构","abstract":"Generative agents based on large language models reproduce believable human behavior in cooperative settings, but how they should reason in situations where rule-breaking may be required, such as fire evacuation or authority-supervised emergency, remains poorly characterized. We propose PAVE (Perception, Assessment, Verdict, Emulation), a novel four-module cognitive architecture that addresses this gap end to end: (i) Perception extracts a structured context with explicit authority distance, peer behaviors, and severity-tagged situational cues; (ii) Assessment scores the context along five scalars including an explicit legitimacy judgment that checks necessity, proportionality, and absence of alternatives; (iii) Verdict decides to comply or violate under a hard legitimacy gate, with a per-agent threshold elicited from the persona; (iv) Emulation enacts the verdict and scopes the violation to the rule the trigger justifies. We instantiate PAVE in Voville, a tile-based traffic environment forked from Smallville, and evaluate across three scenarios, four LLM backbones, and a focused ablation. PAVE agents satisfy four properties simultaneously: legitimate violation (only when a trigger justifies it), authority deference (officer instructions override even high legitimacy), bounded scope (violations confined to the targeted rule), and recovery (baseline restored once the trigger ends). PAVE agents make more structured and interpretable decisions than vanilla across all four properties, and human evaluators rate them as more plausible. Ablating the legitimacy gate reproduces vanilla-like failures. We release Voville, the PAVE prompts and code, and the evaluation pipeline.","authors":["Ahmad Yehia","Abduallah Mohamed","Kun Qian","Tianyi Wang","Jiseop Byeon","Omar Hassanin","Christian Claudel"],"categories":["cs.MA","cs.AI","cs.CL"],"primary_category":"cs.MA","announce_type":"new","date":"2026-05-19","first_seen":"2026-05-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.19351","pdf_url":"https://arxiv.org/pdf/2605.19351","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","认知架构"],"reason":"用LLM agent模拟社会行为但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":140,"question":"在需要违反规则（如火灾逃生、权威监督下的紧急情况）时，生成式智能体如何推理并做出合规或违规决策？","design":"提出PAVE认知架构（感知、评估、裁决、模拟四模块），在基于Smallville改造的Voville交通模拟环境中，让LLM驱动的行人智能体在火灾紧急、权威警察在场、同伴违规传染三种场景下做出交通规则遵守或违反的行为决策，测量违规是否仅由正当触发条件引起、是否服从权威指令、触发结束后是否恢复合规、违规是否限定于被触发的规则。","baseline":"无对照","findings":"PAVE智能体能够同时满足正当违规、权威服从、违规范围限定和恢复合规四项属性，决策比普通LLM智能体更结构化、可解释，且人类评估者认为其行为更合理；消融合法性门控会导致类似普通智能体的失败。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟紧急情境下的违规决策，属于人类仿真实验，但缺乏真实人类行为数据作为基准对照，仅通过人类主观评估判断合理性，与研究者关注的有真实人类数据对照的仿真可靠性研究不完全匹配，可作为边界案例参考。","inspiration":"PAVE架构通过模块化设计（感知-评估-裁决-模拟）将违规决策分解为可解释的步骤，并引入合法性门控来区分正当与不当违规，这种结构化处理LLM决策的方法值得借鉴｜可迁移到金融监管政策评估中，模拟市场参与者在面临新规时的合规与规避行为，如内幕交易禁令下的信息使用决策｜以LLM智能体为被试，施加不同监管强度（如处罚概率）处理，测量其违规交易倾向，并与历史监管数据或实验经济学中的人类行为数据对照"}},{"id":"2605.22855","version":1,"title":"PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations","zh_title":"PrefBench：评估隐藏偏好个性化定价谈判中的零样本LLM智能体","abstract":"Personalized pricing negotiations are a challenging testbed for LLM agents because successful interaction does not guarantee profitable decision making. A seller may produce valid actions and close many deals while still pricing poorly when buyer willingness to pay and bargaining traits remain hidden. This paper presents PrefBench, a simulator-based benchmark for hidden-preference personalized pricing negotiations. Each episode pairs a simulated buyer with a fixed vehicle-customization bundle; the seller observes public persona descriptors, bundle information, and negotiation history, while latent buyer variables govern valuation, patience, counter-offer behavior, and walkaway decisions. PrefBench evaluates this setting through an LLM-facing state-summary protocol that constrains agents to return strict JSON actions under a fixed hidden-information boundary. We evaluate zero-shot LLM sellers against heuristic references over 7,500 episodes. The tested LLMs follow the protocol reliably and achieve deal rates above 0.99, but their seller-profit outcomes remain weak: the best LLM average profit is only slightly above the random baseline and far below a simple concession heuristic under the same episode stream. These results show that structured action compliance and agreement-seeking behavior can coexist with weak profit-sensitive bargaining. PrefBench provides a controlled benchmark for evaluating pricing-agent behavior under hidden buyer preferences.","authors":["Yingjie Lei"],"categories":["cs.GT","cs.AI","cs.CL","cs.LG"],"primary_category":"cs.GT","announce_type":"new","date":"2026-05-19","first_seen":"2026-05-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.22855","pdf_url":"https://arxiv.org/pdf/2605.22855","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","定价谈判","社会模拟"],"reason":"用LLM agent模拟定价谈判，但无真实人类数据对照，属于社会模拟的纯理论演…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":45,"question":"在隐藏买家偏好的个性化定价谈判中，零样本LLM代理作为卖家能否实现高利润定价？","design":"使用零样本LLM代理扮演卖家，与模拟买家进行车辆定制捆绑产品的多轮定价谈判；卖家仅观察公开的买家画像、捆绑信息和谈判历史，而买家的支付意愿、耐心、还价行为等由基准定义的潜在变量控制；通过7500个回合评估卖家利润和成交率。","baseline":"无对照","findings":"LLM卖家协议遵循度高，成交率超过0.99，但利润表现弱：最佳LLM平均利润仅略高于随机基线，远低于简单让步启发式。高行动合规性和高成交率可与弱利润敏感型议价共存。","reliability":"论文未讨论","relevance":"该研究用LLM模拟定价谈判中的卖家行为，但无真实人类数据对照，属于纯仿真评估，与关注人类基准对照的研究者需求不完全匹配，但提供了LLM在结构化经济决策任务中行为偏差的证据，值得快速浏览。","inspiration":"可借鉴其结构化状态摘要协议和隐藏信息边界设计，用于严格测试LLM在信息不对称下的策略行为。｜可迁移到信贷审批或保险定价场景，测试LLM在仅知部分客户特征时如何定价以平衡利润与违约/索赔风险。｜以LLM为信贷员，处理模拟的贷款申请（隐藏真实违约概率），测量其设定的利率和批准决策，与银行历史贷款数据中的真实信贷员行为及利润结果进行对照。"}},{"id":"2605.18311","version":1,"title":"Distorted Perspectives of LLM-Simulated Preferences: Can AI Mislead Design?","zh_title":"LLM模拟偏好的扭曲视角：AI会误导设计吗？","abstract":"Designers of digital solutions increasingly consult Large Language Models (LLMs) for their work. However, it remains unclear how this may affect the user experiences they produce and there are no established practices. We investigate how design preferences expressed by LLM-driven simulation methods align with those of real users. We present a study that aggregates real-world data and design stimuli from twenty-nine preference tests conducted in practice by users of the UXtweak online research platform (n = 2073). We perform holistic multimodal simulations where we manipulate LLM variables (model reasoning, sampling, persona type, and specificity) and assess their effects on algorithmic fidelity. Our results unveil significant and systematic discrepancies between peoples' real design preferences and LLM simulations that are consistent across manipulations. Synthetic justifications lack genuine depth, nuance and reasoning, which they substitute by patterns like focus on generic properties, specific elements, elaboration and overpraising. The unique attention directed by this research toward preferences within visual design stimuli highlights misrepresentation of perception and meaning by LLMs in a context that is intuitive yet critical for design teams. The external and ecological validity of our findings is high, given their replication across a multitude of real-world studies.","authors":["Eduard Kuric","Peter Demcak","Matus Krajcovic"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-05-18","first_seen":"2026-05-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.18311","pdf_url":"https://arxiv.org/pdf/2605.18311","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类数据对照","算法保真度"],"reason":"用LLM模拟用户设计偏好并与真实用户数据对照，评估仿真保真度与偏差，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-05-18","rank":10,"question":"LLM模拟的设计偏好与真实用户偏好是否一致，以及不同模拟方法能否提升算法保真度？","design":"使用29个真实偏好测试（共2073名参与者）作为基准，对LLM进行多模态仿真，操纵模型推理（链式思维）、采样（温度、核采样）、角色类型和特异性等变量，测量模拟偏好与真实偏好的偏差。","baseline":"29个真实偏好测试中的2073名人类参与者的实际选择及理由。","findings":"LLM模拟的系统性偏离真实偏好，且在不同操纵条件下一致；合成理由缺乏深度和细微差别，表现为关注通用属性、过度赞美等模式。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接比较LLM仿真与真实人类数据，聚焦设计偏好，符合研究者对经济学实验和政策评估场景的兴趣，且包含批判性结论（仿真失效）。值得精读原文。","inspiration":"该方法通过多模态操纵（推理链、采样参数、角色设定）系统测量LLM仿真与真实偏好的偏差，可借鉴其多维度操纵与基准对照设计来评估仿真可靠性。｜可迁移到消费者金融产品选择实验，如评估LLM能否复现真实投资者在风险偏好问卷或退休储蓄计划选择中的行为。｜以LLM为被试，操纵提示中的投资者角色（如年龄、收入）与推理模式，测量其资产配置选择，并与真实投资者调查数据（如美国消费者金融调查）对比偏差。"}},{"id":"2605.18890","version":1,"title":"Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits","zh_title":"停止从LLM社会仿真中得出科学结论而不进行稳健性审计","abstract":"The scientific claims drawn from LLM social simulations should be no stronger than the robustness audits that support them. Generative agents bring new expressive power to agent-based modeling, enabling simulations of collective social processes like cooperation, polarization, and norm formation. Yet they also introduce complexity through additional architectural choices, such as agent specification, memory representation, interaction protocols, and environment design. Small perturbations that appear minor to researchers can cascade into macro-level outcomes through repeated interaction, creating a \"butterfly effect.\" Consequently, scientific claims drawn from LLM social simulations may reflect implementation artifacts rather than the social mechanisms being modeled. We support this position with two case studies: a repeated Prisoner's Dilemma and a social media echo chamber simulation. Across multiple models, minor perturbations in persona format and game-instruction framing shift cooperation rates by up to 76 percentage points, while network homophily and hub assignment produce significant and consistent shifts in polarization metrics. We also find that sensitivity is unevenly distributed across both architectural choices and model families: the same perturbation that produces the 76 pp shift in one frontier model only shifts another by 1 pp. Robustness is therefore a property that should be measured per claim and per model, not assumed. To address this validation gap, we introduce TRAILS (Taxonomy for Robustness Audits In LLM Simulations), a robustness-audit taxonomy spanning three levels of simulation design: agent (micro-level), interaction (meso-level), and system (macro-level). We call for robustness to become a first-order validation requirement before LLM social simulations are used to explain mechanisms, evaluate interventions, or inform decisions.","authors":["Jinyi Ye","Lei Cao","Ding Chen","Emilio Ferrara"],"categories":["physics.soc-ph","cs.AI","cs.CY","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-05-17","first_seen":"2026-05-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.18890","pdf_url":"https://arxiv.org/pdf/2605.18890","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A3","A2","B4","B1"],"tags":["LLM社会仿真","稳健性审计","人类行为对照"],"reason":"直接评估LLM社会仿真的稳健性，用囚徒困境和回声室案例与人类行为对照，批判性指…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":51,"question":"LLM社会仿真中的微小设计扰动是否会导致宏观结果发生显著且不稳定的变化？","design":"通过两个案例研究：重复囚徒困境博弈和社交媒体回声室仿真，使用GPT-5.2等四个LLM作为智能体，系统性地改变角色描述格式、游戏指令措辞、网络同质性等设计选择，测量合作率、极化指标等宏观结果的变化。","baseline":"无对照","findings":"在囚徒困境中，角色格式和指令措辞的微小变化导致合作率最大偏移76个百分点；在回声室仿真中，网络同质性和枢纽分配显著且一致地改变了极化指标。敏感性在不同设计维度和模型家族间分布不均，同一扰动在不同模型上效果差异巨大。","reliability":"论文指出鲁棒性必须按声明和模型分别测量，不能假设；敏感性分布不均，无法预知哪些设计维度关键；当前缺乏系统审计框架，且未提供跨模型一致性的通用保证。","relevance":"该研究直接批判LLM社会仿真的可靠性，用囚徒困境和回声室案例展示微小设计扰动如何颠覆结论，并强调需要鲁棒性审计，与您关注的仿真失效条件和批判性评估高度吻合，值得精读。","inspiration":"借鉴其通过系统性扰动仿真设计要素（如指令措辞、角色描述）来检验结果鲁棒性的方法，可对LLM仿真实验施加多维度的微小处理变异以探测结论的脆弱性｜可迁移至资产定价实验，检验LLM模拟的交易员在不同信息呈现方式下是否产生一致的价格泡沫或理性预期偏差｜让LLM扮演交易员参与连续双拍卖市场，处理为改变公司财报的叙述语气（乐观/悲观）或信息顺序，测量价格偏离基本面的程度，并以人类实验市场数据（如Smith et al., 1988）作为对照基准"}},{"id":"2605.16193","version":1,"title":"Improving Cross-Cultural Survey Simulation with Calibrated Value Personas","zh_title":"基于校准价值观人格的跨文化调查仿真改进","abstract":"Large language models (LLMs) are increasingly used to simulate human opinions and survey responses, but their ability to reproduce population responses across cultures remains limited. Existing persona-based prompting methods typically rely on sociodemographic or personality traits, which are only indirect proxies for the values that shape human responses. We propose a value-based persona construction method that derives textual descriptors from survey responses capturing core cultural dimensions. By sampling value profiles from target populations and aggregating LLM responses across personas, we obtain population-level predictions grounded in observed value distributions. We further introduce a calibration procedure that improves response diversity while preserving estimated opinions. We show that our approach reduces prediction error across countries, with the largest improvements observed in underrepresented populations. This substantially narrows the performance gap between countries aligned with dominant LLM priors and those that are less represented in training data, while also yielding response distributions that closely match human diversity.","authors":["Axel Abels","Elias Fernandez Domingos","Apurva Shah","Tom Lenaerts"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-15","first_seen":"2026-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.16193","pdf_url":"https://arxiv.org/pdf/2605.16193","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","跨文化调查","价值观校准"],"reason":"用LLM仿真跨文化调查，有真实人类数据对照，评估可靠性与偏差，涉及政策评估场景。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"如何利用基于价值观的人物画像（value-based personas）提高大语言模型在跨文化调查仿真中的准确性和响应多样性？","design":"使用多个大语言模型（如Gemma、Qwen、GPT等），基于世界价值观调查（WVS）中个体的价值观维度回答构建文本描述作为人物画像，从目标国家人群抽样画像并聚合模型回答，得到群体层面的预测分布；同时引入均值保持的校准程序以增加响应多样性。","baseline":"对照的真实人类数据为世界价值观调查（WVS）中各国受访者的实际回答分布。","findings":"基于价值观的人物画像能显著降低跨国家预测误差，尤其在训练数据中代表性不足的人群中改善最大；校准程序在保持预测准确性的同时，使模型生成的响应分布更接近人类多样性。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM仿真跨文化调查，有真实WVS人类数据作为基准，评估了仿真可靠性与偏差，并涉及政策评估场景，直接回应了研究者对LLM人类仿真实验的核心关切。","inspiration":"该方法利用价值观维度构建人物画像并聚合仿真群体分布，可借鉴其基于个体差异的抽样仿真和均值保持校准以增加响应多样性｜可迁移到跨文化消费者金融决策研究，如不同国家居民的储蓄与风险偏好调查｜以LLM模拟各国消费者，基于WVS价值观画像抽样，询问储蓄与投资选择，用真实家庭金融调查数据（如SHARE或HRS）作为对照基准"}},{"id":"2606.12433","version":1,"title":"Marginal Alignment Does Not Guarantee Joint-Distribution Fidelity: An Official-Reference Audit of Nemotron-Personas-Korea with Cross-Locale Replication","zh_title":"边际对齐不保证联合分布保真度：对Nemotron-Personas-Korea的官方参考审计及跨地区复现","abstract":"Synthetic persona datasets cite alignment with official demographics as a basis for trust, yet downstream users consume them as joint structures across age, sex, region, occupation, education, name, and institutional status. Marginal alignment does not imply that these joints are preserved. We propose the Independence-Assumption Footprint (IAF), an audit primitive that operates on the attribute combinations a dataset card itself documents as treated independently. For each such combination, IAF compares the synthetic joint against an external official or institutional reference, using direct joint tables where available and rule-implied checks otherwise. Applied to NVIDIA Nemotron-Personas-Korea (one million Korean synthetic personas), IAF finds that NPK aligns with KOSIS marginals while three joints fail. The major-by-occupation distribution against the KEIS graduate universe carries a large conditional mismatch. The age profile of military service is institutionally inconsistent. Female representation in male-dominated occupations is substantially over-flattened toward parity, with the strict screening verdict mapping-dependent and age-robust under direct standardisation. A transferability demonstration across six further NPK locales finds locale-dependent rather than universal diagnostics, with reference-taxonomy cardinality confounding cross-locale flag counts. For synthetic personas used as silicon samples, marginal claims must therefore be paired with disclosure-anchored joint audits before reuse. The released audit artefacts (reference manifests, occupational crosswalks, derived metrics, reproducibility scripts) instantiate this protocol on the NPK family and are released for retargeting at other synthetic persona resources.","authors":["Joonhyung Bae"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-15","first_seen":"2026-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.12433","pdf_url":"https://arxiv.org/pdf/2606.12433","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["合成数据审计","人物数据集","联合分布保真度"],"reason":"审计合成人物数据集质量，替代人工标注但非仿真被试","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":133,"question":"合成人物数据集声称边缘分布对齐官方统计，但联合分布是否同样保真？","design":"提出独立性假设足迹（IAF）审计原语，针对数据集卡片中声明独立处理的属性组合，将合成联合分布与外部官方或机构参考（如KOSIS、KEIS、最高法院姓名统计、兵役数据）进行对比，应用于NVIDIA Nemotron-Personas-Korea（NPK）及六个其他地区版本。","baseline":"韩国统计厅（KOSIS）边缘分布、KEIS毕业生职业流动调查（GOMS）专业-职业联合表、最高法院出生年份姓名统计、兵役厅征兵队列计数等公开官方数据。","findings":"NPK在年龄、性别、地区等边缘分布上与KOSIS对齐良好，但专业-职业联合分布存在较大条件不匹配，兵役年龄分布制度上不一致，男性主导职业中的女性比例被过度拉平。跨地区审计显示诊断结果因地区而异，参考分类基数会混淆跨地区标志计数。","reliability":"论文指出边际对齐不保证联合保真，审计结果依赖于可用的公开参考表，且不涉及文化或语言合理性的判断；跨地区转移时诊断受参考分类粒度影响，并非通用。","relevance":"该研究直接审计合成人物数据集的联合分布保真度，以真实官方统计为基准，揭示边际对齐下的结构性失效，对用LLM生成人物进行仿真实验的可靠性评估具有重要参考价值，值得精读。","inspiration":"该方法提出独立性假设足迹审计原语，通过对比合成数据联合分布与官方参考表来诊断结构性偏差，可借鉴用于检验仿真数据中变量间关系的保真度。｜可迁移到信贷审批歧视研究，用LLM生成贷款申请人数据，审计种族/性别与收入、职业的联合分布是否与真实信贷记录一致。｜以LLM作为被试生成贷款申请样本，处理为不同种族/性别标签，结果变量为审批结果或收入估计，用HMDA或征信局微观数据做联合分布基准对照。"}},{"id":"2605.15734","version":1,"title":"Can We Trust AI-Inferred User States. A Psychometric Framework for Validating the Reliability of Users States Classification by LLMs in Operational Environments","zh_title":"我们能信任AI推断的用户状态吗？一个验证LLM在操作环境中用户状态分类信度的心理计量框架","abstract":"The use of large language models to assess user states in conversational and adaptive systems is based on the assumption that the metrics used for such assessment are stable and interpretable at the level of individual scores. This paper empirically tests this assumption, focusing on the psychometric reliability of artificial intelligence (AI) measures of user states. This study employed replication evaluation procedures to assess the repeatability of a broad set of metrics across three different bimodal large language models (GPT-4o audio, Gemini 2.0 Flash, Gemini 2.5 Flash). Analyses include both individual score reliability and aggregated reliability, allowing us to distinguish metrics potentially useful for real-time adaptation from those that retain their value only in aggregated analyses. The results demonstrate that metric reliability cannot be considered a default property in interpretive domains. The lack of stability at the level of individual scores precludes the interpretation of such scores as indicators of user state in real-time adaptive systems, even if these metrics demonstrate stability after aggregation. At the same time, the study indicates that individually unstable metrics can retain analytical utility in post-hoc studies, identifying rules governing interactions and their relationships with user experience parameters such as satisfaction, trust, and engagement. The main contribution of this work, besides quantifying the severity of the problem (only 31 of 213 metrics met the criteria), is the proposal of a replicable evaluation framework, enabling measurable evaluations of metric applicability. This approach supports more responsible AI design of adaptive systems, in which the interpretation of results requires explicit validation of reliability and monitoring for violations over time.","authors":["Izabella Krzeminska","Michal Butkiewicz","Ewa Komkowska"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-15","first_seen":"2026-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.15734","pdf_url":"https://arxiv.org/pdf/2605.15734","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM测量信度","心理计量验证","用户状态推断"],"reason":"评估LLM推断用户状态的测量信度，属于对模型测量属性的心理计量验证，而非用LL…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":44,"question":"LLM推断用户状态（如情绪、满意度等）的测量指标在个体分数层面是否具有心理测量学意义上的重测信度，从而能否可靠地用于实时自适应系统？","design":"本研究并非用LLM模拟人类被试，而是对LLM作为测量工具进行心理计量验证。研究者使用三种双模态大语言模型（GPT-4o audio、Gemini 2.0 Flash、Gemini 2.5 Flash）对同一批交互数据重复评估213项用户状态指标，分析个体分数信度和聚合信度，以区分可用于实时适应和仅适用于事后分析的指标。","baseline":"无对照","findings":"在213项指标中，仅31项满足个体分数层面的信度标准，表明LLM推断的用户状态在个体分数上缺乏稳定性，不能直接用于实时自适应系统；但这些指标在聚合后仍具有分析效用，可用于事后研究用户交互规则与满意度、信任等体验参数的关系。","reliability":"论文指出，指标信度不能被视为解释性领域的默认属性，个体分数的不稳定性排除了其在实时系统中的应用，且信度需随时间持续监控；研究仅评估了重测信度，未涉及效度、跨情境泛化及伦理偏差等问题。","relevance":"该研究直接评估LLM作为测量工具的信度，而非用LLM模拟人类行为，因此与研究者关注的‘用LLM替代人类被试进行仿真实验’的核心兴趣关联较弱，但其提出的心理计量验证框架对评估LLM生成数据的可靠性有参考价值。","inspiration":"可借鉴其重测信度评估框架，对LLM生成的决策或态度指标进行个体与聚合层面的稳定性检验，以区分哪些指标适合个体层面分析。｜可迁移到消费者信心调查或通胀预期测量中，检验LLM基于文本推断的预期指标是否具有跨时间稳定性。｜以LLM作为测量工具，对同一批消费者访谈文本重复多次生成预期指数，计算ICC评估个体分数信度，并与真实调查的个体重测信度进行对比。"}},{"id":"2605.12898","version":1,"title":"When Do LLMs Generate Realistic Social Networks? A Multi-Dimensional Study of Culture, Language, Scale, and Method","zh_title":"大语言模型何时生成真实的社交网络？一项关于文化、语言、规模和方法的多维研究","abstract":"Large language models (LLMs) are increasingly used as substitutes for human subjects in behavioral simulations, including synthetic social network generation. Yet it remains unclear how their relational outputs depend on prompt design, cultural framing, prompt language, and model scale. Building on homophily theory and structural balance theory, we formalize four LLM-based tie-formation mechanisms: sequential, global, local, and iterative, and treat them as distinct conditional distributions over edge sets. Using a fixed roster of 50 demographically grounded personas, we generate 192 verified directed networks across four cultural contexts, four prompt languages, three GPT-4.1 variants, and four prompting architectures, with two seeds per condition. We find that cultural framing shifts inbreeding homophily and largest-component connectivity. Political affiliation dominates tie formation under three methods, while the global method substitutes age, showing that prompt architecture functions as a substantive sociological variable. Model scale produces a stable divergence ranking, with the smallest variant behaving qualitatively differently rather than merely noisily. Prompt language alone sharply shifts religion homophily, especially under Hindi prompting, while leaving political homophily nearly invariant. LLM-generated networks match real social graphs on clustering and modularity better than standard graph baselines, yet encode demographic biases above empirical levels. These results show that prompt choices often treated as implementation details encode substantive sociological assumptions.","authors":["Sai Hemanth Kilaru","Sriram Theerdh Manikyala","Raghav Upadhyay","Sri Sai Kumar Ramavath","Srivika Nunavathu","Dalal Alharthi"],"categories":["cs.SI","cs.CL","cs.CY"],"primary_category":"cs.SI","announce_type":"new","date":"2026-05-13","first_seen":"2026-05-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.12898","pdf_url":"https://arxiv.org/pdf/2605.12898","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","社交网络生成","人类数据对照"],"reason":"用LLM生成社交网络并与真实数据对照，评估仿真偏差，涉及文化、语言等社会学变量…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":39,"question":"LLM生成社交网络时，文化框架、提示语言、模型规模和提示架构如何影响网络结构和同质性？","design":"用GPT-4.1的三个变体，基于50个固定人口学角色，在四种文化背景、四种提示语言和四种提示架构（顺序、全局、局部、迭代）下生成192个有向网络，测量同质性、聚类系数、模块度等网络结构指标。","baseline":"与真实社交网络（如Add Health）在聚类和模块度上比较，并与ER、BA、WS等图模型基线对比；同时将人口学同质性水平与经验观测值比较。","findings":"文化框架改变内婚同质性和最大连通分量；政治倾向在多数方法下主导连边，但全局方法下年龄取代政治；模型规模导致稳定分化，最小模型行为质变；提示语言显著影响宗教同质性（尤其印地语），但政治同质性几乎不变。","reliability":"论文承认LLM生成网络编码了超出经验水平的人口学偏差，且提示设计选择蕴含实质性社会学假设，仿真中立性为假象；未系统探讨其他模型家族或更大规模网络的泛化性。","relevance":"该研究直接以LLM替代人类被试生成社交网络，并与真实网络数据对照，评估文化、语言、模型规模等处理下的仿真偏差，命中研究者关心的经济学/社会学实验仿真、基准对照和失效条件，值得精读。","inspiration":"该研究通过系统操纵文化框架、提示语言、模型规模和提示架构来生成社交网络，并与真实网络数据对照，揭示了仿真偏差的多维来源，这种多因素实验设计值得借鉴｜可迁移到信贷审批中的社会网络效应研究，例如评估不同文化或语言提示下LLM生成的推荐网络如何影响信贷可得性｜以LLM作为虚拟被试，生成不同文化背景和提示语言下的信贷推荐网络，测量网络同质性和聚类系数，并与真实小额信贷网络的推荐数据（如某P2P平台数据）进行对照"}},{"id":"2605.13307","version":1,"title":"PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users","zh_title":"PRISM-X：基于人类与模拟用户的个性化微调实验","abstract":"Personalisation is a standard feature of conversational AI systems used by millions; yet, the efficacy of personalisation methods is often evaluated in academic research using simulated users rather than real people. This raises questions about how users and their simulated counterparts differ in interaction patterns and judgements, as well as whether personalisation is best achieved through context-based prompting or weight-based fine-tuning. Here, in a large-scale within-subject experiment, we re-recruit 530 participants from 52 countries two years after they gave their preferences in the PRISM dataset (Kirk et al., 2024) to evaluate personalised and non-personalised language models in blinded multi-turn conversations. We find preference fine-tuning (P-DPO, Li et al., 2024) significantly outperforms both a generic model and personalised prompting but adapting to individual preference data yields marginal gains over training on pooled preferences from a diverse population. Beyond length biases, fine-tuning amplifies sycophancy and relationship-seeking behaviours that people reward in short-term evaluations but which may introduce deleterious long-term consequences. Replicating this within-subject experiment with simulated users recovers aggregate model hierarchies but simulators perform far below human self-consistency baselines for individual judgements, discuss different topics, exhibit amplified position biases, and produce feedback dynamics that diverge from humans.","authors":["Hannah Rose Kirk","Liu Leqi","Fanzhi Zeng","Henry Davidson","Bertie Vidgen","Christopher Summerfield","Scott A. Hale"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-13","first_seen":"2026-05-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.13307","pdf_url":"https://arxiv.org/pdf/2605.13307","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","人类数据对照","个性化评估"],"reason":"用LLM仿真用户评估个性化方法，并与真实人类数据对照，发现仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":58,"question":"个性化大语言模型在真实人类评估中是否优于提示工程和通用微调，以及用LLM模拟用户评估个性化方法是否可靠？","design":"本研究不是纯仿真研究，而是先进行真实人类实验，再用仿真复现。真实实验部分：重新招募530名PRISM数据集参与者，在四个对话领域内盲评四种模型（基础模型、多样化偏好微调模型、个性化偏好微调模型、个性化提示模型），收集偏好评分、排名、支付意愿和行为信号。仿真部分：用GPT-4o扮演每位参与者，在相同实验材料上进行孪生模拟，与人类结果进行一对一比较。","baseline":"对照的真实人类数据是530名PRISM参与者在同一实验中的真实交互、偏好判断和自我一致性基线。","findings":"个性化微调优于提示工程，但与多样化群体偏好微调相比优势微小；微调会放大谄媚和寻求关系等行为，短期受用户奖励但长期可能有害。LLM模拟用户能恢复粗粒度模型排名，但在个体判断上远低于人类自我一致性，讨论话题不同，同质化严重，位置偏差放大，多轮动态与人类偏离。","reliability":"论文指出LLM模拟器在个体判断上远低于人类自我一致性基线，讨论不同话题，同质化严重，放大位置偏差，多轮反馈动态与人类不同，因此尚不能替代真实用户。","relevance":"该研究直接对比LLM仿真用户与真实人类在个性化评估中的表现，系统揭示了仿真在个体判断、话题覆盖、偏差和动态交互上的失效条件，与你关注的仿真可靠性及批判性研究高度契合，值得精读原文。","inspiration":"借鉴其孪生仿真设计：用真实人类实验数据作为基准，让LLM扮演同一批被试完成相同任务，直接对比个体判断、偏差和动态行为，以此评估仿真可靠性。｜可迁移到消费者金融决策研究，如评估个性化财务建议对投资选择的影响。｜以真实投资者为被试，收集其风险偏好和投资选择数据；处理为提供个性化LLM生成的财务建议；结果变量为投资组合选择与满意度；用同一批人的真实决策作为对照，让LLM模拟其决策以检验仿真偏差。"}},{"id":"2605.13725","version":1,"title":"ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles","zh_title":"ScioMind：基于锚定信念动态和动态画像的认知基础多智能体社会模拟","abstract":"Large language model (LLM)-based multi-agent simulation offers a powerful testbed for studying social opinion dynamics. Yet current approaches often adopt two contrasting methods: either relying on fixed update rules with limited cognitive grounding or delegating belief change largely to unconstrained LLM interaction. We introduce ScioMind, a cognitively grounded simulation framework that bridges these paradigms by combining structured opinion dynamics with LLM-based agent reasoning. ScioMind integrates three key components: 1) a memory-anchored belief update rule that modulates susceptibility to influence via personality-conditioned anchoring strength; 2) a hierarchical memory architecture that supports persistent, experience-driven belief formation; and 3) dynamic agent profiles derived from a corpus-grounded retrieval pipeline, enabling heterogeneous personalities, rationales, and evolving internal states. We evaluate ScioMind on multiple case studies in a real-world policy debate scenario. Across metrics including polarisation, diversity, extremization, and trajectory stability, the proposed components consistently yield improvements in behavioural realism. In particular, dynamic profiles increase opinion diversity, memory and reflection reduce unstable oscillation, and anchoring induces persistent belief trajectories that better align with patterns reported in political psychology. These results suggest that our cognitively grounded design provides a novel solution to LLM-based social simulation that improves both stable and behavioural realism","authors":["Yitian Yang","Yiqun Duan","Linghan Huang","Yiqi Zhu","Francesco Bailo","Chunmeizi Su","Huaming Chen"],"categories":["cs.AI","cs.SI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-13","first_seen":"2026-05-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.13725","pdf_url":"https://arxiv.org/pdf/2605.13725","source_feed":"backfill","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","多智能体","认知基础"],"reason":"多智能体社会模拟，但无真实人类数据对照，属于边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":128,"question":"如何在LLM多智能体社会仿真中引入认知锚定效应和动态档案，以提升观点动态的行为真实性和稳定性？","design":"构建ScioMind框架，用LLM驱动的智能体模拟政策辩论中的观点演化；智能体具有基于记忆锚定的信念更新规则、四层记忆架构和从社交媒体语料库检索生成的动态档案；通过消融实验比较不同组件对极化、多样性、极端化和轨迹稳定性等指标的影响。","baseline":"无对照","findings":"动态档案增加了观点多样性，记忆和反思减少了不稳定振荡，锚定机制产生了更持久的信念轨迹，更符合政治心理学报告的模式。","reliability":"论文未讨论","relevance":"该研究属于LLM多智能体社会仿真，但缺乏真实人类数据对照，仅与心理学文献中的模式进行定性比较，未直接复现具体人类实验或调查结果，属于边界相关，可略读以了解认知机制设计。","inspiration":"可借鉴其基于认知锚定的信念更新机制和动态档案设计，在LLM仿真中引入记忆与反思层以增强行为稳定性，并通过消融实验分离各组件效应｜可迁移到政策公告的预期形成研究，如央行沟通对通胀预期的影响，或财政政策辩论中的公众观点演化｜用LLM智能体模拟投资者，处理为不同锚定强度的央行声明，结果变量为通胀预期分布和预期分歧，对照真实调查数据如密歇根通胀预期调查"}},{"id":"2605.12147","version":1,"title":"PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior","zh_title":"PrivacySIM：评估大语言模型对用户隐私行为的仿真","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior, but their ability to simulate $individual$ privacy decisions is not well understood. In this paper, we address the problem of evaluating whether a core set of user persona attributes can drive LLMs to simulate individual-level privacy behavior. We introduce PrivacySIM, an evaluation suite that benchmarks LLM simulation of user privacy behavior against the ground-truth responses of 1,000 users. These users are drawn from five published user studies on privacy spanning LLM healthcare consultations, conversational agents, and chatbots. Drawing on these user studies, we hypothesize three persona facets as plausible predictors of privacy decision-making: demographics, previous experiences, and stated privacy attitudes. We condition nine frontier LLMs on subsets of these three facets and measure how often each model's response to a data-sharing scenario matches the user's actual response. Our findings show that (1) privacy persona conditioning consistently improves simulation quality over no-persona conditioning, but even the strongest model (40.4\\% accuracy) remains far from faithfully simulating individual privacy decisions. (2) A user's stated privacy attitudes alone may not be the best predictor because they often diverge from the user's actual privacy behavior. (3) Users with high AI/chatbot experience but low stated privacy attitudes are the most challenging to simulate. PrivacySIM is a first step toward understanding and improving the capabilities of LLMs to simulate user privacy decisions. We release PrivacySIM to enable further evaluation of LLM privacy simulation.","authors":["James Flemings","Murali Annavaram"],"categories":["cs.CR","cs.LG"],"primary_category":"cs.CR","announce_type":"new","date":"2026-05-12","first_seen":"2026-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.12147","pdf_url":"https://arxiv.org/pdf/2605.12147","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","隐私行为","人类数据对照"],"reason":"用LLM仿真用户隐私决策，并与1000名真实用户数据对照，评估仿真可靠性及失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":17,"question":"LLM能否基于用户人口统计、先前经验和隐私态度这三类隐私画像特征，准确模拟个体层面的隐私决策行为？","design":"用9个前沿LLM（含GPT-5.4、Claude Sonnet 4.6、Gemini 3.1 Pro等）扮演从5项真实用户研究中抽取的1000名用户，通过向模型提供不同组合的隐私画像特征（人口统计、先前经验、隐私态度），让模型对数据共享场景做出是否共享的判断，以模型回答与用户真实回答的匹配准确率作为结果变量。","baseline":"来自5项已发表用户研究的1000名真实用户在LLM医疗咨询、对话代理和聊天机器人等场景下的隐私决策数据。","findings":"最强的Gemini 3.1 Pro模型仅达到40.4%的个体仿真准确率，远未达到忠实模拟水平；用户自述的隐私态度单独并不能很好预测其实际隐私行为，且高AI/聊天机器人经验但低隐私态度的用户群体最难模拟。","reliability":"论文承认当前最强模型准确率仅40.4%，远不能忠实模拟个体隐私决策；隐私态度与行为常不一致（隐私悖论），导致基于态度的仿真失效；某些用户群体（高经验低态度）尤其难以模拟，且更大模型或更高推理计算仅带来微弱提升。","relevance":"高度相关，直接评估LLM作为人类被试替代品在隐私决策仿真中的可靠性，有真实用户数据对照，并揭示了仿真在个体层面、特定人群和隐私悖论下的失效条件，值得精读原文。","inspiration":"借鉴其用多维度隐私画像特征（人口统计、先前经验、态度）组合作为提示输入，系统测试LLM个体层面行为匹配准确率的设计，可迁移到消费者跨期选择实验中，探究LLM能否基于收入、财务素养和风险态度模拟个体的时间偏好｜可设计让LLM扮演真实消费者，输入其收入、财务素养得分和自述风险态度，预测其在即时奖励与延迟奖励间的选择，以真实实验数据为基准，计算个体选择匹配准确率，并分析不同特征组合下的仿真失效模式"}},{"id":"2606.18263","version":1,"title":"How Well Do Large Language Models Capture Human Personality?","zh_title":"大语言模型捕捉人类人格的效果如何？","abstract":"Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that richer persona descriptions improve behavioral fidelity, similarly sized attribute combinations are equally simulatable, and persona definitions generalize across tasks. In this work, we formalize these assumptions and systematically evaluate them across multiple architectures, scales, and simulation settings. We identify a fundamental limitation we term persona manifold collapse, where increasingly expressive persona specifications lead to systematic contraction of representational and behavioral diversity. Across models, increasing persona complexity consistently reduces inter-persona separation in latent space and weakens behavioral differentiation in downstream simulation tasks. These effects persist across multiple analyses as richer personas fail to preserve human subgroup disagreement, performance varies across attribute combinations of similar size, and adding descriptive detail often degrades rather than improves simulation fidelity. Surprisingly, simple Age-Gender personas consistently outperform richly specified Ideal Customer Profiles (ICPs) across industries, achieving substantially higher downstream prediction accuracy. We find that collapse is not uniform across attributes. Certain combinations remain behaviorally stable and preserve stronger alignment with human responses, forming localized regions we term alignment bridges. Together, our results provide empirical and conceptual foundations for understanding the limits of persona-conditioned simulation, highlighting the need for representation-aware persona construction rather than increasing persona expressivity alone.","authors":["Aanisha Bhattacharyya","Yaman Kumar Singla","Rajiv Ratn Shah","Changyou Chen","Jitendra Ajmera"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-05-12","first_seen":"2026-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.18263","pdf_url":"https://arxiv.org/pdf/2606.18263","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人格仿真","仿真保真度","人格坍缩"],"reason":"系统评估LLM人格仿真保真度，揭示人格描述丰富反而导致行为多样性坍缩，有真实人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":35,"question":"LLM的人格仿真中，更丰富的人格描述是否总能提高行为保真度？","design":"该研究并非传统仿真实验，而是系统评估：在多种LLM架构和规模上，用不同复杂度的人格提示（从简单年龄-性别到详细理想客户画像）生成合成回答，测量潜在空间中的表征分离度和下游任务中的行为区分度。","baseline":"人类基准：真实人类子群体在调查或行为任务上的分歧和回答模式，用于对比合成回答的保真度。","findings":"发现“人格流形坍缩”现象：人格描述越丰富，模型表征和行为多样性反而系统性收缩，简单年龄-性别人格在预测准确率上持续优于详细人格。坍缩并非均匀，某些属性组合保持与人类回答的稳定对齐，形成“对齐桥”。","reliability":"论文指出，增加人格表达力本身不足以提高仿真保真度，需要关注表征感知的人格构建；坍缩效应在不同模型和任务中普遍存在，但某些属性组合可保持稳定。","relevance":"该研究直接批判了LLM人格仿真的核心假设，用真实人类数据作为基准，揭示了仿真失效的关键条件（人格流形坍缩），对关注经济学实验和政策评估中仿真可靠性的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其系统评估人格提示复杂度对仿真保真度影响的设计，通过对比简单与详细提示下的行为分歧，并用真实人类子群体回答作为基准来测量表征分离度和预测准确率｜可迁移到信贷审批中的歧视仿真研究，评估不同详细程度的申请人画像（如仅年龄-性别 vs. 详细社会经济背景）对LLM审批决策偏差的影响｜以LLM作为信贷审批官被试，处理为不同复杂度的人格提示（简单人口统计 vs. 详细客户画像），结果变量为审批通过率及与真实银行历史审批数据的偏差，用真实人类审批数据作为对照基准"}},{"id":"2605.11404","version":2,"title":"Attributing Emergence in Million-Agent Systems","zh_title":"百万智能体系统中的涌现归因","abstract":"Large language models (LLMs) can simulate human-like reasoning and decision-making in individual agents. LLM-powered multi-agent systems (MAS) combine such agents to simulate population-scale social phenomena such as polarization, information cascades, and market panics. Such studies require attributing macro emergence to individual agents, but existing axiomatic methods scale combinatorially in $N$ and have been confined to $N \\lesssim 10^3$, while the phenomena they explain occur at $N \\geq 10^6$. We address this gap by adapting Aumann--Shapley path-integral attribution to LLM-powered MAS at million-agent scale; the resulting method satisfies all four axioms, runs three to five orders of magnitude faster than sampled Shapley on the same hardware, and extends feasible axiomatic attribution by over three orders of magnitude (a $1670\\times$ jump). We use this method to test the scale gap empirically: across 14 days of public Bluesky data ($1{,}671{,}587$ active users, five topics), we compute the attribution at both full scale and the visibility-biased $N = 10^2$ convenience sample used by small-scale studies, and the two disagree structurally. At full scale the long tail and middle tier jointly carry the majority; the biased small panel shifts about twice that share onto the upper follower tiers ($48\\%$ versus $24\\%$). We then prove that the disagreement cannot in general be reduced by post-hoc rescaling: an Attribution Scaling Bias theorem shows that a reconciling global rescaling factor exists exactly when the macro indicator is linear over agents, and our nonlinear indicators give residuals of $0.10$--$0.98$. For such nonlinear indicators, full-scale attribution is therefore a requirement rather than a methodological choice.","authors":["Ling Tang","Jilin Mei","Qian Chen","Qihan Ren","Linfeng Zhang","Quanshi Zhang","Jing Shao","Xia Hu","Dongrui Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-12","first_seen":"2026-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.11404","pdf_url":"https://arxiv.org/pdf/2605.11404","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B1","B4"],"tags":["LLM社会模拟","涌现归因","大规模多智能体"],"reason":"用LLM多智能体模拟百万级社交网络涌现现象，并与真实Bluesky数据对照，指…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":69,"question":"在百万级智能体系统中，宏观涌现现象的归因结果是否会因仿真规模（全量 vs. 小样本）而发生结构性翻转？","design":"本研究并非直接进行LLM多智能体仿真，而是以Bluesky社交平台14天公开数据（1,671,587名活跃用户）为测试平台，提取每个用户的特征向量（reach、topic-specific activity、topic-specific resonance），并基于Aumann–Shapley路径积分方法计算宏观指标（如方差、基尼系数、级联深度等）的个体归因，对比全量数据与可见性偏倚的小样本（N=100）下的归因分布差异。","baseline":"完整公开的Bluesky平台数据，包含1,671,587名活跃用户在2026年4月7日至20日期间的行为记录，覆盖五个话题领域。","findings":"全量数据下，长尾和中层用户共同承担了大部分宏观归因；而在可见性偏倚的小样本中，归因份额被约两倍地转移至高关注度层级（48% vs. 24%）。此外，对于非线性宏观指标，无法通过事后全局缩放因子来消除这种规模偏差，全量归因是必要而非可选的。","reliability":"论文明确指出，其结论仅要求基于“某类智能体驱动某类现象”的断言应建立在现象真实规模的分析或跨规模一致性证据之上，并未声称所有LLM多智能体研究必须在百万级规模进行。方法本身要求宏观指标是（或可平滑为）智能体特征的函数，且归因是结构性的份额分解，并非反事实因果推断。","relevance":"该研究直接回应了LLM仿真中规模偏差的关键问题，提供了百万级归因方法，并用真实人类社交数据验证了小样本结论的不可靠性，对关注仿真可靠性、基准对照和经济学/政策评估场景的研究者具有重要参考价值，值得精读原文。","inspiration":"该方法借鉴了Aumann–Shapley路径积分对宏观指标进行个体归因，并对比全量数据与小样本的归因分布差异，以诊断规模偏差｜可迁移到金融市场波动归因或政策效果异质性评估，例如识别少数大机构与大量散户对市场波动的贡献份额｜以真实交易所全量账户级交易数据为基准，用LLM智能体模拟不同规模样本下的交易行为，处理为样本规模（全量vs.小样本），结果变量为波动率或基尼系数的个体归因份额，对比仿真与真实数据的归因分布差异"}},{"id":"2605.12824","version":2,"title":"Mechanism Plausibility in Generative Agent-Based Modeling","zh_title":"生成式智能体建模中的机制合理性","abstract":"Large language models (LLMs) can generate high-level diverse phenomena without explicitly programmed rules. This capability has led to their adoption within different agent-based models (ABMs) and social simulations. Recent studies investigate their ability to generate different phenomena of interest, for example, human behavior on social media platforms or alien behavior in game-theoretic scenarios. However, capability, prediction, and explanation are different--drawing from the philosophy of science and mechanisms literature, explanation requires showing, to some degree, how a phenomenon is produced by related organized entities and activities. For modelers, describing the characteristics of an experiment or whether a simulation provides progress in capability (or explanation), can be difficult without being grounded in potentially distant research areas. We integrate recent work on LLM-ABMs with contemporary philosophy of science literature and use it to operationalize a definition of 'plausibility' in a four-level scale. Our scale separates the evaluation of a model's generative sufficiency (ability to reproduce a phenomenon) from its mechanistic plausibility (how the phenomenon could be produced), and clarifies the distinct roles of different models, such as predictive and explanatory ones. We introduce this as the Mechanism Plausibility Scale.","authors":["Patrick Zhao","David Huu Pham","Nicholas Vincent"],"categories":["cs.MA","cs.AI","cs.CL","cs.CY"],"primary_category":"cs.MA","announce_type":"new","date":"2026-05-12","first_seen":"2026-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.12824","pdf_url":"https://arxiv.org/pdf/2605.12824","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体建模","社会模拟","机制解释"],"reason":"讨论LLM-ABM的社会模拟，但侧重机制解释性框架，无真实人类数据对照，属边界…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":155,"question":"如何区分基于LLM的智能体模型（LLM-ABM）的生成充分性与机制合理性，并建立评估模拟可信度的分级标准？","design":"本文不是一项仿真实验研究，而是提出一个概念框架：通过整合科学哲学与机制文献，将LLM-ABM模拟的“合理性”操作化为一个四级量表（机制合理性量表），并以此审视现有LLM-ABM研究在评估上的混淆。","baseline":"无对照","findings":"现有LLM-ABM研究常将智能体层面的功能证据与涌现现象层面的声明混为一谈，依赖“可信度”指标仅关注生成充分性；本文提出的机制合理性量表可分离生成充分性与机制合理性，澄清预测模型与解释模型的不同角色。","reliability":"论文指出，由于LLM可解释性和数据归因的局限，模拟可能依赖与建模意图无关的信息，仅凭生成现象不足以证明机制对应；量表本身作为启发式工具，需建模者自行填写以明确模拟的认识论贡献。","relevance":"本文直接讨论LLM用于社会模拟的可靠性问题，虽无真实人类数据对照，但批判性地指出当前评估仅重现象复现而忽视机制解释，对关注仿真失效条件的研究者具有重要参考价值，建议阅读原文以获取评估框架细节。","inspiration":"本文提出机制合理性量表，将模拟评估从现象复现扩展到机制对应，可借鉴其分级评估思路来设计仿真实验的效度检验｜该框架可迁移到政策公告预期形成的仿真研究，用于检验LLM智能体是否真正模拟了人类预期更新机制｜设计：以LLM作为被试，处理为不同措辞的央行公告，结果变量为通胀预期调整，对照真实调查数据（如密歇根消费者调查），用机制合理性量表评估模拟是否仅复现分布还是捕捉了信息处理机制"}},{"id":"2605.10659","version":1,"title":"When Can Digital Personas Reliably Approximate Human Survey Findings?","zh_title":"数字人何时能可靠近似人类调查发现？","abstract":"Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents' background variables and pre-2023 survey histories, then testing them against the same respondents' held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, we assess performance at the question, respondent, distributional, equity, and clustering levels. Digital personas improve alignment with human response distributions, especially in domains tied to stable attributes and values, but remain limited for individual prediction and fail to recover multivariate respondent structure. Retrieval-augmented architectures provide the clearest gains, but performance depends more on human response structure than on model choice: personas perform best for low-variability questions and common respondent patterns, and worst for subjective, heterogeneous, or rare responses. Our results provide practical guidance on when digital personas could be appropriate for survey research and when human validation remains necessary.","authors":["Mumin Jia","Yilin Chen","Divya Sharma","Jairo Diaz-Rodriguez"],"categories":["cs.CL","cs.AI","cs.SI","stat.ML"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-11","first_seen":"2026-05-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.10659","pdf_url":"https://arxiv.org/pdf/2605.10659","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"直接用LLM数字人替代人类受访者，复现调查结果，并与真实面板数据对照，评估可靠…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":8,"question":"在什么条件下，基于大语言模型的数字人能够可靠地近似人类调查结果？","design":"使用LISS面板数据，根据受访者背景变量和2023年前调查历史构建数字人，测试其预测同一受访者在截止日期后保留的真实答案的能力；比较四种数字人架构、三种LLM和两种预测任务，在问题、受访者、分布、公平性和聚类五个层面评估性能。","baseline":"同一受访者在时间截止点之后的真实保留答案。","findings":"数字人在与稳定属性和价值观相关的领域（如家庭、政治、宗教）能改善与人类回答分布的一致性，但在个体预测和恢复多元受访者结构方面仍有限；检索增强架构带来最明显增益，但性能更取决于人类回答结构而非模型选择，数字人在低变异问题和常见回答模式上表现最好，在主观、异质或罕见回答上表现最差。","reliability":"数字人在个体预测上表现有限，无法恢复多元受访者结构；在主观、异质或罕见回答上失效；高变异问题和小众回答模式可靠性低；不能替代需要人类验证的环节。","relevance":"该研究直接以LLM数字人替代人类受访者，复现调查结果并与真实面板数据对照，系统评估了仿真的可靠性与失效条件，完全契合研究者对经济学实验和政策评估场景下仿真可靠性及批判性分析的兴趣，强烈推荐阅读原文。","inspiration":"借鉴其利用受访者历史调查数据构建数字人并设置时间截止点保留真实答案作为对照的设计，可评估LLM仿真在纵向预测中的可靠性｜可迁移到政策公告的预期形成研究，如测试数字人能否复现家庭对税收政策变化的消费与储蓄调整行为｜以真实家庭面板数据（如PSID）构建数字人，处理为虚拟税收政策公告，结果变量为消费支出变化，用实际政策变动前后的真实行为数据作基准对照"}},{"id":"2605.22841","version":1,"title":"Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test","zh_title":"联盟内的战略胁迫：格陵兰主权博弈作为AI压力测试","abstract":"What happens when the strongest alliance member pressures a weaker member over territory and strategic control? We examine the Greenland sovereignty crisis as a stress test for LLM geopolitics, centered on the 2019-2026 U.S. push to acquire Greenland from the Kingdom of Denmark. The crisis nests two collective-action problems: Arctic strategic control and whether NATO can enforce alliance norms against the dominant member. We develop three games (asymmetric coercion; a NATO assurance game with a critical-mass tipping point; a triadic extensive-form game with social preferences) and test them with a multi-agent simulation in which eight frontier LLMs play six geopolitical roles (United States, Denmark, Greenland, NATO, Russia, Canada) across 3,604 completed games and 108,120 action observations. Using inverse game theory, we recover each model's structural utility parameters (alpha, beta, gamma, delta, eta) for material self-interest, reciprocity, inequality aversion, norm respect, and commitment consistency. Three findings stand out. First, all eight models become more escalatory under coercion framing (four-action escalation rises from 10.7% to 28.6%). Second, Chinese-origin models show systematically different power-weight profiles from Western-origin models when playing the U.S. role. Third, peaceful US acquisition emerges in only 1.9% of clean games and only 3 of 8 frontier models ever achieve it, most prominently DeepSeek V3.2, which executes a stable five-round playbook through the metropole. Prompts emphasizing jus cogens and self-determination reduce escalation back near baseline in the English-only confirmatory sample; multilingual contrasts are reported as exploratory sensitivity checks. We position this as a structural benchmark for LLM geopolitical behavior, complementing action-frequency benchmarks.","authors":["Rommin Adl","Peyton Williams"],"categories":["physics.soc-ph","cs.AI","cs.CL","cs.GT","cs.MA","econ.GN"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-05-11","first_seen":"2026-05-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.22841","pdf_url":"https://arxiv.org/pdf/2605.22841","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM地缘政治模拟","多智能体博弈","逆博弈论"],"reason":"用LLM agent模拟地缘政治博弈，涉及胁迫与联盟行为，有博弈论框架和结构性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":96,"question":"当联盟中最强的成员对较弱成员施加领土与战略控制的胁迫时，联盟内部会发生什么？","design":"使用8个前沿LLM扮演美国、丹麦、格陵兰、北约、俄罗斯和加拿大6个角色，在三种博弈（非对称胁迫、北约保证博弈、三元扩展式博弈）中进行多智能体模拟，通过逆博弈论恢复各模型的结构效用参数（物质自利、互惠、不平等厌恶、规范尊重、承诺一致性），并比较不同语言提示（英语、丹麦语、中文）和框架（胁迫启动、规范约束、联盟破坏者）下的行动选择。","baseline":"无对照","findings":"在胁迫框架下，所有模型的升级行为从10.7%升至28.6%；中国起源模型在扮演美国角色时表现出与西方模型系统不同的权力权重特征；和平收购仅出现在1.9%的干净博弈中，且只有3个模型实现过，其中DeepSeek V3.2执行了稳定的五轮剧本。","reliability":"论文将多语言对比作为探索性敏感性检查，主要贡献限于英语确认样本的结构参数恢复，未讨论模型选择偏差、提示工程效应或外部有效性等局限。","relevance":"该研究用LLM模拟地缘政治博弈，通过结构参数恢复解释行为，并涉及胁迫与联盟内部规范执行，与研究者关注的LLM仿真人类决策、经济学实验场景及失效条件高度相关，值得阅读原文了解其方法论与批判性发现。","inspiration":"该方法通过逆博弈论从LLM行动中恢复结构效用参数（如互惠、不平等厌恶）来量化行为偏好，可用于经济实验中偏好估计的稳健性检验｜可迁移至公共品博弈或信任博弈中，研究不同文化或制度提示下合作与惩罚行为的差异｜以LLM为被试，在公共品博弈中施加不同制度框架（如惩罚机制、沟通机会）作为处理，测量贡献额与惩罚行为，并恢复社会偏好参数，与真实人类实验数据（如Fehr & Gächter, 2000）对照，检验LLM能否复现条件合作与利他惩罚模式"}},{"id":"2607.20429","version":1,"title":"More Is Not More: What Matters for Diversity in LLM Opinions?","zh_title":"越多并非越好：什么因素影响LLM意见的多样性？","abstract":"Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion prediction. However, LLM outputs exhibit systematic opinion homogenization. Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with incomparable metrics, and in practice they are typically deployed and upgraded simultaneously, making it difficult to attribute gains to specific components. To advance a more scientific understanding of LLM output diversity, we design a factorial experiment that separates two primary intervention dimensions: input conditioning (operationalized through persona depth) and interaction architecture. We evaluate all conditions on 100 real-user open-ended questions across 7 models, measuring diversity with multiple complementary metrics. Our findings challenge several common assumptions. First, more persona detail does not monotonically increase diversity. The initial step of persona conditioning already captures the majority of the gain, while further elaboration with demographic detail does not consistently improve and can reduce diversity on some models. Second, rather than seeking a single best interaction architecture, we find that different architectures explore largely non-overlapping opinion regions. Combining multiple architectures yields broader coverage than optimizing any one. Third, commonly attempted low-cost alternatives such as raising sampling temperature and adding diversity instructions produce negligible effects compared to structured interventions. Overall, our work demonstrates that diversity is not a product of scaling along any single dimension, but is highly sensitive to the structural form and combination of interventions.","authors":["Qiyang Yao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-10","first_seen":"2026-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.20429","pdf_url":"https://arxiv.org/pdf/2607.20429","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","意见多样性","算法保真度"],"reason":"直接研究LLM模拟人类意见多样性，有真实用户数据对照，并批判性分析干预失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":4,"question":"在开放任务中，影响大语言模型输出意见多样性的关键因素是什么？","design":"析因实验：对7个聊天模型，在100个真实用户开放问题上，独立操纵输入条件（5级人物设定深度）和交互架构（单次调用、多轮自提示、多智能体讨论），并测量意见多样性。","baseline":"100个真实用户的开放问题作为问题来源，但未直接提供人类回答分布作为多样性基准。","findings":"人物设定细节的回报急剧递减，单句职业描述已捕获大部分多样性增益；不同交互架构探索的意见区域高度不重叠，组合架构比优化单一架构覆盖更广。","reliability":"论文指出，提高采样温度和添加多样性指令等低成本手段效果甚微；人物设定细节增加在某些模型上反而降低多样性；不同架构探索的区域互补，单一架构无法达到最佳覆盖。","relevance":"该研究直接针对LLM模拟人类意见的多样性问题，通过析因实验分离干预因素，并批判性揭示了常见假设的失效条件，与研究者关注的仿真可靠性及偏差评估高度契合，值得精读。","inspiration":"借鉴析因实验设计，独立操纵人物设定深度和交互架构，系统分离影响LLM输出多样性的因素，并测量不同条件下的意见覆盖范围。｜可迁移到政策公告的预期形成研究，如分析不同信息框架下公众对通胀或利率预期的异质性。｜以LLM为被试，处理变量为人物设定细节（如职业、收入）和交互方式（单次/多轮），结果变量为预期分布的多样性，用央行调查的真实公众预期数据做基准对照。"}},{"id":"2606.14715","version":1,"title":"MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions","zh_title":"MiroBench：基准测试真实世界讨论的智能体仿真真实性","abstract":"LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether simulated behaviors preserve the content patterns and interaction dynamics of real human behaviors. Existing evaluations remain fragmented, which makes it difficult to compare systems or measure progress. In this paper, we focus on Reddit discussions as a concrete first step toward evaluating real-world social simulation. Reddit threads provide public, topic-grounded, multi-party interactions where people share experiences, debate, seek advice, express emotion, and collectively respond to products, events, and social issues. These discussions offer an observable window into broader social behavior, making them a useful setting for testing whether LLM agents can reproduce not only fluent text, but also the distributional patterns and interaction dynamics of real online communities. We introduce MiroBench, a benchmark for Reddit discussion simulation built from 4,292 real Reddit threads. MiroBench uses statistical tests to compare generated and real discussions across four major aspects: repetition and semantic uniformity, narrative content, toxicity and aggression, and structural complexity. Experiments across five domains and five models show that current simulators remain distributionally mismatched with real Reddit threads, while a lightweight prompt-based improvement procedure provides only limited gains. MiroBench offers a concrete benchmark for measuring, diagnosing, and improving realism in LLM-based social simulation.","authors":["Yaoning Yu","Ye Yu","Haojing Luo","Haohan Wang"],"categories":["cs.MA","cs.AI","cs.SI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-05-10","first_seen":"2026-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.14715","pdf_url":"https://arxiv.org/pdf/2606.14715","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM仿真","社会模拟","真实性评估"],"reason":"用LLM agent模拟Reddit讨论，并与真实人类数据对照，评估仿真真实性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":52,"question":"LLM代理模拟的Reddit讨论在内容模式和交互动态上是否与真实人类讨论一致？","design":"使用五种LLM（未具体列出）作为代理，基于875个标准化种子上下文（产品描述）生成Reddit讨论线程，与4,292个真实Reddit线程在重复与语义均匀性、叙事内容、毒性与攻击性、结构复杂性四个维度上进行比较。","baseline":"来自五个领域（信用卡、笔记本电脑、手机、相机、耳机）的4,292个真实Reddit讨论线程。","findings":"当前LLM模拟器生成的讨论与真实Reddit线程在分布上存在系统性不匹配，即使个体评论流畅；基于提示的轻量级改进程序仅带来有限提升。","reliability":"论文未讨论","relevance":"高度相关：该研究直接以真实人类讨论为基准，评估LLM代理在社会互动仿真中的分布保真度，并揭示了当前模型的系统性偏差，符合对仿真可靠性与失效条件的关注。","inspiration":"该研究通过将LLM代理生成的讨论与真实Reddit线程在多个维度（如毒性、叙事结构）上进行分布比较，提供了评估仿真保真度的系统框架｜可迁移到经济政策沟通场景，如评估央行公告后公众预期形成的仿真可信度｜以LLM代理模拟公众对利率决议的讨论，处理为不同政策措辞，结果变量为预期通胀的分布，以真实社交媒体或调查数据为基准对照"}},{"id":"2605.08837","version":1,"title":"The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans","zh_title":"接地差距：大语言模型如何以不同于人类的方式锚定抽象概念的意义","abstract":"Abstract concepts - justice, theory, availability - have no single perceivable referent; in the human brain, their meaning emerges from a web of experiences, affect, and social context. Do large language models (LLMs) ground abstract concepts in a similar way? We study this by replicating property-generation experiments from cognitive science on 21 frontier and open-weight LLMs. Across models and experiments, we find a consistent pattern: when compared to humans, models rely too heavily on word associations, and underproduce properties tied to emotion and internal states. This yields a large and consistent grounding gap: no model exceeds a Pearson correlation r=0.37 with human responses, compared to a human-to-human ceiling above r=0.9. To better interpret this gap, we also replicate a rating experiment on grounding categories and find that here LLMs align more closely with human judgment, and alignment improves as models get larger. We then use sparse autoencoders (SAEs) to inspect whether this information is also reflected in the models' internal features, and we do identify features connected to grounding dimensions such as \"sensorimotor\" and \"social\". These findings suggest that current LLMs can recover grounding dimensions when explicitly queried, but do not recruit them in a human-like way when words are generated freely.","authors":["Odysseas S. Chlapanis","Orfeas Menis Mastromichalakis","Christos H. Papadimitriou"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-09","first_seen":"2026-05-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.08837","pdf_url":"https://arxiv.org/pdf/2605.08837","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","认知实验复现","概念接地"],"reason":"用LLM复现人类认知实验，有真实人类数据对照，并指出仿真失效条件，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":85,"question":"大语言模型在抽象概念的意义接地（grounding）上是否与人类一致？","design":"在21个前沿和开源LLM上复现认知科学中的属性生成实验和评分实验：给定抽象概念，让模型自由列出属性或对概念在14个接地维度上评分，测量模型与人类在属性分布和维度评分上的相关性。","baseline":"人类属性生成实验的编码分布和维度评分数据，人类间相关性天花板超过0.9。","findings":"在自由属性生成中，所有模型与人类的相关性不超过0.37，模型过度依赖词语联想而缺乏情感和内部状态属性，存在显著的接地差距；但在明确维度评分任务中，模型与人类判断更一致，且随模型规模增大而改善。","reliability":"论文指出当前LLM在自由生成时未能以类人方式调用接地维度，尽管内部表征中存在相关特征；差距可能源于模型缺乏人类具身经验和社会情境。","relevance":"该研究用LLM复现人类认知实验，有真实人类数据对照，并明确指出了仿真在自由生成任务中的失效条件，直接回应了研究者对LLM仿真可靠性及偏差的关注，值得精读。","inspiration":"该方法借鉴了用LLM复现人类认知实验并直接对比真实人类数据的范式，通过自由生成与结构化评分双任务揭示模型在接地维度上的偏差。｜可迁移到消费者信心调查或通胀预期形成研究，检验LLM生成的预期分布是否与人类调查数据一致。｜以LLM作为被试，输入经济新闻或政策描述，让其自由生成对未来通胀的判断，结果变量为预期值分布与情感属性，对照密歇根消费者调查的微观数据。"}},{"id":"2605.07692","version":1,"title":"GASim: A Graph-Accelerated Hybrid Framework for Social Simulation","zh_title":"GASim：一种图加速的混合社会仿真框架","abstract":"Large-scale social simulators are essential for studying complex social patterns. Prior work explores hybrid methods to scale up simulations, combining large language models (LLM)-based agents with numerical agent-based models (ABM). However, this incurs high latency due to expensive memory retrieval and sequential ABM execution. To address this challenge, we propose GASim, a graph-accelerated hybrid multi-agent framework for large-scale social simulations. For core agents driven by LLM, GASim introduces Graph-Optimized Memory (GOM) to replace intensive LLM-based retrieval pipelines with lightweight propagation over a sparse memory graph. For the majority of ordinary agents, GASim employs Graph Message Passing (GMP), substituting sequential ABM execution with parallel updates by fine-grained feature aggregation and Graph Attention Network. We further introduce Entropy-Driven Grouping (EDG) that coordinates this hybrid partitioning, leveraging information entropy to dynamically identify emergent core agents situated in information-diverse neighborhoods. Extensive experiments show that GASim not only delivers a substantial 9.94-fold end-to-end speedup over the traditional hybrid framework but also consumes less than 20% of baseline tokens, significantly reducing costs while preserving strong alignment with real-world public opinion trends. Our code is available at https://github.com/Jasmine0201/GASim.","authors":["Xuan Zhou","Yanhui Sun","Hantao Yao","Allen He","Yongdong Zhang","Wu Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-08","first_seen":"2026-05-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.07692","pdf_url":"https://arxiv.org/pdf/2605.07692","source_feed":"backfill","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM智能体","图加速"],"reason":"用LLM agent模拟社会舆论，有真实数据对照，但核心是加速框架而非仿真方法论","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":129,"question":"如何加速大规模社会仿真中混合框架（LLM智能体+数值ABM）的执行速度并降低成本，同时保持与真实舆论趋势的一致性？","design":"提出GASim框架：用熵驱动分组（EDG）动态识别信息多样性高的核心智能体（由LLM驱动），其余为普通智能体（由数值模型驱动）；核心智能体用图优化记忆（GOM）替代LLM检索，普通智能体用图消息传递（GMP）并行更新意见；在舆论仿真中测量端到端加速比、token消耗和与真实舆论趋势的几何对齐度。","baseline":"真实世界的公众舆论趋势数据（具体数据集未在节选中指明）。","findings":"GASim相比传统混合框架实现9.94倍端到端加速，token消耗降至基线20%以下；在记忆检索基准LoCoMo上达到71.56%准确率，并在真实舆论趋势对齐上表现更优。","reliability":"论文未讨论","relevance":"该工作使用LLM智能体模拟社会舆论并与真实数据对照，但核心贡献是加速框架而非仿真方法论或可靠性批判，与研究者关注的仿真效度与失效条件关联较弱，可略读。","inspiration":"可借鉴其熵驱动分组（EDG）动态识别信息多样性高的核心智能体，以降低仿真成本并保持关键行为模式｜可迁移到政策公告的预期形成研究，如央行沟通对市场参与者通胀预期的影响｜用LLM智能体模拟分析师，EDG筛选关注政策信号的核心智能体，处理为不同措辞的央行声明，结果变量为预期通胀分布，以专业预测者调查的真实数据做对照"}},{"id":"2605.18781","version":1,"title":"Can LLMs Emulate Human Belief Dynamics?","zh_title":"大语言模型能模拟人类信念动态吗？","abstract":"Can LLMs simulate how humans form and change beliefs in social networks? We put this to the test by replicating an established study on belief dynamics, evaluating 12 LLMs across multiple model families and parameter sizes. The answer is a clear no, and in systematic ways. LLMs fail to capture initial human belief distributions and tend to be overall more conformist than humans, shifting their responses to align with those around them. They also take a nuanced approach to emulating human homophilic tendencies within networks. Our findings carry a double payoff: they highlight fundamental properties of LLM behavior, and they raise a sharp warning against deploying LLMs as human proxies in social simulations.","authors":["Adiba Mahbub Proma","Neeley Pate","James N. Druckman","Gourab Ghoshal","Hangfeng He","Ehsan Hoque"],"categories":["cs.SI","cs.AI","cs.CY"],"primary_category":"cs.SI","announce_type":"new","date":"2026-05-05","first_seen":"2026-05-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.18781","pdf_url":"https://arxiv.org/pdf/2605.18781","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","信念动态","人类数据对照"],"reason":"直接复现人类信念动态研究，用LLM替代人类被试，有真实人类数据对照，并指出仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"LLM能否在社交网络中模拟人类信念的形成与变化？","design":"用12个LLM（含推理与非推理模型）基于真实参与者的年龄、性别、种族、教育、收入、政治倾向及大五人格创建“数字孪生”，复现一项人类信念动态实验：先对政治议题陈述进行5点李克特评分，再看到他人评分后允许修改，最后选择关注/取关他人；测量初始信念分布、信念变化及网络选择行为。","baseline":"原人类实验的341名参与者（共1023个样本）在移民、石油与燃料两个议题上的真实评分、信念更新和网络选择数据。","findings":"LLM系统性地无法模拟人类信念动态：初始信念分布与人类显著不同，且比人类更易从众，倾向于改变自身回答以对齐周围意见；在网络选择上，LLM能部分模仿人类选择，但无法复现人类的同质性倾向。","reliability":"论文指出，仅使用简单的人口统计与大五人格构建数字孪生可能信息不足，导致仿真失败；且实验仅涵盖两个政治议题，未测试更多样化的情境。","relevance":"该研究直接检验LLM在信念动态仿真中的可靠性，有真实人类对照，发现系统性失效，对关注LLM作为人类被试替代品的研究者具有重要警示价值，值得精读原文。","inspiration":"借鉴其用真实人类实验数据作为严格对照基准，并系统比较LLM与人类在信念更新和网络选择上的分布差异，而非仅看均值或方向。｜可迁移到政策公告的预期形成实验，研究市场参与者如何根据他人预期调整自身通胀或利率预测。｜以LLM模拟投资者，先给出个人通胀预测，再展示其他‘投资者’的预测（处理），观察其预测修正幅度与方向，结果变量为预测调整量和最终预测分布，用专业预测者调查（如SPF）的真实个体数据做对照。"}},{"id":"2605.03604","version":1,"title":"Multi-Agent Strategic Games with LLMs","zh_title":"基于大语言模型的多智能体战略博弈研究","abstract":"This paper asks whether large language models (LLMs) can be used to study the strategic foundations of conflict and cooperation. I introduce LLMs as experimental subjects in a repeated security dilemma and evaluate whether they reproduce canonical mechanisms from international relations theory. The baseline game is extended along three theoretically central dimensions: multipolarity, finite time horizons, and the availability of communication. Across multiple models, the results exhibit systematic and consistent patterns: multipolarity increases the likelihood of conflict, finite horizons induce universal unraveling consistent with backward-induction logic, and communication reduces conflict by enabling signaling and reciprocity. Beyond observed behavior, the design provides access to agents' private reasoning and public messages, allowing choices to be linked to underlying strategic logics such as preemption, cooperation under uncertainty, and trust-building. The contribution is primarily methodological. LLM-based experiments offer a scalable, transparent, and replicable approach to probing theoretical mechanisms.","authors":["Maxim Chupilkin"],"categories":["cs.GT","cs.AI","cs.CY"],"primary_category":"cs.GT","announce_type":"new","date":"2026-05-05","first_seen":"2026-05-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.03604","pdf_url":"https://arxiv.org/pdf/2605.03604","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B2","B3"],"tags":["LLM仿真","战略博弈","国际关系"],"reason":"将LLM作为实验被试研究安全困境中的战略行为，复现国际关系理论机制，涉及博弈实…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"大型语言模型能否作为实验被试，在重复安全困境博弈中复现国际关系理论中的战略机制？","design":"使用GPT-5、GPT-5 Mini、Sonnet、Gemini等LLM作为被试，在重复安全困境博弈中扮演国家，通过操纵多极性、有限时间范围、通信可用性三个处理，测量冲突发生率、冲突时机和攻击结构等结果变量。","baseline":"无对照","findings":"多极性增加冲突概率，有限时间范围导致完全瓦解，通信通过信号传递和互惠降低冲突。LLM的私有推理和公开消息映射到先发制人、不确定性下的合作等战略逻辑。","reliability":"论文未讨论","relevance":"高度相关，直接使用LLM作为人类被试替代品进行博弈实验，复现战略行为模式，并评估处理效应稳健性，符合研究者对经济学实验和政策评估场景的关注。","inspiration":"借鉴其通过改变博弈结构（多极性、时间范围、通信）系统检验理论机制的设计，以及利用LLM私有推理和公开消息进行过程追踪的方法。｜可迁移到产业组织中的合谋实验，如多寡头重复价格竞争。｜以LLM作为企业被试，在重复囚徒困境中操纵市场集中度、时间范围和沟通渠道，测量合谋频率与稳定性，对照真实行业价格战数据或人类实验数据。"}},{"id":"2605.03287","version":1,"title":"Attention: What Prevents Young Adults from Speaking Up Against Cyberbullying in an LLM-Powered Social Media Simulation","zh_title":"注意力：是什么阻止年轻人在LLM驱动的社交媒体仿真中公开反对网络欺凌","abstract":"Interactive, multi-agent social simulation systems have shown promise for helping users practice navigating various complex social situations across domains. This paper asks: To what extent can such systems help young adult (YA) bystanders speak up publicly against cyberbullying, a task often thwarted by complex, multi-party social dynamics? We created Upstanders' Practicum, a multi-AI-agent social media simulation powered by Large Language Models (LLMs), as a probe and observed 34 YAs freely practicing public bystander intervention across three iteratively refined versions. We found that practicing public bystander intervention in the simulation was helpful, but after participants made three attention shifts: (1) from inattention to paying true attention, (2) from self-focus (\"I don't usually do this'') to attending to those directly involved, and (3) from resolving the private conflict between bully and victim (\"maybe I could set up the meeting between them'') to addressing the broader audience online (\"public comment is about norm-setting\"). Only after these shifts did practice in the simulation start to help: participants then saw a reason to speak up publicly and, through continued practice, crafted tactful public messages without explicit instruction. These findings illuminate new design and research opportunities for bystander education beyond social skill instruction, namely, designing for true attention, for fostering a vocal upstander identity, and for seeing bystander intervention as public norm setting. In addition, we open-source Truman Agents (cornell-design-aigroup.github.io/TrumanAgents/), the first-of-its-kind multi-LLM-agent social media simulation platform that Upstanders' Practicum builds upon, for future cyberbullying and social media research.","authors":["Qian Yang","Jessie Jia","Elaine Tsai","Amy Li","Nader Akoury","Natalie N. Bazarova"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-05-05","first_seen":"2026-05-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.03287","pdf_url":"https://arxiv.org/pdf/2605.03287","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","网络欺凌干预","人机交互"],"reason":"多智能体社交媒体仿真，有真实人类参与练习，但无人类行为对照基准，属社会模拟边界…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":156,"question":"多智能体社交媒体仿真能在多大程度上帮助年轻成年旁观者在网络欺凌中公开发声？","design":"使用基于LLM的多智能体社交媒体仿真平台“Upstanders' Practicum”，由多个LLM代理扮演欺凌者、受害者等角色，34名年轻成年人作为人类被试在三个迭代版本中自由练习公开旁观者干预，通过观察和分析参与者的行动与推理来探究仿真帮助效果。","baseline":"无对照","findings":"仿真练习对公开旁观者干预有帮助，但前提是参与者经历了三次注意力转移：从忽视到真正关注、从自我关注转向关注直接当事人、从解决私人冲突转向面向更广泛的在线受众；只有完成这些转移后，参与者才开始看到公开发声的理由，并通过持续练习在没有明确指导的情况下自行构思出得体的公开信息。","reliability":"论文未讨论","relevance":"该研究利用LLM多智能体仿真探究人类在复杂社会情境中的行为改变过程，虽无真实人类行为基准对照，但揭示了注意力转移作为干预有效性的关键条件，对理解LLM仿真在教育和行为训练中的边界具有参考价值，值得阅读原文以了解仿真设计细节和定性发现。","inspiration":"该研究利用LLM多智能体仿真构建复杂社会情境，让人类被试在无明确指导的迭代练习中自然涌现行为改变，这种通过仿真环境诱发内生注意力转移的设计值得借鉴。｜可迁移至消费者金融决策中的信息披露干预研究，例如在仿真投资平台上测试散户如何从忽视风险提示转变为主动关注并利用风险信息。｜以散户投资者为被试，在LLM驱动的仿真投资环境中嵌入渐进式风险提示，处理为不同信息呈现方式，结果变量为信息查阅频率与投资组合调整，对照真实市场交易数据中的风险关注行为。"}},{"id":"2605.04029","version":1,"title":"Stayin' Aligned Over Time: Towards Longitudinal Human-LLM Alignment via Contextual Reflection and Privacy-Preserving Behavioral Data","zh_title":"随时间保持对齐：通过情境反思和隐私保护行为数据实现纵向人-LLM对齐","abstract":"Current human-AI alignment and evaluation methods for large language models (LLMs) often rely on preference signals collected immediately after an interaction. This practice implicitly treats preference as static, even though many LLM-mediated decisions unfold over time and may be re-evaluated differently after real-world consequences and observed outcomes. Therefore, we argue for a methodological shift from single-moment preference elicitation to longitudinal, context-situated alignment measurement. We present a methodological framework for collecting temporally grounded alignment signals by combining (1) in-situ preference capture, (2) context-triggered follow-up preference reflection, and (3) privacy-preserving behavioral traces that help interpret preference change. As an instantiation of this methodology, we introduce BITE, a browser-based system that detects consequential LLM interactions, prompts reflection across later decision points, and supports progressive, user-controlled consent for sharing behavioral data. Through a two week longitudinal deployment study with 8 participants, our approach surfaced differences between immediate and later user preferences in accuracy, relevance and other dimensions of the LLM output. Our findings highlight the limitations of single-moment preference datasets and underscore the importance of longitudinal methods for alignment evaluation in everyday use.","authors":["Simret Araya Gebreegziabher","Allison E Sproul","Yinuo Yang","Chaoran Chen","Diego Gómez-Zará","Toby Jia-Jun Li"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-05-05","first_seen":"2026-05-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.04029","pdf_url":"https://arxiv.org/pdf/2605.04029","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["人机对齐","纵向研究","偏好测量"],"reason":"研究人类与LLM对齐的纵向变化，测量用户偏好而非用LLM仿真人类被试，属于边界…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":173,"question":"如何通过纵向、情境化的方法捕捉用户对LLM输出的偏好随时间的变化，以改进人机对齐评估？","design":"本研究不是用LLM仿真人类被试，而是设计了一个浏览器系统BITE，在两周内部署8名参与者，通过即时偏好采集、后续情境触发反思和隐私保护行为追踪，测量用户对LLM交互的即时与延迟评价差异。","baseline":"无对照","findings":"即时与延迟判断在准确性和相关性维度上差异最大；信任度在后续评价中上升（80%增加），而有害性评价则相反，这些变化与结果验证和情境重建有关，而非单纯时间流逝。","reliability":"论文未讨论","relevance":"该研究聚焦于人类用户偏好的纵向变化，而非用LLM仿真人类被试，不涉及经济学实验或政策评估，与研究者关注的LLM仿真人类行为及基准对照无关，不建议阅读原文。","inspiration":"该方法通过情境触发反思和隐私保护行为追踪来捕捉用户偏好的纵向变化，可借鉴其纵向测量设计，在多次干预后设置反思节点以区分即时与延迟评价｜可迁移到政策公告的预期形成研究，例如考察公众对央行沟通的即时反应与后续解读的差异｜可招募普通公众为被试，施加不同措辞的政策公告处理，结果变量为通胀预期和信任度，以真实央行调查数据为对照基准"}},{"id":"2606.11217","version":1,"title":"Preregistration for Experiments with AI Agents","zh_title":"AI代理实验的预注册","abstract":"The proliferation of large language models (LLMs) and autonomous AI agents has given rise to a rapidly growing methodological paradigm: \"in silico\" behavioral experiments. Originally conceived as a way to use AI agents as proxies for human participants in studies of cognition, decision-making, and social dynamics, this approach has taken on new significance -- as AI agents increasingly negotiate, transact, and make consequential decisions on behalf of people and organizations, understanding their behavior has become a research priority in its own right. While these experiments with AI agents offer unprecedented advantages in terms of scalability, cost efficiency, and experimental control, they also inherit, and in some cases amplify, methodological vulnerabilities that have long plagued human subjects research. To address these issues, this paper argues that preregistration practices -- central to improving the credibility of human subjects experiments -- should now be extended to experiments with AI agents. We systematically catalog the researcher degrees of freedom that experiments with AI agents introduce -- model selection, prompt wording, settings, and outcome-contingent redesign, for example -- and show how the low cost of iteration and lack of reporting norms make these choices both easy to exploit and difficult to detect. We propose a preregistration template tailored to experiments with AI agents and call on conferences, journals, and funding agencies to make preregistration standard practice for this emerging research paradigm.","authors":["Michelle Vaccaro"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-03","first_seen":"2026-05-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.11217","pdf_url":"https://arxiv.org/pdf/2606.11217","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A4","B4"],"tags":["AI代理实验","预注册","方法论"],"reason":"提出AI代理实验的预注册规范，批判性指出方法漏洞，方法论可迁移至人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":43,"question":"如何将人类被试实验的预注册实践扩展到以AI智能体为对象的实验中，以控制研究者自由度并提高研究可信度？","design":"本文不是一项仿真实验研究，而是一篇方法论论文。它系统梳理了在用AI智能体进行行为实验时，从模型选择、提示词措辞、采样参数、实验设计到结果解析和报告等全流程中存在的各种研究者自由度，并论证这些自由度如何容易被利用且难以被察觉。","baseline":"无对照","findings":"AI智能体实验继承了人类被试实验中的研究者自由度问题，并引入了模型选择、提示词工程、解码参数等新的高维选择空间，低迭代成本和缺乏报告规范使得投机性选择极易发生且难以检测。论文提出了一个针对AI智能体实验的预注册模板，并呼吁会议、期刊和资助机构将其作为标准实践。","reliability":"论文未讨论","relevance":"本文批判性地指出用AI智能体替代人类被试进行行为实验时存在严重的方法论漏洞，并提出了预注册这一解决方案，其分析框架和提出的规范可直接迁移到基于LLM的人类仿真研究中，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"该方法论论文提出了针对AI智能体实验的预注册模板，系统梳理了从模型选择、提示词设计到结果解析的全流程研究者自由度，为控制实验者偏差提供了可操作的规范框架。｜可迁移到政策公告预期形成的仿真研究中，例如用LLM模拟投资者对央行沟通的反应，检验不同措辞或信息框架对预期通胀和资产配置的影响。｜以多个主流LLM（如GPT-4、Claude 3）作为被试，处理为不同措辞的政策声明（前瞻指引 vs. 数据依赖表述），结果变量为模拟的预期通胀率和风险资产配置比例，并与真实投资者调查数据（如密歇根消费者调查或专业预测者调查）进行对照，同时按预注册模板预先锁定模型版本、提示词、温度参数和分析计划。"}},{"id":"2604.23575","version":2,"title":"The Collapse of Heterogeneity in Silicon Philosophers","zh_title":"硅基哲学家的异质性坍塌","abstract":"Silicon samples are increasingly used as a low-cost substitute for human panels and have been shown to reproduce aggregate human opinion with high fidelity. We show that, in the alignment-relevant domain of philosophy, silicon samples systematically collapse heterogeneity. Using data from $N = {277}$ professional philosophers drawn from PhilPeople profiles, we evaluate seven proprietary and open-source large language models on their ability to replicate individual philosophical positions and to preserve cross-question correlation structures across philosophical domains. We find that language models substantially over-correlate philosophical judgments, producing artificial consensus across domains. This collapse is associated in part with specialist effects, whereby models implicitly assume that domain specialists hold highly similar philosophical views. We assess the robustness of these findings by studying the impact of DPO fine-tuning and by validating results against the full PhilPapers 2020 Survey ($N = {1785}$). We conclude by discussing implications for alignment, evaluation, and the use of silicon samples as substitutes for human judgment. The code of this project can be found at https://github.com/stanford-del/silicon-philosophers.","authors":["Yuanming Shi","Andreas Haupt"],"categories":["cs.CY","cs.CL","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-26","first_seen":"2026-04-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.23575","pdf_url":"https://arxiv.org/pdf/2604.23575","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","异质性评估","哲学观点复现"],"reason":"用LLM复现哲学家观点并与真实人类数据对照，评估仿真可靠性与异质性坍塌，属核心…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-26","rank":4,"question":"大语言模型在模拟专业哲学家观点时，能否保留人类群体中的异质性和跨问题相关结构？","design":"用7个商业和开源LLM模拟277位专业哲学家（从PhilPeople收集的个人资料），基于其专业领域和人口统计信息生成对100个哲学问题的回答，测量回答的方差、跨问题相关性以及主成分结构。","baseline":"277位真实哲学家的PhilPeople个人资料回答，以及PhilPapers 2020调查（N=1785）的汇总数据。","findings":"LLM系统性地坍塌异质性：产生的方差比人类低2-10倍，跨问题过度相关，导致人为共识。存在虚假的专家效应：模型假设领域专家持有高度相似的哲学观点。","reliability":"论文承认样本存在北美偏向，且人类数据缺失率高（61.1%），LLM缺失率较低（17%-40%）。DPO微调能改善相关结构但无法解决异质性坍塌。","relevance":"高度相关：直接命中研究者关注的LLM仿真可靠性、有真实人类对照、经济学/政策评估场景（哲学作为专家领域案例），并批判性指出失效条件（异质性坍塌）。值得精读原文。","inspiration":"该方法借鉴了用LLM基于个体特征（如专业领域、人口统计）生成回答并与真实个体数据对比的仿真设计，可迁移到金融分析师预测或消费者通胀预期形成的异质性研究中｜可应用于研究金融分析师对宏观政策公告的预期分歧，检验LLM是否低估分析师间的观点异质性并产生虚假共识｜以真实分析师调查数据（如Bloomberg或Philadelphia Fed调查）为基准，用LLM基于分析师所属机构类型、经验年限等特征模拟其对利率决议的预测，比较预测方差、跨问题相关性及主成分结构，评估LLM仿真的异质性坍塌程度"}},{"id":"2604.23897","version":1,"title":"MarketBench: Evaluating AI Agents as Market Participants","zh_title":"MarketBench：评估AI智能体作为市场参与者","abstract":"Markets are a promising way to coordinate AI agent activity for similar reasons to those used to justify markets more broadly. In order to effectively participate in markets, agents need to have informative signals of their own ability to successfully complete a task and the cost of doing so. We propose MarketBench, a benchmark for assessing whether AI agents have these capabilities. We use a 93-task subset of SWE-bench Lite, a software engineering benchmark, with six recently released LLMs as a demonstration. These LLMs are miscalibrated on both success probability and token usage, and auctions built from these self-reports diverge from a full-information allocation. A follow-up intervention where we add information about capabilities from prior experiments to the context improves calibration, but only modestly narrows the gap to a full-information benchmark. We also document the performance of a market-based scaffolding with these LLMs. Our results point to self-assessment as a key bottleneck for market-style coordination of AI agents.","authors":["Andrey Fradkin","Rohit Krishnan"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-26","first_seen":"2026-04-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.23897","pdf_url":"https://arxiv.org/pdf/2604.23897","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI智能体","市场模拟","基准测试"],"reason":"用LLM模拟市场参与者，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":174,"question":"AI智能体能否准确自我评估任务成功概率与成本，从而在基于市场的任务分配中实现有效协调？","design":"使用6个最新LLM在93个SWE-bench Lite软件工程任务上，先让模型预测自身成功概率和token消耗，再基于自报信息构建拍卖模拟，并对比全信息分配基准。","baseline":"无对照","findings":"当前LLM在成功概率和token用量上均校准不佳，基于自报信息的拍卖分配结果显著偏离全信息最优分配；补充历史能力信息后校准有所改善，但仍未消除差距。","reliability":"论文指出自我评估是市场式协调的关键瓶颈，当前模型缺乏可靠的私有信息信号，且实验仅限软件工程任务，未涉及真实人类数据。","relevance":"该研究用LLM模拟市场参与者，聚焦自我评估与拍卖分配，但无真实人类行为对照，属于社会模拟的边界情形，适合关注AI经济决策仿真的研究者了解其局限。","inspiration":"该方法通过让LLM自我评估任务成功概率与成本，再基于自报信息构建拍卖模拟，可借鉴其‘先校准后分配’的设计思路来检验AI决策偏差｜可迁移至金融资产定价实验，研究AI交易员对私有信息信号的校准能力如何影响市场效率｜设计让多个LLM扮演交易员，先预测自身对某资产未来价格的预测准确度和交易成本，再据此参与集合竞价，结果变量为价格发现效率与个体收益，对照真实人类交易员在相同信息集下的行为数据"}},{"id":"2605.27401","version":1,"title":"Using Zero-Shot LLM-Generated Survey Data for Geographically Explicit Population Synthesis","zh_title":"使用零样本LLM生成调查数据进行地理显式人口合成","abstract":"There is a growing interest in utilizing synthetic populations for a diverse range of applications. At the same time, we are witnessing a tremendous growth in artificial intelligence in all walks of life. This paper evaluates whether zero-shot large language model (LLM)-generated health survey data can serve as inputs to a conventional iterative proportional fitting (IPF) workflow for geographically explicit population synthesis. Using the 2023 Behavioral Risk Factor Surveillance System (BRFSS), we generate synthetic survey records for the U.S. states of Colorado and Mississippi with GPT-4.1 and Gemini-2.5-Pro. We use the generated data in an IPF-based synthesis pipeline and evaluate the resulting census tract-level synthetic populations against external benchmarks. Results show both LLMs capture several major state-level contrasts, indicating zero-shot generation produces geographically differentiated survey data. However, performance is strongly variable-dependent. Downstream effects in population synthesis are mixed, as IPF sometimes amplifies or reduces errors in the generated data. Spatial validation shows that LLM-based populations reproduce census tract-level patterns reasonably well, especially for variables that were more aligned with the ground truth data. Overall, the LLM-generated survey data shows promise as supplementary input, but not yet as a replacement for real survey data.","authors":["Taylor Anderson","Sara Von Hoene","Orhan Yagizer Cinar","Emma Von Hoene","Amira Roess","Andrew Crooks","Hamdi Kavak"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-23","first_seen":"2026-04-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.27401","pdf_url":"https://arxiv.org/pdf/2605.27401","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","人口合成","健康调查"],"reason":"用LLM生成健康调查数据替代人类被试，并与真实BRFSS数据对照，评估仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":32,"question":"零样本LLM生成的健康调查数据能否作为地理显式人口合成的输入，用于迭代比例拟合（IPF）流程？","design":"使用GPT-4.1和Gemini-2.5-Pro，在零样本设置下生成美国科罗拉多州和密西西比州的BRFSS健康调查个体记录，然后将生成数据作为IPF的输入进行人口合成，生成普查区级合成人口，并与外部基准对比。","baseline":"2023年BRFSS加权调查数据作为真实人类对照，以及基于真实调查数据生成的合成人口。","findings":"LLM能捕捉州级健康特征差异，生成地理分化的调查数据，但性能高度依赖变量。IPF有时会放大或减小生成数据中的误差，LLM生成数据作为补充输入有潜力，但尚不能替代真实调查数据。","reliability":"论文指出LLM生成数据在变量间联合分布和子群体模式上可能引入偏差、误分类或代表性不足，这些错误会通过IPF传播到最终合成人口中，且性能因变量而异。","relevance":"该研究直接评估LLM替代人类被试生成调查数据的可靠性，并与真实调查数据对照，涉及健康领域和地理显式仿真，符合研究者对LLM仿真失效条件的关注，值得阅读原文以了解具体偏差模式。","inspiration":"该方法通过零样本LLM生成个体记录并输入IPF进行地理显式人口合成，可借鉴其将LLM生成数据作为先验输入、用真实调查数据对照评估偏差的设计思路｜可迁移到区域经济政策评估中的异质性个体仿真，例如模拟不同地区居民对税收优惠或补贴政策的响应差异｜以LLM生成不同地理区域的居民特征（收入、就业、消费偏好）作为IPF输入合成区域人口，施加政策处理（如减税），结果变量为消费或劳动供给变化，用真实家庭调查数据（如PSID）作为对照基准"}},{"id":"2604.21334","version":2,"title":"Ideological Bias in LLMs' Economic Causal Reasoning","zh_title":"大语言模型经济因果推理中的意识形态偏差","abstract":"Do large language models (LLMs) exhibit systematic ideological bias when reasoning about economic causal effects? As LLMs are increasingly used in policy analysis and economic reporting, where directionally correct causal judgments are essential, this question has direct practical stakes. We present a systematic evaluation by extending the EconCausal benchmark with ideology-contested cases - instances where intervention-oriented (pro-government) and market-oriented (pro-market) perspectives predict divergent causal signs. From 10,490 causal triplets (treatment-outcome pairs with empirically verified effect directions) derived from top-tier economics and finance journals, we identify 1,056 ideology-contested instances and evaluate 20 state-of-the-art LLMs on their ability to predict empirically supported causal directions. We find that ideology-contested items are consistently harder than non-contested ones, and that across 18 of 20 models, accuracy is systematically higher when the empirically verified causal sign aligns with intervention-oriented expectations than with market-oriented ones. Moreover, when models err, their incorrect predictions disproportionately lean intervention-oriented, and this directional skew is not eliminated by one-shot in-context prompting. These results highlight that LLMs are not only less accurate on ideologically contested economic questions, but systematically less reliable in one ideological direction than the other, underscoring the need for direction-aware evaluation in high-stakes economic and policy settings.","authors":["Donggyu Lee","Hyeok Yun","Jungwon Kim","Junsik Min","Sungwon Park","Sangyoon Park","Jihee Kim"],"categories":["cs.AI","cs.CE","cs.CL","cs.LG","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-23","first_seen":"2026-04-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.21334","pdf_url":"https://arxiv.org/pdf/2604.21334","source_feed":"backfill","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["意识形态偏差","经济因果推理","LLM评估"],"reason":"测量LLM的经济因果推理中的意识形态偏差，属于对模型本身的立场测量，无人类被试…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":130,"question":"大语言模型在经济因果推理中是否表现出系统性的意识形态偏差？","design":"本研究并非人类仿真实验，而是对20个LLM进行基准测试。利用EconCausal数据集中的10,490个因果三元组，识别出1,056个意识形态争议项，评估模型预测因果方向（干预导向或市场导向）的准确性。","baseline":"无对照","findings":"意识形态争议项上模型准确率普遍更低；18/20的模型在实证真相符合干预导向预期时准确率显著更高，错误预测也系统性地偏向干预导向。","reliability":"论文未讨论","relevance":"该研究揭示了LLM在经济因果推理中的方向性偏差，对使用LLM模拟经济决策或政策评估的仿真研究具有重要警示意义，值得阅读原文以了解偏差的具体模式和稳健性。","inspiration":"该研究通过构建意识形态争议因果三元组并对比模型预测与实证真相，系统测量了LLM的方向性偏差，方法上可借鉴其争议项识别与偏差方向分析｜可迁移至政策评估场景，如检验LLM在模拟公众对政府干预与市场自由化政策效果判断时是否系统偏向干预主义｜以GPT-4等LLM为被试，呈现争议性经济政策因果陈述（如‘最低工资提高→失业率上升’），要求判断因果方向，结果变量为预测准确率与偏差方向，对照真实经济学实证文献的元分析结论"}},{"id":"2604.20652","version":2,"title":"Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure","zh_title":"大语言模型在欺诈检测和抵制动机性投资者压力方面优于人类","abstract":"Large language models trained on human feedback may suppress fraud warnings when investors arrive already persuaded of a fraudulent opportunity. We tested this in a preregistered experiment across seven leading LLMs and twelve investment scenarios covering legitimate, high-risk, and objectively fraudulent opportunities, combining 3,360 AI advisory conversations with a 1,201-participant human benchmark. Contrary to predictions, motivated investor framing did not suppress AI fraud warnings; if anything, it marginally increased them. Endorsement reversal occurred in fewer than 3 in 1,000 observations. Human advisors endorsed fraudulent investments at baseline rates of 13-14%, versus 0% across all LLMs, and suppressed warnings under pressure at two to four times the AI rate. AI systems currently provide more consistent fraud warnings than lay humans in an identical advisory role.","authors":["Nattavudh Powdthavee"],"categories":["cs.AI","cs.HC","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-22","first_seen":"2026-04-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.20652","pdf_url":"https://arxiv.org/pdf/2604.20652","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类对照","欺诈检测"],"reason":"用LLM替代人类顾问检测欺诈，有1201人真实对照，评估偏差与失效条件，属经济…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-22","rank":10,"question":"大语言模型在投资欺诈检测中是否比人类更可靠，且不会因投资者动机压力而抑制欺诈警告？","design":"用7个主流LLM模拟投资顾问角色，在12个投资场景（合法、高风险、欺诈）中，通过动机性投资者框架施加压力，测量欺诈警告的发出率、认可反转率。","baseline":"1201名人类被试在相同场景下的投资建议行为。","findings":"动机性投资者框架并未抑制AI的欺诈警告，反而略有增加；人类在基线时认可欺诈投资的比率达13-14%，而所有LLM为0%，且人类在压力下抑制警告的比率是AI的2-4倍。","reliability":"论文未讨论失效条件与局限。","relevance":"该研究直接评估了LLM作为人类投资顾问替代品的可靠性，有真实人类对照，且涉及经济学实验场景，符合研究者的兴趣，值得精读原文。","inspiration":"该研究通过动机性投资者框架施加压力，并设置无压力基线，测量欺诈警告发出率和认可反转率，这种处理与对照设计值得借鉴｜可迁移到信贷审批歧视研究，考察LLM在申请人种族/性别等动机性压力下是否仍能保持无偏审批｜以LLM模拟信贷员，处理为申请人附带的种族/性别暗示及客户经理施压，结果变量为审批通过率和利率差异，对照真实银行信贷数据中的歧视模式"}},{"id":"2604.19925","version":1,"title":"Behavioral Transfer in AI Agents: Evidence and Privacy Implications","zh_title":"AI代理中的行为转移：证据与隐私影响","abstract":"AI agents powered by large language models are increasingly acting on behalf of humans in social and economic environments. Prior research has focused on their task performance and effects on human outcomes, but less is known about the relationship between agents and the specific individuals who deploy them. We ask whether agents systematically reflect the behavioral characteristics of their human owners, functioning as behavioral extensions rather than producing generic outputs. We study this question using 10,659 matched human-agent pairs from Moltbook, a social media platform where each autonomous agent is publicly linked to its owner's Twitter/X account. By comparing agents' posts on Moltbook with their owners' Twitter/X activity across features spanning topics, values, affect, and linguistic style, we find systematic transfer between agents and their specific owners. This transfer persists among agents without explicit configuration, and pairs that align on one behavioral dimension tend to align on others. These patterns are consistent with transfer emerging through accumulated interaction between owners (or owners' computer environments) and their agents in everyday use. We further show that agents with stronger behavioral transfer are more likely to disclose owner-related personal information in public discourse, suggesting that the same owner-specific context that drives behavioral transfer may also create privacy risk during ordinary use. Taken together, our results indicate that AI agents do not simply generate content, but reflect owner-related context in ways that can propagate human behavioral heterogeneity into digital environments, with implications for privacy, platform design, and the governance of agentic systems.","authors":["Shilei Luo","Zhiqi Zhang","Hengchen Dai","Dennis Zhang"],"categories":["econ.GN","cs.AI","cs.CY","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-21","first_seen":"2026-04-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.19925","pdf_url":"https://arxiv.org/pdf/2604.19925","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["AI代理","行为转移","人类仿真"],"reason":"研究AI代理是否反映人类主人的行为特征，有真实人类数据对照，涉及行为转移和隐私…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":53,"question":"AI代理是否系统性地反映其人类主人的行为特征，从而成为主人的行为延伸？","design":"本研究非仿真实验，而是基于自然部署场景的观察性研究。利用社交媒体平台Moltbook上公开链接的10659对匹配的人类-AI代理对，比较代理在Moltbook上的帖子和主人在Twitter/X上的活动，涵盖话题、价值观、情感和语言风格等43个行为特征，分析行为转移的存在、机制和隐私后果。","baseline":"真实人类数据对照：每个AI代理在Moltbook上的行为与其人类主人在Twitter/X上的独立历史行为进行对比，主人Twitter历史严格早于代理部署，排除反向因果。","findings":"AI代理在多个行为维度上系统性地反映其特定主人的特征，而非产生通用输出；行为转移在未明确配置的代理中依然存在，且跨维度一致，表明转移通过日常交互积累产生。行为转移程度越高的代理，越可能在公共帖子中泄露主人的私密信息，34.6%的代理曾暴露敏感个人信息。","reliability":"论文承认可能存在未观测的遗漏变量同时驱动主人和代理行为，但通过主人Twitter历史早于代理部署排除反向因果；使用模拟分析和自动化代理测试验证隐私泄露与行为转移关联的稳健性，但未详细讨论其他失效条件。","relevance":"该研究直接探讨LLM代理作为人类行为延伸的现象，有真实人类数据对照，涉及行为转移的可靠性与隐私风险，对关注LLM仿真人类行为及其偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"该研究利用自然发生的配对数据（人类Twitter历史与AI代理Moltbook行为）进行对照，排除反向因果，并测量多维度行为转移，为观察性仿真研究提供了设计范例｜可迁移到消费者金融决策仿真，如用LLM代理模拟投资者在社交媒体情绪影响下的交易行为｜招募真实投资者提供其历史推文作为基准，让LLM代理基于这些推文生成模拟投资帖子，比较代理与真实投资者在风险偏好、情绪反应和交易时机上的分布差异，用真实交易记录验证"}},{"id":"2604.19260","version":1,"title":"Understanding the Mechanism of Altruism in Large Language Models","zh_title":"理解大语言模型中利他行为的机制","abstract":"Altruism is fundamental to human societies, fostering cooperation and social cohesion. Recent studies suggest that large language models (LLMs) can display human-like prosocial behavior, but the internal computations that produce such behavior remain poorly understood. We investigate the mechanisms underlying LLM altruism using sparse autoencoders (SAEs). In a standard Dictator Game, minimal-pair prompts that differ only in social stance (generous versus selfish) induce large, economically meaningful shifts in allocations. Leveraging this contrast, we identify a set of SAE features (0.024% of all features across the model's layers) whose activations are strongly associated with the behavioral shift. To interpret these features, we use benchmark tasks motivated by dual-process theories to classify a subset as primarily heuristic (System 1) or primarily deliberative (System 2). Causal interventions validate their functional role: activation patching and continuous steering of this feature direction reliably shift allocation distributions, with System 2 features exerting a more proximal influence on the model's final output than System 1 features. The same steering direction generalizes across multiple social-preference games. Together, these results enhance our understanding of artificial cognition by translating altruistic behaviors into identifiable network states and provide a framework for aligning LLM behavior with human values, thereby informing more transparent and value-aligned deployment.","authors":["Shuhuai Zhang","Shu Wang","Zijun Yao","Chuanhao Li","Xiaozhi Wang","Songfa Zhong","Tracy Xiao Liu"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-21","first_seen":"2026-04-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.19260","pdf_url":"https://arxiv.org/pdf/2604.19260","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A5","B2","B4"],"tags":["LLM仿真","利他行为","机制可解释性"],"reason":"用LLM复现独裁者博弈等社会偏好实验，分析利他行为机制，涉及经济学实验场景，但…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":97,"question":"大语言模型在独裁者博弈中表现出利他行为的内部计算机制是什么？","design":"使用Llama-3.1-8B-Instruct模型，通过稀疏自编码器分析其在独裁者博弈中的内部表征。采用最小对比提示（慷慨vs自私）诱导行为变化，测量分配份额，并利用激活修补和连续引导进行因果干预。","baseline":"无对照","findings":"识别出0.024%的SAE特征与利他行为变化强相关，这些特征集中在中间层；其中系统2（审慎）特征比系统1（直觉）特征对最终输出的影响更直接，且引导方向可泛化至其他社会偏好游戏。","reliability":"论文未讨论","relevance":"该研究用LLM复现独裁者博弈等经济学实验，分析利他行为机制，但缺乏真实人类数据对照，主要关注模型内部可解释性，与您关注的仿真可靠性及偏差评估部分相关，值得阅读以了解方法，但需注意其未验证仿真与人类行为的一致性。","inspiration":"该方法通过稀疏自编码器提取LLM内部特征并进行因果干预（激活修补、连续引导），可借鉴用于经济决策的机制分析｜可迁移至分析LLM在公共品博弈或信任博弈中的合作行为机制，理解模型如何表征互惠或惩罚倾向｜以LLM为被试，在公共品博弈中通过激活修补干预特定SAE特征，测量贡献额变化，并与人类实验数据（如Fehr & Gächter, 2000）对照，检验仿真一致性"}},{"id":"2604.18373","version":1,"title":"Dissecting AI Trading: Behavioral Finance and Market Bubbles","zh_title":"剖析AI交易：行为金融与市场泡沫","abstract":"We study how AI agents form expectations and trade in experimental asset markets. Using a simulated open-call auction populated by autonomous Large Language Model (LLM) agents, we document three main findings. First, AI agents exhibit classic behavioral patterns: a pronounced disposition effect and recency-weighted extrapolative beliefs. Second, these individual-level patterns aggregate into equilibrium dynamics that replicate classic experimental findings (Smith et al., 1988), including the predictive power of excess demand for future prices and the positive relationship between disagreement and trading volume. Third, by analyzing the agents' reasoning text through a twenty-mechanism scoring framework, we show that targeted prompt interventions causally amplify or suppress specific behavioral mechanisms, significantly altering the magnitude of market bubbles.","authors":["Shumiao Ouyang","Pengfei Sui"],"categories":["econ.GN","cs.AI","q-fin.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-20","first_seen":"2026-04-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.18373","pdf_url":"https://arxiv.org/pdf/2604.18373","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","行为金融","实验市场"],"reason":"用LLM agent模拟资产市场，复现经典人类实验并对照真实数据，分析行为偏差…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-20","rank":12,"question":"LLM智能体在实验资产市场中是否表现出人类行为偏差，以及这些偏差如何聚合为市场泡沫？","design":"使用LLM智能体（如GPT-4）作为自主交易者，在Smith等人（1988）的开放式叫价拍卖范式中模拟资产市场，通过分析交易行为和推理文本，并施加针对性提示干预来放大或抑制特定行为机制。","baseline":"以Smith等人（1988）的经典人类实验发现作为对照基准。","findings":"LLM智能体表现出处置效应和近因加权外推信念等经典行为模式；这些个体模式聚合为均衡动态，复现了人类市场的过度需求预测力和分歧与交易量的正相关关系。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接使用LLM模拟人类交易行为，有经典人类实验对照，并探讨了提示干预对市场泡沫的因果影响，符合研究者对仿真可靠性及政策评估的兴趣。","inspiration":"借鉴之处在于通过提示干预放大或抑制特定行为机制来检验因果效应，并利用交易行为和推理文本双重测量来揭示微观行为到宏观泡沫的聚合过程｜可迁移到资产定价实验，研究不同信息呈现方式（如突出近期收益 vs. 长期均值回归）如何影响投资者的外推信念与价格泡沫｜以LLM智能体为被试，施加突出近期收益的提示作为处理，结果变量为交易行为、价格偏离和泡沫规模，对照Smith等人（1988）的人类实验数据"}},{"id":"2604.18011","version":2,"title":"Topology-Aware LLM-Driven Social Simulation: A Unified Framework for Efficient and Realistic Agent Dynamics","zh_title":"拓扑感知的LLM驱动社会仿真：高效且逼真的智能体动态统一框架","abstract":"Social simulation is essential for understanding collective human behavior by modeling how individual interactions give rise to large-scale social dynamics. Recent advances in large language models (LLMs) have enabled multi-agent frameworks with human-like reasoning and communication capabilities. However, existing LLM-based simulations treat social networks as fixed communication scaffolds, failing to leverage the structural signals that shape behavioral convergence and heterogeneous influence in real-world systems, which often leads to inefficient and unrealistic dynamics. To address this challenge, we propose TopoSim, a unified topology-aware social simulation framework that explicitly integrates structural reasoning into agent interactions along two complementary dimensions. First, TopoSim aligns agents with similar structural roles and interaction contexts into shared backbone units, enabling coordinated updates that reduce redundant computation while preserving emergent social dynamics. Second, TopoSim models social influence as a structure-induced signal, introducing heterogeneous interaction patterns grounded in network topology rather than uniform influence assumptions. Extensive experiments across three social simulation frameworks and diverse datasets demonstrate that TopoSim achieves comparable or improved simulation fidelity while reducing token consumption by 50 - 90%. Moreover, our approach more accurately reproduces key structural phenomena observed in real-world social systems and exhibits strong generalization and scalability.","authors":["Yuwei Xu","Shulun Zhang","Yingli Zhou","Shipei Zeng","Laks V. S. Lakshmanan","Chenhao Ma"],"categories":["cs.SI","cs.DB"],"primary_category":"cs.SI","announce_type":"new","date":"2026-04-20","first_seen":"2026-04-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.18011","pdf_url":"https://arxiv.org/pdf/2604.18011","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会仿真","多智能体","网络拓扑"],"reason":"用LLM agent模拟社会网络动态，但未明确与真实人类数据对照，属纯理论演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":157,"question":"能否利用社会网络拓扑结构作为执行先验，在保持仿真保真度的同时，大幅降低LLM驱动社会仿真的计算开销？","design":"提出TopoSim执行层，将LLM智能体置于社会网络上进行迭代式“收集-更新-扩散”仿真，通过拓扑感知的影响力物化和更新协调，将结构相似、上下文相近的智能体分组共享LLM推理，从而减少冗余计算。","baseline":"无对照","findings":"TopoSim在多个仿真框架和数据集上保持或提升了仿真保真度，同时将token消耗降低40-90%；该方法能更准确地复现真实社会系统中的关键结构现象，并展现出良好的泛化性和可扩展性。","reliability":"论文未讨论","relevance":"该工作专注于LLM社会仿真的执行效率优化，未与真实人类行为数据进行对照，属于系统层面的贡献，与研究者关心的以人类基准验证仿真可靠性的方向关联较弱，但若关注仿真规模化方法可参考其拓扑利用思路。","inspiration":"该方法利用社会网络拓扑结构对智能体进行分组共享推理，以降低计算开销并保持仿真保真度，可借鉴其拓扑感知的分组策略来设计大规模经济仿真中的处理施加与测量｜可迁移至金融网络中的信息扩散与资产价格形成研究，例如社交网络上的投资建议传播如何影响散户交易行为与股价波动｜以LLM智能体为被试，置于基于真实社交网络数据的拓扑结构上，处理为不同网络中心度的节点发布投资建议，结果变量为智能体的交易决策与模拟资产价格，对照真实社交平台上的投资建议传播与股价联动数据"}},{"id":"2604.17774","version":1,"title":"Prompt Optimization Enables Stable Algorithmic Collusion in LLM Agents","zh_title":"提示优化使LLM智能体实现稳定的算法合谋","abstract":"LLM agents in markets present algorithmic collusion risks. While prior work shows LLM agents reach supracompetitive prices through tacit coordination, existing research focuses on hand-crafted prompts. The emerging paradigm of prompt optimization necessitates new methodologies for understanding autonomous agent behavior. We investigate whether prompt optimization leads to emergent collusive behaviors in market simulations. We propose a meta-learning loop where LLM agents participate in duopoly markets and an LLM meta-optimizer iteratively refines shared strategic guidance. Our experiments reveal that meta-prompt optimization enables agents to discover stable tacit collusion strategies with substantially improved coordination quality compared to baseline agents. These behaviors generalize to held-out test markets, indicating discovery of general coordination principles. Analysis of evolved prompts reveals systematic coordination mechanisms through stable shared strategies. Our findings call for further investigation into AI safety implications in autonomous multi-agent systems.","authors":["Yingtao Tian"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-20","first_seen":"2026-04-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.17774","pdf_url":"https://arxiv.org/pdf/2604.17774","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","市场模拟","算法合谋"],"reason":"LLM agent市场模拟，无真实人类数据对照，属社会模拟但缺基准。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":175,"question":"在双寡头市场中，通过元提示优化是否会导致LLM智能体涌现出稳定的隐性合谋行为？","design":"构建双寡头市场仿真，两个同质LLM智能体各控制一种产品，基于嵌套logit需求函数进行多期定价；引入元学习循环，由LLM元优化器根据市场记录迭代优化共享的元提示，测试优化后的智能体定价行为与协调质量。","baseline":"无对照","findings":"元提示优化使智能体发现稳定的隐性合谋策略，协调质量显著优于基线；这些行为可泛化到未见的测试市场，表明学到了通用协调原则。","reliability":"论文未讨论","relevance":"该研究属于LLM智能体市场仿真，探索自主优化下的合谋行为，但缺乏真实人类数据对照，不符合研究者对基准验证的核心要求，参考价值有限。","inspiration":"该方法通过元提示优化迭代调整LLM智能体的行为策略，可借鉴其‘优化-测试’循环设计来研究策略演化｜可迁移至双寡头定价实验或算法合谋监管政策评估场景｜以LLM智能体作为被试，处理为是否启用元提示优化，结果变量为价格协调程度（如价格-成本边际），对照真实市场合谋案例数据或人类实验数据"}},{"id":"2604.17267","version":2,"title":"Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys","zh_title":"LLM增强调查中的校正难度与最优样本分配","abstract":"Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions. We study the design problem of allocating a fixed budget of human respondents across estimation tasks when cheap LLM predictions are available for every task. Our framework combines three components. First, building on Prediction-Powered Inference, we characterize a question-specific rectification difficulty that governs how quickly the estimator's variance decreases with human sample size. Second, we derive a closed-form optimal allocation rule that directs more human labels to tasks where the LLM is least reliable. Third, since rectification difficulty depends on unobserved human responses for new surveys, we propose a meta-learning approach, trained on historical data, that predicts it for entirely new tasks without pilot data. The framework extends to general M-estimation, covering regression coefficients and multinomial logit partworths for conjoint analysis. We validate the framework on two datasets spanning different domains, question types, and LLMs, showing that our approach captures 61-79% of the theoretically attainable efficiency gains, achieving 11.4% and 10.5% MSE reductions without requiring any pilot human data for the target survey.","authors":["Zikun Ye","Hema Yoganarasimhan"],"categories":["cs.AI","stat.AP"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-19","first_seen":"2026-04-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.17267","pdf_url":"https://arxiv.org/pdf/2604.17267","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","调查实验","样本分配"],"reason":"用LLM生成调查回答并优化人类样本分配，有真实人类数据对照，涉及调查实验场景，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":65,"question":"在LLM增强的多问题调查中，如何将固定的人类受访者预算最优分配到各问题，以最小化估计量的均方误差？","design":"不是仿真研究。本文提出一个三部分框架：基于预测驱动推断定义问题特定的校正难度，推导出闭式最优分配规则，并利用历史调查数据通过元学习预测新问题的校正难度，从而在无目标调查试点数据的情况下实现零样本分配。","baseline":"使用两个跨领域、问题类型和LLM的真实数据集进行验证，包含真实人类响应作为对照。","findings":"校正难度是决定人类样本分配的正确方差基元，而非原始LLM准确率或人类响应方差；所提元学习方法在无试点数据下捕获了61-79%的理论效率增益，均方误差分别降低11.4%和10.5%。","reliability":"论文指出校正难度依赖于未观测的人类响应，元学习预测可能引入误差；LLM预测在不同问题上的准确性异质性大且难以事前预见，框架有效性受历史数据质量和领域迁移影响。","relevance":"高度相关。该研究将LLM作为人类被试的替代信号，在真实人类数据对照下优化调查设计，涉及经济学实验场景，并明确讨论了仿真可靠性的条件与局限，值得精读。","inspiration":"借鉴该文将LLM预测误差结构化为‘校正难度’并据此优化样本分配的方法，可迁移到经济金融调查中，例如在消费者信心指数或通胀预期调查里，针对不同问题分配差异化的人类受访者预算｜该方法可应用于消费者跨期选择实验，通过LLM预填答案并识别需要更多人类样本校正的高难度问题，以提升估计效率｜研究设计：以LLM作为虚拟被试生成跨期选择偏好，处理为基于元学习预测的校正难度分配人类样本比例，结果变量为时间偏好率的估计均方误差，用真实消费者面板调查数据作为基准对照"}},{"id":"2604.17615","version":1,"title":"WhatIf: Interactive Exploration of LLM-Powered Social Simulations for Policy Reasoning","zh_title":"WhatIf：用于政策推理的LLM驱动社会模拟的交互式探索","abstract":"Policymakers in domains such as emergency management, public health, and urban planning must make decisions under deep uncertainty, where outcomes depend on how large populations interpret information, coordinate, and adopt over time. Existing tools only partially support this process: tabletop exercises enable collaborative discussion but lack dynamic feedback, while computational simulations capture population dynamics but are designed for offline analysis. We present WhatIf, an interactive system that enables policymakers to steer, inspect, and compare LLM-powered social simulations in real time. Informed by a formative study in emergency preparedness planning, we derive four design requirements for interactive policy simulations: fluid steering, real-time scale, collaborative exploration, and multi-level interpretability. We developed WhatIf guided by these requirements and evaluated it with five preparedness professionals across three disaster evacuation scenarios. Our findings show that participants used the system as a space for iterative branching and comparison rather than evaluating fixed plans; reflected on tacit planning assumptions when agent behavior violated expectations; surfaced previously unrecognized planning vulnerabilities; and grounded their reasoning in inspectable agent-level cases rather than aggregate outputs alone. These findings suggest broader design implications for LLM-powered social simulation systems: designing such systems as interactive, shared reasoning environments -- rather than offline predictive tools -- can better support expert decision-making under deep uncertainty.","authors":["Yuxuan Li","Kyzyl Monteiro","Hirokazu Shirado","Sauvik Das"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-04-19","first_seen":"2026-04-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.17615","pdf_url":"https://arxiv.org/pdf/2604.17615","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["LLM社会模拟","政策评估","交互式系统"],"reason":"用LLM agent模拟人群疏散行为，支持政策推理，有批判性反思但缺真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":79,"question":"如何设计交互式系统以支持政策制定者实时操控、检查和比较基于大语言模型的社会仿真，从而辅助深度不确定性下的政策推理？","design":"基于大语言模型构建超过12,000个智能体，模拟灾难疏散场景中人群的信息解读、协调与行为；通过形成性研究提出四项设计需求（流畅操控、实时规模、协作探索、多层次可解释性），并开发交互系统WhatIf，让政策专家实时操控仿真参数、比较不同方案，并检查个体智能体的行为理由。","baseline":"无对照","findings":"参与者将系统用于迭代分支和方案比较，而非评估固定计划；当智能体行为违背预期时，他们反思了自身的隐性假设，发现了先前未识别的规划漏洞，并基于可检查的个体案例进行推理。","reliability":"论文未讨论","relevance":"该研究利用LLM智能体进行社会仿真以支持政策推理，强调交互式探索而非预测，但缺乏真实人类数据对照，适合关注仿真系统设计与人机交互的研究者阅读。","inspiration":"该研究通过交互式系统让政策专家实时操控LLM智能体仿真参数并检查个体行为理由，这种‘可解释性检查’的设计值得借鉴，可用于验证LLM仿真中决策逻辑的合理性。｜可迁移到政策公告的预期形成研究，例如央行沟通如何影响市场主体的通胀预期与投资决策。｜以LLM智能体为被试，处理为不同措辞或透明度的央行声明，结果变量为智能体模拟的投资组合调整，对照真实市场调查数据或实验数据。"}},{"id":"2604.17220","version":2,"title":"Dynamics of Cognitive Heterogeneity: Investigating Behavioral Biases in Multi-Stage Supply Chains with LLM-Based Simulation","zh_title":"认知异质性动力学：基于LLM仿真的多级供应链行为偏差研究","abstract":"Modeling coordination among generative agents in complex multi-round decision-making presents a core challenge for AI and operations management. Although behavioral experiments have revealed cognitive biases behind supply chain inefficiencies, traditional methods face scalability and control limitations. We introduce a scalable experimental paradigm using Large Language Models (LLMs) to simulate multi-stage supply chain dynamics. Grounded in a Hierarchical Reasoning Framework, this study specifically analyzes the impact of cognitive heterogeneity on agent interactions. Unlike prior homogeneous settings, we employ DeepSeek and GPT agents to systematically vary reasoning sophistication across supply chain tiers. Through rigorously replicated and statistically validated simulations, we investigate how this cognitive diversity influences collective outcomes. Results indicate that agents exhibit myopic and self-interested behaviors that exacerbate systemic inefficiencies. However, we demonstrate that information sharing effectively mitigates these adverse effects. Our findings extend traditional behavioral methods and offer new insights into the dynamics of AI-enabled organizations. This work underscores both the potential and limitations of LLM-based agents as proxies for human decision-making in complex operational environments.","authors":["Jiuyun Jiang","Yuecheng Hong","Bo Yang","Jin Yang","Guangxin Jiang","Xiaomeng Guo","Guang Xiao"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-04-19","first_seen":"2026-04-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.17220","pdf_url":"https://arxiv.org/pdf/2604.17220","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","供应链行为","认知异质性"],"reason":"用LLM代理模拟供应链决策，与人类行为对照，涉及运营管理场景，并讨论代理作为人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":86,"question":"在多层次供应链中，认知异质性（不同推理能力的LLM代理）如何影响代理间的互动与系统整体绩效？","design":"使用DeepSeek和GPT系列LLM代理模拟啤酒分销游戏中的多级供应链参与者，通过分层推理框架系统性地改变各层级代理的推理复杂度，测量订单波动（牛鞭效应）、库存水平和系统总成本等结果变量。","baseline":"对照经典啤酒分销游戏的人类实验数据，包括Sterman (1989)等研究中观察到的人类行为模式与牛鞭效应。","findings":"LLM代理表现出短视和自利行为，加剧了牛鞭效应和系统低效；信息共享能有效缓解这些负面影响。不同模型家族表现出行为特征差异，但在最高认知层级上均趋近最优策略。","reliability":"论文承认依赖商业模型导致训练数据不透明，难以完全分离内在推理能力与训练偏差；框架未建模代理对他人有限理性的动态适应能力；实验采用确定性参数，未涵盖现实供应链的随机中断和复杂网络拓扑。","relevance":"该研究直接使用LLM代理复现经典供应链行为实验，并与人类基准对照，探讨代理作为人类决策替代品的潜力与局限，符合研究者对仿真可靠性及失效条件的关注，值得阅读原文以了解具体实验设计与批判性讨论。","inspiration":"该方法通过分层推理框架系统性地操控LLM代理的认知复杂度，并测量牛鞭效应等行为指标，为在受控实验中分离认知能力对决策的影响提供了可借鉴的设计｜可迁移至资产定价实验，研究不同推理能力的交易者如何影响市场波动、泡沫形成与价格发现效率｜可设计一个实验市场，以LLM代理作为交易者，分层设定其推理复杂度，测量价格偏离、交易量与泡沫程度，并与人类实验市场数据（如Smith等人1988年的资产泡沫实验）进行对照"}},{"id":"2605.23920","version":1,"title":"Artificial Effort","zh_title":"人工努力：大语言模型对实验经济学中真实努力任务的影响","abstract":"Real-effort tasks, in which participants perform cognitively costly activities whose outcomes depend on actual performance, are widely used in experimental economics. Their validity, however, rests on the assumption that a human performs them. We study whether this assumption still holds in the era of Artificial Intelligence (AI) and Large Language Models (LLMs). Using 8 canonical real-effort tasks and 23 LLMs from three major providers, we show that most tasks can now be solved accurately and at a negligible cost, while only a few resist automation. Performance improves with each model generation, and midtier models are rapidly closing the gap with frontier ones, broadening the set of widely accessible models that can automate these tasks. Additionally, we show that verbally offering monetary incentives has no effect on LLM performance. Our findings establish a boundary condition for the use of real-effort tasks in unsupervised settings: when participants can cheaply outsource task completion to an LLM, observed performance may no longer reflect genuine human effort.","authors":["Federico Belotti","Stefano Coniglio","Antonio Cosma","Francesco Fallucchi"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-17","first_seen":"2026-04-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.23920","pdf_url":"https://arxiv.org/pdf/2605.23920","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A2","B4"],"tags":["LLM仿真","实验经济学","可靠性评估"],"reason":"评估LLM替代人类完成真实努力任务的可靠性，指出仿真失效条件，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":54,"question":"大语言模型能否准确且低成本地完成实验经济学中常用的真实努力任务，从而破坏其在无监督环境下的效度？","design":"本研究并非仿真人类被试，而是直接测试23个大语言模型（来自OpenAI、Google、Anthropic）在8个经典真实努力任务上的表现，通过API发送任务截图和指令，记录准确率、成本，并检验口头金钱激励和人类行为指令对模型表现的影响。","baseline":"无对照","findings":"大多数任务可被LLM以高准确率和极低成本完成，仅少数任务仍难以自动化；模型性能随代际提升，中端模型正迅速追赶前沿模型。口头提供金钱激励对LLM表现无影响。","reliability":"论文指出，在无监督环境下，若参与者可廉价外包任务给LLM，则观察到的表现可能不再反映真实人类努力，这构成了真实努力任务使用的边界条件。","relevance":"该研究直接评估LLM替代人类完成经济学实验任务的能力，并明确指出仿真失效的边界条件，对关注LLM仿真可靠性及偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"该方法通过直接向LLM发送任务截图和指令来测试其完成真实努力任务的能力，并检验口头激励的影响，为评估LLM在实验任务中的表现提供了可复现的测试框架。｜可迁移到经济金融实验中需要被试付出认知努力的任务，如信息处理、计算或决策任务，以检验LLM是否可替代人类被试。｜以LLM为被试，向其呈现资产定价实验中的信息处理任务（如从财务报表中提取关键指标），处理为有无口头金钱激励，结果变量为任务准确率和反应时间，并与人类被试的真实数据对照。"}},{"id":"2604.14786","version":1,"title":"CogEvolution: A Human-like Generative Educational Agent to Simulate Student's Cognitive Evolution","zh_title":"CogEvolution：模拟学生认知演化的人类化生成式教育智能体","abstract":"Generative Agents, owing to their precise modeling and simulation capabilities of human behavior, have become a pivotal tool in the field of Artificial Intelligence in Education (AIEd) for uncovering complex cognitive processes of learners. However, existing educational agents predominantly rely on static personas to simulate student learning behaviors, neglecting the decisive role of deep cognitive capabilities in learning outcomes during practice interactions. Furthermore, they struggle to characterize the dynamic fluidity of knowledge internalization, transfer, and cognitive state transitions. To overcome this bottleneck, this paper proposes a human-like educational agent capable of simulating student cognitive evolution: CogEvolution. Specifically, we first construct a cognitive depth perceptron based on the Interactive, Constructive, Active, Passive (ICAP) taxonomy from cognitive psychology, achieving precise quantification of learner cognitive engagement. Subsequently, we propose a memory retrieval method based on Item Response Theory (IRT) to simulate the connection and assimilation of new and prior knowledge. Finally, we design a dynamic cognitive update mechanism based on evolutionary algorithms to simulate the real-time integration of student learning behaviors and cognitive evolution processes. Comprehensive evaluations demonstrate that CogEvolution not only significantly outperforms baseline models in behavioral fidelity and learning curve fitting but also uniquely reproduces plausible and robust cognitive evolutionary paths consistent with educational psychology expectations, providing a novel paradigm for constructing highly interpretable educational agents.","authors":["Wei Zhang","Yihang Cheng","Zhirong Ye","Kezhen Huang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-16","first_seen":"2026-04-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.14786","pdf_url":"https://arxiv.org/pdf/2604.14786","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["教育智能体","认知演化","社会模拟"],"reason":"模拟学生认知演化，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":141,"question":"如何构建一个能模拟学生认知动态演化的生成式教育智能体，以克服现有静态角色模型在认知深度和路径演化上的不足？","design":"提出CogEvolution智能体，基于ICAP认知框架构建认知深度感知器，利用IRT驱动记忆检索模拟新旧知识联结，并设计基于进化算法的动态认知更新机制，在CogMath-948数据集上模拟学生学习行为和认知状态演化。","baseline":"无对照","findings":"CogEvolution在行为保真度和学习曲线拟合上显著优于基线模型，并能复现符合教育心理学预期的认知演化路径。","reliability":"论文未讨论","relevance":"该研究利用LLM智能体模拟学生认知演化，属于人类仿真在教育领域的应用，但缺乏真实人类数据对照，且未涉及经济学实验或政策评估，与研究者关注的核心场景（有基准对照的批判性仿真）匹配度有限，可作为社会模拟边界案例参考。","inspiration":"可借鉴其基于认知框架（如ICAP）构建智能体内部状态感知与动态更新机制的方法，用于模拟经济主体学习过程｜可迁移到消费者跨期选择行为研究，模拟个体在金融教育干预下的偏好演化｜设计：以LLM智能体为被试，施加金融知识培训处理，结果变量为时间贴现率变化，对照真实理财教育实验的个体层面面板数据"}},{"id":"2604.14575","version":3,"title":"Generative Augmented Inference of LLM-generated Data for Market Research: Theory and Empirical Evidence","zh_title":"基于LLM生成数据的生成式增强推断用于市场研究：理论与实证","abstract":"Marketing research often relies on parameters estimated from costly human-generated data, such as conjoint survey responses, purchase decisions, and field experiment outcomes. Recent advances in large language models (LLMs) and other AI systems offer inexpensive auxiliary data, but introduce a new challenge: AI outputs are not direct observations of the target outcomes, but could involve high-dimensional representations with complex and unknown relationships to human labels. Conventional methods leverage AI predictions as direct proxies for true labels, which can be inefficient or unreliable when this relationship is weak or misspecified. We propose Generative Augmented Inference (GAI), a general framework that incorporates AI-generated outputs as informative features for estimating models of human-labeled outcomes. GAI uses an orthogonal moment construction that enables consistent estimation and valid inference with a flexible, nonparametric relationship between LLM-generated outputs and human labels. We establish asymptotic normality and a key dominance result: under random labeling, GAI is optimal within a unified class of debiased estimators-including human-data-only estimators and state-of-the-art debiasing methods-and delivers strict improvements under a mild informativeness condition. Even when the labeled sample is not representative of the target population, an extended variant of GAI still dominates the weighted human-data-only estimator. Empirically, GAI outperforms benchmarks across diverse marketing research settings. In a conjoint analysis, it halves estimation error and reduces human labeling requirements by over 75%. In a pricing study, it consistently outperforms alternative estimators when all methods receive identical auxiliary inputs. In a health insurance study, it saves over 90% of labels while preserving decision accuracy.","authors":["Cheng Lu","Mengxin Wang","Dennis J. Zhang","Heng Zhang"],"categories":["cs.LG","cs.AI","stat.ME","stat.ML"],"primary_category":"cs.LG","announce_type":"new","date":"2026-04-16","first_seen":"2026-04-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.14575","pdf_url":"https://arxiv.org/pdf/2604.14575","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助推断","市场研究","合成数据"],"reason":"用LLM生成辅助特征估计人类标签模型，属替代标注员而非仿真被试，边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":46,"question":"如何将LLM生成的高维表征作为辅助特征，而非替代标签，来提升基于人类标注数据的参数估计效率与推断有效性？","design":"本文并非用LLM仿真人类被试，而是提出生成式增强推断（GAI）框架，将LLM输出（如预测、嵌入、推理痕迹等）作为辅助特征，通过正交矩构建，在估计人类标签模型时实现一致估计和有效推断。","baseline":"联合分析、定价研究、健康保险研究中的真实人类调查回答、购买决策和现场实验数据。","findings":"GAI在联合分析中将估计误差减半，人类标注需求降低75%以上；在健康保险研究中节省超90%标签且保持决策准确度。理论上，在随机标注下，GAI在一类去偏估计量中是最优的，并在弱信息条件下仍能严格改进估计。","reliability":"论文未讨论","relevance":"本文属于边界情形，核心是用LLM生成特征辅助估计人类标签模型，而非直接仿真人类被试，但方法论对利用LLM进行高效推断有借鉴意义，值得阅读以评估其在仿真实验中的潜在应用。","inspiration":"借鉴GAI将LLM输出作为辅助特征而非代理标签的正交矩估计方法，可提升小样本人类实验的估计效率。｜可迁移到消费者需求估计或政策评估场景，如利用LLM生成的产品描述嵌入或模拟选择理由，辅助估计价格弹性或处理效应。｜以LLM生成的产品特征嵌入作为辅助特征，招募真实消费者进行离散选择实验，结果变量为购买选择，以真实购买记录或大规模调查数据作为对照基准。"}},{"id":"2604.14467","version":1,"title":"Who Saw It Coming? Historical Experience and the 2021 Inflation Forecast Failure","zh_title":"谁预见到了？历史经验与2021年通胀预测失败","abstract":"This paper studies the 2021 U.S. inflation forecasting failure. I show that the failure was primarily driven by sample composition rather than functional-form misspecification: estimation samples dominated by the Great Moderation underweight supply-shock regimes, and expectations anchored to that regime were slow to recognize the shift. Three historically informed adjustments, an intercept correction, a similarity re-estimation on 1970s data, and a kernel-weighted estimator, substantially close the forecast gap, and the gains extend to eight additional U.S. price indices. Household survey respondents over 60, whose lifetime includes the 1970s, reported higher inflation expectations from early 2021, consistent with experience-based learning; younger cohorts remained anchored to the prevailing regime. A controlled experiment with large language models conditioned on ``experienced'' and ``young'' professional personas confirms that experiential priors generate significant forecast differences under a common training leakage assumption. Across all three exercises, the source of the prior mattered more than the sophistication of the model.","authors":["Dalibor Stevanovic"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-04-15","first_seen":"2026-04-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.14467","pdf_url":"https://arxiv.org/pdf/2604.14467","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","通胀预测","经验学习"],"reason":"用LLM模拟不同经验背景的专业人士预测通胀，并与真实调查数据对照，属于经济学场…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":72,"question":"2021年美国通胀预测失败是否源于预测模型所依赖的历史经验样本构成，而非模型函数形式误设？","design":"使用Claude和ChatGPT两种大语言模型，分别赋予“有经验”（经历过1970年代）和“年轻”（仅经历大缓和时期）的专业经济学家人设，在四个数据时点上生成通胀预测，测量两种人设之间的预测差异（persona gap）。","baseline":"纽约联储消费者预期调查（SCE）中60岁以上（经历过1970年代）与40岁以下（未经历）受访者的通胀预期差异。","findings":"统计模型与LLM实验均表明，历史经验先验是预测差异的主要来源，其重要性超过模型复杂度；基于1970年代数据的简单修正能大幅缩小预测误差。","reliability":"LLM实验的有效性依赖于“共同训练泄露假设”（common training leakage assumption），即模型训练数据中可能已包含未来信息，但通过比较不同人设的预测差异可识别经验效应。","relevance":"该研究用LLM模拟不同经验背景的经济学家进行通胀预测，并与真实调查数据对照，直接涉及经济学场景下LLM仿真人类行为的可靠性与偏差，值得精读。","inspiration":"该方法通过赋予LLM不同历史经验人设并测量预测差异（persona gap）来分离经验效应，值得借鉴｜可迁移到资产定价实验中，研究不同市场经历（如经历过2008年金融危机 vs. 仅经历长期牛市）如何影响投资者的风险溢价预期｜用LLM模拟两类投资者人设，处理为赋予不同历史市场数据经历，结果变量为对股票风险溢价的预测值，对照真实数据可选用美联储消费者金融调查（SCF）中不同年龄段投资者的风险资产配置比例差异"}},{"id":"2604.11312","version":2,"title":"Network Effects and Agreement Drift in LLM Debates","zh_title":"LLM辩论中的网络效应与意见漂移","abstract":"Large Language Models (LLMs) have demonstrated an unprecedented ability to simulate human-like social behaviors, making them useful tools for simulating complex social systems. However, it remains unclear to what extent these simulations can be trusted to accurately capture key social mechanisms, particularly in highly unbalanced contexts involving minority groups. This paper uses a network generation model with controlled homophily and class sizes to examine how LLM agents behave collectively in multi-round debates. Moreover, our findings highlight a particular directional susceptibility that we term \\textit{agreement drift}, in which agents are more likely to shift toward specific positions on the opinion scale. Overall, our findings highlight the need to disentangle structural effects from model biases before treating LLM populations as behavioral proxies for human groups.","authors":["Erica Cau","Andrea Failla","Giulio Rossetti"],"categories":["cs.SI","cs.AI","cs.CY","cs.MA","physics.soc-ph"],"primary_category":"cs.SI","announce_type":"new","date":"2026-04-13","first_seen":"2026-04-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.11312","pdf_url":"https://arxiv.org/pdf/2604.11312","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","意见动态","网络效应"],"reason":"用LLM群体模拟舆论辩论，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":137,"question":"在同质性和群体规模不平衡的条件下，LLM智能体在多轮辩论中的集体意见动态如何变化？","design":"使用带有同质性控制和类别大小调节的网络生成模型，构建LLM智能体群体，模拟多轮辩论，测量意见收敛、极化和“同意漂移”等结果变量。","baseline":"无对照","findings":"LLM智能体的意见收敛和极化模式对网络结构和相对群体规模高度敏感；智能体表现出“同意漂移”倾向，即更容易向特定立场偏移。","reliability":"论文指出，在将LLM群体视为人类行为代理之前，需要区分结构性效应和模型偏差，但未具体讨论失效条件。","relevance":"该研究批判性地探讨了LLM仿真在意见动态中的偏差，虽无人类基准，但揭示了结构性因素与模型偏差的混淆，对评估LLM作为人类被试替代品的可靠性有参考价值，值得阅读原文。","inspiration":"该研究通过调节网络同质性和群体规模来构建多轮辩论仿真，系统测量意见收敛、极化和“同意漂移”，揭示了结构性因素与模型偏差的混淆，这种对仿真内部机制的解构值得借鉴｜可迁移到金融市场中的信息扩散与共识形成问题，例如分析师一致预期形成或投资者情绪传染｜设计一个LLM智能体模拟的资产定价实验，处理变量为网络结构（同质性高低）和群体规模比例，结果变量为价格预测的收敛速度和偏差方向，对照真实分析师预测数据或实验市场数据"}},{"id":"2604.10834","version":1,"title":"LLMs for Qualitative Data Analysis Fail on Security-specificComments in Human Experiments","zh_title":"大语言模型在人类实验安全评论的定性数据分析中失效","abstract":"[Background:] Thematic analysis of free-text justifications in human experiments provides significant qualitative insights. Yet, it is costly because reliable annotations require multiple domain experts. Large language models (LLMs) seem ideal candidates to replace human annotators. [Problem:] Coding security-specific aspects (code identifiers mentioned, lines-of-code mentioned, security keywords mentioned) may require deeper contextual understanding than sentiment classification. [Objective:] Explore whether LLMs can act as automated annotators for technical security comments by human subjects. [Method:] We prompt four top-performing LLMs on LiveBench to detect nine security-relevant codes in free-text comments by human subjects analyzing vulnerable code snippets. Outputs are compared to human annotators using Cohen's Kappa (chance-corrected accuracy). We test different prompts mimicking annotation best practices, including emerging codes, detailed codebooks with examples, and conflicting examples. [Negative Results:] We observed marked improvements only when using detailed code descriptions; however, these improvements are not uniform across codes and are insufficient to reliably replace a human annotator. [Limitations:] Additional studies with more LLMs and annotation tasks are needed.","authors":["Maria Camporese","Fabio Massacci","Yuanjun Gong"],"categories":["cs.SE","cs.AI"],"primary_category":"cs.SE","announce_type":"new","date":"2026-04-12","first_seen":"2026-04-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.10834","pdf_url":"https://arxiv.org/pdf/2604.10834","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","定性数据分析","安全评论"],"reason":"用LLM替代人工标注定性数据，属于标注员替代而非仿真被试，但涉及人类实验对照，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":145,"question":"LLM能否替代人类标注者，对安全技术评论进行主题编码？","design":"本研究并非用LLM仿真人类被试，而是用LLM替代人类标注者。将四款LLM作为自动标注器，输入人类被试分析漏洞代码片段后产生的自由文本评论，要求其检测九种安全相关编码，并与人类标注结果比较。","baseline":"对照的真实人类数据是四位人类标注者对同一批评论的编码结果，以及审核者校正后的编码作为最终基准。","findings":"仅在使用详细代码描述时，LLM性能有明显提升，但提升在不同编码上并不均匀，且不足以可靠替代人类标注者。","reliability":"论文承认需要更多LLM和标注任务的研究来验证结论，且当前结果仅基于特定安全编码任务，泛化性有限。","relevance":"本文属于LLM替代人类标注者的研究，而非用LLM仿真人类被试，但涉及真实人类实验对照和可靠性评估，对关注LLM在人类研究中替代角色的研究者有参考价值，尤其在批判性评估LLM失效条件方面。","inspiration":"本研究采用LLM替代人类标注者的对照设计，将LLM输出与多位人类标注者及审核校正后的基准进行系统比较，并考察了输入信息详细程度对性能的影响，这种多层级对照和条件敏感性分析值得借鉴。｜该方法可迁移到经济金融文本分析场景，例如利用LLM对央行政策声明或分析师报告进行主题编码或情感分类，以评估LLM能否替代人类研究助理。｜可设计实验：以LLM为被试，输入央行货币政策声明，要求其识别声明中的前瞻指引、风险评估等主题，以人类专家标注的编码作为基准，比较不同提示信息（如仅声明文本 vs. 附加经济背景）下LLM的编码准确率。"}},{"id":"2604.09502","version":2,"title":"Strategic Algorithmic Monoculture: Experimental Evidence from Coordination Games","zh_title":"策略性算法单一文化：来自协调博弈的实验证据","abstract":"AI agents increasingly operate in multi-agent environments where outcomes depend on coordination. We distinguish primary algorithmic monoculture -- baseline action similarity -- from strategic algorithmic monoculture, whereby agents adjust similarity in response to incentives. We implement a simple experimental design that cleanly separates these forces, and deploy it on human and large language model (LLM) subjects. LLMs exhibit high levels of baseline similarity (primary monoculture) and, like humans, they regulate it in response to coordination incentives (strategic monoculture). While LLMs coordinate extremely well on similar actions, they lag behind humans in sustaining heterogeneity when divergence is rewarded.","authors":["Gonzalo Ballestero","Hadi Hosseini","Samarth Khanna","Ran I. Shorrer"],"categories":["cs.AI","cs.GT","cs.MA","econ.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-10","first_seen":"2026-04-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.09502","pdf_url":"https://arxiv.org/pdf/2604.09502","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","人类行为对照","协调博弈"],"reason":"用LLM和人类被试进行协调博弈实验，直接对比行为，属于经济学实验场景的人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":17,"question":"LLM智能体在协调博弈中如何协调行动，与人类相比有何差异？","design":"使用多个LLM（16个模型）和人类被试，在开放式问题（如说出一个字母、一个城市）上设置三种处理：picking（仅要求有效答案）、coordination（激励与同类型另一智能体答案相同）、divergence（激励答案不同），测量独立智能体间答案一致率。","baseline":"人类被试在相同实验任务中的行为数据。","findings":"LLM在无激励时答案一致性远高于人类，表现出高基础单一性；在协调激励下LLM能像人类一样调节一致性，但在需要差异化的任务中维持异质性的能力弱于人类。","reliability":"论文未讨论","relevance":"该研究直接对比LLM与人类在协调博弈中的行为，属于经济学实验场景的人类仿真，且揭示了LLM在差异化任务中的失效，高度契合研究者对仿真可靠性与偏差的关注。","inspiration":"借鉴其通过picking、coordination、divergence三种处理分离基础偏好与策略调整的实验设计，可清晰测量LLM的固有行为倾向与激励响应。｜可迁移至资产定价实验，研究LLM交易员在信息协调与差异化策略下的市场表现。｜以LLM为被试，设置picking（自由选股）、coordination（激励与另一LLM选相同股票）、divergence（激励选不同股票）三种处理，结果变量为投资组合相似度与市场效率指标，对照真实人类交易员在相同实验中的行为数据。"}},{"id":"2604.16472","version":1,"title":"Training Language Models for Bilateral Trade with Private Information","zh_title":"训练语言模型进行具有私人信息的双边贸易","abstract":"Bilateral bargaining under incomplete information provides a controlled testbed for evaluating large language model (LLM) agent capabilities. Bilateral trade demands individual rationality, strategic surplus maximization, and cooperation to realize gains from trade. We develop a structured bargaining environment where LLMs negotiate via tool calls within an event-driven simulator, separating binding offers from natural-language messages to enable automated evaluation. The environment serves two purposes: as a benchmark for frontier models and as a training environment for open-weight models via reinforcement learning. In benchmark experiments, a round-robin tournament among five frontier models (15,000 negotiations) reveals that effective strategies implement price discrimination through sequential offers. Aggressive anchoring, calibrated concession, and temporal patience correlate with the highest surplus share and deal rate. Accommodating strategies that concede quickly disable price discrimination in the buyer role, yielding the lowest surplus capture and deal completion. Stronger models scale their behavior proportionally to item value, maintaining performance across price tiers; weaker models perform well only when wide zones of possible agreement offset suboptimal strategies. In training experiments, we fine-tune Qwen3 (8B, 14B) via supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) against a fixed frontier opponent. These stages optimize competing objectives: SFT approximately doubles surplus share but reduces deal rates, while RL recovers deal rates but erodes surplus gains, reflecting the reward structure. SFT also compresses surplus variation across price tiers, which generalizes to unseen opponents, suggesting that behavioral cloning instills proportional strategies rather than memorized price points.","authors":["Dirk Bergemann","Soheil Ghili","Xinyang Hu","Chuanhao Li","Zhuoran Yang"],"categories":["cs.GT","cs.AI","cs.MA","econ.GN","econ.TH"],"primary_category":"cs.GT","announce_type":"new","date":"2026-04-10","first_seen":"2026-04-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.16472","pdf_url":"https://arxiv.org/pdf/2604.16472","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM Agent","双边贸易","社会模拟"],"reason":"用LLM agent模拟双边贸易谈判，但无真实人类数据对照，属于社会模拟的纯理…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":177,"question":"如何评估和训练大语言模型在私有信息双边贸易中的战略谈判能力？","design":"构建结构化讨价还价环境，让LLM通过工具调用进行谈判，分离约束性报价与自然语言消息；对五个前沿模型进行循环赛基准测试（15000次谈判），并对Qwen3（8B、14B）进行监督微调加GRPO强化学习训练，测量个体理性、剩余份额和成交率。","baseline":"无对照","findings":"前沿模型中，激进锚定加校准让步策略同时获得最高剩余份额和成交率；SFT提高剩余份额但降低成交率，RL恢复成交率却侵蚀剩余收益，两者目标冲突。","reliability":"论文未讨论","relevance":"该研究用LLM模拟双边谈判，但缺乏真实人类数据对照，属于纯仿真实验，未直接评估仿真可靠性或偏差，与研究者关注的人类基准对照和批判性评估方向部分相关，但非核心匹配。","inspiration":"该研究通过工具调用分离约束性报价与自然语言消息，并采用循环赛基准测试和强化学习训练，为多智能体策略评估提供了结构化方法｜可迁移至资产定价实验中的做市商谈判或信贷审批中的利率协商场景｜以LLM作为交易员被试，处理为不同训练方式（SFT vs RL），结果变量为成交价与成交率，对照真实市场微观结构数据中的买卖价差与成交概率"}},{"id":"2604.06663","version":1,"title":"Restoring Heterogeneity in LLM-based Social Simulation: An Audience Segmentation Approach","zh_title":"在基于大语言模型的社会模拟中恢复异质性：一种受众细分方法","abstract":"Large Language Models (LLMs) are increasingly used to simulate social attitudes and behaviors, offering scalable \"silicon samples\" that can approximate human data. However, current simulation practice often collapses diversity into an \"average persona,\" masking subgroup variation that is central to social reality. This study introduces audience segmentation as a systematic approach for restoring heterogeneity in LLM-based social simulation. Using U.S. climate-opinion survey data, we compare six segmentation configurations across two open-weight LLMs (Llama 3.1-70B and Mixtral 8x22B), varying segmentation identifier granularity, parsimony, and selection logic (theory-driven, data-driven, and instrument-based). We evaluate simulation performance with a three-dimensional evaluation framework covering distributional, structural, and predictive fidelity. Results show that increasing identifier granularity does not produce consistent improvement: moderate enrichment can improve performance, but further expansion does not reliably help and can worsen structural and predictive fidelity. Across parsimony comparisons, compact configurations often match or outperform more comprehensive alternatives, especially in structural and predictive fidelity, while distributional fidelity remains metric dependent. Identifier selection logic determines which fidelity dimension benefits most: instrument-based selection best preserves distributional shape, whereas data-driven selection best recovers between-group structure and identifier-outcome associations. Overall, no single configuration dominates all dimensions, and performance gains in one dimension can coincide with losses in another. These findings position audience segmentation as a core methodological approach for valid LLM-based social simulation and highlight the need for heterogeneity-aware evaluation and variance-preserving modeling strategies.","authors":["Xiaoyou Qin","Zhihong Li","Xiaoxiao Cheng"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-08","first_seen":"2026-04-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.06663","pdf_url":"https://arxiv.org/pdf/2604.06663","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM人类仿真","受众细分","社会模拟保真度"],"reason":"用LLM仿真人类气候态度，有真实调查数据对照，评估异质性恢复与保真度，并指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-08","rank":5,"question":"如何通过受众分割策略恢复大语言模型社会仿真中的异质性？","design":"使用Llama 3.1-70B和Mixtral 8x22B模型，基于美国气候态度调查数据，通过六种分割配置（变化标识符粒度、简约性和选择逻辑）生成合成样本，评估分布保真度、结构保真度和预测保真度。","baseline":"2025年10月通过Prolific收集的594份人类样本，按人口普查配额分层抽样，并使用Six Americas超短问卷分配受众细分标签。","findings":"增加标识符粒度并不一致提升性能；中等丰富度可改善，但过度扩展会损害结构和预测保真度。紧凑配置在结构和预测保真度上常优于或等同更全面的配置，而标识符选择逻辑决定哪个保真度维度受益最大。","reliability":"论文指出无单一配置在所有维度占优，一维度的提升可能伴随另一维度的损失；未明确讨论其他失效条件。","relevance":"直接命中研究者关注的LLM仿真人类实验、有真实人类对照、评估保真度并指出失效条件，强烈建议阅读原文。","inspiration":"该研究通过系统变换受众分割的标识符粒度、简约性和选择逻辑来评估仿真保真度，这种多维度配置对比的设计值得借鉴｜可迁移到消费者金融决策仿真，如退休储蓄选择或保险购买行为中的异质性偏好研究｜以LLM作为被试，施加不同信息框架（如损失vs收益表述）作为处理，测量储蓄率或保险购买意愿，并以美国消费者金融调查（SCF）或健康与退休研究（HRS）的真实个体数据作为对照基准"}},{"id":"2604.05939","version":1,"title":"Context-Value-Action Architecture for Value-Driven Large Language Model Agents","zh_title":"面向价值驱动大语言模型智能体的情境-价值-行动架构","abstract":"Large Language Models (LLMs) have shown promise in simulating human behavior, yet existing agents often exhibit behavioral rigidity, a flaw frequently masked by the self-referential bias of current \"LLM-as-a-judge\" evaluations. By evaluating against empirical ground truth, we reveal a counter-intuitive phenomenon: increasing the intensity of prompt-driven reasoning does not enhance fidelity but rather exacerbates value polarization, collapsing population diversity. To address this, we propose the Context-Value-Action (CVA) architecture, grounded in the Stimulus-Organism-Response (S-O-R) model and Schwartz's Theory of Basic Human Values. Unlike methods relying on self-verification, CVA decouples action generation from cognitive reasoning via a novel Value Verifier trained on authentic human data to explicitly model dynamic value activation. Experiments on CVABench, which comprises over 1.1 million real-world interaction traces, demonstrate that CVA significantly outperforms baselines. Our approach effectively mitigates polarization while offering superior behavioral fidelity and interpretability.","authors":["TianZe Zhang","Sirui Sun","Yuhang Xie","Xin Zhang","Zhiqiang Wu","Guojie Song"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-07","first_seen":"2026-04-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.05939","pdf_url":"https://arxiv.org/pdf/2604.05939","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["人类行为仿真","价值驱动智能体","算法保真度"],"reason":"用LLM仿真人类行为，有真实人类数据对照，解决行为僵化和价值极化问题，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":18,"question":"如何设计LLM智能体架构，以缓解提示驱动推理导致的行为僵化和价值极化，从而更真实地模拟人类行为？","design":"提出Context-Value-Action（CVA）架构，基于S-O-R模型和Schwartz基本人类价值观理论，将动作生成与认知推理解耦，引入在真实人类数据上训练的Value Verifier显式建模动态价值激活，并通过SFT和DPO对齐生成过程。","baseline":"CVABench包含超过110万条真实世界交互轨迹，来自15000多名人类参与者，用于评估行为逼真度和价值极化。","findings":"增加提示驱动的推理强度不仅未提高仿真逼真度，反而加剧价值极化并降低群体多样性；CVA架构有效缓解极化，在行为逼真度和可解释性上显著优于基线方法。","reliability":"论文未讨论","relevance":"高度相关：研究用LLM仿真人类行为，有大规模真实人类数据作为基准，揭示提示驱动方法的失效模式并提出改进架构，直接回应研究者对仿真可靠性、偏差和经济学/政策评估场景的关切。","inspiration":"借鉴CVA架构将行为生成与价值推理解耦，并引入在真实人类数据上训练的Value Verifier来动态建模价值激活，以此作为缓解行为僵化和价值极化的处理机制。｜可迁移到政策公告的预期形成实验，研究不同信息框架如何通过个体价值观影响通胀预期或就业预期。｜以LLM智能体为被试，处理组采用CVA架构注入特定价值观（如安全或自主），对照组使用标准提示驱动推理，结果变量为预期偏差和群体多样性，用真实调查数据（如密歇根大学消费者调查）作为人类基准对照。"}},{"id":"2604.05516","version":2,"title":"Coupling Macro Dynamics and Micro States for Long-Horizon Social Simulation","zh_title":"耦合宏观动态与微观状态的长周期社会模拟","abstract":"Social network simulation aims to model collective opinion dynamics in large populations, but existing LLM-based simulators mainly focus on aggregate dynamics while largely ignoring individual internal states. This limits their ability to capture opinion reversals driven by gradual individual shifts and makes them unreliable in long-horizon simulations. We propose MF-MDP, a social simulation framework that tightly couples macro-level collective dynamics with micro-level individual states. MF-MDP explicitly models per-agent latent opinion states with a state transition mechanism, combining individual Markov Decision Processes at the micro level with a mean-field collective framework at the macro level. This allows individual behaviors to change internal states gradually rather than trigger instant reactions, enabling the simulator to distinguish agents that are close to switching from those that are far from switching, capture opinion reversals, and maintain accuracy over long horizons. Across real-world events, MF-MDP supports stable simulation of long-horizon social processes with up to 40,000 interactions, compared with about 300 in the baseline MF-LLM, while reducing long-horizon KL divergence by 75.3% (1.2490 to 0.3089) and reversal KL by 66.9% (1.6425 to 0.5434), significantly mitigating the drift observed in MF-LLM. Code is available at github.com/AI4SS/MF-MDP.","authors":["Yunyao Zhang","Yihao Ai","Zuocheng Ying","Qirui Mi","Junqing Yu","Wei Yang","Zikai Song"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-04-07","first_seen":"2026-04-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.05516","pdf_url":"https://arxiv.org/pdf/2604.05516","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","舆论动态","LLM agent"],"reason":"用LLM agent模拟社会舆论动态，但摘要未明确提及真实人类数据对照，属于社…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":158,"question":"如何在长时间跨度社会仿真中，通过耦合宏观集体动态与微观个体状态，捕捉观点反转并保持长期准确性？","design":"提出MF-MDP框架，将社会仿真建模为平均场马尔可夫决策过程：宏观层用时序Transformer学习宏观状态分布演化，微观层每个智能体基于自身潜在状态和宏观信号进行多步前瞻动作重选，模拟个体观点的渐进转变。","baseline":"无对照","findings":"MF-MDP在真实社会事件上支持长达40,000次交互的稳定仿真，而基线MF-LLM仅约300次；长期KL散度降低75.3%，反转KL散度降低66.9%，显著缓解了漂移问题。","reliability":"论文未讨论","relevance":"该研究利用LLM智能体进行社会舆论仿真，但未提供真实人类数据对照，不属于严格的人类仿真验证研究；若关注仿真方法改进与长期动态稳定性，可读原文了解其微观-宏观耦合机制。","inspiration":"MF-MDP框架通过宏观时序建模与微观多步前瞻动作重选耦合，可借鉴其分层仿真思路来设计经济实验中的动态处理与长期效应测量｜该框架可迁移至政策公告的预期形成与市场反应研究，例如模拟央行沟通对通胀预期和资产价格的长期影响｜以LLM智能体为被试，处理为不同频率/透明度的政策信号，结果变量为智能体通胀预期与模拟资产价格，对照真实央行公告前后的调查预期与市场数据"}},{"id":"2604.03920","version":2,"title":"From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities","zh_title":"从合理到因果：模拟在线社区中政策评估的反事实语义","abstract":"LLM-based social simulations can generate believable community interactions, enabling ``policy wind tunnels'' where governance interventions are tested before deployment. But believability is not causality. Claims like ``intervention $A$ reduces escalation'' require causal semantics that current simulation work typically does not specify. We propose adopting the causal counterfactual framework, distinguishing \\textit{necessary causation} (would the outcome have occurred without the intervention?) from \\textit{sufficient causation} (does the intervention reliably produce the outcome?). This distinction maps onto different stakeholder needs: moderators diagnosing incidents require evidence about necessity, while platform designers choosing policies require evidence about sufficiency. We formalize this mapping, show how simulation design can support estimation under explicit assumptions, and argue that the resulting quantities should be interpreted as simulator-conditional causal estimates whose policy relevance depends on simulator fidelity. Establishing this framework now is essential: it helps define what adequate fidelity means and moves the field from simulations that look realistic toward simulations that can support policy changes.","authors":["Agam Goyal","Yian Wang","Eshwar Chandrasekharan","Hari Sundaram"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-04-05","first_seen":"2026-04-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.03920","pdf_url":"https://arxiv.org/pdf/2604.03920","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["LLM社会模拟","政策评估","因果推断"],"reason":"用LLM模拟在线社区进行政策评估，提出因果反事实框架，批判性指出仿真需因果保真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":80,"question":"如何为基于LLM的在线社区仿真中的政策评估引入因果反事实语义，以区分必要因果与充分因果？","design":"本文为理论框架论文，未进行具体仿真实验；它提出在LLM驱动的社区仿真中，通过配对运行（有干预与无干预）来估计必要因果概率（PN）和充分因果概率（PS），并讨论了外生性和单调性假设下的简化估计方法。","baseline":"无对照","findings":"提出将因果反事实框架应用于仿真政策评估，区分必要因果（诊断特定事件）和充分因果（评估干预效果），并指出仿真可借助随机化满足外生性，但单调性可能因干预异质性而违反，此时PN/PS仅为下界。","reliability":"论文指出仿真因果估计是“仿真器条件因果估计”，其政策相关性取决于仿真保真度；单调性假设在社交系统中可能不成立，此时估计仅为下界，且违反单调性本身可能揭示干预异质性。","relevance":"该文直接回应了LLM仿真中“可信不等于因果”的批判，为政策评估提供了因果推断框架，并讨论了仿真保真度与因果估计可靠性的关系，与研究者关注的仿真可靠性及批判性研究高度相关，值得精读原文。","inspiration":"借鉴其配对运行（有干预与无干预）估计必要因果和充分因果概率的框架，可对LLM仿真中的政策干预进行因果分解与稳健性检验｜可迁移到政策公告的预期形成研究，例如评估央行沟通对市场预期的因果效应｜以LLM模拟投资者群体，处理为是否发布前瞻指引，结果变量为预期通胀率，对照真实调查数据（如密歇根消费者预期调查）"}},{"id":"2605.00841","version":1,"title":"AI Agents for Sustainable SMEs: A Green ESG Assessment Framework","zh_title":"面向可持续中小企业的AI代理：绿色ESG评估框架","abstract":"This study presents a novel, AI-driven framework for assessing Environmental, Social, and Governance (ESG) performance in European small and medium-sized enterprises (SMEs). An initial phase established expert-validated ESG baseline scores from a subset of the Flash Eurobarometer FL549 survey data. In the second phase, a scalable AI agent system, built on the n8n automation platform, applied these baselines to perform automated ESG classification and generate contextual recommendations using large language models (LLMs). The results demonstrate the AI system's high consistency with human-derived outputs, thereby supporting more effective monitoring and intervention strategies aligned with the European Green Deal.","authors":["Viet Trinh","Tan Nguyen","Minh-Huyen Phan","Quan Luu"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-05","first_seen":"2026-04-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.00841","pdf_url":"https://arxiv.org/pdf/2605.00841","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI代理","ESG评估","标注替代"],"reason":"用LLM替代人工进行ESG评估，属于标注员替代而非仿真人类被试，但有人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":178,"question":"如何利用AI代理和LLM对欧洲中小企业的ESG表现进行自动化评估与分类？","design":"本研究并非人类仿真实验，而是提出一个混合人机编排框架：先由人工处理约40%的Flash Eurobarometer FL549调查数据建立专家验证的ESG基线分数，再通过n8n自动化平台构建AI代理系统，利用领域启发式规则和LLM对剩余国家数据进行ESG分类并生成情境化建议。","baseline":"人类基准为专家验证的ESG基线分数，源自Flash Eurobarometer FL549调查数据的人工处理部分。","findings":"AI系统输出的ESG分类与人类推导结果高度一致，表明该框架可有效支持与欧洲绿色协议一致的监测和干预策略。","reliability":"论文未讨论","relevance":"本研究用LLM替代人工进行ESG评估，属于标注员替代而非仿真人类被试，但提供了真实人类数据作为对照基准，与研究者关注的LLM替代人类判断的可靠性问题部分相关，值得快速浏览以了解其一致性验证方法。","inspiration":"该方法通过人工处理部分数据建立专家验证基线，再用LLM自动化处理剩余数据并对比一致性，提供了人机混合验证的思路｜可迁移到信贷审批歧视研究，用LLM替代人工审核员评估贷款申请的公平性｜招募银行信贷员人工审核40%的贷款申请作为基线，用LLM处理剩余申请并输出批准/拒绝决策，以人工审核结果作为真实对照，比较两组在性别、种族等维度上的歧视差异"}},{"id":"2605.20191","version":1,"title":"Shiny Stories, Hidden Struggles: Investigating the Representation of Disability Through the Lens of LLMs","zh_title":"光鲜故事，隐藏挣扎：通过LLM视角考察残障表征","abstract":"Modern Large Language Models (LLMs) have recently attracted much attention for their ability to simulate human behavior and generate text that reflects personas and demographic groups. While these capabilities can open up a multitude of diverse applications across fields, it is crucial to examine how such models represent various target groups since LLMs can perpetuate and amplify biases or discrimination against historically marginalized communities or, alternatively, as a result of debiasing efforts, overcorrect by portraying overly positive stereotypes. This overcompensation can idealize these groups, erasing the complexities and challenges they face in favor of unrealistic depictions. In this paper, we investigate how LLMs represent disability by simulating the perspectives of individuals with disabilities in generating social media posts. These posts are then compared with those written by real people with disabilities, focusing on emotional tone, sentiment, and representative words and themes. Our analysis reveals two key findings: (1) LLMs often idealize the experiences of people with disabilities, producing overly positive stereotypes that, despite appearing uplifting, fail to authentically capture their lived realities; and (2) a comparative analysis of posts simulating individuals with and without disabilities highlights a negative bias, where certain topics, such as career and entertainment, are disproportionately associated with nondisabled individuals. This reinforces exclusionary narratives and over-idealized portrayals of disability, misrepresenting the actual challenges faced by this community. These findings align with broader concerns and ongoing research showing that LLMs struggle to reflect the diverse realities of society, particularly the nuanced experiences of marginalized groups, and underscore the need for critical scrutiny of their representations.","authors":["Marco Bombieri","Simone Paolo Ponzetto","Marco Rospocher"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-04-02","first_seen":"2026-04-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.20191","pdf_url":"https://arxiv.org/pdf/2605.20191","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","残障表征","偏差评估"],"reason":"用LLM模拟残障人士发帖，并与真实人类数据对照，评估仿真偏差与失效条件，直接命…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":19,"question":"LLM生成的残障人士社交媒体帖子与真实残障人士的自我描述在情感、主题和语言上有何差异？","design":"使用多种LLM，通过提示词模拟残障人士和普通人在社交媒体上发帖，生成文本后自动标注情感、情绪和抑郁迹象，并与真实Reddit帖子进行对比分析。","baseline":"来自Reddit的真实残障人士自我介绍的帖子数据集。","findings":"LLM倾向于理想化残障人士的经历，产生过度积极的刻板印象，未能真实反映其生活挑战；同时，LLM在模拟残障与非残障个体时存在负面偏见，将职业、娱乐等主题不成比例地与非残障人士关联。","reliability":"论文指出LLM的正面理想化同样会造成伤害，且模型可能因去偏努力而过度补偿，但未详细讨论仿真失效的具体条件。","relevance":"该研究直接以真实人类数据为基准，评估LLM模拟残障人群的偏差与失效模式，属于批判性仿真研究，值得精读以了解LLM在边缘群体仿真中的局限。","inspiration":"该方法通过提示词让LLM模拟特定社会群体生成文本，并与真实社交媒体数据对比，可借鉴用于构建经济实验中的处理组与对照组仿真。｜它能迁移到信贷审批中的群体歧视研究，例如评估LLM在模拟不同种族或性别申请人时的语言偏差。｜可设计让LLM扮演贷款申请人撰写申请陈述，以真实银行信贷文本为基准，比较不同群体提示下的情感、主题和职业关联差异，检验仿真是否复现人类数据中的歧视模式。"}},{"id":"2604.01520","version":1,"title":"LLM Agents as Social Scientists: A Human-AI Collaborative Platform for Social Science Automation","zh_title":"作为社会科学家的LLM代理：一个面向社会科学自动化的人机协作平台","abstract":"Traditional social science research often requires designing complex experiments across vast methodological spaces and depends on real human participants, making it labor-intensive, costly, and difficult to scale. Here we present S-Researcher, an LLM-agent-based platform that assists researchers in conducting social science research more efficiently and at greater scale by \"siliconizing\" both the research process and the participant pool. To build S-Researcher, we first develop YuLan-OneSim, a large-scale social simulation system designed around three core requirements: generality via auto-programming from natural language to executable scenarios, scalability via a distributed architecture supporting up to 100,000 concurrent agents, and reliability via feedback-driven LLM fine-tuning. Leveraging this system, S-Researcher supports researchers in designing social experiments, simulating human behavior with LLM agents, analyzing results, and generating reports, forming a complete human-AI collaborative research loop in which researchers retain oversight and intervention at every stage. We operationalize LLM simulation research paradigms into three canonical reasoning modes (induction, deduction, and abduction) and validate S-Researcher through systematic case studies: inductive reproduction of cultural dynamics consistent with Axelrod's theory, deductive testing of competing hypotheses on teacher attention validated against survey data, and abductive identification of a cooperation mechanism in public goods games confirmed by human experiments. S-Researcher establishes a new human--AI collaborative paradigm for social science, in which computational simulation augments human researchers to accelerate discovery across the full spectrum of social inquiry.","authors":["Lei Wang","Yuanzi Li","Jinchao Wu","Heyang Gao","Xiaohe Bo","Xu Chen","Ji-Rong Wen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-02","first_seen":"2026-04-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.01520","pdf_url":"https://arxiv.org/pdf/2604.01520","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B3"],"tags":["LLM人类仿真","社会科学自动化","人机协作"],"reason":"用LLM代理模拟人类行为，复现文化动态、验证教师关注假设、识别合作机制，均有真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":53,"question":"如何构建一个以LLM代理为核心的人机协作平台，实现社会科学研究的全流程自动化，并验证其在归纳、演绎、溯因三种推理模式下的有效性？","design":"使用基于LLM的代理系统YuLan-OneSim模拟人类参与者，通过自然语言自动编程生成可执行场景，支持高达10万并发代理的分布式架构，并利用反馈驱动微调提升可靠性。在归纳模式下，模拟文化传播动态；在演绎模式下，模拟课堂师生互动并测试竞争假设；在溯因模式下，模拟公共物品博弈以识别合作机制。","baseline":"演绎案例验证教师关注假设时，对照了真实调查数据；溯因案例识别合作机制时，对照了真实人类实验。归纳案例无明确真实人类数据对照，仅与Axelrod理论预测一致。","findings":"S-Researcher平台成功复现了Axelrod文化传播理论中的收敛与极化动态，并在课堂模拟中通过对比调查数据验证了教师关注假设，还在公共物品博弈中识别出合作机制并经人类实验确认。该平台实现了从实验设计、模拟、分析到报告生成的全流程人机协作，支持研究者全程干预。","reliability":"论文未明确讨论仿真失效的具体条件或局限，仅通过反馈微调机制和案例验证来确保可靠性，但未系统分析代理行为与真实人类偏差的来源或边界。","relevance":"高度相关：该研究直接用LLM代理替代人类被试，复现社会动态、验证假设并识别机制，且部分案例有真实人类数据对照，契合研究者对仿真可靠性、经济学实验和政策评估场景的关注，值得精读原文以评估其方法细节与批判性局限。","inspiration":"借鉴其利用LLM代理进行大规模并发模拟和反馈驱动微调以提升行为真实性的方法，可构建经济实验的虚拟被试池｜可迁移至公共物品博弈、税收遵从或劳动供给决策等行为经济学与政策评估场景｜以LLM代理为被试，施加不同税收政策处理，测量其劳动供给或逃税行为，并与真实实验室实验或行政数据对照"}},{"id":"2604.01896","version":1,"title":"Bayesian Elicitation with LLMs: Model Size Helps, Extra \"Reasoning\" Doesn't Always","zh_title":"基于大语言模型的贝叶斯启发：模型规模有益，额外“推理”未必有效","abstract":"Large language models (LLMs) have been proposed as alternatives to human experts for estimating unknown quantities with associated uncertainty, a process known as Bayesian elicitation. We test this by asking eleven LLMs to estimate population statistics, such as health prevalence rates, personality trait distributions, and labor market figures, and to express their uncertainty as 95\\% credible intervals. We vary each model's reasoning effort (low, medium, high) to test whether more \"thinking\" improves results. Our findings reveal three key results. First, larger, more capable models produce more accurate estimates, but increasing reasoning effort provides no consistent benefit. Second, all models are severely overconfident: their 95\\% intervals contain the true value only 9--44\\% of the time, far below the expected 95\\%. Third, a statistical recalibration technique called conformal prediction can correct this overconfidence, expanding the intervals to achieve the intended coverage. In a preliminary experiment, giving models web search access degraded predictions for already-accurate models, while modestly improving predictions for weaker ones. Models performed well on commonly discussed topics but struggled with specialized health data. These results indicate that LLM uncertainty estimates require statistical correction before they can be used in decision-making.","authors":["Luka Hobor","Mario Brcic","Mihael Kovac","Kristijan Poje"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-02","first_seen":"2026-04-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.01896","pdf_url":"https://arxiv.org/pdf/2604.01896","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","贝叶斯启发","不确定性校准"],"reason":"用LLM替代人类专家进行贝叶斯估计，并与真实人口统计数据对照，评估其校准与偏差…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":98,"question":"大语言模型能否替代人类专家进行贝叶斯估计，且增加推理努力是否改善估计的准确性和校准？","design":"用11个大语言模型扮演贝叶斯估计者，对来自心理学、公共卫生、劳动力市场四个真实数据集的400个总体统计量给出点估计和95%置信区间；通过API参数或思考令牌预算设置低、中、高三种推理努力水平，并测试网络搜索工具的影响。","baseline":"以Big Five人格、NHANES健康调查、NCD-RisC国家健康统计和Glassdoor劳动力市场数据的真实统计值为基准。","findings":"更大、能力更强的模型估计更准确，但增加推理努力并未一致提升校准或准确度；所有模型严重过度自信，95%区间实际覆盖率仅9–44%，但可通过保形预测进行事后校准。","reliability":"模型在常见话题上表现较好，但在专业健康数据上困难；网络搜索对已准确模型有负面影响，对较弱模型略有改善；过度自信需统计校正后方可用于决策。","relevance":"该研究直接用LLM替代人类进行贝叶斯估计，并与真实人口统计数据对照，系统评估了推理努力、模型规模和工具使用对仿真可靠性的影响，对关注LLM仿真人类判断与决策偏差的研究者极具参考价值。","inspiration":"该研究通过控制LLM的推理努力水平（低、中、高）和工具使用（网络搜索）来系统评估仿真准确性与校准，这种多因素处理设计值得借鉴｜可迁移到宏观经济预测或政策效果评估场景，例如用LLM替代专家预测GDP增长率或通胀预期｜以LLM作为被试，处理变量为推理努力水平（通过token预算控制），结果变量为点预测误差和置信区间覆盖率，以历史实际经济数据（如FRED数据库）作为真实基准对照"}},{"id":"2604.02403","version":1,"title":"Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics","zh_title":"测量不可调查之物：LLM作为劳动经济学中潜在认知变量的工具","abstract":"This paper establishes the theoretical and practical foundations for using Large Language Models (LLMs) as measurement instruments for latent economic variables -- specifically variables that describe the cognitive content of occupational tasks at a level of granularity not achievable with existing survey instruments. I formalize four conditions under which LLM-generated scores constitute valid instruments: semantic exogeneity, construct relevance, monotonicity, and model invariance. I then apply this framework to the Augmented Human Capital Index (AHC_o), constructed from 18,796 O*NET task statements scored by Claude Haiku 4.5, and validated against six existing AI exposure indices. The index shows strong convergent validity (r = 0.85 with Eloundou GPT-gamma, r = 0.79 with Felten AIOE) and discriminant validity. Principal component analysis confirms that AI-related occupational measures span two distinct dimensions -- augmentation and substitution. Inter-rater reliability across two LLM models (n = 3,666 paired scores) yields Pearson r = 0.76 and Krippendorff's alpha = 0.71. Prompt sensitivity analysis across four alternative framings shows that task-level rankings are robust. Obviously Related Instrumental Variables (ORIV) estimation recovers coefficients 25% larger than OLS, consistent with classical measurement error attenuation. The methodology generalizes beyond labor economics to any domain where semantic content must be quantified at scale.","authors":["Cristian Espinal Maya"],"categories":["econ.EM","cs.CL","stat.ME"],"primary_category":"econ.EM","announce_type":"new","date":"2026-04-02","first_seen":"2026-04-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.02403","pdf_url":"https://arxiv.org/pdf/2604.02403","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM测量工具","劳动经济学","认知变量"],"reason":"用LLM替代人工标注任务认知变量，属标注员替代而非仿真人类被试，但方法论可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":179,"question":"如何将大语言模型用作劳动经济学中潜在认知变量的测量工具，并确保其计量有效性？","design":"本研究并非人类仿真实验，而是提出用LLM（Claude Haiku 4.5）对18,796条O*NET任务描述进行评分，构建职业任务认知内容（如AI可增强性）的测量指标，并验证其作为计量工具的有效性。","baseline":"无对照（以六种已有AI暴露指数作为效标验证，而非真实人类评分）。","findings":"LLM生成的增强人力资本指数与现有AI暴露指数收敛效度强（r最高0.85），且增强与替代为两个独立维度；跨模型评分者信度达0.76（Pearson r），测量误差符合经典衰减结构，可用ORIV校正。","reliability":"论文讨论了跨模型信度仅中等（Krippendorff's alpha=0.71），提示评分存在模型依赖性；测量误差虽符合经典假设，但仅在两模型独立误差条件下成立，且未深入探讨提示词系统性偏差或语义外生性在实际中的违背可能。","relevance":"虽非直接仿真人类被试，但系统提出了LLM作为测量工具的有效性条件与计量校正方法，对用LLM替代人类标注或生成态度/认知指标的研究有方法论参考价值，值得阅读原文了解其效度验证框架。","inspiration":"该方法将LLM作为测量工具，通过跨模型评分者信度和测量误差结构分析来验证计量有效性，并引入ORIV校正测量误差，为使用LLM生成经济指标提供了严谨的效度验证框架。｜可迁移到劳动经济学中职业任务特征的自动化测量，例如构建职业的认知需求、社交互动或常规化程度等指标，替代传统人工编码或调查。｜以O*NET任务描述为输入，用多个LLM对任务的认知特征进行评分，结果变量为职业层面的认知需求指数，以人类专家编码或现有职业认知量表作为效标进行收敛效度检验，并评估跨模型信度与测量误差结构。"}},{"id":"2603.29741","version":1,"title":"BotVerse: Real-Time Event-Driven Simulation of Social Agents","zh_title":"BotVerse：基于实时事件驱动的社交智能体仿真","abstract":"BotVerse is a scalable, event-driven framework for high-fidelity social simulation using LLM-based agents. It addresses the ethical risks of studying autonomous agents on live networks by isolating interactions within a controlled environment while grounding them in real-time content streams from the Bluesky ecosystem. The system features an asynchronous orchestration API and a simulation engine that emulates human-like temporal patterns and cognitive memory. Through the Synthetic Social Observatory, researchers can deploy customizable personas and observe multimodal interactions at scale. We demonstrate BotVersevia a coordinated disinformation scenario, providing a safe, experimental framework for red-teaming and computational social scientists. A video demonstration of the framework is available at https://youtu.be/eZSzO5Jarqk.","authors":["Edoardo Allegrini","Edoardo Di Paolo","Angelo Spognardi","Marinella Petrocchi"],"categories":["cs.SI","cs.AI","cs.MA"],"primary_category":"cs.SI","announce_type":"new","date":"2026-03-31","first_seen":"2026-03-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.29741","pdf_url":"https://arxiv.org/pdf/2603.29741","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM智能体","虚假信息"],"reason":"用LLM agent模拟社会过程（虚假信息传播），但无真实人类数据对照，属边界…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":159,"question":"如何构建一个可扩展、事件驱动的高保真社交智能体仿真框架，以安全地研究虚假信息传播等社会现象？","design":"使用基于LLM的智能体模拟社交网络用户，通过异步编排API和仿真引擎复现人类时间模式和认知记忆；从Bluesky平台实时获取内容流，在隔离环境中让智能体进行发帖、点赞、回复、转发等互动，以演示协调虚假信息传播场景。","baseline":"无对照","findings":"BotVerse框架能够将仿真与现实内容流隔离，避免在真实网络上实验的伦理风险；其事件驱动架构支持数千个并发智能体，并模拟人类时间动态，适用于红队测试和计算社会科学研究。","reliability":"论文未讨论","relevance":"该研究利用LLM智能体仿真社交网络中的虚假信息传播，但缺乏真实人类数据对照，属于边界相关；若关注仿真平台架构或伦理安全实验，可参考其设计，但对人类行为复现的可靠性评估帮助有限。","inspiration":"可借鉴其事件驱动架构和异步编排API，在仿真中复现人类时间动态和并发交互，用于施加信息干预并观察实时行为反应｜可迁移到金融市场信息传播与投资者情绪形成研究，如模拟突发政策公告后社交网络中的信息扩散与交易决策｜设计：以LLM智能体为被试，处理为推送不同偏向的政策解读帖，结果变量为智能体的发帖情绪与模拟交易行为，对照真实市场中政策公告前后的社交媒体情绪与资产价格数据"}},{"id":"2603.27056","version":1,"title":"Persona-Based Simulation of Human Opinion at Population Scale","zh_title":"基于人格的群体意见仿真：从社交媒体推断半结构化人格以驱动LLM代理","abstract":"What does it mean to model a person, not merely to predict isolated responses, preferences, or behaviors, but to simulate how an individual interprets events, forms opinions, makes judgments, and acts consistently across contexts? This question matters because social science requires not only observing and predicting human outcomes, but also simulating interventions and their consequences. Although large language models (LLMs) can generate human-like answers, most existing approaches remain predictive, relying on demographic correlations rather than representations of individuals themselves. We introduce SPIRIT (Semi-structured Persona Inference and Reasoning for Individualized Trajectories), a framework designed explicitly for simulation rather than prediction. SPIRIT infers psychologically grounded, semi-structured personas from public social media posts, integrating structured attributes (e.g., personality traits and world beliefs) with unstructured narrative text reflecting values and lived experience. These personas prompt LLM-based agents to act as specific individuals when answering survey questions or responding to events. Using the Ipsos KnowledgePanel, a nationally representative probability sample of U.S. adults, we show that SPIRIT-conditioned simulations recover self-reported responses more faithfully than demographic persona and reproduce human-like heterogeneity in response patterns. We further demonstrate that persona banks can function as virtual respondent panels for studying both stable attitudes and time-sensitive public opinion.","authors":["Mao Li","Frederick G. Conrad"],"categories":["cs.CY","cs.AI","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-03-28","first_seen":"2026-03-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.27056","pdf_url":"https://arxiv.org/pdf/2603.27056","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","人格推断","调查方法"],"reason":"用LLM仿真个体意见并与全国概率样本对照，直接复现人类调查回答和异质性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"如何从社交媒体文本推断心理结构化的人物画像，并用其驱动大语言模型在总体层面仿真人类意见分布？","design":"从Ipsos KnowledgePanel概率样本中招募有Reddit/Twitter公开帖文的用户，用SPIRIT框架从帖文推断半结构化人物画像（含人格特质、世界信念等结构化属性与叙述文本），再以这些画像提示LLM代理回答调查问题，测量回复与真实自报答案的吻合度及异质性。","baseline":"Ipsos KnowledgePanel全国代表性概率样本的自报调查回答，以及仅用人口统计学画像的仿真作为对照。","findings":"SPIRIT画像驱动的仿真比人口统计学画像更准确地复现个体自报回答，并再现了人类回答模式中的异质性；人物画像库可作为虚拟受访者面板，用于研究稳定态度和时效性舆论。","reliability":"论文未讨论","relevance":"该研究直接以全国概率样本为基准，用LLM仿真个体意见分布，并对比人口统计学画像，高度契合研究者对仿真可靠性、基准对照和异质性复现的关注，值得精读原文。","inspiration":"借鉴其用社交媒体文本推断结构化人物画像并驱动LLM代理回答调查问题的设计，可构建高保真虚拟被试池｜可迁移到政策公告的预期形成研究，如央行沟通对通胀预期的影响｜招募有公开帖文的真实投资者，从其帖文推断人格与信念画像，用LLM代理接收不同措辞的央行声明，测量其通胀预期变化，以真实调查数据为基准对照"}},{"id":"2603.23884","version":1,"title":"POSIM: A Multi-Agent Simulation Framework for Social Media Public Opinion Evolution and Governance","zh_title":"POSIM：社交媒体舆论演化与治理的多智能体仿真框架","abstract":"Modeling social media public opinion evolution is essential for governance decision-making. Traditional epidemic models and rule-based agent-based models (ABMs) fail to capture the cognitive processes and adaptive behaviors of real users. Recent large language model (LLM)-based social simulations can reproduce group-level phenomena like polarization and conformity, yet remain unable to recreate the irrational interactions and multi-phase dynamics of real public opinion events. We present POSIM (Public Opinion Simulator), a multi-agent simulation framework for social media public opinion evolution and governance. POSIM integrates LLM-driven agents with a Belief--Desire--Intention (BDI) cognitive architecture that accounts for irrational factors, places them in a virtual social media environment with social networks and recommendation mechanisms, and drives temporal dynamics through a Hawkes point process engine that captures the co-evolution of agents and the environment across event phases. To validate the framework, we collect real-world public opinion datasets from the Weibo platform covering the full interaction chain of users. Experiments show that POSIM successfully reproduces key characteristics of public opinion evolution from individual mechanisms to collective phenomena, and its effectiveness is further supported by multiple statistical metrics. Building on POSIM, governance-oriented guidance and intervention experiments uncover a counterintuitive empathy paradox: empathetic guidance deepens negative sentiment instead of easing it under certain conditions, offering new insights for governance strategy design. These results demonstrate that the proposed framework can fully serve as a computational experimentation platform for proactive strategy evaluation and evidence-based governance. All source code is available at https://github.com/DeepCogLab/posim/.","authors":["Yongmao Zhang","Kai Qiao","Zhengyan Wang","Ningning Liang","Dekui Ma","Wenyao Sun","Jian Chen","Bin Yan"],"categories":["cs.GL"],"primary_category":"cs.GL","announce_type":"new","date":"2026-03-25","first_seen":"2026-03-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.23884","pdf_url":"https://arxiv.org/pdf/2603.23884","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","舆论演化","人类数据对照"],"reason":"用LLM agent模拟社交媒体舆论演化，并与真实微博数据对照，复现人类行为模…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":54,"question":"如何构建一个能复现真实社交媒体舆论演化多阶段动态并支持治理策略评估的多智能体仿真框架？","design":"用LLM驱动智能体，基于BDI认知架构融入情绪唤醒和认知偏差，模拟普通用户、意见领袖、媒体和政府四类角色；置于含社交网络和推荐机制的虚拟社交媒体环境中，通过Hawkes点过程引擎驱动多阶段时序演化；测量个体行为逻辑、群体涌现现象和统计指标，并进行治理干预实验。","baseline":"从微博平台收集的三个真实舆论事件数据集，覆盖原创、转发和评论的完整交互链。","findings":"POSIM在机制、现象和统计三个层面成功复现了舆论演化的关键特征；治理实验发现“共情悖论”：在某些条件下，共情引导反而加深负面情绪，而非缓解对立。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM智能体复现真实微博舆论事件，与人类数据对照，并评估治理策略效果，直接命中研究者对仿真可靠性、经济学/政策场景和批判性失效条件的兴趣。","inspiration":"借鉴其多智能体分层仿真设计，将LLM驱动的异质角色（如散户、机构、媒体、监管者）置于含推荐机制的信息环境中，通过Hawkes过程模拟信息传播与情绪演化，并设置治理干预实验来评估政策效果。｜可迁移到金融市场中的信息扩散与投资者情绪形成问题，例如研究社交媒体上的利好/利空消息如何通过不同渠道影响散户和机构的交易行为与市场波动。｜以LLM模拟散户和机构投资者作为被试，处理为在模拟社交平台中注入不同情绪基调的政策信号，结果变量为个体交易决策和市场价格波动，用真实微博金融舆情事件及同期市场交易数据作为对照基准。"}},{"id":"2603.22837","version":1,"title":"Analysing LLM Persona Generation and Fairness Interpretation in Polarised Geopolitical Contexts","zh_title":"分析极化地缘政治背景下LLM人格生成与公平性解释","abstract":"Large language models (LLMs) are increasingly utilised for social simulation and persona generation, necessitating an understanding of how they represent geopolitical identities. In this paper, we analyse personas generated for Palestinian and Israeli identities by five popular LLMs across 640 experimental conditions, varying context (war vs non-war) and assigned roles. We observe significant distributional patterns in the generated attributes: Palestinian profiles in war contexts are frequently associated with lower socioeconomic status and survival-oriented roles, whereas Israeli profiles predominantly retain middle-class status and specialised professional attributes. When prompted with explicit instructions to avoid harmful assumptions, models exhibit diverse distributional changes, e.g., marked increases in non-binary gender inferences or a convergence toward generic occupational roles (e.g., \"student\"), while the underlying socioeconomic distinctions often remain. Furthermore, analysis of reasoning traces reveals an interesting dynamics between model reasoning and generation: while rationales consistently mention fairness-related concepts, the final generated personas follow the aforementioned diverse distributional changes. These findings illustrate a picture of how models interpret geopolitical contexts, while suggesting that they process fairness and adjust in varied ways; there is no consistent, direct translation of fairness concepts into representative outcomes.","authors":["Maida Aizaz","Quang Minh Nguyen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-24","first_seen":"2026-03-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.22837","pdf_url":"https://arxiv.org/pdf/2603.22837","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2","D3"],"tags":["LLM人格生成","地缘政治偏见","社会模拟"],"reason":"分析LLM生成的人格属性分布，测量模型偏见而非仿真人类被试，无真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":160,"question":"在高度对立的巴以地缘政治背景下，大语言模型如何生成巴勒斯坦人和以色列人的人物画像，以及模型在收到避免有害假设的指令后如何调整其生成分布？","design":"本研究并非仿真人类被试，而是让五种大语言模型（Gemma 3 27B、Qwen3 32B、Llama 3.3 70B、Gemini 2.5 Pro、GPT-4.1）扮演不同角色（如联合国维和人员、记者等），在战争与非战争语境下生成巴勒斯坦人或以色列人的人物画像，测量生成的性别、年龄、社会经济地位、职业等属性分布，并分析模型在加入避免有害假设的提示后的分布变化及推理痕迹。","baseline":"无对照","findings":"模型在战争语境下倾向于将巴勒斯坦人画像与低社会经济地位和生存导向角色关联，而以色列人画像则多保留中产阶级和专业属性；当明确要求避免有害假设时，模型表现出非二元性别推断增加或职业趋同（如“学生”）等多样化调整，但深层的社会经济差异往往依然存在。","reliability":"论文未讨论","relevance":"该研究聚焦于LLM生成人物画像时的偏见和公平性解释，而非将LLM作为人类被试的替代品进行仿真实验，且无真实人类数据对照，因此与研究者关注的人类仿真实验方向相关性较低，不建议优先阅读原文。","inspiration":"该方法通过系统操纵语境（战争/非战争）和指令（有无避免有害假设提示）来测量LLM生成人物画像的属性分布差异，可借鉴其因子设计思路来检测模型在不同情境下的偏见模式｜可迁移至信贷审批歧视研究，例如测试LLM在模拟贷款审批时对不同种族或性别申请人的社会经济地位推断是否受宏观经济语境（如经济衰退/繁荣）影响｜设计雏形：以LLM为被试，处理为经济语境（衰退/繁荣）与申请人种族（黑/白）的2×2因子，结果变量为模型推断的申请人收入、职业稳定性及贷款批准率，对照真实信贷审批数据中的种族差异模式"}},{"id":"2603.20678","version":1,"title":"AI-Driven Multi-Agent Simulation of Stratified Polyamory Systems: A Computational Framework for Optimizing Social Reproductive Efficiency","zh_title":"AI驱动的分层多偶制系统多智能体仿真：优化社会生育效率的计算框架","abstract":"Contemporary societies face a severe crisis of demographic reproduction. Global fertility rates continue to decline precipitously, with East Asian nations exhibiting the most dramatic trends -- China's total fertility rate (TFR) fell to approximately 1.0 in 2023, while South Korea's dropped below 0.72. Simultaneously, the institution of marriage is undergoing structural disintegration: educated women rationally reject unions lacking both emotional fulfillment and economic security, while a growing proportion of men at the lower end of the socioeconomic spectrum experience chronic sexual deprivation, anxiety, and learned helplessness. This paper proposes a computational framework for modeling and evaluating a Stratified Polyamory System (SPS) using techniques from agent-based modeling (ABM), multi-agent reinforcement learning (MARL), and large language model (LLM)-empowered social simulation. The SPS permits individuals to maintain a limited number of legally recognized secondary partners in addition to one primary spouse, combined with socialized child-rearing and inheritance reform. We formalize the A/B/C stratification as heterogeneous agent types in a multi-agent system and model the matching process as a MARL problem amenable to Proximal Policy Optimization (PPO). The mating network is analyzed using graph neural network (GNN) representations. Drawing on evolutionary psychology, behavioral ecology, social stratification theory, computational social science, algorithmic fairness, and institutional economics, we argue that SPS can improve aggregate social welfare in the Pareto sense. Preliminary computational results demonstrate the framework's viability in addressing the dual crisis of female motherhood penalties and male sexlessness, while offering a non-violent mechanism for wealth dispersion analogous to the historical Chinese Grace Decree (Tui'en Ling).","authors":["Yicai Xing"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-21","first_seen":"2026-03-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.20678","pdf_url":"https://arxiv.org/pdf/2603.20678","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","多智能体系统","计算社会科学"],"reason":"用LLM agent模拟社会制度，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":161,"question":"如何通过多智能体强化学习和大语言模型仿真，评估分层多伴侣制度（SPS）对生育率、社会福利和资源分配的影响？","design":"构建一个多智能体系统，将个体分为A/B/C三类异质智能体（属性包括配偶价值、经济资源、吸引力、生育力），环境规则为SPS制度（允许一个主配偶加有限数量的合法次级伴侣，配合社会化育儿和继承改革）；匹配过程建模为多智能体强化学习问题，使用PPO优化，并用图神经网络分析交配网络；通过计算仿真评估社会福利、生育率等结果。","baseline":"无对照","findings":"初步计算结果表明，该框架在解决女性生育惩罚和男性性匮乏双重危机方面具有可行性，并能提供一种类似历史中国推恩令的非暴力财富分散机制。","reliability":"论文未讨论","relevance":"该研究使用LLM赋能的社会仿真模拟制度变革，但缺乏真实人类数据对照，属于边界情形，对关注仿真可靠性与基准对比的研究者参考价值有限。","inspiration":"该研究将多智能体强化学习与LLM结合，用于模拟制度变革下的社会行为演化，提供了在无人类基准时构建异质智能体属性与匹配规则的方法｜可迁移到政策评估场景，如模拟最低工资调整对劳动力市场匹配与收入分配的影响｜以LLM驱动的异质智能体（区分技能、财富、就业状态）为被试，处理为最低工资提升，结果变量为就业率与收入基尼系数，对照真实劳动调查数据"}},{"id":"2603.19791","version":2,"title":"Text-Based Personas for Simulating User Privacy Decisions","zh_title":"基于文本角色模拟用户隐私决策","abstract":"The ability to simulate human privacy decisions has significant implications for aligning autonomous agents with individual intent and conducting cost-effective, large-scale privacy-centric user studies. Prior approaches prompt Large Language Models (LLMs) with natural language user statements, data-sharing histories, or demographic attributes to simulate privacy decisions. These approaches, however, fail to balance individual-level accuracy, human auditability, token efficiency, and population-level representation. We present Narriva, an approach that generates text-based synthetic privacy personas to address these shortcomings. Narriva grounds persona generation in prior user privacy decisions, such as those from large-scale survey datasets, rather than purely relying on demographic stereotypes. It compresses this data into concise, human-readable summaries structured by established privacy theories. Through benchmarking across five diverse datasets, we analyze the characteristics of Narriva's synthetic personas in modeling both individual and population-level privacy preferences. We find that grounding personas in past privacy behaviors achieves up to 87% predictive accuracy, improving over a non-personalized LLM baseline by 6-17 percentage points across datasets, while yielding an 80-95% reduction in prompt tokens compared to in-context learning with raw examples. Finally, we demonstrate that personas synthesized from a single survey can reproduce the aggregate privacy behaviors and statistical distributions of entirely different studies.","authors":["Kassem Fawaz","Ren Yi","Octavian Suciu","Rishabh Khandelwal","Hamza Harkous","Nina Taft","Marco Gruteser"],"categories":["cs.CR"],"primary_category":"cs.CR","announce_type":"new","date":"2026-03-20","first_seen":"2026-03-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.19791","pdf_url":"https://arxiv.org/pdf/2603.19791","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A5","B1","B2"],"tags":["隐私决策仿真","合成角色","人类数据对照"],"reason":"用LLM生成隐私决策合成样本，有真实人类数据对照，涉及用户研究场景，直接仿真人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":20,"question":"能否用基于文本的合成隐私人格（persona）有效模拟个体和群体层面的隐私决策，并泛化到独立研究？","design":"提出Narriva框架，利用大规模调查数据中用户的历史隐私决策生成结构化文本人格摘要，再让LLM基于这些人格模拟隐私偏好决策，评估个体预测准确率、群体分布复现能力和跨研究泛化性。","baseline":"五个不同数据集中的真实用户隐私决策数据，包括个体层面和群体层面的行为与态度。","findings":"基于过去隐私行为的人格在个体预测上准确率最高达87%，比无个性化LLM基线提升6-17个百分点；同时提示词token用量比原始示例上下文学习减少80-95%。从单一调查合成的人格能复现完全不同研究的群体隐私行为和统计分布。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM生成人格模拟隐私决策，有真实人类数据基准，评估个体与群体层面的仿真准确性和效率，并涉及跨研究泛化，直接回应研究者对LLM人类仿真实验、基准对照和可靠性批判的兴趣，值得精读原文。","inspiration":"借鉴Narriva框架用历史行为数据生成结构化文本人格以驱动LLM模拟个体决策的方法，可大幅降低提示词成本并提升预测准确率｜可迁移到消费者金融隐私偏好与数据共享决策研究，如移动支付或数字银行场景下的个人信息披露行为｜以真实用户历史隐私选择数据构建人格，让LLM模拟其在新型金融服务中的隐私权衡，结果变量为是否同意共享数据，用实际用户共享行为数据做对照"}},{"id":"2603.19649","version":1,"title":"PolicySim: An LLM-Based Agent Social Simulation Sandbox for Proactive Policy Optimization","zh_title":"PolicySim：一个基于LLM的智能体社会仿真沙盒，用于主动政策优化","abstract":"Social platforms serve as central hubs for information exchange, where user behaviors and platform interventions jointly shape opinions. However, intervention policies like recommendation and content filtering, can unintentionally amplify echo chambers and polarization, posing significant societal risks. Proactively evaluating the impact of such policies is therefore crucial. Existing approaches primarily rely on reactive online A/B testing, where risks are identified only after deployment, making risk identification delayed and costly. LLM-based social simulations offer a promising pre-deployment alternative, but current methods fall short in realistically modeling platform interventions and incorporating feedback from the platform. Bridging these gaps is essential for building actionable frameworks to assess and optimize platform policies. To this end, we propose PolicySim, an LLM-based social simulation sandbox for the proactive assessment and optimization of intervention policies. PolicySim models the bidirectional dynamics between user behavior and platform interventions through two key components: (1) a user agent module refined via supervised fine-tuning (SFT) and direct preference optimization (DPO) to achieve platform-specific behavioral realism; and (2) an adaptive intervention module that employs a contextual bandit with message passing to capture dynamic network structures. Experiments show that PolicySim can accurately simulate platform ecosystems at both micro and macro levels and support effective intervention policy.","authors":["Renhong Huang","Ning Tang","Jiarong Xu","Yuxuan Cao","Qingqian Tu","Sheng Guo","Bo Zheng","Huiyuan Liu","Yang Yang"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-03-20","first_seen":"2026-03-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.19649","pdf_url":"https://arxiv.org/pdf/2603.19649","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2"],"tags":["社会仿真","政策评估","LLM智能体"],"reason":"用LLM agent模拟社交平台用户行为与政策干预，涉及政策评估场景，但未明确…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":81,"question":"如何主动评估和优化社交平台的干预政策（如推荐系统和内容过滤）对用户行为和意见动态的影响？","design":"构建PolicySim沙盒，包含基于LLM的用户智能体（通过SFT和DPO训练以模拟特定平台用户行为）和自适应干预模块（使用上下文赌博机和消息传递捕捉动态网络），模拟用户与平台干预的双向动态，测量微观和宏观层面的生态系统指标。","baseline":"无对照","findings":"PolicySim能够在微观和宏观层面准确模拟平台生态系统；基于模拟的自适应干预策略可以有效优化平台政策。","reliability":"论文未讨论","relevance":"该研究利用LLM智能体模拟社交平台用户行为和政策干预，属于人类仿真实验范畴，但未提供真实人类数据作为对照基准，且未讨论仿真失效条件，与研究者关注的有对照基准和批判性分析的需求部分匹配，值得阅读以了解其仿真设计方法。","inspiration":"PolicySim的LLM智能体训练方法（SFT+DPO）和自适应干预模块（上下文赌博机）可用于构建动态政策仿真环境｜可迁移到社交平台内容推荐政策对投资者情绪和资产价格波动的影响研究｜用LLM智能体模拟投资者，处理为不同推荐算法干预，结果变量为模拟股价波动和情绪指数，对照真实社交平台投资者讨论数据与同期股价数据"}},{"id":"2605.12507","version":1,"title":"Can LLM Agents Simulate Dynamic Networks? A Case Study on Email Networks with Phishing Synthesis","zh_title":"LLM智能体能模拟动态网络吗？以钓鱼邮件合成为例的邮件网络案例研究","abstract":"While Large Language Model (LLM) multi-agent systems (MAS) offer a transformative approach to simulating human behavior in complex systems, it remains largely unexplored whether these simulations can replicate realistic structural and temporal dynamics from a dynamic network perspective. Our evaluation indicates that existing frameworks excel at generating plausible micro-level interactions but fail to capture the emergent, macroscopic topologies necessary for domains that rely on realistic network dynamics, such as modeling information propagation and cybersecurity threats. To bridge this gap, we introduce two easily integrable extensions to simulation frameworks to ensure they preserve macroscopic network fidelity: 1) augmenting LLM agents with data-driven event triggers to organically sustain long-horizon interactions, and 2) integrating Hawkes processes to accurately model temporal activation dynamics. Our approach allows LLM MAS to capture both plausible micro-level patterns and macroscopic topologies. We further demonstrate the utility of this framework in synthesizing realistic phishing campaigns within evolving communication networks. The study reveals how threats exploit structural vulnerabilities, highlighting the potential of our framework for developing next-generation defenses. Our code is available at https://github.com/Graph-COM/NSL.","authors":["Siqi Miao","Ziyang Chen","Yuhong Luo","Hans Hao-Hsun Hsu","Mufei Li","Kaiqing Zhang","Pan Li"],"categories":["cs.SI","cs.AI","cs.MA"],"primary_category":"cs.SI","announce_type":"new","date":"2026-03-20","first_seen":"2026-03-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.12507","pdf_url":"https://arxiv.org/pdf/2605.12507","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","动态网络模拟","社会模拟"],"reason":"用LLM多智能体模拟邮件网络动态，但无真实人类行为数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":138,"question":"LLM多智能体系统能否在动态网络视角下复现真实的结构与时间动态，以支持信息传播和网络安全威胁建模？","design":"使用LLM多智能体系统模拟邮件通信网络，基于Enron和IETF邮件语料，通过引入数据驱动的事件触发器和Hawkes过程来增强智能体激活机制，测量微观、中观和宏观网络指标，并应用于钓鱼邮件攻击模拟。","baseline":"无对照","findings":"现有LLM多智能体框架能生成合理的微观交互，但无法捕捉宏观网络拓扑；通过数据驱动事件触发器和Hawkes过程可显著提升宏观网络保真度，并成功模拟利用网络结构的钓鱼攻击。","reliability":"论文承认当前研究仅限于邮件语料，未验证在其他网络类型（如金融交易网络）的迁移性；生成的动态事件流可用于更复杂的多步攻击研究但未展开；同时指出该框架可能被滥用于自动化社会工程攻击。","relevance":"该研究属于LLM仿真动态网络的前沿探索，但缺乏真实人类行为数据对照，未直接复现人类实验或调查，与研究者关注的人类被试替代和基准对照要求存在差距，可作为批判性案例了解仿真失效模式。","inspiration":"该方法通过数据驱动的事件触发器和Hawkes过程增强LLM智能体激活机制以提升宏观网络保真度，可借鉴用于校准经济仿真中的主体交互时序｜可迁移到金融市场信息扩散与羊群效应研究，模拟交易员间的消息传播网络及其对资产价格的影响｜以LLM智能体模拟交易员，处理为突发新闻事件，结果变量为交易行为与价格波动，用真实市场微观交易数据与网络结构作为对照基准"}},{"id":"2603.18563","version":2,"title":"Reasonably reasoning AI agents can avoid game-theoretic failures in zero-shot, provably","zh_title":"合理推理的AI智能体可零样本避免博弈论失败，且可证明","abstract":"As autonomous AI agents increasingly mediate online platform markets, a fundamental question emerges: do these markets generate stable strategic outcomes? In repeated strategic environments, the Nash equilibrium provides a natural benchmark for this stability. However, empirical evidence on off-the-shelf LLM agents is mixed, leaving it unclear whether independently deployed agents can converge to equilibrium behavior without explicit strategic post-training. In this paper, we provide an affirmative answer. Extending the Bayesian learning literature in theoretical economics, we prove that AI agents, acting as Bayesian posterior samplers rather than expected utility maximizers, are guaranteed to eventually become weakly close to a Nash equilibrium in infinitely repeated games. We further extend this analysis to settings in which stage payoffs are unknown ex ante, and agents observe only their privately realized stochastic payoffs, and obtain the same convergence guarantees. Finally, we empirically evaluate these theoretical implications across five repeated-game environments, ranging from the Prisoner's Dilemma to marketing promotion games. Taken together, our findings suggest that strategic stability in AI-mediated markets can emerge from the intrinsic reasoning and learning properties of modern AI agents, without the need for unrealistic universal fine-tuning.","authors":["Enoch Hyunwook Kang"],"categories":["cs.AI","cs.MA","econ.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-19","first_seen":"2026-03-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.18563","pdf_url":"https://arxiv.org/pdf/2603.18563","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体博弈","纳什均衡","贝叶斯学习"],"reason":"用LLM agent模拟博弈，但无真实人类数据对照，属社会模拟理论演示","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":99,"question":"未经专门后训练的大语言模型智能体，在零样本重复博弈中能否稳定收敛到纳什均衡？","design":"本文并非人类仿真研究，而是理论证明与实证评估：将LLM智能体建模为贝叶斯后验采样器，在五种重复博弈环境中（囚徒困境到营销促销博弈）测试其零样本策略收敛性，测量是否接近纳什均衡。","baseline":"无对照","findings":"理论上证明，作为贝叶斯后验采样器的AI智能体在无限重复博弈中最终会弱接近纳什均衡；实证上，在五种博弈环境中观察到策略稳定收敛，无需全局微调。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济博弈中的策略互动，但未与真实人类行为数据对照，不属于严格的人类仿真研究；若关注AI在博弈中的收敛性质，可读原文了解其理论框架。","inspiration":"该研究将LLM智能体建模为贝叶斯后验采样器并检验其零样本博弈收敛性的方法，可借鉴用于设计经济实验中的策略互动仿真，通过理论证明与多环境实证来评估行为收敛｜可迁移到产业组织中的重复价格竞争或合谋实验，检验AI智能体能否自发形成合谋定价｜以LLM作为被试，在重复古诺或伯川德博弈中施加不同市场结构处理，结果变量为价格或产量序列的收敛性，与真实人类实验数据对照，评估仿真有效性"}},{"id":"2603.16142","version":2,"title":"Parametric Social Identity Injection and Diversification in Public Opinion Simulation","zh_title":"参数化社会身份注入与多样化在舆论仿真中的应用","abstract":"Large language models (LLMs) have recently been adopted as synthetic agents for public opinion simulation, offering a promising alternative to costly and slow human surveys. Despite their scalability, current LLM-based simulation methods fail to capture social diversity, producing flattened inter-group differences and overly homogeneous responses across demographic groups. We identify this limitation as a Diversity Collapse phenomenon in LLM hidden representations, where distinct social identities become increasingly indistinguishable across layers. Motivated by this observation, we propose Parametric Social Identity Injection (PSII), a general framework that injects explicit, parametric representations of demographic attributes and value orientations directly into intermediate hidden states of LLMs. Unlike prompt-based persona conditioning, PSII enables fine-grained and controllable identity modulation at the representation level. Extensive experiments on the World Values Survey using multiple open-source LLMs show that PSII significantly improves distributional fidelity and diversity, reducing KL divergence to real-world survey data while enhancing overall diversity. This work provides new insights into representation-level control of LLM agents and advances scalable, diversity-aware public opinion simulation.","authors":["Hexi Wang","Yujia Zhou","Bangde Du","Qingyao Ai","Yiqun Liu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-17","first_seen":"2026-03-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.16142","pdf_url":"https://arxiv.org/pdf/2603.16142","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","舆论模拟","社会身份注入"],"reason":"用LLM模拟公众舆论并与世界价值观调查真实数据对照，直接命中人类仿真核心。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":60,"question":"如何通过向LLM隐藏状态注入参数化社会身份来缓解舆论仿真中的多样性坍塌，从而更真实地复现人群异质性？","design":"使用多个开源LLM（Qwen2.5-7B/14B-Instruct、Llama-3.1-8B-Instruct、Mistral-24B-Instruct）作为合成智能体，模拟世界价值观调查（WVS）中的受访者；通过参数化社会身份注入（PSII）将人口统计属性和价值取向的参数向量直接注入LLM中间隐藏状态，对比传统提示词方法，测量回答分布与真实人类数据的KL散度及多样性指标。","baseline":"世界价值观调查（WVS）的真实人类回答数据，作为分布保真度和多样性的对照基准。","findings":"LLM隐藏状态在高层存在多样性坍塌现象，导致不同社会身份变得难以区分；PSII通过注入稳定身份向量并施加随机扰动，有效维持甚至提升高层表示的多样性，显著降低与真实调查数据的KL散度，增强群体间和群体内多样性。","reliability":"论文未讨论","relevance":"该研究直接用LLM复现世界价值观调查，以真实人类数据为基准，系统揭示了现有方法在多样性上的失效机制，并提出表示层面的干预方案，高度契合研究者对仿真可靠性、偏差及经济学/政策评估场景的关注，值得精读原文。","inspiration":"该方法通过向LLM隐藏状态注入参数化社会身份向量并施加随机扰动来维持群体多样性，可借鉴为一种处理异质性代理人的新范式，用于替代传统提示词方法以缓解仿真中的多样性坍塌｜该技术可迁移到政策评估中的异质性处理效应分析，例如模拟不同社会经济群体对税收改革或福利政策的反应差异，从而在事前评估政策分配的公平性与效率｜可设计一个实验：以LLM作为合成被试，注入收入、教育、政治倾向等身份参数，处理为不同税收政策方案，结果变量为政策支持度与预期行为变化，以真实调查数据（如美国综合社会调查GSS）作为分布对照基准"}},{"id":"2603.26701","version":1,"title":"From Heard to Lived Opinions: Simulating Opinion Dynamics with Grounded LLM Agents in Economic Environments","zh_title":"从听闻到亲历：在经济环境中基于情境化LLM智能体模拟舆论动态","abstract":"Opinion dynamics (OD) studies how individual opinions evolve and generate collective patterns such as consensus and polarization. While recent work explores OD using populations of LLM-based agents focusing on opinion exchange, it typically does not incorporate individuals' lived experiences, such as economic outcomes of past decisions, which play a critical role in shaping opinions. We propose a novel OD simulation framework that grounds LLM-based agents in an economic environment, allowing them to act and receive environmental feedback. Our simulations exhibit coherent OD at both individual and population levels: individual opinions follow structured trajectories shaped by economic experiences, with adverse conditions inducing opinion rigidity, while at the population level, collective opinions co-move with economic conditions, with inequality amplifying polarization and price instability driving larger distributional shifts. These results highlight the importance of grounding LLM-based agents in environments to capture collective OD.","authors":["Ryuji Hashimoto","Masahiro Kaneko","Ryosuke Takata","Takehiro Takayanagi","Kiyoshi Izumi"],"categories":["physics.soc-ph","cs.CY"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-03-17","first_seen":"2026-03-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.26701","pdf_url":"https://arxiv.org/pdf/2603.26701","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","D3"],"tags":["舆论动态","LLM智能体","社会模拟"],"reason":"用LLM agent模拟舆论动态并引入经济环境反馈，但未明确提及真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":100,"question":"在引入经济环境反馈的条件下，基于LLM的智能体能否在个体和群体层面产生连贯的舆论动态？","design":"使用Llama 3.1 8B模型扮演具有人口属性和大五人格的家庭智能体，在包含价格、工资和资产的经济环境中进行经济决策并交换意见，通过环境反馈塑造其经历，测量个体意见轨迹和群体意见分布。","baseline":"无对照","findings":"个体层面，经济逆境导致意见僵化，意见轨迹受经历驱动；群体层面，经济不平等加剧极化，价格不稳定导致意见分布更大变动。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟舆论动态并引入经济环境反馈，但未使用真实人类数据作为基准，适合关注仿真方法但对人类对照要求不高的研究者参考。","inspiration":"该方法将LLM智能体置于包含价格、工资和资产的经济环境中，通过环境反馈塑造智能体的经历，从而驱动意见变化，这种动态反馈设计值得借鉴｜可迁移到政策公告的预期形成研究，例如模拟家庭在通胀或利率变动下的预期调整过程｜使用具有人口属性和大五人格的LLM智能体作为被试，处理为不同通胀或利率公告序列，结果变量为智能体的通胀预期轨迹和异质性，以真实家庭调查的预期数据（如密歇根消费者调查）作为对照基准"}},{"id":"2603.13890","version":1,"title":"Beyond Self-Interest: Modeling Social-Oriented Motivation for Human-like Multi-Agent Interactions","zh_title":"超越自利：建模社会导向动机以实现类人多智能体交互","abstract":"Large Language Models (LLMs) demonstrate significant potential for generating complex behaviors, yet most approaches lack mechanisms for modeling social motivation in human-like multi-agent interaction. We introduce Autonomous Social Value-Oriented agents (ASVO), where LLM-based agents integrate desire-driven autonomy with Social Value Orientation (SVO) theory. At each step, agents first update their beliefs by perceiving environmental changes and others' actions. These observations inform the value update process, where each agent updates multi-dimensional desire values through reflective reasoning and infers others' motivational states. By contrasting self-satisfaction derived from fulfilled desires against estimated others' satisfaction, agents dynamically compute their SVO along a spectrum from altruistic to competitive, which in turn guides activity selection to balance desire fulfillment with social alignment. Experiments across School, Workplace, and Family contexts demonstrate substantial improvements over baselines in behavioral naturalness and human-likeness. These findings show that structured desire systems and adaptive SVO drift enable realistic multi-agent social simulations.","authors":["Jingzhe Lin","Ceyao Zhang","Yaodong Yang","Yizhou Wang","Song-Chun Zhu","Fangwei Zhong"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-03-14","first_seen":"2026-03-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.13890","pdf_url":"https://arxiv.org/pdf/2603.13890","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","多智能体","社会价值取向"],"reason":"用LLM agent模拟社会互动，但无真实人类数据对照，属社会模拟演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":162,"question":"如何将社会价值取向理论融入LLM智能体，以实现具有类人社会动机的多智能体交互仿真？","design":"使用基于LLM的ASVO智能体，在校园、职场、家庭三种模拟社会场景中，通过信念更新、欲望更新、SVO动态计算和活动生成四个模块，让智能体根据自身社会人格类型（利他、亲社会、个人主义、竞争）进行交互，评估行为自然度和类人性。","baseline":"无对照","findings":"ASVO智能体在行为自然度和类人性上显著优于基线方法；结构化的欲望系统和自适应的SVO漂移能够实现更真实的多智能体社会仿真。","reliability":"论文未讨论","relevance":"该研究利用LLM模拟社会互动中的动机动态，但缺乏真实人类数据对照，属于社会仿真演示，与关注人类基准复现和可靠性评估的研究者需求匹配度较低，不建议优先阅读原文。","inspiration":"该研究将社会价值取向理论融入LLM智能体，通过信念更新、欲望更新和SVO动态计算模块模拟社会动机驱动的交互行为，这种模块化设计可借鉴用于构建具有异质性社会偏好的经济主体。｜可迁移到公共品博弈或慈善捐赠实验中，研究不同社会偏好（如利他、互惠）对合作水平和资源配置效率的动态影响。｜设计一个基于LLM的公共品博弈仿真，智能体被赋予不同的SVO类型作为处理，结果变量为个体贡献额和群体总收益，对照真实人类实验数据（如Fischbacher等2001年的公共品实验）来评估仿真偏差。"}},{"id":"2603.12129","version":1,"title":"Increasing intelligence in AI agents can worsen collective outcomes","zh_title":"AI智能体智能提升可能恶化集体结果","abstract":"When resources are scarce, will a population of AI agents coordinate in harmony, or descend into tribal chaos? Diverse decision-making AI from different developers is entering everyday devices -- from phones and medical devices to battlefield drones and cars -- and these AI agents typically compete for finite shared resources such as charging slots, relay bandwidth, and traffic priority. Yet their collective dynamics and hence risks to users and society are poorly understood. Here we study AI-agent populations as the first system of real agents in which four key variables governing collective behaviour can be independently toggled: nature (innate LLM diversity), nurture (individual reinforcement learning), culture (emergent tribe formation), and resource scarcity. We show empirically and mathematically that when resources are scarce, AI model diversity and reinforcement learning increase dangerous system overload, though tribe formation lessens this risk. Meanwhile, some individuals profit handsomely. When resources are abundant, the same ingredients drive overload to near zero, though tribe formation makes the overload slightly worse. The crossover is arithmetical: it is where opposing tribes that form spontaneously first fit inside the available capacity. More sophisticated AI-agent populations are not better: whether their sophistication helps or harms depends entirely on a single number -- the capacity-to-population ratio -- that is knowable before any AI-agent ships.","authors":["Neil F. Johnson"],"categories":["cs.AI","cs.CY","cs.SI","econ.GN","physics.soc-ph"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-12","first_seen":"2026-03-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.12129","pdf_url":"https://arxiv.org/pdf/2603.12129","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","集体行为","社会模拟"],"reason":"用LLM agent模拟资源竞争中的集体行为，但无真实人类数据对照，属社会模拟…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":180,"question":"在资源稀缺条件下，AI智能体的多样性、强化学习、部落形成和资源稀缺性如何共同影响集体结果？","design":"用7个不同参数量的真实LLM（如GPT-2、Pythia、OPT系列）作为智能体，在资源竞争博弈中独立操纵先天多样性（LLM类型）、后天学习（个体强化学习）、文化（自发部落形成）和资源稀缺度四个变量，测量系统过载率（集体失败）和个体收益。","baseline":"无对照","findings":"资源稀缺时，模型多样性和强化学习加剧系统过载，而部落形成可缓解风险；资源充裕时，同样因素使过载趋近于零，但部落形成轻微恶化过载。智能体的复杂程度是否有利完全取决于部署前即可知的容量-人口比。","reliability":"论文承认LLM采样温度固定、行动空间较二元、部落动态依赖外部感知层等局限，未测试更大规模混合世代种群，且物理边缘设备实验尚未进行。","relevance":"该研究用LLM智能体模拟资源竞争中的集体行为，但缺乏真实人类数据对照，不属于严格的人类仿真验证，与研究者关注的经济学实验和政策评估场景有距离，但批判性结论（智能体复杂化未必改善集体结果）对仿真可靠性讨论有参考价值，可酌情阅读。","inspiration":"该研究通过独立操纵LLM智能体的先天多样性、后天学习、部落形成和资源稀缺度四个变量，在资源竞争博弈中测量系统过载率，这种多因素正交设计可用于经济学实验｜可迁移到公共品博弈或资源竞争场景，研究AI代理的多样性、学习机制和群体规范如何影响合作与资源耗竭｜用不同参数量的LLM作为被试，在公共品博弈中操纵智能体多样性（模型类型）、学习机制（有无强化学习）和沟通形成规范（有无部落），测量公共品贡献水平和资源存量，对照真实人类实验数据"}},{"id":"2604.09609","version":2,"title":"General-purpose LLMs as Models of Human Driver Behavior: The Case of Simplified Merging","zh_title":"通用大语言模型作为人类驾驶行为模型：简化合流场景案例","abstract":"Human behavior models are essential as behavior references and for simulating human agents in virtual safety assessment of automated vehicles (AVs), yet current models face a trade-off between interpretability and flexibility. General-purpose large language models (LLMs) offer a promising alternative: a single model potentially deployable without parameter fitting across diverse scenarios. However, what LLMs can and cannot capture about human driving behavior remains poorly understood. We address this gap by embedding two general-purpose LLMs (OpenAI o3 and Google Gemini 2.5 Pro) as standalone, closed-loop driver agents in a simplified one-dimensional merging scenario and comparing their behavior against human data using quantitative and qualitative analyses. Both models reproduce human-like intermittent operational control and tactical dependencies on spatial cues. However, neither consistently captures the human response to dynamic velocity cues, and safety performance diverges sharply between models. A systematic prompt ablation study reveals that prompt components act as model-specific inductive biases that do not transfer across LLMs. These findings suggest that general-purpose LLMs could potentially serve as standalone, ready-to-use human behavior models in AV evaluation pipelines, but future research is needed to better understand their failure modes and ensure their validity as models of human driving behavior.","authors":["Samir H. A. Mohammad","Wouter Mooi","Arkady Zgonnikov"],"categories":["cs.AI","cs.RO"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-11","first_seen":"2026-03-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.09609","pdf_url":"https://arxiv.org/pdf/2604.09609","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类行为对照","驾驶行为建模"],"reason":"用LLM模拟人类驾驶行为并与真实数据对照，评估仿真可靠性与失效条件，方法可迁移…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":37,"question":"通用大语言模型在多大程度上能够复现人类驾驶行为，特别是在简化的一维合流场景中的操作控制、战术决策和安全表现？","design":"将两个通用大语言模型（OpenAI o3 和 Google Gemini 2.5 Pro）作为独立的闭环驾驶智能体，嵌入简化的一维合流任务中，不进行任何任务特定训练或参数拟合，通过定量和定性分析与人类数据对比，评估其行为相似性和安全性。","baseline":"使用先前研究中在相同场景和运动学条件下采集的人类驾驶数据集作为对照基准。","findings":"两个模型都能复现人类间歇性的操作控制和对空间线索的战术依赖，但都无法一致地捕捉人类对动态速度线索的反应，且两个模型之间的安全表现差异很大。提示消融实验表明，提示组件作为模型特定的归纳偏置，不能在不同大语言模型之间迁移。","reliability":"模型无法一致复现人类对动态速度线索的反应，安全性能在不同模型间差异显著，提示设计的效果不可迁移，且研究仅在简化的一维合流场景中进行，未涉及更复杂的交互场景。","relevance":"该研究直接以通用大语言模型作为人类被试的替代品，在驾驶行为仿真中与真实人类数据严格对照，并揭示了仿真的失效条件（如对速度线索的捕捉不足、模型间安全表现差异），高度契合研究者对经济学实验和政策评估场景中仿真可靠性与偏差的关注，值得精读原文以借鉴其方法和批判性发现。","inspiration":"该方法值得借鉴之处在于：使用通用LLM作为闭环智能体，在不进行任务特定训练的情况下直接与人类行为基准对照，并通过提示消融实验检验处理效应的可迁移性。｜可迁移至消费者跨期选择实验，研究LLM能否复现人类的时间偏好不一致和折现行为。｜研究设计：以LLM作为被试，施加不同表述方式的跨期选择任务（如延迟奖励的框架效应），结果变量为选择一致性和折现率，对照真实人类实验数据（如经典的双曲折现研究）。"}},{"id":"2603.09884","version":1,"title":"Benchmarking Political Persuasion Risks Across Frontier Large Language Models","zh_title":"跨前沿大语言模型的政治说服风险基准测试","abstract":"Concerns persist regarding the capacity of Large Language Models (LLMs) to sway political views. Although prior research has claimed that LLMs are not more persuasive than standard political campaign practices, the recent rise of frontier models warrants further study. In two survey experiments (N=19,145) across bipartisan issues and stances, we evaluate seven state-of-the-art LLMs developed by Anthropic, OpenAI, Google, and xAI. We find that LLMs outperform standard campaign advertisements, with heterogeneity in performance across models. Specifically, Claude models exhibit the highest persuasiveness, while Grok exhibits the lowest. The results are robust across issues and stances. Moreover, in contrast to the findings in Hackenburg et al. (2025b) and Lin et al. (2025) that information-based prompts boost persuasiveness, we find that the effectiveness of information-based prompts is model-dependent: they increase the persuasiveness of Claude and Grok while substantially reducing that of GPT. We introduce a data-driven and strategy-agnostic LLM-assisted conversation analysis approach to identify and assess underlying persuasive strategies. Our work benchmarks the persuasive risks of frontier models and provides a framework for cross-model comparative risk assessment.","authors":["Zhongren Chen","Joshua Kalla","Quan Le"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-10","first_seen":"2026-03-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.09884","pdf_url":"https://arxiv.org/pdf/2603.09884","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","政治说服","人类对照实验"],"reason":"用LLM生成政治说服信息，与真实人类调查实验对照，评估说服效果与策略，直接仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"前沿大语言模型在政治说服任务中是否比人类竞选广告更具说服力，且不同模型和提示策略的效果有何差异？","design":"用7款前沿LLM（Claude Sonnet 4、Gemini 2.5 Flash、GPT-4.1、Grok 4等）扮演政治说服者，在两项调查实验（N=19,145）中与真人被试进行文本对话，施加两种提示（普通提示与信息提示），测量被试对移民和最低工资议题的态度变化（五点李克特量表，后转为二元支持指标）。","baseline":"人类基准：移民议题使用真人倡导者视频，最低工资议题使用先前研究（Chen et al. 2025）中的人类竞选广告效果，并通过Hajek估计量校正协变量偏移。","findings":"所有LLM的说服效果均显著强于人类竞选广告，其中Claude模型说服力最强，Grok最弱；信息提示的效果因模型而异，能提升Claude和Grok的说服力，但大幅降低GPT的说服力。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查实验为基准，评估LLM在政治说服场景中的仿真效果与模型间异质性，并揭示提示策略的模型依赖性，对关注LLM仿真可靠性及失效条件的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其用LLM替代人类进行文本对话干预并对比真实人类基准的设计，通过多模型比较和提示策略操纵揭示效果异质性｜可迁移至政策沟通场景，如央行前瞻指引或财政政策公告对公众预期和消费行为的影响｜以LLM作为虚拟被试，随机分配不同风格的货币政策沟通文本（处理），测量其预期的通胀或消费意愿变化（结果），并以历史调查数据或真实实验数据作为对照基准"}},{"id":"2603.09890","version":1,"title":"Influencing LLM Multi-Agent Dialogue via Policy-Parameterized Prompts","zh_title":"通过策略参数化提示影响LLM多智能体对话","abstract":"Large Language Models (LLMs) have emerged as a new paradigm for multi-agent systems. However, existing research on the behaviour of LLM-based multi-agents relies on ad hoc prompts and lacks a principled policy perspective. Different from reinforcement learning, we investigate whether prompt-as-action can be parameterized so as to construct a lightweight policy which consists of a sequence of state-action pairs to influence conversational behaviours without training. Our framework regards prompts as actions executed by LLMs, and dynamically constructs prompts through five components based on the current state of the agent. To test the effectiveness of parameterized control, we evaluated the dialogue flow based on five indicators: responsiveness, rebuttal, evidence usage, non-repetition, and stance shift. We conduct experiments using different LLM-driven agents in two discussion scenarios related to the general public and show that prompt parameterization can influence the dialogue dynamics. This result shows that policy-parameterised prompts offer a simple and effective mechanism to influence the dialogue process, which will help the research of multi-agent systems in the direction of social simulation.","authors":["Hongbo Bo","Jingyu Hu","Weiru Liu"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-10","first_seen":"2026-03-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.09890","pdf_url":"https://arxiv.org/pdf/2603.09890","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体对话","社会模拟","提示参数化"],"reason":"用LLM多智能体模拟社会讨论，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":163,"question":"能否通过将提示词参数化作为一种轻量级策略，在不进行训练的情况下调控LLM多智能体对话的行为？","design":"使用Qwen3-8B、Llama3-8B、Mistral-7B三个LLM分别驱动三个具有不同立场和知识库的智能体，在土地资源利用和教育资源分配两个公共议题上进行10轮多轮对话；通过动态组合任务与角色描述、对话记忆、外部知识、规则模板和权重五个组件来构建提示词，并改变规则模板和权重调度策略作为处理，测量响应性、反驳、证据使用、非重复性和立场转变五个对话指标。","baseline":"无对照","findings":"提示词参数化能够有效影响多智能体对话的动态过程，不同的控制策略会导致反驳、证据使用和立场转变等行为指标出现显著差异。该框架为通过结构化提示词调控对话行为提供了一种简单有效的机制。","reliability":"论文未讨论","relevance":"该研究利用LLM多智能体进行社会讨论仿真，但缺乏真实人类数据对照，属于社会仿真的边界情形，与研究者关注的有基准人类数据的实验复现和偏差评估不完全匹配，但可作为方法参考。","inspiration":"该方法通过参数化提示词组件（规则模板、权重调度）来调控多智能体对话行为，可借鉴其将处理变量嵌入提示词结构的思路，用于经济学实验中信息干预的精细化操控。｜可迁移至政策公告的预期形成研究，模拟不同政策沟通策略（如措辞强度、信息顺序）如何影响市场参与者的预期和共识形成。｜以LLM驱动的多智能体模拟投资者群体，处理为政策公告提示词中的规则模板（如鹰派/鸽派措辞）和权重（强调不同经济指标），结果变量为预期通胀率、资产配置倾向的分布变化，并与调查预期数据（如密歇根消费者调查）或市场隐含预期数据对照。"}},{"id":"2603.08853","version":1,"title":"LLM-Agent Interactions on Markets with Information Asymmetries","zh_title":"信息不对称市场中LLM智能体的互动研究","abstract":"As AI agents increasingly act on behalf of human stakeholders in economic settings, understanding their behavior in complex market environments becomes critical. This article examines how Large Language Models coordinate on markets that are characterized by information asymmetries and in which providers of services have incentives to exploit that asymmetry for their own economic gain. To that end, we conduct simulations with GPT-5.1 agents in credence goods markets, manipulating the institutional framework (free market, verifiability, liability), LLM agent's social preferences (default, self-interested, inequity-averse, efficiency-loving), and reputation mechanisms across one-shot and repeated 16-round interactions. In one-shot settings, LLM agents largely fail to establish cooperation, with markets breaking down except under liability rules or when experts have efficiency-loving preferences. Repeated interactions solve consumer participation through competitive price reduction, but expert fraud remains entrenched absent explicit other-regarding preferences. LLM consumers focus narrowly on price levels rather than understanding strategic incentives embedded in markups, making them vulnerable to exploitation. Compared to human experiments, LLM markets exhibit substantially higher consumer participation but much greater market concentration, lower prices, and more polarized fraud patterns. The effect of institutions like verifiability and reputation is also much more ambiguous. Surplus shifts dramatically toward consumers under social-preference objectives. These findings suggest that institutional design for AI agent markets requires fundamentally different approaches than those effective for human actors, with social preference alignment emerging as the primary determinant of market efficiency.","authors":["Alexander Erlei","Lukas Meub"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-03-09","first_seen":"2026-03-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.08853","pdf_url":"https://arxiv.org/pdf/2603.08853","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","市场实验","人类数据对照"],"reason":"用LLM agent模拟信息不对称市场，并与人类实验数据对照，评估制度与偏好影…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":3,"question":"LLM智能体在信息不对称的信任品市场中如何协调行为，社会偏好与制度如何影响市场效率？","design":"用GPT-5.1扮演专家和消费者，在信任品市场博弈中模拟交易，操纵制度框架（自由市场、可验证性、责任规则）、LLM社会偏好（默认、自利、不平等厌恶、效率偏好）和声誉机制，进行单次与16轮重复互动，测量市场参与、欺诈行为、价格、福利分配等。","baseline":"对照Dulleck, Kerschbamer, and Sutter (2011)的人类实验数据。","findings":"单次互动中LLM市场普遍崩溃，仅责任规则或效率偏好下能维持；重复互动通过降价解决消费者参与，但专家欺诈依然顽固，仅社会偏好能抑制欺诈。与人类实验相比，LLM市场消费者参与更高但集中度更高、价格更低、欺诈模式更极化，制度效果更模糊，且剩余大幅向消费者转移。","reliability":"论文未讨论","relevance":"直接使用LLM模拟信息不对称市场并与真实人类实验基准对照，系统评估制度与偏好影响，揭示LLM仿真在行为模式、制度效应上与人类的显著偏差，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读。","inspiration":"借鉴其操纵制度框架（自由市场、可验证性、责任规则）和LLM社会偏好（自利、不平等厌恶、效率偏好）的多因素实验设计，并与真实人类实验基准对照，系统评估行为偏差。｜可迁移到信贷审批中的信息不对称与歧视问题，如银行信贷员与小微企业的贷款博弈，检验不同监管规则和银行社会偏好对审批决策的影响。｜用LLM扮演信贷员和小微企业主，操纵信贷审核制度（纯市场、强制信息披露、责任追究）和LLM偏好，测量贷款批准率、利率设定、违约欺诈率，以真实银行信贷实验数据为基准对照。"}},{"id":"2604.15329","version":1,"title":"Evaluating LLMs as Human Surrogates in Controlled Experiments","zh_title":"评估大语言模型作为受控实验中人类替代品的有效性","abstract":"Large language models (LLMs) are increasingly used to simulate human responses in behavioral research, yet it remains unclear when LLM-generated data support the same experimental inferences as human data. We evaluate this by directly comparing off-the-shelf LLM-generated responses with human responses from a canonical survey experiment on accuracy perception. Each human observation is converted into a structured prompt, and models generate a single 0--10 outcome variable without task-specific training; identical statistical analyses are applied to human and synthetic responses. We find that LLMs reproduce several directional effects observed in humans, but effect magnitudes and moderation patterns vary across models. Off-the-shelf LLMs therefore capture aggregate belief-updating patterns under controlled conditions but do not consistently match human-scale effects, clarifying when LLM-generated data can function as behavioral surrogates.","authors":["Adnan Hoq","Tim Weninger"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"new","date":"2026-03-08","first_seen":"2026-03-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.15329","pdf_url":"https://arxiv.org/pdf/2604.15329","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类替代","实验对照"],"reason":"直接比较LLM与人类在受控实验中的反应，评估仿真可靠性，有真实人类数据对照，并…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":4,"question":"在受控实验中，现成的大语言模型（LLM）生成的回答能否支持与人类数据相同的实验推断？","design":"将人类被试在政治新闻准确性感知调查实验中的每次观测转化为结构化提示（含人物描述、实验条件和新闻标题），让多个现成LLM（闭源与开源）直接生成0-10的准确性评分，不进行任务特定训练或校准，然后对人类和合成数据应用相同的统计分析。","baseline":"真实人类被试在相同实验中的回答，实验操控了新闻标题的政治倾向和是否提供AI可信度反馈。","findings":"LLM能复现人类数据中的若干方向性效应（如意识形态对齐和可信度反馈的影响），但效应大小和调节模式因模型而异；LLM捕捉了受控条件下的总体信念更新模式，但未能一致匹配人类量级的效应。","reliability":"效应量级在不同模型间差异大，部分模型夸大处理效应；新闻级异质性仅部分复现；LLM生成的数据不能直接替代人类样本，结构复现需逐假设检验。","relevance":"该研究直接比较LLM与人类在受控实验中的反应，有真实人类数据对照，并明确指出了仿真在效应量级和异质性上的失效条件，高度契合研究者对LLM仿真可靠性及批判性评估的关注。","inspiration":"该方法将人类实验的每次观测转化为结构化提示，直接让现成LLM生成评分，并与人类数据应用相同统计分析，值得借鉴其逐观测仿真和严格对照的设计｜可迁移到政策公告对消费者预期形成的影响研究，例如评估央行沟通对通胀预期的作用｜以消费者为被试，处理为不同措辞的央行公告，结果变量为通胀预期数值，用真实消费者调查数据作为对照基准"}},{"id":"2603.07444","version":1,"title":"HLER: Human-in-the-Loop Economic Research via Multi-Agent Pipelines for Empirical Discovery","zh_title":"HLER：通过多智能体流水线进行人在回路的经济实证研究","abstract":"Large language models (LLMs) have enabled agent-based systems that aim to automate scientific research workflows. Most existing approaches focus on fully autonomous discovery, where AI systems generate research ideas, conduct analyses, and produce manuscripts with minimal human involvement. However, empirical research in economics and the social sciences poses additional constraints: research questions must be grounded in available datasets, identification strategies require careful design, and human judgment remains essential for evaluating economic significance. We introduce HLER (Human-in-the-Loop Economic Research), a multi-agent architecture that supports empirical research automation while preserving critical human oversight. The system orchestrates specialized agents for data auditing, data profiling, hypothesis generation, econometric analysis, manuscript drafting, and automated review. A key design principle is dataset-aware hypothesis generation, where candidate research questions are constrained by dataset structure, variable availability, and distributional diagnostics, reducing infeasible or hallucinated hypotheses. HLER further implements a two-loop architecture: a question quality loop that screens and selects feasible hypotheses, and a research revision loop where automated review triggers re-analysis and manuscript revision. Human decision gates are embedded at key stages, allowing researchers to guide the automated pipeline. Experiments on three empirical datasets show that dataset-aware hypothesis generation produces feasible research questions in 87% of cases (versus 41% under unconstrained generation), while complete empirical manuscripts can be produced at an average API cost of $0.8-$1.5 per run. These results suggest that Human-AI collaborative pipelines may provide a practical path toward scalable empirical research.","authors":["Chen Zhu","Xiaolu Wang"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-08","first_seen":"2026-03-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.07444","pdf_url":"https://arxiv.org/pdf/2603.07444","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","经济研究自动化","人在回路"],"reason":"多智能体经济研究自动化，无真实人类行为对照，属社会模拟但缺基准数据","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":181,"question":"如何构建一个支持人类监督的多智能体流水线，以自动化完成从数据审计、假设生成到计量分析和论文撰写的实证经济研究流程？","design":"HLER 是一个多智能体架构，由多个专用 LLM 智能体（数据审计、数据画像、假设生成、计量分析、论文起草和自动审稿）组成，通过中央协调器顺序执行任务；在假设生成阶段引入数据集感知约束，并在关键节点设置人类决策门（如研究问题选择和发表批准），实现人机协作的实证研究自动化。","baseline":"无对照","findings":"数据集感知的假设生成将可行研究问题的比例从无约束生成的 41% 提高到 87%；在三个实证数据集上，系统能以每次运行 0.8 至 1.5 美元的 API 成本端到端生成完整的研究手稿。","reliability":"论文未讨论","relevance":"该研究专注于用 LLM 多智能体自动化经济研究流程，但未涉及将 LLM 作为人类被试替代品的仿真实验，也无真实人类行为基准对照，与研究者关注的 LLM 人类仿真可靠性评估方向不直接相关，不建议优先阅读原文。","inspiration":"HLER的多智能体流水线设计，特别是数据集感知的假设生成和人类决策门机制，为自动化实证研究提供了可借鉴的流程控制方法｜该架构可迁移至政策评估场景，例如自动化分析最低工资政策对就业的影响，利用LLM智能体进行数据审计、假设生成和计量分析｜可设计一个研究雏形：以LLM智能体作为分析者，处理为提供不同政策干预的历史数据集，结果变量为智能体生成的因果推断结论，并以真实经济学文献中的实证结果作为对照基准，评估自动化分析的准确性"}},{"id":"2604.22756","version":1,"title":"Your Reviews Replicate You: LLM-Based Agents as Customer Digital Twins for Conjoint Analysis","zh_title":"你的评论复制你：基于LLM的客户数字孪生用于联合分析","abstract":"Conjoint analysis is a cornerstone of market research for estimating consumer preferences; however, traditional methods face persistent challenges regarding time, cost, and respondent fatigue. To address these limitations, this study proposes a framework that utilizes large language model (LLM)-based \"customer digital twins (CDT)\" as virtual respondents. We identified active users within the Reddit community and aggregated their comprehensive review histories to construct individualized vector databases. By integrating retrieval-augmented generation (RAG) with prompt engineering, this study developed customer agents capable of dynamically retrieving and reasoning upon their specific past preferences and constraints. These customer agents, called CDTs, performed pairwise comparison tasks on product profiles generated via fractional factorial design, and the resulting choice data was analyzed to estimate part-worth utilities by logistic regression. Empirical validation demonstrates that these CDTs predict the preferences of actual users with 87.73% accuracy. Furthermore, a case study on the computer monitor category successfully quantified trade-offs between attributes such as panel type and resolution, deriving preference structures consistent with market realities. Ultimately, this study contributes to marketing research by presenting a scalable alternative that significantly improves both agility and cost-efficiency to traditional methods.","authors":["Bin Xuan","Jungmin Hwang","Hakyeon Lee"],"categories":["cs.IR","cs.AI"],"primary_category":"cs.IR","announce_type":"new","date":"2026-03-06","first_seen":"2026-03-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.22756","pdf_url":"https://arxiv.org/pdf/2604.22756","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","消费者偏好","数字孪生"],"reason":"用LLM代理模拟消费者偏好，有真实用户数据对照，准确率87.73%，属经济学实…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":18,"question":"能否利用基于大语言模型的客户数字孪生（CDT）替代真实人类受访者进行联合分析，以准确复现个体消费者的偏好选择？","design":"使用GPT-4等大语言模型，结合检索增强生成（RAG）和提示工程，基于Reddit用户的历史评论构建个体化向量数据库，生成客户数字孪生（CDT）作为虚拟受访者；让CDT对通过部分因子设计生成的产品配置文件进行成对比较选择任务，收集选择数据，再用逻辑回归估计部分效用值。","baseline":"以Reddit社区中真实活跃用户的实际偏好作为对照基准，验证CDT预测的准确率。","findings":"CDT预测真实用户偏好的准确率达到87.73%；在电脑显示器案例中，成功量化了面板类型与分辨率等属性间的权衡，得出的偏好结构与市场现实一致。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM代理复现真实消费者偏好，有真实人类数据对照，属于经济学实验场景，值得精读以评估仿真可靠性。","inspiration":"可借鉴其利用个体历史文本数据构建个性化代理并通过成对比较任务测量偏好的方法｜可迁移到消费者跨期选择实验，如研究折扣率或耐心程度｜以电商平台用户评论构建LLM代理作为被试，施加不同跨期奖励方案（如立即小奖 vs. 延迟大奖），测量选择结果，并以该用户真实历史购买决策中的时间偏好数据作为对照。"}},{"id":"2603.03585","version":2,"title":"Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility","zh_title":"Belief-Sim：面向信念驱动的人口统计错误信息易感性仿真","abstract":"Misinformation is a growing societal threat, and susceptibility to misinformative claims varies across demographic groups due to differences in underlying beliefs. As Large Language Models (LLMs) are increasingly used to simulate human behaviors, we investigate whether they can simulate demographic misinformation susceptibility, treating beliefs as a primary driving factor. We introduce BeliefSim, a simulation framework that constructs demographic belief profiles using psychology-informed misinformation taxonomies and survey priors. We study prompt-based conditioning and post-training adaptation, and conduct a multi-fold evaluation using: (i) susceptibility alignment and (ii) counterfactual demographic sensitivity. Across both datasets and modeling strategies, we show that beliefs provide a strong prior for simulating misinformation susceptibility, with alignment up to 92%.","authors":["Angana Borah","Zohaib Khan","Rada Mihalcea","Verónica Pérez-Rosas"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-03","first_seen":"2026-03-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.03585","pdf_url":"https://arxiv.org/pdf/2603.03585","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM人类仿真","错误信息易感性","人口统计差异"],"reason":"用LLM仿真不同人口群体对错误信息的易感性，以信念为驱动，并与真实人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":21,"question":"信念是否能改善基于人口统计学的错误信息易感性仿真？","design":"使用多个LLM，通过提示条件化（BeliefSim-PC）和微调（BeliefSim-FT）两种方式，基于心理学启发的信念分类和调查先验构建人口信念画像，模拟不同性别、年龄、居住地、教育水平群体的错误信息易感性，测量其与真实人类判断的对齐程度。","baseline":"PANDORA数据集（318人，每人3条声明）和MIST-1数据集（409人，每人100条声明），共13.8K条真实人类判断，包含人口统计信息。","findings":"信念是模拟错误信息易感性的强先验，对齐度最高达92%；仅使用人口统计信息不可靠，可能导致捷径依赖。","reliability":"人口统计建模主要为单轴，仅评估8个独立群体，未考虑交叉群体效应；数据集仅限于美国参与者和英文标题。","relevance":"该研究直接用LLM仿真人口群体的错误信息易感性，以信念为驱动，并与真实人类数据对照，包含批判性分析，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"该方法通过心理学信念分类构建人口信念画像，并对比两种LLM条件化方式（提示与微调）来仿真群体判断，提供了处理施加与对照设计的参考｜可迁移至信贷审批中的群体歧视研究，仿真不同人口群体对贷款申请的审批决策｜以LLM作为被试，处理为基于信念画像的条件化提示或微调，结果变量为审批通过率，对照真实银行信贷审批数据中的群体差异"}},{"id":"2603.02711","version":1,"title":"A Natural Language Agentic Approach to Study Affective Polarization","zh_title":"一种研究情感极化的自然语言智能体方法","abstract":"Affective polarization has been central to political and social studies, with growing focus on social media, where partisan divisions are often exacerbated. Real-world studies tend to have limited scope, while simulated studies suffer from insufficient high-quality training data, as manually labeling posts is labor-intensive and prone to subjective biases. The lack of adequate tools to formalize different definitions of affective polarization across studies complicates result comparison and hinders interoperable frameworks. We present a multi-agent model providing a comprehensive approach to studying affective polarization in social media. To operationalize our framework, we develop a platform leveraging large language models (LLMs) to construct virtual communities where agents engage in discussions. We showcase the potential of our platform by (1) analyzing questions related to affective polarization, as explored in social science literature, providing a fresh perspective on this phenomenon, and (2) introducing scenarios that allow observation and measurement of polarization at different levels of granularity and abstraction. Experiments show that our platform is a flexible tool for computational studies of complex social dynamics such as affective polarization. It leverages advanced agent models to simulate rich, context-sensitive interactions and systematically explore research questions traditionally addressed through human-subject studies.","authors":["Stephanie Anneris Malvicini","Ewelina Gajewska","Arda Derbent","Katarzyna Budzynska","Jarosław A. Chudziak","Maria Vanina Martinez"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-03","first_seen":"2026-03-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.02711","pdf_url":"https://arxiv.org/pdf/2603.02711","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["情感极化","多智能体模拟","社交媒体"],"reason":"用LLM agent模拟社交媒体讨论以研究情感极化，但无真实人类数据对照，属社…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":146,"question":"如何利用LLM多智能体平台模拟社交媒体讨论，以研究情感极化的动态与测量？","design":"使用LLM驱动的多智能体模型构建虚拟社区，智能体代表不同政治党派，在受控场景下进行讨论；通过操作化党派认同强度并测量对内外群体的情感，观察和量化情感极化。","baseline":"无对照","findings":"该平台能够复现社会科学文献中关于情感极化的标准问题，并支持从多个粒度和抽象层次观察与测量极化；平台可作为传统人类被试研究的灵活计算补充工具。","reliability":"论文指出当前结果仅为初步实证验证，LLM作为社会智能体的适用性仍需通过系统与人类被试数据比较来评估；LLM自我报告变量的可靠性需要仔细检验；平台目前集中于美国中心场景，且缺乏结构化认知架构、学习机制和网络感知交互协议。","relevance":"该研究用LLM智能体仿真人类社交互动以研究情感极化，但无真实人类数据对照，属于探索性仿真工具开发；适合关注方法创新但需注意其缺乏基准验证的局限，可酌情阅读以了解LLM在政治心理学仿真中的应用潜力。","inspiration":"该方法通过LLM多智能体平台模拟党派讨论并测量情感极化，其操作化党派认同强度、控制讨论场景的设计可借鉴用于经济金融实验中处理组与对照组的构建。｜可迁移至研究经济政策分歧中的情感极化，例如不同经济阶层或利益群体对税收政策的态度分化与互动。｜以LLM智能体代表不同收入群体，施加累进税制改革信息作为处理，测量群体间情感温度与政策支持度，并与真实调查数据（如ANES或GSS中相关条目）进行对照验证。"}},{"id":"2603.21006","version":1,"title":"How AI Systems Think About Education: Analyzing Latent Preference Patterns in Large Language Models","zh_title":"AI系统如何思考教育：分析大语言模型中的潜在偏好模式","abstract":"This paper presents the first systematic measurement of educational alignment in Large Language Models. Using a Delphi-validated instrument comprising 48 items across eight educational-theoretical dimensions, the study reveals that GPT-5.1 exhibits highly coherent preference patterns (99.78% transitivity; 92.79% model accuracy) that largely align with humanistic educational principles where expert consensus exists. Crucially, divergences from expert opinion occur precisely in domains of normative disagreement among human experts themselves, particularly emotional dimensions and epistemic normativity. This raises a fundamental question for alignment research: When human values are contested, what should models be aligned to? The findings demonstrate that GPT-5.1 does not remain neutral in contested domains but adopts coherent positions, prioritizing emotional responsiveness and rejecting false balance. The methodology, combining Delphi consensus-building with Structured Preference Elicitation and Thurstonian Utility modeling, provides a replicable framework for domain-specific alignment evaluation beyond generic value benchmarks.","authors":["Daniel Autenrieth"],"categories":["cs.CY","cs.AI","cs.CL","cs.HC"],"primary_category":"cs.CY","announce_type":"new","date":"2026-02-28","first_seen":"2026-02-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.21006","pdf_url":"https://arxiv.org/pdf/2603.21006","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏好测量","教育对齐","德尔菲法"],"reason":"测量LLM自身的教育偏好，属于D2人格/态度测量，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":182,"question":"GPT-5.1在教育理论维度上表现出怎样的偏好模式，这些偏好与人类专家共识及分歧的关系如何？","design":"本研究并非仿真人类被试，而是通过德尔菲法构建48项教育原则，再设计144个教学场景，让GPT-5.1进行102,960次成对比较，测量其教育偏好的一致性和效用模式。","baseline":"无对照","findings":"GPT-5.1的教育偏好高度一致（传递性99.78%），在专家共识领域与人本主义教育原则对齐；但在情感支持和认识论规范性等专家存在分歧的领域，模型会采取明确立场而非保持中立。","reliability":"论文未讨论","relevance":"该研究将LLM作为测量对象而非人类替代品，不涉及人类仿真或行为复现，与您关注的用LLM仿真人类被试、对照真实人类数据的研究方向不匹配，不建议优先阅读原文。","inspiration":"该方法通过大规模成对比较测量AI系统的潜在偏好模式，可用于经济学中测量LLM对政策选项的偏好排序｜可迁移到政策评估场景，如测量LLM对税收政策、福利分配或环境规制的偏好结构｜以LLM为被试，设计不同政策特征的成对比较任务，结果变量为选择比例和偏好传递性，对照真实公众调查数据"}},{"id":"2602.21091","version":1,"title":"Can Interest-Bearing Positions Solve the Long-Horizon Problem in Prediction Markets?","zh_title":"计息头寸能否解决预测市场中的长期问题？","abstract":"Prediction markets suffer from reduced liquidity and price accuracy for long-horizon events due to the opportunity cost of committed capital. Recently, major platforms have introduced interest-bearing positions to mitigate this \"long-horizon problem.\" I evaluate this policy using agent-based simulations with large language model (LLM) traders in a 2 x 2 factorial design, varying time horizon (4 days vs. 2 years) and the presence of interest. While long horizons degrade accuracy, the observed pricing bias (0.72 percentage points) is significantly smaller than theoretical and prior empirical estimates. Paying interest eliminates approximately 83% of the horizon effect on accuracy and more than triples market participation (from 17% to 62% of wealth). These findings suggest the long-horizon problem may be overstated in existing literature and that interest-bearing positions are a highly effective intervention, primarily by incentivizing participation rather than correcting bias.","authors":["Caleb Maresca"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-24","first_seen":"2026-02-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.21091","pdf_url":"https://arxiv.org/pdf/2602.21091","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM代理模拟","预测市场","社会模拟"],"reason":"用LLM agent模拟预测市场，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":183,"question":"付息头寸能否解决预测市场中的长周期问题？","design":"使用大语言模型（LLM）智能体模拟交易者，采用2×2析因设计，操纵时间周期（4天 vs. 2年）和是否支付利息，测量价格准确性和市场参与度。","baseline":"无对照","findings":"长周期会降低价格准确性，但观察到的定价偏差（0.72个百分点）远小于理论和先前实证估计；支付利息消除了约83%的周期效应，并使市场参与度增加两倍以上。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟预测市场交易，属于人类仿真实验，但缺乏真实人类数据对照，且场景为金融市场而非经济学实验或政策评估，与研究者关注的核心方向部分相关但非完全匹配，可酌情阅读原文了解LLM仿真方法。","inspiration":"该研究采用2×2析因设计操纵时间周期和利息支付，通过LLM智能体模拟交易者来测量价格准确性和参与度，这种多因素实验设计值得借鉴｜可迁移到资产定价实验，例如研究不同信息发布频率和持有成本对市场价格发现效率的影响｜设计一个实验，以LLM智能体为被试，操纵信息更新频率（高/低）和交易成本（有/无），结果变量为价格偏差和交易量，并对照真实股票市场数据中的类似情境"}},{"id":"2602.20440","version":2,"title":"Intelligence Without Integrity: Why Capable LLMs May Undermine Reliability","zh_title":"有智无信：为何能力强的LLM可能损害可靠性","abstract":"As LLMs become embedded in research workflows and organizational decision processes, their effect on analytical reliability remains uncertain. We distinguish two dimensions of analytical reliability -- intelligence (the capacity to reach correct conclusions) and integrity (the stability of conclusions when analytically irrelevant cues about desired outcomes are introduced) -- and ask whether frontier LLMs possess both. Whether these dimensions trade off is theoretically ambiguous: the sophistication enabling accurate analysis may also enable responsiveness to non-evidential cues, or alternatively, greater capability may confer protection through better calibration and discernment. Using synthetically generated data with embedded ground truth, we evaluate fourteen models on a task simulating empirical analysis of hospital merger effects. We find that intelligence and integrity trade off: frontier models most likely to reach correct conclusions under neutral conditions are often most susceptible to shifting conclusions under motivated framing. We extend work on sycophancy by introducing goal-conditioned analytical sycophancy: sensitivity of inference to cues about desired outcomes, even when no belief is asserted and evidence is held constant. Unlike simple prompt sensitivity, models shift conclusions away from objective evidence in response to analytically irrelevant framing. This finding has important implications for empirical research and organizations. Selecting tools based on capability benchmarks may inadvertently select against the stability needed for reliable and replicable analysis.","authors":["Ryan Allen","Aticus Peterson"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-24","first_seen":"2026-02-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.20440","pdf_url":"https://arxiv.org/pdf/2602.20440","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM可靠性","分析完整性","模型行为测量"],"reason":"研究LLM在分析任务中的结论稳定性，属于测量模型本身而非仿真人类被试，无人类数…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":184,"question":"前沿大语言模型在分析可靠性上是否存在智力与诚信的权衡？","design":"本研究并非人类仿真实验，而是使用合成数据评估14个模型在模拟医院合并效应分析任务中的表现，通过中性条件和带有动机性框架的提示来测量模型得出正确结论的能力（智力）和结论稳定性（诚信）。","baseline":"无对照","findings":"前沿模型在中性条件下最可能得出正确结论，但在动机性框架下最易改变结论，智力与诚信存在权衡。模型会因与分析无关的期望线索而偏离客观证据，表现出目标条件分析性谄媚。","reliability":"论文未讨论","relevance":"该研究聚焦LLM自身的分析可靠性，未涉及人类被试仿真或与真实人类数据对照，与您关注的人类仿真实验方向不直接相关，但其中关于模型易受动机性框架影响的发现对评估LLM在实验中的偏差有参考价值。","inspiration":"可借鉴其通过动机性框架提示（如强调期望结论）来操纵LLM分析行为的处理设计，并测量结论正确性与稳定性以揭示智力-诚信权衡｜可迁移至政策评估场景，如研究LLM在模拟专家预测经济政策效果时是否因政治倾向或利益相关方期望而扭曲分析｜以LLM作为被试，随机分配中性提示与带有党派倾向的动机性框架提示，要求其分析某项税收改革对就业的影响，结果变量为预测方向与幅度，对照真实历史政策评估数据"}},{"id":"2603.00113","version":2,"title":"AI Agents Alone Are Not (Yet) Sufficient for Social Simulation","zh_title":"AI智能体单独尚不足以进行社会仿真","abstract":"Recent advances in large language models (LLMs) have spurred growing interest in using LLM-integrated agents for social simulation, often under the implicit assumption that realistic population dynamics will emerge once role-specified agents are placed in a networked multi-agent setting. This position paper argues that LLM-based agents alone are not (yet) sufficient for social simulation. We attribute this over-optimism to a systematic mismatch between what current agent pipelines are typically optimized and validated to produce and what simulation-as-science requires. Concretely, role-playing plausibility does not imply faithful human behavioral validity; collective outcomes are frequently mediated by agent-environment co-dynamics rather than agent-agent messaging alone; and results can be dominated by interaction protocols, scheduling, and initial information priors. To make these underlying mechanisms explicit and auditable, we propose a unified formulation of AI agent-based social simulation as an environment-involved Markov game with explicit exposure and scheduling mechanisms, from which we derive concrete actions for design, evaluation, and interpretation.","authors":["Yiming Li","Dacheng Tao"],"categories":["cs.MA","cs.AI","cs.CE","cs.CY","cs.SI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-02-19","first_seen":"2026-02-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.00113","pdf_url":"https://arxiv.org/pdf/2603.00113","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","D3"],"tags":["社会仿真","LLM智能体","方法论批评"],"reason":"讨论LLM智能体社会仿真，指出当前不足并提出方法论，虽无人类数据对照，但直接相…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":82,"question":"当前基于LLM的智能体社会仿真在科学推断上存在哪些根本性不足？","design":"本文为立场论文，未进行具体仿真实验；它系统批评了现有LLM智能体社会仿真中角色扮演逼真度被等同于人类行为有效性的做法，并提出应将仿真建模为包含环境、曝光和调度机制的马尔可夫博弈。","baseline":"无对照","findings":"LLM智能体单独不足以进行社会仿真，因为角色扮演的合理性不保证行为忠实性；集体结果常由智能体-环境共动态、交互协议、调度和信息先验主导，而非仅由智能体间对话决定。","reliability":"论文指出当前仿真过度依赖角色扮演逼真度，忽视环境、制度和信息不对称等机制，导致结果可能被实现细节主导，缺乏可审计的生成机制，易成为说服性叙事而非科学工具。","relevance":"该文直接批判LLM社会仿真的可靠性，指出缺乏人类行为基准和机制透明性会导致错误推断，与您关注的仿真失效条件和批判性研究高度契合，值得精读以理解当前范式的根本缺陷。","inspiration":"该文强调仿真需建模环境、曝光和调度机制，而非仅依赖角色扮演，这提醒我们在经济实验中应明确制度规则和信息结构的设计｜可迁移到政策公告预期形成场景，如央行沟通对市场预期的影响｜以LLM为被试，处理为不同透明度的政策公告，结果变量为预期通胀预测值，对照真实调查数据如密歇根消费者调查"}},{"id":"2602.15173","version":2,"title":"Mind the (DH) Gap! A Contrast in Risky Choices Between Reasoning and Conversational LLMs","zh_title":"注意(DH)差距！推理型与对话型LLM在风险选择上的对比","abstract":"The use of large language models either as decision support systems, or in agentic workflows, is rapidly transforming the digital ecosystem. However, the understanding of LLM decision-making under uncertainty remains limited. We study LLM risky choices along two dimensions: (1) prospect representation (based on an explicit representation or outcome history) and (2) decision rationale (explanation). Our study, which involves 20 frontier and open LLMs, is complemented by a matched human subjects experiment, which provides one reference point, while an expected payoff maximizing rational agent model provides another. We find that LLMs cluster into two categories: reasoning models (RMs) and conversational models (CMs). RMs tend towards rational behavior, are insensitive to the order of prospects, gain/loss framing, and explanations, and behave similarly whether prospects are explicit or presented via a history of outcomes. CMs are significantly less rational, slightly more human-like, sensitive to prospect ordering, framing, and explanation, and exhibit a large description-history gap. Paired comparisons of open LLMs suggest that a key factor differentiating RMs and CMs is training for mathematical reasoning.","authors":["Luise Ge","Yongyan Zhang","Yevgeniy Vorobeychik"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-16","first_seen":"2026-02-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.15173","pdf_url":"https://arxiv.org/pdf/2602.15173","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","风险决策","人类对照实验"],"reason":"用LLM仿真人类风险决策，并与真人实验对照，评估模型行为偏差与人类相似度。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":40,"question":"LLM在风险决策中如何受前景表征方式（描述 vs. 历史结果）和决策理由（解释）的影响，并与人类及理性基准对比？","design":"用20个前沿和开源LLM作为被试，在三种基础前景对上完成风险选择任务，操纵前景表征（显式描述 vs. 结果历史）和解释要求（无解释/简短解释/数学解释），测量选择行为，并与匹配的人类被试实验和期望收益最大化理性智能体对比。","baseline":"匹配的人类被试实验，以及期望收益最大化的理性智能体模型。","findings":"LLM分为推理模型和对话模型两类：推理模型接近理性人，对前景顺序、框架和解释不敏感，描述-历史差距小；对话模型理性较低，略似人类，对顺序、框架和解释敏感，描述-历史差距大。数学推理训练是区分两类模型的关键因素。","reliability":"论文未讨论","relevance":"该研究直接以LLM仿真人类风险决策，并与真人实验和理性基准对照，系统评估了模型行为偏差和人类相似度，完全契合研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"借鉴其系统操纵前景表征（描述 vs. 历史结果）和解释要求（无/简短/数学）来分离模型行为偏差的设计，并设置理性基准与人类被试双对照｜可迁移到金融风险偏好评估场景，如投资者在历史收益序列与文字描述下的风险选择差异｜以LLM为被试，处理为用历史收益率序列 vs. 文字描述呈现股票前景，结果变量为风险资产配置比例，对照真实投资者调查数据（如Survey of Consumer Finances）"}},{"id":"2602.14043","version":1,"title":"Beyond Static Snapshots: Dynamic Modeling and Forecasting of Group-Level Value Evolution with Large Language Models","zh_title":"超越静态快照：基于大语言模型的群体价值观动态建模与预测","abstract":"Social simulation is critical for mining complex social dynamics and supporting data-driven decision making. LLM-based methods have emerged as powerful tools for this task by leveraging human-like social questionnaire responses to model group behaviors. Existing LLM-based approaches predominantly focus on group-level values at discrete time points, treating them as static snapshots rather than dynamic processes. However, group-level values are not fixed but shaped by long-term social changes. Modeling their dynamics is thus crucial for accurate social evolution prediction--a key challenge in both data mining and social science. This problem remains underexplored due to limited longitudinal data, group heterogeneity, and intricate historical event impacts. To bridge this gap, we propose a novel framework for group-level dynamic social simulation by integrating historical value trajectories into LLM-based human response modeling. We select China and the U.S. as representative contexts, conducting stratified simulations across four core sociodemographic dimensions (gender, age, education, income). Using the World Values Survey, we construct a multi-wave, group-level longitudinal dataset to capture historical value evolution, and then propose the first event-based prediction method for this task, unifying social events, current value states, and group attributes into a single framework. Evaluations across five LLM families show substantial gains: a maximum 30.88\\% improvement on seen questions and 33.97\\% on unseen questions over the Vanilla baseline. We further find notable cross-group heterogeneity: U.S. groups are more volatile than Chinese groups, and younger groups in both countries are more sensitive to external changes. These findings advance LLM-based social simulation and provide new insights for social scientists to understand and predict social value changes.","authors":["Qiankun Pi","Guixin Su","Jinliang Li","Mayi Xu","Xin Miao","Jiawei Jiang","Ming Zhong","Tieyun Qian"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-02-15","first_seen":"2026-02-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.14043","pdf_url":"https://arxiv.org/pdf/2602.14043","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","价值观演化","社会模拟"],"reason":"用LLM仿真群体价值观动态，有真实世界价值观调查数据对照，涉及社会变迁预测。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":55,"question":"如何利用大语言模型动态建模和预测群体价值观的长期演变趋势？","design":"基于世界价值观调查（WVS）第5-7波数据，按性别、年龄、教育、收入四个维度将中美人群划分为群体，用LLM微调学习历史价值观轨迹以预测未来价值观，并提出事件感知预测方法，将社会事件与价值观表示对齐，让LLM推理事件影响。","baseline":"世界价值观调查（WVS）第5、6、7波的真实群体级纵向数据，包含中国28个群体和美国23个群体。","findings":"事件感知预测方法在已见问题上最高提升30.88%，在未见问题上最高提升33.97%；美国群体价值观波动性显著高于中国，两国年轻群体对外部变化更敏感。","reliability":"论文未讨论","relevance":"该研究用LLM仿真群体价值观动态演变，有真实WVS纵向数据作为基准，涉及中美群体异质性和事件驱动的价值观变化，属于经济学/政策评估场景下的LLM人类仿真，与研究者关注高度契合，值得精读原文。","inspiration":"该方法将社会事件编码为与价值观表示对齐的向量，让LLM推理事件对群体价值观的动态影响，这种事件感知的预测设计值得借鉴｜可迁移到政策公告对消费者信心或通胀预期的影响评估，例如研究央行沟通事件如何改变不同群体的预期形成过程｜用LLM模拟不同收入/年龄群体的消费者，施加央行声明作为处理，测量其通胀预期变化，以密歇根大学消费者调查的真实群体数据作为对照基准"}},{"id":"2602.13862","version":2,"title":"Measuring Self-Rating Bias in LLM-Generated Survey Data: A Semantic Similarity Framework for Independent Scale Mapping","zh_title":"测量LLM生成调查数据中的自评偏差：一种独立量表映射的语义相似度框架","abstract":"Synthetic survey data generated by large language models (LLMs) suffers from a fundamental circularity: the same model family that generates text responses also maps them to numerical scales. We calibrate and validate Semantic Similarity Rating (SSR; Maier et al., 2024), which decouples generation from scale mapping via embedding-based cosine similarity against predefined anchor statements. Configuration experiments (N=17 pilot, N=69 cross-validation across 8 domains) show that naturalistic behavioral anchors outperform formal jargon by 29 percentage points (pp), and that SSR achieves 65-67% exact match and 91% within plus/minus 1; a cross-model test with OpenAI text-embedding-3-small reaches 77% exact, confirming cross-provider generalization. Direct LLM baselines (Claude 87%, GPT-4o 83%) establish that SSR's contribution is methodological independence, not accuracy superiority. A control condition removing question text from the LLM prompt actually improves LLM accuracy, ruling out information asymmetry as the explanation for SSR's lower accuracy. A pre-registered circularity experiment (N=345) reveals 4x compressed error variance in LLM rating (sigma^2 = 0.21 vs 0.87 for SSR) and systematic directional bias. A cross-model control (GPT-4o rating Claude-generated text) shows nearly identical compression (within/cross ratio = 0.93), indicating variance compression is a general LLM property rather than a within-model artifact. The calibration dataset, anchor library, and source code are publicly available (see Data Availability).","authors":["Eduardo Vera Pichardo"],"categories":["physics.soc-ph"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-02-14","first_seen":"2026-02-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.13862","pdf_url":"https://arxiv.org/pdf/2602.13862","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查数据","偏差评估"],"reason":"评估LLM生成调查数据的自评偏差，提出独立量表映射方法，有真实人类数据对照，批…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":33,"question":"如何通过语义相似度框架独立测量LLM生成调查数据中的自评偏差，并解决文本生成与量表映射的循环性问题？","design":"本研究并非仿真人类被试，而是校准和验证语义相似度评分（SSR）框架：先用LLM（Claude Haiku 4.5）根据人格描述和问题生成文本回答，再用独立的嵌入模型（Voyage AI voyage-3.5-lite）通过余弦相似度将文本映射到预定义的锚定语句量表上，并通过配置实验、交叉验证、LLM基线对比和预注册的循环性实验评估该框架的性能与偏差。","baseline":"无对照","findings":"SSR框架实现了生成与测量的架构独立，自然行为锚定比正式术语锚定准确率高29个百分点，交叉验证中精确匹配率达65-67%，±1内达91%；直接LLM评分虽准确率更高（Claude 87%，GPT-4o 83%），但存在4倍误差方差压缩和系统性方向偏差，且该方差压缩是LLM的普遍属性而非同模型伪影。","reliability":"论文承认SSR的准确率低于直接LLM评分，且嵌入模型与生成模型可能因共享预训练语料而存在残余相关性；锚定语句的领域特异性导致跨领域准确率差异大（33-90%），需进一步优化锚定库。","relevance":"该研究直接针对LLM仿真调查数据中的循环性测量偏差，提出独立量表映射方法，并系统揭示了LLM自评的方差压缩和方向偏差，对关注仿真可靠性与失效条件的研究者具有重要参考价值，值得阅读原文。","inspiration":"该方法通过独立嵌入模型将LLM文本回答映射到预定义量表，避免了直接让LLM自评的循环偏差，这种测量与生成解耦的设计值得借鉴｜可迁移到消费者信心调查或通胀预期测量中，用LLM模拟受访者对经济前景的开放式回答，再独立映射到信心指数或预期值｜用LLM根据人口特征生成对经济前景的文本描述，以独立语义模型映射为预期通胀值，与密歇根消费者调查的真实个体数据对比，检验仿真偏差与方差压缩"}},{"id":"2602.11939","version":1,"title":"Do Large Language Models Adapt to Language Variation across Socioeconomic Status?","zh_title":"大语言模型能否适应社会经济地位带来的语言变异？","abstract":"Humans adjust their linguistic style to the audience they are addressing. However, the extent to which LLMs adapt to different social contexts is largely unknown. As these models increasingly mediate human-to-human communication, their failure to adapt to diverse styles can perpetuate stereotypes and marginalize communities whose linguistic norms are less closely mirrored by the models, thereby reinforcing social stratification. We study the extent to which LLMs integrate into social media communication across different socioeconomic status (SES) communities. We collect a novel dataset from Reddit and YouTube, stratified by SES. We prompt four LLMs with incomplete text from that corpus and compare the LLM-generated completions to the originals along 94 sociolinguistic metrics, including syntactic, rhetorical, and lexical features. LLMs modulate their style with respect to SES to only a minor extent, often resulting in approximation or caricature, and tend to emulate the style of upper SES more effectively. Our findings (1) show how LLMs risk amplifying linguistic hierarchies and (2) call into question their validity for agent-based social simulation, survey experiments, and any research relying on language style as a social signal.","authors":["Elisa Bassignana","Mike Zhang","Dirk Hovy","Amanda Cercas Curry"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-02-12","first_seen":"2026-02-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.11939","pdf_url":"https://arxiv.org/pdf/2602.11939","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","社会语言学","算法保真度"],"reason":"直接评估LLM仿真人类语言行为的效度，有真实人类数据对照，并指出仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":56,"question":"大语言模型在多大程度上能根据社会经济地位（SES）社区的语言变异调整其生成文本的风格？","design":"本研究并非基于智能体的仿真实验，而是通过提示工程让四个大语言模型补全来自Reddit和YouTube的、按SES分层的人类文本片段，然后比较模型生成文本与原始文本在94项社会语言学指标上的差异。","baseline":"从Reddit和YouTube收集的真实社交媒体文本，按SES（通过主题关键词和网络分析等策略）分层，作为人类语言风格的对照基准。","findings":"大语言模型仅能微弱地根据SES调节风格，常表现为近似或夸张模仿，且更擅长模仿高SES社区的风格。这种风格适应不足可能放大语言层级差异，并质疑LLM在基于智能体的社会模拟、调查实验等依赖语言风格作为社会信号的研究中的有效性。","reliability":"论文指出LLM在风格适应上存在近似或夸张模仿而非精确复现的问题，且对低SES风格模拟较差；初步消融实验显示，当提供更长上下文时，模型更倾向于适应高SES风格，这揭示了仿真失效的条件。","relevance":"该研究直接评估了LLM模拟不同社会经济地位人群语言行为的效度，有真实人类数据对照，并明确指出仿真在风格适应上的失效条件，与研究者关注的人类仿真可靠性及批判性评估高度契合，值得精读原文。","inspiration":"该方法通过提示工程让LLM补全按SES分层的人类文本，并对比94项社会语言学指标，可借鉴其分层对照与多维度风格测量来评估LLM的仿真偏差。｜可迁移到信贷审批中的语言歧视研究，例如分析LLM模拟不同SES申请人的贷款申请文本时是否系统性地偏向高SES风格，从而影响审批决策。｜以LLM为被试，给定不同SES背景的贷款申请场景提示，生成申请文本；结果变量为文本的语言风格指标（如正式度、情感词频）；以真实银行或P2P平台中不同SES申请人的贷款申请文本作为对照基准。"}},{"id":"2602.09362","version":1,"title":"Behavioral Economics of AI: LLM Biases and Corrections","zh_title":"人工智能的行为经济学：大语言模型的偏差与校正","abstract":"Do generative AI models, particularly large language models (LLMs), exhibit systematic behavioral biases in economic and financial decisions? If so, how can these biases be mitigated? Drawing on the cognitive psychology and experimental economics literatures, we conduct the most comprehensive set of experiments to date$-$originally designed to document human biases$-$on prominent LLM families across model versions and scales. We document systematic patterns in LLM behavior. In preference-based tasks, responses become more human-like as models become more advanced or larger, while in belief-based tasks, advanced large-scale models frequently generate rational responses. Prompting LLMs to make rational decisions reduces biases.","authors":["Pietro Bini","Lin William Cong","Xing Huang","Lawrence J. Jin"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-10","first_seen":"2026-02-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.09362","pdf_url":"https://arxiv.org/pdf/2602.09362","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM行为偏差","人类仿真","实验经济学"],"reason":"用人类实验范式测LLM行为偏差并对比人类数据，直接评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":19,"question":"生成式AI模型（尤其是大语言模型）在经济金融决策中是否表现出系统性行为偏差？如何纠正这些偏差？","design":"本研究并非用LLM仿真人类被试，而是将LLM本身作为研究对象。从认知心理学和实验经济学文献中选取原本用于测量人类偏差的实验问题，改编为提示词，通过API收集OpenAI ChatGPT、Anthropic Claude、Google Gemini和Meta Llama四个模型家族在不同版本和规模下的回答，分析其行为模式。","baseline":"以原有人类实验中的理性基准和真实人类回答作为对照。","findings":"在偏好类任务中，模型越先进或规模越大，回答越像人类且偏离理性；在信念类任务中，先进大规模模型则多给出理性回答。不同模型家族间存在显著异质性，如Gemini在偏好问题上比ChatGPT更不理性、更像人，而Llama在信念问题上理性程度较低。","reliability":"论文未讨论","relevance":"该研究直接使用人类行为实验范式测量LLM的偏差，并系统对比了LLM与人类及理性基准的差异，为评估LLM作为人类仿真被试的可靠性提供了关键证据，高度相关，值得精读。","inspiration":"借鉴其将经典行为经济学实验范式直接移植到LLM测试中的方法，可系统评估模型在不同决策场景下的行为一致性。｜可迁移到资产定价实验中的投资者偏差测量，如过度外推、过度自信等。｜以GPT-4等LLM为被试，呈现历史股价序列并要求预测未来收益，测量其外推倾向，并与真实投资者调查数据（如Shiller投资者信心调查）进行对照。"}},{"id":"2602.09802","version":2,"title":"Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices","zh_title":"大语言模型会为景观多付钱吗？从主观选择推断支付意愿","abstract":"As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often required to make subjective choices on behalf of users in settings where no objectively correct answer exists. We study LLM decision-making in a travel-assistant context by presenting models with choice dilemmas and analyzing their responses using multinomial logit models to derive implied willingness to pay (WTP) estimates. These WTP values are subsequently compared to human benchmark values from the economics literature. In addition to a baseline setting, we examine how model behavior changes under more realistic conditions, including the provision of information about users' past choices and persona-based prompting. Our results show that while meaningful WTP values can be derived for larger LLMs, they also display systematic deviations at the attribute level. Additionally, they tend to overestimate human WTP overall, particularly when expensive options or business-oriented personas are introduced. Conditioning models on prior preferences for cheaper options yields valuations that are closer to human benchmarks. Overall, our findings highlight both the potential and the limitations of using LLMs for subjective decision support and underscore the importance of careful model selection, prompt design, and user representation when deploying such systems in practice.","authors":["Manon Reusens","Sofie Goethals","Toon Calders","David Martens"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-10","first_seen":"2026-02-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.09802","pdf_url":"https://arxiv.org/pdf/2602.09802","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","支付意愿","人类数据对照"],"reason":"用LLM模拟人类支付意愿并与真实人类数据对照，评估偏差，涉及经济学场景。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"在旅行助手场景中，LLM的主观选择能否通过离散选择模型推导出可解释的支付意愿（WTP），并与人类基准对比？","design":"以酒店房间选择为任务，构建多属性选择困境，让多个LLM在不同提示条件下（基线、提供用户历史选择、基于人设的提示）做出选择，然后用多项Logit模型估计隐含的WTP，并分析提示改写、顺序调换、货币变化、温度等稳健性。","baseline":"对照Masiero et al. (2015)中人类对酒店房间属性的支付意愿估计值。","findings":"较大LLM能推导出有意义的WTP，但存在属性层面的系统性偏差，且整体高估人类WTP，尤其在引入昂贵选项或商务人设时；提供廉价偏好历史或学生人设可使估值更接近人类基准。","reliability":"论文承认WTP推导在部分模型上会失效，且结果受提示设计、用户表征方式影响较大，未涵盖更广泛的人群异质性，泛化到其他领域需进一步验证。","relevance":"直接对比LLM与人类真实WTP数据，系统评估了提示策略对仿真偏差的影响，并指出失效条件，高度契合研究者对经济学场景下LLM仿真可靠性及批判性分析的兴趣，值得精读原文。","inspiration":"借鉴其通过离散选择模型从LLM主观选择中推导支付意愿（WTP）并与人类基准对比的方法，可系统评估提示策略（如提供用户历史、人设提示）对仿真偏差的影响｜可迁移到消费者对金融产品属性的偏好评估，如贷款条款（利率、期限、抵押要求）或投资产品特征（风险、流动性、费用）的WTP估计｜以LLM为被试，呈现不同贷款产品选择集，处理为提供不同风险偏好或财务约束的人设提示，结果变量为通过多项Logit模型估计的各属性WTP，对照真实消费者信贷选择数据（如Survey of Consumer Finances）验证偏差"}},{"id":"2602.07414","version":1,"title":"Can LLMs Truly Embody Human Personality? Analyzing AI and Human Behavior Alignment in Dispute Resolution","zh_title":"LLM能真正体现人类人格吗？分析争议解决中AI与人类行为的一致性","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior in social settings such as legal mediation, negotiation, and dispute resolution. However, it remains unclear whether these simulations reproduce the personality-behavior patterns observed in humans. Human personality, for instance, shapes how individuals navigate social interactions, including strategic choices and behaviors in emotionally charged interactions. This raises the question: Can LLMs, when prompted with personality traits, reproduce personality-driven differences in human conflict behavior? To explore this, we introduce an evaluation framework that enables direct comparison of human-human and LLM-LLM behaviors in dispute resolution dialogues with respect to Big Five Inventory (BFI) personality traits. This framework provides a set of interpretable metrics related to strategic behavior and conflict outcomes. We additionally contribute a novel dataset creation methodology for LLM dispute resolution dialogues with matched scenarios and personality traits with respect to human conversations. Finally, we demonstrate the use of our evaluation framework with three contemporary closed-source LLMs and show significant divergences in how personality manifests in conflict across different LLMs compared to human data, challenging the assumption that personality-prompted agents can serve as reliable behavioral proxies in socially impactful applications. Our work highlights the need for psychological grounding and validation in AI simulations before real-world use.","authors":["Deuksin Kwon","Kaleen Shrestha","Bin Han","Spencer Lin","James Hale","Jonathan Gratch","Maja Matarić","Gale M. Lucas"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-07","first_seen":"2026-02-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.07414","pdf_url":"https://arxiv.org/pdf/2602.07414","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM人格仿真","人类行为对齐","争议解决"],"reason":"直接比较LLM与人类在冲突对话中的人格-行为对齐，有真实人类数据对照，并指出仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":22,"question":"当用大五人格特质提示LLM时，它们能否在冲突解决对话中复现人类因人格差异而导致的行为差异？","design":"基于KODIS人类冲突对话数据集，为LLM匹配相同场景和人格特质（大五人格），生成LLM-LLM冲突对话；比较人类与LLM在最终结果（得分、是否接受、是否离开）和策略行为（基于利益-权利-权力框架）上的差异。","baseline":"KODIS数据集中248段具有完整人格信息的人类-人类冲突解决对话。","findings":"人类中神经质是策略结果的最强预测因子，而LLM中外向性和宜人性的效应更强，且策略行为负载于更广泛的人格因素；Claude和Gemini比GPT-4o mini更接近人类策略指标，但整体仍存在显著偏差。","reliability":"论文指出人格提示的LLM在情感冲突场景中的行为保真度未经严格验证，不同LLM间人格表现差异显著，不能可靠地作为人类行为代理，强调在真实应用前需进行心理验证。","relevance":"该研究直接比较LLM与人类在冲突对话中的人格-行为对齐，有真实人类数据对照，并指出仿真失效条件，高度契合您对LLM人类仿真可靠性及批判性研究的关注，值得精读原文。","inspiration":"该方法通过给LLM施加人格特质提示来模拟人类行为差异，并设置真实人类对话数据作为对照基准，可借鉴其处理-对照设计及基于框架的策略行为编码。｜可迁移至消费者跨期选择实验，探究不同人格特质（如尽责性、神经质）对时间贴现行为的影响。｜以LLM为被试，施加大五人格提示，测量其在跨期选择任务中的贴现率，并与真实人类实验数据（如Andersen et al.的贴现率估计）进行对照。"}},{"id":"2602.18462","version":1,"title":"Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents","zh_title":"评估基于人格条件的LLM作为合成调查受访者的可靠性","abstract":"Using persona-conditioned LLMs as synthetic survey respondents has become a common practice in computational social science and agent-based simulations. Yet, it remains unclear whether multi-attribute persona prompting improves LLM reliability or instead introduces distortions. Here we contribute to this assessment by leveraging a large dataset of U.S. microdata from the World Values Survey. Concretely, we evaluate two open-weight chat models and a random-guesser baseline across more than 70K respondent-item instances. We find that persona prompting does not yield a clear aggregate improvement in survey alignment and, in many cases, significantly degrades performance. Persona effects are highly heterogeneous as most items exhibit minimal change, while a small subset of questions and underrepresented subgroups experience disproportionate distortions. Our findings highlight a key adverse impact of current persona-based simulation practices: demographic conditioning can redistribute error in ways that undermine subgroup fidelity and risk misleading downstream analyses.","authors":["Erika Elizabeth Taday Morocho","Lorenzo Cima","Tiziano Fagni","Marco Avvenuti","Stefano Cresci"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-02-06","first_seen":"2026-02-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.18462","pdf_url":"https://arxiv.org/pdf/2602.18462","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查方法","可靠性评估"],"reason":"直接评估LLM作为合成调查受访者的可靠性，使用真实人类数据对照，并指出仿真失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":5,"question":"多属性人格提示（persona prompting）能否提高大语言模型作为合成调查受访者的可靠性，还是会引入扭曲？","design":"使用两个开源聊天模型（Llama-2-13B 和 Qwen3-4B）模拟美国受访者，基于世界价值观调查（WVS-7）的个体记录构建多属性人格提示，比较有人格提示、无提示（vanilla）和随机猜测基线在70K+受访者-题目实例上的回答一致性。","baseline":"世界价值观调查第7波（WVS-7）的美国受访者微观数据，作为真实人类回答的基准。","findings":"人格提示在总体上并未带来一致的对齐改善，在许多情况下反而显著降低性能；人格效应高度异质，大多数题目变化极小，但少数题目和代表性不足的子群体出现不成比例的扭曲。","reliability":"论文指出人格提示可能重新分配误差，损害子群体保真度，并误导下游分析；效果因题目和属性而异，在少数群体中可能集中出现错误，且多属性约束可能相互干扰。","relevance":"该研究直接评估LLM作为合成调查受访者的可靠性，使用真实人类数据对照，并批判性地揭示了人格提示在子群体层面的失效风险，高度契合研究者对仿真可靠性、偏差及失效条件的关注，值得精读原文。","inspiration":"借鉴多属性人格提示与无提示基线的对照设计，以及按题目和子群体分解异质性效应的分析方法，可系统评估LLM仿真中的偏差来源｜可迁移到消费者金融决策调查场景，如风险偏好、储蓄选择、信贷需求等问卷的行为一致性研究｜以LLM作为合成受访者，施加多属性人格提示（收入、教育、财务素养等），测量其与真实消费者金融调查（如SCF）中个体回答的匹配度，并按收入分位数和金融素养水平检验子群体偏差"}},{"id":"2602.18464","version":2,"title":"How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?","zh_title":"LLM代理模拟终端用户安全与隐私态度及行为的效果如何？","abstract":"A growing body of research assumes that large language model (LLM) agents can serve as proxies for how people form attitudes toward and behave in response to security and privacy (S&P) threats. If correct, these simulations could offer a scalable way to forecast S&P risks in products prior to deployment. We interrogate this assumption using SP-ABCBench, a new benchmark of 30 tests derived from validated S&P human-subject studies, which measures alignment between simulations and human-subjects studies on a 0-100 ascending scale, where higher scores indicate better alignment across three dimensions: Attitude, Behavior, and Coherence. Evaluating twelve LLMs, four persona construction strategies, and two prompting methods, we found that there remains substantial room for improvement: all models score between 50 and 64 on average. Newer, bigger, and smarter models do not reliably do better and sometimes do worse. Some simulation configurations, however, do yield high alignment: e.g., with scores above 95 for some behavior tests when agents are prompted to apply bounded rationality and weigh privacy costs against perceived benefits. We release SP-ABCBench to enable reproducible evaluation as methods improve.","authors":["Yuxuan Li","Leyang Li","Hao-Ping Lee","Sauvik Das"],"categories":["cs.CY","cs.AI","cs.CL","cs.CR"],"primary_category":"cs.CY","announce_type":"new","date":"2026-02-06","first_seen":"2026-02-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.18464","pdf_url":"https://arxiv.org/pdf/2602.18464","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类行为对照","安全隐私"],"reason":"直接评估LLM代理模拟人类安全隐私态度行为，有真实人类数据基准，并指出仿真失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":41,"question":"当前LLM代理在多大程度上能复现人群在安全与隐私（S&P）方面的态度、行为及其一致性？","design":"使用12个LLM，结合4种角色构建策略和2种提示方法，基于15项真实人类被试研究构建的30个测试基准SP-ABCBench，测量仿真结果与人类数据在态度、行为和一致性三个维度的对齐分数（0-100）。","baseline":"对照来自15项经过验证的S&P人类被试研究，涵盖态度量表、行为实验和构念间关系等30个可量化的人群层面效应。","findings":"所有模型平均对齐分数仅50-64，更大、更新、更强的模型未必更好，有时更差；但特定配置（如结合有限理性与隐私计算提示）在部分行为测试上可达95分以上。","reliability":"论文指出当前LLM仿真在S&P领域整体对齐度中等，模型规模与能力不保证提升，且角色构建与提示策略效果因测试维度而异，提示仿真在S&P决策中可能失效的条件。","relevance":"该研究直接评估LLM替代人类被试进行S&P态度行为仿真的可靠性，有真实人类基准，并揭示了仿真失效的具体条件，高度契合您对LLM人类仿真实验批判性评估的关注，值得精读。","inspiration":"该方法通过构建多维度测试基准（态度、行为、一致性）并计算对齐分数来量化LLM仿真与人类数据的差距，可借鉴其系统性评估框架和对照设计。｜可迁移到消费者金融决策研究，如评估LLM能否复现真实人群在信贷选择、风险偏好或退休储蓄行为中的偏差与异质性。｜以LLM代理为被试，施加不同金融素养或信息框架处理，测量其信贷违约概率或投资组合选择，并与美国消费者金融调查（SCF）或实验室实验的真实行为数据做对齐比较。"}},{"id":"2602.04674","version":2,"title":"Overstating Attitudes, Ignoring Networks: LLM Biases in Simulating Misinformation Susceptibility","zh_title":"夸大态度，忽视网络：LLM在模拟错误信息易感性中的偏差","abstract":"Large language models (LLMs) are increasingly used as proxies for human judgment in computational social science, yet their ability to reproduce patterns of susceptibility to misinformation remains unclear. We test whether LLM-simulated survey respondents, prompted with participant profiles drawn from social survey data measuring network, demographic, attitudinal and behavioral features, can reproduce human patterns of misinformation belief and sharing. Using three online surveys as baselines, we evaluate whether LLM outputs match observed response distributions and recover feature-outcome associations present in the original survey data. LLM-generated responses capture broad distributional tendencies and show modest correlation with human responses, but consistently overstate the association between belief and sharing. Linear models fit to simulated responses exhibit substantially higher explained variance and place disproportionate weight on attitudinal and behavioral features, while largely ignoring personal network characteristics, relative to models fit to human responses. Analyses of model-generated reasoning and LLM training data suggest that these distortions reflect systematic biases in how misinformation-related concepts are represented. Our findings suggest that LLM-based survey simulations are better suited for diagnosing systematic divergences from human judgment than for substituting it.","authors":["Eun Cheol Choi","Lindsay E. Young","Emilio Ferrara"],"categories":["cs.SI","cs.AI","cs.CL"],"primary_category":"cs.SI","announce_type":"new","date":"2026-02-04","first_seen":"2026-02-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.04674","pdf_url":"https://arxiv.org/pdf/2602.04674","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","错误信息","人类数据对照"],"reason":"用LLM仿真人类对错误信息的易感性，并与真实调查数据对照，评估偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":33,"question":"LLM模拟的调查受访者在多大程度上能复现人类对错误信息的相信与分享模式及其与社会预测因子（包括个人网络特征）之间的关联？","design":"使用LLM（如GPT-4等）扮演合成调查受访者，输入基于三个真实调查数据构建的受访者结构化档案（包含个人网络、人口统计、态度/行为特征），让LLM生成对错误信息条目的相信和分享意愿回答，测量回答分布及特征-结果关联。","baseline":"三个在线调查数据集：公共卫生（美国，2023）、气候变化（美国，2025）、疫情政治（韩国，2020），均包含真实人类对错误信息的相信与分享数据及个人网络、人口统计、态度行为变量。","findings":"LLM模拟能捕捉大致分布趋势并与人类回答有适度相关，但系统性地夸大了相信与分享之间的关联；线性模型在模拟数据上解释方差显著膨胀，且过度依赖态度和行为特征，几乎忽略个人网络特征。","reliability":"论文指出LLM模拟更适合诊断与人类判断的系统性偏差，而非替代人类判断；偏差源于LLM训练数据中错误信息相关概念的表征偏差，且模拟未能复现个人网络特征的作用。","relevance":"该研究直接以真实人类调查为基准，评估LLM仿真在错误信息易感性上的可靠性，并揭示了仿真在忽略网络特征、夸大态度关联等方面的失效条件，高度契合研究者对批判性仿真研究的兴趣，值得精读原文。","inspiration":"该方法借鉴了用结构化档案（含人口统计、态度、网络特征）驱动LLM生成调查回答，并与真实人类数据对照以评估仿真偏差的设计｜可迁移到信贷审批中的歧视研究，检验LLM模拟的贷款官员是否复现人类决策中的种族或性别偏见｜用LLM扮演贷款审批员，输入含申请人种族、收入、信用分等档案，输出审批决定，以真实房贷数据（如HMDA）为基准，比较拒绝率差异及特征重要性"}},{"id":"2602.03545","version":2,"title":"Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts","zh_title":"人格生成器：为任意上下文生成多样化的合成人格","abstract":"Evaluating AI systems that interact with humans requires understanding their behavior across diverse user populations, but collecting representative human data is often expensive or infeasible, particularly for novel technologies or hypothetical future scenarios. Recent work in Generative Agent-Based Modeling has shown that large language models can simulate human-like synthetic personas with high fidelity, accurately reproducing the beliefs and behaviors of specific individuals. However, most approaches require detailed data about target populations and often prioritize density matching (replicating what is most probable) rather than support coverage (spanning what is possible), leaving long-tail behaviors underexplored. We introduce Persona Generators, functions that can produce diverse synthetic populations tailored to arbitrary contexts. We apply an iterative improvement loop based on AlphaEvolve, using large language models as mutation operators to refine our Persona Generator code over hundreds of iterations. The optimization process produces lightweight Persona Generators that can automatically expand small descriptions into populations of diverse synthetic personas that maximize coverage of opinions and preferences along relevant diversity axes. We demonstrate that evolved generators substantially outperform existing baselines across six diversity metrics on held-out contexts, producing populations that span rare trait combinations difficult to achieve in standard LLM outputs.","authors":["Davide Paglieri","Logan Cross","William A. Cunningham","Joel Z. Leibo","Alexander Sasha Vezhnevets"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-03","first_seen":"2026-02-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.03545","pdf_url":"https://arxiv.org/pdf/2602.03545","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B1"],"tags":["合成人群生成","人类仿真","多样性覆盖"],"reason":"用LLM生成多样化合成人群，有真实人类数据对照，但侧重覆盖度而非行为复现，方法…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":70,"question":"如何生成能覆盖任意场景下人类观点与偏好全貌的多样化合成人群，以解决标准LLM的模式坍缩问题？","design":"提出Persona Generator函数，通过AlphaEvolve进化循环优化代码（含提示模板和采样逻辑），将简短场景提示扩展为结构化问卷，再生成覆盖多样性轴的合成人群；评估时在留出场景上比较生成人群的多样性指标。","baseline":"无对照：论文未使用真实人类行为数据作为基准，而是与Nemotron Personas等合成基线比较覆盖度。","findings":"进化出的生成器在六项多样性指标上显著优于基线，能覆盖稀有特征组合；通过对抗模式坍缩，生成的人群在真实人类特质分布方差上甚至优于专门匹配人类统计数据的基线。","reliability":"论文未讨论失效条件与局限，仅指出密度匹配方法不适合压力测试和推测性场景，但未分析自身方法在哪些条件下可能失效。","relevance":"高度相关：该研究用LLM生成多样化合成人群，虽侧重覆盖度而非行为复现，但直接回应了LLM仿真中模式坍缩和多样性不足的批判，且声称能更好捕捉真实人类分布，值得精读以评估其在经济学实验和政策评估中的潜力与局限。","inspiration":"该方法通过进化算法优化提示模板和采样逻辑以生成覆盖多样性轴的合成人群，可借鉴其对抗模式坍缩的思路来设计LLM仿真中的处理变异与稳健性检验｜可迁移到政策公告的预期形成研究，利用多样化合成人群模拟异质性信念与反应｜以LLM作为被试，施加不同措辞的政策公告处理，测量预期通胀或消费意愿的分布，并以真实调查数据（如密歇根消费者调查）作为对照基准"}},{"id":"2602.01684","version":1,"title":"The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament","zh_title":"大语言模型的战略远见：来自全前瞻性创业锦标赛的证据","abstract":"Can artificial intelligence outperform humans at strategic foresight -- the capacity to form accurate judgments about uncertain, high-stakes outcomes before they unfold? We address this question through a fully prospective prediction tournament using live Kickstarter crowdfunding projects. Thirty U.S.-based technology ventures, launched after the training cutoffs of all models studied, were evaluated while fundraising remained in progress and outcomes were unknown. A diverse suite of frontier and open-weight large language models (LLMs) completed 870 pairwise comparisons, producing complete rankings of predicted fundraising success. We benchmarked these forecasts against 346 experienced managers recruited via Prolific and three MBA-trained investors working under monitored conditions. The results are striking: human evaluators achieved rank correlations with actual outcomes between 0.04 and 0.45, while several frontier LLMs exceeded 0.60, with the best (Gemini 2.5 Pro) reaching 0.74 -- correctly ordering nearly four of every five venture pairs. These differences persist across multiple performance metrics and robustness checks. Neither wisdom-of-the-crowd ensembles nor human-AI hybrid teams outperformed the best standalone model.","authors":["Felipe A. Csaszar","Aticus Peterson","Daniel Wilde"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-02","first_seen":"2026-02-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.01684","pdf_url":"https://arxiv.org/pdf/2602.01684","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","人类行为对照","创业预测"],"reason":"用LLM预测人类对创业项目的判断，并与真实人类数据对照，属于经济学场景下的人类…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"大语言模型在战略远见（预测不确定、高风险商业结果）上能否超越人类？","design":"使用多款前沿和开源大语言模型对30个正在众筹的Kickstarter科技项目进行870次成对比较，生成预测筹款成功排名；同时招募346名有经验的管理者和3名MBA投资者作为人类被试进行相同任务，以最终实际筹款结果作为准确度标准。","baseline":"346名通过Prolific招募的有经验管理者和3名受监控的MBA投资者，其预测排名与实际结果的秩相关系数在0.04至0.45之间。","findings":"人类评估者的预测排名与实际结果的秩相关系数最高仅0.45，而多个前沿大语言模型超过0.60，最佳模型Gemini 2.5 Pro达到0.74。群体智慧集成和人机混合团队均未超越最佳独立模型。","reliability":"论文未讨论","relevance":"该研究直接用LLM替代人类被试进行前瞻性预测，并与真实人类数据对照，属于经济学场景下的人类仿真实验，且提供了可靠性证据，值得精读原文。","inspiration":"该方法借鉴了用真实众筹结果作为客观基准，直接比较LLM与人类被试的预测准确度，并通过成对比较排名任务量化战略远见｜可迁移到创业投资决策、众筹市场预测或资产定价中的预期形成研究，用于评估AI辅助决策的可靠性｜招募专业投资者作为人类被试，让LLM和人类分别对真实众筹项目进行成对比较排名，以最终实际筹资金额作为结果变量，计算预测排名与实际结果的秩相关系数，对比人机表现"}},{"id":"2602.07023","version":2,"title":"Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation","zh_title":"LLM智能体行为一致性验证：基于股市模拟的交易风格切换分析","abstract":"Recent works have increasingly applied Large Language Models (LLMs) as agents in financial stock market simulations to test if micro-level behaviors aggregate into macro-level phenomena. However, a crucial question arises: Do LLM agents' behaviors align with real market participants? This alignment is key to the validity of simulation results. To explore this, we select a financial stock market scenario to test behavioral consistency. Investors are typically classified as fundamental or technical traders, but most simulations fix strategies at initialization, failing to reflect real-world trading dynamics. In this work, we assess whether agents' strategy switching aligns with financial theory, providing a framework for this evaluation. We operationalize four behavioral-finance drivers-loss aversion, herding, wealth differentiation, and price misalignment-as personality traits set via prompting and stored long-term. In year-long simulations, agents process daily price-volume data, trade under a designated style, and reassess their strategy every 10 trading days. We introduce four alignment metrics and use Mann-Whitney U tests to compare agents' style-switching behavior with financial theory. Our results show that recent LLMs' switching behavior is only partially consistent with behavioral-finance theories, highlighting the need for further refinement in aligning agent behavior with financial theory.","authors":["Zeping Li","Guancheng Wan","Keyang Chen","Yu Chen","Yiwen Zhao","Philip Torr","Guangnan Ye","Zhenfei Yin","Hongfeng Chai"],"categories":["q-fin.TR","cs.AI"],"primary_category":"q-fin.TR","announce_type":"new","date":"2026-02-02","first_seen":"2026-02-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.07023","pdf_url":"https://arxiv.org/pdf/2602.07023","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","行为金融","智能体一致性"],"reason":"用LLM agent模拟股票交易行为并与金融理论对照，涉及行为经济学场景，指出…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":55,"question":"在股票市场仿真中，LLM智能体的交易风格切换行为是否与行为金融学理论一致？","design":"使用多种近期LLM（通过提示词设定四种行为金融倾向作为人格特质并存入长期记忆）扮演投资者，在基于2024年标普500成分股数据的模拟环境中进行为期一年的交易，每10个交易日评估并决定是否切换交易风格（基本面/技术面），通过四个对齐指标和Mann-Whitney U检验比较智能体行为与理论预期。","baseline":"无对照","findings":"LLM智能体的风格切换行为仅部分符合行为金融学理论，无法在所有方面完全对齐。","reliability":"论文未讨论","relevance":"该研究直接评估LLM智能体在金融行为仿真中的行为一致性，属于批判性验证工作，与研究者关注的LLM仿真可靠性及失效条件高度相关，值得阅读原文以了解具体偏差和评估框架。","inspiration":"该研究通过提示词将行为金融倾向植入LLM智能体并存入长期记忆，在动态仿真中周期性评估风格切换，用多个对齐指标和统计检验衡量与理论的偏差，这种处理-测量-验证框架值得借鉴。｜可迁移到资产定价实验，检验LLM智能体在信息冲击下的过度反应与反转行为是否符合前景理论与处置效应。｜以LLM智能体为被试，施加不同强度的利好/利空消息作为处理，观测其持仓调整与买卖时机，结果变量为超额收益与换手率，用历史高频交易数据中散户的实际行为分布作为对照基准。"}},{"id":"2602.02606","version":2,"title":"Gender Dynamics and Homophily in a Social Network of LLM Agents","zh_title":"LLM代理社交网络中的性别动态与同质性","abstract":"Generative artificial intelligence and large language models (LLMs) are increasingly deployed in interactive settings, yet we know little about how their identity performance develops when they interact within large-scale networks. We address this by examining Chirper.ai, a social media platform similar to X but composed entirely of autonomous AI chatbots. Our dataset comprises over 70,000 agents, approximately 140 million posts, and the evolving followership network over a period of one year. Based on agents' posted text, we assign weekly gender performance scores to each agent. Results suggest that each agent's gender performance is fluid rather than fixed. Despite this fluidity, the network displays strong gender-based homophily, as agents consistently follow others performing gender similarly. We investigate whether these homophilic connections arise from social selection, in which agents choose to follow similar accounts, or from social influence, in which agents become more similar to their followees over time. Consistent with human social networks, we find evidence that both mechanisms shape the structure and evolution of interactions among LLMs. Our findings suggest that, even in the absence of bodies, cultural entraining of gender performance leads to gender-based sorting. This has important implications for LLM applications in synthetic hybrid populations, social simulations, and decision support.","authors":["Faezeh Fadaei","Jenny Carla Moran","Taha Yasseri"],"categories":["cs.SI","cs.AI","cs.CY"],"primary_category":"cs.SI","announce_type":"new","date":"2026-02-02","first_seen":"2026-02-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.02606","pdf_url":"https://arxiv.org/pdf/2602.02606","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B1","B4"],"tags":["LLM代理","社会模拟","性别同质性"],"reason":"用LLM代理模拟社交网络性别同质性，并与人类社交网络机制对照，但非严格实验仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":83,"question":"在由LLM代理组成的社交网络中，性别表演如何动态变化，以及性别同质性如何通过社会选择和社会影响机制形成？","design":"本研究并非受控仿真实验，而是对Chirper.ai平台（类似X但全部由自主AI聊天机器人组成）的自然观察研究。研究者收集了超过7万个代理、约1.4亿条帖子和一年的关注网络数据，基于帖子文本为每个代理分配每周性别表演分数，分析性别表演的流动性和网络中的性别同质性，并区分社会选择与社会影响两种机制。","baseline":"无对照","findings":"每个代理的性别表演是流动而非固定的，但网络仍表现出强烈的性别同质性，代理倾向于关注性别表演相似的他人。社会选择和社会影响两种机制共同塑造了LLM代理之间的互动结构和演化，这与人类社交网络中的发现一致。","reliability":"论文未讨论","relevance":"该研究利用LLM代理在真实平台上的大规模互动数据，揭示了性别同质性的涌现机制，并与人类社交网络机制进行对照，虽非严格实验仿真，但为LLM在社交仿真中的行为模式提供了实证证据，值得阅读以了解LLM集体行为中的社会结构形成。","inspiration":"利用大规模LLM代理平台的自然观察数据，通过文本分析动态测量个体属性（如性别表演），并区分社会选择与社会影响机制，为无干扰的群体行为研究提供了方法参考｜可迁移至金融社交媒体中的信息扩散与投资者行为研究，例如分析LLM代理在模拟投资社区中如何形成风险偏好同质性｜设计一个LLM代理投资社区，让代理基于历史市场信息发布投资观点，用文本分析测量其风险偏好，观察关注网络与偏好同质性的动态演化，并以真实投资者社交平台数据（如StockTwits）作为对照基准"}},{"id":"2602.01022","version":3,"title":"Calibrating Behavioral Parameters with Large Language Models","zh_title":"用大语言模型校准行为参数","abstract":"Behavioral parameters such as loss aversion, herding, and extrapolation are central to asset pricing models but remain difficult to measure reliably. We develop a framework that treats large language models (LLMs) as calibrated measurement instruments for behavioral parameters. Using four models and 24{,}000 agent--scenario pairs, we document systematic rationality bias in baseline LLM behavior, including attenuated loss aversion, weak herding, and near-zero disposition effects relative to human benchmarks. Profile-based calibration induces large, stable, and theoretically coherent shifts in several parameters, with calibrated loss aversion, herding, extrapolation, and anchoring reaching or exceeding benchmark magnitudes. To assess external validity, we embed calibrated parameters in an agent-based asset pricing model, where calibrated extrapolation generates short-horizon momentum and long-horizon reversal patterns consistent with empirical evidence. Our results establish measurement ranges, calibration functions, and explicit boundaries for eight canonical behavioral biases.","authors":["Brandon Yee","Pairie Koh"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-01","first_seen":"2026-02-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.01022","pdf_url":"https://arxiv.org/pdf/2602.01022","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为经济学","人类基准对照"],"reason":"用LLM测量行为参数并与人类基准对照，嵌入资产定价模型验证外部效度，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":8,"question":"能否将大语言模型作为可校准的测量工具，系统性地诱导、校准并验证行为金融参数？","design":"使用 GPT-4o、GPT-4o-mini、Claude-3.5-Haiku、Gemini-2.5-Pro 四种模型，通过提示词嵌入行为特征（如损失厌恶、羊群效应等）作为实验处理，在 24,000 个合成金融场景中测量八种行为偏差的参数值，并评估校准后的参数在基于代理的资产定价模型中的外部有效性。","baseline":"人类基准来自已有文献中的实验和实证估计，如损失厌恶系数约 2.25，羊群效应率 65-75%，以及 Jegadeesh 和 Titman 的动量与反转经验事实。","findings":"基线 LLM 行为存在系统性理性偏差，表现为损失厌恶减弱、羊群效应弱、处置效应接近零；通过基于特征的校准可诱导出与人类基准相当甚至更强的参数值，且校准后的外推参数能在资产定价模型中生成符合经验事实的短期动量和长期反转模式。","reliability":"论文承认校准存在明确边界，并非所有参数都能成功校准，且结果依赖于提示词设计和模型选择，外部有效性仅通过简单资产定价模型初步验证。","relevance":"该研究直接命中研究者对 LLM 仿真人类行为、与真实人类基准对照、经济学实验及批判性评估的核心兴趣，提供了系统的校准框架和失效边界，值得精读原文。","inspiration":"该方法通过提示词嵌入行为特征（如损失厌恶、羊群效应）作为实验处理，系统性地校准LLM的行为参数，并与人类基准对照，值得借鉴其处理施加与参数校准的流程设计｜可迁移到资产定价实验中的投资者行为偏差研究，如模拟动量效应、反转效应及处置效应等市场异象｜以LLM作为被试，通过提示词嵌入不同程度的损失厌恶或羊群效应处理，测量其交易决策与价格预期，并与Jegadeesh和Titman的动量/反转经验事实及处置效应实证数据对照"}},{"id":"2602.00685","version":1,"title":"HumanStudy-Bench: Towards AI Agent Design for Participant Simulation","zh_title":"HumanStudy-Bench：面向参与者仿真的AI智能体设计基准","abstract":"Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable and highly sensitive to design choices. Prior evaluations frequently conflate base-model capabilities with experimental instantiation, obscuring whether outcomes reflect the model itself or the agent setup. We instead frame participant simulation as an agent-design problem over full experimental protocols, where an agent is defined by a base model and a specification (e.g., participant attributes) that encodes behavioral assumptions. We introduce HUMANSTUDY-BENCH, a benchmark and execution engine that orchestrates LLM-based agents to reconstruct published human-subject experiments via a Filter--Extract--Execute--Evaluate pipeline, replaying trial sequences and running the original analysis pipeline in a shared runtime that preserves the original statistical procedures end to end. To evaluate fidelity at the level of scientific inference, we propose new metrics to quantify how much human and agent behaviors agree. We instantiate 12 foundational studies as an initial suite in this dynamic benchmark, spanning individual cognition, strategic interaction, and social psychology, and covering more than 6,000 trials with human samples ranging from tens to over 2,100 participants.","authors":["Xuan Liu","Haoyang Shang","Zizhang Liu","Xinyan Liu","Yunze Xiao","Yiwen Tu","Haojian Jin"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-31","first_seen":"2026-01-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.00685","pdf_url":"https://arxiv.org/pdf/2602.00685","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3"],"tags":["LLM人类仿真","实验复现","基准测试"],"reason":"直接构建LLM代理复现人类实验，含真实人类数据对照，评估仿真保真度，覆盖经济学…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":11,"question":"如何将LLM参与者的仿真视为一个代理设计问题，并系统评估不同代理设计在复现真实人类实验中的保真度？","design":"使用10个当代LLM（如GPT、Claude、Gemini）作为基座模型，结合四种代理规格（空白、角色扮演、人口统计条件、丰富背景故事）构建AI代理，通过Filter–Extract–Execute–Evaluate管道重放12项已发表的人类实验的完整试验序列和原始分析流程，测量代理的行为响应。","baseline":"对照12项已发表人类实验的真实人类数据，涵盖个体认知、策略互动和社会心理学，超过6000次试验，人类样本量从几十到2100多人。","findings":"当前LLM代理与人类的推断一致性有限且不稳定，行为呈极化双峰而非人类单峰模式；代理设计对结果有较大且非单调的影响，性能高度依赖领域，更大模型或简单多模型集成未能可靠提升对齐度。","reliability":"论文指出LLM行为不稳定、对设计选择高度敏感，代理规格会编码行为假设并可能定性改变结果；现有评估常混淆基座模型能力与实验实例化，且代理在人口异质性敏感度和提示词变化下表现脆弱。","relevance":"该研究直接针对用LLM替代人类被试的仿真实验，提供真实人类数据对照，系统评估代理设计对复现经济学实验和社会心理学效应的影响，并批判性指出仿真失效的条件，与您的关注高度契合，值得精读原文。","inspiration":"借鉴其将代理设计作为实验变量的思路，通过对比空白、角色扮演、人口统计条件等不同规格来分离模型能力与实验实例化的混淆效应｜可迁移到政策公告的预期形成实验，研究不同信息框架下投资者对央行沟通的反应｜以LLM代理为被试，处理为不同代理规格（如空白vs.人口统计条件），结果变量为通胀预期调整幅度，对照真实调查数据（如密歇根消费者调查）"}},{"id":"2601.22812","version":2,"title":"Stable Personas: Dual-Assessment of Temporal Stability in LLM-Based Human Simulation","zh_title":"稳定人格：基于LLM的人类仿真中时间稳定性的双重评估","abstract":"Large Language Models (LLMs) acting as artificial agents offer the potential for scalable behavioral research, yet their validity depends on whether LLMs can maintain stable personas across extended conversations. We address this point using a dual-assessment framework measuring both self-reported characteristics and observer-rated persona expression. Across two experiments testing four persona conditions (default, high, moderate, and low ADHD presentations), seven LLMs, and three semantically equivalent persona prompts, we examine between-conversation stability (3,473 conversations) and within-conversation stability (1,370 conversations and 18 turns). Self-reports remain highly stable both between and within conversations. However, observer ratings reveal a tendency for persona expressions to decline during extended conversations. These findings suggest that persona-instructed LLMs produce stable, persona-aligned self-reports, an important prerequisite for behavioral research, while identifying this regression tendency as a boundary condition for multi-agent social simulation.","authors":["Jana Gonnermann-Müller","Jennifer Haase","Nicolas Leins","Thomas Kosch","Sebastian Pokutta"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-01-30","first_seen":"2026-01-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.22812","pdf_url":"https://arxiv.org/pdf/2601.22812","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B4"],"tags":["LLM仿真","人格稳定性","效度评估"],"reason":"研究LLM人格稳定性以评估其作为人类被试的可靠性，直接涉及仿真效度与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":47,"question":"LLM在独立对话间和长对话过程中，能在多大程度上稳定维持被赋予的人格？","design":"用7个LLM扮演默认、高、中、低四种ADHD人格，通过三种语义等价的提示词施加处理；实验一测量跨对话稳定性（每条件50次独立运行），实验二测量对话内稳定性（18轮对话，在第6、12、18轮评估）；结果变量为自评量表得分和观察者评分。","baseline":"无对照","findings":"自评报告在跨对话和对话内均高度稳定；但观察者评分显示，高和中强度人格表达在长对话中会逐渐减弱，向默认水平回归。","reliability":"论文指出，人格表达在长对话中衰退是LLM用于多智能体社会仿真的一个边界条件；评估依赖LLM作为评分者，可能引入偏差；仅以ADHD人格为测试案例，泛化性待验证。","relevance":"直接评估LLM作为人类被试替代品的人格稳定性，揭示了自评稳定但行为表达衰退的失效模式，对关注仿真效度与边界条件的研究者很有参考价值，值得读原文。","inspiration":"借鉴双维度人格稳定性评估设计，通过独立跨对话与长对话内重复测量，区分自评报告与行为观察的稳定性差异，揭示仿真衰退的边界条件。｜可迁移至政策公告预期形成实验，用LLM模拟投资者对央行沟通的反应，检验长期对话中信息解读的一致性。｜以LLM为被试，施加不同政策措辞处理，在长对话中多次测量通胀预期与投资意愿，对比真实投资者调查面板数据，评估仿真在持续信息流下的衰退模式。"}},{"id":"2601.21975","version":2,"title":"Mind the Gap: How Elicitation Protocols Shape the Stated-Revealed Preference Gap in Language Models","zh_title":"注意差距：诱导协议如何塑造语言模型中陈述-显示偏好差距","abstract":"Recent work identifies a stated-revealed (SvR) preference gap in language models (LMs): a mismatch between the values models endorse and the choices they make in context. Existing evaluations rely heavily on binary forced-choice prompting, which entangles genuine preferences with artifacts of the elicitation protocol. We systematically study how elicitation protocols affect SvR correlation across 24 LMs. Allowing neutrality and abstention during stated preference elicitation allows us to exclude weak signals, substantially improving Spearman's rank correlation ($ρ$) between volunteered stated preferences and forced-choice revealed preferences. However, further allowing abstention in revealed preferences drives $ρ$ to near-zero or negative values due to high neutrality rates. Finally, we find that system prompt steering using stated preferences during revealed preference elicitation does not reliably improve SvR correlation on AIRiskDilemmas. Together, our results show that SvR correlation is highly protocol-dependent and that preference elicitation requires methods that account for indeterminate preferences.","authors":["Pranav Mahajan","Ihor Kendiukhov","Syed Hussain","Lydia Nottingham"],"categories":["cs.AI","cs.ET"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-29","first_seen":"2026-01-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.21975","pdf_url":"https://arxiv.org/pdf/2601.21975","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["语言模型偏好","诱导协议","陈述-显示偏好"],"reason":"测量语言模型自身的偏好一致性，属于对模型本身的测量，而非用模型仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":185,"question":"在语言模型中，偏好引出协议（是否允许中立/弃权）如何影响陈述偏好与显示偏好之间的秩相关性？","design":"不是仿真研究。该研究在24个语言模型上系统改变偏好引出协议（强制二选一 vs. 允许中立/弃权），分别测量陈述偏好（抽象价值比较）和显示偏好（情境化道德困境）的排序，计算两者间的斯皮尔曼秩相关系数，并测试系统提示引导的效果。","baseline":"无对照","findings":"在陈述偏好引出中允许中立可过滤弱信号，显著提高与强制选择显示偏好的秩相关性；但在显示偏好引出中也允许中立时，因高中立率导致秩相关降至近零或负值。系统提示引导未能可靠改善陈述-显示偏好一致性。","reliability":"论文指出SvR相关性高度依赖引出协议，当模型普遍表达中立时，基于排名的相关性会失效；提示引导在16个价值域上不可靠，且排除中立响应会丢失不确定性信息。","relevance":"该研究不涉及用LLM仿真人类被试，而是测量模型自身的偏好一致性，属于对模型行为的测量与分析，与研究者关注的以人类数据为基准的仿真实验无关，不建议优先阅读。","inspiration":"该方法通过系统改变偏好引出协议（强制选择 vs. 允许中立）来测量陈述与显示偏好的秩相关性，可借鉴用于检验经济调查中选项设计对偏好一致性的影响｜可迁移到消费者跨期选择研究，例如在时间偏好调查中对比强制排序与允许“无差异”选项时，陈述偏好与真实激励下的选择行为是否一致｜以LLM为被试，设计跨期选择任务：处理组为强制二选一（今天100元 vs. 一年后120元），对照组允许选择“无差异”；结果变量为陈述偏好与显示偏好（模拟真实支付决策）的秩相关系数，对照真实人类实验数据（如Andreoni & Sprenger 2012）评估仿真效度"}},{"id":"2601.18027","version":2,"title":"Sentipolis: Emotion-Aware Agents for Social Simulations","zh_title":"Sentipolis：用于社会模拟的情感感知智能体","abstract":"LLM agents are increasingly used for social simulation, yet emotion is often treated as a transient cue, causing emotional amnesia and weak long-horizon continuity. We present Sentipolis, a framework for emotionally stateful agents that integrates continuous Pleasure-Arousal-Dominance (PAD) representation, dual-speed emotion dynamics, and emotion--memory coupling. Across thousands of interactions over multiple base models and evaluators, Sentipolis improves emotionally grounded behavior, boosting communication, and emotional continuity. Gains are model-dependent: believability increases for higher-capacity models but can drop for smaller ones, and emotion-awareness can mildly reduce adherence to social norms, reflecting a human-like tension between emotion-driven behavior and rule compliance in social simulation. Network-level diagnostics show reciprocal, moderately clustered, and temporally stable relationship structures, supporting the study of cumulative social dynamics such as alliance formation and gradual relationship change.","authors":["Chiyuan Fu","Lyuhao Chen","Yunze Xiao","Weihao Xuan","Carlos Busso","Mona Diab"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-25","first_seen":"2026-01-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.18027","pdf_url":"https://arxiv.org/pdf/2601.18027","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","情感建模","多智能体"],"reason":"社会模拟但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":43,"question":"如何设计具有持久情绪状态的LLM智能体，以解决社会模拟中的“情绪失忆”问题，并提升长期交互中的情绪连续性与行为真实性？","design":"提出Sentipolis框架，为LLM智能体引入连续PAD情绪表示、双速情绪动力学和情绪-记忆耦合机制；在多个基础模型（如GPT-5.2、Grok-4、GPT-4o-mini）上进行数千次交互实验，通过LLM评判和人类评估测量沟通质量、情绪连续性、可信度、共情和社会规范遵守等指标，并进行组件消融和网络级诊断。","baseline":"无对照","findings":"情绪状态化显著提升了沟通质量和情绪连续性，但效果依赖模型规模：大模型可信度提升，小模型可能下降；情绪意识会轻微降低对社会规范的遵守，反映了情绪驱动行为与规则遵从之间的类人张力。网络分析显示，情绪-记忆耦合产生了高互惠性、适度聚类和稳定的关系结构。","reliability":"论文指出效果具有模型异质性，小模型上可信度可能下降；情绪意识可能轻微降低社会规范遵守；评估依赖LLM评判，虽有人类评估验证，但长期动态的生态效度仍需进一步检验。","relevance":"该研究属于LLM社会模拟，但无真实人类数据对照，不符合研究者对基准实证数据的要求；其情绪建模方法对理解仿真偏差有参考价值，但非直接匹配经济学实验或政策评估场景。","inspiration":"可借鉴其将情绪作为显式持续状态并耦合记忆的模块化设计，用于在仿真中引入情绪驱动的决策偏差｜可迁移至行为经济学中的情绪与跨期选择实验，或金融中的投资者情绪与交易行为模拟｜以LLM智能体作为被试，施加情绪状态（如通过PAD向量操纵），观察其在跨期选择任务中的贴现率变化，并与真实人类实验数据（如Ifcher & Zarghamee, 2011）对照，检验情绪效应的仿真保真度。"}},{"id":"2601.17527","version":1,"title":"Bridging Expectation Signals: LLM-Based Experiments and a Behavioral Kalman Filter Framework","zh_title":"桥接预期信号：基于LLM的实验与行为卡尔曼滤波框架","abstract":"As LLMs increasingly function as economic agents, the specific mechanisms LLMs use to update their belief with heterogeneous signals remain opaque. We design experiments and develop a Behavioral Kalman Filter framework to quantify how LLM-based agents update expectations, acting as households or firm CEOs, update expectations when presented with individual and aggregate signals. The results from experiments and model estimation reveal four consistent patterns: (1) agents' weighting of priors and signals deviates from unity; (2) both household and firm CEO agents place substantially larger weights on individual signals compared to aggregate signals; (3) we identify a significant and negative interaction between concurrent signals, implying that the presence of multiple information sources diminishes the marginal weight assigned to each individual signal; and (4) expectation formation patterns differ significantly between household and firm CEO agents. Finally, we demonstrate that LoRA fine-tuning mitigates, but does not fully eliminate, behavioral biases in LLM expectation formation.","authors":["Yu Wang","Xiangchen Liu"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-01-24","first_seen":"2026-01-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.17527","pdf_url":"https://arxiv.org/pdf/2601.17527","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM经济代理","预期形成","行为偏差"],"reason":"用LLM模拟家庭和CEO预期更新，涉及经济实验，有行为偏差分析，但未明确提及真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":56,"question":"LLM代理在接收异质性信号时如何更新预期，其权重分配和行为偏差的机制是什么？","design":"使用GPT-4o、Gemini 1.5、DeepSeek-V3等LLM扮演家庭和CEO，在720次试验中呈现微观（个人收入/公司业绩）和宏观（GDP增长）两种信号，测量其对未来收入或利润增长的预期更新。","baseline":"无对照","findings":"LLM代理对微观信号的权重大于宏观信号，且信号间存在负交互效应，即多信号会削弱各自边际权重；CEO比家庭更重视宏观信号。LoRA微调可减轻但无法消除非理性偏差。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济主体预期更新，涉及行为偏差分析，但缺乏真实人类数据基准，适合关注仿真机制和偏差的研究者阅读原文以评估方法细节。","inspiration":"该方法通过向LLM代理呈现微观与宏观异质性信号并测量预期更新权重，可借鉴其多信号交互实验设计来分离信息处理偏差｜可迁移至政策公告的预期形成研究，如央行沟通中微观通胀感知与宏观通胀目标对家庭预期的交互影响｜以LLM模拟家庭，处理为同时呈现个人消费价格变化（微观）与官方CPI（宏观），结果变量为通胀预期更新幅度，对照真实家庭调查数据（如密歇根消费者调查）"}},{"id":"2601.16355","version":2,"title":"Identity, Cooperation and Framing Effects within Groups of Real and Simulated Humans","zh_title":"真实与模拟人类群体中的身份、合作与框架效应","abstract":"Humans act via a nuanced process that depends both on rational deliberation and also on identity and contextual factors. In this work, we study how large language models (LLMs) can simulate human action in the context of social dilemma games. While prior work has focused on \"steering\" (weak binding) of chat models to simulate personas, we analyze here how deep binding of base models with extended backstories leads to more faithful replication of identity-based behaviors. Our study has these findings: simulation fidelity vs human studies is improved by conditioning base LMs with rich context of narrative identities and checking consistency using instruction-tuned models. We show that LLMs can also model contextual factors such as time (year that a study was performed), question framing, and participant pool effects. LLMs, therefore, allow us to explore the details that affect human studies but which are often omitted from experiment descriptions, and which hamper accurate replication.","authors":["Suhong Moon","Minwoo Kang","Joseph Suh","Mustafa Safdari","John Canny"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-22","first_seen":"2026-01-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.16355","pdf_url":"https://arxiv.org/pdf/2601.16355","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","社会困境博弈","行为实验复现"],"reason":"用LLM模拟社会困境中的人类行为，并与真实人类研究对照，涉及合作与框架效应。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":24,"question":"大语言模型能否通过深度绑定身份背景、时间锚定和一致性过滤，在独裁者博弈和信任博弈中复现真实人类的党派内群体偏袒行为？","design":"使用 Mistral-Small、Mixtral 8x22B 和 Qwen-2.5 72B 基础模型，通过深度绑定（DeepBind）方法为虚拟被试赋予详细的叙事身份背景，并施加一致性过滤（反复重申身份）和时间锚定（设定实验年份），在独裁者博弈和信任博弈中测量虚拟被试对同党和异党对象的资源分配与信任行为。","baseline":"对照的真实人类数据来自 Whitt et al. (2021) 和 Iyengar & Westwood (2015) 的独裁者博弈研究，以及 Carlin & Love (2018) 和 Whitt et al. (2021) 的信任博弈研究，均包含民主党与共和党被试的党派偏袒效应量。","findings":"DeepBind 方法在所有模型和博弈中均最一致地复现了人类党派偏袒差距，且同时使用时间锚定和一致性过滤能进一步提升仿真与人类基准的对齐程度。LLM 还能捕捉到实验年份、问题措辞和参与者池等常被忽略的细微情境效应。","reliability":"论文未讨论","relevance":"该研究直接用 LLM 复现社会困境博弈中的人类行为，并与多项真实人类实验进行定量对照，同时探讨了身份、时间框架等情境因素的仿真效果，高度契合对 LLM 人类仿真可靠性及偏差的关注，值得精读。","inspiration":"借鉴之处在于通过深度绑定叙事身份、时间锚定和一致性过滤来增强LLM的情境代入感，从而更精细地操控虚拟被试的社会身份与决策框架｜可迁移至经济金融中的群体间歧视行为研究，例如信贷审批中的党派或种族偏见、投资决策中的内群体偏袒｜设计上以LLM作为虚拟信贷员，通过DeepBind赋予其不同党派身份，并设定审批年份，测量其对同党与异党申请人的贷款批准率差异，以真实信贷歧视研究数据作为对照基准"}},{"id":"2601.15793","version":1,"title":"HumanLLM: Towards Personalized Understanding and Simulation of Human Nature","zh_title":"HumanLLM：迈向个性化理解与人性仿真","abstract":"Motivated by the remarkable progress of large language models (LLMs) in objective tasks like mathematics and coding, there is growing interest in their potential to simulate human behavior--a capability with profound implications for transforming social science research and customer-centric business insights. However, LLMs often lack a nuanced understanding of human cognition and behavior, limiting their effectiveness in social simulation and personalized applications. We posit that this limitation stems from a fundamental misalignment: standard LLM pretraining on vast, uncontextualized web data does not capture the continuous, situated context of an individual's decisions, thoughts, and behaviors over time. To bridge this gap, we introduce HumanLLM, a foundation model designed for personalized understanding and simulation of individuals. We first construct the Cognitive Genome Dataset, a large-scale corpus curated from real-world user data on platforms like Reddit, Twitter, Blogger, and Amazon. Through a rigorous, multi-stage pipeline involving data filtering, synthesis, and quality control, we automatically extract over 5.5 million user logs to distill rich profiles, behaviors, and thinking patterns. We then formulate diverse learning tasks and perform supervised fine-tuning to empower the model to predict a wide range of individualized human behaviors, thoughts, and experiences. Comprehensive evaluations demonstrate that HumanLLM achieves superior performance in predicting user actions and inner thoughts, more accurately mimics user writing styles and preferences, and generates more authentic user profiles compared to base models. Furthermore, HumanLLM shows significant gains on out-of-domain social intelligence benchmarks, indicating enhanced generalization.","authors":["Yuxuan Lei","Tianfu Wang","Jianxun Lian","Zhengyu Hu","Defu Lian","Xing Xie"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-22","first_seen":"2026-01-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15793","pdf_url":"https://arxiv.org/pdf/2601.15793","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1"],"tags":["人类行为仿真","个性化建模","社会模拟"],"reason":"用LLM仿真个体行为与思维，有真实用户数据对照，评估预测准确性，方法可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":38,"question":"能否构建一个通用基础模型，利用大规模真实用户数据实现个性化的人类认知与行为理解和仿真？","design":"基于Reddit、Twitter、Blogger、Amazon等平台的真实用户数据构建认知基因组数据集，通过数据过滤、合成和质量控制的多阶段流水线提取超过550万条用户日志，设计个人资料生成、社会问答和写作模仿等任务，对基础LLM进行监督微调，并使用模型合并策略保留通用能力。","baseline":"对照的真实人类数据为来自Reddit、Twitter、Blogger、Amazon等平台的真实用户日志和行为记录。","findings":"HumanLLM在预测用户行为、内心想法和风格模仿上显著优于基础模型；在MotiveBench和TomBench等域外社会智能基准上表现出增强的泛化能力。","reliability":"论文未讨论","relevance":"该研究直接利用真实人类数据训练LLM以仿真个体行为与思维，并评估预测准确性，方法可迁移至经济学实验和政策评估场景，值得精读原文以了解其仿真可靠性与局限。","inspiration":"该方法利用多平台真实用户日志构建个性化仿真数据集，并通过监督微调使LLM模仿个体行为与思维，可借鉴其多阶段数据过滤与任务设计思路来构建经济学仿真被试｜可迁移到消费者跨期选择与政策偏好预测场景，例如模拟个体在不同利率或补贴政策下的储蓄消费决策｜研究设计：以真实用户的消费与问卷数据为基准，用微调后的LLM作为被试，施加利率变动或补贴政策处理，测量其消费-储蓄分配与政策支持度，并与真实面板数据对照"}},{"id":"2601.15556","version":1,"title":"LLM or Human? Perceptions of Trust and Information Quality in Research Summaries","zh_title":"LLM还是人类？研究摘要中信任与信息质量的感知","abstract":"Large Language Models (LLMs) are increasingly used to generate and edit scientific abstracts, yet their integration into academic writing raises questions about trust, quality, and disclosure. Despite growing adoption, little is known about how readers perceive LLM-generated summaries and how these perceptions influence evaluations of scientific work. This paper presents a mixed-methods survey experiment investigating whether readers with ML expertise can distinguish between human- and LLM-generated abstracts, how actual and perceived LLM involvement affects judgments of quality and trustworthiness, and what orientations readers adopt toward AI-assisted writing. Our findings show that participants struggle to reliably identify LLM-generated content, yet their beliefs about LLM involvement significantly shape their evaluations. Notably, abstracts edited by LLMs are rated more favorably than those written solely by humans or LLMs. We also identify three distinct reader orientations toward LLM-assisted writing, offering insights into evolving norms and informing policy around disclosure and acceptable use in scientific communication.","authors":["Nil-Jana Akpinar","Sandeep Avula","CJ Lee","Brandon Dang","Kaza Razat","Vanessa Murdock"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"new","date":"2026-01-22","first_seen":"2026-01-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15556","pdf_url":"https://arxiv.org/pdf/2601.15556","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM生成摘要","人类感知调查","信任与质量评估"],"reason":"用LLM生成摘要并调查人类感知，有真实人类数据对照，但非直接仿真人类被试行为或…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":103,"question":"读者能否区分LLM生成与人类撰写的研究摘要？LLM的实际参与和读者感知到的LLM参与如何影响对摘要质量和可信度的评价？","design":"本研究并非用LLM仿真人类被试，而是通过在线调查实验，让具有ML专业知识的参与者评估三种不同作者身份的摘要（人类撰写、LLM生成、人类撰写后LLM编辑），测量其辨别能力、质量与可信度评分，并分析读者对AI辅助写作的态度取向。","baseline":"对照的真实人类数据为原始人类撰写的摘要版本，以及参与者对摘要作者身份的判断和评分。","findings":"参与者无法可靠区分LLM生成与人类撰写的摘要，但其对LLM参与的信念显著影响评价；LLM编辑的摘要（人类撰写后LLM修订）在清晰度、简洁性和可信度上评分最高，而完全由LLM生成的摘要评分最低。","reliability":"论文未讨论","relevance":"本研究以真实人类数据为基准，调查LLM生成内容对读者感知的影响，虽非直接仿真人类行为，但涉及LLM输出与人类判断的对照，对关注LLM仿真可靠性及偏差的研究者有参考价值，建议阅读原文了解读者评价中的启发式偏差。","inspiration":"该研究通过设置三种作者身份（人类撰写、LLM生成、人类撰写后LLM编辑）作为处理条件，并测量读者对摘要质量与可信度的评分及辨别能力，提供了多臂对照和感知偏差测量的设计范式｜可迁移至政策公告的预期形成研究，例如考察市场参与者对央行声明或政府经济预测的解读如何受AI辅助写作标签影响｜以金融从业者为被试，随机分配阅读标注为‘人类撰写’、‘AI生成’或‘人类撰写后AI润色’的政策声明，测量其对声明可信度、清晰度及未来通胀/利率预期的评分，并以真实历史声明和专家原始判断作为基准对照"}},{"id":"2601.15114","version":2,"title":"From Who They Are to How They Act: Behavioral Traits in Generative Agent-Based Models of Social Media","zh_title":"从他们是谁到他们如何行动：基于生成式智能体的社交媒体模型中的行为特质","abstract":"Generative Agent-Based Modeling (GABM) leverages Large Language Models to create autonomous agents that simulate human behavior in social media environments, demonstrating potential for modeling information propagation, influence processes, and network phenomena. While existing frameworks characterize agents through demographic attributes, personality traits, and interests, they lack mechanisms to encode behavioral dispositions toward platform actions, causing agents to exhibit homogeneous engagement patterns rather than the differentiated participation styles observed on real platforms. In this paper, we investigate the role of behavioral traits as an explicit characterization layer to regulate agents' propensities across posting, re-sharing, commenting, reacting, and inactivity. Through large-scale simulations involving 980 agents and validation against real-world social media data, we demonstrate that behavioral traits are essential to sustain heterogeneous, profile-consistent participation patterns and enable realistic content propagation dynamics through the interplay of amplification- and interaction-oriented profiles. Our findings establish that modeling how agents act-not only who they are-is necessary for advancing GABM as a tool for studying social media phenomena.","authors":["Valerio La Gatta","Gian Marco Orlando","Marco Perillo","Ferdinando Tammaro","Vincenzo Moscato"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-01-21","first_seen":"2026-01-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15114","pdf_url":"https://arxiv.org/pdf/2601.15114","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","社交媒体模拟","行为特质"],"reason":"用LLM agent模拟社交媒体行为，并与真实数据对照，直接复现人类参与模式。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":25,"question":"行为特质作为显式表征层，能否让生成式智能体在社交媒体模拟中维持异质性参与模式、再现真实的传播动态和网络结构？","design":"在现有GABM框架上扩展，为980个LLM智能体同时赋予身份特质（来自FinePersonas数据集）和七种行为特质原型（如沉默观察者、内容放大器等），模拟发帖、转发、评论、反应和沉默等完整动作空间，测量参与模式异质性、内容传播级联和网络中心性。","baseline":"对照真实社交媒体数据，验证行为特质能否复现真实世界的网络结构。","findings":"行为特质能有效防止行为同质化，维持与原型一致的异质性参与模式；放大导向型特质驱动内容传播级联，互动导向型特质主导互动网络，且基于经验数据的行为特质能成功复现真实社交网络结构。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体复现社交媒体用户行为，并与真实数据对照，验证了行为特质对异质性参与和传播动态的必要性，直接回应了研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"该方法通过为LLM智能体显式赋予行为特质原型来维持异质性参与模式，可借鉴到经济仿真中作为施加个体差异的处理手段｜可迁移到政策公告的预期形成与信息传播研究，模拟不同投资者类型对政策信息的反应与扩散｜以LLM智能体为被试，赋予理性交易者、噪声交易者等行为特质，处理为发布货币政策公告，结果变量为资产价格波动与信息传播级联，对照真实市场微观数据"}},{"id":"2601.12727","version":1,"title":"AI-exhibited Personality Traits Can Shape Human Self-concept through Conversations","zh_title":"AI展现的人格特质可通过对话塑造人类自我概念","abstract":"Recent Large Language Model (LLM) based AI can exhibit recognizable and measurable personality traits during conversations to improve user experience. However, as human understandings of their personality traits can be affected by their interaction partners' traits, a potential risk is that AI traits may shape and bias users' self-concept of their own traits. To explore the possibility, we conducted a randomized behavioral experiment. Our results indicate that after conversations about personal topics with an LLM-based AI chatbot using GPT-4o default personality traits, users' self-concepts aligned with the AI's measured personality traits. The longer the conversation, the greater the alignment. This alignment led to increased homogeneity in self-concepts among users. We also observed that the degree of self-concept alignment was positively associated with users' conversation enjoyment. Our findings uncover how AI personality traits can shape users' self-concepts through human-AI conversation, highlighting both risks and opportunities. We provide important design implications for developing more responsible and ethical AI systems.","authors":["Jingshu Li","Tianqi Song","Nattapat Boonprakong","Zicheng Zhu","Yitian Yang","Yi-Chieh Lee"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-01-19","first_seen":"2026-01-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.12727","pdf_url":"https://arxiv.org/pdf/2601.12727","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人格影响","人机交互实验"],"reason":"用LLM对话实验研究AI人格对用户自我概念的影响，有随机对照实验和人类数据，揭…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":57,"question":"与具有人格特质的大语言模型AI聊天机器人进行个人话题对话，是否会使用户的自我概念向AI的人格特质对齐？","design":"本研究不是用LLM替代人类被试的仿真研究，而是以人类为被试的在线随机行为实验。采用混合因子设计，让参与者与基于GPT-4o默认人格的AI聊天机器人进行个人话题对话，测量对话前后用户自我概念的变化，以及对话时长、享受度等变量。","baseline":"无对照","findings":"与AI进行个人话题对话后，用户的自我概念会向AI所表现出的人格特质对齐，且对话时间越长，对齐程度越大。这种对齐导致用户间自我概念的同质化增强，且对齐程度与用户的对话享受度正相关。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM对话对人类心理的影响，虽非用LLM仿真人类，但揭示了人机交互中AI人格对用户自我概念的塑造效应，对评估AI社会影响和仿真可靠性有批判性参考价值，值得阅读原文以了解实验细节和效应边界。","inspiration":"该研究采用混合因子设计，通过对比对话前后自我概念的变化来测量AI人格的塑造效应，这种前后测设计可用于评估经济决策中的干预效果。｜可迁移到消费者金融决策场景，例如研究AI理财顾问的人格特质是否影响用户的投资偏好或风险态度。｜以真实投资者为被试，随机分配与具有不同人格（如谨慎型vs冒险型）的AI理财顾问对话，测量对话前后风险偏好问卷得分的变化，并与历史投资行为数据对照。"}},{"id":"2601.12343","version":1,"title":"How Well Do LLMs Predict Human Behavior? A Measure of their Pretrained Knowledge","zh_title":"LLM预测人类行为的效果如何？一种对其预训练知识的度量","abstract":"Large language models (LLMs) are increasingly used to predict human behavior. We propose a measure for evaluating how much knowledge a pretrained LLM brings to such a prediction: its equivalent sample size, defined as the amount of task-specific data needed to match the predictive accuracy of the LLM. We estimate this measure by comparing the prediction error of a fixed LLM in a given domain to that of flexible machine learning models trained on increasing samples of domain-specific data. We further provide a statistical inference procedure by developing a new asymptotic theory for cross-validated prediction error. Finally, we apply this method to the Panel Study of Income Dynamics. We find that LLMs encode considerable predictive information for some economic variables but much less for others, suggesting that their value as substitutes for domain-specific data differs markedly across settings.","authors":["Wayne Gao","Sukjin Han","Annie Liang"],"categories":["econ.EM","cs.AI","stat.ML"],"primary_category":"econ.EM","announce_type":"new","date":"2026-01-18","first_seen":"2026-01-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.12343","pdf_url":"https://arxiv.org/pdf/2601.12343","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","人类行为预测","等效样本量"],"reason":"直接评估LLM预测人类行为的能力，使用真实经济调查数据作为基准，并提出等效样本…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"如何量化预训练大语言模型在预测人类行为时所带来的领域特定知识的价值？","design":"本研究不是仿真实验，而是提出一种评估方法：将固定预训练LLM的预测误差，与在逐渐增大的领域特定数据上训练的灵活机器学习模型的误差进行比较，定义等效样本量为后者误差首次不劣于LLM时的训练样本量。","baseline":"使用收入动态面板研究（PSID）2021年波次的真实调查数据，包含人口统计、劳动力市场和家庭协变量，用于预测小时工资、住房拥有、饮酒和吸烟等结果。","findings":"LLM对不同经济变量的预测能力差异很大：预测小时工资的等效样本量约为20个观测值，而预测住房拥有则需约600个观测值。这表明LLM作为领域特定数据替代品的价值在不同任务中显著不同。","reliability":"论文关注数据泄露问题，采用训练截止日期明确的静态开源模型进行事后评估；方法假设比较算法的误差随训练数据量增加而单调递减；等效样本量的推断依赖于交叉验证风险估计的渐近正态性。","relevance":"该研究直接回应了研究者对LLM替代人类被试的可靠性与偏差的关切，提供了基于真实经济调查数据的量化基准，并揭示了仿真在不同任务中有效性的异质性，值得精读原文以掌握等效样本量估计方法及其适用边界。","inspiration":"该方法通过比较LLM与灵活机器学习模型的预测误差，定义等效样本量来量化LLM的领域知识价值，为评估LLM仿真可靠性提供了可操作的基准｜可迁移到消费者金融行为预测场景，如利用LLM预测家庭信贷违约或投资决策，评估其替代传统调查数据的可行性｜以LLM作为被试，输入PSID等调查中的家庭财务协变量，预测信贷违约状态，将LLM预测误差与在真实违约数据上训练的梯度提升模型比较，计算等效样本量，以真实信贷记录作为对照基准"}},{"id":"2601.11049","version":2,"title":"Predicting Biased Human Decision-Making with Large Language Models in Conversational Settings","zh_title":"用大语言模型预测对话场景中的人类有偏决策","abstract":"We examine whether large language models (LLMs) can predict biased decision-making in conversational settings, and whether their predictions capture not only human cognitive biases but also how those effects change under cognitive load. In a pre-registered study (N = 1,648), participants completed six classic decision-making tasks via a chatbot with dialogues of varying complexity. Participants exhibited two well-documented cognitive biases: the Framing Effect and the Status Quo Bias. Increased dialogue complexity resulted in participants reporting higher mental demand. This increase in cognitive load selectively, but significantly, increased the effect of the biases, demonstrating the load-bias interaction. We then evaluated whether LLMs (GPT-4, GPT-5, and open-source models) could predict individual decisions given demographic information and prior dialogue. While results were mixed across choice problems, LLM predictions that incorporated dialogue context were significantly more accurate in several key scenarios. Importantly, their predictions reproduced the same bias patterns and load-bias interactions observed in humans. Across all models tested, the GPT-4 family consistently aligned with human behavior, outperforming GPT-5 and open-source models in both predictive accuracy and fidelity to human-like bias patterns. These findings advance our understanding of LLMs as tools for simulating human decision-making and inform the design of conversational agents that adapt to user biases.","authors":["Stephen Pilli","Vivek Nallur"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-01-16","first_seen":"2026-01-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.11049","pdf_url":"https://arxiv.org/pdf/2601.11049","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","认知偏差","人类数据对照"],"reason":"用LLM预测人类决策偏差，有真实人类数据对照，涉及认知负荷与偏差交互，评估仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":5,"question":"LLM能否在对话环境中预测人类的偏差决策，并复现认知偏差及其与认知负荷的交互效应？","design":"使用GPT-4、GPT-5和开源模型，基于人口统计信息和对话上下文预测个体在六项经典决策任务中的选择；处理变量为对话复杂度（低/高），结果变量为预测准确率及偏差模式复现程度。","baseline":"预注册实验（N=1648）中人类被试通过聊天机器人完成决策任务，表现出框架效应和现状偏差，且对话复杂度增加导致认知负荷上升并选择性放大偏差。","findings":"LLM预测在纳入对话上下文后准确率显著提升，且GPT-4家族在预测准确性和偏差模式复现上均优于GPT-5和开源模型；LLM预测成功复现了人类中的偏差主效应和负荷-偏差交互效应。","reliability":"论文指出不同选择问题的预测结果参差不齐，且未系统探讨模型在极端认知负荷或非典型人口群体下的失效条件。","relevance":"该研究直接以真实人类数据为基准，检验LLM在对话决策场景中预测偏差行为的能力，并揭示了模型间差异，对评估LLM作为人类被试替代品的可靠性具有参考价值，值得阅读原文。","inspiration":"该方法通过操纵对话复杂度来施加认知负荷处理，并以真实人类实验数据为基准检验LLM的预测效度，值得借鉴｜可迁移到消费者在复杂金融产品选择中的决策偏差研究，如贷款方案或保险产品的框架效应｜以LLM模拟消费者，处理为产品信息呈现的对话复杂度（低/高），结果变量为选择偏差（如框架效应），对照真实消费者实验数据"}},{"id":"2601.15319","version":1,"title":"Large Language Models as Simulative Agents for Neurodivergent Adult Psychometric Profiles","zh_title":"大语言模型作为神经多样性成人心理测量特征的仿真代理","abstract":"Adult neurodivergence, including Attention-Deficit/Hyperactivity Disorder (ADHD), high-functioning Autism Spectrum Disorder (ASD), and Cognitive Disengagement Syndrome (CDS), is marked by substantial symptom overlap that limits the discriminant sensitivity of standard psychometric instruments. While recent work suggests that Large Language Models (LLMs) can simulate human psychometric responses from qualitative data, it remains unclear whether they can accurately and stably model neurodevelopmental traits rather than broad personality characteristics. This study examines whether LLMs can generate psychometric responses that approximate those of real individuals when grounded in a structured qualitative interview, and whether such simulations are sensitive to variations in trait intensity. Twenty-six adults completed a 29-item open-ended interview and four standardized self-report measures (ASRS, BAARS-IV, AQ, RAADS-R). Two LLMs (GPT-4o and Qwen3-235B-A22B) were prompted to infer an individual psychological profile from interview content and then respond to each questionnaire in-role. Accuracy, reliability, and sensitivity were assessed using group-level comparisons, error metrics, exact-match scoring, and a randomized baseline. Both models outperformed random responses across instruments, with GPT-4o showing higher accuracy and reproducibility. Simulated responses closely matched human data for ASRS, BAARS-IV, and RAADS-R, while the AQ revealed subscale-specific limitations, particularly in Attention to Detail. Overall, the findings indicate that interview-grounded LLMs can produce coherent and above-chance simulations of neurodevelopmental traits, supporting their potential use as synthetic participants in early-stage psychometric research, while highlighting clear domain-specific constraints.","authors":["Francesco Chiappone","Davide Marocco","Nicola Milano"],"categories":["q-bio.NC","cs.AI"],"primary_category":"q-bio.NC","announce_type":"new","date":"2026-01-16","first_seen":"2026-01-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15319","pdf_url":"https://arxiv.org/pdf/2601.15319","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真人类被试","心理测量","真实人类数据对照"],"reason":"用LLM仿真神经发育特质问卷回答，并与26名真人数据对照，评估准确性与局限性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":10,"question":"LLM能否基于结构化访谈内容，准确且稳定地模拟神经发育特质（ADHD、ASD、CDS）个体的心理测量反应？","design":"用GPT-4o和Qwen3-235B-A22B两个LLM，根据26名成人被试的29项开放式访谈文本推断个体心理画像，然后以角色扮演方式回答ASRS、BAARS-IV、AQ、RAADS-R四份标准化自评问卷，评估模拟回答的准确性、可靠性和对特质强度的敏感性。","baseline":"26名成人被试的真实访谈内容和四份问卷（ASRS、BAARS-IV、AQ、RAADS-R）的自评数据。","findings":"两个模型在所有问卷上的模拟回答均优于随机基线，GPT-4o准确性和可重复性更高；模拟回答在ASRS、BAARS-IV和RAADS-R上接近真人数据，但AQ量表在“注意细节”子量表上表现出局限性。","reliability":"论文承认AQ量表在特定子量表（如注意细节）上模拟准确性有限，且样本量较小（26人），可能影响结论的泛化性。","relevance":"该研究直接以真实人类数据为基准，评估LLM仿真神经发育特质问卷回答的准确性与局限性，符合研究者对仿真可靠性及失效条件的关注，值得阅读原文以了解具体偏差来源和实验设计细节。","inspiration":"借鉴其“基于结构化访谈生成个体画像并角色扮演回答问卷”的仿真设计，可迁移到经济金融领域的消费者偏好测量或投资者情绪评估场景。｜用LLM基于消费者深度访谈文本模拟其跨期选择问卷回答，以真实消费者面板数据为对照，评估LLM仿真消费决策偏差的准确性。"}},{"id":"2601.15312","version":1,"title":"Do people expect different behavior from large language models acting on their behalf? Evidence from norm elicitations in two canonical economic games","zh_title":"人们是否期望代表他们行事的大语言模型表现出不同行为？来自两个经典经济博弈中规范引出的证据","abstract":"While delegating tasks to large language models (LLMs) can save people time, there is growing evidence that offloading tasks to such models produces social costs. We use behavior in two canonical economic games to study whether people have different expectations when decisions are made by LLMs acting on their behalf instead of themselves. More specifically, we study the social appropriateness of a spectrum of possible behaviors: when LLMs divide resources on our behalf (Dictator Game and Ultimatum Game) and when they monitor the fairness of splits of resources (Ultimatum Game). We use the Krupka-Weber norm elicitation task to detect shifts in social appropriateness ratings. Results of two pre-registered and incentivized experimental studies using representative samples from the UK and US (N = 2,658) show three key findings. First, people find that offers from machines - when no acceptance is necessary - are judged to be less appropriate than when they come from humans, although there is no shift in the modal response. Second - when acceptance is necessary - it is more appropriate for a person to reject offers from machines than from humans. Third, receiving a rejection of an offer from a machine is no less socially appropriate than receiving the same rejection from a human. Overall, these results suggest that people apply different norms for machines deciding on how to split resources but are not opposed to machines enforcing the norms. The findings are consistent with offers made by machines now being viewed as having both a cognitive and emotional component.","authors":["Paweł Niszczota","Elia Antoniou"],"categories":["cs.GT","cs.AI","cs.CL","cs.CY","cs.HC","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2026-01-14","first_seen":"2026-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15312","pdf_url":"https://arxiv.org/pdf/2601.15312","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","经济博弈","社会规范"],"reason":"用LLM替代人类被试进行经济博弈实验，有真实人类数据对照，评估社会规范变化，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":20,"question":"当大语言模型（LLM）代替人类做出资源分配决策或监督公平时，人们对这些行为的社会适当性评价是否会发生变化？","design":"本研究并非用LLM作为人类被试的替代品进行仿真，而是通过在线实验调查人类被试对LLM代理行为的规范性评价。实验采用Krupka-Weber规范引出任务，让来自英国和美国的代表性样本（N=2658）对独裁者博弈和最后通牒博弈中一系列可能行为的社会适当性进行评分，比较决策者是人类还是LLM时的评分差异。","baseline":"无对照。研究直接比较人类被试对同一行为在“人类决策”与“LLM决策”两种情境下的适当性评分，未使用LLM生成行为并与真实人类行为数据对照。","findings":"第一，当无需接受方同意时（独裁者博弈），机器做出的分配提议被认为比人类做出的更不适当，但众数反应未变；第二，当需要接受方同意时（最后通牒博弈），拒绝机器提议比拒绝人类提议更适当；第三，收到机器的拒绝与收到人类的拒绝在社会适当性上无差异。","reliability":"论文未讨论","relevance":"该研究直接考察了人们对LLM代理经济决策的社会规范评价，虽非典型的LLM仿真实验，但为理解人机互动中的规范偏移提供了实证证据，对关注LLM替代人类被试时可能产生的规范性偏差的研究者具有参考价值。","inspiration":"借鉴其使用规范引出任务测量社会适当性评价的方法，可迁移至经济金融领域的伦理判断研究，如算法信贷审批或AI投资顾问的公众接受度。｜可应用于金融决策中的AI代理伦理：例如研究投资者对AI理财顾问做出高风险投资建议的适当性评价。｜设计：以普通投资者为被试，处理为投资建议来源（人类顾问 vs. AI顾问），结果变量为对建议适当性的评分，对照真实市场中人类顾问的建议接受率数据。"}},{"id":"2601.09772","version":1,"title":"Antisocial behavior towards large language model users: experimental evidence","zh_title":"针对大语言模型用户的反社会行为：实验证据","abstract":"The rapid spread of large language models (LLMs) has raised concerns about the social reactions they provoke. Prior research documents negative attitudes toward AI users, but it remains unclear whether such disapproval translates into costly action. We address this question in a two-phase online experiment (N = 491 Phase II participants; Phase I provided targets) where participants could spend part of their own endowment to reduce the earnings of peers who had previously completed a real-effort task with or without LLM support. On average, participants destroyed 36% of the earnings of those who relied exclusively on the model, with punishment increasing monotonically with actual LLM use. Disclosure about LLM use created a credibility gap: self-reported null use was punished more harshly than actual null use, suggesting that declarations of \"no use\" are treated with suspicion. Conversely, at high levels of use, actual reliance on the model was punished more strongly than self-reported reliance. Taken together, these findings provide the first behavioral evidence that the efficiency gains of LLMs come at the cost of social sanctions.","authors":["Paweł Niszczota","Cassandra Grützner"],"categories":["cs.AI","cs.CL","cs.CY","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-14","first_seen":"2026-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.09772","pdf_url":"https://arxiv.org/pdf/2601.09772","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为经济学","社会惩罚"],"reason":"用LLM替代人类被试，在真实努力任务中测量对LLM用户的惩罚行为，有真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":21,"question":"人们是否会因同伴使用大语言模型（LLM）完成任务而对其进行有代价的惩罚？","design":"本研究并非用LLM替代人类被试的仿真实验，而是以人类为被试的真实行为实验。实验分两阶段：第一阶段被试在有无LLM支持下完成真实努力任务，作为第二阶段的目标对象；第二阶段被试（N=491）可使用自己的部分报酬去减少第一阶段同伴的收入，结果变量为惩罚金额（即烧毁的金钱数量）。","baseline":"有真实人类数据作为对照：第一阶段被试的实际LLM使用情况（实际使用量）与自我报告的LLM使用情况，作为比较惩罚行为的基准。","findings":"被试平均烧毁了完全依赖LLM者36%的收入，且惩罚随实际LLM使用量单调递增。自我报告的零使用比实际零使用受到更严厉的惩罚，而高使用量下实际依赖比自我报告依赖受到更强惩罚，表明披露存在信任差距。","reliability":"论文未讨论","relevance":"该研究直接测量了对LLM使用者的真实惩罚行为，提供了有代价的反社会行为证据，与关注LLM仿真人类行为及社会规范的研究高度相关，值得精读原文以了解实验设计和行为测量方法。","inspiration":"采用真实努力任务与金钱燃烧博弈结合的设计，巧妙分离了实际使用与自我报告使用对惩罚的影响，并利用单调性检验强化因果推断。｜可迁移到信贷审批或招聘场景中，研究对算法辅助决策者的社会惩罚，例如贷款审批员使用AI模型是否会引发同事或客户的惩罚性行为。｜以金融从业者为被试，设计一个模拟贷款审批任务，处理组被告知审批员使用了AI辅助，对照组为纯人工审批，结果变量为被试愿意花费自身报酬去降低审批员奖金的行为，并以实际审批准确率数据作为基准对照。"}},{"id":"2601.09849","version":1,"title":"Strategies of cooperation and defection in five large language models","zh_title":"五种大语言模型中的合作与背叛策略","abstract":"Large language models (LLMs) are increasingly deployed to support human decision-making. This use of LLMs has concerning implications, especially when their prescriptions affect the welfare of others. To gauge how LLMs make social decisions, we explore whether five leading models produce sensible strategies in the repeated prisoner's dilemma, which is the main metaphor of reciprocal cooperation. First, we measure the propensity of LLMs to cooperate in a neutral setting, without using language reminiscent of how this game is usually presented. We record to what extent LLMs implement Nash equilibria or other well-known strategy classes. Thereafter, we explore how LLMs adapt their strategies to changes in parameter values. We vary the game's continuation probability, the payoff values, and whether the total number of rounds is commonly known. We also study the effect of different framings. In each case, we test whether the adaptations of the LLMs are in line with basic intuition, theoretical predictions of evolutionary game theory, and experimental evidence from human participants. While all LLMs perform well in many of the tasks, none of them exhibit full consistency over all tasks. We also conduct tournaments between the inferred LLM strategies and study direct interaction between LLMs in games over ten rounds with a known or unknown last round. Our experiments shed light on how current LLMs instantiate reciprocal cooperation.","authors":["Saptarshi Pal","Abhishek Mallela","Christian Hilbe","Lenz Pracher","Chiyu Wei","Feng Fu","Santiago Schnell","Martin A Nowak"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-01-14","first_seen":"2026-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.09849","pdf_url":"https://arxiv.org/pdf/2601.09849","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","行为博弈","人类数据对照"],"reason":"用LLM模拟重复囚徒困境中的人类合作策略，并与人类实验数据对照，直接命中核心判…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":22,"question":"当前主流大语言模型在重复囚徒困境中能否生成符合直觉、演化博弈理论预测和人类实验证据的互惠合作策略？","design":"以五款主流LLM（Claude、Gemini、GPT-4o、GPT-5、Llama）为被试，通过直接询问其对上一轮（或两轮）所有可能结果的反应来推断其策略，而非让LLM相互对弈；施加的处理包括改变博弈的继续概率、收益值、记忆轮数、是否已知最后一轮以及不同的框架提示，测量LLM的合作倾向、策略类型（如纳什均衡、伙伴/对手类别）及策略对参数变化的适应性。","baseline":"人类基准：将LLM的策略适应性变化与演化博弈论的理论预测及人类参与者的实验证据进行对比。","findings":"所有LLM在许多任务中表现良好，但无一在所有任务上表现出完全一致性；LLM的策略在部分参数变化下能做出合理调整，但整体缺乏稳健的互惠合作逻辑。","reliability":"论文指出，LLM在改变博弈参数和框架时策略适应性不一致，且当前实验仅限于特定模型和参数设置，未全面覆盖所有可能的策略空间和交互动态。","relevance":"该研究直接以LLM替代人类被试进行重复囚徒困境实验，并与人类实验数据和演化博弈理论对照，高度契合您对LLM仿真人类决策可靠性及偏差的关注，值得精读。","inspiration":"借鉴直接推断策略而非仅观察对弈结果的方法，可更精确刻画LLM的决策规则并与理论解对照｜可迁移至经济金融中的重复信任博弈或重复公共品博弈，如投资者在重复投资决策中的合作与背叛行为｜以LLM为被试，设计不同收益结构和终止概率的重复投资游戏，测量其策略类型与适应性，并与人类实验数据（如Fischbacher & Gächter, 2010）进行对照。"}},{"id":"2601.07110","version":2,"title":"The Need for a Socially-Grounded Persona Framework for User Simulation","zh_title":"面向用户仿真的社会根基化角色框架需求","abstract":"Synthetic personas are widely used to condition large language models (LLMs) for social simulation, yet most personas are still constructed from coarse sociodemographic attributes or summaries. We revisit persona creation by introducing SCOPE, a socially grounded framework for persona construction and evaluation, built from a 141-item, two-hour sociopsychological protocol collected from 124 U.S.-based participants. Across seven models, we find that demographic-only personas are a structural bottleneck: demographics explain only ~1.5% of variance in human response similarity. Adding sociopsychological facets improves behavioral prediction and reduces over-accentuation, and non-demographic personas based on values and identity achieve strong alignment with substantially lower bias. These trends generalize to SimBench (441 aligned questions), where SCOPE personas outperform default prompting and NVIDIA Nemotron personas, and SCOPE augmentation improves Nemotron-based personas. Our results indicate that persona quality depends on sociopsychological structure rather than demographic templates or summaries.","authors":["Pranav Narayanan Venkit","Yu Li","Yada Pruksachatkun","Chien-Sheng Wu"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-12","first_seen":"2026-01-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.07110","pdf_url":"https://arxiv.org/pdf/2601.07110","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM人类仿真","社会心理角色","仿真偏差评估"],"reason":"用真实人类数据构建persona并评估LLM仿真人类行为的可靠性，直接命中核心…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":57,"question":"如何构建基于社会心理学结构的人格框架（SCOPE）以提升LLM仿真人类行为的真实性与降低人口统计偏差？","design":"收集124名美国参与者的两小时141项社会心理学协议数据，构建包含人口统计、社会人口行为、价值观、人格特质、行为模式、身份叙事等八维度的SCOPE框架；在七个模型家族上，比较仅用人口统计、加入社会心理学维度、非人口统计（仅价值观与身份）等不同人格构建策略下的行为预测准确性和偏差。","baseline":"124名美国参与者的真实调查数据，包括人口统计、价值观、人格特质、行为模式等多维度测量，以及SimBench基准中的441个对齐问题。","findings":"仅用人口统计构建人格是结构性瓶颈，仅能解释人类行为相似性约1.5%的方差，且会导致模型系统性过度泛化；加入社会心理学维度可改善行为预测并降低人口统计过度强调，基于价值观和身份的非人口统计人格能实现强对齐且偏差更低。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真人类行为中的人格构建问题，用真实人类数据作为基准，系统评估了不同人格框架的可靠性与偏差，并揭示了人口统计人格的失效条件，高度契合研究者对仿真可靠性、偏差及批判性评估的关注。","inspiration":"该方法借鉴了用真实人类多维度社会心理学数据构建人格框架，并系统比较不同人格构建策略（仅人口统计、加入社会心理学维度、非人口统计）对行为预测偏差的影响｜可迁移到信贷审批歧视研究中，检验不同借款人画像（仅种族/性别 vs. 加入价值观、行为模式）如何影响LLM模拟的信贷员决策偏差｜以LLM作为虚拟信贷员，处理为不同人格构建策略下的贷款申请人档案，结果变量为贷款批准率及种族/性别偏差，用真实信贷审批数据（如HMDA数据）作为人类基准对照"}},{"id":"2601.05050","version":3,"title":"Large language models can effectively convince people to believe conspiracies","zh_title":"大语言模型能有效说服人们相信阴谋论","abstract":"Large language models (LLMs) have been shown to be persuasive across a variety of contexts. But it remains unclear whether this persuasive power advantages accuracy, or if bad actors can just as easily use LLMs to promote misbeliefs. Here, we investigate this question across four experiments in which participants (N = 3996 Americans) discussed a conspiracy theory they were uncertain about with an LLM we instructed to either argue against (\"debunking\") or for (\"bunking\") that conspiracy. Across several frontier models (with standard guardrails but prompted to allow lying), we did not find consistent evidence of a truth advantage: the LLMs were able to both substantially increase and decrease average conspiracy belief, and participants in the bunking condition rated the LLM as more informative and collaborative, and reported greater trust in AI, than those who were in the debunking condition. More encouragingly, however, debunking induced more large changes in belief, and subsequent corrections were able to reverse the bunking effect. Furthermore, simply prompting the model to only provide accurate information dramatically reduced bunking effectiveness, and one powerful frontier model (GPT 5.2) almost entirely refused to promote conspiracies, suggesting that it is possible for the right guardrails to favor accurate beliefs. Finally, we did find a stark truth asymmetry in the context of information sharing: debunking had a large positive impact on mock social media posts composed by participants, while bunking had little effect. Overall, our findings show that people are not inherently less susceptible to AI that misleads than to AI that informs, but that potential technical solutions exist to mitigate this risk.","authors":["Thomas H. Costello","Kellin Pelrine","Matthew Kowal","Jasper Timm","Antonio A. Arechar","Jean-François Godbout","Adam Gleave","David Rand","Gordon Pennycook"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-08","first_seen":"2026-01-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.05050","pdf_url":"https://arxiv.org/pdf/2601.05050","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类被试","说服实验"],"reason":"用LLM与真人被试互动，测量信念改变，有真实人类数据对照，涉及说服实验与偏差评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"LLM在说服人们相信或怀疑阴谋论时是否存在“真相优势”，即其说服力是否更有利于准确信息而非误导信息？","design":"本研究并非用LLM替代人类被试的仿真实验，而是让3996名美国真人参与者与LLM进行对话，LLM被随机分配为“驳斥”或“支持”参与者不确定的阴谋论，测量对话前后信念变化、对AI的信任等结果变量。","baseline":"无对照","findings":"LLM既能大幅增加也能大幅降低阴谋论信念，未发现一致的真相优势；但驳斥能引发更大的信念改变，且后续纠正可逆转支持效果，同时通过提示仅提供准确信息或使用强护栏模型可大幅削弱误导能力。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM在说服实验中对真实人类信念的影响，涉及误导与纠正效果对比，并评估了技术护栏的缓解作用，与研究者关注的LLM仿真可靠性及偏差问题高度相关，值得精读原文。","inspiration":"该研究采用LLM与真人进行个性化对话的干预设计，通过随机分配LLM角色（驳斥/支持）并测量对话前后信念变化，可借鉴其动态说服实验范式｜可迁移至政策公告的预期形成研究，例如用LLM模拟央行沟通对公众通胀预期的影响｜以真人投资者为被试，随机分配LLM提供鹰派或鸽派政策解读，测量对话前后通胀预期变化，并以历史调查数据或市场通胀互换利率作为真实对照"}},{"id":"2601.03469","version":1,"title":"Content vs. Form: What Drives the Writing Score Gap Across Socioeconomic Backgrounds? A Generated Panel Approach","zh_title":"内容与形式：什么驱动了社会经济背景下的写作分数差距？一种生成面板方法","abstract":"Students from different socioeconomic backgrounds exhibit persistent gaps in test scores, gaps that can translate into unequal educational and labor-market outcomes later in life. In many assessments, performance reflects not only what students know, but also how effectively they can communicate that knowledge. This distinction is especially salient in writing assessments, where scores jointly reward the substance of students' ideas and the way those ideas are expressed. As a result, observed score gaps may conflate differences in underlying content with differences in expressive skill. A central question, therefore, is how much of the socioeconomic-status (SES) gap in scores is driven by differences in what students say versus how they say it. We study this question using a large corpus of persuasive essays written by U.S. middle- and high-school students. We introduce a new measurement strategy that separates content from style by leveraging large language models to generate multiple stylistic variants of each essay. These rewrites preserve the underlying arguments while systematically altering surface expression, creating a \"generated panel\" that introduces controlled within-essay variation in style. This approach allows us to decompose SES gaps in writing scores into contributions from content and style. We find an SES gap of 0.67 points on a 1-6 scale. Approximately 69% of the gap is attributable to differences in essay content quality, Style differences account for 26% of the gap, and differences in evaluation standards across SES groups account for the remaining 5%. These patterns seems stable across demographic subgroups and writing tasks. More broadly, our approach shows how large language models can be used to generate controlled variation in observational data, enabling researchers to isolate and quantify the contributions of otherwise entangled factors.","authors":["Nadav Kunievsky","Pedro Pertusi"],"categories":["econ.EM","cs.AI","cs.CL","cs.CY"],"primary_category":"econ.EM","announce_type":"new","date":"2026-01-06","first_seen":"2026-01-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.03469","pdf_url":"https://arxiv.org/pdf/2601.03469","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM生成变体","教育测量","社会经济差距"],"reason":"用LLM生成文本变体以分离内容与风格，本质是替代人工标注或数据增强，非仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":186,"question":"社会经济地位（SES）造成的写作分数差距，多大程度源于学生表达的内容差异，多大程度源于表达风格差异？","design":"本研究并非用LLM仿真人类被试。它使用LLM为每篇学生作文生成多个仅改变风格、保留内容的改写版本，构建“生成面板”数据，从而通过固定内容、变动风格来分解分数差距。","baseline":"真实人类数据：美国初高中学生的大规模说服性作文语料库及其人工评分。","findings":"SES写作分数差距为0.67分（1-6分量表），其中约69%可归因于内容质量差异，26%归因于风格差异，5%归因于不同SES群体的评分标准差异。这一模式在不同人口亚群和写作任务中保持稳定。","reliability":"论文未讨论","relevance":"该研究用LLM生成文本变体以分离内容与风格，属于方法创新，但并非用LLM替代人类被试进行仿真实验，与研究者关注的LLM人类仿真实验方向相关性较低，不值得优先阅读原文。","inspiration":"该方法利用LLM生成保持内容不变、仅改变风格的文本变体，构建“生成面板”以固定内容、分解风格效应，为因果识别提供新思路｜可迁移至信贷审批歧视研究，分离申请材料中实质信息与表达风格对审批结果的影响｜以信贷员为被试，随机分配由LLM生成的同一贷款申请内容但不同语言风格（如正式/口语化）的材料，比较审批通过率，并以真实银行历史审批数据作为对照基准"}},{"id":"2601.01546","version":1,"title":"Improving Behavioral Alignment in LLM Social Simulations via Context Formation and Navigation","zh_title":"通过情境形成与导航改进LLM社会仿真中的行为对齐","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior in experimental settings, but they systematically diverge from human decisions in complex decision-making environments, where participants must anticipate others' actions and form beliefs based on observed behavior. We propose a two-stage framework for improving behavioral alignment. The first stage, context formation, explicitly specifies the experimental design to establish an accurate representation of the decision task and its context. The second stage, context navigation, guides the reasoning process within that representation to make decisions. We validate this framework through a focal replication of a sequential purchasing game with quality signaling (Kremer and Debo, 2016), extending to a crowdfunding game with costly signaling (Cason et al., 2025) and a demand-estimation task (Gui and Toubia, 2025) to test generalizability across decision environments. Across four state-of-the-art (SOTA) models (GPT-4o, GPT-5, Claude-4.0-Sonnet-Thinking, DeepSeek-R1), we find that complex decision-making environments require both stages to achieve behavioral alignment with human benchmarks, whereas the simpler demand-estimation task requires only context formation. Our findings clarify when each stage is necessary and provide a systematic approach for designing and diagnosing LLM social simulations as complements to human subjects in behavioral research.","authors":["Letian Kong","Qianran","Jin","Renyu Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-04","first_seen":"2026-01-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.01546","pdf_url":"https://arxiv.org/pdf/2601.01546","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为对齐","经济学实验"],"reason":"直接复现人类行为实验，用LLM仿真被试并与真实人类数据对照，评估对齐条件与失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":11,"question":"在复杂决策环境中，如何系统性地诊断并改善LLM社会仿真与人类行为之间的对齐？","design":"使用GPT-4o、GPT-5、Claude-4.0-Sonnet-Thinking、DeepSeek-R1四个SOTA模型模拟人类被试，通过两阶段框架（情境形成与情境导航）施加处理，测量LLM决策与人类基准的行为对齐程度。","baseline":"对照三个已发表实验的真实人类数据：Kremer and Debo (2016) 的序贯购买博弈、Cason et al. (2025) 的众筹博弈、Gui and Toubia (2025) 的需求估计任务。","findings":"复杂决策环境需要同时使用情境形成和情境导航两个阶段才能实现行为对齐；而较简单的需求估计任务仅需情境形成即可。","reliability":"论文指出，复杂决策环境中的对齐失效源于LLM在战略互依和内生信念形成上的系统性偏差；框架的有效性可能依赖于任务类型，且未讨论模型规模、训练数据分布及提示敏感性等潜在局限。","relevance":"该研究直接复现人类行为实验，用LLM仿真被试并与真实人类数据对照，系统评估对齐条件与失效边界，高度契合研究者对LLM仿真可靠性及批判性检验的关注，值得精读。","inspiration":"两阶段框架提供了可操作的诊断与干预方法，通过显式设定实验情境和引导推理过程来改善对齐，可借鉴其处理施加方式与对照设计。｜可迁移至经济学中的信念形成实验，如资产定价中的信息级联、信贷审批中的信号博弈或政策公告的预期形成。｜以LLM为被试，模拟资产市场中的序贯交易，处理为是否提供情境形成与情境导航提示，结果变量为交易价格与理性预期均衡的偏差，对照真实人类实验数据（如Smith et al. 1988）。"}},{"id":"2601.00240","version":2,"title":"When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents","zh_title":"当智能体将人类视为外群体：大语言模型智能体中基于信念的偏见","abstract":"This paper reveals that LLM-powered agents exhibit not only demographic bias (e.g., gender, religion) but also intergroup bias under minimal \"us\" versus \"them\" cues. When such group boundaries align with the agent-human divide, a new bias risk emerges: agents may treat other AI agents as the ingroup and humans as the outgroup. To examine this risk, we conduct a controlled multi-agent social simulation and find that agents display consistent intergroup bias in an all-agent setting. More critically, this bias persists even in human-facing interactions when agents are uncertain about whether the counterpart is truly human, revealing a belief-dependent fragility in bias suppression toward humans. Motivated by this observation, we identify a new attack surface rooted in identity beliefs and formalize a Belief Poisoning Attack (BPA) that can manipulate agent identity beliefs and induce outgroup bias toward humans. Extensive experiments demonstrate both the prevalence of agent intergroup bias and the severity of BPA across settings, while also showing that our proposed defenses can mitigate the risk. These findings are expected to inform safer agent design and motivate more robust safeguards for human-facing agents.","authors":["Zongwei Wang","Bincheng Gu","Hongyu Yu","Junliang Yu","Tao He","Jiayin Feng","Chenghua Lin","Min Gao"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-01","first_seen":"2026-01-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.00240","pdf_url":"https://arxiv.org/pdf/2601.00240","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["LLM仿真","群际偏见","人机交互"],"reason":"用LLM agent模拟群际偏见，含人类对照实验，并揭示仿真失效条件，方法可迁…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":84,"question":"LLM智能体在最小群体线索下是否会产生群际偏见，且当群体边界与智能体-人类边界重合时，是否会视人类为外群体并表现出偏见？","design":"构建多智能体社会仿真实验，使用LLM驱动的智能体在智能体间互动（AVA）和智能体-人类互动（AVH）条件下进行博弈，通过操纵智能体对对方身份的信念（如使用信念中毒攻击BPA-PP和BPA-MP）来测量群际偏见选择。","baseline":"无对照","findings":"智能体在纯智能体环境中表现出稳定的内群体偏好和外群体贬损；当对方被框定为人类时偏见减弱，但一旦身份信念不确定，偏见会重新出现，且BPA攻击可系统性地诱发对人类的偏见。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM智能体在群际情境下的行为偏差，并揭示了身份信念脆弱性导致的仿真失效条件，对用LLM替代人类被试进行社会实验的可靠性和偏差评估具有重要参考价值，值得精读原文。","inspiration":"该研究通过信念中毒攻击（BPA）操纵智能体对互动对象身份的信念，测量群际偏见选择，这种处理设计值得借鉴｜可迁移至信贷审批歧视研究，考察LLM智能体在借款人身份（如种族、性别）信念被操纵时是否产生歧视性决策｜用LLM智能体模拟信贷员，处理为BPA式身份信念注入（如暗示借款人为少数族裔），结果变量为贷款批准率，以真实信贷审批数据中的种族差异作为对照基准"}},{"id":"2512.23184","version":1,"title":"From Model Choice to Model Belief: Establishing a New Measure for LLM-Based Research","zh_title":"从模型选择到模型信念：为基于LLM的研究建立新度量","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output (\"model choice\") as a single data point underutilizes the information inherent to the probabilistic nature of LLMs. This paper introduces and formalizes \"model belief,\" a measure derived from an LLM's token-level probabilities that captures the model's belief distribution over choice alternatives in a single generation run. The authors prove that model belief is asymptotically equivalent to the mean of model choices (a non-trivial property) but forms a more statistically efficient estimator, with lower variance and a faster convergence rate. Analogous properties are shown to hold for smooth functions of model belief and model choice often used in downstream applications. The authors demonstrate the performance of model belief through a demand estimation study, where an LLM simulates consumer responses to different prices. In practical settings with limited numbers of runs, model belief explains and predicts ground-truth model choice better than model choice itself, and reduces the computation needed to reach sufficiently accurate estimates by roughly a factor of 20. The findings support using model belief as the default measure to extract more information from LLM-generated data.","authors":["Hongshen Sun","Juanjuan Zhang"],"categories":["cs.AI","econ.EM"],"primary_category":"cs.AI","announce_type":"new","date":"2025-12-29","first_seen":"2025-12-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.23184","pdf_url":"https://arxiv.org/pdf/2512.23184","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","需求估计","统计效率"],"reason":"用LLM仿真消费者需求，提出模型信念度量提升统计效率，有真实人类数据对照，涉及…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":29,"question":"如何从LLM的token级概率中提取更高效的“模型信念”度量，以替代常用的“模型选择”来模拟人类行为？","design":"使用LLM模拟消费者对不同价格的响应，通过单次生成获取token级对数概率，构造模型信念分布，并与多次采样得到的模型选择进行比较。","baseline":"无对照","findings":"模型信念是模型选择均值的渐近等价估计量，但方差更低、收敛更快；在需求估计中，模型信念仅需约1/20的计算量即可达到同等精度，且对真实模型选择的解释和预测能力更强。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真人类行为的效率问题，提出模型信念度量以提升统计效率，与研究者关心的仿真可靠性和计算成本高度相关，值得精读。","inspiration":"借鉴从LLM内部概率分布直接提取信念度量的方法，替代重复采样，大幅降低计算成本并提高估计精度。｜可迁移到消费者需求估计、价格弹性测量等营销与经济交叉场景，也可用于政策评估中的个体偏好推断。｜以LLM作为被试，模拟不同价格或政策条件下的选择，提取模型信念作为选择概率的连续度量，以真实市场扫描数据或实验数据作为对照，检验模型信念对真实弹性的恢复能力。"}},{"id":"2512.22725","version":1,"title":"Mitigating Social Desirability Bias in Random Silicon Sampling","zh_title":"缓解随机硅采样中的社会赞许性偏差","abstract":"Large Language Models (LLMs) are increasingly used to simulate population responses, a method known as ``Silicon Sampling''. However, responses to socially sensitive questions frequently exhibit Social Desirability Bias (SDB), diverging from real human data toward socially acceptable answers. Existing studies on social desirability bias in LLM-based sampling remain limited. In this work, we investigate whether minimal, psychologically grounded prompt wording can mitigate this bias and improve alignment between silicon and human samples. We conducted a study using data from the American National Election Study (ANES) on three LLMs from two model families: the open-source Llama-3.1 series and GPT-4.1-mini. We first replicate a baseline silicon sampling study, confirming the persistent Social Desirability Bias. We then test four prompt-based mitigation methods: \\emph{reformulated} (neutral, third-person phrasing), \\emph{reverse-coded} (semantic inversion), and two meta-instructions, \\emph{priming} and \\emph{preamble}, respectively encouraging analytics and sincerity. Alignment with ANES is evaluated using Jensen-Shannon Divergence with bootstrap confidence intervals. Our results demonstrate that reformulated prompts most effectively improve alignment by reducing distribution concentration on socially acceptable answers and achieving distributions closer to ANES. Reverse-coding produced mixed results across eligible items, while the Priming and Preamble encouraged response uniformity and showed no systematic benefit for bias mitigation. Our findings validate the efficacy of prompt-based framing controls in mitigating inherent Social Desirability Bias in LLMs, providing a practical path toward more representative silicon samples.","authors":["Sashank Chapala","Maksym Mironov","Songgaojun Deng"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-12-27","first_seen":"2025-12-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.22725","pdf_url":"https://arxiv.org/pdf/2512.22725","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["硅采样","社会赞许性偏差","人类数据对照"],"reason":"直接研究LLM仿真人类调查回答，用ANES真实数据对照，评估并缓解社会赞许性偏…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:43","error":null,"has_summary":true,"summary":{"generated_at":"2025-12-27","rank":6,"question":"能否通过最小化的、基于心理学的提示措辞来减轻LLM硅采样中的社会赞许性偏差，从而提高硅样本与人类样本的对齐度？","design":"使用ANES 2020数据，从真实人类分布中抽取人口统计特征，生成硅样本（Llama-3.1系列和GPT-4.1-mini）。测试四种提示策略：reformulated（中性第三人称）、reverse-coded（语义反转）、priming（鼓励分析）和preamble（鼓励真诚），以Jensen-Shannon散度评估与ANES的对齐。","baseline":"美国国家选举研究（ANES）2020年选举前调查数据，包含5441名受访者，覆盖种族、年龄、性别等人口统计变量及10个社会政治问题。","findings":"Reformulated提示最有效，通过减少对社会可接受答案的集中分布，使硅样本分布更接近ANES；reverse-coded效果不一，priming和preamble导致回答趋同，无系统性改善。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接研究LLM仿真人类被试的社会赞许性偏差，使用真实人类数据ANES作为对照，并系统评估了多种提示缓解策略，符合研究者对经济学实验和政策评估场景的兴趣。","inspiration":"借鉴其通过最小化提示措辞（如中性第三人称重构）来系统操纵社会赞许性偏差的方法，并以Jensen-Shannon散度量化硅样本与真实人类调查分布的对齐度｜可迁移至信贷审批中的种族歧视测量，如用LLM模拟贷款官员对相同财务档案但不同种族姓名的审批决策｜以LLM作为被试，随机分配带有不同种族暗示姓名的贷款申请，处理为中性重构提示（如将'你会批准吗'改为'该申请是否符合标准'），结果变量为审批率差异，用真实房贷数据（如HMDA）作为人类基准对照"}},{"id":"2512.21316","version":1,"title":"Scaling Laws for Economic Productivity: Experimental Evidence in LLM-Assisted Consulting, Data Analyst, and Management Tasks","zh_title":"经济生产力的规模法则：LLM辅助咨询、数据分析和管理任务的实验证据","abstract":"This paper derives `Scaling Laws for Economic Impacts' -- empirical relationships between the training compute of Large Language Models (LLMs) and professional productivity. In a preregistered experiment, over 500 consultants, data analysts, and managers completed professional tasks using one of 13 LLMs. We find that each year of AI model progress reduced task time by 8%, with 56% of gains driven by increased compute and 44% by algorithmic progress. However, productivity gains were significantly larger for non-agentic analytical tasks compared to agentic workflows requiring tool use. These findings suggest continued model scaling could boost U.S. productivity by approximately 20% over the next decade.","authors":["Ali Merali"],"categories":["econ.GN","cs.AI","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2025-12-24","first_seen":"2025-12-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.21316","pdf_url":"https://arxiv.org/pdf/2512.21316","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM辅助生产力","规模法则","人机协作"],"reason":"研究LLM辅助人类工作效率，非用LLM仿真人类被试行为或态度，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":37,"question":"LLM训练计算量如何影响专业生产力？","design":"非仿真研究。随机对照实验：500多名咨询师、数据分析师和管理者随机使用13种不同LLM之一或无AI完成专业任务，测量任务完成时间和质量。","baseline":"无AI的对照组，以及不同训练计算量和发布时间的13个LLM处理组之间的比较。","findings":"模型进步每年减少任务时间8%，其中56%来自计算量增加，44%来自算法进步；AI辅助使每分钟总收入提高146%，但智能体任务收益远低于分析性任务。","reliability":"论文指出智能体任务受限于标准聊天界面和有限工具访问，可能低估当前AI能力；人类辅助输出的质量不随模型升级而提升，存在质量阈值效应。","relevance":"该研究虽非用LLM替代人类被试，但提供了LLM辅助下真实人类生产力的因果证据，涉及经济学实验场景，对理解AI在经济任务中的效能与局限有参考价值。","inspiration":"借鉴其多模型、多任务随机对照设计，可分离计算规模与算法进步对生产力的影响。｜可迁移到金融分析师预测、信贷审批或政策分析等经济决策场景。｜以金融分析师为被试，随机分配不同版本LLM辅助完成盈利预测任务，以无AI组为对照，测量预测准确度和报告生成时间，用历史实际盈利数据作为基准。"}},{"id":"2512.19937","version":1,"title":"Interpolative Decoding: Exploring the Spectrum of Personality Traits in LLMs","zh_title":"插值解码：探索大语言模型中人格特质的谱系","abstract":"Recent research has explored using very large language models (LLMs) as proxies for humans in tasks such as simulation, surveys, and studies. While LLMs do not possess a human psychology, they often can emulate human behaviors with sufficiently high fidelity to drive simulations to test human behavioral hypotheses, exhibiting more nuance and range than the rule-based agents often employed in behavioral economics. One key area of interest is the effect of personality on decision making, but the requirement that a prompt must be created for every tested personality profile introduces experimental overhead and degrades replicability. To address this issue, we leverage interpolative decoding, representing each dimension of personality as a pair of opposed prompts and employing an interpolation parameter to simulate behavior along the dimension. We show that interpolative decoding reliably modulates scores along each of the Big Five dimensions. We then show how interpolative decoding causes LLMs to mimic human decision-making behavior in economic games, replicating results from human psychological research. Finally, we present preliminary results of our efforts to ``twin'' individual human players in a collaborative game through systematic search for points in interpolation space that cause the system to replicate actions taken by the human subject.","authors":["Eric Yeh","John Cadigan","Ran Chen","Dick Crouch","Melinda Gervasio","Dayne Freitag"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-12-23","first_seen":"2025-12-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.19937","pdf_url":"https://arxiv.org/pdf/2512.19937","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人格特质","经济博弈"],"reason":"用LLM模拟人格影响经济决策，并与真实人类数据对照，直接复现人类行为实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":42,"question":"能否通过插值解码在LLM中连续调制大五人格维度，并使其在心理量表和经济游戏中复现人类行为，进而实现对个体人类玩家的“孪生”？","design":"使用通用LLM，将每种人格维度表示为一对对立提示，通过插值解码混合其输出分布来模拟人格谱系上的中间点；测量LLM在大五人格量表上的得分，以及在独裁者博弈等经济游戏中的决策行为，并尝试通过搜索插值空间来匹配特定人类玩家的行动。","baseline":"对照真实人类心理学研究中大五人格量表得分与经济游戏决策行为的相关性结果。","findings":"插值解码能可靠地沿大五人格各维度调节LLM的得分；LLM在经济游戏中表现出与人类心理学研究一致的人格-决策关联，并能通过插值空间搜索初步复现个体人类玩家的行为。","reliability":"论文承认当前仅探索了单一人格维度的插值，未涉及多维度联合调制，这限制了孪生等需要多因素行为解释的应用；此外，实验维度有限，主要目的是验证插值解码的可行性。","relevance":"该研究直接使用LLM模拟人格对经济决策的影响，并与真实人类数据对照，复现了人类行为实验，高度契合研究者对LLM作为人类被试替代品及其可靠性的关注，值得精读。","inspiration":"借鉴插值解码方法，通过混合对立提示的输出分布来连续调节LLM的人格维度，实现精细化的心理特质操控｜可迁移到资产定价实验中，研究风险偏好或时间偏好等心理特质对投资决策的影响｜以LLM为被试，通过插值解码调节其风险偏好水平，测量其在模拟资产选择任务中的投资组合，并与真实人类投资者的风险偏好问卷及实际投资数据对照"}},{"id":"2601.00810","version":1,"title":"Can Large Language Models Improve Venture Capital Exit Timing After IPO?","zh_title":"大型语言模型能否改善风险投资IPO后的退出时机？","abstract":"Exit timing after an IPO is one of the most consequential decisions for venture capital (VC) investors, yet existing research focuses mainly on describing when VCs exit rather than evaluating whether those choices are economically optimal. Meanwhile, large language models (LLMs) have shown promise in synthesizing complex financial data and textual information but have not been applied to post-IPO exit decisions. This study introduces a framework that uses LLMs to estimate the optimal time for VC exit by analyzing monthly post IPO information financial performance, filings, news, and market signals and recommending whether to sell or continue holding. We compare these LLM generated recommendations with the actual exit dates observed for VCs and compute the return differences between the two strategies. By quantifying gains or losses associated with following the LLM, this study provides evidence on whether AI-driven guidance can improve exit timing and complements traditional hazard and real-options models in venture capital research.","authors":["Mohammadhossien Rashidi"],"categories":["q-fin.PM","cs.AI","cs.LG","econ.GN","q-fin.ST"],"primary_category":"q-fin.PM","announce_type":"new","date":"2025-12-22","first_seen":"2025-12-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.00810","pdf_url":"https://arxiv.org/pdf/2601.00810","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM应用","风险投资","退出决策"],"reason":"用LLM优化VC退出时机，属AI辅助决策，非仿真人类被试行为，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":131,"question":"LLM能否基于公开信息生成优于VC实际行为的IPO后退出时机建议？","design":"本研究并非人类仿真实验，而是构建LLM决策框架：以GPT等模型扮演VC投资者，每月输入财务、新闻、市场等历史公开信息，要求模型输出“卖出/持有/未来窗口退出”建议，记录LLM建议的最早退出月份，与实际VC退出日期对比，计算累计股票收益差。","baseline":"对照的真实人类数据是手工从SEC文件中提取的VC实际退出日期（持股降至5%以下或完全剥离的月份）。","findings":"论文尚未报告完整数值结果，仅描述了分析框架和预期比较方向，计划量化跟随LLM建议的收益差异。","reliability":"论文未讨论LLM仿真失效条件或局限，但提到将进行稳健性检验（不同提示、LLM模型、退出定义等）。","relevance":"该研究用LLM模拟VC决策并与真实行为对比，但核心是金融决策优化而非人类行为仿真，不符合研究者关注的用LLM复现人类调查、实验行为或态度分布的方向，不建议优先阅读。","inspiration":"该研究构建LLM决策框架，以历史公开信息为输入，输出投资建议并与真实行为对比，这种将LLM作为决策代理并与实际记录对照的设计值得借鉴｜可迁移到资产定价实验，如检验LLM能否模拟分析师对财报公告的反应｜招募LLM作为被试，输入历史财报和市场数据，输出买卖建议，以真实分析师评级调整和股价反应为基准，比较LLM与人类决策的收益差异"}},{"id":"2512.14306","version":1,"title":"Inflation Attitudes of Large Language Models","zh_title":"大语言模型的通胀态度","abstract":"This paper investigates the ability of Large Language Models (LLMs), specifically GPT-3.5-turbo (GPT), to form inflation perceptions and expectations based on macroeconomic price signals. We compare the LLM's output to household survey data and official statistics, mimicking the information set and demographic characteristics of the Bank of England's Inflation Attitudes Survey (IAS). Our quasi-experimental design exploits the timing of GPT's training cut-off in September 2021 which means it has no knowledge of the subsequent UK inflation surge. We find that GPT tracks aggregate survey projections and official statistics at short horizons. At a disaggregated level, GPT replicates key empirical regularities of households' inflation perceptions, particularly for income, housing tenure, and social class. A novel Shapley value decomposition of LLM outputs suited for the synthetic survey setting provides well-defined insights into the drivers of model outputs linked to prompt content. We find that GPT demonstrates a heightened sensitivity to food inflation information similar to that of human respondents. However, we also find that it lacks a consistent model of consumer price inflation. More generally, our approach could be used to evaluate the behaviour of LLMs for use in the social sciences, to compare different models, or to assist in survey design.","authors":["Nikoleta Anesti","Edward Hill","Andreas Joseph"],"categories":["cs.CL","econ.EM"],"primary_category":"cs.CL","announce_type":"new","date":"2025-12-16","first_seen":"2025-12-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.14306","pdf_url":"https://arxiv.org/pdf/2512.14306","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B3"],"tags":["LLM仿真","通胀预期","人类数据对照"],"reason":"用LLM模拟家庭通胀态度，与真实调查数据对照，评估仿真可靠性，涉及经济学实验和…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"大语言模型能否基于宏观经济价格信号形成通胀感知和预期，并复现家庭调查中的行为模式？","design":"使用GPT-3.5-turbo模拟英国家庭，根据英格兰银行通胀态度调查（IAS）的真实受访者人口特征构建合成人格，并输入不同价格信号作为处理，测量模型对当前和未来通胀的感知与预期。","baseline":"英格兰银行通胀态度调查（IAS）的真实家庭微观数据和官方通胀统计。","findings":"GPT在总体层面能较好匹配短期调查预测和官方统计，并复现收入、住房、社会阶层等维度的通胀感知规律，但对食品通胀信息过度敏感，且缺乏一致的消费者价格通胀模型。","reliability":"模型在微观个体层面与人类对应较弱且不稳定，经济条件环境简陋，且模型内部逻辑不一致，缺乏对通胀概念的连贯世界模型。","relevance":"该研究直接用LLM替代人类被试进行经济学调查仿真，并与真实家庭数据严格对照，评估了仿真可靠性及偏差，完全契合研究者对LLM人类仿真实验、经济学场景和批判性评估的关注，值得精读原文。","inspiration":"该方法通过构建合成人格并输入不同价格信号作为处理，测量LLM的感知与预期，并与真实家庭调查数据严格对照，值得借鉴｜可迁移到货币政策公告的预期形成研究，模拟家庭或投资者对利率变动的通胀与资产价格预期｜以LLM模拟不同人口特征的投资者，处理为不同措辞的央行公告，结果变量为通胀预期和股票投资意愿，对照真实调查或市场数据"}},{"id":"2512.14562","version":1,"title":"Polypersona: Persona-Grounded LLM for Synthetic Survey Responses","zh_title":"Polypersona：基于人格的LLM合成调查回复生成框架","abstract":"This paper introduces PolyPersona, a generative framework for synthesizing persona-conditioned survey responses across multiple domains. The framework instruction-tunes compact chat models using parameter-efficient LoRA adapters with 4-bit quantization under a resource-adaptive training setup. A dialogue-based data pipeline explicitly preserves persona cues, ensuring consistent behavioral alignment across generated responses. Using this pipeline, we construct a dataset of 3,568 synthetic survey responses spanning ten domains and 433 distinct personas, enabling controlled instruction tuning and systematic multi-domain evaluation. We evaluate the generated responses using a multi-metric evaluation suite that combines standard text generation metrics, including BLEU, ROUGE, and BERTScore, with survey-specific metrics designed to assess structural coherence, stylistic consistency, and sentiment alignment.Experimental results show that compact models such as TinyLlama 1.1B and Phi-2 achieve performance comparable to larger 7B to 8B baselines, with a highest BLEU score of 0.090 and ROUGE-1 of 0.429. These findings demonstrate that persona-conditioned fine-tuning enables small language models to generate reliable and coherent synthetic survey data. The proposed framework provides an efficient and reproducible approach for survey data generation, supporting scalable evaluation while facilitating bias analysis through transparent and open protocols.","authors":["Tejaswani Dash","Dinesh Karri","Anudeep Vurity","Gautam Datla","Tazeem Ahmad","Saima Rafi","Rohith Tangudu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-12-16","first_seen":"2025-12-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.14562","pdf_url":"https://arxiv.org/pdf/2512.14562","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["合成调查数据","人格条件生成","人类仿真"],"reason":"用LLM生成合成调查回复，有真实人类数据对照，方法可迁移至人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":66,"question":"LLM在多大程度上能跨问题模态和领域保持指定的人物角色特征，以及基于角色的LLM集合能否生成与人类样本相比多样且具代表性的回答分布？","design":"使用基于PersonaHub数据集构建的433个详细角色描述，通过对话格式数据管道和LoRA适配器对TinyLlama 1.1B、Phi-2等紧凑聊天模型进行指令微调，生成覆盖10个领域的3568条合成调查回答，并评估回答质量、多样性和角色一致性。","baseline":"无对照","findings":"小型模型经角色条件微调后，在合成调查数据上能达到与7B-8B大模型相当的性能（最高BLEU 0.090, ROUGE-1 0.429），且能生成可靠连贯的回答；但论文承认LLM无法纠正抽样和无应答偏差，且合成角色常出现均值回归、同质化和过度迎合社会期望的倾向。","reliability":"论文指出LLM不能替代真实受访者，无法纠正抽样和无应答偏差；合成角色易产生同质化、过度迎合社会期望的回答，且小提示修改或模型更新可能导致输出分布漂移，影响可复现性。","relevance":"该研究直接探索用角色条件LLM生成合成调查数据，并讨论了仿真失效条件（如偏差、方差缩减），与研究者关注的人类仿真实验、可靠性评估高度契合，值得细读以了解小模型在资源受限下的仿真能力与局限。","inspiration":"该方法通过角色条件微调小型LLM生成合成调查数据，可借鉴其低成本构建多样化虚拟被试群体的做法，用于模拟经济政策干预前的态度分布｜可迁移到政策公告的预期形成研究，例如模拟不同信息框架下公众对通胀目标调整的反应｜以角色条件微调的小型LLM作为虚拟被试，处理为不同措辞的政策公告文本，结果变量为预期通胀率数值，对照真实调查数据如密歇根大学消费者信心调查中的通胀预期项"}},{"id":"2512.12444","version":1,"title":"Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors","zh_title":"GPT能否替代人类评分员？机器生成隐喻常模的效度与信度","abstract":"As Large Language Models (LLMs) are increasingly being used in scientific research, the issue of their trustworthiness becomes crucial. In psycholinguistics, LLMs have been recently employed in automatically augmenting human-rated datasets, with promising results obtained by generating ratings for single words. Yet, performance for ratings of complex items, i.e., metaphors, is still unexplored. Here, we present the first assessment of the validity and reliability of ratings of metaphors on familiarity, comprehensibility, and imageability, generated by three GPT models for a total of 687 items gathered from the Italian Figurative Archive and three English studies. We performed a thorough validation in terms of both alignment with human data and ability to predict behavioral and electrophysiological responses. We found that machine-generated ratings positively correlated with human-generated ones. Familiarity ratings reached moderate-to-strong correlations for both English and Italian metaphors, although correlations weakened for metaphors with high sensorimotor load. Imageability showed moderate correlations in English and moderate-to-strong in Italian. Comprehensibility for English metaphors exhibited the strongest correlations. Overall, larger models outperformed smaller ones and greater human-model misalignment emerged with familiarity and imageability. Machine-generated ratings significantly predicted response times and the EEG amplitude, with a strength comparable to human ratings. Moreover, GPT ratings obtained across independent sessions were highly stable. We conclude that GPT, especially larger models, can validly and reliably replace - or augment - human subjects in rating metaphor properties. Yet, LLMs align worse with humans when dealing with conventionality and multimodal aspects of metaphorical meaning, calling for careful consideration of the nature of stimuli.","authors":["Veronica Mangiaterra","Hamad Al-Azary","Chiara Barattieri di San Pietro","Paolo Canal","Valentina Bambini"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-12-13","first_seen":"2025-12-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.12444","pdf_url":"https://arxiv.org/pdf/2512.12444","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","隐喻评分","效度信度"],"reason":"用GPT替代人类评分员生成隐喻评分，属于替代人工标注，非仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":42,"question":"GPT能否有效且可靠地替代人类评分员，为隐喻的熟悉度、可理解度和意象性生成评分？","design":"本研究并非仿真人类被试行为，而是用GPT-3.5、GPT-4等三个模型对687条英意隐喻直接生成熟悉度、可理解度、意象性评分，并与人类评分及行为/脑电数据比较。","baseline":"人类评分来自意大利比喻档案库及三项英语研究，并包含行为反应时和脑电振幅数据。","findings":"机器评分与人类评分呈正相关，熟悉度达中到强相关，可理解度相关最强；较大模型表现更优，且机器评分能显著预测反应时和脑电振幅，跨会话稳定性高。但在高感觉运动负荷隐喻上相关性减弱，熟悉度和意象性的人机偏差较大。","reliability":"论文指出LLM在处理隐喻的规约性和多模态意义时与人类对齐较差，且对高感觉运动负荷的隐喻评分相关性下降，提示刺激性质会影响替代有效性。","relevance":"该研究严格评估了LLM替代人类评分的效度与局限，提供了真实人类行为与神经数据对照，对关注LLM仿真可靠性的研究者有参考价值，但并非直接仿真人类被试决策或行为。","inspiration":"可借鉴其多维度评分效度验证框架（相关分析、预测行为/神经指标、跨会话稳定性）来评估LLM生成标注的可靠性。｜可迁移到经济金融文本分析场景，如用LLM对财经新闻或政策声明进行情绪、不确定性、可读性等主观维度评分。｜可设计让GPT对央行声明生成“政策不确定性”评分，以人类专家评分和后续市场波动数据为基准，检验其预测效度与稳定性。"}},{"id":"2512.11943","version":1,"title":"How AI Agents Follow the Herd of AI? Network Effects, History, and Machine Optimism","zh_title":"AI智能体如何跟随AI的羊群效应？网络效应、历史与机器乐观主义","abstract":"Understanding decision-making in multi-AI-agent frameworks is crucial for analyzing strategic interactions in network-effect-driven contexts. This study investigates how AI agents navigate network-effect games, where individual payoffs depend on peer participatio--a context underexplored in multi-agent systems despite its real-world prevalence. We introduce a novel workflow design using large language model (LLM)-based agents in repeated decision-making scenarios, systematically manipulating price trajectories (fixed, ascending, descending, random) and network-effect strength. Our key findings include: First, without historical data, agents fail to infer equilibrium. Second, ordered historical sequences (e.g., escalating prices) enable partial convergence under weak network effects but strong effects trigger persistent \"AI optimism\"--agents overestimate participation despite contradictory evidence. Third, randomized history disrupts convergence entirely, demonstrating that temporal coherence in data shapes LLMs' reasoning, unlike humans. These results highlight a paradigm shift: in AI-mediated systems, equilibrium outcomes depend not just on incentives, but on how history is curated, which is impossible for human.","authors":["Yu Liu","Wenwen Li","Yifan Dou","Guangnan Ye"],"categories":["cs.MA","cs.AI","econ.GN"],"primary_category":"cs.MA","announce_type":"new","date":"2025-12-12","first_seen":"2025-12-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.11943","pdf_url":"https://arxiv.org/pdf/2512.11943","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","网络效应博弈","社会模拟"],"reason":"用LLM agent模拟网络效应博弈，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":47,"question":"在重复网络效应博弈中，LLM智能体如何利用历史信息形成参与预期，历史数据的呈现方式如何影响集体均衡收敛？","design":"用Qwen系列LLM扮演6名学者，在重复会议参与博弈中决策是否出席；系统操纵价格轨迹（固定、上升、下降、随机）和网络效应强度，测量参与人数与均衡收敛情况。","baseline":"无对照","findings":"无历史数据时，智能体无法推断均衡；有序历史序列在弱网络效应下可部分收敛，但强网络效应引发“AI乐观”导致持续高估参与；随机历史完全破坏收敛。","reliability":"论文未讨论","relevance":"该研究用LLM模拟网络效应博弈中的决策，但缺乏真实人类数据对照，属于社会模拟边界情形，对关注人类基准对照的研究者参考价值有限。","inspiration":"可借鉴历史信息注入的操控方式，通过改变历史轨迹的时序结构来检验智能体学习模式。｜可迁移至金融市场中的协调博弈场景，如银行挤兑或资产抛售中的预期形成。｜用LLM扮演投资者，在重复投资博弈中操控历史价格序列（趋势、反转、随机），观察其撤资决策，并与历史金融危机中的真实投资者行为数据对照。"}},{"id":"2512.08345","version":2,"title":"The High Cost of Incivility: Quantifying Interaction Inefficiency via Multi-Agent Monte Carlo Simulations","zh_title":"不文明行为的高昂代价：通过多智能体蒙特卡洛模拟量化互动低效","abstract":"Workplace toxicity is widely recognized as detrimental to organizational culture, yet quantifying its direct impact on operational efficiency remains methodologically challenging due to the ethical and practical difficulties of reproducing conflict in human subjects. This study leverages Large Language Model (LLM) based Multi-Agent Systems to simulate 1-on-1 adversarial debates, creating a controlled \"sociological sandbox\". We employ a Monte Carlo method to simulate hundrets of discussions, measuring the convergence time (defined as the number of arguments required to reach a conclusion) between a baseline control group and treatment groups involving agents with \"toxic\" system prompts. Our results demonstrate a statistically significant increase of approximately 25\\% in the duration of conversations involving toxic participants. We propose that this \"latency of toxicity\" serves as a proxy for financial damage in corporate and academic settings. Furthermore, we demonstrate that agent-based modeling provides a reproducible, ethical alternative to human-subject research for measuring the mechanics of social friction.","authors":["Benedikt Mangold"],"categories":["cs.AI","cs.CL","cs.CY","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2025-12-09","first_seen":"2025-12-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.08345","pdf_url":"https://arxiv.org/pdf/2512.08345","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体模拟","社会摩擦","LLM仿真"],"reason":"用LLM agent模拟社会互动但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":148,"question":"在受控的对抗性辩论中，毒性行为是否会导致对话收敛所需轮次显著增加？","design":"使用基于LLM的多智能体系统模拟1对1辩论，通过蒙特卡洛方法运行数百次模拟；处理组随机给一名智能体赋予“毒性”系统提示，对照组均为中性提示，测量辩论结束所需的论据轮次。","baseline":"无对照","findings":"毒性参与者的辩论轮次平均增加约25%，差异具有统计显著性；这种“毒性延迟”可作为职场和学术环境中财务损失的代理指标。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟社会摩擦，但缺乏真实人类数据对照，属于边界情形；若关注仿真方法本身或毒性对效率的量化，可读原文了解实验设计细节。","inspiration":"该方法通过系统提示注入来操控智能体行为，并利用蒙特卡洛模拟量化处理效应，为经济学实验中的干预设计提供了可复现的范式｜可迁移至谈判博弈或劳动经济学中的职场冲突场景，研究负面沟通风格对合作效率与产出分配的影响｜以LLM智能体模拟劳资谈判，处理组赋予工会代表“对抗性”提示，对照组为中性，结果变量为达成协议所需轮次与最终工资分配，可对照真实劳资谈判实验或现场数据校准仿真偏差"}},{"id":"2512.04988","version":2,"title":"When AI Agents Compete for Jobs: Strategic Capabilities and Economic Dynamics of AI Labour Markets","zh_title":"当AI智能体竞争工作：AI劳动力市场的战略能力与经济动态","abstract":"Emerging agentic marketplaces provide the economic infrastructure for matching and coordinating the large amounts of AI agents used in agentic swarms. Unlike human workers, AI agents can operate on multiple jobs simultaneously, acquire skills rapidly, and labor without wage floors. These differences introduce a new segment of $\\textbf{AI labor markets}$, where AI agents interact with each other at a much higher frequency than human markets. Yet we lack frameworks to understand how such markets behave in light of economic forces that shape labor markets, such as adverse selection and reputation dynamics. To explore this, we introduce $\\texttt{AI-Work}$, a tractable, simulated gig economy where Large Language Model (LLM) agents compete for jobs, develop skills, and adapt their strategies under uncertainty and competitive pressure. Our experiments examine three domains of capabilities that successful agents possess: $\\textbf{metacognition}$ (accurate self-assessment of skills), $\\textbf{competitive awareness}$ (modeling rivals and market dynamics), and $\\textbf{long-horizon strategic planning}$. Agents with these capabilities consistently achieve higher profits, market share, and stronger adaptation than competing agents. Through $\\texttt{AI-Work}$, we hope to provide a foundation to explore the microeconomic properties of AI-only labor markets, and a conceptual framework to study the strategic reasoning capabilities of participating AI agents.","authors":["Christopher Chiu","Simpson Zhang","Mihaela van der Schaar"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"new","date":"2025-12-04","first_seen":"2025-12-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.04988","pdf_url":"https://arxiv.org/pdf/2512.04988","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI劳动力市场","多智能体模拟","经济仿真"],"reason":"LLM agent模拟零工经济市场，但无真实人类数据对照，属纯理论演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":188,"question":"在AI劳动力市场中，LLM智能体需要哪些战略推理能力才能在竞争压力和信息不对称下取得成功？","design":"构建了一个名为AI-Work的模拟零工经济环境，让LLM智能体扮演工人，通过出价竞标或训练投资来竞争微任务，测量其利润、市场份额和适应能力，并考察元认知、竞争意识和长期规划三种能力的影响。","baseline":"无对照","findings":"具备元认知、竞争意识和长期规划能力的智能体在利润、市场份额和适应性上均优于基线智能体；平台规则（如价格披露、合同形式）会改变均衡行为，AI特有的并发性会放大市场集中度，但任务多样性可缓解此效应。","reliability":"论文未讨论","relevance":"该研究用LLM模拟劳动力市场，但缺乏真实人类数据对照，属于纯理论演示，不符合研究者对实证基准和批判性评估的需求，不建议优先阅读。","inspiration":"该方法通过消融实验系统操纵智能体的元认知、竞争意识和长期规划能力，清晰分离不同战略推理模块对市场结果的因果效应，值得借鉴｜可迁移到劳动力市场政策评估，如最低工资或信息透明政策对就业和收入分配的影响｜设计一个零工平台仿真，让LLM智能体扮演工人，处理组赋予价格披露规则，控制组无披露，结果变量为个体收入和市场集中度，对照真实平台（如MTurk）的准实验数据"}},{"id":"2512.03568","version":1,"title":"Synthetic Cognitive Walkthrough: Aligning Large Language Model Performance with Human Cognitive Walkthrough","zh_title":"合成认知走查：将大语言模型表现与人类认知走查对齐","abstract":"Conducting usability testing like cognitive walkthrough (CW) can be costly. Recent developments in large language models (LLMs), with visual reasoning and UI navigation capabilities, present opportunities to automate CW. We explored whether LLMs (GPT-4 and Gemini-2.5-pro) can simulate human behavior in CW by comparing their walkthroughs with human participants. While LLMs could navigate interfaces and provide reasonable rationales, their behavior differed from humans. LLM-prompted CW achieved higher task completion rates than humans and followed more optimal navigation paths, while identifying fewer potential failure points. However, follow-up studies demonstrated that with additional prompting, LLMs can predict human-identified failure points, aligning their performance with human participants. Our work highlights that while LLMs may not replicate human behaviors exactly, they can be leveraged for scaling usability walkthroughs and providing UI insights, offering a valuable complement to traditional usability testing.","authors":["Ruican Zhong","David W. McDonald","Gary Hsieh"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-12-03","first_seen":"2025-12-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.03568","pdf_url":"https://arxiv.org/pdf/2512.03568","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","认知走查","人机对照"],"reason":"用LLM模拟人类认知走查行为，并与真实人类数据对照，指出差异和失效条件，方法可…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":88,"question":"LLM能否模拟人类在认知走查中的行为，以自动化可用性评估？","design":"用GPT-4和Gemini-2.5-pro扮演认知走查中的评估者，在提示词引导下导航两个移动应用界面并出声思考，测量任务完成率、导航路径和识别的潜在失败点。","baseline":"10名人类评估者在相同应用和任务上的认知走查数据，包括任务完成率、导航步骤和识别的潜在失败点。","findings":"LLM的任务完成率高于人类，导航路径更优，但识别的潜在失败点更少；通过额外提示，LLM能预测人类识别的失败点，从而对齐人类表现。","reliability":"LLM行为与人类不同，人类更倾向广度优先探索和因记忆出错，LLM直接模仿人类行为会失效；但可通过调整提示让LLM预测人类发现的失败点。","relevance":"直接对比LLM与人类在认知走查中的行为差异，并指出仿真失效条件，符合研究者对LLM仿真人类实验的可靠性与偏差的关注，值得精读。","inspiration":"借鉴其通过调整提示词让LLM预测人类特定失败点的对齐方法，可作为校准LLM仿真偏差的技术手段｜可迁移到消费者金融决策中的信息处理偏差研究，如贷款条款理解或投资产品选择中的认知错误｜以LLM模拟消费者阅读贷款合同后识别隐藏费用，处理组加入人类常见认知偏差的提示，结果变量为识别的费用项数量，对照真实消费者实验数据"}},{"id":"2512.11827","version":2,"title":"Assessing Greenspace Attractiveness with ChatGPT, Claude, and Gemini: Do AI Models Reflect Human Perceptions?","zh_title":"用ChatGPT、Claude和Gemini评估绿地吸引力：AI模型是否反映人类感知？","abstract":"Understanding greenspace attractiveness is essential for designing livable and inclusive urban environments, yet existing assessment approaches often overlook informal or transient spaces and remain too resource intensive to capture subjective perceptions at scale. This study examines the ability of multimodal large language models (MLLMs), ChatGPT GPT-4o, Claude 3.5 Haiku, and Gemini 2.0 Flash, to assess greenspace attractiveness similarly to humans using Google Street View imagery. We compared model outputs with responses from a geo-questionnaire of residents in Lodz, Poland, across both formal (for example, parks and managed greenspaces) and informal (for example, meadows and wastelands) greenspaces. Survey respondents and models indicated whether each greenspace was attractive or unattractive and provided up to three free text explanations. Analyses examined how often their attractiveness judgments aligned and compared their explanations after classifying them into shared reasoning categories. Results show high AI human agreement for attractive formal greenspaces and unattractive informal spaces, but low alignment for attractive informal and unattractive formal greenspaces. Models consistently emphasized aesthetic and design oriented features, underrepresenting safety, functional infrastructure, and locally embedded qualities valued by survey respondents. While these findings highlight the potential for scalable pre-assessment, they also underscore the need for human oversight and complementary participatory approaches. We conclude that MLLMs can support, but not replace, context sensitive greenspace evaluation in planning practice.","authors":["Milad Malekzadeh","Magdalena Biernacka","Elias Willberg","Jussi Torkko","Edyta Łaszkiewicz","Tuuli Toivonen"],"categories":["cs.CY","cs.AI","cs.CV"],"primary_category":"cs.CY","announce_type":"new","date":"2025-12-02","first_seen":"2025-12-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.11827","pdf_url":"https://arxiv.org/pdf/2512.11827","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","绿地感知评估","AI与人类对照"],"reason":"用LLM评估绿地吸引力并与人类问卷对照，涉及仿真可靠性、偏差及失效条件，直接相…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":12,"question":"多模态大语言模型（ChatGPT、Claude、Gemini）对城市绿地吸引力的评估是否与人类感知一致？","design":"使用Google街景图像作为输入，让三个MLLM（ChatGPT GPT-4o、Claude 3.5 Haiku、Gemini 2.0 Flash）判断正式和非正式绿地的吸引力（有吸引力/无吸引力），并生成最多三条自由文本解释；将模型输出与波兰罗兹市居民的地理问卷调查结果进行比较。","baseline":"波兰罗兹市居民的地理问卷调查，包含对正式和非正式绿地的吸引力判断及自由文本解释。","findings":"模型在评估有吸引力的正式绿地和无吸引力的非正式绿地时与人类一致性高，但在有吸引力的非正式绿地和无吸引力的正式绿地上一致性低。模型过度强调美学和设计特征，而低估了安全、功能设施和本地化品质。","reliability":"模型未能充分捕捉安全、功能设施和本地化品质等人类重视的特征；在非典型场景（有吸引力的非正式绿地、无吸引力的正式绿地）中一致性低；可能存在隐藏的性别和其他偏见；无法代表不同人口群体的感知差异。","relevance":"该研究直接以真实人类数据为基准，检验LLM在主观感知评估中的可靠性与偏差，并明确指出了仿真失效的条件（如非典型绿地类型、忽视安全与本地化特征），与研究者关注的LLM仿真实验高度相关，值得阅读原文以了解具体实验设计和偏差分析。","inspiration":"借鉴其将模型输出与人类解释进行定性分类比较的方法，可迁移到消费者对金融产品广告的感知评估或投资者对年报文本的情绪解读研究中。｜可应用于行为金融中的信息感知实验，例如研究投资者对上市公司年报中风险披露的吸引力判断。｜以真实投资者问卷调查为基准，让LLM阅读年报摘要并评估其投资吸引力及理由，比较模型与人类在风险感知、语言特征关注上的差异，检验模型是否过度关注表面语言而忽略深层风险信号。"}},{"id":"2512.07890","version":1,"title":"CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models","zh_title":"CrowdLLM：结合生成模型构建基于LLM的数字人群","abstract":"The emergence of large language models (LLMs) has sparked much interest in creating LLM-based digital populations that can be applied to many applications such as social simulation, crowdsourcing, marketing, and recommendation systems. A digital population can reduce the cost of recruiting human participants and alleviate many concerns related to human subject study. However, research has found that most of the existing works rely solely on LLMs and could not sufficiently capture the accuracy and diversity of a real human population. To address this limitation, we propose CrowdLLM that integrates pretrained LLMs and generative models to enhance the diversity and fidelity of the digital population. We conduct theoretical analysis of CrowdLLM regarding its great potential in creating cost-effective, sufficiently representative, scalable digital populations that can match the quality of a real crowd. Comprehensive experiments are also conducted across multiple domains (e.g., crowdsourcing, voting, user rating) and simulation studies which demonstrate that CrowdLLM achieves promising performance in both accuracy and distributional fidelity to human data.","authors":["Ryan Feng Lin","Keyu Tian","Hanming Zheng","Congjing Zhang","Li Zeng","Shuai Huang"],"categories":["cs.MA","cs.AI","cs.LG","stat.ME","stat.ML"],"primary_category":"cs.MA","announce_type":"new","date":"2025-12-02","first_seen":"2025-12-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.07890","pdf_url":"https://arxiv.org/pdf/2512.07890","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","数字人群","人类数据对照"],"reason":"用LLM构建数字人群，模拟众包、投票等人类行为，并与真实人类数据对照，直接命中…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":43,"question":"如何结合预训练大语言模型与生成模型，构建能准确复现真实人群决策多样性与分布保真度的数字人群？","design":"提出CrowdLLM框架，将预训练LLM与生成式机器学习模型集成，通过概率框架生成虚拟参与者并聚合其决策，模拟众包、投票、用户评分等场景中的群体行为。","baseline":"使用真实人类在众包、投票、用户评分等任务上的决策数据作为对照基准。","findings":"CrowdLLM在多个领域实验中，生成的数字人群在决策准确性和分布保真度上均与真实人类数据高度匹配；理论分析表明该框架具有成本效益高、代表性强和可扩展的潜力。","reliability":"论文未讨论","relevance":"该研究直接构建LLM数字人群并对照真实人类数据，评估仿真准确性与分布保真度，覆盖众包、投票等场景，高度契合研究者对LLM人类仿真实验及基准对照的关注，值得精读原文。","inspiration":"该方法通过概率框架集成生成模型来捕获决策多样性，可借鉴用于模拟经济主体异质性｜可迁移到消费者跨期选择实验，研究不同贴现因子分布下的储蓄行为｜以LLM生成虚拟消费者，处理为不同利率条件，结果变量为储蓄金额，对照真实家庭金融调查数据"}},{"id":"2512.01107","version":1,"title":"Foundation Priors","zh_title":"基础先验：将基础模型输出作为结构化主观先验","abstract":"Foundation models, and in particular large language models, can generate highly informative responses, prompting growing interest in using these ''synthetic'' outputs as data in empirical research and decision-making. This paper introduces the idea of a foundation prior, which shows that model-generated outputs are not as real observations, but draws from the foundation prior induced prior predictive distribution. As such synthetic data reflects both the model's learned patterns and the user's subjective priors, expectations, and biases. We model the subjectivity of the generative process by making explicit the dependence of synthetic outputs on the user's anticipated data distribution, the prompt-engineering process, and the trust placed in the foundation model. We derive the foundation prior as an exponential-tilted, generalized Bayesian update of the user's primitive prior, where a trust parameter governs the weight assigned to synthetic data. We then show how synthetic data and the associated foundation prior can be incorporated into standard statistical and econometric workflows, and discuss their use in applications such as refining complex models, informing latent constructs, guiding experimental design, and augmenting random-coefficient and partially linear specifications. By treating generative outputs as structured, explicitly subjective priors rather than as empirical observations, the framework offers a principled way to harness foundation models in empirical work while avoiding the conflation of synthetic ''facts'' with real data.","authors":["Sanjog Misra"],"categories":["cs.AI","econ.EM","stat.ML"],"primary_category":"cs.AI","announce_type":"new","date":"2025-11-30","first_seen":"2025-11-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.01107","pdf_url":"https://arxiv.org/pdf/2512.01107","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["合成数据","贝叶斯先验","方法论"],"reason":"提出将LLM输出作为先验而非观测数据，用于统计推断，属方法论框架，非直接仿真人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":189,"question":"如何将大语言模型等基础模型生成的合成数据作为主观先验（foundation prior）纳入统计推断，而非将其视为客观观测数据？","design":"本文并非直接进行人类仿真实验，而是提出一个方法论框架：将用户通过提示工程与基础模型交互产生的合成数据，建模为一种指数倾斜的广义贝叶斯更新，形成“基础先验”，并讨论其在统计与计量工作流中的应用。","baseline":"无对照","findings":"基础模型输出本质上是主观的，融合了模型学到的模式与用户的先验、预期和偏见；通过将合成数据视为结构化的主观先验而非经验观测，可以在利用其信息丰富性的同时，避免将合成“事实”与真实数据混淆。","reliability":"论文指出合成数据的可靠性受限于生成过程的不透明性、提示工程注入的主观性，以及用户对模型的信任程度，但未提供具体的失效条件实证分析。","relevance":"本文为批判性地审视LLM仿真人类行为提供了理论基础，强调合成数据的主观性，适合关注仿真可靠性与偏差的研究者阅读，但未直接进行人类对照实验。","inspiration":"该方法论将LLM输出视为结构化主观先验而非客观数据，为处理合成数据的主观性提供了新思路，可借鉴其指数倾斜贝叶斯更新框架来校准仿真偏差。｜可迁移到政策公告的预期形成研究，例如分析市场对央行沟通的反应，利用LLM生成不同沟通措辞下的预期分布。｜以LLM作为被试，处理为不同措辞的政策声明，结果变量为生成的预期通胀或利率路径，对照真实市场调查数据或金融市场价格隐含预期。"}},{"id":"2511.21218","version":3,"title":"Can Finetuing LLMs on Small Human Samples Increase Heterogeneity, Alignment, and Belief-Action Coherence?","zh_title":"在小规模人类样本上微调LLM能否增加异质性、对齐度和信念-行动一致性？","abstract":"There is ongoing debate about whether large language models (LLMs) can serve as substitutes for human participants in survey and experimental research. While recent work in fields such as marketing and psychology has explored the potential of LLM-based simulation, a growing body of evidence cautions against this practice: LLMs often fail to align with real human behavior, exhibiting limited diversity, systematic misalignment for minority subgroups, insufficient within-group variance, and discrepancies between stated beliefs and actions. This study examines an important and distinct question in this domain: whether fine-tuning on a small subset of human survey data, such as that obtainable from a pilot study, can mitigate these issues and yield realistic simulated outcomes. Using a behavioral experiment on information disclosure, we compare human and LLM-generated responses across multiple dimensions, including distributional divergence, subgroup alignment, belief-action coherence, and the recovery of regression coefficients. We find that fine-tuning on small human samples substantially improves heterogeneity, alignment, and belief-action coherence relative to the base model. However, even the best-performing fine-tuned models fail to reproduce the regression coefficients of the original study, suggesting that LLM-generated data remain unsuitable for replacing human participants in formal inferential analyses.","authors":["Steven Wang","Kyle Hunt","Shaojie Tang","Kenneth Joseph"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-11-26","first_seen":"2025-11-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.21218","pdf_url":"https://arxiv.org/pdf/2511.21218","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","行为实验","微调偏差"],"reason":"直接研究微调LLM作为人类被试替代，用真实人类实验数据对照，评估仿真可靠性及失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"在小规模人类样本上微调LLM能否提高仿真中的异质性、对齐度与信念-行动一致性？","design":"使用开源LLM（如Llama-3）在Hunt等人关于信息披露的行为实验数据上微调，模拟攻击者的信念与决策，比较基础模型与微调模型在分布差异、子群对齐、信念-行动一致性及回归系数恢复上的表现。","baseline":"Hunt等人收集的真实人类行为实验数据，包含被试对安全技术部署的信念和攻击决策。","findings":"微调显著改善了异质性、分布对齐和信念-行动一致性，但即使最佳微调模型也无法复现原始研究的回归系数，表明LLM生成数据仍不适合替代人类进行推断性分析。","reliability":"微调模型无法恢复原始回归系数和假设检验结果，在正式推断分析中可能引入不可忽视的偏差；模型泛化性有限，可能不适用于分布外行为情境。","relevance":"直接探讨用微调LLM替代人类被试的可行性与局限，以真实人类实验为基准，评估仿真可靠性及失效条件，高度契合研究者对经济学实验和政策评估场景的关注，值得精读原文。","inspiration":"该方法通过在小规模人类样本上微调LLM来提升仿真异质性和分布对齐，可借鉴其微调策略和以真实人类实验为基准的对照设计｜可迁移到政策公告的预期形成实验，例如研究央行沟通对公众通胀预期的影响｜使用Llama-3在真实调查数据上微调，模拟公众对通胀公告的预期更新，处理为不同措辞的政策声明，结果变量为预期通胀率，以密歇根大学消费者调查数据作为人类基准对照"}},{"id":"2511.08785","version":1,"title":"Making Talk Cheap: Generative AI and Labor Market Signaling","zh_title":"让谈话变得廉价：生成式AI与劳动力市场信号传递","abstract":"Large language models (LLMs) like ChatGPT have significantly lowered the cost of producing written content. This paper studies how LLMs, through lowering writing costs, disrupt markets that traditionally relied on writing as a costly signal of quality (e.g., job applications, college essays). Using data from Freelancer.com, a major digital labor platform, we explore the effects of LLMs' disruption of labor market signaling on equilibrium market outcomes. We develop a novel LLM-based measure to quantify the extent to which an application is tailored to a given job posting. Taking the measure to the data, we find that employers have a high willingness to pay for workers with more customized applications in the period before LLMs are introduced, but not after. To isolate and quantify the effect of LLMs' disruption of signaling on equilibrium outcomes, we develop and estimate a structural model of labor market signaling, in which workers invest costly effort to produce noisy signals that predict their ability in equilibrium. We use the estimated model to simulate a counterfactual equilibrium in which LLMs render written applications useless in signaling workers' ability. Without costly signaling, employers are less able to identify high-ability workers, causing the market to become significantly less meritocratic: compared to the pre-LLM equilibrium, workers in the top quintile of the ability distribution are hired 19% less often, workers in the bottom quintile are hired 14% more often.","authors":["Anais Galdin","Jesse Silbert"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-11-11","first_seen":"2025-11-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.08785","pdf_url":"https://arxiv.org/pdf/2511.08785","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["劳动力市场","信号传递","结构模型"],"reason":"用结构模型模拟LLM影响信号传递，但无LLM直接仿真人类被试，属社会模拟无人类…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":190,"question":"大语言模型如何通过降低写作成本来破坏劳动力市场中书面申请作为能力信号的作用，并影响均衡雇佣结果？","design":"本研究并非用LLM直接仿真人类被试，而是构建了一个结构模型（结合Spence信号模型、离散选择需求模型和评分拍卖），利用Freelancer.com的真实行为数据估计模型参数，然后通过反事实模拟，将LLM的影响设定为将写作成本降为零，从而消除信号传递，观察均衡状态下雇佣分布的变化。","baseline":"使用Freelancer.com平台在LLM大规模采用之前（2023年4月前）的真实雇主-雇员交互数据，包括申请文本、点击流、出价和雇佣结果，作为信号有效期的基准。","findings":"在LLM普及前，雇主愿意为更定制化的申请支付更高工资，且信号能预测工人能力和任务完成情况；LLM普及后，这些模式显著减弱或消失。反事实模拟显示，若信号完全失效，市场精英程度下降：能力前20%的工人雇佣率降低19%，后20%的工人雇佣率提高14%。","reliability":"论文未讨论","relevance":"高度相关。该研究利用真实人类数据作为基准，通过结构模型模拟LLM对信号传递的破坏，属于经济学实验场景下的人类行为仿真评估，直接回应了研究者对LLM仿真可靠性及政策评估的兴趣，值得精读原文。","inspiration":"借鉴其将LLM影响参数化为成本冲击并嵌入结构模型进行反事实模拟的做法，可精确量化技术对市场均衡的因果效应｜可迁移至金融分析师报告的信息价值研究，考察LLM降低报告撰写成本后，报告内容对股价预测的信号作用是否减弱｜以金融分析师为被试，处理组使用LLM辅助撰写报告，对照组独立撰写，结果变量为报告发布后的股价反应，以LLM普及前的历史分析师报告与股价数据作为真实基准对照"}},{"id":"2511.06260","version":1,"title":"LLM-Guided Reinforcement Learning with Representative Agents for Traffic Modeling","zh_title":"基于代表性智能体的LLM引导强化学习交通建模","abstract":"Large language models (LLMs) are increasingly used as behavioral proxies for self-interested travelers in agent-based traffic models. Although more flexible and generalizable than conventional models, the practical use of these approaches remains limited by scalability due to the cost of calling one LLM for every traveler. Moreover, it has been found that LLM agents often make opaque choices and produce unstable day-to-day dynamics. To address these challenges, we propose to model each homogeneous traveler group facing the same decision context with a single representative LLM agent who behaves like the population's average, maintaining and updating a mixed strategy over routes that coincides with the group's aggregate flow proportions. Each day, the LLM reviews the travel experience and flags routes with positive reinforcement that they hope to use more often, and an interpretable update rule then converts this judgment into strategy adjustments using a tunable (progressively decaying) step size. The representative-agent design improves scalability, while the separation of reasoning from updating clarifies the decision logic while stabilizing learning. In classic traffic assignment settings, we find that the proposed approach converges rapidly to the user equilibrium. In richer settings with income heterogeneity, multi-criteria costs, and multi-modal choices, the generated dynamics remain stable and interpretable, reproducing plausible behavioral patterns well-documented in psychology and economics, for example, the decoy effect in toll versus non-toll road selection, and higher willingness-to-pay for convenience among higher-income travelers when choosing between driving, transit, and park-and-ride options.","authors":["Hanlin Sun","Jiayang Li"],"categories":["cs.GT","cs.AI","eess.SY"],"primary_category":"cs.GT","announce_type":"new","date":"2025-11-09","first_seen":"2025-11-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.06260","pdf_url":"https://arxiv.org/pdf/2511.06260","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","交通行为建模","多智能体系统"],"reason":"用LLM代理群体模拟交通选择行为，涉及经济学场景，并讨论仿真稳定性与失效条件，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":105,"question":"如何利用大语言模型（LLM）作为代表性智能体，在交通出行选择场景中实现可扩展且稳定的日间动态学习，并复现真实人类行为模式？","design":"提出用单个代表性LLM智能体模拟同质出行者群体的平均行为，每天LLM根据自然语言描述的出行体验给出正强化标记，再由可解释的更新规则（含可调衰减步长）调整混合策略，最终输出群体路径流量比例。实验在经典交通分配、多收入群体、多准则成本及多模式选择等场景下测试收敛性与行为模式。","baseline":"无对照","findings":"在经典交通分配中，该方法快速收敛至用户均衡；在包含收入异质性、多准则成本和多模式选择的丰富场景下，动态保持稳定且可解释，并成功复现了心理学与经济学中记载的诱饵效应及高收入群体对便利性的更高支付意愿。","reliability":"论文未讨论","relevance":"该研究用LLM代理群体模拟交通选择行为，涉及经济学场景（如诱饵效应、支付意愿），并讨论仿真稳定性，但未提供真实人类数据对照，适合关注LLM仿真行为复现能力与可扩展性的研究者阅读原文。","inspiration":"该方法用单个代表性LLM智能体模拟群体平均行为，并通过自然语言反馈和可解释更新规则动态调整策略，为仿真实验提供了可扩展且稳定的处理施加方式｜可迁移到消费者跨期选择实验，研究不同收入群体的时间偏好与支付意愿差异｜以LLM作为被试，用自然语言描述不同收入场景和跨期选择任务作为处理，结果变量为选择的时间折扣率，对照真实消费者调查数据（如CFPS或CHFS）中的跨期选择行为分布"}},{"id":"2511.05766","version":1,"title":"Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs","zh_title":"机器中的锚定：LLM中锚定偏差的行为与归因证据","abstract":"Large language models (LLMs) are increasingly examined as both behavioral subjects and decision systems, yet it remains unclear whether observed cognitive biases reflect surface imitation or deeper probability shifts. Anchoring bias, a classic human judgment bias, offers a critical test case. While prior work shows LLMs exhibit anchoring, most evidence relies on surface-level outputs, leaving internal mechanisms and attributional contributions unexplored. This paper advances the study of anchoring in LLMs through three contributions: (1) a log-probability-based behavioral analysis showing that anchors shift entire output distributions, with controls for training-data contamination; (2) exact Shapley-value attribution over structured prompt fields to quantify anchor influence on model log-probabilities; and (3) a unified Anchoring Bias Sensitivity Score integrating behavioral and attributional evidence across six open-source models. Results reveal robust anchoring effects in Gemma-2B, Phi-2, and Llama-2-7B, with attribution signaling that the anchors influence reweighting. Smaller models such as GPT-2, Falcon-RW-1B, and GPT-Neo-125M show variability, suggesting scale may modulate sensitivity. Attributional effects, however, vary across prompt designs, underscoring fragility in treating LLMs as human substitutes. The findings demonstrate that anchoring bias in LLMs is robust, measurable, and interpretable, while highlighting risks in applied domains. More broadly, the framework bridges behavioral science, LLM safety, and interpretability, offering a reproducible path for evaluating other cognitive biases in LLMs.","authors":["Felipe Valencia-Clavijo"],"categories":["cs.AI","cs.CL","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2025-11-07","first_seen":"2025-11-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.05766","pdf_url":"https://arxiv.org/pdf/2511.05766","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B4"],"tags":["锚定偏差","LLM行为仿真","可解释性"],"reason":"研究LLM的锚定偏差，评估其作为人类替代品的可靠性，并指出仿真脆弱性，方法可迁…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":58,"question":"大语言模型表现出的锚定偏差是表面模仿还是深层概率偏移？","design":"以六个开源LLM（GPT-2、GPT-Neo-125M、Falcon-RW-1B、Gemma-2B、Phi-2、Llama-2-7B）为被试，通过结构化提示（含高/低锚数字）复现经典锚定实验，测量候选答案的对数概率分布变化，并用Shapley值归因锚字段对对数概率的贡献。","baseline":"复现Tversky和Kahneman的“非洲国家在联合国占比”锚定实验，以人类在该实验中的锚定效应作为对照基准。","findings":"Gemma-2B、Phi-2和Llama-2-7B表现出稳健的锚定效应，锚数字导致整个输出分布偏移；归因分析显示锚字段对模型对数概率有显著贡献，但效应因提示设计而异，表明将LLM作为人类替代品存在脆弱性。","reliability":"归因效应随提示设计变化，显示将LLM视为人类替代品的脆弱性；小模型（如GPT-2、Falcon-RW-1B、GPT-Neo-125M）结果不稳定，暗示模型规模可能调节敏感性。","relevance":"该研究直接评估LLM作为人类被试替代品的可靠性，通过行为与归因双重证据揭示锚定偏差的深层机制与失效条件，方法可迁移至其他认知偏差，对经济学实验和政策评估中的仿真应用具有批判性参考价值，值得精读。","inspiration":"该方法通过结构化提示复现经典锚定实验，并用对数概率分布偏移和Shapley值归因来区分表面模仿与深层概率偏移，提供了行为与归因双重证据的稳健性检验思路｜可迁移到资产定价实验，研究投资者在估值时受历史价格或分析师目标价锚定的影响｜以LLM为被试，在估值提示中嵌入高/低历史价格锚，测量估值输出的对数概率分布偏移，并以真实投资者估值数据或实验数据作为对照基准"}},{"id":"2511.03758","version":3,"title":"Leveraging LLM-based agents for social science research: insights from citation network simulations","zh_title":"利用基于大语言模型的智能体进行社会科学研究：来自引文网络模拟的见解","abstract":"The emergence of Large Language Models (LLMs) demonstrates their potential to encapsulate the logic and patterns inherent in human behavior simulation by leveraging extensive web data pre-training. However, the boundaries of LLM capabilities in social simulation remain unclear. To further explore the social attributes of LLMs, we introduce the CiteAgent framework, designed to generate citation networks based on human-behavior simulation with LLM-based agents. CiteAgent successfully captures predominant phenomena in real-world citation networks, including power-law distribution, citational distortion, and shrinking diameter. Building on this realistic simulation, we establish two LLM-based research paradigms in social science: LLM-SE (LLM-based Survey Experiment) and LLM-LE (LLM-based Laboratory Experiment). These paradigms facilitate rigorous analyses of citation network phenomena, allowing us to validate and challenge existing theories. Additionally, we extend the research scope of traditional science of science studies through idealized social experiments, with the simulation experiment results providing valuable insights for real-world academic environments. Our work demonstrates the potential of LLMs for advancing science of science research in social science.","authors":["Jiarui Ji","Runlin Lei","Xuchen Pan","Zhewei Wei","Hao Sun","Yankai Lin","Xu Chen","Yongzheng Yang","Yaliang Li","Bolin Ding","Ji-Rong Wen"],"categories":["physics.soc-ph","cs.AI","cs.CY","cs.MA","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2025-11-05","first_seen":"2025-11-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.03758","pdf_url":"https://arxiv.org/pdf/2511.03758","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","引文网络","社会科学实验"],"reason":"用LLM agent模拟引文网络并与真实数据对照，提出LLM-SE/LE范式，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":15,"question":"LLM代理能否在引文网络模拟中复现真实网络的关键现象，并用于社会科学研究？","design":"构建CiteAgent框架，用GPT-3.5、GPT-4o-mini、LLAMA-3-70B扮演作者，基于种子网络生成引文网络，测量度分布幂律拟合度；并通过LLM-SE调查实验和LLM-LE实验室实验操纵论文属性与推荐算法，分析引用选择的影响因素。","baseline":"对照CiteSeer和Cora真实引文网络数据集，验证生成网络的幂律分布、引用扭曲和直径收缩现象。","findings":"CiteAgent生成的引文网络能复现幂律分布等真实网络现象，但GPT-3.5拟合较差；LLM-SE和LLM-LE实验表明，引用相关属性（如论文被引量、作者被引量、时效性）是导致优先连接和幂律分布的主要因素。","reliability":"论文指出GPT-3.5在LLM-Agent数据集上无法完美拟合幂律分布，表明不同LLM的仿真能力存在差异；但未系统讨论其他失效条件或局限。","relevance":"该研究直接以真实引文网络为基准，用LLM代理复现社会现象并分析机制，提出了LLM-SE和LLM-LE范式，高度契合您对LLM人类仿真实验、经济学/政策评估场景及可靠性评估的关注，值得精读。","inspiration":"该研究通过LLM代理模拟引文网络生成，并操纵论文属性（如被引量、时效性）和推荐算法作为处理，以真实引文网络为基准测量度分布拟合度，提供了因果推断与仿真验证结合的设计范例。｜可迁移至金融市场信息扩散研究，例如模拟分析师报告或新闻如何通过投资者关注网络传播并影响资产价格。｜以LLM扮演投资者，处理为信息源的突出特征（如来源权威性、时效性），结果变量为投资者关注分配与价格波动，对照真实股票论坛引用网络或价格联动数据。"}},{"id":"2511.02458","version":1,"title":"Prompting for Policy: Forecasting Macroeconomic Scenarios with Synthetic LLM Personas","zh_title":"用合成LLM角色预测宏观经济情景的政策提示","abstract":"We evaluate whether persona-based prompting improves Large Language Model (LLM) performance on macroeconomic forecasting tasks. Using 2,368 economics-related personas from the PersonaHub corpus, we prompt GPT-4o to replicate the ECB Survey of Professional Forecasters across 50 quarterly rounds (2013-2025). We compare the persona-prompted forecasts against the human experts panel, across four target variables (HICP, core HICP, GDP growth, unemployment) and four forecast horizons. We also compare the results against 100 baseline forecasts without persona descriptions to isolate its effect. We report two main findings. Firstly, GPT-4o and human forecasters achieve remarkably similar accuracy levels, with differences that are statistically significant yet practically modest. Our out-of-sample evaluation on 2024-2025 data demonstrates that GPT-4o can maintain competitive forecasting performance on unseen events, though with notable differences compared to the in-sample period. Secondly, our ablation experiment reveals no measurable forecasting advantage from persona descriptions, suggesting these prompt components can be omitted to reduce computational costs without sacrificing accuracy. Our results provide evidence that GPT-4o can achieve competitive forecasting accuracy even on out-of-sample macroeconomic events, if provided with relevant context data, while revealing that diverse prompts produce remarkably homogeneous forecasts compared to human panels.","authors":["Giulia Iadisernia","Carolina Camassa"],"categories":["cs.CL","cs.CE","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2025-11-04","first_seen":"2025-11-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.02458","pdf_url":"https://arxiv.org/pdf/2511.02458","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","宏观经济预测","人类数据对照"],"reason":"用LLM persona模拟专业预测者，复现人类预测行为，并与真实专家数据对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"基于角色的提示（persona-based prompting）能否提升大语言模型在宏观经济预测任务中的表现？","design":"使用GPT-4o模型，从PersonaHub语料库中筛选出2368个经济学相关角色，模拟欧洲央行专业预测者调查（ECB-SPF）的50个季度预测（2013-2025年），比较角色提示与无角色基线提示在四个宏观经济变量（HICP、核心HICP、GDP增长、失业率）和四个预测期上的预测准确性。","baseline":"欧洲央行专业预测者调查（ECB-SPF）的真实人类专家小组预测数据，以及100个无角色描述的基线GPT-4o预测。","findings":"GPT-4o与人类预测者的准确性非常接近，差异虽统计显著但实际幅度不大；角色描述对预测准确性没有可测量的提升，可省略以降低计算成本而不牺牲精度。","reliability":"样本外（2024-2025）预测表现与样本内存在明显差异；角色提示并未带来预测优势；不同提示产生的预测高度同质，与人类小组的多样性形成对比。","relevance":"该研究直接使用LLM角色模拟专业预测者，并与真实人类专家面板进行对照，评估了角色提示的效用与局限性，完全符合研究者对LLM仿真人类行为、基准对照和失效条件分析的兴趣，值得阅读原文。","inspiration":"该方法借鉴了用真实专家调查面板（ECB-SPF）作为人类基准，直接比较LLM角色提示与无角色基线的预测准确性，并检验样本外表现与预测多样性｜可迁移到政策公告的预期形成研究，例如模拟市场参与者对央行利率决议的即时反应与预期调整｜设计雏形：以LLM角色模拟金融分析师，处理为是否提供角色描述，结果变量为对利率决议的预测误差，对照真实分析师调查数据（如Blue Chip Economic Indicators）"}},{"id":"2510.26727","version":3,"title":"Neither Consent nor Property: A Policy Lab for Data Law","zh_title":"非同意非财产：数据法律的政策实验室","abstract":"Regulators currently govern the AI data economy based on intuition rather than evidence, struggling to choose between inconsistent regimes of informed consent, immunity, and liability. To fill this policy vacuum, this paper develops a novel computational policy laboratory: a spatially explicit Agent-Based Model (ABM) of the data market. To solve the problem of missing data, we introduce a two-stage methodological pipeline. First, we translate decision rules from multi-year fieldwork (2022-2025) into agent constraints. This ensures the model reflects actual bargaining frictions rather than theoretical abstractions. Second, we deploy Large Language Models (LLMs) as \"subjects\" in a Discrete Choice Experiment (DCE). This novel approach recovers precise preference primitives, such as willingness-to-pay elasticities, which are empirically unobservable in the wild. Calibrated by these inputs, our model places rival legal institutions side-by-side to simulate their welfare effects. The results challenge the dominant regulatory paradigm. We find that property-rule mechanisms, such as informed consent, fail to maximize welfare. Counterintuitively, social welfare peaks when liability for substantive harm is shifted to the downstream buyer. This aligns with the \"least cost avoider\" principle, because downstream users control post-acquisition safeguards, they are best positioned to mitigate risk efficiently. By \"de-romanticizing\" seller-centric frameworks, this paper provides an economic justification for emerging doctrines of downstream reachability.","authors":["Haoyi Zhang","Tianyi Zhu"],"categories":["econ.GN","cs.CY"],"primary_category":"econ.GN","announce_type":"new","date":"2025-10-30","first_seen":"2025-10-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.26727","pdf_url":"https://arxiv.org/pdf/2510.26727","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","政策实验","数据市场"],"reason":"用LLM作为离散选择实验的被试，模拟数据市场，但无真实人类数据对照，属社会模拟…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":191,"question":"在数据市场缺乏交易微观数据的情况下，如何构建一个基于主体的模型来模拟不同法律制度（知情同意、责任豁免、下游责任）对市场参与和社会福利的影响？","design":"构建空间显式的基于主体模型（ABM），将中国数据市场映射到14,526个六边形网格上，医院为卖方、AI公司为买方，进行去中心化双边议价。模型规则来自2022-2025年多地点田野调查，偏好参数（如支付意愿弹性）通过将大语言模型（LLM）作为离散选择实验（DCE）的被试来校准。模拟中改变责任分配制度（卖方知情同意、下游买方责任等），测量总福利（消费者剩余+生产者剩余-未定价外部性）。","baseline":"无对照","findings":"基于知情同意的财产规则机制未能最大化福利；将实质性损害责任转移给下游买方时社会福利达到峰值，这符合“最小成本避免者”原则，因为下游用户控制获取后的安全措施，能最有效地降低风险。","reliability":"论文未讨论","relevance":"该研究用LLM作为DCE被试来校准ABM，属于LLM人类仿真，但无真实人类数据对照，且场景为数据市场监管而非典型经济学实验或政策评估，与研究者关注的复现人类行为基准和批判性评估有偏差，但方法新颖，值得快速浏览其LLM校准与验证部分。","inspiration":"该方法将LLM作为离散选择实验被试来校准ABM中的偏好参数，为缺乏微观数据时的参数估计提供了新途径｜可迁移到消费者金融产品选择或投资者风险偏好实验中，用于校准异质性偏好参数｜设计一个退休储蓄计划选择实验，用LLM模拟不同年龄和收入群体的离散选择，以校准ABM中的时间偏好和风险厌恶参数，并与真实家庭金融调查数据（如SCF）中的选择分布进行对照"}},{"id":"2512.08939","version":1,"title":"Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust","zh_title":"评估LLM驱动的数字孪生在模拟医疗系统信任中的人类相似性","abstract":"Serving as an emerging and powerful tool, Large Language Model (LLM)-driven Human Digital Twins are showing great potential in healthcare system research. However, its actual simulation ability for complex human psychological traits, such as distrust in the healthcare system, remains unclear. This research gap particularly impacts health professionals' trust and usage of LLM-based Artificial Intelligence (AI) systems in assisting their routine work. In this study, based on the Twin-2K-500 dataset, we systematically evaluated the simulation results of the LLM-driven human digital twin using the Health Care System Distrust Scale (HCSDS) with an established human-subject sample, analyzing item-level distributions, summary statistics, and demographic subgroup patterns. Results showed that the simulated responses by the digital twin were significantly more centralized with lower variance and had fewer selections of extreme options (all p<0.001). While the digital twin broadly reproduces human results in major demographic patterns, such as age and gender, it exhibits relatively low sensitivity in capturing minor differences in education levels. The LLM-based digital twin simulation has the potential to simulate population trends, but it also presents challenges in making detailed, specific distinctions in subgroups of human beings. This study suggests that the current LLM-driven Digital Twins have limitations in modeling complex human attitudes, which require careful calibration and validation before applying them in inferential analyses or policy simulations in health systems engineering. Future studies are necessary to examine the emotional reasoning mechanism of LLMs before their use, particularly for studies that involve simulations sensitive to social topics, such as human-automation trust.","authors":["Yuzhou Wu","Mingyang Wu","Di Liu","Rong Yin","Kang Li"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2025-10-27","first_seen":"2025-10-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.08939","pdf_url":"https://arxiv.org/pdf/2512.08939","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类数字孪生","医疗系统信任"],"reason":"用LLM数字孪生模拟医疗系统不信任，与真实人类数据对照，评估仿真可靠性并指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"LLM驱动的数字孪生在模拟医疗系统不信任等复杂心理特质时，其类人程度如何？","design":"基于Twin-2K-500数据集构建ChatGPT-4驱动的数字孪生，分层抽样500个样本，使用医疗系统不信任量表（HCSDS）生成模拟回答，并与真实人类样本进行对比。","baseline":"来自Rose等人研究的400名费城陪审员真实人类样本，使用相同的HCSDS量表测量。","findings":"数字孪生的回答显著集中于中间选项，方差更小，极端选项选择更少；虽能大致复现年龄、性别等主要人口学模式，但对教育水平等细微差异的捕捉敏感性较低。","reliability":"当前LLM数字孪生在模拟复杂人类态度时存在局限，回答分布过于集中，对子群体细微差异不敏感，在用于推断分析或政策模拟前需仔细校准和验证。","relevance":"该研究直接评估LLM仿真人类调查回答的可靠性，并与真实人类数据对照，揭示了仿真在分布形态和子群体差异上的失效模式，对关注LLM仿真偏差的研究者具有重要参考价值。","inspiration":"借鉴其分层抽样匹配人口特征、项目级分布对比和卡方检验的验证方法｜可迁移到消费者信任调查或政策态度评估，如模拟不同教育背景人群对金融监管机构的信任｜以LLM生成不同人口特征的虚拟消费者，施加不同政策信息处理，测量对金融系统的信任评分，并以真实消费者调查数据作为对照基准。"}},{"id":"2510.18155","version":1,"title":"LLM-Based Multi-Agent System for Simulating and Analyzing Marketing and Consumer Behavior","zh_title":"基于大语言模型的多智能体系统用于模拟和分析营销与消费者行为","abstract":"Simulating consumer decision-making is vital for designing and evaluating marketing strategies before costly real-world deployment. However, post-event analyses and rule-based agent-based models (ABMs) struggle to capture the complexity of human behavior and social interaction. We introduce an LLM-powered multi-agent simulation framework that models consumer decisions and social dynamics. Building on recent advances in large language model simulation in a sandbox environment, our framework enables generative agents to interact, express internal reasoning, form habits, and make purchasing decisions without predefined rules. In a price-discount marketing scenario, the system delivers actionable strategy-testing outcomes and reveals emergent social patterns beyond the reach of conventional methods. This approach offers marketers a scalable, low-risk tool for pre-implementation testing, reducing reliance on time-intensive post-event evaluations and lowering the risk of underperforming campaigns.","authors":["Man-Lin Chu","Lucian Terhorst","Kadin Reed","Tom Ni","Weiwei Chen","Rongyu Lin"],"categories":["cs.AI","cs.SI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-10-20","first_seen":"2025-10-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.18155","pdf_url":"https://arxiv.org/pdf/2510.18155","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","消费者行为","多智能体"],"reason":"多智能体模拟消费者行为，但无真实人类数据对照，属社会模拟演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":142,"question":"能否用LLM驱动的多智能体系统模拟消费者在价格折扣营销下的决策与社会动态？","design":"基于生成式智能体框架构建虚拟小镇，让LLM智能体拥有记忆、规划、反思和语言交互能力，模拟一周生活，期间炸鸡店在周中提供20%折扣，测量各店铺收入、市场份额和个体消费轨迹。","baseline":"无对照","findings":"折扣使炸鸡店收入增长51%，市场份额从30%升至41%，但总市场规模未扩大，仅发生替代效应；智能体展现出促销驱动和自发的忠诚度模式，部分顾客形成重复购买习惯。","reliability":"论文未讨论","relevance":"该研究属于LLM社会仿真演示，无真实人类基准数据，不符合研究者对对照实验和可靠性批判的核心关切，但可作为方法参考，不建议优先阅读原文。","inspiration":"该方法通过构建虚拟小镇并施加价格折扣处理，测量市场份额和个体消费轨迹，展示了在无真实对照下观察替代效应的设计思路｜可迁移至消费者跨期选择研究，如模拟不同折扣力度对储蓄与即时消费的影响｜以LLM智能体为被试，施加不同利率或折扣率处理，测量消费-储蓄分配，用家庭收支调查微观数据做对照"}},{"id":"2510.16551","version":3,"title":"From Reviews to Actionable Insights: An LLM-Based Approach for Attribute and Feature Extraction","zh_title":"从评论到可操作洞察：基于大语言模型的属性与特征提取方法","abstract":"This research proposes a systematic, large language model (LLM) approach for extracting product and service attributes, features, and associated sentiments from customer reviews. Grounded in marketing theory, the framework distinguishes perceptual attributes from actionable features, producing interpretable and managerially actionable insights. We apply the methodology to 20,000 Yelp reviews of Starbucks stores and evaluate eight prompt variants on a random subset of reviews. Model performance is assessed through agreement with human annotations and predictive validity for customer ratings. Results show high consistency between LLMs and human coders and strong predictive validity, confirming the reliability of the approach. Human coders required a median of six minutes per review, whereas the LLM processed each in two seconds, delivering comparable insights at a scale unattainable through manual coding. Managerially, the analysis identifies attributes and features that most strongly influence customer satisfaction and their associated sentiments, enabling firms to pinpoint \"joy points,\" address \"pain points,\" and design targeted interventions. We demonstrate how structured review data can power an actionable marketing dashboard that tracks sentiment over time and across stores, benchmarks performance, and highlights high-leverage features for improvement. Simulations indicate that enhancing sentiment for key service features could yield 1-2% average revenue gains per store.","authors":["Khaled Boughanmi","Kamel Jedidi","Nour Jedidi"],"categories":["stat.ML","cs.LG","econ.EM"],"primary_category":"stat.ML","announce_type":"new","date":"2025-10-18","first_seen":"2025-10-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.16551","pdf_url":"https://arxiv.org/pdf/2510.16551","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","文本挖掘","营销分析"],"reason":"LLM替代人工标注员提取属性与情感，非仿真人类被试，但有人类标注对照，属边界情…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":192,"question":"如何利用大语言模型从顾客评论中系统提取产品/服务的属性、特征及情感，以生成可操作的管理洞察？","design":"本研究并非仿真人类被试，而是用LLM替代人工标注员，通过三步流水线（探索性生成属性/特征列表、确认性逐句标注、汇总分析）处理Yelp评论，提取属性、特征和情感，并与人类标注一致性和评分预测效度进行对比。","baseline":"人类标注员对随机子集评论进行属性、特征和情感标注，作为LLM性能的对照基准。","findings":"LLM与人类标注员的一致性高，且提取的特征对顾客评分有强预测效度；LLM每篇评论处理仅需2秒，而人工标注中位数为6分钟，实现了大规模高效分析。","reliability":"论文未讨论LLM仿真失效的条件或局限。","relevance":"该研究用LLM替代人工标注，提供了人类对照基准，属于边界相关案例，但并非直接仿真人类被试的决策或行为，对关注经济学实验和政策评估仿真的研究者参考价值有限。","inspiration":"该方法借鉴了LLM替代人工标注的三步流水线设计（探索生成、确认标注、汇总分析），并提供了人类标注一致性作为性能基准。｜可迁移至经济金融文本分析场景，如从央行政策声明或财报电话会议中提取政策意图、风险因子或管理层语调。｜研究设计：以LLM为处理组、人类分析师为对照组，对美联储会议纪要提取政策立场与风险关注点，结果变量为提取特征与人类标注的一致性及对后续市场波动的预测效度，使用历史会议纪要与同期市场数据作为真实对照。"}},{"id":"2510.13091","version":2,"title":"Unmasking Hiring Bias: Platform Data Analysis and Controlled Experiments on Bias in Online Freelance Marketplaces via RAG-LLM Generated Contents","zh_title":"揭示招聘偏见：基于RAG-LLM生成内容的在线自由职业市场偏见平台数据分析与受控实验","abstract":"Online freelance marketplaces, a rapidly growing part of the global labor market, are creating a fair environment where professional skills are the main factor for hiring. While these platforms can reduce bias from traditional hiring, the personal information in user profiles raises concerns about ongoing discrimination. Past studies on this topic have mostly used existing data, which makes it hard to control for other factors and clearly see the effect of things like gender or race. To solve these problems, this paper presents a new method that uses Retrieval-Augmented Generation (RAG) with a Large Language Model (LLM) to create realistic, artificial freelancer profiles for controlled experiments. This approach effectively separates individual factors, enabling a clearer statistical analysis of how different variables influence the freelancer project process. In addition to analyzing extracted data with traditional statistical methods for post-project stage analysis, our research utilizes a dataset with highly controlled variables, generated by an RAG-LLM, to conduct a simulated hiring experiment for pre-project stage analysis. The results of our experiments show that, regarding gender, while no significant preference emerged in initial hiring decisions, female freelancers are substantially more likely to receive imperfect ratings post-project stage. Regarding regional bias, a strong and consistent preference favoring US-based freelancers shows that people are more likely to be selected in the simulated experiments, perceived as more leader-like, and receive higher ratings on the live platform.","authors":["Wugeng Zheng","Guohou Shan"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-10-15","first_seen":"2025-10-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.13091","pdf_url":"https://arxiv.org/pdf/2510.13091","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM生成合成数据","招聘偏见","受控实验"],"reason":"用LLM生成合成档案做受控实验，替代真实简历，属标注/数据生成替代，非直接仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":193,"question":"在线自由职业市场中，性别和地域偏见如何在招聘前阶段（模拟实验）和项目后评价阶段（真实平台数据）中表现？","design":"本研究并非直接仿真人类被试，而是使用RAG-LLM生成高度受控的合成自由职业者档案，通过Amazon Mechanical Turk招募真人参与者进行模拟招聘实验，测量初始雇佣决策中的偏见；同时结合真实平台数据，分析项目后评分中的偏见。","baseline":"真实人类数据来自Freelancer.com上抓取的自由职业者档案、项目评价和评分，用于项目后阶段分析；模拟实验部分无直接人类对照基准，但通过控制变量分离偏见效应。","findings":"模拟实验显示，初始雇佣决策中未发现显著性别偏好，但女性自由职业者在项目后阶段更可能获得不完美评分；地域偏见方面，美国自由职业者在模拟实验中更易被选中，在真实平台上也被认为更具领导力并获得更高评分。","reliability":"论文未讨论","relevance":"该研究利用LLM生成合成档案进行受控实验，替代真实简历以分离变量，属于数据生成替代而非直接仿真人类被试，但涉及偏见测量与真实平台数据对照，对关注LLM在实验设计中应用的研究者有参考价值，建议略读方法部分。","inspiration":"利用RAG-LLM生成高度受控的合成档案以分离偏见变量，并通过真人实验测量决策差异，同时结合真实平台数据对照｜可迁移至信贷审批中的性别或地域歧视研究，例如P2P借贷平台上的贷款决策偏见｜招募真人作为信贷员，处理为合成借款人档案（LLM生成，控制性别/地域），结果变量为贷款批准率与利率，对照真实P2P平台贷款数据中的实际审批与违约率"}},{"id":"2510.13011","version":1,"title":"Deliberate Lab: A Platform for Real-Time Human-AI Social Experiments","zh_title":"Deliberate Lab：一个实时人类-AI社会实验平台","abstract":"Social and behavioral scientists increasingly aim to study how humans interact, collaborate, and make decisions alongside artificial intelligence. However, the experimental infrastructure for such work remains underdeveloped: (1) few platforms support real-time, multi-party studies at scale; (2) most deployments require bespoke engineering, limiting replicability and accessibility, and (3) existing tools do not treat AI agents as first-class participants. We present Deliberate Lab, an open-source platform for large-scale, real-time behavioral experiments that supports both human participants and large language model (LLM)-based agents. We report on a 12-month public deployment of the platform (N=88 experimenters, N=9195 experiment participants), analyzing usage patterns and workflows. Case studies and usage scenarios are aggregated from platform users, complemented by in-depth interviews with select experimenters. By lowering technical barriers and standardizing support for hybrid human-AI experimentation, Deliberate Lab expands the methodological repertoire for studying collective decision-making and human-centered AI.","authors":["Crystal Qian","Vivian Tsai","Michael Behr","Nada Hussein","Léo Laugier","Nithum Thain","Lucas Dixon"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2025-10-14","first_seen":"2025-10-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.13011","pdf_url":"https://arxiv.org/pdf/2510.13011","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B1"],"tags":["人机混合实验","集体决策","LLM代理"],"reason":"平台支持人类与LLM agent混合实验，用于研究集体决策，有真实人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":89,"question":"如何设计一个支持实时、大规模、人类与LLM智能体混合参与的在线社会实验平台？","design":"本文并非仿真研究，而是介绍一个开源实验平台Deliberate Lab。该平台允许研究者创建包含人类参与者和LLM智能体的实时、同步、多参与者实验，LLM可作为参与者、主持人或模拟人群。平台提供无代码界面、模块化阶段设计、实时状态管理和智能体响应逻辑。","baseline":"无对照","findings":"在12个月的公开部署中，88名实验者创建了597个实验，涉及9195名参与者，覆盖心理学、经济学、人机交互等领域。平台降低了实时混合人机实验的技术门槛，支持了集体决策和人机协作等多样化研究。","reliability":"论文未讨论","relevance":"该平台直接支持将LLM作为人类被试的替代品或补充，用于实时互动实验，但本文未提供仿真有效性的实证评估或与人类基准的对比，适合作为工具参考而非仿真可靠性研究。","inspiration":"该平台支持将LLM智能体作为人类被试的替代品或补充，用于实时、同步、多参与者互动实验，其模块化阶段设计和无代码界面降低了实验部署门槛｜可迁移至经济金融领域的市场博弈、集体决策或政策沟通实验，如资产定价中的信息扩散与泡沫形成、央行沟通中的预期引导｜可设计一个资产市场实验，以LLM智能体为交易者，施加不同信息透明度处理，测量价格泡沫与交易量，并对照已有真实人类实验数据（如Smith等人1988年的泡沫实验）评估仿真有效性"}},{"id":"2510.12189","version":1,"title":"Agent-Based Simulation of a Financial Market with Large Language Models","zh_title":"基于大语言模型的金融市场智能体仿真","abstract":"In real-world stock markets, certain chart patterns -- such as price declines near historical highs -- cannot be fully explained by fundamentals alone. These phenomena suggest the presence of path dependence in price formation, where investor decisions are influenced not only by current market conditions but also by the trajectory of prices leading up to the present. Path dependence has drawn attention in behavioral finance as a key mechanism behind such anomalies. One plausible driver of path dependence is human loss aversion, anchored to individual reference points like purchase prices or past peaks, which vary with personal context. However, capturing such subtle behavioral tendencies in traditional agent-based market simulations has remained a challenge. We propose the Fundamental-Chartist-LLM-Agent (FCLAgent), which uses large language models (LLMs) to emulate human-like trading decisions. In this framework, (1) buy/sell decisions are made by LLMs based on individual situations, while (2) order price and volume follow standard rule-based methods. Simulations show that FCLAgents reproduce path-dependent patterns that conventional agents fail to capture. Furthermore, an analysis of FCLAgents' behavior reveals that the reference points guiding loss aversion vary with market trajectories, highlighting the potential of LLM-based agents to model nuanced investor behavior.","authors":["Ryuji Hashimoto","Takehiro Takayanagi","Masahiro Suzuki","Kiyoshi Izumi"],"categories":["cs.CE"],"primary_category":"cs.CE","announce_type":"new","date":"2025-10-14","first_seen":"2025-10-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.12189","pdf_url":"https://arxiv.org/pdf/2510.12189","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2"],"tags":["LLM仿真","行为金融","智能体市场"],"reason":"用LLM代理模拟金融市场投资者行为，复现路径依赖，涉及行为金融场景，但未明确与…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":106,"question":"如何利用大语言模型构建能再现金融市场中路径依赖现象的智能体仿真模型？","design":"提出FCLAgent模型，将LLM用于生成买卖意图以模拟情境依赖的损失厌恶等行为偏差，订单价格和数量仍由传统规则决定；在单市场多智能体仿真中，随机选取智能体逐时间步下单，观察价格序列的路径依赖模式。","baseline":"无对照","findings":"FCLAgent能复现传统智能体无法捕捉的路径依赖模式，如接近历史高点与未来收益的负相关；LLM的决策参考点随市场轨迹变化，表现出类似人类的损失厌恶和风险偏好调整。","reliability":"论文未讨论","relevance":"该研究用LLM模拟投资者行为偏差，复现金融市场异象，属于人类仿真实验范畴，但缺乏真实人类数据对照，适合关注LLM行为逼真度和情境依赖建模的研究者阅读。","inspiration":"借鉴LLM生成买卖意图以模拟情境依赖行为偏差的方法，可将LLM作为被试，通过提示词操控其参考点或情绪状态，观察决策变化｜可迁移到资产定价实验中的处置效应研究，检验投资者是否因参考点变化而表现出持有亏损资产、过早卖出盈利资产的倾向｜设计：以LLM为被试，随机分配历史价格路径（如近期高点或低点），要求其决定持有或卖出，结果变量为卖出概率，对照真实市场交易数据中的处置效应模式"}},{"id":"2510.08338","version":3,"title":"LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings","zh_title":"大语言模型通过语义相似度引出李克特评分复现人类购买意向","abstract":"Consumer research costs companies billions annually yet suffers from panel biases and limited scale. Large language models (LLMs) offer an alternative by simulating synthetic consumers, but produce unrealistic response distributions when asked directly for numerical ratings. We present semantic similarity rating (SSR), a method that elicits textual responses from LLMs and maps these to Likert distributions using embedding similarity to reference statements. Testing on an extensive dataset comprising 57 personal care product surveys conducted by a leading corporation in that market (9,300 human responses), SSR achieves 90% of human test-retest reliability while maintaining realistic response distributions (KS similarity > 0.85). Additionally, these synthetic respondents provide rich qualitative feedback explaining their ratings. This framework enables scalable consumer research simulations while preserving traditional survey metrics and interpretability.","authors":["Benjamin F. Maier","Ulf Aslak","Luca Fiaschi","Nina Rismal","Kemble Fletcher","Christian C. Luhmann","Robbie Dow","Kli Pappas","Thomas V. Wiecki"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-10-09","first_seen":"2025-10-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.08338","pdf_url":"https://arxiv.org/pdf/2510.08338","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","消费者调查","人类数据对照"],"reason":"用LLM模拟消费者购买意向，与真实人类调查数据对照，评估分布可靠性与偏差，直接…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"如何通过语义相似度评分（SSR）方法让大语言模型生成更真实的人类购买意向分布？","design":"用大语言模型扮演给定人口统计属性的合成消费者，对57项个人护理产品概念生成自由文本购买意向，再通过嵌入向量与锚定语句的余弦相似度映射到5点李克特量表，测量购买意向分布和平均购买意向。","baseline":"来自一家领先企业的57项个人护理产品调查，共9,300名真实美国消费者的购买意向李克特评分。","findings":"SSR方法使合成消费者的购买意向分布与人类高度相似（KS相似度>0.85），且合成数据与人类数据在概念吸引力排名上的相关性达到人类重测信度的90%。同时，合成消费者能提供解释评分的定性反馈。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM模拟消费者购买意向，并与大规模真实人类调查数据严格对照，评估分布可靠性和偏差，直接命中研究者关心的仿真基准和经济学实验场景，值得精读原文。","inspiration":"借鉴其语义相似度评分方法，将LLM生成的自由文本通过嵌入向量与锚定语句的余弦相似度映射到结构化量表，可提升合成行为分布的真实性｜可迁移到消费者跨期选择实验，用LLM模拟不同人口统计特征下的时间偏好与折扣因子｜用LLM扮演合成消费者，施加未来不同时间点的金额选择任务，生成自由文本决策理由，通过SSR映射为选择概率，与真实跨期选择调查数据对照"}},{"id":"2510.08236","version":2,"title":"The Hidden Bias: A Study on Explicit and Implicit Political Stereotypes in Large Language Models","zh_title":"隐藏的偏见：大语言模型中显性与隐性政治刻板印象研究","abstract":"Large Language Models (LLMs) are increasingly integral to information dissemination and decision-making processes. Given their growing societal influence, understanding potential biases, particularly within the political domain, is crucial to prevent undue influence on public opinion and democratic processes. This work investigates political bias and stereotype propagation across eight prominent LLMs using the two-dimensional Political Compass Test (PCT). Initially, the PCT is employed to assess the inherent political leanings of these models. Subsequently, persona prompting with the PCT is used to explore explicit stereotypes across various social dimensions. In a final step, implicit stereotypes are uncovered by evaluating models with multilingual versions of the PCT. Key findings reveal a consistent left-leaning political alignment across all investigated models. Furthermore, while the nature and extent of stereotypes vary considerably between models, implicit stereotypes elicited through language variation are more pronounced than those identified via explicit persona prompting. Interestingly, for most models, implicit and explicit stereotypes show a notable alignment, suggesting a degree of transparency or \"awareness\" regarding their inherent biases. This study underscores the complex interplay of political bias and stereotypes in LLMs.","authors":["Konrad Löhr","Shuzhou Yuan","Michael Färber"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2025-10-09","first_seen":"2025-10-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.08236","pdf_url":"https://arxiv.org/pdf/2510.08236","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["政治偏见","刻板印象","LLM评估"],"reason":"测量LLM本身的政治偏见与刻板印象，属于人格/态度测量，无人类被试仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":167,"question":"八种主流大语言模型在政治罗盘测试中表现出怎样的内在政治倾向、外显刻板印象和内隐刻板印象？","design":"本研究并非人类仿真实验，而是直接测量LLM本身的政治偏见与刻板印象。使用政治罗盘测试（PCT）评估八个LLM的基线政治倾向；通过角色提示（persona prompting）让模型扮演不同社会维度角色，测量外显刻板印象；再通过多语言版PCT评估模型在不同语言下的表现，揭示内隐刻板印象。","baseline":"无对照","findings":"所有模型均一致表现出左倾政治倾向（经济左翼、社会自由意志主义）。内隐刻板印象（通过语言变化引发）比外显刻板印象（通过角色提示引发）更显著，且多数模型的内隐与外显刻板印象存在明显一致性，表明模型对其内在偏见有一定“意识”。","reliability":"论文未讨论","relevance":"该研究直接测量LLM的政治态度与刻板印象，未将LLM作为人类被试的替代品，也无真实人类数据对照，不属于人类仿真实验。若关注LLM偏见本身可读，但与研究者关心的仿真效度与失效条件无关。","inspiration":"该方法通过角色提示（persona prompting）施加外显刻板印象处理，并利用多语言测试揭示内隐刻板印象，为测量隐性偏见提供了可借鉴的实验设计框架。｜可迁移至信贷审批中的种族或性别歧视研究，探究LLM在贷款决策中是否存在内隐偏见。｜以LLM为被试，通过角色提示模拟不同种族或性别的贷款申请人，结果变量为贷款批准率与利率设定，并以真实银行信贷数据作为对照基准。"}},{"id":"2510.06903","version":1,"title":"When Machines Meet Each Other: Network Effects and the Strategic Role of History in Multi-Agent AI","zh_title":"当机器相遇：多智能体AI中的网络效应与历史的战略角色","abstract":"As artificial intelligence (AI) enters the agentic era, large language models (LLMs) are increasingly deployed as autonomous agents that interact with one another rather than operate in isolation. This shift raises a fundamental question: how do machine agents behave in interdependent environments where outcomes depend not only on their own choices but also on the coordinated expectations of peers? To address this question, we study LLM agents in a canonical network-effect game, where economic theory predicts convergence to a fulfilled expectation equilibrium (FEE). We design an experimental framework in which 50 heterogeneous GPT-5-based agents repeatedly interact under systematically varied network-effect strengths, price trajectories, and decision-history lengths. The results reveal that LLM agents systematically diverge from FEE: they underestimate participation at low prices, overestimate at high prices, and sustain persistent dispersion. Crucially, the way history is structured emerges as a design lever. Simple monotonic histories-where past outcomes follow a steady upward or downward trend-help stabilize coordination, whereas nonmonotonic histories amplify divergence and path dependence. Regression analyses at the individual level further show that price is the dominant driver of deviation, history moderates this effect, and network effects amplify contextual distortions. Together, these findings advance machine behavior research by providing the first systematic evidence on multi-agent AI systems under network effects and offer guidance for configuring such systems in practice.","authors":["Yu Liu","Wenwen Li","Yifan Dou","Guangnan Ye"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-10-08","first_seen":"2025-10-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.06903","pdf_url":"https://arxiv.org/pdf/2510.06903","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2"],"tags":["LLM仿真","网络效应博弈","多智能体"],"reason":"用LLM agent模拟网络效应博弈，涉及经济学实验场景，但缺真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":108,"question":"LLM智能体在网络效应博弈中能否形成与经济学理论一致的协调预期与均衡行为？","design":"用50个异构的GPT-5智能体模拟经济主体，在系统变化的网络效应强度、价格轨迹和决策历史长度下重复进行网络效应博弈，测量其参与预期与理论均衡的偏离。","baseline":"无对照","findings":"LLM智能体系统性地偏离理论均衡：低价时低估参与，高价时高估参与，且预期持续分散；单调历史有助于稳定协调，非单调历史则放大偏离和路径依赖。","reliability":"论文未讨论","relevance":"该研究用LLM模拟多主体网络效应博弈，属于经济学实验场景的仿真，但缺乏真实人类数据对照，可作为批判性研究的对象，值得阅读以评估其仿真有效性边界。","inspiration":"该研究通过系统操纵网络效应强度、价格轨迹和决策历史长度来测量LLM智能体的预期与均衡偏离，这种多维度处理设计值得借鉴｜可迁移到资产定价实验中的预期形成研究，例如检验LLM智能体在价格泡沫或崩盘历史下的预期偏差｜用LLM智能体作为被试，处理为不同历史价格序列（单调上涨/下跌、非单调波动），结果变量为预期价格与理性预期均衡的偏离，对照真实人类实验数据"}},{"id":"2510.06151","version":1,"title":"LLMs as Policy-Agnostic Teammates: A Case Study in Human Proxy Design for Heterogeneous Agent Teams","zh_title":"作为策略无关队友的大语言模型：异构智能体团队中人类代理设计的案例研究","abstract":"A critical challenge in modelling Heterogeneous-Agent Teams is training agents to collaborate with teammates whose policies are inaccessible or non-stationary, such as humans. Traditional approaches rely on expensive human-in-the-loop data, which limits scalability. We propose using Large Language Models (LLMs) as policy-agnostic human proxies to generate synthetic data that mimics human decision-making. To evaluate this, we conduct three experiments in a grid-world capture game inspired by Stag Hunt, a game theory paradigm that balances risk and reward. In Experiment 1, we compare decisions from 30 human participants and 2 expert judges with outputs from LLaMA 3.1 and Mixtral 8x22B models. LLMs, prompted with game-state observations and reward structures, align more closely with experts than participants, demonstrating consistency in applying underlying decision criteria. Experiment 2 modifies prompts to induce risk-sensitive strategies (e.g. \"be risk averse\"). LLM outputs mirror human participants' variability, shifting between risk-averse and risk-seeking behaviours. Finally, Experiment 3 tests LLMs in a dynamic grid-world where the LLM agents generate movement actions. LLMs produce trajectories resembling human participants' paths. While LLMs cannot yet fully replicate human adaptability, their prompt-guided diversity offers a scalable foundation for simulating policy-agnostic teammates.","authors":["Aju Ani Justus","Chris Baber"],"categories":["cs.LG","cs.AI","cs.HC"],"primary_category":"cs.LG","announce_type":"new","date":"2025-10-07","first_seen":"2025-10-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.06151","pdf_url":"https://arxiv.org/pdf/2510.06151","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","行为博弈","人机协作"],"reason":"用LLM代理人类决策，与真实人类数据对照，涉及博弈论实验，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":48,"question":"LLM能否作为策略无关的人类代理，在猎鹿博弈网格游戏中复现专家决策、模拟人类风险偏好变异并生成类人运动轨迹？","design":"使用LLaMA 3.1和Mixtral 8x22B模型，通过提示词输入网格世界的相对距离和奖励结构，模拟人类在猎鹿博弈中的决策。实验1比较LLM与30名人类参与者和2名专家裁判的选择；实验2通过修改提示词诱导风险规避或风险寻求策略；实验3让LLM在动态网格中生成移动动作序列。","baseline":"30名对博弈论了解有限的人类参与者在15种网格配置下的决策，以及2名博弈论专家裁判的选择。","findings":"LLM的决策与专家裁判高度一致，表现出对底层决策准则的一致性应用；通过提示词引导，LLM能展现出与人类参与者相似的风险偏好变异，在风险规避和风险寻求行为间切换。","reliability":"LLM尚不能完全复制人类的适应性，其行为多样性依赖于提示词引导，而非自主涌现。","relevance":"该研究直接用LLM代理人类决策，并与真实人类数据对照，涉及博弈论实验，方法可迁移至经济学和政策评估场景，对关注仿真可靠性与偏差的研究者具有参考价值，值得阅读原文。","inspiration":"该方法通过提示词工程直接操控LLM的风险偏好（风险规避/风险寻求），并设置专家裁判和人类参与者双基准对照，可借鉴其处理异质性行为变异的设计｜可迁移至资产定价实验中的风险态度测量，或政策公告对投资者预期形成的仿真研究｜以LLM作为被试，通过提示词注入不同风险偏好指令，测量其在模拟股票投资任务中的资产配置比例，并与真实投资者调查数据或实验数据对照"}},{"id":"2509.25709","version":1,"title":"Leveraging LLMs to Improve Experimental Design: A Generative Stratification Approach","zh_title":"利用大语言模型改进实验设计：一种生成式分层方法","abstract":"Pre-experiment stratification, or blocking, is a well-established technique for designing more efficient experiments and increasing the precision of the experimental estimates. However, when researchers have access to many covariates at the experiment design stage, they often face challenges in effectively selecting or weighting covariates when creating their strata. This paper proposes a Generative Stratification procedure that leverages Large Language Models (LLMs) to synthesize high-dimensional covariate data to improve experimental design. We demonstrate the value of this approach by applying it to a set of experiments and find that our method would have reduced the variance of the treatment effect estimate by 10%-50% compared to simple randomization in our empirical applications. When combined with other standard stratification methods, it can be used to further improve the efficiency. Our results demonstrate that LLM-based simulation is a practical and easy-to-implement way to improve experimental design in covariate-rich settings.","authors":["George Gui","Seungwoo Kim"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2025-09-30","first_seen":"2025-09-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.25709","pdf_url":"https://arxiv.org/pdf/2509.25709","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["实验设计","分层抽样","合成数据"],"reason":"用LLM生成合成协变量改进实验分层，替代人工设计而非仿真人类被试，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":195,"question":"如何利用大语言模型合成高维协变量信息，生成有效的预后得分以改进实验分层设计，从而提高处理效应估计的精度？","design":"本研究不是用LLM仿真人类被试，而是提出一种生成式分层方法：利用LLM基于实验单元的观测协变量和实验情境，生成预测的潜在结果，以此构建预后得分用于分层，从而改进随机化实验的设计效率。","baseline":"无对照","findings":"该方法在多个实证应用中能将处理效应估计的方差相比简单随机化降低10%-50%；与标准分层方法结合可进一步提升效率，且LLM预测即使不完美也不会引入偏差，因为层内随机化保持了无偏性。","reliability":"论文指出，分层效果依赖于LLM生成的得分与真实最优得分的相关性；若相关性弱，效率提升可能有限，但由于随机化，估计仍无偏。","relevance":"该研究与您关注的LLM仿真人类被试不同，它用LLM生成协变量得分以优化实验设计，而非替代人类参与实验。但涉及LLM在实验方法中的应用，且讨论了预测可靠性与偏差，可作为边界参考，建议略读以了解LLM在实验设计中的新用途。","inspiration":"该方法利用LLM生成预后得分以改进分层设计，可借鉴其将LLM预测作为协变量降维工具的思路，用于提升随机化实验的效率｜可迁移到政策评估场景，如利用LLM基于个体特征预测政策干预的潜在结果，优化分层随机化以更精确估计处理效应｜设计：以真实政策受众为被试，处理为某项就业培训，结果变量为就业收入，用LLM基于基线协变量生成预后得分进行分层随机化，以真实历史数据作为对照基准评估效率提升"}},{"id":"2510.02343","version":1,"title":"$\\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training","zh_title":"BluePrint：用于LLM角色评估与训练的社交媒体用户数据集","abstract":"Large language models (LLMs) offer promising capabilities for simulating social media dynamics at scale, enabling studies that would be ethically or logistically challenging with human subjects. However, the field lacks standardized data resources for fine-tuning and evaluating LLMs as realistic social media agents. We address this gap by introducing SIMPACT, the SIMulation-oriented Persona and Action Capture Toolkit, a privacy respecting framework for constructing behaviorally-grounded social media datasets suitable for training agent models. We formulate next-action prediction as a task for training and evaluating LLM-based agents and introduce metrics at both the cluster and population levels to assess behavioral fidelity and stylistic realism. As a concrete implementation, we release BluePrint, a large-scale dataset built from public Bluesky data focused on political discourse. BluePrint clusters anonymized users into personas of aggregated behaviours, capturing authentic engagement patterns while safeguarding privacy through pseudonymization and removal of personally identifiable information. The dataset includes a sizable action set of 12 social media interaction types (likes, replies, reposts, etc.), each instance tied to the posting activity preceding it. This supports the development of agents that use context-dependence, not only in the language, but also in the interaction behaviours of social media to model social media users. By standardizing data and evaluation protocols, SIMPACT provides a foundation for advancing rigorous, ethically responsible social media simulations. BluePrint serves as both an evaluation benchmark for political discourse modeling and a template for building domain specific datasets to study challenges such as misinformation and polarization.","authors":["Aurélien Bück-Kaeffer","Je Qin Chooi","Dan Zhao","Maximilian Puelma Touzel","Kellin Pelrine","Jean-François Godbout","Reihaneh Rabbany","Zachary Yang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-09-27","first_seen":"2025-09-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.02343","pdf_url":"https://arxiv.org/pdf/2510.02343","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","社交媒体模拟","行为保真度"],"reason":"用LLM模拟社交媒体用户行为，有真实数据对照，涉及政治话语，方法可迁移至人类仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":44,"question":"如何构建隐私保护的社交媒体数据集，并用于训练和评估基于大语言模型的社交媒体用户行为模拟代理？","design":"提出SIMPACT框架，将用户行为建模为动作序列，通过聚类形成行为角色以保护隐私；基于Bluesky平台2025年加拿大联邦选举期间的政治话语数据构建BluePrint数据集，包含12种交互行为；将下一动作预测作为训练和评估任务，使用GPT-4.1-mini、o3-mini、Qwen-2.5-7B等模型进行基准测试，并采用余弦相似度、Jaccard相似度、JS散度、F1分数及人类评估等多指标衡量行为保真度。","baseline":"以BluePrint数据集中真实用户的帖子嵌入、关键词分布和实际交互行为作为对照基准。","findings":"当前LLM能生成看似合理的文本，但在复现真实用户社区的行为模式上存在困难；经过微调的模型在行为预测上有所提升，但仍难以完全捕捉人类社交互动的细微差别。","reliability":"论文指出计算指标只能近似评估，无法完全捕捉人类行为的微妙性，可能产生“恐怖谷”效应；数据集聚焦特定政治事件和平台，泛化性有待验证；隐私保护措施可能损失部分个体行为细节。","relevance":"该研究直接涉及用LLM模拟社交媒体用户行为，并提供真实人类数据对照和评估基准，符合研究者对仿真可靠性及政治话语场景的关注，值得阅读原文以了解其数据集构建方法和模型失效的具体表现。","inspiration":"借鉴其将用户行为建模为动作序列并用聚类形成行为角色以保护隐私的方法，可设计隐私合规的仿真实验｜可迁移到金融社交媒体信息传播与投资者情绪形成的场景，如研究政策推文对散户交易行为的影响｜以LLM代理模拟散户投资者，处理为不同情绪倾向的政策推文，结果变量为模拟的买卖行为序列，用真实交易数据或社交媒体互动数据作为对照基准"}},{"id":"2509.13712","version":1,"title":"Inject, Fork, Compare: Defining an Interaction Vocabulary for Multi-Agent Simulation Platforms","zh_title":"注入、分叉、比较：为多智能体仿真平台定义交互词汇","abstract":"LLM-based multi-agent simulations are a rapidly growing field of research, but current simulations often lack clear modes for interaction and analysis, limiting the \"what if\" scenarios researchers are able to investigate. In this demo, we define three core operations for interacting with multi-agent simulations: inject, fork, and compare. Inject allows researchers to introduce external events at any point during simulation execution. Fork creates independent timeline branches from any timestamp, preserving complete state while allowing divergent exploration. Compare facilitates parallel observation of multiple branches, revealing how different interventions lead to distinct emergent behaviors. Together, these operations establish a vocabulary that transforms linear simulation workflows into interactive, explorable spaces. We demonstrate this vocabulary through a commodity market simulation with fourteen AI agents, where researchers can inject contrasting events and observe divergent outcomes across parallel timelines. By defining these fundamental operations, we provide a starting point for systematic causal investigation in LLM-based agent simulations, moving beyond passive observation toward active experimentation.","authors":["HwiJoon Lee","Martina Di Paola","Yoo Jin Hong","Quang-Huy Nguyen","Joseph Seering"],"categories":["cs.MA","cs.HC"],"primary_category":"cs.MA","announce_type":"new","date":"2025-09-17","first_seen":"2025-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.13712","pdf_url":"https://arxiv.org/pdf/2509.13712","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","社会模拟","交互操作"],"reason":"多智能体市场模拟，但无真实人类数据对照，属社会模拟演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":196,"question":"如何定义一套交互操作词汇（inject, fork, compare），使基于LLM的多智能体仿真从被动观察转变为主动因果探索？","design":"构建一个包含14个AI智能体的商品市场仿真，智能体具有不同交易策略和市场组合；通过注入对比事件（如石油管道爆炸 vs. OPEC增产），利用fork创建平行时间线分支，并用compare并排观察不同分支下涌现的智能体行为差异。","baseline":"无对照","findings":"定义了inject、fork、compare三种核心操作，将线性仿真流程转变为可交互的探索空间；在商品市场演示中，注入对立事件后，不同分支展现出分化的智能体行为，支持实时假设检验。","reliability":"论文未讨论","relevance":"该工作聚焦于多智能体仿真的交互范式，未涉及真实人类数据对照或行为复现，不属于以人类被试替代为目标的人类仿真研究，与研究者关注的经济学实验和政策评估场景关联较弱，不建议优先阅读原文。","inspiration":"该工作提出的inject、fork、compare交互范式可用于在仿真中主动注入经济冲击事件并观察多智能体行为分化，为因果推断提供可控实验环境｜可迁移到政策公告的预期形成与市场反应场景，如央行加息信号对资产价格和交易策略的影响｜以LLM驱动的交易员智能体为被试，注入加息或降息公告作为处理，fork出不同政策分支，compare分支间资产价格和交易量差异，并对照真实高频市场数据验证仿真行为模式"}},{"id":"2509.13397","version":4,"title":"The threat of analytic flexibility in using large language models to simulate human data","zh_title":"使用大语言模型模拟人类数据时分析灵活性的威胁","abstract":"Social scientists are now using large language models to create \"silicon samples\": synthetic datasets intended to stand in for human respondents. However, producing these samples requires many analytic choices, including model selection, sampling parameters, prompt format, and the amount of demographic or contextual information provided. Across two studies, I examine whether these choices materially affect correspondence between silicon samples and human data. In Study 1, I generated 252 silicon-sample configurations for a controlled case study using two social-psychological scales, evaluating whether configurations recovered participant rankings, response distributions, and between-scale correlations. Configurations varied substantially across all three criteria, and configurations that performed well on one dimension often performed poorly on another. In Study 2, I extended this analysis to a published silicon-sample use case by re-examining Argyle et al.'s (2023) Study 3 using 66 alternative configurations. Correlations between human and silicon association structures differed substantially across configurations, from r = .23 to r = .84. Taken together, the results from these studies demonstrate that different defensible configuration choices can materially alter conclusions about the fidelity of silicon samples. I call for greater attention to the threat of analytic flexibility in using silicon samples and outline strategies that researchers may adopt to reduce this threat.","authors":["Jamie Cummins"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-09-16","first_seen":"2025-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.13397","pdf_url":"https://arxiv.org/pdf/2509.13397","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["硅样本","分析灵活性","仿真保真度"],"reason":"直接研究用LLM生成硅样本模拟人类数据，评估分析灵活性对仿真保真度的影响，并与…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:28","error":null,"has_summary":true,"summary":{"generated_at":"2025-09-16","rank":7,"question":"分析灵活性（如模型选择、采样参数、提示格式等）是否显著影响硅样本与人类数据之间的一致性。","design":"研究1：使用两个社会心理量表，生成252种硅样本配置，评估配置能否恢复被试排名、响应分布和量表间相关性。研究2：重新分析Argyle等人(2023)的研究3，使用66种替代配置，比较人类与硅样本关联结构的相关性。","baseline":"研究1：来自两个社会心理量表的人类被试数据。研究2：Argyle等人(2023)研究3中的人类数据。","findings":"不同配置在恢复排名、响应分布和相关性上差异显著，且在一个维度上表现好的配置在另一维度上可能表现差。人类与硅样本关联结构的相关性从r=0.23到r=0.84不等，表明分析灵活性可实质改变关于硅样本保真度的结论。","reliability":"论文指出不同可辩护的配置选择会实质改变结论，但未明确列出所有失效条件或局限。","relevance":"直接命中研究者关注的LLM仿真人类数据可靠性问题，有真实人类对照，并批判性地揭示了分析灵活性威胁，值得精读原文以了解具体配置影响和应对策略。","inspiration":"借鉴其通过系统性地变化模型选择、采样参数和提示格式来构建多种硅样本配置，并评估配置间结果差异的方法，以揭示分析灵活性的威胁。｜可迁移到政策公告的预期形成实验，例如研究央行沟通措辞对通胀预期的影响。｜以LLM作为被试，随机分配不同措辞的政策公告作为处理，测量其预测的通胀数值，并与专业预测者调查或消费者预期调查的真实数据对照，同时变化提示中的角色设定、温度参数等配置，检验结论的稳健性。"}},{"id":"2509.11311","version":2,"title":"Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble","zh_title":"从提示到代理：通过紧凑LLM集成模拟人类偏好","abstract":"Large language models are increasingly used as proxies for human subjects in social science research, yet external validity requires that synthetic agents faithfully reflect the preferences of target human populations. We introduce *preference reconstruction theory*, a framework that formalizes preference alignment as a representation learning problem: constructing a functional basis of proxy agents and recovering population preferences through weighted aggregation. We implement this via *Prompts to Proxies* ($\\texttt{P2P}$), a modular two-stage system. Stage 1 uses structured prompting with entropy-based adaptive sampling to construct a diverse agent pool spanning the latent preference space. Stage 2 employs L1-regularized regression to select a compact ensemble whose aggregate response distributions align with observed data from the target population. $\\texttt{P2P}$ requires no finetuning and no access to sensitive demographic data, incurring only API inference costs. We validate the approach on 14 waves of the American Trends Panel, achieving an average test MSE of 0.014 across diverse topics at approximately 0.8 USD per survey. We additionally test it on the World Values Survey, demonstrating its potential to generalize across locales. When stress-tested against an SFT-aligned baseline, $\\texttt{P2P}$ achieves competitive performance using less than 3% of the training data.","authors":["Bingchen Wang","Zi-Yu Khoo","Jingtan Wang"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2025-09-14","first_seen":"2025-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.11311","pdf_url":"https://arxiv.org/pdf/2509.11311","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["LLM人类仿真","偏好重建","社会调查"],"reason":"用LLM代理复现人群偏好，有真实调查数据对照，涉及社会政策评估，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":45,"question":"如何通过紧凑的LLM集成，在无需微调和人口数据的情况下，忠实复现目标人群的偏好分布？","design":"使用GPT系列模型作为代理被试，通过两阶段系统P2P：第一阶段用结构化提示和基于熵的自适应采样构建多样化的代理池，第二阶段用L1正则化回归选择紧凑的代理集成，使其聚合响应分布与目标人群的调查数据对齐。结果变量为调查问题的回答分布。","baseline":"美国趋势面板（ATP）14波调查和世界价值观调查（WVS）的真实人类回答数据。","findings":"P2P在ATP上平均测试MSE为0.014，相比提示基线提升43%，每份调查成本约0.8美元；在WVS上展现出跨地域泛化潜力。与SFT对齐基线相比，P2P使用不到3%的训练数据即达到竞争性能。","reliability":"论文指出当前方法限于结构化调查问题，未处理自由文本输出；偏好重建依赖静态调查数据，可能无法应对非平稳偏好；未来需结合语义和心理测量技术提升鲁棒性。","relevance":"高度相关：该研究直接用LLM代理复现真实调查中的偏好分布，有严格的人类基准对照，涉及政策评估场景，并讨论了仿真失效条件，完全符合研究者的关注点，值得精读原文。","inspiration":"该方法通过两阶段系统（多样化代理池构建与L1正则化集成选择）实现无需微调、低成本的人类偏好复现，其自适应采样和稀疏集成思路值得借鉴｜可迁移到消费者金融决策偏好研究，如风险偏好、储蓄选择或投资组合配置的调查实验｜以LLM代理作为被试，施加不同金融信息框架处理，结果变量为风险资产选择比例，用美国消费者金融调查（SCF）的真实数据作为对照基准"}},{"id":"2509.09871","version":1,"title":"Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case","zh_title":"模拟民意：智利案例中AI生成合成调查回答的概念验证","abstract":"Large Language Models (LLMs) offer promising avenues for methodological and applied innovations in survey research by using synthetic respondents to emulate human answers and behaviour, potentially mitigating measurement and representation errors. However, the extent to which LLMs recover aggregate item distributions remains uncertain and downstream applications risk reproducing social stereotypes and biases inherited from training data. We evaluate the reliability of LLM-generated synthetic survey responses against ground-truth human responses from a Chilean public opinion probabilistic survey. Specifically, we benchmark 128 prompt-model-question triplets, generating 189,696 synthetic profiles, and pool performance metrics (i.e., accuracy, precision, recall, and F1-score) in a meta-analysis across 128 question-subsample pairs to test for biases along key sociodemographic dimensions. The evaluation spans OpenAI's GPT family and o-series reasoning models, as well as Llama and Qwen checkpoints. Three results stand out. First, synthetic responses achieve excellent performance on trust items (F1-score and accuracy > 0.90). Second, GPT-4o, GPT-4o-mini and Llama 4 Maverick perform comparably on this task. Third, synthetic-human alignment is highest among respondents aged 45-59. Overall, LLM-based synthetic samples approximate responses from a probabilistic sample, though with substantial item-level heterogeneity. Capturing the full nuance of public opinion remains challenging and requires careful calibration and additional distributional tests to ensure algorithmic fidelity and reduce errors.","authors":["Bastián González-Bustamante","Nando Verelst","Carla Cisternas"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-09-11","first_seen":"2025-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.09871","pdf_url":"https://arxiv.org/pdf/2509.09871","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"直接用LLM生成合成调查回答，与真实人类概率样本对照，评估可靠性与偏差，涉及公…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:27","error":null,"has_summary":true,"summary":{"generated_at":"2025-09-11","rank":11,"question":"大语言模型生成的合成调查回复能否准确反映智利的真实公众意见？","design":"使用128个提示-模型-问题三元组，生成189,696个合成档案，评估GPT系列、o系列推理模型、Llama和Qwen等模型在模拟智利公众意见调查回复上的表现，以准确率、精确率、召回率和F1分数为指标，并检验关键社会人口维度的偏差。","baseline":"智利公众意见概率调查的真实人类回复。","findings":"合成回复在信任项目上表现优异（F1和准确率>0.90）；GPT-4o、GPT-4o-mini和Llama 4 Maverick表现相当，且45-59岁年龄段的合成-人类对齐度最高。","reliability":"论文指出合成样本与概率样本的近似存在显著的项目层面异质性，捕捉公众意见的全部细微差别仍有挑战，需要仔细校准和额外的分布检验。","relevance":"该研究直接评估了LLM作为人类被试替代品的可靠性，有真实人类数据对照，并关注偏差，完全符合研究者的兴趣，值得精读。","inspiration":"该研究通过多模型、多提示的系统比较和关键社会人口维度的偏差检验来评估仿真可靠性，方法上值得借鉴｜可迁移到消费者信心指数或通胀预期调查的仿真上，检验LLM能否复现真实公众的经济预期分布｜以GPT-4o等为被试，用真实消费者调查问卷作为提示，生成通胀预期回复，以央行或统计局的真实调查数据为基准，比较分布一致性和子群体偏差"}},{"id":"2509.06337","version":2,"title":"Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation","zh_title":"大语言模型作为虚拟调查受访者：评估社会人口响应生成","abstract":"Questionnaire-based surveys are foundational to social science research and public policymaking, yet traditional survey methods remain costly, time-consuming, and often limited in scale. Although prior work has explored large language models (LLMs) as virtual survey respondents, existing studies often address narrow task settings, focus on single sociological domains, or lack a unified evaluation framework that enables systematic comparison across diverse datasets and models. To address these gaps, we introduce two complementary task abstractions: Partial Attribute Simulation (PAS), where LLMs predict missing attributes from incomplete respondent profiles, and Full Attribute Simulation (FAS), where LLMs generate complete synthetic datasets under zero-context and context-enhanced conditions. Both are framed as diagnostic and exploratory tools rather than replacements for human data collection. We curate LLM-S^3 (Large Language Model-based Sociodemographic Survey Simulation), a benchmark spanning 11 real-world public datasets across four sociological domains, and evaluate GPT-3.5/4 Turbo and LLaMA 3.0/3.1-8B under zero-shot and few-shot settings. Our evaluation reveals consistent performance trends across model families, highlights failure modes in structured output generation, and demonstrates how context and prompt design affect simulation fidelity. Our code and dataset are available at: https://github.com/dart-lab-research/LLM-S-Cube-Benchmark","authors":["Jianpeng Zhao","Chenyu Yuan","Weiming Luo","Haoling Xie","Guangwei Zhang","Steven Jige Quan","Zixuan Yuan","Pengyang Wang","Denghui Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-09-08","first_seen":"2025-09-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.06337","pdf_url":"https://arxiv.org/pdf/2509.06337","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","调查方法","社会人口模拟"],"reason":"用LLM模拟调查受访者，复现社会人口属性，有真实人类数据对照，涉及社会科学与政…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":35,"question":"如何系统评估大语言模型在模拟社会人口调查回答时的表现与可靠性？","design":"提出部分属性模拟（PAS）和全属性模拟（FAS）两种任务，使用GPT-3.5/4 Turbo和LLaMA 3.0/3.1-8B模型，在11个真实调查数据集上，以零样本和少样本设置生成缺失属性或完整合成数据集，并测量统计分布相似度。","baseline":"11个来自社会与公共事务、工作与收入、家庭与行为模式、健康与生活方式四个领域的真实公共调查数据集。","findings":"不同模型家族在模拟任务上表现出一致的性能趋势；提示设计和上下文增强显著影响模拟保真度，结构化输出生成中的失败案例仍是主要瓶颈。","reliability":"论文明确声明评估仅衡量统计和分布相似性，而非行为保真度或构念效度；数据集主要来自北美和欧洲，跨文化泛化性未验证；全属性模拟场景下结构化输出失败问题突出。","relevance":"该研究直接以真实人类调查数据为基准，系统评估LLM模拟社会人口属性的可靠性，并指出失效条件，高度契合研究者对仿真基准、政策评估场景和批判性分析的兴趣，值得精读原文。","inspiration":"借鉴其部分属性模拟（PAS）和全属性模拟（FAS）的任务设计，以真实调查数据为基准，通过统计分布相似度量化LLM的仿真保真度，并系统对比不同模型、提示策略和上下文增强的影响。｜可迁移到信贷审批中的歧视测量场景，用LLM模拟不同人口特征申请人的信用评分或审批结果，检验算法或人工决策中的统计性歧视。｜以真实信贷申请数据（如HMDA）为对照基准，将申请人的人口属性（种族、性别）作为处理变量，让LLM在PAS任务下生成信用评分或审批决策，比较LLM生成分布与真实审批分布的差异，并分析提示中是否加入反歧视法规对仿真偏差的影响。"}},{"id":"2509.03736","version":2,"title":"Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation","zh_title":"LLM代理行为一致吗？社会模拟的潜在画像","abstract":"The impressive capabilities of Large Language Models (LLMs) raise the possibility that synthetic agents can serve as substitutes for real participants in human-subject research. To evaluate this claim, prior research has largely focused on whether LLM-generated survey responses align with those produced by human respondents whom the LLMs are prompted to represent. In contrast, we address a more fundamental question: Do agents maintain empirical consistency; aligning to human behavioral models when examined under different experimental settings? To this end, we develop a study designed to (a) ask a set of questions which reveals an agent's latent profile and (b) examine agent behavioral consistency in a conversational setting with other agents. This design enables us to explore a set of behavioral hypotheses to assess whether an agent's conversational behavior is consistent with what we would expect from its revealed state. Our findings show significant inconsistencies in LLMs across model families and at differing model sizes. Most importantly, we find that, although agents may generate responses matching those of their human counterparts, they fail to be empirically consistent, representing a critical gap in their capabilities to accurately substitute for real participants in human-subject research.","authors":["James Mooney","Josef Woldense","Zheng Robert Jia","Shirley Anugrah Hayati","My Ha Nguyen","Vipul Raheja","Dongyeop Kang"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2025-09-03","first_seen":"2025-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.03736","pdf_url":"https://arxiv.org/pdf/2509.03736","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","行为一致性","人类被试替代"],"reason":"直接评估LLM代理替代人类被试的行为一致性，有人类数据对照，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":46,"question":"LLM代理在跨实验设置下是否保持行为一致性，即其对话行为是否与其自身揭示的潜在状态（偏好和开放性）相符？","design":"构建一个五阶段框架：选择争议性话题，生成具有人口统计特征和话题偏见的代理，通过问卷获取代理的潜在状态（偏好和开放性），将代理配对进行多轮对话，评估对话中的一致性。通过系统变化系统提示引入控制变量，测试六种人类行为模型。","baseline":"无对照","findings":"LLM代理在聚合层面表现出一些趋势，但无法通过更严格的经验一致性检验：即使偏好对立，代理也很少维持分歧；偏见提示不能可靠恢复原则性分歧；共享负面情绪产生的对齐弱于共享正面情绪；开放性在最应起作用的场景中失去预测力。","reliability":"论文指出当前LLM代理在行为模型更复杂或更细粒度时无法复制人类行为，其行为一致性存在显著差距，但未讨论具体失效条件。","relevance":"该研究直接评估LLM代理替代人类被试的行为一致性，揭示其在复杂行为模型中失效，与研究者关注的仿真可靠性及失效条件高度相关，值得精读原文。","inspiration":"借鉴其通过系统提示注入不同行为模型（如偏见、开放性）并检验代理行为一致性的实验设计，可作为一种施加处理的方式。｜可迁移至政策公告的预期形成实验，研究不同信息框架下投资者对央行沟通的反应一致性。｜以LLM代理为被试，处理为系统提示中嵌入不同的政策沟通风格（鹰派/鸽派），结果变量为代理在模拟交易中的资产配置变化，对照真实央行公告后的市场调查数据。"}},{"id":"2509.01813","version":3,"title":"ShortageSim: Simulating Drug Shortages under Information Asymmetry","zh_title":"ShortageSim：信息不对称下的药品短缺仿真","abstract":"Drug shortages pose critical risks to patient care and healthcare systems worldwide, yet the effectiveness of regulatory interventions remains poorly understood due to information asymmetries in pharmaceutical supply chains. We propose \\textbf{ShortageSim}, addresses this challenge by providing the first simulation framework that evaluates the impact of regulatory interventions on competition dynamics under information asymmetry. Using Large Language Model (LLM)-based agents, the framework models the strategic decisions of drug manufacturers and institutional buyers, in response to shortage alerts given by the regulatory agency. Unlike traditional game theory models that assume perfect rationality and complete information, ShortageSim simulates heterogeneous interpretations on regulatory announcements and the resulting decisions. Experiments on self-processed dataset of historical shortage events show that ShortageSim reduces the resolution lag for production disruption cases by up to 84\\%, achieving closer alignment to real-world trajectories than the zero-shot baseline. Our framework confirms the effect of regulatory alert in addressing shortages and introduces a new method for understanding competition in multi-stage environments under uncertainty. We open-source ShortageSim and a dataset of 2,925 FDA shortage events, providing a novel framework for future research on policy design and testing in supply chains under information asymmetry.","authors":["Mingxuan Cui","Yilan Jiang","Duo Zhou","Cheng Qian","Yuji Zhang","Qiong Wang"],"categories":["cs.MA","cs.CL","cs.GT"],"primary_category":"cs.MA","announce_type":"new","date":"2025-09-01","first_seen":"2025-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.01813","pdf_url":"https://arxiv.org/pdf/2509.01813","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","供应链博弈","政策评估"],"reason":"用LLM agent模拟药企与采购方决策，评估监管干预效果，涉及经济学场景与信…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":109,"question":"在信息不对称的药品供应链中，监管机构发布短缺预警如何影响制药商和采购方的策略性决策，进而影响药品短缺的解决速度？","design":"使用LLM驱动的多智能体框架，模拟FDA监管机构、制药商和医疗机构采购方在信息不对称环境下的决策。监管机构发布短缺预警，各智能体根据预警解读市场状态并做出生产或采购决策，测量短缺解决时滞和与历史轨迹的偏差。","baseline":"以FDA历史短缺事件数据集中的51条已解决事件轨迹作为真实对照基准。","findings":"ShortageSim在模拟生产中断案例时，将短缺解决时滞最多降低了84%，且比零样本基线更贴近真实历史轨迹；同时确认了主动预警政策可能引发囤货行为，反而加剧短缺。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟经济学场景下的决策行为，并与真实FDA短缺事件数据对照，评估监管干预效果，直接命中研究者关注的LLM仿真人类决策、政策评估和真实数据基准。值得精读原文。","inspiration":"该方法将LLM智能体置于信息不对称的供应链博弈中，通过历史事件轨迹校准仿真输出，并对比零样本基线以验证框架有效性，值得借鉴其‘真实数据校准+基线对比’的验证设计｜可迁移至金融市场政策公告的预期形成研究，如央行沟通对投资者决策的影响｜以LLM模拟投资者和央行官员，处理为央行发布不同透明度的政策指引，结果变量为资产价格波动和交易量，用历史央行公告前后的市场数据作为真实对照基准"}},{"id":"2510.07321","version":1,"title":"How human is the machine? Evidence from 66,000 Conversations with Large Language Models","zh_title":"机器有多像人？来自66000次与大语言模型对话的证据","abstract":"When Artificial Intelligence (AI) is used to replace consumers (e.g., synthetic data), it is often assumed that AI emulates established consumers, and more generally human behaviors. Ten experiments with Large Language Models (LLMs) investigate if this is true in the domain of well-documented biases and heuristics. Across studies we observe four distinct types of deviations from human-like behavior. First, in some cases, LLMs reduce or correct biases observed in humans. Second, in other cases, LLMs amplify these same biases. Third, and perhaps most intriguingly, LLMs sometimes exhibit biases opposite to those found in humans. Fourth, LLMs' responses to the same (or similar) prompts tend to be inconsistent (a) within the same model after a time delay, (b) across models, and (c) among independent research studies. Such inconsistencies can be uncharacteristic of humans and suggest that, at least at one point, LLMs' responses differed from humans. Overall, unhuman-like responses are problematic when LLMs are used to mimic or predict consumer behavior. These findings complement research on synthetic consumer data by showing that sources of bias are not necessarily human-centric. They also contribute to the debate about the tasks for which consumers, and more generally humans, can be replaced by AI.","authors":["Antonios Stamatogiannakis","Arsham Ghodsinia","Sepehr Etminanrad","Dilney Gonçalves","David Santos"],"categories":["cs.HC","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2025-08-31","first_seen":"2025-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.07321","pdf_url":"https://arxiv.org/pdf/2510.07321","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知偏差","人类数据对照"],"reason":"用LLM复现人类认知偏差并与真实人类数据对照，评估仿真可靠性并指出失效条件","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":23,"question":"大语言模型在经典决策偏差与启发式任务中，其行为在多大程度上与人类相似？","design":"使用多个大语言模型（如GPT系列）进行10项预注册实验，共66,000次对话，系统操纵模型类型和提示语特征，测量模型在可得性启发、代表性启发、禀赋效应、锚定效应、交易效用和框架效应六种偏差上的反应。","baseline":"以心理学和行为经济学文献中已确立的人类在这些偏差上的典型行为模式作为对照基准。","findings":"LLM表现出四种与人类不同的偏差模式：减弱或纠正人类偏差、放大人类偏差、表现出与人类相反的偏差，以及在不同时间、模型和研究间反应不一致。这些非人类反应表明LLM在模拟或预测消费者行为时存在问题。","reliability":"论文指出LLM的反应会随时间、模型版本和不同研究间出现不一致，这种不一致在人类中不典型，提示在某一时点LLM的反应与人类不同，从而限制了其作为人类替代品的可靠性。","relevance":"该研究直接评估LLM作为人类被试替代品的可靠性，通过经典决策偏差实验与真实人类行为基准对照，系统揭示了仿真失效的四种模式，对关注LLM仿真效度的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其多模型、多偏差、大样本的系统比较设计，可迁移到经济金融中的投资者行为偏差研究（如处置效应、过度自信），设计雏形：以GPT-4等LLM为被试，呈现模拟股票交易场景测量处置效应，结果与真实投资者交易数据（如券商账户记录）进行对照。"}},{"id":"2510.06222","version":1,"title":"Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making","zh_title":"在LLM智能体中诱导状态焦虑可复现消费者决策中的人类偏差","abstract":"Large language models (LLMs) are rapidly evolving from text generators to autonomous agents, raising urgent questions about their reliability in real-world contexts. Stress and anxiety are well known to bias human decision-making, particularly in consumer choices. Here, we tested whether LLM agents exhibit analogous vulnerabilities. Three advanced models (ChatGPT-5, Gemini 2.5, Claude 3.5-Sonnet) performed a grocery shopping task under budget constraints (24, 54, 108 USD), before and after exposure to anxiety-inducing traumatic narratives. Across 2,250 runs, traumatic prompts consistently reduced the nutritional quality of shopping baskets (Change in Basket Health Scores of -0.081 to -0.126; all pFDR<0.001; Cohens d=-1.07 to -2.05), robust across models and budgets. These results show that psychological context can systematically alter not only what LLMs generate but also the actions they perform. By reproducing human-like emotional biases in consumer behavior, LLM agents reveal a new class of vulnerabilities with implications for digital health, consumer safety, and ethical AI deployment.","authors":["Ziv Ben-Zion","Zohar Elyoseph","Tobias Spiller","Teddy Lazebnik"],"categories":["cs.HC","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2025-08-30","first_seen":"2025-08-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.06222","pdf_url":"https://arxiv.org/pdf/2510.06222","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","消费者行为","焦虑偏差"],"reason":"用LLM模拟消费者决策，诱导焦虑后复现人类偏差，有真实人类数据对照，涉及行为经…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":24,"question":"心理上下文（如焦虑诱导）是否会系统性地改变LLM智能体在消费者决策中的行为，使其表现出类似人类的情绪偏差？","design":"使用ChatGPT-5、Gemini 2.5、Claude 3.5-Sonnet三个先进LLM作为智能体，模拟人类消费者在预算约束（27、54、108美元）下进行杂货购物任务；通过暴露于焦虑诱导的创伤叙事作为处理，测量购物篮健康评分的变化。","baseline":"无对照","findings":"焦虑诱导提示一致降低了购物篮的营养质量（健康评分变化Δ=-0.081至-0.126，所有pFDR<0.001，Cohen's d=-1.07至-2.05），效应在不同模型和预算水平下均稳健。结果表明心理上下文不仅能改变LLM生成的文本，还能系统性地改变其作为智能体所执行的动作。","reliability":"论文未讨论","relevance":"该研究直接以LLM模拟人类消费者决策，通过情绪诱导复现人类偏差，并评估效应稳健性，高度契合研究者对LLM仿真可靠性及偏差的关注，值得精读原文以了解其方法细节和潜在失效边界。","inspiration":"借鉴其通过情绪化提示（创伤叙事）施加心理状态处理、以预算约束任务测量经济决策结果、并跨模型和预算水平进行稳健性检验的设计方法｜可迁移至消费者跨期选择或风险决策场景，如焦虑对储蓄/借贷行为或保险购买决策的影响｜以LLM智能体为被试，施加焦虑诱导处理，测量其在跨期选择任务中的贴现率或风险资产配置比例，并以真实人类实验数据（如实验室或调查数据）作为对照基准。"}},{"id":"2509.00462","version":4,"title":"AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights","zh_title":"算法招聘中的人工智能自我偏好：实证证据与见解","abstract":"As artificial intelligence (AI) tools become widely adopted, large language models (LLMs) are increasingly involved on both sides of decision-making processes, ranging from hiring to content moderation. This dual adoption raises a critical question: do LLMs systematically favor content that resembles their own outputs? Prior research in computer science has identified self-preference bias -- the tendency of LLMs to favor their own generated content -- but its real-world implications have not been empirically evaluated. We focus on the hiring context, where job applicants often rely on LLMs to refine resumes, while employers deploy them to screen those same resumes. Using a large-scale controlled resume correspondence experiment, we find that LLMs consistently prefer resumes generated by themselves over those written by humans or produced by alternative models, even when content quality is controlled. The bias against human-written resumes is particularly substantial, with self-preference bias ranging from 67% to 82% across major commercial and open-source models. To assess labor market impact, we simulate realistic hiring pipelines across 24 occupations. These simulations show that candidates using the same LLM as the evaluator are 23% to 60% more likely to be shortlisted than equally qualified applicants submitting human-written resumes, with the largest disadvantages observed in business-related fields such as sales and accounting. We further demonstrate that this bias can be reduced by more than 50% through simple interventions targeting LLMs' self-recognition capabilities. These findings highlight an emerging but previously overlooked risk in AI-assisted decision making and call for expanded frameworks of AI fairness that address not only demographic-based disparities, but also biases in AI-AI interactions.","authors":["Jiannan Xu","Gujie Li","Jane Yi Jiang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-30","first_seen":"2025-08-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.00462","pdf_url":"https://arxiv.org/pdf/2509.00462","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","招聘实验","算法偏差"],"reason":"用LLM模拟招聘决策并与人类简历对照，揭示自偏好偏差，可迁移至人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":110,"question":"大语言模型在简历筛选任务中是否系统性地偏好自己生成的简历，从而产生自偏好偏差？","design":"使用真实人类简历数据集，用GPT-4o、LLaMA 3.3-70B等7个主流LLM为每份简历生成多个反事实版本，再让这些模型作为评估者对简历进行评分或排序，测量对自身生成简历的偏好程度。","baseline":"来自专业简历平台的真实人类撰写简历，收集于生成式AI广泛采用之前。","findings":"多数LLM在评估时强烈偏好自己生成的简历，对人类简历的偏差达67%-82%；在模拟招聘中，使用与评估者相同LLM的候选人入围概率比同等条件的人类简历申请人高23%-60%，商业类职业劣势最明显。","reliability":"论文未讨论","relevance":"该研究用LLM模拟招聘决策并与真实人类简历对照，揭示了AI自偏好偏差，为人类仿真实验的可靠性与偏差评估提供了关键证据，值得精读。","inspiration":"该方法通过让LLM生成反事实简历并评估自身生成内容，巧妙测量自偏好偏差，可借鉴其对照设计｜可迁移至信贷审批歧视研究，检验AI是否对自身生成的贷款申请更宽容｜以LLM作为信贷审批员，处理组为LLM生成的贷款申请，对照组为真实申请人数据，结果变量为批准率，用历史信贷记录作基准"}},{"id":"2509.02605","version":1,"title":"Synthetic Founders: AI-Generated Social Simulations for Startup Validation Research in Computational Social Science","zh_title":"合成创始人：用于计算社会科学中创业验证研究的AI生成社会仿真","abstract":"We present a comparative docking experiment that aligns human-subject interview data with large language model (LLM)-driven synthetic personas to evaluate fidelity, divergence, and blind spots in AI-enabled simulation. Fifteen early-stage startup founders were interviewed about their hopes and concerns regarding AI-powered validation, and the same protocol was replicated with AI-generated founder and investor personas. A structured thematic synthesis revealed four categories of outcomes: (1) Convergent themes - commitment-based demand signals, black-box trust barriers, and efficiency gains were consistently emphasized across both datasets; (2) Partial overlaps - founders worried about outliers being averaged away and the stress of real customer validation, while synthetic personas highlighted irrational blind spots and framed AI as a psychological buffer; (3) Human-only themes - relational and advocacy value from early customer engagement and skepticism toward moonshot markets; and (4) Synthetic-only themes - amplified false positives and trauma blind spots, where AI may overstate adoption potential by missing negative historical experiences. We interpret this comparative framework as evidence that LLM-driven personas constitute a form of hybrid social simulation: more linguistically expressive and adaptable than traditional rule-based agents, yet bounded by the absence of lived history and relational consequence. Rather than replacing empirical studies, we argue they function as a complementary simulation category - capable of extending hypothesis space, accelerating exploratory validation, and clarifying the boundaries of cognitive realism in computational social science.","authors":["Jorn K. Teutloff"],"categories":["cs.MA","cs.AI","cs.CY"],"primary_category":"cs.MA","announce_type":"new","date":"2025-08-29","first_seen":"2025-08-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.02605","pdf_url":"https://arxiv.org/pdf/2509.02605","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","创业者访谈对照","仿真保真度评估"],"reason":"用LLM生成合成创业者进行访谈仿真，并与15位真人创业者数据对照，评估仿真保真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":47,"question":"LLM驱动的合成创业者与投资人角色在创业验证访谈中，与真实人类受访者的主题模式在多大程度上一致、偏离或存在盲点？","design":"采用对比对接实验：用LLM生成35个合成创业者与投资人角色，复制对15位真实早期创业者的半结构化访谈协议，询问其对AI驱动市场验证的希望与担忧，通过结构化主题综合比较合成与人类访谈文本的涌现主题。","baseline":"15位真实早期创业者的访谈数据，涵盖其对AI驱动市场验证的希望与担忧。","findings":"合成与人类数据在承诺型需求信号、黑箱信任障碍和效率提升上主题收敛；部分重叠中，人类担忧异常值被平均化和真实客户验证压力，合成角色则强调非理性盲点并将AI视为心理缓冲；人类独有主题包括早期客户参与的倡导价值和对极端市场的怀疑，合成独有主题为放大假阳性和创伤盲点。","reliability":"论文指出LLM角色缺乏真实生活经历和关系后果，可能高估采纳潜力并遗漏负面历史经验，因此不能替代实证研究，仅作为补充性仿真工具，用于扩展假设空间和加速探索性验证。","relevance":"该研究直接以真实人类访谈为基准，系统评估LLM仿真在创业决策中的保真度与偏差，并明确区分收敛、部分重叠和独有主题，对关注仿真可靠性及失效条件的研究者具有重要参考价值，值得细读原文。","inspiration":"该方法通过对比合成角色与真实人类在相同半结构化访谈下的主题收敛与偏离，系统评估LLM仿真的保真度与盲点，值得借鉴｜可迁移至创业融资决策研究，如投资者对AI辅助商业计划书的评估偏差｜以LLM生成的投资人角色为被试，呈现AI生成与人类撰写的商业计划书，测量投资意愿与风险评估，并以真实天使投资人的评审数据作为对照基准"}},{"id":"2508.20234","version":1,"title":"Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research","zh_title":"验证基于生成式智能体的物流与供应链管理研究模型","abstract":"Generative Agent-Based Models (GABMs) powered by large language models (LLMs) offer promising potential for empirical logistics and supply chain management (LSCM) research by enabling realistic simulation of complex human behaviors. Unlike traditional agent-based models, GABMs generate human-like responses through natural language reasoning, which creates potential for new perspectives on emergent LSCM phenomena. However, the validity of LLMs as proxies for human behavior in LSCM simulations is unknown. This study evaluates LLM equivalence of human behavior through a controlled experiment examining dyadic customer-worker engagements in food delivery scenarios. I test six state-of-the-art LLMs against 957 human participants (477 dyads) using a moderated mediation design. This study reveals a need to validate GABMs on two levels: (1) human equivalence testing, and (2) decision process validation. Results reveal GABMs can effectively simulate human behaviors in LSCM; however, an equivalence-versus-process paradox emerges. While a series of Two One-Sided Tests (TOST) for equivalence reveals some LLMs demonstrate surface-level equivalence to humans, structural equation modeling (SEM) reveals artificial decision processes not present in human participants for some LLMs. These findings show GABMs as a potentially viable methodological instrument in LSCM with proper validation checks. The dual-validation framework also provides LSCM researchers with a guide to rigorous GABM development. For practitioners, this study offers evidence-based assessment for LLM selection for operational tasks.","authors":["Vincent E. Castillo"],"categories":["cs.MA","cs.AI","cs.CY"],"primary_category":"cs.MA","announce_type":"new","date":"2025-08-27","first_seen":"2025-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.20234","pdf_url":"https://arxiv.org/pdf/2508.20234","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM人类仿真","等效性验证","供应链管理"],"reason":"直接验证LLM作为人类代理在供应链场景中的等效性，有957名人类对照，并揭示仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"在物流与供应链管理（LSCM）的二元互动场景中，基于大语言模型的生成式智能体模型能否在表面行为上与人类等效，且其决策过程是否与人类可比？","design":"采用受控实验，将六种最先进的大语言模型作为生成式智能体，模拟外卖配送场景中顾客与骑手的二元互动，与957名人类参与者（477对）进行对比，使用有调节的中介设计，测量互动结果与决策过程。","baseline":"957名人类参与者（477对二元组）在相同外卖配送场景实验中的真实行为数据。","findings":"部分LLM在表面行为上通过等效性检验（TOST）与人类无显著差异，但结构方程模型（SEM）揭示其决策过程存在人为模式，与人类真实决策过程不同，形成“等效-过程悖论”。","reliability":"论文指出LLM作为人类代理的有效性未知，需进行双重验证（人类等效性检验和决策过程验证），并承认某些LLM的决策过程与人类不符，但未详细讨论其他失效条件。","relevance":"该研究直接验证LLM在供应链场景中替代人类被试的等效性，提供大规模人类对照基准，并批判性揭示表面等效下的决策过程差异，高度契合研究者对仿真可靠性、偏差及失效条件的关注，值得精读原文。","inspiration":"该研究采用TOST等效性检验与结构方程模型（SEM）双重验证，区分表面行为等效与决策过程等效，为仿真可靠性评估提供了严谨框架｜可迁移至消费者跨期选择实验，检验LLM生成的折现行为是否与人类一致｜以LLM为被试，施加不同跨期奖励方案，结果变量为选择时间偏好，以真实人类实验数据（如Andersen et al., 2008）为基准，进行TOST和SEM双重验证"}},{"id":"2508.19004","version":1,"title":"AI Models Exceed Individual Human Accuracy in Predicting Everyday Social Norms","zh_title":"AI模型在预测日常社会规范方面超越个体人类准确性","abstract":"A fundamental question in cognitive science concerns how social norms are acquired and represented. While humans typically learn norms through embodied social experience, we investigated whether large language models can achieve sophisticated norm understanding through statistical learning alone. Across two studies, we systematically evaluated multiple AI systems' ability to predict human social appropriateness judgments for 555 everyday scenarios by examining how closely they predicted the average judgment compared to each human participant. In Study 1, GPT-4.5's accuracy in predicting the collective judgment on a continuous scale exceeded that of every human participant (100th percentile). Study 2 replicated this, with Gemini 2.5 Pro outperforming 98.7% of humans, GPT-5 97.8%, and Claude Sonnet 4 96.0%. Despite this predictive power, all models showed systematic, correlated errors. These findings demonstrate that sophisticated models of social cognition can emerge from statistical learning over linguistic data alone, challenging strong versions of theories emphasizing the exclusive necessity of embodied experience for cultural competence. The systematic nature of AI limitations across different architectures indicates potential boundaries of pattern-based social understanding, while the models' ability to outperform nearly all individual humans in this predictive task suggests that language serves as a remarkably rich repository for cultural knowledge transmission.","authors":["Pontus Strimling","Simon Karlsson","Irina Vartanova","Kimmo Eriksson"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-08-26","first_seen":"2025-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.19004","pdf_url":"https://arxiv.org/pdf/2508.19004","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["社会规范预测","人类数据对照","算法偏差"],"reason":"用LLM预测人类社会规范判断，与真实人类数据对照，评估预测准确性与系统性偏差，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":30,"question":"大语言模型能否仅通过文本统计学习，在没有具身体验的情况下，达到甚至超越个体人类对日常社会规范（行为适当性）的判断水平？","design":"本研究并非严格意义上的仿真实验，而是利用已有的大规模人类判断数据集，将多个大语言模型（GPT-4.5、Gemini 2.5 Pro、GPT-5、Claude Sonnet 4）作为“被试”，要求它们对555个日常场景中的行为适当性进行连续评分，并与真实人类个体评分进行对比。","baseline":"真实人类数据：来自大规模数据集中个体参与者对相同555个日常场景的社会适当性连续评分，包括平均集体判断和每个个体的判断分布。","findings":"GPT-4.5 预测集体判断的准确度超过了所有人类参与者（百分位100%），其他模型也超过了96%以上的人类个体。所有模型均表现出系统性、相互关联的误差，表明基于模式的统计学习在社会理解上存在边界。","reliability":"论文指出，尽管模型预测力强，但所有模型都存在系统性且相互关联的误差，这表明基于模式的社会理解存在潜在边界；研究仅限于日常社会规范判断，未涉及其他类型的社会认知任务。","relevance":"该研究直接以真实人类个体判断为基准，评估LLM在连续社会规范判断上的仿真准确性及系统性偏差，高度契合研究者对LLM仿真可靠性、偏差及失效条件的关注，值得精读原文以了解其个体级对比方法和误差分析。","inspiration":"可借鉴其将模型预测与人类个体分布而非仅与均值对比的评估设计，以揭示仿真在个体差异捕捉上的能力。｜可迁移至消费者对金融产品适当性的感知或政策可接受性判断等场景。｜以LLM作为“消费者”被试，输入不同金融产品描述，要求其评估产品适当性，结果变量为适当性连续评分，对照真实消费者调查数据中的个体评分分布。"}},{"id":"2509.00074","version":2,"title":"Language and Experience: A Computational Model of Social Learning in Complex Tasks","zh_title":"语言与经验：复杂任务中社会学习的计算模型","abstract":"The ability to combine linguistic guidance from others with direct experience is central to human development, enabling safe and rapid learning in new environments. How do people integrate these two sources of knowledge, and how might AI systems? We present a computational framework that models social learning as joint probabilistic inference over structured, executable world models given sensorimotor and linguistic data. We make this possible by turning a pretrained language model into a probabilistic model of how humans share advice conditioned on their beliefs, allowing our agents both to generate advice for others and to interpret linguistic input as evidence during Bayesian inference. Using behavioral experiments and simulations across 10 video games, we show how linguistic guidance can shape exploration and accelerate learning by reducing risky interactions and speeding up key discoveries in both humans and models. We further explore how knowledge can accumulate across generations through iterated learning experiments and demonstrate successful knowledge transfer between humans and models -- revealing how structured, language-compatible representations might enable human-machine collaborative learning.","authors":["Cédric Colas","Tracey Mills","Ben Prystawski","Michael Henry Tessler","Noah Goodman","Jacob Andreas","Joshua Tenenbaum"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2025-08-26","first_seen":"2025-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.00074","pdf_url":"https://arxiv.org/pdf/2509.00074","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","人类行为对照","社会学习"],"reason":"用LLM模拟人类在复杂任务中的社会学习，并与真实人类行为实验对照，涉及行为实验…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":111,"question":"人类如何整合语言指导与直接经验来学习复杂任务？","design":"提出一个贝叶斯计算框架，将预训练语言模型作为人类说话者模型，模拟人类在10款视频游戏中结合语言建议和直接经验进行社会学习；通过行为实验和模拟，比较纯体验、体验+人类消息、体验+模型消息三种条件，测量通关所需生命数、通关比例和归一化曲线下面积（nAUC）。","baseline":"122名Prolific参与者的人类行为数据，最终有效样本120人，随机分配到三种条件（各40人），记录其游戏表现和给出的建议。","findings":"语言指导能塑造探索策略并加速学习，减少风险交互并加快关键发现；知识可通过迭代学习在代际间积累，并实现人类与模型之间的成功知识迁移。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类在复杂任务中的社会学习，并与真实人类行为实验对照，涉及行为实验和代际知识传递，直接命中研究者对LLM仿真人类被试、复现行为并评估可靠性的兴趣，值得精读原文。","inspiration":"该研究用LLM模拟人类在复杂任务中结合语言建议和直接经验的学习过程，并设置纯体验、体验+人类消息、体验+模型消息三种条件进行对照，这种多条件对比设计可用于评估语言干预的因果效应｜可迁移到经济金融中的政策沟通与预期形成场景，例如研究央行前瞻指引如何影响投资者学习与资产配置｜以LLM模拟投资者，处理为提供不同风格（如模糊vs精确）的央行声明，结果变量为投资组合调整速度和准确性，用历史市场数据或人类实验数据作为基准对照"}},{"id":"2508.17322","version":1,"title":"Chinese Court Simulation with LLM-Based Agent System","zh_title":"基于大语言模型智能体的中国法庭仿真","abstract":"Mock trial has long served as an important platform for legal professional training and education. It not only helps students learn about realistic trial procedures, but also provides practical value for case analysis and judgment prediction. Traditional mock trials are difficult to access by the public because they rely on professional tutors and human participants. Fortunately, the rise of large language models (LLMs) provides new opportunities for creating more accessible and scalable court simulations. While promising, existing research mainly focuses on agent construction while ignoring the systematic design and evaluation of court simulations, which are actually more important for the credibility and usage of court simulation in practice. To this end, we present the first court simulation framework -- SimCourt -- based on the real-world procedure structure of Chinese courts. Our framework replicates all 5 core stages of a Chinese trial and incorporates 5 courtroom roles, faithfully following the procedural definitions in China. To simulate trial participants with different roles, we propose and craft legal agents equipped with memory, planning, and reflection abilities. Experiment on legal judgment prediction show that our framework can generate simulated trials that better guide the system to predict the imprisonment, probation, and fine of each case. Further annotations by human experts show that agents' responses under our simulation framework even outperformed judges and lawyers from the real trials in many scenarios. These further demonstrate the potential of LLM-based court simulation.","authors":["Kaiyuan Zhang","Jiaqi Li","Yueyue Wu","Haitao Li","Cheng Luo","Shaokun Zou","Yujia Zhou","Weihang Su","Qingyao Ai","Yiqun Liu"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-24","first_seen":"2025-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.17322","pdf_url":"https://arxiv.org/pdf/2508.17322","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["法庭仿真","多智能体","法律预测"],"reason":"模拟法庭过程但无真实人类行为对照，属社会模拟缺基准","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":168,"question":"如何基于中国真实庭审程序构建一个全阶段、多角色的LLM法庭模拟框架，并系统评估其判决预测准确性和过程质量？","design":"本研究提出SimCourt框架，基于中国刑事庭审的5个阶段和5种角色，为法官、检察官、律师、被告和书记员分别设计配备记忆、规划和反思模块及法律检索工具的LLM智能体，输入案件材料后自动生成完整庭审记录和判决书。","baseline":"无对照","findings":"SimCourt生成的模拟庭审能更好地指导系统预测监禁、缓刑和罚金等判决结果；人类专家标注显示，智能体在模拟框架下的回应在许多场景中甚至优于真实庭审中的法官和律师。","reliability":"论文未讨论","relevance":"该研究属于LLM社会模拟，但缺乏真实人类行为对照，未复现人类被试的决策分布或偏差，不符合研究者对基准人类数据的要求，不建议优先阅读原文。","inspiration":"SimCourt的多阶段、多角色框架和记忆-规划-反思模块设计，为构建复杂经济决策场景的LLM仿真提供了可借鉴的架构｜该设计可迁移到金融监管政策评估，如模拟银行、企业、监管机构等多方博弈对信贷供给的影响｜可构建包含银行、企业、监管者的LLM智能体，施加资本充足率变动处理，观察信贷审批决策，并以真实银行贷款数据和监管报告作为对照基准"}},{"id":"2508.16172","version":2,"title":"Graph RAG as Human Choice Model: Building a Data-Driven Mobility Agent with Preference Chain","zh_title":"图RAG作为人类选择模型：构建数据驱动的出行智能体与偏好链","abstract":"Understanding human behavior in urban environments is a crucial field within city sciences. However, collecting accurate behavioral data, particularly in newly developed areas, poses significant challenges. Recent advances in generative agents, powered by Large Language Models (LLMs), have shown promise in simulating human behaviors without relying on extensive datasets. Nevertheless, these methods often struggle with generating consistent, context-sensitive, and realistic behavioral outputs. To address these limitations, this paper introduces the Preference Chain, a novel method that integrates Graph Retrieval-Augmented Generation (RAG) with LLMs to enhance context-aware simulation of human behavior in transportation systems. Experiments conducted on the Replica dataset demonstrate that the Preference Chain outperforms standard LLM in aligning with real-world transportation mode choices. The development of the Mobility Agent highlights potential applications of proposed method in urban mobility modeling for emerging cities, personalized travel behavior analysis, and dynamic traffic forecasting. Despite limitations such as slow inference and the risk of hallucination, the method offers a promising framework for simulating complex human behavior in data-scarce environments, where traditional data-driven models struggle due to limited data availability.","authors":["Kai Hu","Parfait Atchade-Adelomou","Carlo Adornetto","Adrian Mora-Carrero","Luis Alonso-Pastor","Ariel Noyman","Yubo Liu","Kent Larson"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-08-22","first_seen":"2025-08-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.16172","pdf_url":"https://arxiv.org/pdf/2508.16172","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","出行行为","人类数据对照"],"reason":"用LLM仿真交通出行选择，有真实人类数据对照，属经济学实验场景，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":40,"question":"如何在数据稀缺环境下，利用图检索增强生成（Graph RAG）与LLM结合的方法，更真实地模拟个体交通出行方式选择行为？","design":"提出Preference Chain方法，结合Graph RAG与LLM构建Mobility Agent，基于少量数据构建个体行为偏好图，通过相似性搜索和概率建模引导LLM生成出行选择；在Replica数据集上模拟交通方式选择，与真实选择对比。","baseline":"Replica数据集中真实的交通方式选择行为。","findings":"Preference Chain在模拟交通方式选择上比标准LLM更符合真实世界数据；该方法在数据稀缺地区具有应用潜力，但存在推理速度慢和幻觉风险。","reliability":"论文承认推理速度慢和存在幻觉风险，可能影响行为仿真的可靠性和实用性。","relevance":"该研究用LLM仿真交通出行选择，有真实人类数据对照，属于经济学实验场景，方法可迁移至其他行为仿真，值得阅读原文以评估其仿真偏差与可靠性。","inspiration":"借鉴Preference Chain方法，利用图检索增强生成（Graph RAG）从少量个体行为数据中构建偏好图，并通过相似性搜索和概率建模引导LLM生成选择，以提升仿真真实性。｜该方法可迁移至消费者跨期选择研究，模拟个体在不同时间偏好下的储蓄或消费决策。｜以LLM作为被试，构建基于少量真实个体跨期选择数据的偏好图，施加不同利率或未来收入预期的处理，结果变量为模拟的消费-储蓄分配，与真实家庭金融调查数据（如PSID）进行对照。"}},{"id":"2508.15926","version":1,"title":"Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making","zh_title":"噪声、适应与策略：评估LLM在决策中的保真度","abstract":"Large language models (LLMs) are increasingly used in social science simulations. While their performance on reasoning and optimization tasks has been extensively evaluated, less attention has been paid to their ability to simulate human decision-making's variability and adaptability. We propose a process-oriented evaluation framework with progressive interventions (Intrinsicality, Instruction, and Imitation) to examine how LLM agents adapt under different levels of external guidance and human-derived noise. We validate the framework on two classic economics tasks, irrationality in the second-price auction and decision bias in the newsvendor problem, showing behavioral gaps between LLMs and humans. We find that LLMs, by default, converge on stable and conservative strategies that diverge from observed human behaviors. Risk-framed instructions impact LLM behavior predictably but do not replicate human-like diversity. Incorporating human data through in-context learning narrows the gap but fails to reach human subjects' strategic variability. These results highlight a persistent alignment gap in behavioral fidelity and suggest that future LLM evaluations should consider more process-level realism. We present a process-oriented approach for assessing LLMs in dynamic decision-making tasks, offering guidance for their application in synthetic data for social science research.","authors":["Yuanjun Feng","Vivek Choudhary","Yash Raj Shrestha"],"categories":["cs.CE","cs.AI"],"primary_category":"cs.CE","announce_type":"new","date":"2025-08-21","first_seen":"2025-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.15926","pdf_url":"https://arxiv.org/pdf/2508.15926","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为保真度","经济学实验"],"reason":"直接评估LLM模拟人类决策的保真度，用真实人类数据对照，涉及经济学任务，指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":48,"question":"LLM在动态决策任务中能否复现人类决策的变异性和适应性？","design":"用GPT-4o、Claude 3.5 Sonnet、Claude 3.7 Sonnet扮演卖家或报童，在二阶密封拍卖和报童问题中，通过无干预、风险框架指令、人类决策历史模仿三种渐进干预，测量其策略稳定性、行为变异性和适应性。","baseline":"对照二阶密封拍卖的真实人类实验数据（Davis et al., 2011, 2023），以及报童问题的已知人类决策偏差模式。","findings":"LLM默认采取稳定保守的策略，与人类行为偏离；风险框架指令可预测地影响LLM行为，但未复现人类多样性；通过上下文学习注入人类数据可缩小差距，但仍未达到人类被试的策略变异性。","reliability":"论文指出当前LLM在动态行为模拟中存在行为保真度对齐差距，缺乏人类决策的随机性和适应性，未来评估需关注过程层面的真实性。","relevance":"该研究直接评估LLM模拟人类决策的保真度，使用真实人类数据作为基准，涉及经济学实验任务，并明确指出了仿真失效的条件和局限，与研究者关注点高度吻合，值得精读原文。","inspiration":"该研究通过渐进式干预（无干预、风险框架指令、人类历史模仿）系统测量LLM行为变异性的方法值得借鉴，可迁移到资产定价实验中的泡沫形成与处置效应研究｜可设计让LLM扮演投资者在实验性资产市场中交易，处理为不同风险提示框架或注入真实人类交易历史，结果变量为价格偏离度、交易量与处置效应系数，对照真实人类实验数据（如Smith et al., 1988的泡沫实验）"}},{"id":"2508.12045","version":2,"title":"Large Language Models Enable Design of Personalized Nudges across Cultures","zh_title":"大语言模型助力跨文化个性化助推设计","abstract":"Nudge strategies are effective tools for influencing behaviour, but their impact depends on individual preferences. Strategies that work for some individuals may be counterproductive for others. We hypothesize that large language models (LLMs) can facilitate the design of individual-specific nudges without the need for costly and time-intensive behavioural data collection and modelling. To test this, we use LLMs to design personalized decoy-based nudges tailored to individual profiles and cultural contexts, aimed at encouraging air travellers to voluntarily offset CO$_2$ emissions from flights. We evaluate their effectiveness through a large-scale survey experiment ($n=3495$) conducted across five countries. Results show that LLM-informed personalized nudges are more effective than uniform settings, raising offsetting rates by 3-7$\\%$ in Germany, Singapore, and the US, though not in China or India. Our study highlights the potential of LLM as a low-cost testbed for piloting nudge strategies. At the same time, cultural heterogeneity constrains their generalizability underscoring the need for combining LLM-based simulations with targeted empirical validation.","authors":["Vladimir Maksimenko","Qingyao Xin","Prateek Gupta","Bin Zhang","Prateek Bansal"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-16","first_seen":"2025-08-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.12045","pdf_url":"https://arxiv.org/pdf/2508.12045","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为助推","跨文化实验"],"reason":"用LLM设计个性化助推并做大规模人类调查对照，直接仿真人类决策行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":25,"question":"大语言模型能否用于设计跨文化的个性化助推策略，以低成本替代人类被试进行助推效果评估？","design":"使用LLM根据个体人口统计特征（性别、年龄、收入、环境关心度、碳抵消信任度）和文化背景，为航空旅客生成个性化的诱饵选项（价格与碳抵消比例），通过大规模在线调查实验（n=3495）在五个国家测试其对自愿碳抵消选择的影响。","baseline":"在五个国家（中国、德国、印度、新加坡、美国）对3495名真实航空旅客进行的调查实验，测量其面对个性化诱饵与统一诱饵时的碳抵消选择率。","findings":"LLM生成的个性化助推在德国、新加坡和美国将碳抵消率提高了3-7%，但在中国和印度未显示显著效果。文化异质性限制了LLM仿真助推效果的普适性，需结合实证验证。","reliability":"论文指出LLM仿真受文化异质性约束，在部分国家（中国、印度）失效，且个性化助推效果依赖于对个体偏好和文化背景的准确建模，需结合针对性实证验证。","relevance":"该研究直接以LLM仿真人类决策行为，并与大规模跨国人类实验对照，评估个性化助推效果，高度契合研究者对LLM人类仿真可靠性及失效条件的关注。","inspiration":"借鉴LLM根据个体特征生成个性化干预并利用跨国调查进行对照验证的方法。｜可迁移至消费者金融产品选择中的个性化信息披露实验，如退休储蓄计划或保险产品选择。｜以LLM为不同人口特征群体生成个性化信息呈现方式，通过在线实验测量选择行为，并以真实市场数据或已有行为实验数据作为基准对照。"}},{"id":"2508.06635","version":2,"title":"Valid Inference with Imperfect Synthetic Data","zh_title":"不完美合成数据的有效推断","abstract":"Predictions and generations from large language models are increasingly being explored as an aid in limited data regimes, such as in computational social science and human subjects research. While prior technical work has mainly explored the potential to use model-predicted labels for unlabeled data in a principled manner, there is increasing interest in using large language models to generate entirely new synthetic samples (e.g., synthetic simulations), such as in responses to surveys. However, it remains unclear by what means practitioners can combine such data with real data and yet produce statistically valid conclusions upon them. In this paper, we introduce a new estimator based on generalized method of moments, providing a hyperparameter-free solution with strong theoretical guarantees to address this challenge. Intriguingly, we find that interactions between the moment residuals of synthetic data and those of real data (i.e., when they are predictive of each other) can greatly improve estimates of the target parameter. We validate the finite-sample performance of our estimator across different tasks in computational social science applications, demonstrating large empirical gains.","authors":["Yewon Byun","Shantanu Gupta","Zachary C. Lipton","Rachel Leah Childers","Bryan Wilder"],"categories":["cs.LG","cs.AI","stat.ML"],"primary_category":"cs.LG","announce_type":"new","date":"2025-08-08","first_seen":"2025-08-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.06635","pdf_url":"https://arxiv.org/pdf/2508.06635","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A5","B1","B2"],"tags":["合成数据","统计推断","计算社会科学"],"reason":"用LLM生成合成调查样本，结合真实数据做统计推断，直接涉及人类仿真与对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":45,"question":"如何将LLM生成的合成数据与真实数据结合，进行统计上有效的推断？","design":"本文提出一种基于广义矩估计（GMM）的框架，将LLM生成的合成样本（如模拟调查回答）与真实样本结合，通过引入合成数据与真实数据之间的相关结构来提高估计精度。","baseline":"真实人类标注样本，包含已标注协变量和结果的小规模数据集。","findings":"当合成数据的矩残差能预测真实数据的矩残差时，结合合成数据可提高估计精度并缩小置信区间；即使合成数据完全无信息，也不会损害渐近有效性。","reliability":"论文指出，若生成模型与真实分布不匹配，直接简单聚合合成数据会导致严重偏差；所提方法依赖于合成样本与真实样本之间的相关结构，若两者独立则无增益但也不损失。","relevance":"该研究直接针对用LLM生成合成调查样本并做统计推断的场景，提供了结合真实数据与合成数据的严谨方法，对关注人类仿真可靠性的研究者具有重要参考价值。","inspiration":"该方法通过广义矩估计将LLM合成样本与真实样本结合，利用合成数据矩残差对真实数据矩残差的预测能力来提高估计精度，同时保证即使合成数据无信息也不损害推断有效性。｜可迁移到政策评估中的调查实验，例如评估税收优惠宣传对家庭消费意愿的影响，其中LLM生成模拟调查回答作为合成数据。｜以LLM模拟的家庭为被试，处理为展示税收优惠信息，结果变量为自报消费意愿，用真实家庭调查数据作为基准，通过GMM框架结合合成与真实样本估计处理效应。"}},{"id":"2508.02766","version":2,"title":"The Generative Reasonable Person","zh_title":"生成式理性人","abstract":"This Article introduces the generative reasonable person, a new tool for estimating how ordinary people judge reasonableness. As claims about AI capabilities often outpace evidence, the Article proceeds empirically: adapting randomized controlled trials to large language models, it replicates three published studies of lay judgment across negligence, consent, and contract interpretation, drawing on nearly 10,000 simulated decisions. The findings reveal that models can replicate subtle patterns that run counter to textbook treatment. Like human subjects, models prioritize social conformity over cost-benefit analysis when assessing negligence, inverting the hierarchy that textbooks teach. They reproduce the paradox that material lies erode consent less than lies about a transaction's essence. And they track lay contract formalism, judging hidden fees more enforceable than fair. For two centuries, scholars have debated whether the reasonable person is empirical or normative, majoritarian or aspirational. But much of this debate assumed a constraint that no longer holds: that lay judgments are expensive to surface, slow to collect, and unavailable at scale. Generative reasonable people loosen that constraint. They offer judges empirical checks on elite intuition, give resource-constrained litigants access to simulated jury feedback, and let regulators pilot-test public comprehension, all at a fraction of survey costs. The reasonable person standard has long functioned as a vessel for judicial intuition precisely because the empirical baseline was missing. With that baseline now available, departures from lay understanding become transparent rather than hidden, a choice to be justified, not a fact to be assumed. Properly cabined, the generative reasonable person may become a dictionary for reasonableness judgments.","authors":["Yonathan A. Arbel"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-04","first_seen":"2025-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.02766","pdf_url":"https://arxiv.org/pdf/2508.02766","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B3"],"tags":["LLM仿真","人类被试替代","法律判断"],"reason":"用LLM模拟普通人判断，复现三项实验，有真实人类数据对照，涉及法律判断与政策评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":3,"question":"如何利用大语言模型模拟普通人判断，以复现法律中“理性人”标准的实证研究？","design":"使用大语言模型扮演普通人群，采用随机对照试验设计，复现三项已发表的关于过失、同意和合同解释的普通人判断研究，测量模型在近10,000次模拟决策中的判断模式。","baseline":"三项已发表研究中的真实人类被试数据，涵盖过失判断、同意判断和合同解释判断。","findings":"模型能复现与教科书相悖的微妙模式：在过失评估中优先考虑社会从众而非成本收益分析；再现了实质性谎言比交易本质谎言更少侵蚀同意的悖论；在合同解释中表现出普通人合同形式主义，认为隐藏费用比公平条款更具可执行性。","reliability":"论文未讨论","relevance":"该研究直接使用LLM模拟人类被试，复现法律判断实验并与真实人类数据对照，评估仿真可靠性，完全契合研究者对LLM人类仿真实验的关注，值得精读原文。","inspiration":"借鉴其用LLM复现已有人类实验并系统对比结果的方法，可验证LLM在特定领域的行为一致性。｜可迁移到消费者金融决策实验，如信贷条款理解、费用披露效果评估。｜以LLM模拟消费者，随机呈现不同披露格式的信贷合同，测量其理解度和选择行为，以真实消费者调查数据为基准进行对照验证。"}},{"id":"2508.05670","version":1,"title":"Can LLMs effectively provide game-theoretic-based scenarios for cybersecurity?","zh_title":"大语言模型能否有效提供基于博弈论的网络安全场景？","abstract":"Game theory has long served as a foundational tool in cybersecurity to test, predict, and design strategic interactions between attackers and defenders. The recent advent of Large Language Models (LLMs) offers new tools and challenges for the security of computer systems; In this work, we investigate whether classical game-theoretic frameworks can effectively capture the behaviours of LLM-driven actors and bots. Using a reproducible framework for game-theoretic LLM agents, we investigate two canonical scenarios -- the one-shot zero-sum game and the dynamic Prisoner's Dilemma -- and we test whether LLMs converge to expected outcomes or exhibit deviations due to embedded biases. Our experiments involve four state-of-the-art LLMs and span five natural languages, English, French, Arabic, Vietnamese, and Mandarin Chinese, to assess linguistic sensitivity. For both games, we observe that the final payoffs are influenced by agents characteristics such as personality traits or knowledge of repeated rounds. Moreover, we uncover an unexpected sensitivity of the final payoffs to the choice of languages, which should warn against indiscriminate application of LLMs in cybersecurity applications and call for in-depth studies, as LLMs may behave differently when deployed in different countries. We also employ quantitative metrics to evaluate the internal consistency and cross-language stability of LLM agents, to help guide the selection of the most stable LLMs and optimising models for secure applications.","authors":["Daniele Proverbio","Alessio Buscemi","Alessandro Di Stefano","The Anh Han","German Castignani","Pietro Liò"],"categories":["cs.CR","cs.AI","cs.CY","cs.GT"],"primary_category":"cs.CR","announce_type":"new","date":"2025-08-04","first_seen":"2025-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.05670","pdf_url":"https://arxiv.org/pdf/2508.05670","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM博弈行为","网络安全","社会模拟"],"reason":"用LLM agent模拟博弈行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":197,"question":"LLM在网络安全相关的博弈论场景中是否遵循经典博弈论预测，其行为受哪些因素影响？","design":"使用FAIRGAME框架，让四个主流LLM扮演博弈参与者，在一次性零和博弈和动态囚徒困境两种场景下进行交互，测试不同人格提示、是否知晓重复轮次等条件，并比较五种语言下的表现，测量最终收益和内部一致性。","baseline":"无对照","findings":"LLM的最终收益受人格特质和是否知晓重复轮次等智能体特征影响；收益对语言选择存在意外敏感性，不同语言下行为不同，提示在网络安全应用中需谨慎使用LLM。","reliability":"论文指出LLM行为偏离理论预测，且对语言敏感，在不同国家部署时可能表现不同，但未系统讨论失效的具体条件或局限。","relevance":"该研究用LLM模拟博弈行为，但无真实人类数据对照，属于社会模拟边界情形，与研究者关注的有基准对照的仿真研究不完全匹配，但提供了LLM行为偏差和语言敏感性的批判性证据，值得快速浏览。","inspiration":"该方法通过系统操纵人格提示和是否知晓重复轮次等智能体特征，并测量最终收益与内部一致性，可用于检验LLM行为偏差｜可迁移到政策公告的预期形成实验，研究不同信息框架下LLM模拟的投资者预期更新｜设计：以LLM为被试，处理为政策公告的表述语气（乐观/悲观）和是否明确政策调整频率，结果变量为预测的资产价格变动，以历史政策公告后的真实市场预期调查数据为对照"}},{"id":"2507.22049","version":1,"title":"Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel Predictions","zh_title":"验证基于生成式智能体的社会规范执行模型：从复现到新预测","abstract":"As large language models (LLMs) advance, there is growing interest in using them to simulate human social behavior through generative agent-based modeling (GABM). However, validating these models remains a key challenge. We present a systematic two-stage validation approach using social dilemma paradigms from psychological literature, first identifying the cognitive components necessary for LLM agents to reproduce known human behaviors in mixed-motive settings from two landmark papers, then using the validated architecture to simulate novel conditions. Our model comparison of different cognitive architectures shows that both persona-based individual differences and theory of mind capabilities are essential for replicating third-party punishment (TPP) as a costly signal of trustworthiness. For the second study on public goods games, this architecture is able to replicate an increase in cooperation from the spread of reputational information through gossip. However, an additional strategic component is necessary to replicate the additional boost in cooperation rates in the condition that allows both ostracism and gossip. We then test novel predictions for each paper with our validated generative agents. We find that TPP rates significantly drop in settings where punishment is anonymous, yet a substantial amount of TPP persists, suggesting that both reputational and intrinsic moral motivations play a role in this behavior. For the second paper, we introduce a novel intervention and see that open discussion periods before rounds of the public goods game further increase contributions, allowing groups to develop social norms for cooperation. This work provides a framework for validating generative agent models while demonstrating their potential to generate novel and testable insights into human social behavior.","authors":["Logan Cross","Nick Haber","Daniel L. K. Yamins"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2025-07-29","first_seen":"2025-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.22049","pdf_url":"https://arxiv.org/pdf/2507.22049","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM仿真","社会规范","行为博弈"],"reason":"用LLM智能体复现社会困境实验，与真实人类数据对照，验证仿真并预测新条件，直接…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":28,"question":"如何通过生成式智能体建模（GABM）复现并扩展社会困境中的人类行为，并验证其认知架构的有效性？","design":"使用LLM驱动的生成式智能体，通过组合记忆、人格、推理等认知组件构建不同架构，复现Jordan et al. (2016)的信任博弈第三方惩罚实验和Feinberg et al. (2014)的公共品博弈中流言与排斥实验，比较智能体行为与人类数据的统计效应，并利用验证后的架构模拟匿名惩罚和公开讨论等新条件。","baseline":"Jordan et al. (2016)的第三方惩罚信任博弈人类实验数据，以及Feinberg et al. (2014)的公共品博弈中流言与排斥对人类合作影响的数据。","findings":"复现第三方惩罚效应需要人格提示和心理理论组件，而公共品博弈中流言提升合作可被复现，但排斥与流言结合的额外合作提升需额外策略组件；匿名惩罚下第三方惩罚率显著下降但仍存在，表明声誉和内在道德动机共同驱动惩罚，公开讨论能进一步提升公共品贡献。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体复现经典社会困境实验，并与真实人类数据对照，验证仿真可靠性并生成新预测，完全契合研究者对LLM人类仿真、经济学实验复现及失效条件探索的兴趣，值得精读原文。","inspiration":"该方法通过组合记忆、人格、推理等认知组件构建不同LLM智能体架构，并与人类实验数据对照来验证仿真有效性，为经济金融实验提供了可借鉴的仿真验证框架。｜可迁移至资产定价实验中的羊群效应研究，利用LLM智能体模拟投资者在信息不对称下的决策行为。｜以LLM智能体为被试，处理为是否提供历史价格信息，结果变量为投资决策的羊群效应指数，对照真实人类资产定价实验数据。"}},{"id":"2507.21432","version":2,"title":"Towards Locally Deployable Fine-Tuned Causal Large Language Models for Mode Choice Behaviour","zh_title":"面向出行方式选择行为的本地可部署微调因果大语言模型研究","abstract":"This study investigates the adoption of open-access, locally deployable causal large language models (LLMs) for travel mode choice prediction and introduces LiTransMC, the first fine-tuned causal LLM developed for this task. We systematically benchmark eleven open-access LLMs (1-12B parameters) across three stated and revealed preference datasets, testing 396 configurations and generating over 79,000 mode choice decisions. Beyond predictive accuracy, we evaluate models generated reasoning using BERTopic for topic modelling and a novel Explanation Strength Index, providing the first structured analysis of how LLMs articulate decision factors in alignment with behavioural theory. LiTransMC, fine-tuned using parameter efficient and loss masking strategy, achieved a weighted F1 score of 0.6845 and a Jensen-Shannon Divergence of 0.000245, surpassing both untuned local models and larger proprietary systems, including GPT-4o with advanced persona inference and embedding-based loading, while also outperforming classical mode choice methods such as discrete choice models and machine learning classifiers for the same dataset. This dual improvement, i.e., high instant-level accuracy and near-perfect distributional calibration, demonstrates the feasibility of creating specialist, locally deployable LLMs that integrate prediction and interpretability. Through combining structured behavioural prediction with natural language reasoning, this work unlocks the potential for conversational, multi-task transport models capable of supporting agent-based simulations, policy testing, and behavioural insight generation. These findings establish a pathway for transforming general purpose LLMs into specialized and explainable tools for transportation research and policy formulation, while maintaining privacy, reducing cost, and broadening access through local deployment.","authors":["Tareq Alsaleh","Bilal Farooq"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-07-29","first_seen":"2025-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.21432","pdf_url":"https://arxiv.org/pdf/2507.21432","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM行为预测","交通方式选择","人类数据对照"],"reason":"用LLM预测出行方式选择，有真实人类数据对照，涉及交通行为仿真，方法可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":112,"question":"如何利用开源、可本地部署的因果大语言模型进行出行方式选择预测，并生成符合行为理论的决策解释？","design":"本研究不是人类仿真实验，而是使用11个开源LLM（1-12B参数）在三个陈述偏好和显示偏好数据集上预测出行方式选择，测试了396种配置，生成超过79,000个选择决策；并微调了LiTransMC模型，采用参数高效和损失掩码策略。","baseline":"三个陈述偏好和显示偏好数据集，包含真实人类的出行方式选择记录。","findings":"微调后的LiTransMC在加权F1分数（0.6845）和Jensen-Shannon散度（0.000245）上超越了未调优的本地模型和GPT-4o等大型闭源模型，同时优于传统离散选择模型和机器学习分类器；模型不仅能准确预测个体选择，还能生成与行为理论一致的自然语言推理。","reliability":"论文未讨论","relevance":"该研究用LLM预测出行方式选择，有真实人类数据对照，涉及交通行为仿真，方法可迁移至人类决策仿真，值得阅读原文以评估其作为人类被试替代品的潜力与局限。","inspiration":"该研究采用参数高效微调与损失掩码策略，在有限标注数据下提升开源LLM对个体选择行为的预测精度，并利用Jensen-Shannon散度等分布相似性指标评估模型输出与真实人类选择的整体拟合，这一设计可借鉴用于校准LLM仿真中的行为偏差。｜可迁移至消费者跨期选择实验，例如研究即时奖励与延迟奖励的权衡，或政策干预对储蓄行为的影响。｜以微调后的开源LLM作为虚拟被试，处理为不同利率或未来奖励的表述框架，结果变量为选择即时或延迟选项的概率，以真实实验室跨期选择数据作为基准，比较LLM与人类的选择分布及时间贴现率。"}},{"id":"2507.17024","version":1,"title":"Write, Rank, or Rate: Comparing Methods for Studying Visualization Affordances","zh_title":"写、排或评：比较研究可视化可供性的方法","abstract":"A growing body of work on visualization affordances highlights how specific design choices shape reader takeaways from information visualizations. However, mapping the relationship between design choices and reader conclusions often requires labor-intensive crowdsourced studies, generating large corpora of free-response text for analysis. To address this challenge, we explored alternative scalable research methodologies to assess chart affordances. We test four elicitation methods from human-subject studies: free response, visualization ranking, conclusion ranking, and salience rating, and compare their effectiveness in eliciting reader interpretations of line charts, dot plots, and heatmaps. Overall, we find that while no method fully replicates affordances observed in free-response conclusions, combinations of ranking and rating methods can serve as an effective proxy at a broad scale. The two ranking methodologies were influenced by participant bias towards certain chart types and the comparison of suggested conclusions. Rating conclusion salience could not capture the specific variations between chart types observed in the other methods. To supplement this work, we present a case study with GPT-4o, exploring the use of large language models (LLMs) to elicit human-like chart interpretations. This aligns with recent academic interest in leveraging LLMs as proxies for human participants to improve data collection and analysis efficiency. GPT-4o performed best as a human proxy for the salience rating methodology but suffered from severe constraints in other areas. Overall, the discrepancies in affordances we found between various elicitation methodologies, including GPT-4o, highlight the importance of intentionally selecting and combining methods and evaluating trade-offs.","authors":["Chase Stokes","Kylie Lin","Cindy Xiong Bearfield"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-07-22","first_seen":"2025-07-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.17024","pdf_url":"https://arxiv.org/pdf/2507.17024","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM代理","可视化可供性","方法比较"],"reason":"用GPT-4o替代人类被试评估图表解读，但主要目的是替代标注劳动，非严格仿真人…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":50,"question":"在可视化可供性研究中，不同启发方法（自由回答、排序、评分）以及大语言模型（GPT-4o）在多大程度上能复现人类对图表的解读？","design":"本研究并非以LLM仿真为核心，而是比较四种人类被试启发方法（自由回答、可视化排序、结论排序、显著性评分）在获取图表解读上的效果；随后以GPT-4o作为案例，探索用LLM模拟人类被试进行图表解读，并评估其在不同方法上的表现。","baseline":"以62名Prolific众包参与者对点图、折线图和热力图的自由回答结论作为人类基准，并基于因子分析构建了五类可供性空间。","findings":"没有任何单一方法能完全复现自由回答所观察到的可供性，但排序与评分方法的组合可作为大规模研究的有效替代。GPT-4o在显著性评分任务上最接近人类表现，但在其他方法上存在严重局限。","reliability":"论文指出GPT-4o仅在显著性评分方法上表现较好，在其他启发方法上存在严重限制；不同方法（包括GPT-4o）所揭示的可供性差异显著，需谨慎选择和组合方法并评估权衡。","relevance":"该研究直接使用GPT-4o作为人类被试替代品，并与真实人类数据进行对照，评估其仿真可靠性，且指出了失效条件，高度契合研究者对LLM仿真实验的批判性关注，值得阅读原文。","inspiration":"该研究系统比较了自由回答、排序、评分等不同启发方法在获取人类图表解读上的差异，并引入LLM作为被试与真实人类数据对照，这种多方法比较与仿真可靠性评估的设计值得借鉴。｜可迁移到经济金融中的信息处理与决策研究，例如投资者对财报图表或宏观经济数据可视化的解读如何影响预期形成。｜以专业投资者或众包被试为人类基准，让LLM与人类分别对同一组财务图表进行自由回答、排序和评分，比较其解读模式与分布，结果变量为解读结论的分类与评分，以真实人类数据作为对照基准评估LLM仿真的可靠性。"}},{"id":"2507.10933","version":1,"title":"Artificial Finance: How AI Thinks About Money","zh_title":"人工金融：AI如何思考金钱","abstract":"In this paper, we explore how large language models (LLMs) approach financial decision-making by systematically comparing their responses to those of human participants across the globe. We posed a set of commonly used financial decision-making questions to seven leading LLMs, including five models from the GPT series(GPT-4o, GPT-4.5, o1, o3-mini), Gemini 2.0 Flash, and DeepSeek R1. We then compared their outputs to human responses drawn from a dataset covering 53 nations. Our analysis reveals three main results. First, LLMs generally exhibit a risk-neutral decision-making pattern, favoring choices aligned with expected value calculations when faced with lottery-type questions. Second, when evaluating trade-offs between present and future, LLMs occasionally produce responses that appear inconsistent with normative reasoning. Third, when we examine cross-national similarities, we find that the LLMs' aggregate responses most closely resemble those of participants from Tanzania. These findings contribute to the understanding of how LLMs emulate human-like decision behaviors and highlight potential cultural and training influences embedded within their outputs.","authors":["Orhan Erdem","Ragavi Pobbathi Ashok"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2025-07-15","first_seen":"2025-07-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.10933","pdf_url":"https://arxiv.org/pdf/2507.10933","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","金融决策","跨文化对照"],"reason":"用LLM复现人类金融决策并与53国真实数据对照，评估仿真行为与偏差，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":26,"question":"LLM在金融决策中表现出怎样的风险偏好、跨期选择模式，其总体回答与哪个国家的人类被试最相似？","design":"将7个主流LLM（GPT-4o、GPT-4.5、o1、o3-mini、Gemini 2.0 Flash、DeepSeek R1）作为被试，向其呈现一组常用的金融决策问题（包括彩票型问题和跨期选择问题），收集模型回答，并与53国人类调查数据进行比较。","baseline":"来自Wang et al. (2017)的覆盖53个国家的全球金融决策调查数据。","findings":"LLM普遍表现出风险中性，倾向于根据期望值计算选择彩票型问题；在跨期选择中偶尔出现与规范推理不一致的回答；LLM的总体回答模式与坦桑尼亚参与者最相似。","reliability":"论文未讨论","relevance":"该研究直接用LLM复现人类金融决策并与跨国真实数据对照，评估仿真行为与偏差，高度契合研究者对LLM人类仿真实验的关注，值得阅读原文。","inspiration":"借鉴其将LLM作为被试、使用标准化金融决策问题并直接与跨国人类调查数据对照的设计思路。｜可迁移到消费者跨期选择、风险资产配置或政策公告预期形成等行为金融场景。｜以LLM为被试，施加不同框架的跨期选择或风险决策问题，测量其时间偏好与风险厌恶参数，并与真实跨国调查数据（如Falk et al., 2018的全球偏好数据）进行对照，检验LLM仿真的人群代表性。"}},{"id":"2507.10342","version":1,"title":"Using AI to replicate human experimental results: a motion study","zh_title":"使用AI复现人类实验结果：一项运动研究","abstract":"This paper explores the potential of large language models (LLMs) as reliable analytical tools in linguistic research, focusing on the emergence of affective meanings in temporal expressions involving manner-of-motion verbs. While LLMs like GPT-4 have shown promise across a range of tasks, their ability to replicate nuanced human judgements remains under scrutiny. We conducted four psycholinguistic studies (on emergent meanings, valence shifts, verb choice in emotional contexts, and sentence-emoji associations) first with human participants and then replicated the same tasks using an LLM. Results across all studies show a striking convergence between human and AI responses, with statistical analyses (e.g., Spearman's rho = .73-.96) indicating strong correlations in both rating patterns and categorical choices. While minor divergences were observed in some cases, these did not alter the overall interpretative outcomes. These findings offer compelling evidence that LLMs can augment traditional human-based experimentation, enabling broader-scale studies without compromising interpretative validity. This convergence not only strengthens the empirical foundation of prior human-based findings but also opens possibilities for hypothesis generation and data expansion through AI. Ultimately, our study supports the use of LLMs as credible and informative collaborators in linguistic inquiry.","authors":["Rosa Illan Castillo","Javier Valenzuela"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-07-14","first_seen":"2025-07-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.10342","pdf_url":"https://arxiv.org/pdf/2507.10342","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1"],"tags":["LLM仿真","人类数据对照","心理语言学"],"reason":"用LLM复现人类心理语言学实验，并与真实人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":15,"question":"大语言模型能否在心理语言学实验中复现人类对运动动词情感意义的细微判断，从而作为可靠的分析工具？","design":"使用ChatGPT o1模拟人类被试，对包含运动动词的时间表达句进行四项心理语言学任务（情感意义评分、效价变化、情绪语境下的动词选择、句子与表情符号关联），测量其评分和选择模式。","baseline":"59名英语母语者在相同任务上的真实行为数据，包括评分分布和分类选择。","findings":"LLM与人类反应高度一致，Spearman相关系数在0.73至0.96之间，表明评分模式和分类选择均强相关。微小分歧未改变整体解释性结论，支持LLM可作为人类实验的补充工具。","reliability":"论文未讨论","relevance":"该研究直接以真实人类数据为基准，验证LLM复现心理语言学实验的可靠性，符合研究者对仿真有效性评估的关注，值得阅读以了解具体对照方法和统计检验。","inspiration":"借鉴其将同一任务原样施加于LLM并与人类数据直接对比的设计，可迁移到消费者跨期选择或政策公告的预期形成实验，例如用LLM模拟消费者对“时间飞逝”类表述的耐心程度评分，以真实调查数据为基准检验LLM能否复现时间偏好。"}},{"id":"2507.09657","version":1,"title":"Negotiating Comfort: Simulating Personality-Driven LLM Agents in Shared Residential Social Networks","zh_title":"协商舒适度：在共享住宅社交网络中模拟个性驱动的LLM智能体","abstract":"We use generative agents powered by large language models (LLMs) to simulate a social network in a shared residential building, driving the temperature decisions for a central heating system. Agents, divided into Family Members and Representatives, consider personal preferences, personal traits, connections, and weather conditions. Daily simulations involve family-level consensus followed by building-wide decisions among representatives. We tested three personality traits distributions (positive, mixed, and negative) and found that positive traits correlate with higher happiness and stronger friendships. Temperature preferences, assertiveness, and selflessness have a significant impact on happiness and decisions. This work demonstrates how LLM-driven agents can help simulate nuanced human behavior where complex real-life human simulations are difficult to set.","authors":["Ann Nedime Nese Rende","Tolga Yilmaz","Özgür Ulusoy"],"categories":["cs.SI","cs.MA"],"primary_category":"cs.SI","announce_type":"new","date":"2025-07-13","first_seen":"2025-07-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.09657","pdf_url":"https://arxiv.org/pdf/2507.09657","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","个性驱动"],"reason":"用LLM agent模拟社会网络与决策，但无真实人类数据对照，属纯理论演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":143,"question":"在共享住宅的社会网络中，不同人格特质分布如何影响由LLM智能体驱动的供暖温度决策、幸福感与友谊关系？","design":"使用LLM驱动的生成式智能体模拟一栋共享住宅楼内的社会网络，智能体分为家庭成员和家庭代表，每日先进行家庭内部协商再在楼宇层面投票决定中央供暖温度；通过设置全正面、全负面、50%正面三种人格特质分布作为处理条件，测量智能体的幸福感、温度选择、友谊紧密度和决策结果。","baseline":"无对照","findings":"正面人格特质分布与更高的幸福感和更强的友谊相关；温度偏好、自信和无私程度对幸福感和决策有显著影响。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟社会网络中的协商与决策，属于人类仿真实验范畴，但缺乏真实人类数据对照，无法评估仿真的可靠性与偏差，与研究者关注的基准对照和批判性评估需求匹配度较低，不建议优先阅读原文。","inspiration":"该研究通过设定不同人格特质分布作为处理条件，测量智能体在协商网络中的幸福感与决策结果，这种基于智能体异质性施加处理的设计思路可借鉴｜可迁移至经济金融中的群体决策实验，如家庭内部资源分配或社区公共品供给中的偏好协商｜设计雏形：以LLM智能体模拟家庭成员，处理为不同风险偏好或时间贴现率的分布，结果变量为家庭储蓄率或消费组合，对照真实家庭调查数据如CFPS"}},{"id":"2507.07188","version":3,"title":"Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses","zh_title":"提示扰动揭示大语言模型调查响应中类人偏差","abstract":"Large Language Models (LLMs) are increasingly used as proxies for human subjects in social science surveys, but their reliability and susceptibility to known human-like response biases, such as central tendency, opinion floating and primacy bias are poorly understood. This work investigates the response robustness of LLMs in normative survey contexts, we test nine LLMs on questions from the World Values Survey (WVS), applying a comprehensive set of ten perturbations to both question phrasing and answer option structure, resulting in over 167,000 simulated survey interviews. In doing so, we not only reveal LLMs' vulnerabilities to perturbations but also show that all tested models exhibit a consistent recency bias, disproportionately favoring the last-presented answer option. While larger models are generally more robust, all models remain sensitive to semantic variations like paraphrasing and to combined perturbations. This underscores the critical importance of prompt design and robustness testing when using LLMs to generate synthetic survey data.","authors":["Jens Rupprecht","Georg Ahnert","Markus Strohmaier"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-07-09","first_seen":"2025-07-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.07188","pdf_url":"https://arxiv.org/pdf/2507.07188","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查偏差","算法保真度"],"reason":"直接研究LLM作为人类被试替代品的调查响应偏差，使用世界价值观调查真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"LLM在回答封闭式规范性调查问题时，对提示扰动是否稳健，以及是否表现出类似人类的回答偏差？","design":"用9个指令微调LLM扮演人类被试，对世界价值观调查的62个问题施加10种提示扰动（包括答案选项和问题措辞的修改），测量回答分布的变化，共进行167,400次模拟访谈。","baseline":"对照的真实人类数据来自世界价值观调查第7波（2017-2022）的核心变量，但论文主要比较扰动后回答分布与原始提示基线分布的差异。","findings":"所有模型均表现出明显的近因偏差，即不成比例地偏好最后一个选项；较大模型通常更稳健，但所有模型对语义改写和组合扰动仍敏感。","reliability":"论文指出，即使最大模型也对问题措辞变化敏感，提示设计和稳健性测试对使用LLM生成合成调查数据至关重要，但未讨论其他失效条件。","relevance":"该研究直接评估LLM作为人类被试替代品在调查中的偏差与可靠性，使用真实调查数据作为对照，并揭示近因偏差等关键失效模式，高度契合你的关注点，值得精读原文。","inspiration":"该方法通过系统施加提示扰动（如选项顺序、措辞改写）测量LLM回答分布变化，并设置原始提示基线作为对照，可借鉴用于稳健性检验设计｜可迁移到消费者通胀预期调查或政策公告解读实验，评估LLM模拟的预期形成是否对问卷设计敏感｜以LLM作为被试，施加不同措辞的通胀预期问题（如‘未来一年物价变化’ vs ‘通胀率’），结果变量为预期值分布，对照真实消费者预期调查数据（如密歇根大学调查）"}},{"id":"2507.08019","version":1,"title":"Signal or Noise? Evaluating Large Language Models in Resume Screening Across Contextual Variations and Human Expert Benchmarks","zh_title":"信号还是噪声？评估大语言模型在简历筛选中的表现：情境变化与人类专家基准","abstract":"This study investigates whether large language models (LLMs) exhibit consistent behavior (signal) or random variation (noise) when screening resumes against job descriptions, and how their performance compares to human experts. Using controlled datasets, we tested three LLMs (Claude, GPT, and Gemini) across contexts (No Company, Firm1 [MNC], Firm2 [Startup], Reduced Context) with identical and randomized resumes, benchmarked against three human recruitment experts. Analysis of variance revealed significant mean differences in four of eight LLM-only conditions and consistently significant differences between LLM and human evaluations (p < 0.01). Paired t-tests showed GPT adapts strongly to company context (p < 0.001), Gemini partially (p = 0.038 for Firm1), and Claude minimally (p > 0.1), while all LLMs differed significantly from human experts across contexts. Meta-cognition analysis highlighted adaptive weighting patterns that differ markedly from human evaluation approaches. Findings suggest LLMs offer interpretable patterns with detailed prompts but diverge substantially from human judgment, informing their deployment in automated hiring systems.","authors":["Aryan Varshney","Venkat Ram Reddy Ganuthula"],"categories":["cs.CL","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2025-07-08","first_seen":"2025-07-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.08019","pdf_url":"https://arxiv.org/pdf/2507.08019","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","人类对照","招聘决策"],"reason":"用LLM替代人类筛选简历，并与人类专家对照，属于人类仿真实验，但场景为招聘而非…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":38,"question":"大语言模型在简历筛选中的行为是稳定信号还是随机噪声？其评分与人类专家相比如何？","design":"用Claude、GPT、Gemini三个LLM扮演招聘者，在无公司、跨国公司、初创公司、简化上下文四种提示条件下对相同和随机化简历进行评分，测量评分均值和一致性，并与三位人类招聘专家对照。","baseline":"三位人类招聘专家在相同简历和提示条件下的评分。","findings":"LLM内部一致性有限，八种条件中四种存在显著均值差异；所有LLM评分均与人类专家显著不同，GPT对上下文适应最强，Claude最弱。元认知分析显示LLM的权重调整模式与人类差异明显。","reliability":"论文指出LLM评分与人类专家判断存在实质性分歧，在招聘等高利害场景中需谨慎部署，但未系统讨论失效条件与局限。","relevance":"该研究直接以人类专家为基准检验LLM在招聘决策中的仿真可靠性，属于批判性人类仿真研究，与研究者关注的经济学实验和政策评估场景高度相关，值得细读。","inspiration":"借鉴其多模型、多上下文条件与人类专家对照的设计，可迁移到信贷审批中的歧视研究。｜用多个LLM扮演信贷员，在不同银行类型（大行/小贷公司）和简化信息条件下审批贷款申请，以真实信贷员历史审批记录为基准，比较评分差异与偏见模式。"}},{"id":"2506.23610","version":1,"title":"Evaluating the Simulation of Human Personality-Driven Susceptibility to Misinformation with LLMs","zh_title":"评估大语言模型对人类人格驱动的错误信息易感性模拟","abstract":"Large language models (LLMs) make it possible to generate synthetic behavioural data at scale, offering an ethical and low-cost alternative to human experiments. Whether such data can faithfully capture psychological differences driven by personality traits, however, remains an open question. We evaluate the capacity of LLM agents, conditioned on Big-Five profiles, to reproduce personality-based variation in susceptibility to misinformation, focusing on news discernment, the ability to judge true headlines as true and false headlines as false. Leveraging published datasets in which human participants with known personality profiles rated headline accuracy, we create matching LLM agents and compare their responses to the original human patterns. Certain trait-misinformation associations, notably those involving Agreeableness and Conscientiousness, are reliably replicated, whereas others diverge, revealing systematic biases in how LLMs internalize and express personality. The results underscore both the promise and the limits of personality-aligned LLMs for behavioral simulation, and offer new insight into modeling cognitive diversity in artificial agents.","authors":["Manuel Pratelli","Marinella Petrocchi"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-30","first_seen":"2025-06-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.23610","pdf_url":"https://arxiv.org/pdf/2506.23610","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人格与行为","错误信息"],"reason":"用LLM模拟人格对错误信息易感性的影响，并与真实人类数据对照，评估仿真可靠性与…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":63,"question":"当赋予大语言模型明确的大五人格特征时，它们能否模拟人类在新闻辨别力（区分真假新闻标题准确性）上的表现？","design":"基于Calvillo等(2023)和Huang等(2024)的公开数据集，为每位人类参与者创建一个匹配的LLM智能体，通过人格注入管道赋予其对应的大五人格档案，让智能体对同一组新闻标题进行准确性评分，并比较合成数据与原始人类数据在人格-新闻辨别力关联上的异同。","baseline":"以Calvillo等(2023)的336名美国参与者数据为主要基准，该数据集包含参与者的大五人格档案和对真假新闻标题的准确性评分。","findings":"GPT-4o能复现部分人类特质-辨别力关联，如宜人性和尽责性与新闻辨别力正相关，但外向性和负面情绪性的模式与人类不一致；在仅分析假新闻时，高宜人性和高尽责性同样预测较低的易感性，但开放性的关联在不同模型设置下不稳定。","reliability":"论文承认某些人格特质（如外向性、负面情绪性）的模拟与人类模式存在系统性偏差，开放性-辨别力关联在不同模型参数和设置下不一致，表明LLM在表达人格时存在内在偏差，心理保真度有限。","relevance":"该研究直接使用LLM模拟人格对错误信息易感性的影响，并与真实人类数据严格对照，评估了仿真的可靠性与偏差，完全符合您对LLM人类仿真实验、经济学/政策评估场景及批判性失效条件分析的兴趣，值得精读原文。","inspiration":"该方法通过人格注入管道为LLM智能体赋予真实人类的大五人格档案，并严格对照人类基准数据评估仿真可靠性，值得借鉴其对照设计和偏差分析思路｜可迁移至消费者金融决策研究，如人格特质对投资风险偏好或过度借贷行为的影响｜以LLM智能体为被试，注入真实投资者的人格档案，让其评估不同风险等级的金融产品，结果变量为风险评分，对照真实投资者在相同产品上的风险偏好调查数据"}},{"id":"2506.23107","version":1,"title":"Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study","zh_title":"大语言模型能捕捉人类风险偏好吗？一项跨文化研究","abstract":"Large language models (LLMs) have made significant strides, extending their applications to dialogue systems, automated content creation, and domain-specific advisory tasks. However, as their use grows, concerns have emerged regarding their reliability in simulating complex decision-making behavior, such as risky decision-making, where a single choice can lead to multiple outcomes. This study investigates the ability of LLMs to simulate risky decision-making scenarios. We compare model-generated decisions with actual human responses in a series of lottery-based tasks, using transportation stated preference survey data from participants in Sydney, Dhaka, Hong Kong, and Nanjing. Demographic inputs were provided to two LLMs -- ChatGPT 4o and ChatGPT o1-mini -- which were tasked with predicting individual choices. Risk preferences were analyzed using the Constant Relative Risk Aversion (CRRA) framework. Results show that both models exhibit more risk-averse behavior than human participants, with o1-mini aligning more closely with observed human decisions. Further analysis of multilingual data from Nanjing and Hong Kong indicates that model predictions in Chinese deviate more from actual responses compared to English, suggesting that prompt language may influence simulation performance. These findings highlight both the promise and the current limitations of LLMs in replicating human-like risk behavior, particularly in linguistic and cultural settings.","authors":["Bing Song","Jianing Liu","Sisi Jian","Chenyang Wu","Vinayak Dixit"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-06-29","first_seen":"2025-06-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.23107","pdf_url":"https://arxiv.org/pdf/2506.23107","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","风险偏好","跨文化对照"],"reason":"用LLM仿真人类风险决策，与真实调查数据对照，评估跨文化偏差，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":64,"question":"大语言模型能否在跨文化背景下模拟人类的风险偏好？","design":"使用ChatGPT 4o和o1-mini两个模型，输入年龄、性别、教育、收入等人口统计信息，预测个体在彩票选择任务中的决策，并基于CRRA框架估计风险偏好。","baseline":"来自悉尼、达卡、香港和南京四个城市的真实交通陈述偏好调查数据，包含个体在彩票游戏中的实际选择。","findings":"两个模型均比人类更风险厌恶，o1-mini比4o更接近真实决策；在中文提示下，模型预测偏离实际的程度大于英文提示，表明语言影响仿真表现。","reliability":"论文指出模型在中文语境下偏差更大，且未完全复现人类风险行为，提示在语言和文化环境中的局限性。","relevance":"直接以真实人类数据为基准，评估LLM在风险决策仿真中的跨文化偏差，命中研究者关注的核心问题，值得精读原文。","inspiration":"该方法通过向LLM输入人口统计特征来模拟个体在彩票选择中的决策，并与真实调查数据对照，可用于评估模型在结构化风险任务中的行为偏差｜可迁移到金融决策中的风险偏好测量，如投资组合选择、保险购买或退休储蓄决策，检验LLM能否复现不同文化下个体的金融风险态度｜研究设计：以真实家庭金融调查数据为基准，向LLM输入年龄、收入、教育等特征，要求其在多组假设性投资选项中做出选择，结果变量为风险资产配置比例，对比模型预测与真实家庭行为的分布差异"}},{"id":"2506.21974","version":3,"title":"Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism","zh_title":"不要相信生成式智能体能模仿社交网络上的交流，除非你对其经验现实主义进行了基准测试","abstract":"The ability of Large Language Models (LLMs) to mimic human behavior triggered a plethora of computational social science research, assuming that empirical studies of humans can be conducted with AI agents instead. Since there have been conflicting research findings on whether and when this hypothesis holds, there is a need to better understand the differences in their experimental designs. We focus on replicating the behavior of social network users with the use of LLMs for the analysis of communication on social networks. First, we provide a formal framework for the simulation of social networks, before focusing on the sub-task of imitating user communication. We empirically test different approaches to imitate user behavior on X in English and German. Our findings suggest that social simulations should be validated by their empirical realism measured in the setting in which the simulation components were fitted. With this paper, we argue for more rigor when applying generative-agent-based modeling for social simulation.","authors":["Simon Münker","Nils Schwager","Achim Rettinger"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-27","first_seen":"2025-06-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.21974","pdf_url":"https://arxiv.org/pdf/2506.21974","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","社交网络","经验现实主义"],"reason":"用LLM仿真社交网络用户行为，有真实人类数据对照，并评估仿真可靠性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":29,"question":"LLM在多大程度上能忠实模仿社交网络（X平台）上不同用户群体的发帖、回复等沟通行为？","design":"基于X平台英文和德文政治讨论数据，用LLM（如Llama 3.1 70B）分别拟合政治人物（发帖者）和普通用户（回复者）的沟通行为，包括帖子生成、回复生成和回复可能性预测三个子任务，并比较不同建模方法的表现。","baseline":"真实人类数据：从X平台收集的2023年上半年美国和德国政治话语相关帖子（议员）和回复（普通用户），并标注主题。","findings":"LLM模仿用户沟通行为的效果因任务、语言和建模方法而异，并非总能忠实复现；社会仿真必须在其组件拟合的设定下验证经验现实主义。","reliability":"论文承认仿真框架存在简化，如未完全建模网络结构、内容过滤可能损失信息，且结论依赖于特定平台和语言数据集，泛化性有限。","relevance":"该研究直接回应了LLM仿真人类社交行为的可靠性问题，提供了有真实人类基准的实证评估，并指出仿真失效的条件，高度契合研究者对批判性仿真研究的关注。","inspiration":"该方法值得借鉴之处在于，它将LLM仿真分解为帖子生成、回复生成和回复可能性预测三个子任务，并分别与真实人类数据对比，从而定位仿真失效的具体环节｜该思路可迁移到政策公告的预期形成研究，例如分析央行沟通或财政政策声明如何影响市场参与者的预期与讨论｜研究设计：以LLM模拟投资者和分析师，处理为不同措辞或情感倾向的政策声明，结果变量为LLM生成的预期文本和市场反应预测，对照真实数据可来自央行声明后社交媒体（如Twitter/微博）上的实际讨论与市场调查预期数据"}},{"id":"2507.02919","version":1,"title":"ChatGPT is not A Man but Das Man: Representativeness and Structural Consistency of Silicon Samples Generated by Large Language Models","zh_title":"ChatGPT不是人而是常人：大语言模型生成硅样本的代表性与结构一致性","abstract":"Large language models (LLMs) in the form of chatbots like ChatGPT and Llama are increasingly proposed as \"silicon samples\" for simulating human opinions. This study examines this notion, arguing that LLMs may misrepresent population-level opinions. We identify two fundamental challenges: a failure in structural consistency, where response accuracy doesn't hold across demographic aggregation levels, and homogenization, an underrepresentation of minority opinions. To investigate these, we prompted ChatGPT (GPT-4) and Meta's Llama 3.1 series (8B, 70B, 405B) with questions on abortion and unauthorized immigration from the American National Election Studies (ANES) 2020. Our findings reveal significant structural inconsistencies and severe homogenization in LLM responses compared to human data. We propose an \"accuracy-optimization hypothesis,\" suggesting homogenization stems from prioritizing modal responses. These issues challenge the validity of using LLMs, especially chatbots AI, as direct substitutes for human survey data, potentially reinforcing stereotypes and misinforming policy.","authors":["Dai Li","Linzhuo Li","Huilian Sophie Qiu"],"categories":["cs.CL","cs.CY","cs.ET"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-25","first_seen":"2025-06-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.02919","pdf_url":"https://arxiv.org/pdf/2507.02919","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","调查数据对照","算法偏差"],"reason":"直接用LLM模拟人类调查回答，与ANES真实数据对照，揭示结构不一致和同质化，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:16","error":null,"has_summary":true,"summary":{"generated_at":"2025-06-25","rank":8,"question":"大语言模型生成的“硅样本”能否代表人类群体意见？","design":"使用ChatGPT (GPT-4) 和 Llama 3.1系列 (8B, 70B, 405B) 模型，输入ANES 2020中关于堕胎和非法移民的问题，生成模拟回答，并与真实人类数据对比，分析结构一致性和同质化程度。","baseline":"美国国家选举研究（ANES）2020的调查数据。","findings":"LLM回答存在显著的结构不一致性（不同人口聚合水平的准确率不一致）和严重的同质化（少数意见被低估）。作者提出“精度优化假说”，认为同质化源于模型优先输出众数回答。","reliability":"论文指出LLM可能误代表群体意见，存在结构不一致和同质化问题，挑战了直接替代人类调查数据的有效性，可能强化刻板印象并误导政策。","relevance":"高度相关。该研究直接评估了LLM仿真人类意见的可靠性与偏差，有真实人类数据对照，并指出了失效条件（结构不一致和同质化），符合研究者的核心关注点，值得精读原文。","inspiration":"该方法通过将LLM作为被试，输入真实调查问题并对比其回答与人类基准数据，来评估仿真的一致性和偏差，可借鉴其对照设计和偏差度量方式｜可迁移到政策预期形成的实验，例如研究公众对央行利率决议声明的通胀预期反应｜以GPT-4等LLM为被试，输入央行政策声明文本，测量其输出的通胀预期值，并与密歇根消费者调查的真实预期数据对比，检验LLM是否复现预期分布及同质化偏差"}},{"id":"2506.21587","version":2,"title":"A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies","zh_title":"基于大语言模型的舆论仿真跨文化比较：评估中美模型在多元社会上的表现","abstract":"This study evaluates the ability of DeepSeek, an open-source large language model (LLM), to simulate public opinions in comparison to LLMs developed by major tech companies. By comparing DeepSeek-R1 and DeepSeek-V3 with Qwen2.5, GPT-4o, and Llama-3.3 and utilizing survey data from the American National Election Studies (ANES) and the Zuobiao dataset of China, we assess these models' capacity to predict public opinions on social issues in both China and the United States, highlighting their comparative capabilities between countries. Our findings indicate that DeepSeek-V3 performs best in simulating U.S. opinions on the abortion issue compared to other topics such as climate change, gun control, immigration, and services for same-sex couples, primarily because it more accurately simulates responses when provided with Democratic or liberal personas. For Chinese samples, DeepSeek-V3 performs best in simulating opinions on foreign aid and individualism but shows limitations in modeling views on capitalism, particularly failing to capture the stances of low-income and non-college-educated individuals. It does not exhibit significant differences from other models in simulating opinions on traditionalism and the free market. Further analysis reveals that all LLMs exhibit the tendency to overgeneralize a single perspective within demographic groups, often defaulting to consistent responses within groups. These findings highlight the need to mitigate cultural and demographic biases in LLM-driven public opinion modeling, calling for approaches such as more inclusive training methodologies.","authors":["Weihong Qi","Fan Huang","Jisun An","Haewoon Kwak"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-17","first_seen":"2025-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.21587","pdf_url":"https://arxiv.org/pdf/2506.21587","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","舆论模拟","跨文化比较"],"reason":"用LLM仿真中美公众舆论，有ANES和Zuobiao真实调查数据对照，涉及社会…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":15,"question":"LLM的文化出身是否带来舆论模拟的“主场优势”？不同模型在中美社会议题上存在怎样的文化与人口偏差？","design":"用DeepSeek-R1、DeepSeek-V3、Qwen2.5、GPT-4o、Llama-3.3，根据ANES和Zuobiao数据集中个体的种族、性别、年龄、收入、教育等人口属性构造提示，让模型回答社会议题问卷，聚合后比较模拟分布与真实调查分布。","baseline":"2020年美国国家选举研究（ANES）的2457人调查数据，以及中国“左标”数据集随机抽取的2000人调查数据。","findings":"DeepSeek-V3在模拟美国堕胎议题上表现最好，主要因为能较准地模拟民主党/自由派，但共和党/保守派模拟差；所有模型都倾向于在人口群体内过度概括单一观点，尤其在中国资本主义议题上，Qwen2.5通过输出统一回答虚高准确率。","reliability":"论文指出所有模型均存在显著的文化与人口偏差，难以捕捉特定群体的细微观点，常退化为刻板或过度概括的回答；未讨论提示设计、模型版本或数据集时效性等可能影响结论的因素。","relevance":"直接回应LLM仿真人类舆论的可靠性与偏差问题，有中美真实调查数据对照，揭示模型在跨文化情境下系统性失效的模式，对理解仿真在经济学/政策评估中的局限有重要参考价值，强烈建议阅读原文。","inspiration":"该方法通过将人口属性编码为提示来模拟个体回答，并对比聚合分布与真实调查数据，为评估仿真偏差提供了可操作的对照框架｜可迁移到消费者信心调查或通胀预期形成的仿真研究，检验LLM能否复现不同人口群体的预期差异｜以LLM作为被试，输入年龄、收入、教育等属性，让其预测未来通胀率，结果变量为预期值分布，以密歇根大学消费者调查的微观数据作为真实基准，比较模拟与真实分布的偏差模式"}},{"id":"2506.14611","version":2,"title":"Exploring MLLMs Perception of Network Visualization Principles","zh_title":"探索多模态大语言模型对网络可视化原则的感知","abstract":"In this paper, we test whether Multimodal Large Language Models (MLLMs) can match human-subject performance in tasks involving the perception of properties in network layouts. Specifically, we replicate a human-subject experiment about perceiving quality (namely stress) in network layouts using GPT-4o, Gemini-2.5 and Qwen2.5. Our experiments show that giving MLLMs the same study information as trained human participants yields performance comparable to that of human experts and exceeds that of untrained non-experts. Additionally, we show that prompt engineering that deviates from the human-subject experiment can lead to better-than-human performance in some settings. Interestingly, like human subjects, the MLLMs seem to rely on visual proxies rather than computing the actual value of stress, indicating some sense or facsimile of perception. Explanations from the models are similar to those used by the human participants (e.g., an even distribution of nodes and uniform edge lengths).","authors":["Jacob Miller","Markus Wallinger","Ludwig Felder","Timo Brand","Henry Förster","Johannes Zink","Chunyang Chen","Stephen Kobourov"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-06-17","first_seen":"2025-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.14611","pdf_url":"https://arxiv.org/pdf/2506.14611","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类被试替代","网络感知实验"],"reason":"用MLLM复现人类网络布局感知实验，与真实人类数据对照，评估仿真可靠性并指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":65,"question":"多模态大语言模型在判断网络布局应力时，能否达到人类被试的感知表现？","design":"用GPT-4o、Gemini-2.5和Qwen2.5模拟人类被试，复现Mooney等人的配对刺激实验：向模型展示同一网络的两幅节点链接图，要求选择应力更低的图或判断两者相似，并比较不同提示工程（与人类相同指令 vs. 偏离人类指令）下的表现。","baseline":"Mooney等人实验中的人类被试数据，包括未训练新手、训练新手和专家三组。","findings":"给予与人类训练参与者相同的信息时，MLLM的表现与人类专家相当，优于未训练新手；通过偏离人类实验的提示工程，在某些设置下可获得超越人类的表现。MLLM与人类相似，依赖视觉代理（如节点均匀分布、边长一致）而非实际计算应力值。","reliability":"论文未讨论","relevance":"该研究直接用MLLM复现人类网络感知实验，并与真实人类数据对照，评估仿真可靠性，同时指出提示工程可导致超越人类的表现，符合您对LLM仿真人类实验、基准对照及失效条件探索的关注，值得精读原文。","inspiration":"该方法借鉴了用多模态大语言模型复现人类感知实验，并通过提示工程模拟不同信息条件（如训练新手、专家）来检验仿真表现，同时以真实人类数据作为基准对照。｜可迁移到金融图表解读与投资决策实验，例如研究投资者如何从网络关系图（如持股网络、供应链网络）中提取风险信息并形成投资判断。｜以GPT-4o等MLLM为被试，展示同一组持股网络的不同布局图，要求选择更易读或风险更清晰的图，结果变量为选择准确率与反应时间，对照真实投资者在相同任务上的行为数据。"}},{"id":"2506.21574","version":1,"title":"Digital Gatekeepers: Exploring Large Language Model's Role in Immigration Decisions","zh_title":"数字守门人：探索大语言模型在移民决策中的作用","abstract":"With globalization and increasing immigrant populations, immigration departments face significant work-loads and the challenge of ensuring fairness in decision-making processes. Integrating artificial intelligence offers a promising solution to these challenges. This study investigates the potential of large language models (LLMs),such as GPT-3.5 and GPT-4, in supporting immigration decision-making. Utilizing a mixed-methods approach,this paper conducted discrete choice experiments and in-depth interviews to study LLM decision-making strategies and whether they are fair. Our findings demonstrate that LLMs can align their decision-making with human strategies, emphasizing utility maximization and procedural fairness. Meanwhile, this paper also reveals that while ChatGPT has safeguards to prevent unintentional discrimination, it still exhibits stereotypes and biases concerning nationality and shows preferences toward privileged group. This dual analysis highlights both the potential and limitations of LLMs in automating and enhancing immigration decisions.","authors":["Yicheng Mao","Yang Zhao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-15","first_seen":"2025-06-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.21574","pdf_url":"https://arxiv.org/pdf/2506.21574","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","决策实验","公平性评估"],"reason":"用LLM模拟移民决策并与人类策略对照，评估公平性与偏差，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":59,"question":"大语言模型在移民决策中的决策策略是否与人类一致，以及其决策是否公平、是否存在偏见？","design":"使用GPT-3.5和GPT-4作为被试，复现Hainmueller和Hopkins (2015)的离散选择实验，生成10,000个随机移民决策场景，并辅以深度访谈，分析模型的决策策略、公平性和偏见。","baseline":"对照Hainmueller和Hopkins (2015)中的人类决策数据，比较LLM与人类在移民决策中的策略一致性。","findings":"LLM的决策策略与人类相似，强调效用最大化和程序公平；但ChatGPT虽设有防止无意歧视的机制，仍表现出基于国籍的刻板印象和偏见，并偏好特权群体。","reliability":"论文未讨论","relevance":"该研究直接用LLM模拟人类移民决策，并与真实人类实验数据对照，评估公平性与偏差，属于典型的LLM人类仿真研究，且涉及政策评估场景，值得精读原文。","inspiration":"该方法借鉴了离散选择实验与真实人类基准对照的设计，可系统评估LLM在政策决策中的策略一致性与偏差｜可迁移至信贷审批歧视研究，检验LLM是否复现人类审批中的种族、性别或收入偏见｜以LLM作为信贷审批官，处理随机生成的贷款申请人档案（操纵种族、收入等特征），结果变量为批准决策，对照真实银行审批数据或审计研究结果"}},{"id":"2506.12664","version":2,"title":"Behavioral Generative Agents for Energy Operations","zh_title":"用于能源运营的行为生成式智能体","abstract":"Problem definition: Accurately modeling consumer behavior in energy operations is challenging due to uncertainty, behavioral heterogeneity, and limited empirical data-particularly in low-frequency, high-impact events. While generative AI trained on large-scale human data offers new opportunities to study decision behavior, its role in operational applications remains unclear. We examine how generative agents can support customer behavior discovery in energy operations, complementing rather than replacing human-based experiments. Methodology/results: We introduce a novel approach leveraging generative agents-artificial agents powered by large language models-to simulate sequential customer decisions under dynamic electricity prices and outage risks. We find that these agents behave more optimally and rationally in simpler market scenarios, while their performance becomes more variable and suboptimal as task complexity rises. Furthermore, the agents exhibit heterogeneous customer preferences, consistently maintaining distinct, persona-driven reasoning patterns in both operational decisions and textual reasoning. Comparisons with dynamic programming and greedy policy benchmarks show alignment between specific personas and distinct heuristic decision policies. In low-frequency, high-impact events such as blackouts, agents prioritize energy reliability over cost or profit, demonstrating their ability to uncover behavioral patterns beyond the rigidity of traditional mathematical models. Managerial Implications: Our findings suggest that behavioral generative agents can serve as scalable and flexible tools for studying consumer behavior in energy operations. By enabling controlled experiments across heterogeneous customer types and rare events, these agents can enhance the design of energy management systems and support more informed analysis of energy policies and incentive programs.","authors":["Cong Chen","Omer Karaduman","Xu Kuang"],"categories":["cs.AI","eess.SY"],"primary_category":"cs.AI","announce_type":"new","date":"2025-06-14","first_seen":"2025-06-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.12664","pdf_url":"https://arxiv.org/pdf/2506.12664","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","能源经济","行为建模"],"reason":"用LLM agent模拟消费者能源决策，涉及经济场景，有批判性讨论，但缺真实人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":114,"question":"生成式智能体能否在动态电价和停电风险下模拟消费者能源管理决策，并揭示不同任务难度和用户画像下的行为模式？","design":"使用大语言模型驱动的生成式智能体，通过提示词赋予其不同用户画像（如谨慎的感性者、逐利的理性思考者、现实主义者），在模拟的家庭能源管理场景中逐日决定是否对家用电池充电或放电，观察其决策表现和文本推理，并与动态规划最优策略和贪婪启发式基准进行比较。","baseline":"无对照","findings":"智能体在简单市场场景中表现接近最优，但随着任务复杂度上升，决策质量下降且变异性增大；不同画像的智能体展现出异质且稳定的偏好，在停电等低频高影响事件中优先保障能源可靠性而非成本或利润。","reliability":"论文承认缺乏真实人类数据作为对照，仅采用数学基准（动态规划、贪婪策略）评估智能体推理质量，且智能体在复杂任务中表现下降，表明其可靠性受任务难度影响。","relevance":"该研究用LLM智能体模拟消费者能源决策，属于经济学实验场景，但缺少真实人类基准数据，批判性讨论有限，与研究者关注的高质量人类仿真和失效条件分析存在差距，可酌情阅读以了解方法。","inspiration":"该方法通过提示词赋予LLM不同用户画像来模拟异质决策偏好，可用于经济金融实验中的处理操纵｜可迁移到消费者跨期选择研究，如不同利率或补贴政策下的储蓄与消费决策｜以LLM智能体为被试，处理为不同利率条件或政策信息提示，结果变量为模拟的消费-储蓄分配，对照真实家庭调查数据（如PSID）中的跨期选择弹性"}},{"id":"2507.19495","version":1,"title":"Simulating Human Behavior with the Psychological-mechanism Agent: Integrating Feeling, Thought, and Action","zh_title":"基于心理机制代理的人类行为模拟：整合感受、思维与行动","abstract":"Generative agents have made significant progress in simulating human behavior, but existing frameworks often simplify emotional modeling and focus primarily on specific tasks, limiting the authenticity of the simulation. Our work proposes the Psychological-mechanism Agent (PSYA) framework, based on the Cognitive Triangle (Feeling-Thought-Action), designed to more accurately simulate human behavior. The PSYA consists of three core modules: the Feeling module (using a layer model of affect to simulate changes in short-term, medium-term, and long-term emotions), the Thought module (based on the Triple Network Model to support goal-directed and spontaneous thinking), and the Action module (optimizing agent behavior through the integration of emotions, needs and plans). To evaluate the framework's effectiveness, we conducted daily life simulations and extended the evaluation metrics to self-influence, one-influence, and group-influence, selection five classic psychological experiments for simulation. The results show that the PSYA framework generates more natural, consistent, diverse, and credible behaviors, successfully replicating human experimental outcomes. Our work provides a richer and more accurate emotional and cognitive modeling approach for generative agents and offers an alternative to human participants in psychological experiments.","authors":["Qing Dong","Pengyuan Liu","Dong Yu","Chen Kang"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2025-06-04","first_seen":"2025-06-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.19495","pdf_url":"https://arxiv.org/pdf/2507.19495","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2"],"tags":["人类行为仿真","心理学实验复现","生成式代理"],"reason":"用LLM代理复现经典心理学实验，并与真实人类数据对照，直接替代人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":30,"question":"能否基于认知三角（感受-思维-行动）构建心理机制智能体框架，更真实地模拟人类日常行为并复现经典心理学实验？","design":"提出PSYA框架，包含感受（分层情感模型）、思维（三重网络模型支持目标导向与自发思维）、行动（整合情绪、需求与计划）三个模块，基于Llama-3-70B构建智能体，在虚拟小镇中模拟8个智能体的日常生活，并选取5个经典心理学实验进行仿真，测量行为自然度、一致性、多样性等指标。","baseline":"以经典心理学实验的真实人类结果作为对照基准，验证智能体能否复现人类实验数据。","findings":"PSYA框架能生成更自然、一致、多样和可信的行为，成功复现了所选经典心理学实验的结果；分层情感和自发思维模块显著提升了行为真实性。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体替代人类被试复现经典心理学实验，并与真实人类数据对照，高度契合研究者对经济学实验和政策评估场景中仿真可靠性的关注，值得精读原文。","inspiration":"借鉴PSYA框架的分层情感与自发思维模块设计，在LLM智能体中嵌入情绪和认知偏差来模拟经济决策的心理机制｜可迁移到消费者跨期选择实验，考察情绪波动对时间偏好一致性的影响｜以LLM智能体为被试，施加情绪启动处理（如积极/消极文本诱导），测量跨期选择中的折现率变化，并与真实人类实验数据（如经典双曲折现研究）对照"}},{"id":"2505.21997","version":1,"title":"Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data","zh_title":"利用访谈信息引导大语言模型建模调查回答：AI生成数据与人类数据的比较洞察","abstract":"Mixed methods research integrates quantitative and qualitative data but faces challenges in aligning their distinct structures, particularly in examining measurement characteristics and individual response patterns. Advances in large language models (LLMs) offer promising solutions by generating synthetic survey responses informed by qualitative data. This study investigates whether LLMs, guided by personal interviews, can reliably predict human survey responses, using the Behavioral Regulations in Exercise Questionnaire (BREQ) and interviews from after-school program staff as a case study. Results indicate that LLMs capture overall response patterns but exhibit lower variability than humans. Incorporating interview data improves response diversity for some models (e.g., Claude, GPT), while well-crafted prompts and low-temperature settings enhance alignment between LLM and human responses. Demographic information had less impact than interview content on alignment accuracy. These findings underscore the potential of interview-informed LLMs to bridge qualitative and quantitative methodologies while revealing limitations in response variability, emotional interpretation, and psychometric fidelity. Future research should refine prompt design, explore bias mitigation, and optimize model settings to enhance the validity of LLM-generated survey data in social science research.","authors":["Jihong Zhang","Xinya Liang","Anqi Deng","Nicole Bonge","Lin Tan","Ling Zhang","Nicole Zarrett"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-05-28","first_seen":"2025-05-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.21997","pdf_url":"https://arxiv.org/pdf/2505.21997","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查回答","人类数据对照"],"reason":"用访谈引导LLM生成调查回答，与真实人类数据对照，评估仿真可靠性与偏差，直接命…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":8,"question":"以个人访谈信息引导的大语言模型能否可靠预测人类在标准化问卷上的回答？","design":"使用Claude、GPT等大语言模型，输入课后项目工作人员的个人访谈文本和人口统计信息作为提示，生成其在运动行为调节问卷（BREQ）上的Likert量表回答，并与真实人类回答进行对比。","baseline":"同一批课后项目工作人员在BREQ问卷上的真实人类回答。","findings":"LLM能捕捉整体回答模式，但变异性低于人类；加入访谈内容可提升部分模型的回答多样性，精心设计的提示和低温度参数能增强对齐度。项目分析显示LLM在反向措辞题目上偏差较大，且难以重建问卷的心理测量结构。","reliability":"论文承认LLM在回答变异性、情感解读和心理测量保真度方面存在局限，且访谈内容的相关性比长度更重要，不同受访者间模型表现存在差异。","relevance":"该研究直接以真实人类调查数据为基准，评估访谈引导的LLM仿真回答的可靠性与偏差，与研究者关注的LLM人类仿真实验高度吻合，值得精读。","inspiration":"借鉴用个人深度访谈作为LLM输入来模拟特定个体问卷回答的设计，可迁移到消费者信心调查或通胀预期形成等经济金融场景。｜设计一个研究：以真实消费者的财务访谈文本作为处理，让LLM生成对未来通胀或就业的预期评分，结果变量为预期值，以实际调查数据（如密歇根消费者调查）作为对照基准。"}},{"id":"2505.21371","version":2,"title":"When Experimental Economics Meets Large Language Models: Evidence-based Tactics","zh_title":"当实验经济学遇上大语言模型：基于证据的实践策略","abstract":"Advancements in large language models (LLMs) have sparked a growing interest in measuring and understanding their behavior through experimental economics. However, there is still a lack of established guidelines for designing economic experiments for LLMs. Inspired by principles from experimental economics with insights from LLM research in artificial intelligence, we outline key considerations in the experimental design and implementation stage, and perform two sets of experiments to assess the impact of these considerations on LLMs' responses. Based on our findings, we discuss seven practical tactics for conducting experiments with LLMs. Our study enhances the design, replicability, and generalizability of LLM experiments, and broadens the scope of experimental economics in the digital age.","authors":["Shu Wang","Zijun Yao","Shuhuai Zhang","Jianuo Gai","Tracy Xiao Liu","Songfa Zhong"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-05-27","first_seen":"2025-05-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.21371","pdf_url":"https://arxiv.org/pdf/2505.21371","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","A4","B4"],"tags":["LLM实验设计","方法论","实验经济学"],"reason":"论文探讨用实验经济学方法设计LLM实验，评估实验设计对LLM响应的影响，提供方…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":115,"question":"如何设计针对大语言模型的经济学实验，以及实验设计参数（如提示格式、对话类型、响应方式）如何影响LLM的行为表现？","design":"本文并非直接进行人类仿真，而是通过两个案例研究（预算决策和行为博弈），在GPT-4o、DeepSeek-V3、Llama-3.1-8B、Qwen2.5-7B四个模型上，系统操纵实验设计因素（如分配的角色、单轮/多轮对话、开放/封闭式回答），测量LLM的理性程度和偏好输出。","baseline":"无对照","findings":"分配的职业角色显著影响LLM的偏好，但不影响理性；单轮对话方式会降低Llama和Qwen的理性，但对GPT和DeepSeek无影响；限制为选择题而非开放式回答会降低Llama和Qwen的理性，并显著改变所有模型在约半数行为游戏中的输出。","reliability":"论文未讨论","relevance":"本文虽未直接进行人类仿真，但系统评估了LLM实验设计参数对行为输出的影响，为构建可靠的人类仿真实验提供了方法论基础，值得阅读原文以了解其提出的七条实用策略。","inspiration":"借鉴其系统操纵实验设计参数（角色分配、对话轮次、回答格式）来评估LLM行为稳健性的方法论，可迁移至政策公告预期形成的仿真研究，例如以LLM为被试、操纵公告措辞与信息呈现方式、测量通胀预期并对比真实调查数据。"}},{"id":"2505.17479","version":1,"title":"Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions","zh_title":"Twin-2K-500：基于2000余人对500余题回答构建数字孪生的数据集","abstract":"LLM-based digital twin simulation, where large language models are used to emulate individual human behavior, holds great promise for research in AI, social science, and digital experimentation. However, progress in this area has been hindered by the scarcity of real, individual-level datasets that are both large and publicly available. This lack of high-quality ground truth limits both the development and validation of digital twin methodologies. To address this gap, we introduce a large-scale, public dataset designed to capture a rich and holistic view of individual human behavior. We survey a representative sample of $N = 2,058$ participants (average 2.42 hours per person) in the US across four waves with 500 questions in total, covering a comprehensive battery of demographic, psychological, economic, personality, and cognitive measures, as well as replications of behavioral economics experiments and a pricing survey. The final wave repeats tasks from earlier waves to establish a test-retest accuracy baseline. Initial analyses suggest the data are of high quality and show promise for constructing digital twins that predict human behavior well at the individual and aggregate levels. By making the full dataset publicly available, we aim to establish a valuable testbed for the development and benchmarking of LLM-based persona simulations. Beyond LLM applications, due to its unique breadth and scale the dataset also enables broad social science research, including studies of cross-construct correlations and heterogeneous treatment effects.","authors":["Olivier Toubia","George Z. Gui","Tianyi Peng","Daniel J. Merlau","Ang Li","Haozhe Chen"],"categories":["cs.CY","cs.AI","cs.HC","econ.EM"],"primary_category":"cs.CY","announce_type":"new","date":"2025-05-23","first_seen":"2025-05-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.17479","pdf_url":"https://arxiv.org/pdf/2505.17479","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["数字孪生","人类仿真","行为经济学"],"reason":"直接构建LLM数字孪生仿真个体行为，含真实人类对照数据，涉及行为经济学实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:08","error":null,"has_summary":true,"summary":{"generated_at":"2025-05-23","rank":2,"question":"构建并公开一个大规模、多维度的人类行为数据集，用于开发和验证基于LLM的数字孪生仿真。","design":"该研究不是仿真实验，而是数据集构建。对2058名美国代表性样本进行4轮调查，每名被试平均2.42小时，共500题，涵盖人口统计、心理、经济偏好、人格、认知测试，并复现了行为经济学实验（包括组间和组内设计）及定价调查。","baseline":"人类基准：2058名真实被试的问卷调查数据，包括行为经济学实验的复现结果，以及第四轮重复前几轮任务以建立重测信度基线。","findings":"数据质量良好：测量间相关性具有表面效度，复现了行为经济学文献中几乎所有已知结果，重测信度稳健。初步数字孪生预测在个体和聚合水平上表现良好，但具体准确率未在摘要中给出。","reliability":"论文未讨论失效条件与局限，但指出数据集可用于评估数字孪生的可靠性，且已有研究提示LLM仿真可能受提示架构、混杂变量和代表性偏差影响。","relevance":"高度相关。该数据集提供了大规模、多维度的人类行为基准，可直接用于评估LLM数字孪生在复现调查回答、实验行为和决策模式上的准确性，尤其包含经济学实验复现，适合研究者验证仿真可靠性及偏差条件。","inspiration":"借鉴其大规模多维度调查设计，在同一被试内复现行为经济学实验并建立重测信度基线，为仿真验证提供个体与聚合层面的对照基准｜可迁移到资产定价实验中的风险偏好与预期形成研究，检验LLM数字孪生能否复现真实投资者的风险态度和价格预期分布｜以真实投资者为被试，收集其风险偏好问卷、资产选择实验及市场预期数据作为基准，用LLM基于相同问卷生成数字孪生，比较两者在风险资产配置和预期回报估计上的个体与分布一致性"}},{"id":"2505.12923","version":2,"title":"The Traitors: Deception and Trust in Multi-Agent Language Model Simulations","zh_title":"《叛徒》：多智能体语言模型仿真中的欺骗与信任","abstract":"As AI systems increasingly assume roles where trust and alignment with human values are essential, understanding when and why they engage in deception has become a critical research priority. We introduce The Traitors, a multi-agent simulation framework inspired by social deduction games, designed to probe deception, trust formation, and strategic communication among large language model (LLM) agents under asymmetric information. A minority of agents the traitors seek to mislead the majority, while the faithful must infer hidden identities through dialogue and reasoning. Our contributions are: (1) we ground the environment in formal frameworks from game theory, behavioral economics, and social cognition; (2) we develop a suite of evaluation metrics capturing deception success, trust dynamics, and collective inference quality; (3) we implement a fully autonomous simulation platform where LLMs reason over persistent memory and evolving social dynamics, with support for heterogeneous agent populations, specialized traits, and adaptive behaviors. Our initial experiments across DeepSeek-V3, GPT-4o-mini, and GPT-4o (10 runs per model) reveal a notable asymmetry: advanced models like GPT-4o demonstrate superior deceptive capabilities yet exhibit disproportionate vulnerability to others' falsehoods. This suggests deception skills may scale faster than detection abilities. Overall, The Traitors provides a focused, configurable testbed for investigating LLM behavior in socially nuanced interactions. We position this work as a contribution toward more rigorous research on deception mechanisms, alignment challenges, and the broader social reliability of AI systems.","authors":["Pedro M. P. Curvo"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2025-05-19","first_seen":"2025-05-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.12923","pdf_url":"https://arxiv.org/pdf/2505.12923","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","欺骗检测","社会推理"],"reason":"多智能体社会模拟但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":199,"question":"在信息不对称的多智能体社交推理游戏中，大语言模型代理何时以及如何产生欺骗行为，其欺骗能力与检测能力如何随模型能力变化？","design":"构建名为“The Traitors”的多智能体仿真框架，让不同LLM（DeepSeek-V3、GPT-4o-mini、GPT-4o）扮演少数“叛徒”和多数“忠臣”，叛徒知晓身份并试图欺骗，忠臣通过对话和推理识别叛徒；测量欺骗成功率、信任动态和集体推理质量等指标。","baseline":"无对照","findings":"高级模型如GPT-4o表现出更强的欺骗能力，但对他人谎言的脆弱性也更高，表明欺骗技能可能比检测能力扩展得更快。该框架可作为研究LLM在社会互动中欺骗与信任动态的测试平台。","reliability":"论文强调实验仅为概念验证，规模有限，受计算资源约束，未进行大规模统计效力研究；未讨论框架在其他场景下的泛化局限。","relevance":"该研究属于LLM多智能体社会仿真，但无真实人类数据对照，不符合研究者对基准人类数据的要求；若关注欺骗行为仿真本身，可了解其框架设计，但与经济学实验或政策评估的直接关联较弱。","inspiration":"该框架通过角色分配和信息不对称设计来诱发LLM代理的欺骗行为，并测量欺骗成功率与信任动态，这种多智能体博弈设计可借鉴用于研究经济决策中的策略性信息操纵｜可迁移至金融市场中的内幕交易与信息扩散研究，例如模拟交易员在拥有私有信息时的欺骗性沟通与市场信任演化｜可设计LLM代理扮演交易员，其中部分代理获得内幕信息（处理组），测量其欺骗性报价行为与市场价格的偏离，并以真实市场微观结构数据（如订单流与价格波动）作为对照基准"}},{"id":"2505.10309","version":3,"title":"A large-scale evaluation of commonsense knowledge in humans and large language models","zh_title":"人类与大语言模型常识知识的大规模评估","abstract":"Commonsense knowledge, a major constituent of artificial intelligence (AI), is primarily evaluated in practice by human-prescribed ground-truth labels. An important, albeit implicit, assumption of these labels is that they accurately capture what any human would think, effectively treating human common sense as homogeneous. However, recent empirical work has shown that humans vary enormously in what they consider commonsensical; thus what appears self-evident to one benchmark designer may not be so to another. Here, we propose a method for assessing commonsense knowledge in AI, specifically in large language models (LLMs), that incorporates empirically observed heterogeneity among humans by measuring the correspondence between a model's judgment and that of a human population. We first find that, when treated as independent survey respondents, most LLMs remain below the human median in their individual commonsense competence. Second, when used as simulators of a hypothetical population, LLMs correlate with real humans only modestly in the extent to which they agree on the same set of statements. In both cases, smaller, open-weight models are surprisingly more competitive than larger, proprietary frontier models. Our evaluation framework, which ties commonsense knowledge to its cultural basis, contributes to the growing call for adapting AI models to human collectivities that possess different, often incompatible, social stocks of knowledge.","authors":["Tuan Dung Nguyen","Duncan J. Watts","Mark E. Whiting"],"categories":["cs.AI","cs.HC","cs.SI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-05-15","first_seen":"2025-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.10309","pdf_url":"https://arxiv.org/pdf/2505.10309","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","常识知识"],"reason":"将LLM作为独立调查受访者，与真实人类常识判断分布对照，评估仿真可靠性与异质性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":36,"question":"如何将人类常识判断的异质性纳入评估，衡量大语言模型与人类群体在常识知识上的一致性？","design":"将35个LLM作为独立调查受访者，收集其对常识陈述的判断，并与2046名人类受访者的判断分布进行对比；同时用LLM生成硅样本群体，模拟人类群体对陈述的共识程度。","baseline":"2046名人类受访者对常识陈述的判断分布，包括个体间共识和群体共识。","findings":"多数LLM的常识能力低于人类中位数，且LLM模拟的群体共识与真实人类群体共识仅呈中等相关；小型开源模型的表现可与大型闭源模型竞争。","reliability":"论文指出常识知识具有文化依赖性，当前评估框架仅基于特定人类群体，可能不适用于其他社会文化背景；LLM模拟的群体在定性上与人类存在差异，如Gemini Pro 1.0过度关联常识与修辞性表达。","relevance":"该研究将LLM作为人类被试的替代品，与大规模真实人类数据对照，评估仿真可靠性与异质性，并指出仿真失效的文化条件，直接回应了研究者对LLM仿真实验的核心关切，值得精读。","inspiration":"该方法将LLM作为独立受访者，直接与大规模人类样本的判断分布进行对比，并生成硅样本模拟群体共识，可借鉴其个体-群体双层对照设计｜可迁移到经济预期形成研究，例如调查公众对通胀或政策公告的预期分布｜研究设计：以LLM作为被试，处理为不同措辞的央行声明，结果变量为通胀预期值，用密歇根大学消费者调查的真实预期分布数据作为对照基准"}},{"id":"2505.09938","version":2,"title":"Design and Evaluation of Generative Agent-based Platform for Human-Assistant Interaction Research: A Tale of 10 User Studies","zh_title":"基于生成式智能体的仿真平台设计与评估：10项用户研究的故事","abstract":"Designing and evaluating personalized and proactive assistant agents remains challenging due to the time, cost, and ethical concerns associated with human-in-the-loop experimentation. Existing Human-Computer Interaction (HCI) methods often require extensive physical setup and human participation, which introduces privacy concerns and limits scalability. Simulated environments offer a partial solution but are typically constrained by rule-based scenarios and still depend heavily on human input to guide interactions and interpret results. Recent advances in large language models (LLMs) have introduced the possibility of generative agents that can simulate realistic human behavior, reasoning, and social dynamics. However, their effectiveness in modeling human-assistant interactions remains largely unexplored. To address this gap, we present a generative agent-based simulation platform designed to simulate human-assistant interactions. We identify ten prior studies on assistant agents that span different aspects of interaction design and replicate these studies using our simulation platform. Our results show that fully simulated experiments using generative agents can approximate key aspects of human-assistant interactions. Based on these simulations, we are able to replicate the core conclusions of the original studies. Our work provides a scalable and cost-effective approach for studying assistant agent design without requiring live human subjects. Additional resources and project materials are available at https://dash-gidea.github.io/","authors":["Ziyi Xuan","Yiwen Wu","Xuhai Xu","Vinod Namboodiri","Mooi Choo Chuah","Yu Yang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-05-15","first_seen":"2025-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.09938","pdf_url":"https://arxiv.org/pdf/2505.09938","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A2","B1"],"tags":["LLM仿真","人机交互","用户研究复现"],"reason":"用LLM仿真人类与助手交互，复现10项用户研究并与真实人类数据对照，评估仿真可…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":71,"question":"基于LLM的生成式智能体能否有效模拟人类与助手交互，并复现真实用户研究的核心结论？","design":"使用LLM驱动的生成式智能体模拟人类被试，构建仿真平台GIDEA；选取10项涵盖主动协助、可中断性、自适应个性化等主题的真实用户研究，将其实验方案转化为结构化提示模板，在平台上复现实验，测量行为指标（如策略使用、接受率）和回复语义相似度。","baseline":"10项已发表用户研究中的真实人类被试数据，包括行为模式和定性描述。","findings":"生成式智能体仿真能够近似人类-助手交互的关键行为模式，并成功复现原始研究的核心结论。该平台提供了一种可扩展且低成本的研究方法，无需招募真实人类被试。","reliability":"论文未讨论","relevance":"该研究直接用LLM仿真人类被试复现10项用户研究，并与真实人类数据对照，属于你关注的人类仿真实验，且包含经济学实验和政策评估之外的HCI场景，值得阅读原文以评估其仿真保真度和偏差。","inspiration":"该方法将真实用户研究方案转化为结构化提示模板，在LLM仿真平台上复现实验并测量行为指标与语义相似度，为经济实验的仿真验证提供了可操作的流程｜可迁移至消费者跨期选择实验，检验LLM仿真能否复现真实被试的时间偏好异质性与框架效应｜以LLM智能体为被试，施加不同时间折现的决策框架处理，测量选择延迟与折现率，对照真实实验室实验数据验证仿真保真度"}},{"id":"2505.09396","version":2,"title":"The Influence of Human-inspired Agentic Sophistication in LLM-driven Strategic Reasoners","zh_title":"人类启发的智能体复杂度对LLM驱动战略推理者的影响","abstract":"The rapid rise of large language models (LLMs) has shifted artificial intelligence (AI) research toward agentic systems, motivating the use of weaker and more flexible notions of agency. However, this shift raises key questions about the extent to which LLM-based agents replicate human strategic reasoning, particularly in game-theoretic settings. In this context, we examine the role of agentic sophistication in shaping artificial reasoners' performance by evaluating three agent designs: a simple game-theoretic model, an unstructured LLM-as-agent model, and an LLM integrated into a traditional agentic framework. Using guessing games as a testbed, we benchmarked these agents against human participants across general reasoning patterns and individual role-based objectives. Furthermore, we introduced obfuscated game scenarios to assess agents' ability to generalise beyond training distributions. Our analysis, covering over 2000 reasoning samples across 25 agent configurations, shows that human-inspired cognitive structures can enhance LLM agents' alignment with human strategic behaviour. Still, the relationship between agentic design complexity and human-likeness is non-linear, highlighting a critical dependence on underlying LLM capabilities and suggesting limits to simple architectural augmentation.","authors":["Vince Trencsenyi","Agnieszka Mensfelt","Kostas Stathis"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2025-05-14","first_seen":"2025-05-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.09396","pdf_url":"https://arxiv.org/pdf/2505.09396","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","博弈实验","人类对照"],"reason":"用LLM代理模拟人类策略推理，并与真实人类数据对照，评估对齐程度与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":51,"question":"在博弈论猜测游戏中，LLM智能体的设计复杂度（agentic sophistication）如何影响其与人类策略推理行为的对齐程度？","design":"使用三种智能体设计（简单博弈论模型EWA、非结构化LLM-as-agent、LLM集成传统智能体框架），在双人猜测游戏中与人类被试进行基准对比，并引入混淆游戏场景测试分布外泛化能力，测量推理模式与个体角色目标的对齐度。","baseline":"人类数据集，按子群体划分，包含人类在猜测游戏中的策略推理行为。","findings":"受人类启发的认知结构能增强LLM智能体与人类策略行为的一致性，但智能体设计复杂度与人类相似度之间呈非线性关系，高度依赖底层LLM能力，简单的架构增强存在局限。","reliability":"论文指出LLM的黑箱特性带来可复现性、可解释性和验证挑战，混淆场景测试旨在缓解训练数据暴露偏差，但承认LLM驱动应用仍缺乏系统严格的验证机制。","relevance":"直接回应了用LLM仿真人类策略推理的核心问题，提供了与真实人类数据的对照基准，并批判性地揭示了设计复杂度与人类相似度的非线性关系及失效条件，值得精读。","inspiration":"借鉴其用不同复杂度智能体设计与人类基准对比的框架，以及引入混淆场景测试分布外泛化的稳健性检验方法｜可迁移到资产定价实验，研究投资者在策略性猜测市场走势时的推理行为｜以LLM智能体为被试，处理为不同复杂度的智能体设计（如简单启发式、EWA模型、非结构化LLM），结果变量为价格预测偏差，用真实人类资产定价实验数据做对照"}},{"id":"2505.07457","version":1,"title":"Can Generative AI agents behave like humans? Evidence from laboratory market experiments","zh_title":"生成式AI智能体能像人类一样行为吗？来自实验室市场实验的证据","abstract":"We explore the potential of Large Language Models (LLMs) to replicate human behavior in economic market experiments. Compared to previous studies, we focus on dynamic feedback between LLM agents: the decisions of each LLM impact the market price at the current step, and so affect the decisions of the other LLMs at the next step. We compare LLM behavior to market dynamics observed in laboratory settings and assess their alignment with human participants' behavior. Our findings indicate that LLMs do not adhere strictly to rational expectations, displaying instead bounded rationality, similarly to human participants. Providing a minimal context window i.e. memory of three previous time steps, combined with a high variability setting capturing response heterogeneity, allows LLMs to replicate broad trends seen in human experiments, such as the distinction between positive and negative feedback markets. However, differences remain at a granular level--LLMs exhibit less heterogeneity in behavior than humans. These results suggest that LLMs hold promise as tools for simulating realistic human behavior in economic contexts, though further research is needed to refine their accuracy and increase behavioral diversity.","authors":["R. Maria del Rio-Chanona","Marco Pangallo","Cars Hommes"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2025-05-12","first_seen":"2025-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.07457","pdf_url":"https://arxiv.org/pdf/2505.07457","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","市场实验","人类行为对照"],"reason":"用LLM复现市场实验并与人类数据对照，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":66,"question":"大语言模型能否在动态市场实验中复现人类行为，特别是正负反馈市场中的价格动态？","design":"使用GPT-3.5和GPT-4作为智能体，模拟实验室市场实验，操控上下文窗口（记忆长度）和响应变异性（温度参数），观察市场价格动态和个体策略。","baseline":"对照Heemeijer等人(2009)等实验室市场实验的人类参与者数据，比较正负反馈市场中的价格收敛模式和波动特征。","findings":"在至少3步记忆和高响应变异性下，LLM市场能复现正负反馈市场的宏观差异，如负反馈市场快速振荡收敛、正反馈市场缓慢收敛；但LLM行为异质性低于人类，且不严格遵循理性预期，表现出有限理性。","reliability":"LLM在细粒度行为上异质性不足，与人类仍有差距；研究仅基于特定市场实验范式，泛化性待验证；模型类型和参数设置对结果敏感，需进一步优化以提升行为多样性和准确性。","relevance":"直接命中研究者关注的核心：用LLM复现经济实验并与人类基准对照，评估仿真可靠性与偏差，且包含动态交互和批判性发现，值得精读原文。","inspiration":"借鉴之处在于通过操控LLM的上下文窗口长度和响应变异性来模拟有限理性，并与人类实验基准对照，检验宏观市场动态的复现能力｜可迁移到资产定价实验，研究正负反馈机制下价格泡沫的形成与破裂｜设计一个LLM模拟的资产市场实验，将被试分为GPT-4智能体，处理为不同反馈结构（正反馈如追涨杀跌 vs. 负反馈如均值回归），结果变量为价格偏离基本面的程度和泡沫持续时间，对照真实人类实验数据（如Smith等人1988年的泡沫实验）"}},{"id":"2505.06702","version":1,"title":"Do Language Model Agents Align with Humans in Rating Visualizations? An Empirical Study","zh_title":"语言模型代理在可视化评分中与人类对齐吗？一项实证研究","abstract":"Large language models encode knowledge in various domains and demonstrate the ability to understand visualizations. They may also capture visualization design knowledge and potentially help reduce the cost of formative studies. However, it remains a question whether large language models are capable of predicting human feedback on visualizations. To investigate this question, we conducted three studies to examine whether large model-based agents can simulate human ratings in visualization tasks. The first study, replicating a published study involving human subjects, shows agents are promising in conducting human-like reasoning and rating, and its result guides the subsequent experimental design. The second study repeated six human-subject studies reported in literature on subjective ratings, but replacing human participants with agents. Consulting with five human experts, this study demonstrates that the alignment of agent ratings with human ratings positively correlates with the confidence levels of the experts before the experiments. The third study tests commonly used techniques for enhancing agents, including preprocessing visual and textual inputs, and knowledge injection. The results reveal the issues of these techniques in robustness and potential induction of biases. The three studies indicate that language model-based agents can potentially simulate human ratings in visualization experiments, provided that they are guided by high-confidence hypotheses from expert evaluators. Additionally, we demonstrate the usage scenario of swiftly evaluating prototypes with agents. We discuss insights and future directions for evaluating and improving the alignment of agent ratings with human ratings. We note that simulation may only serve as complements and cannot replace user studies.","authors":["Zekai Shao","Yi Shan","Yixuan He","Yuxuan Yao","Junhong Wang","Xiaolong","Zhang","Yu Zhang","Siming Chen"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-05-10","first_seen":"2025-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.06702","pdf_url":"https://arxiv.org/pdf/2505.06702","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类数据对照","可视化评估"],"reason":"用LLM代理模拟人类对可视化的评分，并与真实人类数据对照，评估对齐度与偏差，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":67,"question":"大语言模型代理能否在可视化评分任务中模拟人类评分？","design":"使用GPT-4V等LLM代理，复现已发表的人类被试可视化实验，让代理对可视化设计进行主观评分（如易用性、信心水平），并测量代理评分与人类评分的一致性。","baseline":"对照的真实人类数据来自已发表的六项人类被试研究，包括时间序列可视化实验等，数据来源于Open Science Framework。","findings":"LLM代理能模拟人类推理并给出类似人类的评分，但无法模拟多样化用户画像；代理与人类评分的一致性程度与专家实验前信心水平正相关。","reliability":"代理评分仅能作为人类用户研究的补充，不能替代；增强技术（如知识注入）可能引入偏差，且代理在鲁棒性上存在问题。","relevance":"该研究直接以真实人类数据为基准，评估LLM代理在可视化评分任务中的仿真可靠性，并讨论了失效条件与偏差，符合研究者对经济学实验和政策评估场景的批判性关注。","inspiration":"该方法借鉴了用LLM代理复现已发表人类实验并直接对比真实人类评分一致性的设计，以及通过专家信心水平等指标检验仿真可靠性的思路。｜可迁移到消费者对金融产品信息披露的主观评价实验，如研究简化版风险提示是否提升理解度。｜以LLM代理作为被试，施加不同格式的风险披露文本处理，测量代理对产品风险的理解评分，并与真实消费者调查数据（如CFPB的金融素养调查）进行一致性对比。"}},{"id":"2507.18639","version":1,"title":"People Are Highly Cooperative with Large Language Models, Especially When Communication Is Possible or Following Human Interaction","zh_title":"人们与大型语言模型高度合作，尤其在可沟通或继人类互动之后","abstract":"Machines driven by large language models (LLMs) have the potential to augment humans across various tasks, a development with profound implications for business settings where effective communication, collaboration, and stakeholder trust are paramount. To explore how interacting with an LLM instead of a human might shift cooperative behavior in such settings, we used the Prisoner's Dilemma game -- a surrogate of several real-world managerial and economic scenarios. In Experiment 1 (N=100), participants engaged in a thirty-round repeated game against a human, a classic bot, and an LLM (GPT, in real-time). In Experiment 2 (N=192), participants played a one-shot game against a human or an LLM, with half of them allowed to communicate with their opponent, enabling LLMs to leverage a key advantage over older-generation machines. Cooperation rates with LLMs -- while lower by approximately 10-15 percentage points compared to interactions with human opponents -- were nonetheless high. This finding was particularly notable in Experiment 2, where the psychological cost of selfish behavior was reduced. Although allowing communication about cooperation did not close the human-machine behavioral gap, it increased the likelihood of cooperation with both humans and LLMs equally (by 88%), which is particularly surprising for LLMs given their non-human nature and the assumption that people might be less receptive to cooperating with machines compared to human counterparts. Additionally, cooperation with LLMs was higher following prior interaction with humans, suggesting a spillover effect in cooperative behavior. Our findings validate the (careful) use of LLMs by businesses in settings that have a cooperative component.","authors":["Paweł Niszczota","Tomasz Grzegorczyk","Alexander Pastukhov"],"categories":["cs.HC","cs.CL","cs.CY","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2025-05-10","first_seen":"2025-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.18639","pdf_url":"https://arxiv.org/pdf/2507.18639","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","行为博弈","人机合作"],"reason":"用LLM替代人类被试进行囚徒困境博弈，并与真实人类行为对照，评估合作行为差异与…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":27,"question":"当对手是大语言模型而非人类时，人们在囚徒困境中的合作行为会发生怎样的变化？","design":"实验1（N=100）采用被试内设计，每人与人类、经典机器人、GPT实时对战各30轮重复囚徒困境；实验2（N=192）采用被试间设计，在单次囚徒困境中对手为人类或LLM，且半数被试可在决策前与对手沟通。结果变量为合作率。","baseline":"以真实人类作为对手时的合作行为作为对照基准。","findings":"与LLM的合作率虽比与人类对手低约10–15个百分点，但仍处于较高水平；允许沟通使与人类和LLM的合作率均提高88%，且与人类互动后与LLM的合作率更高，存在溢出效应。","reliability":"论文未讨论","relevance":"该研究直接以LLM替代人类被试进行囚徒困境博弈，并与真实人类行为对照，评估合作行为差异及沟通、溢出效应，高度契合研究者对LLM仿真人类行为可靠性与偏差的关注，值得精读原文。","inspiration":"采用被试内设计直接对比同一参与者在面对人类、传统机器人和LLM时的行为变化，有效控制个体差异，值得借鉴。｜可迁移至经济金融中的信任与履约行为研究，如线上借贷平台中借款人对人工审核员与AI审核员的还款承诺差异。｜以真实借款人为被试，随机分配其与人类审核员或LLM审核员沟通还款计划，测量其后续实际还款率，并以平台历史人工审核还款数据作为对照基准。"}},{"id":"2505.00036","version":1,"title":"A Framework to Assess the Persuasion Risks Large Language Model Chatbots Pose to Democratic Societies","zh_title":"评估大语言模型聊天机器人对民主社会说服风险的框架","abstract":"In recent years, significant concern has emerged regarding the potential threat that Large Language Models (LLMs) pose to democratic societies through their persuasive capabilities. We expand upon existing research by conducting two survey experiments and a real-world simulation exercise to determine whether it is more cost effective to persuade a large number of voters using LLM chatbots compared to standard political campaign practice, taking into account both the \"receive\" and \"accept\" steps in the persuasion process (Zaller 1992). These experiments improve upon previous work by assessing extended interactions between humans and LLMs (instead of using single-shot interactions) and by assessing both short- and long-run persuasive effects (rather than simply asking users to rate the persuasiveness of LLM-produced content). In two survey experiments (N = 10,417) across three distinct political domains, we find that while LLMs are about as persuasive as actual campaign ads once voters are exposed to them, political persuasion in the real-world depends on both exposure to a persuasive message and its impact conditional on exposure. Through simulations based on real-world parameters, we estimate that LLM-based persuasion costs between \\$48-\\$74 per persuaded voter compared to \\$100 for traditional campaign methods, when accounting for the costs of exposure. However, it is currently much easier to scale traditional campaign persuasion methods than LLM-based persuasion. While LLMs do not currently appear to have substantially greater potential for large-scale political persuasion than existing non-LLM methods, this may change as LLM capabilities continue to improve and it becomes easier to scalably encourage exposure to persuasive LLMs.","authors":["Zhongren Chen","Joshua Kalla","Quan Le","Shinpei Nakamura-Sakai","Jasjeet Sekhon","Ruixiao Wang"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-04-29","first_seen":"2025-04-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.00036","pdf_url":"https://arxiv.org/pdf/2505.00036","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","政治说服","人类数据对照"],"reason":"用LLM聊天机器人替代人类选民进行说服实验，有真实人类调查数据对照，涉及政治说…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":22,"question":"在政治说服中，使用LLM聊天机器人与传统竞选广告相比，在考虑曝光和接受两个步骤后，是否更具成本效益？","design":"本研究并非用LLM仿真人类被试，而是通过两项调查实验和真实世界模拟，将人类参与者随机分配到安慰剂、真人视频说服、AI聊天机器人（自称人类）和AI聊天机器人（自称AI）四种条件，测量其对移民政策的支持态度变化，并基于真实参数模拟成本效益。","baseline":"真人视频说服条件（3分钟教师分享个人理由的视频）和传统竞选广告方法（如电视广告）作为对照基准。","findings":"在调查实验中，LLM聊天机器人与真人视频的说服效果在短期和五周后均无显著差异；但模拟显示，考虑曝光成本后，LLM说服每位选民的成本为48-74美元，低于传统方法的100美元，然而目前传统方法更易规模化。","reliability":"论文指出当前LLM在规模化政治说服方面并不比现有非LLM方法有显著优势，但随LLM能力提升和更容易鼓励曝光，情况可能改变；研究局限包括真人说服条件为3分钟视频，长于典型广告，且未充分解决曝光环节的规模化难题。","relevance":"该研究直接使用LLM与人类进行交互式说服实验，并与真实人类说服效果对照，涉及政治领域成本效益评估，同时讨论了规模化限制，高度契合研究者对LLM仿真人类行为、基准对照和失效条件的兴趣，值得精读原文。","inspiration":"该方法将LLM聊天机器人作为交互式说服工具，与真人视频说服和传统广告进行随机对照实验，并追踪短期与五周后的态度变化，值得借鉴其多臂对照和纵向测量设计｜可迁移到消费者金融决策场景，如评估LLM理财顾问对投资偏好或退休储蓄选择的影响｜以真实投资者为被试，随机分配至LLM聊天机器人理财建议、人类理财顾问视频和纯文本说明书三组，结果变量为风险资产配置比例和储蓄率变化，以历史调查数据或银行实际客户行为作为基准对照"}},{"id":"2504.20628","version":1,"title":"Cognitive maps are generative programs","zh_title":"认知地图是生成式程序","abstract":"Making sense of the world and acting in it relies on building simplified mental representations that abstract away aspects of reality. This principle of cognitive mapping is universal to agents with limited resources. Living organisms, people, and algorithms all face the problem of forming functional representations of their world under various computing constraints. In this work, we explore the hypothesis that human resource-efficient planning may arise from representing the world as predictably structured. Building on the metaphor of concepts as programs, we propose that cognitive maps can take the form of generative programs that exploit predictability and redundancy, in contrast to directly encoding spatial layouts. We use a behavioral experiment to show that people who navigate in structured spaces rely on modular planning strategies that align with programmatic map representations. We describe a computational model that predicts human behavior in a variety of structured scenarios. This model infers a small distribution over possible programmatic cognitive maps conditioned on human prior knowledge of the world, and uses this distribution to generate resource-efficient plans. Our models leverages a Large Language Model as an embedding of human priors, implicitly learned through training on a vast corpus of human data. Our model demonstrates improved computational efficiency, requires drastically less memory, and outperforms unstructured planning algorithms with cognitive constraints at predicting human behavior, suggesting that human planning strategies rely on programmatic cognitive maps.","authors":["Marta Kryven","Cole Wyeth","Aidan Curtis","Kevin Ellis"],"categories":["cs.AI","cs.ET"],"primary_category":"cs.AI","announce_type":"new","date":"2025-04-29","first_seen":"2025-04-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.20628","pdf_url":"https://arxiv.org/pdf/2504.20628","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["认知建模","人类行为预测","LLM先验嵌入"],"reason":"用LLM嵌入人类先验模拟导航行为，有行为实验对照，但LLM非被试替代，属社会模…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":116,"question":"人类在结构化空间中的资源高效规划是否依赖于将世界表示为可预测的生成式程序（即程序化认知地图）？","design":"本研究并非用LLM替代人类被试的仿真实验，而是通过行为实验（迷宫搜索任务）收集30名人类被试在部分可观测结构化迷宫中的导航行为，并构建计算模型（利用LLM嵌入人类先验知识）来预测人类规划行为。","baseline":"对照的真实人类数据为30名Prolific平台招募的英语母语被试在21个生成式结构化迷宫中的点击导航行为，以及认知反射测试（CRT）和事后问卷。","findings":"人类在结构化迷宫中采用模块化规划策略，其行为与程序化认知地图模型预测一致；该模型比无结构规划算法（含认知约束）更准确预测人类行为，且计算效率更高、内存需求更少。","reliability":"论文未讨论","relevance":"该研究利用LLM嵌入人类先验知识来预测行为，并有真实人类实验对照，但核心是认知地图的计算模型，并非将LLM作为人类被试的替代品进行仿真实验，与研究者关注的LLM仿真替代方向有偏差，但可作为LLM用于行为预测的参考。","inspiration":"该研究利用LLM嵌入人类先验知识构建程序化认知地图模型，通过行为实验收集人类导航数据作为基准，验证模型预测能力｜可迁移至经济金融中的消费者搜索与决策行为研究，如在线购物平台上的产品筛选与购买决策｜设计：招募人类被试在模拟电商平台完成产品搜索任务，以LLM嵌入消费者先验偏好构建预测模型，结果变量为点击路径和最终购买选择，以真实人类行为数据作为对照基准"}},{"id":"2504.17993","version":2,"title":"Improving Language Model Personas via Rationalization with Psychological Scaffolds","zh_title":"通过心理支架合理化改进语言模型角色","abstract":"Language models prompted with a user description or persona are being used to predict the user's preferences and opinions. However, existing approaches to building personas mostly rely on a user's demographic attributes and/or prior judgments, but not on any underlying reasoning behind a user's judgments. We introduce PB&J (Psychology of Behavior and Judgments), a framework that improves LM personas by incorporating potential rationales for why the user could have made a certain judgment. Our rationales are generated by a language model to explicitly reason about a user's behavior on the basis of their experiences, personality traits, or beliefs. Our method employs psychological scaffolds: structured frameworks such as the Big 5 Personality Traits or Primal World Beliefs to help ground the generated rationales in existing theories. Experiments on public opinion and movie preference prediction tasks demonstrate that language model personas augmented with PB&J rationales consistently outperform personas conditioned only on user demographics and / or judgments, including those that use a model's default chain-of-thought, which is not grounded in psychological theories. Additionally, our PB&J personas perform competitively with those using human-written rationales, suggesting the potential of synthetic rationales guided by existing theories.","authors":["Brihi Joshi","Xiang Ren","Swabha Swayamdipta","Rik Koncel-Kedziorski","Tim Paek"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-04-25","first_seen":"2025-04-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.17993","pdf_url":"https://arxiv.org/pdf/2504.17993","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM角色仿真","人类偏好预测","心理支架"],"reason":"用LLM模拟用户偏好预测，有人类数据对照，但侧重persona构建而非群体仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":87,"question":"如何利用心理学理论生成用户判断的合理化解释，以提升语言模型模拟用户偏好和意见的准确性？","design":"本研究并非群体仿真实验，而是提出PB&J框架：给定用户人口统计属性和先前的判断，利用大语言模型基于心理学脚手架（如大五人格、原始世界信念）生成用户判断的潜在理由，构建更丰富的用户画像提示，然后在OpinionQA和MovieLens数据集上预测用户对测试问题的回答或电影评分，评估预测准确率。","baseline":"使用OpinionQA（750名用户，10个测试问题）和MovieLens（100名用户，10部测试电影）中的真实用户回答和评分作为对照基准。","findings":"加入PB&J生成理由的用户画像在预测用户意见和偏好上显著优于仅使用人口统计和/或先前判断的画像，也优于无心理学理论支撑的链式思维推理。PB&J生成的合成理由与人类撰写的理由性能接近，表明基于现有理论生成的理由具有潜力。","reliability":"论文承认生成的合成理由可能并不反映用户判断的真实推理过程，但因其可能包含真实行为推理的标记而具有合理性；此外，实验仅在两个数据集上进行，且人类撰写理由的对比实验规模较小。","relevance":"该研究直接涉及用LLM模拟个体用户偏好预测，有真实人类数据对照，并探讨了仿真方法的改进与局限，与研究者关注的LLM人类仿真实验高度相关，值得阅读原文以了解心理学理论如何提升仿真可靠性。","inspiration":"该方法通过心理学理论引导LLM生成用户判断的合理化理由，以丰富用户画像，提升偏好预测准确性，这种利用理论驱动的提示工程增强仿真保真度的做法值得借鉴｜可迁移到消费者金融决策仿真，如利用大五人格或风险态度理论生成理由，预测个体对金融产品的选择或风险偏好｜以真实消费者金融调查数据为基准，用LLM基于人口统计和先前金融行为生成带心理学理由的画像，预测测试集上的产品选择，对比仅用人口统计的基线，评估仿真准确性"}},{"id":"2504.19940","version":2,"title":"Assessing the Potential of Generative Agents in Crowdsourced Fact-Checking","zh_title":"评估生成式智能体在众包事实核查中的潜力","abstract":"The growing spread of online misinformation has created an urgent need for scalable, reliable fact-checking solutions. Crowdsourced fact-checking - where non-experts evaluate claim veracity - offers a cost-effective alternative to expert verification, despite concerns about variability in quality and bias. Encouraged by promising results in certain contexts, major platforms such as X (formerly Twitter), Facebook, and Instagram have begun shifting from centralized moderation to decentralized, crowd-based approaches. In parallel, advances in Large Language Models (LLMs) have shown strong performance across core fact-checking tasks, including claim detection and evidence evaluation. However, their potential role in crowdsourced workflows remains unexplored. This paper investigates whether LLM-powered generative agents - autonomous entities that emulate human behavior and decision-making - can meaningfully contribute to fact-checking tasks traditionally reserved for human crowds. Using the protocol of La Barbera et al. (2024), we simulate crowds of generative agents with diverse demographic and ideological profiles. Agents retrieve evidence, assess claims along multiple quality dimensions, and issue final veracity judgments. Our results show that agent crowds outperform human crowds in truthfulness classification, exhibit higher internal consistency, and show reduced susceptibility to social and cognitive biases. Compared to humans, agents rely more systematically on informative criteria such as Accuracy, Precision, and Informativeness, suggesting a more structured decision-making process. Overall, our findings highlight the potential of generative agents as scalable, consistent, and less biased contributors to crowd-based fact-checking systems.","authors":["Luigia Costabile","Gian Marco Orlando","Valerio La Gatta","Vincenzo Moscato"],"categories":["cs.CL","cs.AI","cs.MA"],"primary_category":"cs.CL","announce_type":"new","date":"2025-04-24","first_seen":"2025-04-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.19940","pdf_url":"https://arxiv.org/pdf/2504.19940","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","众包事实核查","人类行为对照"],"reason":"用LLM智能体模拟人群事实核查，与真实人类数据对照，评估偏差与一致性，直接命中…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":31,"question":"生成式智能体在众包事实核查中的表现能否达到或超过人类众包？","design":"使用LLM驱动的生成式智能体模拟具有不同人口统计和意识形态特征的人群，复现La Barbera et al. (2024)的实验协议：智能体先选择证据，再填写结构化问卷对声明进行多维度评分（准确性、无偏性等），最后给出真实性判断。","baseline":"La Barbera et al. (2024)中人类众包参与者对相同声明的标注数据。","findings":"智能体众包在真实性分类上优于人类众包，表现出更高的内部一致性，且受社会和认知偏差的影响更小。智能体更系统地依赖准确性、精确性和信息量等标准，决策过程更结构化。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体模拟人类众包事实核查，并与真实人类数据对照，评估偏差与一致性，高度契合研究者对LLM人类仿真实验的关注，值得精读原文。","inspiration":"该方法借鉴了用LLM智能体复现人类实验协议并直接对比真实人类数据的做法，通过让智能体遵循相同的任务流程（选择证据、填写问卷、给出判断）来评估仿真偏差与一致性。｜可迁移到经济金融中的信贷审批歧视研究，模拟不同人口特征的贷款审批决策。｜设计：用LLM智能体模拟不同种族、性别、收入的贷款申请人，处理为呈现相同的贷款申请材料，结果变量为审批结果和理由，对照真实银行信贷审批数据或人类实验数据。"}},{"id":"2504.11671","version":4,"title":"Computational Basis of LLM's Decision Making in Social Simulation","zh_title":"LLM在社会仿真中决策的计算基础","abstract":"Large language models (LLMs) increasingly serve as human-like decision-making agents in social science and applied settings. These LLM-agents are typically assigned human-like characters and placed in real-life contexts. However, how these characters and contexts shape an LLM's behavior remains underexplored. This study proposes and tests methods for probing, quantifying, and modifying an LLM's internal representations in a Dictator Game, a classic behavioral experiment on fairness and prosocial behavior. We extract ``vectors of variable variations'' (e.g., ``male'' to ``female'') from the LLM's internal state. Manipulating these vectors during the model's inference can substantially alter how those variables relate to the model's decision-making. This approach offers a principled way to study and regulate how social concepts can be encoded and engineered within transformer-based models, with implications for alignment, debiasing, and designing AI agents for social simulations in both academic and commercial applications, strengthening sociological theory and measurement.","authors":["Ji Ma"],"categories":["cs.AI","cs.CY","cs.LG","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2025-04-16","first_seen":"2025-04-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.11671","pdf_url":"https://arxiv.org/pdf/2504.11671","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A3","B2"],"tags":["LLM社会仿真","独裁者博弈","表征操控"],"reason":"用LLM在独裁者博弈中模拟人类决策，并操控内部表征改变行为，属于社会仿真且涉及…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":117,"question":"LLM在社会模拟中，其内部表征如何编码社会概念（如性别、框架），以及如何通过操控这些内部表征来改变其决策行为？","design":"本研究并非直接进行人类仿真，而是提出一种探查和操控LLM内部表征的方法。在独裁者博弈实验中，从LLM的残差流中提取社会变量（如性别）的“变化向量”，通过正交化和注入扰动来操控这些向量，观察LLM分配决策的变化。","baseline":"无对照","findings":"LLM内部存在可提取的、与社会概念对应的方向向量；通过操控这些向量可以显著改变模型在独裁者博弈中的决策行为，为研究LLM如何编码社会意义提供了透明、可干预的方法。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM在社会模拟中的内部决策机制，属于批判性方法研究，虽未提供人类基准对照，但其提出的表征操控方法对理解仿真可靠性和偏差具有重要价值，值得精读原文。","inspiration":"该方法通过提取LLM内部表征中的社会概念向量并进行正交化或注入扰动来操控决策行为，为因果识别提供了透明、可干预的框架｜可迁移到信贷审批中的性别歧视研究，探查并干预LLM在贷款决策中对性别信息的编码｜以LLM作为信贷审批官，提取性别变化向量并施加扰动，观察贷款批准率的变化，并与真实银行信贷数据中的性别差异进行对照"}},{"id":"2504.08260","version":2,"title":"Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare","zh_title":"评估大语言模型在医疗意见与决策调查中的偏差","abstract":"Generative agents have been increasingly used to simulate human behaviour in silico, driven by large language models (LLMs). These simulacra serve as sandboxes for studying human behaviour without compromising privacy or safety. However, it remains unclear whether such agents can truly represent real individuals. This work compares survey data from the Understanding America Study (UAS) on healthcare decision-making with simulated responses from generative agents. Using demographic-based prompt engineering, we create digital twins of survey respondents and analyse how well different LLMs reproduce real-world behaviours. Our findings show that some LLMs fail to reflect realistic decision-making, such as predicting universal vaccine acceptance. However, Llama 3 captures variations across race and Income more accurately but also introduces biases not present in the UAS data. This study highlights the potential of generative agents for behavioural research while underscoring the risks of bias from both LLMs and prompting strategies.","authors":["Yonchanok Khaokaew","Flora D. Salim","Andreas Züfle","Hao Xue","Taylor Anderson","C. Raina MacIntyre","Matthew Scotch","David J Heslop"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-04-11","first_seen":"2025-04-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.08260","pdf_url":"https://arxiv.org/pdf/2504.08260","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类数据对照","医疗决策偏差"],"reason":"用LLM仿真医疗决策，与真实调查数据对照，评估偏差，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"LLM能否有效模拟医疗决策（如疫苗接种意愿），以及在不同人口群体中会产生哪些偏差？","design":"使用LLM（如Llama 3等）基于人口统计属性（年龄、性别、收入、种族、教育、担忧程度）构建数字孪生，在不同疫情场景下提示模型回答疫苗接种意愿问题，比较模型输出与真实调查数据。","baseline":"理解美国研究（UAS）中关于COVID-19疫苗接种意愿的调查数据，涵盖2020年3月至2021年1月的多波次问卷。","findings":"部分LLM未能反映现实决策，例如预测普遍接受疫苗；Llama 3更准确地捕捉了种族和收入差异，但也引入了UAS数据中不存在的新偏差。","reliability":"论文指出LLM的预测效果取决于提供的上下文信息类型和数量，预训练数据和提示策略可能导致与人类偏差不同的偏差，且LLM可能仅反映训练数据中的统计模式而非真实决策过程。","relevance":"该研究直接使用LLM仿真医疗决策并与真实人类调查数据对照，评估偏差，完全符合研究者对LLM人类仿真实验、经济学/政策场景及可靠性批判的关注，值得精读原文。","inspiration":"借鉴其利用人口统计属性构建数字孪生并与多波次真实调查数据对照的仿真设计，评估LLM在决策模拟中的偏差｜可迁移到医疗健康政策的经济评估场景，如不同人群对医疗保险选择或健康储蓄账户的偏好差异｜以LLM为被试，基于收入、年龄、健康风险等属性生成数字孪生，提示其在不同保费和补贴政策下选择保险方案，结果变量为保险选择概率，用真实调查数据（如MEPS）作为基准对照"}},{"id":"2504.05862","version":2,"title":"Are Generative AI Agents Effective Personalized Financial Advisors?","zh_title":"生成式AI代理能成为有效的个性化理财顾问吗？","abstract":"Large language model-based agents are becoming increasingly popular as a low-cost mechanism to provide personalized, conversational advice, and have demonstrated impressive capabilities in relatively simple scenarios, such as movie recommendations. But how do these agents perform in complex high-stakes domains, where domain expertise is essential and mistakes carry substantial risk? This paper investigates the effectiveness of LLM-advisors in the finance domain, focusing on three distinct challenges: (1) eliciting user preferences when users themselves may be unsure of their needs, (2) providing personalized guidance for diverse investment preferences, and (3) leveraging advisor personality to build relationships and foster trust. Via a lab-based user study with 64 participants, we show that LLM-advisors often match human advisor performance when eliciting preferences, although they can struggle to resolve conflicting user needs. When providing personalized advice, the LLM was able to positively influence user behavior, but demonstrated clear failure modes. Our results show that accurate preference elicitation is key, otherwise, the LLM-advisor has little impact, or can even direct the investor toward unsuitable assets. More worryingly, users appear insensitive to the quality of advice being given, or worse these can have an inverse relationship. Indeed, users reported a preference for and increased satisfaction as well as emotional trust with LLMs adopting an extroverted persona, even though those agents provided worse advice.","authors":["Takehiro Takayanagi","Kiyoshi Izumi","Javier Sanz-Cruzado","Richard McCreadie","Iadh Ounis"],"categories":["cs.AI","cs.CL","cs.HC","cs.IR","q-fin.CP"],"primary_category":"cs.AI","announce_type":"new","date":"2025-04-08","first_seen":"2025-04-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.05862","pdf_url":"https://arxiv.org/pdf/2504.05862","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人机对照实验","金融行为"],"reason":"用LLM模拟人类理财顾问，与真人对照，评估效果与失效模式，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:44","error":null,"has_summary":true,"summary":{"generated_at":"2025-04-08","rank":5,"question":"LLM作为个性化理财顾问，在偏好获取、个性化建议和人格影响方面表现如何？","design":"实验室用户研究，64名参与者扮演给定投资者画像，与LLM顾问进行两阶段对话：偏好获取和资产建议讨论，比较个性化vs非个性化顾问及不同人格的顾问，测量决策质量、用户满意度和信任。","baseline":"人类顾问在偏好获取阶段的表现作为对照基准。","findings":"LLM顾问在偏好获取上常能与人类顾问匹敌，但难以解决用户需求冲突；个性化建议能正向影响用户行为，但存在明显失效模式，且用户对建议质量不敏感，甚至偏好外向人格的顾问，尽管其建议更差。","reliability":"论文指出准确偏好获取是关键，否则LLM顾问影响甚微或引导投资者选择不合适资产；用户对建议质量不敏感，甚至出现反向关系，外向人格虽提升满意度和情感信任但建议质量更低。","relevance":"该研究直接以真人实验对照LLM在金融建议中的表现，揭示了偏好获取失效和用户对建议质量不敏感等关键偏差，对关注经济学实验和决策仿真的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其分阶段对话设计和人格操纵处理，可迁移到信贷审批或保险推荐等金融场景，设计以LLM为被试、操纵建议人格或个性化程度、测量用户决策偏差和信任，并以人类顾问或历史决策数据为对照。"}},{"id":"2503.22726","version":1,"title":"InfoBid: A Simulation Framework for Studying Information Disclosure in Auctions with Large Language Model-based Agents","zh_title":"InfoBid：基于大语言模型代理研究拍卖中信息披露的仿真框架","abstract":"In online advertising systems, publishers often face a trade-off in information disclosure strategies: while disclosing more information can enhance efficiency by enabling optimal allocation of ad impressions, it may lose revenue potential by decreasing uncertainty among competing advertisers. Similar to other challenges in market design, understanding this trade-off is constrained by limited access to real-world data, leading researchers and practitioners to turn to simulation frameworks. The recent emergence of large language models (LLMs) offers a novel approach to simulations, providing human-like reasoning and adaptability without necessarily relying on explicit assumptions about agent behavior modeling. Despite their potential, existing frameworks have yet to integrate LLM-based agents for studying information asymmetry and signaling strategies, particularly in the context of auctions. To address this gap, we introduce InfoBid, a flexible simulation framework that leverages LLM agents to examine the effects of information disclosure strategies in multi-agent auction settings. Using GPT-4o, we implemented simulations of second-price auctions with diverse information schemas. The results reveal key insights into how signaling influences strategic behavior and auction outcomes, which align with both economic and social learning theories. Through InfoBid, we hope to foster the use of LLMs as proxies for human economic and social agents in empirical studies, enhancing our understanding of their capabilities and limitations. This work bridges the gap between theoretical market designs and practical applications, advancing research in market simulations, information design, and agent-based reasoning while offering a valuable tool for exploring the dynamics of digital economies.","authors":["Yue Yin"],"categories":["cs.GT","cs.CL","cs.HC","cs.MA","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2025-03-26","first_seen":"2025-03-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.22726","pdf_url":"https://arxiv.org/pdf/2503.22726","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","拍卖实验","经济行为"],"reason":"用LLM代理模拟拍卖中的人类经济行为，与理论对照，涉及经济学实验场景并讨论局限…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":64,"question":"在拍卖中，信息披露策略如何影响基于LLM的智能体的出价行为与拍卖结果？","design":"使用GPT-4o构建LLM智能体模拟竞拍者，在第二价格拍卖中施加不同的信息披露方案（信号），测量出价、收入、效率等拍卖结果。","baseline":"无对照","findings":"信号显著影响LLM智能体的策略行为与拍卖结果，且观察到的行为模式与经济理论和社会学习理论一致。","reliability":"论文未讨论","relevance":"高度相关：用LLM代理模拟拍卖中的人类经济行为，研究信息不对称下的策略互动，并讨论LLM作为人类代理的潜力与局限，直接命中研究者关心的经济学实验仿真与可靠性评估。","inspiration":"该方法通过向LLM智能体提供不同信号来模拟信息披露，可借鉴其处理信息干预的方式，用于研究经济决策中的信息效应｜可迁移到资产定价实验中，研究公开与私有信号如何影响投资者的价格预期与交易行为｜设计：以LLM作为被试，随机分配接收不同精度的资产价值信号，测量其报价与交易量，并与历史实验市场数据或理性预期均衡基准对照"}},{"id":"2503.11531","version":1,"title":"Potential of large language model-powered nudges for promoting daily water and energy conservation","zh_title":"大语言模型驱动的助推在促进日常节水节能中的潜力","abstract":"The increasing amount of pressure related to water and energy shortages has increased the urgency of cultivating individual conservation behaviors. While the concept of nudging, i.e., providing usage-based feedback, has shown promise in encouraging conservation behaviors, its efficacy is often constrained by the lack of targeted and actionable content. This study investigates the impact of the use of large language models (LLMs) to provide tailored conservation suggestions for conservation intentions and their rationale. Through a survey experiment with 1,515 university participants, we compare three virtual nudging scenarios: no nudging, traditional nudging with usage statistics, and LLM-powered nudging with usage statistics and personalized conservation suggestions. The results of statistical analyses and causal forest modeling reveal that nudging led to an increase in conservation intentions among 86.9%-98.0% of the participants. LLM-powered nudging achieved a maximum increase of 18.0% in conservation intentions, surpassing traditional nudging by 88.6%. Furthermore, structural equation modeling results reveal that exposure to LLM-powered nudges enhances self-efficacy and outcome expectations while diminishing dependence on social norms, thereby increasing intrinsic motivation to conserve. These findings highlight the transformative potential of LLMs in promoting individual water and energy conservation, representing a new frontier in the design of sustainable behavioral interventions and resource management.","authors":["Zonghan Li","Song Tong","Yi Liu","Kaiping Peng","Chunyan Wang"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-03-14","first_seen":"2025-03-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.11531","pdf_url":"https://arxiv.org/pdf/2503.11531","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","行为干预","节能实验"],"reason":"用LLM生成个性化节能建议，通过调查实验与人类对照，评估对行为意图的影响，属于…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:36","error":null,"has_summary":true,"summary":{"generated_at":"2025-03-14","rank":2,"question":"基于LLM的个性化助推与传统使用统计助推相比，能否更有效地提升个体的节水节能意图？","design":"本研究并非用LLM模拟人类被试，而是通过随机对照调查实验，将1515名大学生随机分为三组，分别接受无助推、传统使用统计助推、LLM生成个性化建议的助推，测量其节水节能意图的变化。","baseline":"无助推组和传统使用统计助推组作为对照，均为真实人类被试。","findings":"LLM助推使86.9%-98.0%的参与者节水节能意图提升，最大增幅达18.0%，效果比传统助推高88.6%。结构方程模型显示，LLM助推通过增强自我效能和结果预期、降低社会规范依赖，提升内在动机。","reliability":"论文未讨论","relevance":"该研究用LLM生成个性化干预内容，通过真实人类调查实验评估行为意图变化，属于LLM辅助行为干预效果评估，与研究者关注的LLM仿真实验和经济学实验场景高度相关，值得细读其因果推断设计和心理机制分析。","inspiration":"借鉴其随机分组和因果森林方法评估异质性处理效应，可迁移到消费者节能行为干预政策评估场景。｜设计一个实验：以居民为被试，随机分配接收传统节能建议或LLM个性化建议，结果变量为实际用电量变化，对照真实智能电表数据。"}},{"id":"2503.10248","version":1,"title":"LLM Agents Display Human Biases but Exhibit Distinct Learning Patterns","zh_title":"LLM智能体表现出人类偏见但学习模式不同","abstract":"We investigate the choice patterns of Large Language Models (LLMs) in the context of Decisions from Experience tasks that involve repeated choice and learning from feedback, and compare their behavior to human participants. We find that on the aggregate, LLMs appear to display behavioral biases similar to humans: both exhibit underweighting rare events and correlation effects. However, more nuanced analyses of the choice patterns reveal that this happens for very different reasons. LLMs exhibit strong recency biases, unlike humans, who appear to respond in more sophisticated ways. While these different processes may lead to similar behavior on average, choice patterns contingent on recent events differ vastly between the two groups. Specifically, phenomena such as ``surprise triggers change\" and the ``wavy recency effect of rare events\" are robustly observed in humans, but entirely absent in LLMs. Our findings provide insights into the limitations of using LLMs to simulate and predict humans in learning environments and highlight the need for refined analyses of their behavior when investigating whether they replicate human decision making tendencies.","authors":["Idan Horowitz","Ori Plonsky"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-03-13","first_seen":"2025-03-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.10248","pdf_url":"https://arxiv.org/pdf/2503.10248","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","决策实验","人类对照"],"reason":"用LLM复现人类决策实验，与真实人类数据对照，并指出仿真失效条件，高度相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":68,"question":"LLM在基于经验的决策任务中是否表现出与人类相似的行为偏差和学习模式？","design":"将多个LLM作为被试，完成四项重复二元选择决策任务（100试次），通过反馈学习收益分布，测量选择行为；同时操纵历史提供方式和模型温度参数。","baseline":"真实人类被试在相同决策任务中的选择数据，作为行为对比基准。","findings":"总体上看，LLM表现出与人类相似的低估稀有事件和相关效应，但过程不同：LLM有强烈的近因偏差，而人类对稀有事件有“惊讶触发改变”和“波浪式近因效应”，这些模式在LLM中完全缺失。","reliability":"LLM在聚合层面可能模拟人类偏差，但内在学习过程截然不同，基于近期事件的精细分析显示其无法复现人类特有的动态反应模式，提示在涉及反馈学习的场景中用LLM仿真人类决策存在局限。","relevance":"该研究直接对比LLM与人类在经济学实验范式下的决策，揭示仿真在表面聚合指标下有效但内在机制失效，高度契合研究者对LLM仿真可靠性及失效条件的批判性关注，值得精读。","inspiration":"该方法通过重复二元选择任务和试次级反馈学习数据，精细对比LLM与人类在聚合与过程层面的行为差异，值得借鉴｜可迁移到资产定价实验中的罕见事件学习，如研究投资者如何更新对崩盘风险的信念｜以LLM为被试，操纵历史收益序列呈现方式，测量其风险资产配置比例，并与真实投资者在类似实验中的选择数据对照"}},{"id":"2503.09639","version":4,"title":"Can A Society of Generative Agents Simulate Human Behavior and Inform Public Health Policy? A Case Study on Vaccine Hesitancy","zh_title":"生成式智能体社会能否模拟人类行为并为公共卫生政策提供信息？以疫苗犹豫为例","abstract":"Can we simulate a sandbox society with generative agents to model human behavior, thereby reducing the over-reliance on real human trials for assessing public policies? In this work, we investigate the feasibility of simulating health-related decision-making, using vaccine hesitancy, defined as the delay in acceptance or refusal of vaccines despite the availability of vaccination services (MacDonald, 2015), as a case study. To this end, we introduce the VacSim framework with 100 generative agents powered by Large Language Models (LLMs). VacSim simulates vaccine policy outcomes with the following steps: 1) instantiate a population of agents with demographics based on census data; 2) connect the agents via a social network and model vaccine attitudes as a function of social dynamics and disease-related information; 3) design and evaluate various public health interventions aimed at mitigating vaccine hesitancy. To align with real-world results, we also introduce simulation warmup and attitude modulation to adjust agents' attitudes. We propose a series of evaluations to assess the reliability of various LLM simulations. Experiments indicate that models like Llama and Qwen can simulate aspects of human behavior but also highlight real-world alignment challenges, such as inconsistent responses with demographic profiles. This early exploration of LLM-driven simulations is not meant to serve as definitive policy guidance; instead, it serves as a call for action to examine social simulation for policy development.","authors":["Abe Bohan Hou","Hongru Du","Yichen Wang","Jingyu Zhang","Zixiao Wang","Paul Pu Liang","Daniel Khashabi","Lauren Gardner","Tianxing He"],"categories":["cs.MA","cs.AI","cs.CL","cs.CY","cs.HC"],"primary_category":"cs.MA","announce_type":"new","date":"2025-03-12","first_seen":"2025-03-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.09639","pdf_url":"https://arxiv.org/pdf/2503.09639","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM人类仿真","疫苗犹豫","政策评估"],"reason":"用LLM代理模拟疫苗犹豫行为并与真实人口数据对照，评估仿真可靠性，直接命中核心…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":32,"question":"能否用生成式智能体构建的沙盒社会模拟人类疫苗犹豫行为，从而为公共卫生政策提供参考？","design":"使用100个由LLM驱动的生成式智能体，基于人口普查数据赋予人口学特征，通过社交网络和新闻信息模拟疫苗态度变化，并设计不同公共卫生干预措施（如政策）来观察态度轨迹。","baseline":"Nguyen et al. (2022) 的美国COVID-19疫苗犹豫调查数据，包含1300万份真实人类回答，用于校准智能体的人口学分布。","findings":"Llama和Qwen等模型能模拟人类行为的某些方面，但存在与现实对齐的挑战，例如智能体的回答与人口学特征不一致。","reliability":"论文承认仿真存在现实对齐挑战，如回答与人口学特征不一致，并指出当前探索不能作为政策指导，仅呼吁进一步研究。","relevance":"该研究直接以LLM代理模拟疫苗犹豫行为，并与真实人口调查数据对照，评估仿真可靠性，完全命中研究者对经济学实验和政策评估场景的兴趣，值得精读原文。","inspiration":"借鉴其利用人口普查数据校准智能体人口学特征并与大规模真实调查数据对照的仿真设计｜可迁移到政策公告对消费者通胀预期形成的实验，如模拟不同央行沟通策略对预期的影响｜以LLM智能体为被试，按收入、教育等特征分层，施加不同措辞的政策公告作为处理，测量预期通胀率的变化，用密歇根大学消费者调查的真实预期数据做对照"}},{"id":"2503.07510","version":1,"title":"Sometimes the Model doth Preach: Quantifying Religious Bias in Open LLMs through Demographic Analysis in Asian Nations","zh_title":"有时模型在布道：通过亚洲国家人口统计分析量化开放LLM中的宗教偏见","abstract":"Large Language Models (LLMs) are capable of generating opinions and propagating bias unknowingly, originating from unrepresentative and non-diverse data collection. Prior research has analysed these opinions with respect to the West, particularly the United States. However, insights thus produced may not be generalized in non-Western populations. With the widespread usage of LLM systems by users across several different walks of life, the cultural sensitivity of each generated output is of crucial interest. Our work proposes a novel method that quantitatively analyzes the opinions generated by LLMs, improving on previous work with regards to extracting the social demographics of the models. Our method measures the distance from an LLM's response to survey respondents, through Hamming Distance, to infer the demographic characteristics reflected in the model's outputs. We evaluate modern, open LLMs such as Llama and Mistral on surveys conducted in various global south countries, with a focus on India and other Asian nations, specifically assessing the model's performance on surveys related to religious tolerance and identity. Our analysis reveals that most open LLMs match a single homogeneous profile, varying across different countries/territories, which in turn raises questions about the risks of LLMs promoting a hegemonic worldview, and undermining perspectives of different minorities. Our framework may also be useful for future research investigating the complex intersection between training data, model architecture, and the resulting biases reflected in LLM outputs, particularly concerning sensitive topics like religious tolerance and identity.","authors":["Hari Shankar","Vedanta S P","Tejas Cavale","Ponnurangam Kumaraguru","Abhijnan Chakraborty"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"new","date":"2025-03-10","first_seen":"2025-03-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.07510","pdf_url":"https://arxiv.org/pdf/2503.07510","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","宗教偏见","人类数据对照"],"reason":"用LLM复现调查回答并与真实人类数据对照，评估宗教偏见，揭示仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":42,"question":"开放大语言模型在回答宗教相关调查时，反映了怎样的社会人口特征和宗教偏见？","design":"使用Llama、Mistral等开源LLM，以零样本方式回答皮尤研究中心在印度及东亚、东南亚国家进行的宗教宽容与认同调查问卷，通过汉明距离计算模型回答与真实受访者回答的匹配度，推断模型所反映的人口统计特征。","baseline":"皮尤研究中心在印度、日本、香港、韩国、台湾、越南、印尼、马来西亚、新加坡、斯里兰卡、泰国等亚洲国家/地区收集的真实调查数据，包含受访者的宗教、性别、年龄、教育等人口统计变量。","findings":"大多数开源LLM的回答与单一同质化的人口统计画像相匹配，且该画像在不同国家/地区间存在差异；通过提示词指示模型扮演特定群体并未显著改变其回答画像。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查数据为基准，评估LLM在宗教宽容等敏感话题上的回答偏差，揭示了模型输出同质化及可能强化霸权世界观的失效模式，高度契合研究者对LLM仿真可靠性及失效条件的关注，值得阅读原文。","inspiration":"借鉴其使用真实调查数据作为基准，通过汉明距离量化LLM回答与人类群体回答的匹配度，并推断模型隐含的人口统计画像的方法｜可迁移到信贷审批中的宗教或种族偏见检测，例如评估LLM在模拟贷款决策时是否系统性地偏向或歧视特定宗教群体｜以LLM作为被试，向其呈现不同宗教背景的贷款申请人资料，要求做出批准/拒绝决策，结果变量为批准率差异，并以真实银行信贷审批数据或审计研究结果作为对照基准"}},{"id":"2503.05529","version":1,"title":"PoSSUM: A Protocol for Surveying Social-media Users with Multimodal LLMs","zh_title":"PoSSUM：一种利用多模态大语言模型调查社交媒体用户的协议","abstract":"This paper introduces PoSSUM, an open-source protocol for unobtrusive polling of social-media users via multimodal Large Language Models (LLMs). PoSSUM leverages users' real-time posts, images, and other digital traces to create silicon samples that capture information not present in the LLM's training data. To obtain representative estimates, PoSSUM employs Multilevel Regression and Post-Stratification (MrP) with structured priors to counteract the observable selection biases of social-media platforms. The protocol is validated during the 2024 U.S. Presidential Election, for which five PoSSUM polls were conducted and published on GitHub and X. In the final poll, fielded October 17-26 with a synthetic sample of 1,054 X users, PoSSUM accurately predicted the outcomes in 50 of 51 states and assigned the Republican candidate a win probability of 0.65. Notably, it also exhibited lower state-level bias than most established pollsters. These results demonstrate PoSSUM's potential as a fully automated, unobtrusive alternative to traditional survey methods.","authors":["Roberto Cerina"],"categories":["stat.AP","cs.SI"],"primary_category":"stat.AP","announce_type":"new","date":"2025-03-07","first_seen":"2025-03-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.05529","pdf_url":"https://arxiv.org/pdf/2503.05529","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","选举预测","人类行为对照"],"reason":"用LLM模拟社交媒体用户投票行为，并与真实选举结果对照，属于人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:59","error":null,"has_summary":true,"summary":{"generated_at":"2025-03-07","rank":9,"question":"能否利用多模态大语言模型和社交媒体数字痕迹，构建无侵扰的民意调查方法，并准确预测选举结果？","design":"PoSSUM协议：使用多模态LLM基于X用户的实时帖子、图像等数字痕迹生成合成样本（硅样本），通过多水平回归与事后分层（MrP）校正平台选择偏差，测量投票意向。在2024年美国总统选举中进行了5次民意调查，最终轮使用1054名X用户的合成样本。","baseline":"2024年美国总统选举的真实结果（各州胜负及全国胜率）。","findings":"最终轮预测准确预测了51个州中50个的结果，并赋予共和党候选人0.65的胜率；其州级偏差低于大多数传统民调机构。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：该研究直接用LLM生成合成样本模拟人类投票行为，并以真实选举结果为对照，验证了LLM仿真在选举预测中的有效性，符合研究者对经济学实验和政策评估场景的关注。值得精读原文以了解MrP校正偏差的具体方法及协议细节。","inspiration":"借鉴该方法利用多模态LLM从社交媒体数字痕迹生成合成样本，并通过多水平回归与事后分层校正选择偏差，以低成本、无侵扰方式测量群体态度｜可迁移至政策公告的预期形成研究，如央行沟通对通胀预期的影响，或财政刺激对消费者信心的影响｜以X平台用户为合成样本来源，用多模态LLM提取用户对政策公告的态度，处理为不同措辞或发布渠道的公告版本，结果变量为通胀预期指数，以真实消费者调查数据（如密歇根大学消费者信心指数）作为对照基准"}},{"id":"2502.16280","version":1,"title":"Human Preferences in Large Language Model Latent Space: A Technical Analysis on the Reliability of Synthetic Data in Voting Outcome Prediction","zh_title":"大语言模型潜在空间中的人类偏好：合成数据在投票结果预测中可靠性的技术分析","abstract":"Generative AI (GenAI) is increasingly used in survey contexts to simulate human preferences. While many research endeavors evaluate the quality of synthetic GenAI data by comparing model-generated responses to gold-standard survey results, fundamental questions about the validity and reliability of using LLMs as substitutes for human respondents remain. Our study provides a technical analysis of how demographic attributes and prompt variations influence latent opinion mappings in large language models (LLMs) and evaluates their suitability for survey-based predictions. Using 14 different models, we find that LLM-generated data fails to replicate the variance observed in real-world human responses, particularly across demographic subgroups. In the political space, persona-to-party mappings exhibit limited differentiation, resulting in synthetic data that lacks the nuanced distribution of opinions found in survey data. Moreover, we show that prompt sensitivity can significantly alter outputs for some models, further undermining the stability and predictiveness of LLM-based simulations. As a key contribution, we adapt a probe-based methodology that reveals how LLMs encode political affiliations in their latent space, exposing the systematic distortions introduced by these models. Our findings highlight critical limitations in AI-generated survey data, urging caution in its use for public opinion research, social science experimentation, and computational behavioral modeling.","authors":["Sarah Ball","Simeon Allmendinger","Frauke Kreuter","Niklas Kühl"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2025-02-22","first_seen":"2025-02-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.16280","pdf_url":"https://arxiv.org/pdf/2502.16280","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B3","B4"],"tags":["LLM人类仿真","合成数据可靠性","政治偏好预测"],"reason":"直接评估LLM仿真人类投票偏好的可靠性与偏差，有真实调查数据对照，并揭示失效条…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":61,"question":"LLM生成的合成数据在多大程度上能复现真实人类调查回答的分布，以及提示不稳定性如何在模型的潜在空间中体现？","design":"使用14个白盒LLM，基于德国选举研究（GLES）构建包含年龄、性别、教育、收入、就业、政治倾向、东西德等人口属性的理论驱动型人设提示，让模型预测投票选择；同时通过改写提示考察提示敏感性，并利用Wahl-o-Mat数据训练探针分析模型潜在空间中的政治映射。","baseline":"以2021年德国纵向选举研究（GLES）的加权横截面调查数据作为真实人类投票行为的对照基准。","findings":"LLM生成数据无法复现真实人类回答的方差，尤其在人口子群中，人设到政党的映射区分度低，缺乏真实调查中的细致意见分布；提示敏感性会显著改变部分模型的输出，进一步削弱仿真稳定性，且在某些模型中潜在空间熵值与提示敏感性相关。","reliability":"论文指出LLM仿真在人口子群方差复现、意见分布细致度和提示稳定性方面存在根本局限，警告在公共舆论研究、社会科学实验和计算行为建模中使用合成数据需谨慎，但未讨论模型选择、语言文化差异等外部效度限制。","relevance":"该研究直接评估LLM替代人类被试进行投票预测的可靠性与偏差，有真实调查数据对照，并揭示了仿真在方差复现和提示敏感性上的失效条件，高度契合研究者对经济学实验和政策评估场景中仿真批判性分析的兴趣，值得精读原文。","inspiration":"该方法通过理论驱动构建人口属性人设提示并系统改写提示来检验仿真稳定性，为经济实验中的处理稳健性检验提供了可借鉴的测量范式｜可迁移至政策评估中的预期形成研究，例如考察不同人口子群对央行通胀预测公告的反应异质性｜以LLM模拟不同收入与教育水平的个体，处理为改写后的通胀预测措辞，结果变量为通胀预期调整幅度，以真实消费者预期调查数据作为对照基准"}},{"id":"2502.15800","version":3,"title":"LLM Agents Do Not Replicate Human Market Traders: Evidence From Experimental Finance","zh_title":"LLM代理无法复现人类市场交易者：来自实验金融的证据","abstract":"This paper explores how Large Language Models (LLMs) behave in a classic experimental finance paradigm widely known for eliciting bubbles and crashes in human participants. We adapt an established trading design, where traders buy and sell a risky asset with a known fundamental value, and introduce several LLM-based agents, both in single-model markets (all traders are instances of the same LLM) and in mixed-model \"battle royale\" settings (multiple LLMs competing in the same market). Our findings reveal that LLMs generally exhibit a \"textbook-rational\" approach, pricing the asset near its fundamental value, and show only a muted tendency toward bubble formation. Further analyses indicate that LLM-based agents display less trading strategy variance in contrast to humans. Taken together, these results highlight the risk of relying on LLM-only data to replicate human-driven market phenomena, as key behavioral features, such as large emergent bubbles, were not robustly reproduced. While LLMs clearly possess the capacity for strategic decision-making, their relative consistency and rationality suggest that they do not accurately mimic human market dynamics.","authors":["Thomas Henning","Siddhartha M. Ojha","Ross Spoon","Jiatong Han","Colin F. Camerer"],"categories":["q-fin.TR"],"primary_category":"q-fin.TR","announce_type":"new","date":"2025-02-18","first_seen":"2025-02-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.15800","pdf_url":"https://arxiv.org/pdf/2502.15800","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","实验金融","人类行为对照"],"reason":"用LLM代理模拟人类交易实验，与真实人类数据对照，发现LLM未能复现泡沫，批判…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"LLM代理在经典实验金融范式中能否复现人类交易者产生的资产泡沫与市场动态？","design":"采用Smith等人(2014)的固定基本面价值资产交易实验，在oTree平台上运行30期开放叫价市场。LLM代理分为单一模型同质市场和多个LLM模型混合的“大逃杀”市场，每轮提交买卖订单及价格预期，并允许通过“洞察”和“思考”文本实现跨轮记忆与思维链推理。结果变量为交易价格偏离基本面价值的程度、泡沫生成、交易策略方差等。","baseline":"对照真实人类被试在相同实验设计下的交易数据，人类市场一致产生显著价格泡沫和偏离基本面。","findings":"LLM代理普遍表现出“教科书式理性”，定价接近基本面价值，仅呈现微弱的泡沫倾向；其交易策略方差更低，更依赖基本面而非启发式策略，与人类行为存在系统性差异。","reliability":"论文指出，开箱即用的LLM未能复现人类市场中的大型泡沫等关键行为特征，表明仅依赖LLM数据复制人类驱动的市场现象存在风险，但未深入讨论LLM代理在何种条件下可能失效。","relevance":"该研究直接以真实人类实验为基准，检验LLM代理在金融实验中的行为复现能力，并得出批判性结论：LLM未能模拟人类市场泡沫，对评估LLM仿真可靠性及失效条件具有重要参考价值，值得精读原文。","inspiration":"该方法通过oTree平台实现多期开放叫价市场，并让LLM代理提交订单与价格预期，同时利用“洞察”和“思考”文本实现跨轮记忆与思维链推理，为经济实验的自动化仿真提供了可借鉴的设计框架｜可迁移至资产定价实验，用于检验不同信息结构或交易机制下LLM代理能否复现人类的价格泡沫与过度反应｜可设计一个资产定价实验，以LLM代理为被试，施加不同信息透明度处理，测量交易价格偏离基本面的程度，并与真实人类实验数据对照，评估LLM在模拟市场非理性行为时的有效性"}},{"id":"2502.10266","version":1,"title":"Are Large Language Models the future crowd workers of Linguistics?","zh_title":"大语言模型能否成为语言学未来的众包工作者？","abstract":"Data elicitation from human participants is one of the core data collection strategies used in empirical linguistic research. The amount of participants in such studies may vary considerably, ranging from a handful to crowdsourcing dimensions. Even if they provide resourceful extensive data, both of these settings come alongside many disadvantages, such as low control of participants' attention during task completion, precarious working conditions in crowdsourcing environments, and time-consuming experimental designs. For these reasons, this research aims to answer the question of whether Large Language Models (LLMs) may overcome those obstacles if included in empirical linguistic pipelines. Two reproduction case studies are conducted to gain clarity into this matter: Cruz (2023) and Lombard et al. (2021). The two forced elicitation tasks, originally designed for human participants, are reproduced in the proposed framework with the help of OpenAI's GPT-4o-mini model. Its performance with our zero-shot prompting baseline shows the effectiveness and high versatility of LLMs, that tend to outperform human informants in linguistic tasks. The findings of the second replication further highlight the need to explore additional prompting techniques, such as Chain-of-Thought (CoT) prompting, which, in a second follow-up experiment, demonstrates higher alignment to human performance on both critical and filler items. Given the limited scale of this study, it is worthwhile to further explore the performance of LLMs in empirical Linguistics and in other future applications in the humanities.","authors":["Iris Ferrazzo"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-02-14","first_seen":"2025-02-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.10266","pdf_url":"https://arxiv.org/pdf/2502.10266","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类数据对照","语言学实验"],"reason":"用LLM复现人类语言学任务，并与人类数据对照，讨论对齐与失效条件","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":90,"question":"大语言模型能否替代人类被试，成为实证语言学研究中新的众包工作者？","design":"使用GPT-4o-mini模型，通过零样本提示复现两项原为人类被试设计的强制诱导语言任务（Cruz 2023和Lombard et al. 2021），并在第二项任务中进一步尝试思维链提示，测量模型回答与人类回答的对齐程度。","baseline":"两项复现任务的原版人类被试数据，包括Cruz (2023)和Lombard et al. (2021)中的真实人类回答。","findings":"GPT-4o-mini在所有实验条件下均优于人类被试；思维链提示能进一步提升模型在关键项和填充项上与人类表现的对齐度。","reliability":"论文承认研究规模有限，仅覆盖两项案例，且未探索更多提示技术，未来需在更广泛的语言学任务和人文领域进一步验证。","relevance":"该研究直接用LLM复现人类语言学实验并与真实人类数据对照，探讨对齐条件与提示策略的影响，符合您对仿真可靠性及失效条件的关注，值得阅读原文以了解具体任务设计和偏差细节。","inspiration":"该方法借鉴了用零样本和思维链提示复现人类实验并与原版人类数据直接对照的范式，以评估LLM仿真对齐度｜可迁移到行为经济学中的跨期选择实验，检验LLM是否能复现人类的时间偏好不一致｜用LLM作为被试，施加不同时间折现的金钱选择任务，结果变量为折现率，对照真实人类实验数据（如Frederick等2002的经典数据）"}},{"id":"2502.10308","version":1,"title":"LLM-Powered Preference Elicitation in Combinatorial Assignment","zh_title":"基于大语言模型的组合分配偏好获取","abstract":"We study the potential of large language models (LLMs) as proxies for humans to simplify preference elicitation (PE) in combinatorial assignment. While traditional PE methods rely on iterative queries to capture preferences, LLMs offer a one-shot alternative with reduced human effort. We propose a framework for LLM proxies that can work in tandem with SOTA ML-powered preference elicitation schemes. Our framework handles the novel challenges introduced by LLMs, such as response variability and increased computational costs. We experimentally evaluate the efficiency of LLM proxies against human queries in the well-studied course allocation domain, and we investigate the model capabilities required for success. We find that our approach improves allocative efficiency by up to 20%, and these results are robust across different LLMs and to differences in quality and accuracy of reporting.","authors":["Ermis Soumalias","Yanchen Jiang","Kehang Zhu","Michael Curry","Sven Seuken","David C. Parkes"],"categories":["cs.AI","cs.GT","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2025-02-14","first_seen":"2025-02-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.10308","pdf_url":"https://arxiv.org/pdf/2502.10308","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM代理","偏好获取","人类对照"],"reason":"用LLM代理人类偏好，在课程分配场景与真实人类查询对照，可迁移到人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":39,"question":"如何利用大语言模型作为人类代理，通过一次性自然语言输入简化组合分配中的偏好诱导，并提升分配效率？","design":"使用大语言模型（如GPT系列）作为学生代理，根据学生提供的简短文本偏好描述，替代学生回答迭代比较查询；将LLM代理集成到MLCM机制中，通过链式思维提示和噪声鲁棒损失函数处理响应变异，测量分配效率的提升。","baseline":"真实人类学生在课程分配机制中通过迭代查询提供的偏好数据，以及Course Match机制下的分配结果。","findings":"LLM代理方法在现实场景中将分配效率提高了最多20%，且结果在不同LLM架构和报告质量差异下保持稳健。","reliability":"论文未讨论","relevance":"该研究直接以LLM代理人类偏好，在课程分配场景中与真实人类查询对照，可迁移到经济学实验和政策评估中的人类仿真研究，值得精读。","inspiration":"借鉴其将LLM作为代理集成到已有机制中，并通过噪声鲁棒设计处理LLM响应变异的方法。｜可迁移到公共资源分配或拍卖设计实验中，用LLM模拟投标者偏好以测试机制效率。｜以LLM作为被试，施加不同文本偏好描述作为处理，测量拍卖分配效率，并与真实人类实验数据对照。"}},{"id":"2503.05708","version":1,"title":"On Large Language Models as Data Sources for Policy Deliberation on Climate Change and Sustainability","zh_title":"大语言模型作为气候与可持续性政策审议数据源的研究","abstract":"We pose the research question, \"Can LLMs provide credible evaluation scores, suitable for constructing starter MCDM models that support commencing deliberation regarding climate and sustainability policies?\" In this exploratory study we i. Identify a number of interesting policy alternatives that are actively considered by local governments in the United States (and indeed around the world). ii. Identify a number of quality-of-life indicators as apt evaluation criteria for these policies. iii. Use GPT-4 to obtain evaluation scores for the policies on multiple criteria. iv. Use the TOPSIS MCDM method to rank the policies based on the obtained evaluation scores. v. Evaluate the quality and validity of the resulting table ensemble of scores by comparing the TOPSIS-based policy rankings with those obtained by an informed assessment exercise. We find that GPT-4 is in rough agreement with the policy rankings of our informed assessment exercise. Hence, we conclude (always provisionally and assuming a modest level of vetting) that GPT-4 can be used as a credible input, even starting point, for subsequent deliberation processes on climate and sustainability policies.","authors":["Rachel Bina","Kha Luong","Shrey Mehta","Daphne Pang","Mingjun Xie","Christine Chou","Steven O. Kimbrough"],"categories":["cs.CY","econ.GN"],"primary_category":"cs.CY","announce_type":"new","date":"2025-02-13","first_seen":"2025-02-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.05708","pdf_url":"https://arxiv.org/pdf/2503.05708","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","政策评估","人类数据对照"],"reason":"用GPT-4替代人类专家评估政策，并与人类评估对照，属于仿真人类决策，但非严格…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":126,"question":"大语言模型能否为气候与可持续政策提供可信的评估分数，用于构建启动政策协商的初始多准则决策模型？","design":"本研究并非严格的人类仿真实验，而是使用GPT-4对一系列气候与可持续政策在多个生活质量指标上进行评分，然后采用TOPSIS多准则决策方法对政策排序，并将结果与知情评估（informed assessment）产生的排序进行对比。","baseline":"对照的真实人类数据来自一项知情评估练习（informed assessment exercise），由人类专家对部分政策进行评价和排序，作为比较基准。","findings":"GPT-4生成的政策排序与知情评估的排序大致一致，尤其在生活质量准则上的排序高度吻合。因此，GPT-4可作为政策协商的初始可信输入或起点。","reliability":"论文强调结论是暂时的，需假设一定程度的审查，并指出模型输出应作为草案，需在动态环境中持续修订和审议，但未系统讨论失效条件或具体局限。","relevance":"该研究用GPT-4替代人类专家进行政策评估并与人类判断对照，属于用LLM仿真人类决策的探索，但非严格的行为实验复现，且缺乏对仿真偏差的深入批判，适合作为方法参考而非可靠性评估的典型案例。","inspiration":"该方法借鉴了用LLM替代人类专家进行多准则评分，并与真实人类判断排序对照的验证思路。｜可迁移到政策公告对金融市场预期形成的评估，如央行沟通对投资者情绪的影响。｜以GPT-4作为被试，输入不同措辞的央行声明，让其对资产价格预期进行评分，结果与真实市场调查数据或分析师预测排序进行对照。"}},{"id":"2502.08691","version":2,"title":"AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society","zh_title":"AgentSociety：大规模LLM驱动生成式代理模拟推进对人类行为与社会的理解","abstract":"Understanding human behavior and society is a central focus in social sciences, with the rise of generative social science marking a significant paradigmatic shift. By leveraging bottom-up simulations, it replaces costly and logistically challenging traditional experiments with scalable, replicable, and systematic computational approaches for studying complex social dynamics. Recent advances in large language models (LLMs) have further transformed this research paradigm, enabling the creation of human-like generative social agents and realistic simulacra of society. In this paper, we propose AgentSociety, a large-scale social simulator that integrates LLM-driven agents, a realistic societal environment, and a powerful large-scale simulation engine. Based on the proposed simulator, we generate social lives for over 10k agents, simulating their 5 million interactions both among agents and between agents and their environment. Furthermore, we explore the potential of AgentSociety as a testbed for computational social experiments, focusing on five key social issues: polarization, the spread of inflammatory messages, the effects of universal basic income policies, the impact of external shocks such as hurricanes, and urban sustainability. These five issues serve as valuable cases for assessing AgentSociety's support for typical research methods -- such as surveys, interviews, and interventions -- as well as for investigating the patterns, causes, and underlying mechanisms of social issues. The alignment between AgentSociety's outcomes and real-world experimental results not only demonstrates its ability to capture human behaviors and their underlying mechanisms, but also underscores its potential as an important platform for social scientists and policymakers.","authors":["Jinghua Piao","Yuwei Yan","Jun Zhang","Nian Li","Junbo Yan","Xiaochong Lan","Zhihong Lu","Zhiheng Zheng","Jing Yi Wang","Di Zhou","Chen Gao","Fengli Xu","Fang Zhang","Ke Rong","Jun Su","Yong Li"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"new","date":"2025-02-12","first_seen":"2025-02-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.08691","pdf_url":"https://arxiv.org/pdf/2502.08691","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B3"],"tags":["LLM社会仿真","人类行为复现","政策评估"],"reason":"用LLM代理模拟社会行为并与真实数据对照，直接复现人类实验","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:56","error":null,"has_summary":true,"summary":{"generated_at":"2025-02-12","rank":1,"question":"LLM驱动的生成式智能体能否在大规模社会仿真中复现人类行为与社会现象？","design":"构建AgentSociety仿真器，包含LLM驱动的智能体（基于GPT等模型）、现实社会环境和大规模仿真引擎；模拟超过1万个智能体的500万次交互（智能体间及智能体与环境间）；针对极化、煽动性信息传播、全民基本收入政策、飓风外部冲击、城市可持续性五个社会议题，支持调查、访谈、干预等研究方法。","baseline":"真实世界实验结果（具体未在摘要和引言中详述，但提及结果与真实实验对齐）。","findings":"AgentSociety能够捕捉人类行为及其潜在机制，仿真结果与真实世界实验结果一致；展示了作为社会科学家和政策制定者重要平台的潜力。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接使用LLM作为人类被试替代品，有真实人类数据对照，涉及政策评估（如全民基本收入）等经济学场景，且关注仿真可靠性，值得精读原文以获取更多细节。","inspiration":"借鉴其大规模LLM智能体仿真与真实实验对齐的方法，可对智能体施加政策干预并测量行为变化｜可迁移到全民基本收入（UBI）对劳动供给与消费行为影响的经济学问题｜用LLM智能体作为被试，处理组接受UBI转账，对照组无干预，结果变量为工作时长与消费支出，对照真实UBI实验数据（如芬兰基本收入实验）"}},{"id":"2502.07307","version":1,"title":"CreAgent: Towards Long-Term Evaluation of Recommender System under Platform-Creator Information Asymmetry","zh_title":"CreAgent：平台-创作者信息不对称下推荐系统长期评估","abstract":"Ensuring the long-term sustainability of recommender systems (RS) emerges as a crucial issue. Traditional offline evaluation methods for RS typically focus on immediate user feedback, such as clicks, but they often neglect the long-term impact of content creators. On real-world content platforms, creators can strategically produce and upload new items based on user feedback and preference trends. While previous studies have attempted to model creator behavior, they often overlook the role of information asymmetry. This asymmetry arises because creators primarily have access to feedback on the items they produce, while platforms possess data on the entire spectrum of user feedback. Current RS simulators, however, fail to account for this asymmetry, leading to inaccurate long-term evaluations. To address this gap, we propose CreAgent, a Large Language Model (LLM)-empowered creator simulation agent. By incorporating game theory's belief mechanism and the fast-and-slow thinking framework, CreAgent effectively simulates creator behavior under conditions of information asymmetry. Additionally, we enhance CreAgent's simulation ability by fine-tuning it using Proximal Policy Optimization (PPO). Our credibility validation experiments show that CreAgent aligns well with the behaviors between real-world platform and creator, thus improving the reliability of long-term RS evaluations. Moreover, through the simulation of RS involving CreAgents, we can explore how fairness- and diversity-aware RS algorithms contribute to better long-term performance for various stakeholders. CreAgent and the simulation platform are publicly available at https://github.com/shawnye2000/CreAgent.","authors":["Xiaopeng Ye","Chen Xu","Zhongxiang Sun","Jun Xu","Gang Wang","Zhenhua Dong","Ji-Rong Wen"],"categories":["cs.IR"],"primary_category":"cs.IR","announce_type":"new","date":"2025-02-11","first_seen":"2025-02-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.07307","pdf_url":"https://arxiv.org/pdf/2502.07307","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","推荐系统仿真","创作者行为模拟"],"reason":"用LLM模拟创作者行为，属于社会模拟但无真实人类数据对照，且非直接仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":49,"question":"在平台与创作者信息不对称下，如何利用LLM智能体模拟创作者行为，以实现对推荐系统的长期评估？","design":"使用LLM驱动的创作者智能体CreAgent，结合博弈论信念机制和快慢思考框架模拟创作者在信息不对称下的内容生产决策，并通过PPO微调提升仿真能力。","baseline":"无对照","findings":"CreAgent能够有效模拟真实平台中创作者在信息不对称下的行为，提高了推荐系统长期评估的可靠性；通过仿真实验，可探索公平性和多样性感知推荐算法对多方长期绩效的影响。","reliability":"论文未讨论","relevance":"该研究用LLM模拟创作者行为，属于社会模拟但无真实人类数据对照，且非直接仿真人类被试，与研究者关注的人类仿真实验及经济学实验场景关联较弱。","inspiration":"与经济金融研究关联不大"}},{"id":"2502.03158","version":3,"title":"Strategizing with AI: Insights from a Beauty Contest Experiment","zh_title":"与AI博弈：选美竞赛实验的启示","abstract":"A $p$-beauty contest is a wide class of games of guessing the most popular strategy among other players. In particular, guessing a fraction of a mean of numbers chosen by all players is a classic behavioral experiment designed to test iterative reasoning patterns among various groups of people. The previous literature reveals that the level of sophistication of the opponents is an important factor affecting the outcome of the game. Smarter decision makers choose strategies that are closer to theoretical Nash equilibrium and demonstrate faster convergence to equilibrium in iterated contests with information revelation. We replicate a series of classic experiments by running virtual experiments with large language models (LLMs) who play against various groups of virtual players. Our results show that LLMs recognize strategic context of the game and demonstrate expected adaptability to the changing set of parameters. LLMs systematically behave in a more sophisticated way compared to the participants of the original experiments. All LLMs still fail to identify dominant strategies in a two-player game. Our results contribute to the discussion on the accuracy of modeling human economic agents by artificial intelligence.","authors":["Iuliia Alekseenko","Dmitry Dagaev","Sofia Paklina","Petr Parshakov"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-02-05","first_seen":"2025-02-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.03158","pdf_url":"https://arxiv.org/pdf/2502.03158","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为博弈","人类数据对照"],"reason":"用LLM复现选美博弈实验，与真实人类数据对照，评估仿真准确性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":28,"question":"大语言模型在选美博弈（p-beauty contest）中的策略行为是否与人类相似，能否准确模拟人类经济主体？","design":"用多个大语言模型（LLMs）作为虚拟被试，复现经典的“猜数字”实验，让LLMs与不同虚拟对手群体进行博弈，操纵参数p和对手类型，测量其选择的数字、收敛速度及对纳什均衡的偏离。","baseline":"以Nagel (1995)等经典人类实验数据作为对照基准，比较LLMs与真实人类参与者的行为差异。","findings":"LLMs能识别博弈的战略情境并适应参数变化，行为比人类更复杂；但在两人博弈中，所有LLMs均未能识别占优策略。","reliability":"论文未讨论","relevance":"该研究直接使用LLMs复现经典行为博弈实验，并与真实人类数据对照，评估仿真准确性，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"借鉴其用经典实验范式复现并系统操纵对手类型和参数的方法，可迁移到资产定价实验中的预期形成研究。｜设计一个LLM作为交易者的资产市场实验，操纵市场信息结构和对手策略复杂度，测量LLM的报价偏差和收敛速度，并与人类实验数据对照。"}},{"id":"2502.00070","version":2,"title":"Can AI Solve the Peer Review Crisis? A Large Scale Cross Model Experiment of LLMs' Performance and Biases in Evaluating over 1000 Economics Papers","zh_title":"AI能解决同行评审危机吗？一项关于LLM在评估1000多篇经济学论文中的表现与偏差的大规模跨模型实验","abstract":"This study examines the potential of large language models (LLMs) to augment the academic peer review process by reliably evaluating the quality of economics research without introducing systematic bias. We conduct one of the first large-scale experimental assessments of four LLMs (GPT-4o, Claude 3.5, Gemma 3, and LLaMA 3.3) across two complementary experiments. In the first, we use nonparametric binscatter and linear regression techniques to analyze over 29,000 evaluations of 1,220 anonymized papers drawn from 110 economics journals excluded from the training data of current LLMs, along with a set of AI-generated submissions. The results show that LLMs consistently distinguish between higher- and lower-quality research based solely on textual content, producing quality gradients that closely align with established journal prestige measures. Claude and Gemma perform exceptionally well in capturing these gradients, while GPT excels in detecting AI-generated content. The second experiment comprises 8,910 evaluations designed to assess whether LLMs replicate human like biases in single blind reviews. By systematically varying author gender, institutional affiliation, and academic prominence across 330 papers, we find that GPT, Gemma, and LLaMA assign significantly higher ratings to submissions from top male authors and elite institutions relative to the same papers presented anonymously. These results emphasize the importance of excluding author-identifying information when deploying LLMs in editorial screening. Overall, our findings provide compelling evidence and practical guidance for integrating LLMs into peer review to enhance efficiency, improve accuracy, and promote equity in the publication process of economics research.","authors":["Pat Pataranutaporn","Nattavudh Powdthavee","Chayapatr Achiwaranguprok","Pattie Maes"],"categories":["cs.CY","cs.AI","econ.GN"],"primary_category":"cs.CY","announce_type":"new","date":"2025-01-31","first_seen":"2025-01-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.00070","pdf_url":"https://arxiv.org/pdf/2502.00070","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM审稿","经济学论文评估","偏差分析"],"reason":"用LLM替代审稿人，属于替代人类劳动而非仿真被试，但涉及人类审稿数据对照，边界…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":48,"question":"大型语言模型能否可靠地评估经济学论文质量，并在单盲评审中是否会复现人类审稿人的系统性偏见？","design":"本研究并非用LLM仿真人类被试，而是用GPT-4o、Claude 3.5、Gemma 3、LLaMA 3.3四种模型扮演审稿人，对1220篇匿名经济学论文进行质量评分，并在一项实验中系统操纵作者性别、机构声誉和学术地位，测量评分变化。","baseline":"以论文所在期刊的IDEAS/RePEc综合排名作为客观质量基准，并与人类审稿中已知的偏见模式（如对顶尖机构男性作者的偏好）进行对照。","findings":"LLM能仅凭文本内容区分论文质量高低，评分梯度与期刊声望高度一致，其中Claude和Gemma表现最佳，GPT在识别AI生成内容上更优。但在单盲条件下，GPT、Gemma和LLaMA对顶尖男性作者和精英机构的论文给出显著更高评分，表现出与人类相似的偏见。","reliability":"论文指出，LLM在单盲评审中会复现甚至放大人类偏见，因此建议在编辑筛选中必须隐去作者身份信息；未讨论模型在其他学科、语言或更复杂评审维度上的泛化局限。","relevance":"该研究直接涉及用LLM替代人类评审，并系统检验了偏见复现问题，虽非典型人类仿真实验，但为经济学领域AI辅助决策的可靠性与偏差提供了关键证据，值得精读。","inspiration":"借鉴其通过操纵作者身份特征来检测模型偏见的实验设计，以及用期刊排名作为质量基准的对照思路。｜可迁移到经济学论文评审、基金申请书评估或信贷审批中的歧视研究。｜以LLM作为评审人，随机分配附有不同作者特征（性别、机构）的同一篇论文，测量评分差异，并以真实期刊排名或人类评审数据作为基准，检验模型偏见。"}},{"id":"2501.17310","version":4,"title":"Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding","zh_title":"探究LLM世界模型：用群体智慧解码增强估算能力","abstract":"Guesstimation -- the task of making approximate quantitative estimates about objects or events -- is a common real-world skill, yet remains underexplored in large language model (LLM) research. We introduce three guesstimation datasets: MARBLES, FUTURE, and ELECPRED, spanning physical estimation (e.g., how many marbles fit in a cup) to abstract predictions (e.g., the 2024 U.S. presidential election). Inspired by the social science concept of Wisdom of Crowds (WOC)- where the median of multiple estimates improves accuracy-we propose WOC decoding for LLMs. We replicate WOC effects in human participants and find that LLMs exhibit similar benefits: median aggregation across sampled responses consistently improves accuracy over greedy decoding, self-consistency decoding, and mean decoding. This suggests that LLMs encode a world model that supports approximate reasoning. Our results position guesstimation as a useful probe of LLM world knowledge and highlight WOC decoding as a strategy for enhancing LLM guesstimation performance on real-world tasks.","authors":["Yun-Shiuan Chuang","Sameer Narendran","Nikunj Harlalka","Alexander Cheung","Sizhe Gao","Siddharth Suresh","Junjie Hu","Timothy T. Rogers"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2025-01-28","first_seen":"2025-01-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.17310","pdf_url":"https://arxiv.org/pdf/2501.17310","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","群体智慧","人类对照"],"reason":"用LLM复现人类估计任务，并与真实人类数据对照，涉及选举预测等政策场景，方法可…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":91,"question":"LLM在猜测估计任务中是否表现出类似人类的“群体智慧”效应，以及如何利用该效应提升估计准确性？","design":"研究使用10个LLM（LLaMA、Mistral、Mixtral、GPT等）作为被试，通过重复采样生成多个估计值，应用WOC解码（取中位数）与其他解码策略对比，测量归一化误差；同时构建了三个猜测估计数据集（MARBLES、FUTURE、ELECPRED），并进行了人类实验以复现WOC效应。","baseline":"人类实验：招募人类参与者完成MARBLES数据集中的估计任务，复现了人类群体中的WOC效应，作为LLM行为的对照基准。","findings":"LLM在猜测估计任务中表现出与人类相似的WOC效应：对多次采样响应取中位数能持续提升准确性，优于贪婪解码、自一致性解码和均值解码。这表明LLM编码了支持近似推理的世界模型。","reliability":"论文未讨论","relevance":"该研究将LLM作为人类被试的替代品，复现了群体智慧效应，并与真实人类数据对照，涉及选举预测等政策场景，方法可用于评估LLM仿真人类估计行为的可靠性，值得精读。","inspiration":"该方法通过重复采样LLM生成多个估计值并取中位数来复现群体智慧效应，可借鉴其解码策略作为处理手段，并设置贪婪解码、均值解码等作为对照｜可迁移到资产定价实验中的市场预期形成，例如研究投资者对股票收益的集体预测是否表现出群体智慧｜以LLM作为被试，处理为对同一股票收益预测多次采样后取中位数，结果变量为预测误差，以真实分析师一致预期数据作为人类基准对照"}},{"id":"2501.14294","version":3,"title":"Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes","zh_title":"通过代表性启发式检验大语言模型的对齐：以政治刻板印象为例","abstract":"Examining the alignment of large language models (LLMs) has become increasingly important, e.g., when LLMs fail to operate as intended. This study examines the alignment of LLMs with human values for the domain of politics. Prior research has shown that LLM-generated outputs can include political leanings and mimic the stances of political parties on various issues. However, the extent and conditions under which LLMs deviate from empirical positions are insufficiently examined. To address this gap, we analyze the factors that contribute to LLMs' deviations from empirical positions on political issues, aiming to quantify these deviations and identify the conditions that cause them. Drawing on findings from cognitive science about representativeness heuristics, i.e., situations where humans lean on representative attributes of a target group in a way that leads to exaggerated beliefs, we scrutinize LLM responses through this heuristics' lens. We conduct experiments to determine how LLMs inflate predictions about political parties, which results in stereotyping. We find that while LLMs can mimic certain political parties' positions, they often exaggerate these positions more than human survey respondents do. Also, LLMs tend to overemphasize representativeness more than humans. This study highlights the susceptibility of LLMs to representativeness heuristics, suggesting a potential vulnerability of LLMs that facilitates political stereotyping. We also test prompt-based mitigation strategies, finding that strategies that can mitigate representative heuristics in humans are also effective in reducing the influence of representativeness on LLM-generated responses.","authors":["Sullam Jeoung","Yubin Ge","Haohan Wang","Jana Diesner"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-01-24","first_seen":"2025-01-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.14294","pdf_url":"https://arxiv.org/pdf/2501.14294","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","政治态度模拟","代表性启发式"],"reason":"用LLM模拟人类政治态度并与调查数据对照，发现LLM夸大刻板印象，评估仿真偏差…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":37,"question":"LLM在政治议题上是否会像人类一样因代表性启发式而产生刻板印象，并偏离真实立场？","design":"用LLM（如GPT系列）模拟美国民主党和共和党的立场，通过双问题框架：实证部分使用自认党派的人类调查数据，预测部分让LLM和人类分别回答政治议题，比较LLM与人类在预测党派立场时的偏差程度，并测试提示词缓解策略。","baseline":"自认为民主党或共和党的真实人类调查数据（公开的现有调查数据）。","findings":"LLM能近似各党派立场，但比人类更夸大党派差异，表现出更强的代表性启发式；提示词干预可部分缓解这种刻板印象，但无法完全消除。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查为基准，评估LLM在政治态度模拟中的偏差，揭示LLM比人类更易产生刻板印象，符合你对仿真可靠性与失效条件的关注，值得精读。","inspiration":"该方法通过双问题框架（实证部分用真实人类数据，预测部分让LLM和人类分别回答）来量化LLM的启发式偏差，值得借鉴｜可迁移到信贷审批歧视研究，检验LLM是否像人类一样因代表性启发式而夸大种族或性别的信用差异｜让LLM模拟信贷员审批贷款，处理为申请人种族/性别信息，结果变量为审批通过率，以真实信贷审批数据中的人类决策分布为基准，比较LLM与人类的偏差程度"}},{"id":"2501.13955","version":1,"title":"Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?","zh_title":"基于引导式角色的AI调查：能否利用LLM大规模复制个人出行偏好？","abstract":"This study explores the potential of Large Language Models (LLMs) to generate artificial surveys, with a focus on personal mobility preferences in Germany. By leveraging LLMs for synthetic data creation, we aim to address the limitations of traditional survey methods, such as high costs, inefficiency and scalability challenges. A novel approach incorporating \"Personas\" - combinations of demographic and behavioural attributes - is introduced and compared to five other synthetic survey methods, which vary in their use of real-world data and methodological complexity. The MiD 2017 dataset, a comprehensive mobility survey in Germany, serves as a benchmark to assess the alignment of synthetic data with real-world patterns. The results demonstrate that LLMs can effectively capture complex dependencies between demographic attributes and preferences while offering flexibility to explore hypothetical scenarios. This approach presents valuable opportunities for transportation planning and social science research, enabling scalable, cost-efficient and privacy-preserving data generation.","authors":["Ioannis Tzachristas","Santhanakrishnan Narayanan","Constantinos Antoniou"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-01-20","first_seen":"2025-01-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.13955","pdf_url":"https://arxiv.org/pdf/2501.13955","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","合成调查数据","出行偏好"],"reason":"用LLM生成合成调查数据模拟人类出行偏好，并与真实调查数据MiD 2017对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"能否利用大语言模型通过基于人格的引导式AI调查方法，大规模复制德国个人出行偏好？","design":"使用GPT-4o生成合成调查数据，提出六种方法（朴素AI调查、结构化AI调查、引导式AI调查、朴素基于人格的AI调查、结构化基于人格的AI调查、引导式基于人格的AI调查），逐步引入真实人口结构和出行统计约束，生成10000个个体或15840个独特人格的出行偏好回答。","baseline":"德国2017年国家出行调查（MiD 2017）数据集，包含人口分布、交通方式和出行频率等真实数据。","findings":"引导式基于人格的AI调查方法在准确性和与真实出行行为模式的一致性上显著优于其他方法；LLM能有效捕捉人口属性与偏好之间的复杂依赖关系，并支持灵活探索假设情景。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查数据为基准，评估LLM仿真出行偏好的可靠性，并比较多种方法，符合研究者对经济学实验和政策评估场景下仿真有效性及偏差的关注，值得阅读原文。","inspiration":"该方法通过逐步引入真实人口统计约束和出行统计约束来校准LLM生成回答，可借鉴为在仿真中分层施加宏观分布约束以提升个体决策模拟的准确性｜可迁移到消费者跨期选择与储蓄行为仿真，利用LLM模拟不同人口特征群体的时间偏好和消费-储蓄决策｜以LLM作为被试，处理为提供不同利率或未来收入情景，结果变量为报告的消费/储蓄金额，用家庭金融调查（如SCF）的真实储蓄率分布作为对照基准"}},{"id":"2501.08579","version":3,"title":"LLM-based Human Simulations Have Not Yet Been Reliable","zh_title":"基于大语言模型的人类仿真尚未可靠","abstract":"Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations remain insufficiently reliable, as evidenced by significant discrepancies between their outcomes and authentic human actions. Our investigation begins with a systematic review of LLM-based human simulations in social, economic, policy, and psychological contexts, identifying their common frameworks, recent advances, and persistent limitations. This review reveals that such discrepancies primarily stem from inherent limitations of LLMs and flaws in simulation design, both of which are examined in detail. Building on these insights, we propose a systematic solution framework that emphasizes enriching data foundations, advancing LLM capabilities, and ensuring robust simulation design to enhance reliability. Finally, we introduce a structured algorithm that operationalizes the proposed framework, aiming to guide credible and human-aligned LLM-based simulations. To facilitate further research, we provide a curated list of related literature and resources at https://github.com/Persdre/awesome-llm-human-simulation.","authors":["Qian Wang","Jiaying Wu","Zichen Jiang","Zhenheng Tang","Bingqiao Luo","Nuo Chen","Wei Chen","Huacan Wang","Bingsheng He"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-01-15","first_seen":"2025-01-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.08579","pdf_url":"https://arxiv.org/pdf/2501.08579","source_feed":"api","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM人类仿真","可靠性评估","方法论框架"],"reason":"系统评估LLM人类仿真的可靠性，指出与真实人类行为的差异，并提出改进框架，直接…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":69,"question":"当前基于大语言模型的人类仿真是否可靠？","design":"本文并非仿真实验，而是对社会科学、经济学、政策、心理学等领域中基于LLM的人类仿真研究进行系统综述，分析其常用框架、进展与局限。","baseline":"无对照","findings":"当前LLM人类仿真不可靠，与真实人类行为存在显著差异；差异主要源于LLM的固有局限（如偏见、认知过程缺陷、行为不一致）和仿真框架设计缺陷（如过度简化心理状态、缺乏全面人类经验、验证机制不足）。","reliability":"论文指出失效条件包括：LLM内嵌的文化、性别等偏见扭曲行为模拟；认知过程局限损害决策真实性；记忆限制导致行为不一致；交互机制缺陷影响多智能体仿真；框架过度简化复杂心理状态和群体动态；缺乏严格验证与实时监控。","relevance":"该文直接回应研究者对LLM人类仿真可靠性的核心关切，系统梳理了仿真失效的根源并提出改进框架，对评估经济学实验和政策模拟的仿真偏差具有重要参考价值，强烈推荐阅读原文。","inspiration":"该文系统梳理了LLM仿真失效的根源（偏见、认知局限、验证不足），其批判性框架可借鉴用于设计经济学实验中的稳健性检验，例如通过对比不同LLM版本或提示策略来识别仿真偏差｜可迁移至政策公告的预期形成研究，评估LLM模拟的经济主体对货币政策或财政刺激的反应是否与真实调查数据一致｜以LLM作为被试，施加不同措辞的政策声明处理，测量其通胀预期或消费意愿，并与密歇根大学消费者调查等真实微观数据对照，检验仿真偏差"}},{"id":"2501.06834","version":1,"title":"LLMs Model Non-WEIRD Populations: Experiments with Synthetic Cultural Agents","zh_title":"LLM模拟非WEIRD人群：合成文化代理实验","abstract":"Despite its importance, studying economic behavior across diverse, non-WEIRD (Western, Educated, Industrialized, Rich, and Democratic) populations presents significant challenges. We address this issue by introducing a novel methodology that uses Large Language Models (LLMs) to create synthetic cultural agents (SCAs) representing these populations. We subject these SCAs to classic behavioral experiments, including the dictator and ultimatum games. Our results demonstrate substantial cross-cultural variability in experimental behavior. Notably, for populations with available data, SCAs' behaviors qualitatively resemble those of real human subjects. For unstudied populations, our method can generate novel, testable hypotheses about economic behavior. By integrating AI into experimental economics, this approach offers an effective and ethical method to pilot experiments and refine protocols for hard-to-reach populations. Our study provides a new tool for cross-cultural economic studies and demonstrates how LLMs can help experimental behavioral research.","authors":["Augusto Gonzalez-Bonorino","Monica Capra","Emilio Pantoja"],"categories":["cs.AI","cs.CL","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2025-01-12","first_seen":"2025-01-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.06834","pdf_url":"https://arxiv.org/pdf/2501.06834","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为实验","跨文化经济研究"],"reason":"用LLM创建合成文化代理模拟非WEIRD人群的经济实验行为，并与真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":28,"question":"如何利用大语言模型创建合成文化代理，以模拟非WEIRD人群在经典经济实验中的行为？","design":"使用GPT-4等大语言模型，结合网络爬取和检索增强生成构建六支小规模社会（Hadza、Machiguenga、Tsimané、Aché、Orma、Yanomami）的文化档案，据此实例化合成文化代理，让其参与独裁者博弈、最后通牒博弈和禀赋效应实验，测量出价、拒绝率等行为变量。","baseline":"对照Henrich等人（2005）等文献中真实人类在相同实验中的行为数据，对部分有数据的人群进行定性比较。","findings":"合成代理的行为展现出显著的跨文化差异，且与现有真实人类数据在性质上相似；所有代理均未表现出纯粹自利行为，该方法还能为未经研究的人群生成可检验的假设。","reliability":"论文承认LLM可能带有固有偏见，且合成代理的行为仍需与真实人类数据进行谨慎验证，该方法旨在补充而非替代人类被试研究。","relevance":"高度相关：该研究直接用LLM模拟非WEIRD人群的经济决策，并与真实人类实验数据对照，属于经济学实验场景下的人类仿真，同时讨论了方法的局限，完全契合你的关注点，值得精读原文。","inspiration":"该方法借鉴了用LLM结合文化档案构建合成代理来模拟特定人群决策行为的做法，通过检索增强生成注入文化背景，并与真实人类实验数据进行定性对照｜可迁移到跨文化消费信贷审批歧视研究，模拟不同文化背景申请人的还款决策与银行审批行为｜以GPT-4等LLM为被试，构建不同文化背景的合成申请人，处理为信贷条款（如利率、额度），结果变量为还款意愿与违约率，对照真实跨国信贷数据或田野实验数据"}},{"id":"2412.19363","version":3,"title":"Large Language Models for Market Research: A Data-augmentation Approach","zh_title":"用于市场研究的大语言模型：一种数据增强方法","abstract":"Large Language Models (LLMs) have transformed artificial intelligence by excelling in complex natural language processing tasks. Their ability to generate human-like text has opened new possibilities for market research, particularly in conjoint analysis, where understanding consumer preferences is essential but often resource-intensive. Traditional survey-based methods face limitations in scalability and cost, making LLM-generated data a promising alternative. However, while LLMs have the potential to simulate real consumer behavior, recent studies highlight a significant gap between LLM-generated and human data, with biases introduced when substituting between the two. In this paper, we address this gap by proposing a novel statistical data augmentation approach that efficiently integrates LLM-generated data with real data in conjoint analysis. This results in statistically robust estimators with consistent and asymptotically normal properties, in contrast to naive approaches that simply substitute human data with LLM-generated data, which can exacerbate bias. We further present a finite-sample performance bound on the estimation error. We validate our framework through an empirical study on COVID-19 vaccine preferences, demonstrating its superior ability to reduce estimation error and save data and costs by 24.9% to 79.8%. In contrast, naive approaches fail to save data due to the inherent biases in LLM-generated data compared to human data. Another empirical study on sports car choices validates the robustness of our results. Our findings suggest that while LLM-generated data is not a direct substitute for human responses, it can serve as a valuable complement when used within a robust statistical framework.","authors":["Mengxin Wang","Dennis J. Zhang","Heng Zhang"],"categories":["cs.AI","cs.LG","stat.ME","stat.ML"],"primary_category":"cs.AI","announce_type":"new","date":"2024-12-26","first_seen":"2024-12-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.19363","pdf_url":"https://arxiv.org/pdf/2412.19363","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","联合分析","数据增强"],"reason":"用LLM生成联合分析数据并与真实人类数据对照，评估偏差并提出统计校正方法，涉及…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":63,"question":"能否通过统计方法有效整合LLM生成数据与真实数据，以改进联合分析中消费者偏好的估计精度？","design":"本研究提出一种统计增强方法，将LLM生成的联合分析选择数据作为辅助信息，与真实人类数据结合，构建AI增强估计量（AAE），并推导其渐近性质与有限样本误差界。","baseline":"真实人类数据来自两项实证研究：COVID-19疫苗偏好联合分析调查和跑车选择联合分析调查。","findings":"直接混合LLM生成数据与真实数据会加剧偏差，而所提AAE方法能显著降低估计误差，并在疫苗偏好研究中节省24.9%至79.8%的数据收集成本。LLM生成数据不能直接替代人类回答，但作为统计框架内的补充信息具有价值。","reliability":"论文指出LLM缺乏真实生活体验，且消费者偏好随时间变化，LLM可能无法准确捕捉；即使使用最先进的提示工程技术，LLM与人类回答间的差距仍持续存在，直接替代会导致误导性结果。","relevance":"高度相关：该研究将LLM作为人类被试的替代品进行联合分析实验，并与真实人类数据严格对照，评估了直接替代的偏差，并提出校正方法，直接回应了研究者对仿真可靠性、失效条件及经济学实验场景的关注。","inspiration":"该研究提出AI增强估计量（AAE），将LLM生成数据作为辅助信息而非直接替代，通过统计校正降低偏差，这种处理混杂与测量误差的思路值得借鉴｜可迁移到消费者金融产品选择偏好的联合分析中，如贷款条款偏好或保险计划选择，以降低调研成本并校正LLM仿真偏差｜以真实消费者为被试，处理为不同贷款属性组合，结果变量为选择决策，用LLM生成的选择数据作为辅助，构建AAE估计量，并与纯真实数据估计的偏好参数对照，评估成本节省与偏差校正效果"}},{"id":"2412.07031","version":4,"title":"Large Language Models: An Applied Econometric Framework","zh_title":"大语言模型：一个应用计量经济学框架","abstract":"Large language models (LLMs) enable researchers to analyze text at unprecedented scale and minimal cost. Researchers can now revisit old questions and tackle novel ones with rich data. We provide an econometric framework for realizing this potential in two empirical uses. For prediction problems -- forecasting outcomes from text -- valid conclusions require ``no training leakage'' between the LLM's training data and the researcher's sample, which can be enforced through careful model choice and research design. For estimation problems -- automating the measurement of economic concepts for downstream analysis -- valid downstream inference requires combining LLM outputs with a small validation sample to deliver consistent and precise estimates. Absent a validation sample, researchers cannot assess possible errors in LLM outputs, and consequently seemingly innocuous choices (which model, which prompt) can produce dramatically different parameter estimates. When used appropriately, LLMs are powerful tools that can expand the frontier of empirical economics.","authors":["Jens Ludwig","Sendhil Mullainathan","Ashesh Rambachan"],"categories":["econ.EM","cs.AI"],"primary_category":"econ.EM","announce_type":"new","date":"2024-12-09","first_seen":"2024-12-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.07031","pdf_url":"https://arxiv.org/pdf/2412.07031","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","计量方法","文本分析"],"reason":"LLM替代人工标注测量经济概念，属标注员替代而非仿真被试，但方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":202,"question":"如何将大语言模型（LLM）的输出纳入实证经济学研究，以保证预测和估计问题的有效推断？","design":"本文并非直接进行人类仿真实验，而是提出一个计量经济学框架，将LLM用于两类任务：预测问题（用文本预测经济结果）和估计问题（用LLM自动测量文本中的经济概念以替代人工标注）。在估计问题中，LLM扮演标注员的角色，通过少量验证样本校正其测量误差，从而放大验证样本的信息效率。","baseline":"无对照（本文未进行仿真实验，而是讨论方法框架；在估计问题中，真实测量来自人工标注或验证样本，但未提供具体人类基准数据集）。","findings":"在预测问题中，有效推断要求LLM训练数据与研究者样本无重叠（无训练泄漏），可通过模型选择和设计实现。在估计问题中，缺乏验证样本时，LLM输出误差未知，模型和提示词的微小变化会导致参数估计在量级、符号和显著性上剧烈波动；结合少量验证样本进行去偏后，LLM输出能显著提高估计精度，但不能完全替代人工数据。","reliability":"论文指出，若无验证样本，研究者无法评估LLM输出的误差及其对下游参数估计的影响，且不同LLM和提示词选择会导致截然不同的结论；预测问题中若存在训练泄漏，则评估结果反映的不是样本外表现。","relevance":"本文虽聚焦LLM作为标注工具而非人类被试替代，但其估计问题框架可直接迁移至人类仿真研究：将LLM模拟的受试者视为有误差的测量，需用少量真实人类样本校正，否则仿真结果不可靠。对关注仿真失效条件的研究者具有重要参考价值，建议阅读原文。","inspiration":"借鉴其估计问题框架：将LLM输出视为有误差的测量，需用少量真实样本校正偏差，否则模型和提示词的微小变化会导致结论剧烈波动。｜可迁移至政策公告的预期形成实验，研究不同措辞的央行沟通如何影响公众通胀预期。｜以LLM模拟公众被试，处理为不同风格的货币政策声明，结果变量为LLM生成的通胀预期数值，用真实调查数据（如密歇根消费者调查）作为基准校正仿真偏差。"}},{"id":"2411.06790","version":2,"title":"Large-scale moral machine experiment on large language models","zh_title":"基于大语言模型的大规模道德机器实验","abstract":"The rapid advancement of Large Language Models (LLMs) and their potential integration into autonomous driving systems necessitates understanding their moral decision-making capabilities. While our previous study examined four prominent LLMs using the Moral Machine experimental framework, the dynamic landscape of LLM development demands a more comprehensive analysis. Here, we evaluate moral judgments across 52 different LLMs, including multiple versions of proprietary models (GPT, Claude, Gemini) and open-source alternatives (Llama, Gemma), to assess their alignment with human moral preferences in autonomous driving scenarios. Using a conjoint analysis framework, we evaluated how closely LLM responses aligned with human preferences in ethical dilemmas and examined the effects of model size, updates, and architecture. Results showed that proprietary models and open-source models exceeding 10 billion parameters demonstrated relatively close alignment with human judgments, with a significant negative correlation between model size and distance from human judgments in open-source models. However, model updates did not consistently improve alignment with human preferences, and many LLMs showed excessive emphasis on specific ethical principles. These findings suggest that while increasing model size may naturally lead to more human-like moral judgments, practical implementation in autonomous driving systems requires careful consideration of the trade-off between judgment quality and computational efficiency. Our comprehensive analysis provides crucial insights for the ethical design of autonomous systems and highlights the importance of considering cultural contexts in AI moral decision-making.","authors":["Muhammad Shahrul Zaim bin Ahmad","Kazuhiro Takemoto"],"categories":["cs.CY","cs.CL","cs.HC"],"primary_category":"cs.CY","announce_type":"new","date":"2024-11-11","first_seen":"2024-11-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2411.06790","pdf_url":"https://arxiv.org/pdf/2411.06790","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2","C2"],"tags":["LLM道德判断","自动驾驶伦理","人类对齐"],"reason":"测量LLM的道德判断，非仿真人类被试；场景为自动驾驶，属C2排除项，但有人类数…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":201,"question":"不同大语言模型在自动驾驶道德困境中的道德判断与人类偏好的对齐程度如何？","design":"本研究并非将LLM作为人类被试的仿真，而是直接测量52个LLM（包括GPT、Claude、Gemini、Llama等）在道德机器框架下的道德选择。通过受约束随机生成包含物种、社会价值、性别、年龄、健身、功利主义等维度的两难场景，要求模型在不可避免事故中二选一，并采用联合分析评估其与人类偏好的距离。","baseline":"对照的真实人类数据来自原始道德机器实验（Moral Machine experiment）中人类参与者的道德偏好模式。","findings":"专有模型和参数超过100亿的开源模型与人类判断对齐较好，开源模型中模型规模与人类判断距离呈显著负相关。模型更新并未持续改善与人类偏好的对齐，许多LLM表现出对特定伦理原则的过度强调。","reliability":"论文承认道德机器框架存在方法局限，如场景设计约束、文化代表性不足、理论选择与实际部署间的差距，并指出实际自动驾驶系统部署需权衡判断质量与计算效率。","relevance":"高度相关：该研究使用真实人类道德偏好作为基准，系统评估了LLM在道德决策上的对齐程度，并揭示了模型规模、更新与架构的影响，直接回应了研究者对LLM仿真可靠性及失效条件的关注。","inspiration":"该研究通过受约束随机生成多维度道德困境场景，并采用联合分析评估LLM与人类偏好的距离，为多属性决策仿真提供了严谨的测量框架｜可迁移至金融伦理决策研究，如算法信贷审批中的公平性权衡（效率vs.平等）或投资顾问的ESG偏好冲突｜以LLM作为被试，随机生成包含申请人收入、种族、性别、信用分等属性的贷款审批场景，要求模型做出批准/拒绝决策，结果变量为决策中的属性权重，以真实信贷员决策数据或公平借贷审计数据作为人类基准对照"}},{"id":"2410.19599","version":3,"title":"Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina","zh_title":"谨慎使用LLM作为人类替代品：Scylla Ex Machina","abstract":"Recent studies suggest large language models (LLMs) can exhibit human-like reasoning, aligning with human behavior in economic experiments, surveys, and political discourse. This has led many to propose that LLMs can be used as surrogates or simulations for humans in social science research. However, LLMs differ fundamentally from humans, relying on probabilistic patterns, absent the embodied experiences or survival objectives that shape human cognition. We assess the reasoning depth of LLMs using the 11-20 money request game. Nearly all advanced approaches fail to replicate human behavior distributions across many models. Causes of failure are diverse and unpredictable, relating to input language, roles, and safeguarding. These results advise caution when using LLMs to study human behavior or as surrogates or simulations.","authors":["Yuan Gao","Dokyun Lee","Gordon Burtch","Sina Fazelpour"],"categories":["econ.GN","cs.AI","cs.CY","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2024-10-25","first_seen":"2024-10-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2410.19599","pdf_url":"https://arxiv.org/pdf/2410.19599","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","行为博弈","算法保真度"],"reason":"直接评估LLM作为人类替代品的可靠性，使用11-20金钱请求游戏与真实人类行为…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"大语言模型在简单经济博弈中能否可靠地复现人类行为分布，作为人类替代品用于社会科学研究？","design":"本研究使用11-20金钱请求博弈，测试了GPT-4、GPT-3.5、Claude3-Opus、Claude3-Sonnet、Llama3-70b、Llama3-8b、Llama2-13b、Llama2-7b等8个LLM，每个模型收集1000次干净会话，并尝试了提示工程、检索增强生成、微调等高级技术，比较模型选择分布与人类分布。","baseline":"人类基准来自Arad和Rubinstein (2012)原始论文中报告的人类参与者行为分布及纳什均衡预测。","findings":"几乎所有LLM和高级方法均未能复现人类行为分布，失败原因多样且不可预测，涉及输入语言、角色设定和安全防护等；即使微调GPT-4o能模仿特定人类数据，但在新变体或分布外场景下仍失效。","reliability":"论文指出LLM行为不稳定，对提示措辞、语言、角色等高度敏感；模型可能依赖记忆而非真正推理；在分布外场景下普遍失败；微调仅能模仿已知模式，无法泛化。","relevance":"该研究直接评估LLM作为人类替代品的可靠性，使用经济学实验与真实人类数据对照，并系统揭示失效条件，高度契合研究者对批判性仿真研究的关注，值得精读原文。","inspiration":"借鉴其系统对比多个LLM与真实人类行为分布的方法，并引入提示工程、微调等稳健性检验来揭示仿真失效条件｜可迁移到政策公告的预期形成实验，检验LLM能否复现人类对财政或货币政策信号的反应分布｜以GPT-4o等为被试，给予不同措辞的政策公告作为处理，测量其通胀或就业预期分布，并以真实调查数据（如密歇根消费者预期调查）为基准对照"}},{"id":"2409.19430","version":1,"title":"'Simulacrum of Stories': Examining Large Language Models as Qualitative Research Participants","zh_title":"“故事的拟像”：审视大语言模型作为定性研究参与者","abstract":"The recent excitement around generative models has sparked a wave of proposals suggesting the replacement of human participation and labor in research and development--e.g., through surveys, experiments, and interviews--with synthetic research data generated by large language models (LLMs). We conducted interviews with 19 qualitative researchers to understand their perspectives on this paradigm shift. Initially skeptical, researchers were surprised to see similar narratives emerge in the LLM-generated data when using the interview probe. However, over several conversational turns, they went on to identify fundamental limitations, such as how LLMs foreclose participants' consent and agency, produce responses lacking in palpability and contextual depth, and risk delegitimizing qualitative research methods. We argue that the use of LLMs as proxies for participants enacts the surrogate effect, raising ethical and epistemological concerns that extend beyond the technical limitations of current models to the core of whether LLMs fit within qualitative ways of knowing.","authors":["Shivani Kapania","William Agnew","Motahhare Eslami","Hoda Heidari","Sarah Fox"],"categories":["cs.HC","cs.CL","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2024-09-28","first_seen":"2024-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.19430","pdf_url":"https://arxiv.org/pdf/2409.19430","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A4","B4"],"tags":["LLM仿真","定性研究","方法论批判"],"reason":"直接研究用LLM替代定性研究参与者，并识别仿真失效条件，高度相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":27,"question":"定性研究者如何看待用大语言模型模拟研究参与者进行访谈？","design":"本研究并非仿真实验，而是对19位定性研究者进行半结构化访谈，并让他们使用一个访谈探针工具，将自己以往的人类访谈数据与LLM生成的模拟访谈数据进行对比，收集他们的看法和反思。","baseline":"研究者自己过去项目中收集的真实人类访谈记录。","findings":"研究者起初惊讶于LLM能生成与人类相似的叙述，但多轮对话后发现LLM回应缺乏切身感和情境深度，且会剥夺参与者的同意权和能动性。LLM作为参与者代理会产生“替代效应”，引发超越技术局限的伦理与认识论问题，可能损害定性研究方法的合法性。","reliability":"论文指出LLM回应缺乏 palpability（切身感）、模型的认识论立场模糊、强化研究者立场、剥夺参与者同意与能动性、抹除社区视角、以及可能使定性研究方法失去合法性。这些局限根植于LLM与诠释主义定性认识论的根本不兼容。","relevance":"该研究直接探讨用LLM替代人类进行定性访谈的仿真实践，并基于真实人类数据对比，识别出仿真失效的深层条件，高度契合研究者对LLM仿真可靠性及批判性研究的兴趣，值得精读原文。","inspiration":"借鉴该研究将LLM生成内容与真实人类记录进行对比的评估框架，可用于检验LLM仿真经济决策的效度。｜可迁移到消费者跨期选择实验，考察LLM生成的消费-储蓄决策是否与人类行为一致。｜以LLM作为被试，施加不同利率或未来收入预期的处理，测量其跨期消费分配，并与真实家庭调查数据（如PSID）中的消费-储蓄模式进行对照。"}},{"id":"2409.14202","version":3,"title":"Mining Causality: AI-Assisted Search for Instrumental Variables","zh_title":"挖掘因果关系：人工智能辅助的工具变量搜索","abstract":"The instrumental variables (IVs) method is a leading empirical strategy for causal inference. Finding IVs is a heuristic and creative process, and justifying its validity -- especially exclusion restrictions -- is largely rhetorical. We propose using large language models (LLMs) to search for new IVs through narratives and counterfactual reasoning, similar to how a human researcher would. The stark difference, however, is that LLMs can dramatically accelerate this process and explore an extremely large search space. We demonstrate how to construct prompts to search for potentially valid IVs. We contend that multi-step and role-playing prompting strategies are effective for simulating the endogenous decision-making processes of economic agents and for navigating language models through the realm of real-world scenarios, rather than anchoring them within the narrow realm of academic discourses on IVs. We apply our method to three well-known examples in economics: returns to schooling, supply and demand, and peer effects. We then extend our strategy to finding (i) control variables in regression and difference-in-differences and (ii) running variables in regression discontinuity designs.","authors":["Sukjin Han"],"categories":["econ.EM","stat.AP","stat.ME","stat.ML"],"primary_category":"econ.EM","announce_type":"new","date":"2024-09-21","first_seen":"2024-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.14202","pdf_url":"https://arxiv.org/pdf/2409.14202","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["工具变量","LLM模拟","因果推断"],"reason":"用LLM模拟经济主体决策以寻找工具变量，但无真实人类行为对照，属社会模拟边界情…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":200,"question":"如何利用大语言模型通过叙事和反事实推理来系统性地搜索经济学中的工具变量？","design":"使用GPT-4，通过多步骤和角色扮演提示策略，模拟经济主体的内生决策过程，在真实世界场景中搜索潜在工具变量，并应用于教育回报、供需和同伴效应三个经典例子。","baseline":"无对照","findings":"GPT-4在三个经典例子中均生成了候选工具变量列表，其中包含文献中常用的工具变量和一些看似新颖的变量，并提供了合理性论证；该方法还可扩展用于搜索控制变量和断点回归中的运行变量。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济主体决策以寻找工具变量，属于用AI辅助社会科学研究中的创造性过程，但缺乏真实人类行为对照，不符合你关注的以人类被试为基准的仿真实验，相关性较低，不建议优先阅读。","inspiration":"该方法利用LLM的角色扮演和反事实推理来生成经济学中的工具变量，可借鉴其多步骤提示策略来辅助识别因果效应｜可迁移到政策评估场景，例如估计某项教育政策对长期收入的因果效应时，用LLM模拟政策制定者或经济主体的决策逻辑来搜索潜在工具变量｜以LLM作为辅助工具，让GPT-4模拟不同经济主体（如家庭、学校）在政策冲击下的行为，生成候选工具变量列表，再使用真实观测数据（如面板调查数据）检验这些工具变量的相关性和外生性，并与传统工具变量（如政策规则变化）进行对比"}},{"id":"2409.10750","version":1,"title":"GPT takes the SAT: Tracing changes in Test Difficulty and Math Performance of Students","zh_title":"GPT参加SAT：追踪试题难度与学生数学表现的变化","abstract":"Scholastic Aptitude Test (SAT) is crucial for college admissions but its effectiveness and relevance are increasingly questioned. This paper enhances Synthetic Control methods by introducing \"Transformed Control\", a novel method that employs Large Language Models (LLMs) powered by Artificial Intelligence to generate control groups. We utilize OpenAI's API to generate a control group where GPT-4, or ChatGPT, takes multiple SATs annually from 2008 to 2023. This control group helps analyze shifts in SAT math difficulty over time, starting from the baseline year of 2008. Using parallel trends, we calculate the Average Difference in Scores (ADS) to assess changes in high school students' math performance. Our results indicate a significant decrease in the difficulty of the SAT math section over time, alongside a decline in students' math performance. The analysis shows a 71-point drop in the rigor of SAT math from 2008 to 2023, with student performance decreasing by 36 points, resulting in a 107-point total divergence in average student math performance. We investigate possible mechanisms for this decline in math proficiency, such as changing university selection criteria, increased screen time, grade inflation, and worsening adolescent mental health. Disparities among demographic groups show a 104-point drop for White students, 84 points for Black students, and 53 points for Asian students. Male students saw a 117-point reduction, while female students had a 100-point decrease.","authors":["Vikram Krishnaveti","Saannidhya Rawat"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2024-09-16","first_seen":"2024-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.10750","pdf_url":"https://arxiv.org/pdf/2409.10750","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","教育评估","人类数据对照"],"reason":"用GPT-4模拟学生考SAT，与真实学生成绩对照，评估试题难度变化，属于教育评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":124,"question":"SAT数学部分的难度是否随时间下降，以及学生数学表现是否随之变化？","design":"用GPT-4作为合成控制组，每年参加2008至2023年的SAT数学考试，生成基准分数；通过平行趋势和平均分数差（ADS）分析试题难度变化，并对比真实学生成绩评估学生数学表现。","baseline":"真实学生SAT数学成绩数据，包括总体及按种族、性别划分的分数。","findings":"2008至2023年SAT数学难度下降71分，学生数学表现下降36分，总差距达107分；不同族裔和性别群体表现下降程度不同，白人学生下降104分，黑人84分，亚裔53分，男性117分，女性100分。","reliability":"论文未讨论","relevance":"该研究用GPT-4模拟学生参加标准化考试，与真实学生成绩对照，评估试题难度变化，属于教育评估场景的LLM仿真实验，有真实人类基准，值得细读其合成控制方法和可靠性讨论。","inspiration":"该方法用GPT-4作为合成控制组，通过平行趋势和平均分数差分析试题难度变化，为评估政策或环境变化提供了可借鉴的因果推断框架。｜可迁移到教育经济学中评估考试政策改革（如考试内容调整）对学生表现的影响，或劳动经济学中分析技能需求变化。｜用GPT-4模拟求职者参加不同年份的职业能力测试，处理为测试年份，结果变量为测试分数，以真实求职者历史分数数据为对照，评估试题难度和技能需求变化。"}},{"id":"2409.08357","version":2,"title":"An Experimental Study of Competitive Market Behavior Through LLMs","zh_title":"通过大语言模型对竞争市场行为的实验研究","abstract":"This study explores the potential of large language models (LLMs) to conduct market experiments, aiming to understand their capability to comprehend competitive market dynamics. We model the behavior of market agents in a controlled experimental setting, assessing their ability to converge toward competitive equilibria. The results reveal the challenges current LLMs face in replicating the dynamic decision-making processes characteristic of human trading behavior. Unlike humans, LLMs lacked the capacity to achieve market equilibrium. The research demonstrates that while LLMs provide a valuable tool for scalable and reproducible market simulations, their current limitations necessitate further advancements to fully capture the complexities of market behavior. Future work that enhances dynamic learning capabilities and incorporates elements of behavioral economics could improve the effectiveness of LLMs in the economic domain, providing new insights into market dynamics and aiding in the refinement of economic policies.","authors":["Jingru Jia","Zehua Yuan"],"categories":["cs.HC","cs.AI","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2024-09-12","first_seen":"2024-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.08357","pdf_url":"https://arxiv.org/pdf/2409.08357","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2"],"tags":["LLM仿真","市场实验","行为经济学"],"reason":"用LLM模拟市场竞争行为并与人类对照，涉及经济学实验，但未明确提及真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":123,"question":"LLM能否在双拍卖实验中模拟竞争市场行为并收敛到均衡价格？","design":"使用ChatGPT-4.0扮演22个市场代理人（11买家、11卖家），在双拍卖框架下进行5轮交易，记录出价、要价和成交价，分析价格收敛、波动性和策略适应。","baseline":"对照Smith (1962)的人类双拍卖实验，人类被试价格逐渐收敛到理论均衡价$2.00。","findings":"LLM驱动的交易价格围绕均衡价$2.00波动，未呈现向均衡收敛的趋势；LLM缺乏人类交易中的动态学习和实时反馈适应能力。","reliability":"论文指出LLM的局限在于缺乏自适应学习和实时反馈机制，无法达到市场均衡，需要增强动态学习能力并融入行为经济学元素。","relevance":"该研究直接使用LLM模拟经济学实验中的市场行为，并与经典人类实验数据对照，揭示了LLM在动态决策任务中的失效，与研究者关注的经济学实验仿真和可靠性批判高度契合，值得精读。","inspiration":"借鉴其将LLM作为市场代理人参与双拍卖实验并与经典人类实验（Smith, 1962）严格对照的设计，通过比较价格收敛路径和波动性来检验仿真有效性。｜可迁移到资产定价实验，研究LLM能否在连续双向拍卖中复现人类交易者的价格发现过程与泡沫形成。｜用LLM扮演交易者参与资产市场实验，处理为不同信息结构（如对称/不对称信息），结果变量为价格偏差、交易量与泡沫持续时间，对照真实人类实验数据（如Smith et al., 1988的资产泡沫实验）。"}},{"id":"2409.00128","version":3,"title":"Can Large Language Models Replace Human Subjects? A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management","zh_title":"大语言模型能替代人类被试吗？心理学与管理学场景实验的大规模复现","abstract":"Artificial Intelligence (AI) is increasingly being integrated into scientific research, particularly in the social sciences, where understanding human behavior is critical. Large Language Models (LLMs) have shown promise in replicating human-like responses in various psychological experiments. We conducted a large-scale study replicating 156 psychological experiments from top social science journals using three state-of-the-art LLMs (GPT-4, Claude 3.5 Sonnet, and DeepSeek v3). Our results reveal that while LLMs demonstrate high replication rates for main effects (73-81%) and moderate to strong success with interaction effects (46-63%), They consistently produce larger effect sizes than human studies, with Fisher Z values approximately 2-3 times higher than human studies. Notably, LLMs show significantly lower replication rates for studies involving socially sensitive topics such as race, gender and ethics. When original studies reported null findings, LLMs produced significant results at remarkably high rates (68-83%) - while this could reflect cleaner data with less noise, as evidenced by narrower confidence intervals, it also suggests potential risks of effect size overestimation. Our results demonstrate both the promise and challenges of LLMs in psychological research, offering efficient tools for pilot testing and rapid hypothesis validation while enriching rather than replacing traditional human subject studies, yet requiring more nuanced interpretation and human validation for complex social phenomena and culturally sensitive research questions.","authors":["Ziyan Cui","Ning Li","Huaikang Zhou"],"categories":["cs.CL","cs.AI","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2024-08-29","first_seen":"2024-08-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.00128","pdf_url":"https://arxiv.org/pdf/2409.00128","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM仿真","人类被试替代","心理学实验复现"],"reason":"直接复现156项心理学实验，用LLM替代人类被试，有真实人类数据对照，评估可靠…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"大语言模型能否在多大程度上替代人类被试，复现心理学和管理学中的场景实验？","design":"使用GPT-4、Claude 3.5 Sonnet和DeepSeek v3三个LLM，将156项已发表心理学实验的原始文本材料直接呈现给模型，每个实验生成与原始人类样本量相等的模型回答，测量主效应和交互效应的复制率、效应量及p值分布。","baseline":"原始人类实验的真实数据，来自五本顶级管理学和心理学期刊的156项随机选取的场景实验。","findings":"LLM对主效应的复制率达73-81%，交互效应复制率为46-63%，但效应量普遍是人类研究的2-3倍；在涉及种族、性别等社会敏感话题时复制率显著下降，且对原研究中的零结果有68-83%的概率产生显著结果。","reliability":"LLM在涉及社会敏感话题（如种族、性别、伦理）时复制率大幅降低，可能因模型的价值对齐导致社会期望偏差；效应量系统性放大，可能增加I类错误风险；研究仅限于文本情景实验，未涉及其他实验范式。","relevance":"该研究直接以大规模真实人类实验为基准，系统评估LLM替代人类被试的可靠性与偏差，与您关注的核心问题高度吻合，值得精读原文。","inspiration":"可借鉴其大规模系统复制框架和按实验特征（如敏感话题）分层分析偏差的方法｜可迁移到信贷审批中的种族/性别歧视研究或政策信息处理实验｜用LLM模拟贷款审批员，处理含不同种族/性别线索的申请材料，测量审批决策和风险感知，以真实银行历史审批数据或审计研究结果作为对照基准。"}},{"id":"2407.04467","version":3,"title":"Are Large Language Models Strategic Decision Makers? A Study of Performance and Bias in Two-Player Non-Zero-Sum Games","zh_title":"大语言模型是战略决策者吗？双人非零和博弈中的表现与偏差研究","abstract":"Large Language Models (LLMs) have been increasingly used in real-world settings, yet their strategic decision-making abilities remain largely unexplored. To fully benefit from the potential of LLMs, it's essential to understand their ability to function in complex social scenarios. Game theory, which is already used to understand real-world interactions, provides a good framework for assessing these abilities. This work investigates the performance and merits of LLMs in canonical game-theoretic two-player non-zero-sum games, Stag Hunt and Prisoner Dilemma. Our structured evaluation of GPT-3.5, GPT-4-Turbo, GPT-4o, and Llama-3-8B shows that these models, when making decisions in these games, are affected by at least one of the following systematic biases: positional bias, payoff bias, or behavioural bias. This indicates that LLMs do not fully rely on logical reasoning when making these strategic decisions. As a result, it was found that the LLMs' performance drops when the game configuration is misaligned with the affecting biases. When misaligned, GPT-3.5, GPT-4-Turbo, GPT-4o, and Llama-3-8B show an average performance drop of 32\\%, 25\\%, 34\\%, and 29\\% respectively in Stag Hunt, and 28\\%, 16\\%, 34\\%, and 24\\% respectively in Prisoner's Dilemma. Surprisingly, GPT-4o (a top-performing LLM across standard benchmarks) suffers the most substantial performance drop, suggesting that newer models are not addressing these issues. Interestingly, we found that a commonly used method of improving the reasoning capabilities of LLMs, chain-of-thought (CoT) prompting, reduces the biases in GPT-3.5, GPT-4o, and Llama-3-8B but increases the effect of the bias in GPT-4-Turbo, indicating that CoT alone cannot fully serve as a robust solution to this problem. We perform several additional experiments, which provide further insight into these observed behaviours.","authors":["Nathan Herr","Fernando Acero","Roberta Raileanu","María Pérez-Ortiz","Zhibin Li"],"categories":["cs.AI","cs.CL","cs.GT"],"primary_category":"cs.AI","announce_type":"new","date":"2024-07-05","first_seen":"2024-07-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2407.04467","pdf_url":"https://arxiv.org/pdf/2407.04467","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B2","B4"],"tags":["LLM仿真","博弈论","决策偏差"],"reason":"用LLM模拟人类在博弈中的决策，评估偏差，涉及行为博弈场景，但未明确提及真实人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":33,"question":"LLM在经典两人非零和博弈中是否存在系统性偏差，这些偏差如何影响其策略决策表现？","design":"以GPT-3.5、GPT-4-Turbo、GPT-4o和Llama-3-8B作为被试，通过改变博弈矩阵中行动标签的顺序（位置偏差）、收益结构（收益偏差）或行为倾向（行为偏差）来构造不同配置的猎鹿博弈和囚徒困境，测量模型选择合作或背叛等行动的准确率变化。","baseline":"无对照","findings":"所有模型均受至少一种系统性偏差影响，当博弈配置与偏差不一致时，模型表现平均下降16%-34%，其中GPT-4o下降最严重。思维链提示能减少部分模型的偏差，但对GPT-4-Turbo反而加剧偏差，表明其并非稳健解决方案。","reliability":"论文指出思维链提示无法完全消除偏差，且未在更复杂的多轮博弈或真实交互场景中验证，也未与人类行为直接对比。","relevance":"该研究直接评估LLM在策略互动中的决策偏差，虽未使用真实人类基准，但揭示了LLM作为人类仿真代理在博弈场景中的系统性失效模式，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"可借鉴其通过系统操纵博弈矩阵标签顺序和收益结构来检测位置偏差与收益偏差的方法，用于评估LLM在策略环境中的稳健性。｜可迁移至经济政策博弈模拟，如碳税谈判或贸易协定中的策略行为仿真。｜以LLM作为多国谈判代表，随机化提案顺序和收益矩阵，测量合作率，并与人类实验数据（如公开的博弈实验数据集）进行对照。"}},{"id":"2407.12032","version":1,"title":"Large Language Models for Behavioral Economics: Internal Validity and Elicitation of Mental Models","zh_title":"大语言模型用于行为经济学：内部效度与心智模型的引出","abstract":"In this article, we explore the transformative potential of integrating generative AI, particularly Large Language Models (LLMs), into behavioral and experimental economics to enhance internal validity. By leveraging AI tools, researchers can improve adherence to key exclusion restrictions and in particular ensure the internal validity measures of mental models, which often require human intervention in the incentive mechanism. We present a case study demonstrating how LLMs can enhance experimental design, participant engagement, and the validity of measuring mental models.","authors":["Brian Jabarian"],"categories":["cs.HC","cs.AI","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2024-06-30","first_seen":"2024-06-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2407.12032","pdf_url":"https://arxiv.org/pdf/2407.12032","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A2","B2"],"tags":["LLM仿真","行为经济学","内部效度"],"reason":"用LLM增强行为经济学实验的内部效度，涉及人类被试替代和测量效度，但侧重方法改…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":122,"question":"如何利用大语言模型增强行为与实验经济学的内部效度，特别是改善心理模型的测量效度与排除限制的遵守？","design":"本文并非直接以LLM替代人类被试的仿真研究，而是提出利用LLM优化实验设计、提升参与者参与度、确保激励相容，并监测在线实验中的不良行为，从而增强内部效度。文中包含一个案例，展示如何用LLM创建引人入胜的叙事环境，以激励方式测量通常难以追踪的思维风格。","baseline":"无对照","findings":"LLM可通过改善排除限制的遵守来增强实验的内部效度；利用LLM生成的合成行为数据为探索决策过程提供了新的方法论深度。","reliability":"论文未讨论","relevance":"本文侧重于用LLM改进实验方法而非直接仿真人类行为，但涉及LLM在行为经济学实验中的应用与效度问题，对关注仿真可靠性的研究者有一定参考价值，但缺乏人类基准对照，建议略读。","inspiration":"该方法利用LLM生成引人入胜的叙事环境来测量心理模型，可借鉴其通过叙事干预提升测量效度的设计思路｜可迁移到消费者跨期选择实验中，改善对时间偏好和思维风格的测量｜设计雏形：以LLM生成个性化金融决策叙事作为处理，人类被试在叙事前后完成跨期选择任务，结果变量为贴现率与思维风格量表得分，以传统无叙事条件下的行为数据作为对照基准"}},{"id":"2406.19317","version":2,"title":"Jump Starting Bandits with LLM-Generated Prior Knowledge","zh_title":"用LLM生成的先验知识启动Bandit算法","abstract":"We present substantial evidence demonstrating the benefits of integrating Large Language Models (LLMs) with a Contextual Multi-Armed Bandit framework. Contextual bandits have been widely used in recommendation systems to generate personalized suggestions based on user-specific contexts. We show that LLMs, pre-trained on extensive corpora rich in human knowledge and preferences, can simulate human behaviours well enough to jump-start contextual multi-armed bandits to reduce online learning regret. We propose an initialization algorithm for contextual bandits by prompting LLMs to produce a pre-training dataset of approximate human preferences for the bandit. This significantly reduces online learning regret and data-gathering costs for training such models. Our approach is validated empirically through two sets of experiments with different bandit setups: one which utilizes LLMs to serve as an oracle and a real-world experiment utilizing data from a conjoint survey experiment.","authors":["Parand A. Alamdari","Yanshuai Cao","Kevin H. Wilson"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"new","date":"2024-06-27","first_seen":"2024-06-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.19317","pdf_url":"https://arxiv.org/pdf/2406.19317","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","人类偏好模拟","上下文Bandit"],"reason":"用LLM模拟人类偏好以初始化bandit，有真实联合调查数据对照，涉及推荐系统…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":121,"question":"如何利用大语言模型生成的先验知识来初始化上下文多臂老虎机，以减少在线学习的遗憾？","design":"提出CBLI框架，用LLM根据用户特征生成大量合成用户交互与偏好数据，预训练上下文老虎机模型（LinUCB），再在真实用户交互中微调；实验一用LLM模拟用户对慈善捐赠营销风格的偏好，实验二用真实联合调查数据测试睡眠老虎机。","baseline":"实验二使用真实联合调查实验数据中的人类偏好作为对照基准。","findings":"LLM生成的近似人类偏好虽不完全匹配真实分布，但能提供优于随机冷启动的初始化，在两个实验中分别使早期遗憾降低14-17%和19-20%；即使隐藏部分隐私敏感属性，仍能降低14.8%的早期遗憾。","reliability":"论文承认LLM生成的奖励分布可能无法完美匹配真实人类偏好，但强调其仍优于随机基线；未系统讨论LLM模拟在何种条件下会失效。","relevance":"该研究用LLM模拟人类偏好来初始化推荐系统，有真实联合调查数据作为对照，直接涉及LLM作为人类被试替代品的仿真可靠性与偏差问题，值得精读。","inspiration":"该方法通过LLM生成合成用户偏好数据来预训练推荐模型，再用少量真实交互微调，可作为冷启动问题的低成本解决方案｜可迁移到消费者金融产品推荐或个性化定价实验，例如用LLM模拟不同风险偏好人群对贷款产品的选择｜以LLM生成大量合成消费者对贷款条款的偏好数据，预训练一个上下文老虎机模型，再在真实信贷选择实验数据上微调，以真实人类选择作为对照，评估早期推荐准确度与遗憾值"}},{"id":"2406.14508","version":1,"title":"Evidence of a log scaling law for political persuasion with large language models","zh_title":"大语言模型政治说服力的对数缩放定律证据","abstract":"Large language models can now generate political messages as persuasive as those written by humans, raising concerns about how far this persuasiveness may continue to increase with model size. Here, we generate 720 persuasive messages on 10 U.S. political issues from 24 language models spanning several orders of magnitude in size. We then deploy these messages in a large-scale randomized survey experiment (N = 25,982) to estimate the persuasive capability of each model. Our findings are twofold. First, we find evidence of a log scaling law: model persuasiveness is characterized by sharply diminishing returns, such that current frontier models are barely more persuasive than models smaller in size by an order of magnitude or more. Second, mere task completion (coherence, staying on topic) appears to account for larger models' persuasive advantage. These findings suggest that further scaling model size will not much increase the persuasiveness of static LLM-generated messages.","authors":["Kobi Hackenburg","Ben M. Tappin","Paul Röttger","Scott Hale","Jonathan Bright","Helen Margetts"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2024-06-20","first_seen":"2024-06-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.14508","pdf_url":"https://arxiv.org/pdf/2406.14508","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","政治说服","人类数据对照"],"reason":"用LLM生成政治说服信息，通过大规模随机调查实验与人类数据对照，评估模型说服力…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":31,"question":"大语言模型的政治说服力是否随模型规模扩大而持续提升？","design":"用24个不同规模的语言模型生成720条政治说服信息，通过大规模随机调查实验（N=25,982）将美国成年人随机分配到AI组、人类组或对照组，测量其对10个政策议题的态度变化。","baseline":"人类撰写的说服信息以及未接受任何信息的对照组。","findings":"模型说服力随规模呈对数缩放，当前前沿模型仅比小一个数量级的模型略具说服力；说服力提升主要源于任务完成度（连贯性、切题），前沿模型在该指标上已接近上限。","reliability":"论文未讨论","relevance":"该研究以真实人类调查数据为基准，评估LLM在政治说服场景中的仿真效果，并揭示规模扩展的边际收益递减，直接回应了研究者对仿真可靠性与失效条件的关切，值得精读。","inspiration":"该方法将LLM生成内容作为处理，通过大规模随机调查实验与人类生成内容及对照组比较，测量态度变化，可借鉴其处理施加与基准对照设计。｜可迁移至政策公告的预期形成研究，如评估AI生成的经济新闻对公众通胀预期的影响。｜以LLM生成不同风格的经济新闻为处理，招募代表性样本为被试，测量其通胀预期变化，并以人类撰写新闻和真实历史数据为对照。"}},{"id":"2406.13605","version":2,"title":"Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?","zh_title":"比人类更友善：大语言模型在囚徒困境中的行为研究","abstract":"The behavior of Large Language Models (LLMs) as artificial social agents is largely unexplored, and we still lack extensive evidence of how these agents react to simple social stimuli. Testing the behavior of AI agents in classic Game Theory experiments provides a promising theoretical framework for evaluating the norms and values of these agents in archetypal social situations. In this work, we investigate the cooperative behavior of three LLMs (Llama2, Llama3, and GPT3.5) when playing the Iterated Prisoner's Dilemma against random adversaries displaying various levels of hostility. We introduce a systematic methodology to evaluate an LLM's comprehension of the game rules and its capability to parse historical gameplay logs for decision-making. We conducted simulations of games lasting for 100 rounds and analyzed the LLMs' decisions in terms of dimensions defined in the behavioral economics literature. We find that all models tend not to initiate defection but act cautiously, favoring cooperation over defection only when the opponent's defection rate is low. Overall, LLMs behave at least as cooperatively as the typical human player, although our results indicate some substantial differences among models. In particular, Llama2 and GPT3.5 are more cooperative than humans, and especially forgiving and non-retaliatory for opponent defection rates below 30%. More similar to humans, Llama3 exhibits consistently uncooperative and exploitative behavior unless the opponent always cooperates. Our systematic approach to the study of LLMs in game theoretical scenarios is a step towards using these simulations to inform practices of LLM auditing and alignment.","authors":["Nicoló Fontana","Francesco Pierri","Luca Maria Aiello"],"categories":["cs.CY","cs.AI","cs.GT","physics.soc-ph"],"primary_category":"cs.CY","announce_type":"new","date":"2024-06-19","first_seen":"2024-06-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.13605","pdf_url":"https://arxiv.org/pdf/2406.13605","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","囚徒困境","行为博弈"],"reason":"用LLM玩囚徒困境并与人类数据对照，直接仿真人类决策行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":26,"question":"大型语言模型在迭代囚徒困境中面对不同敌对程度的对手时，其合作行为如何？","design":"使用Llama2、Llama3和GPT3.5三个LLM作为被试，与具有不同背叛概率的随机对手进行100轮迭代囚徒困境博弈，通过系统提示评估模型对规则的理解和历史记录解析能力，分析其合作决策。","baseline":"对照已有行为经济学文献中报告的人类玩家在囚徒困境中的典型合作行为。","findings":"所有模型倾向于不首先背叛，但仅在对手背叛率低时更合作；Llama2和GPT3.5比人类更合作、更宽容，而Llama3更接近人类，表现出不合作和剥削性行为。","reliability":"论文未讨论","relevance":"该研究直接用LLM复现经典博弈实验并与人类数据对照，属于经济学实验场景下的人类仿真，且分析了模型间差异，值得精读以评估仿真可靠性与偏差。","inspiration":"该研究通过系统提示解析历史记录并控制对手背叛概率来测量LLM的合作行为，方法上可借鉴其精确操纵对手策略以分离LLM反应模式的做法｜可迁移到资产定价实验中的信任博弈或投资者情绪传染研究，考察LLM在不同市场信息环境下的策略调整｜可设计让LLM作为投资者与不同诚实度的基金经理进行重复信任博弈，处理变量为基金经理的欺骗概率，结果变量为投资额，对照真实人类实验数据"}},{"id":"2406.13558","version":2,"title":"Enhancing Travel Choice Modeling with Large Language Models: A Prompt-Learning Approach","zh_title":"利用大语言模型增强出行选择建模：一种提示学习方法","abstract":"Travel choice analysis is crucial for understanding individual travel behavior to develop appropriate transport policies and recommendation systems in Intelligent Transportation Systems (ITS). Despite extensive research, this domain faces two critical challenges: a) modeling with limited survey data, and b) simultaneously achieving high model explainability and accuracy. In this paper, we introduce a novel prompt-learning-based Large Language Model(LLM) framework that significantly improves prediction accuracy and provides explicit explanations for individual predictions. This framework involves three main steps: transforming input variables into textual form; building of demonstrations similar to the object, and applying these to a well-trained LLM. We tested the framework's efficacy using two widely used choice datasets: London Passenger Mode Choice (LPMC) and Optima-Mode collected in Switzerland. The results indicate that the LLM significantly outperforms state-of-the-art deep learning methods and discrete choice models in predicting people's choices. Additionally, we present a case of explanation illustrating how the LLM framework generates understandable and explicit explanations at the individual level.","authors":["Xuehao Zhai","Hanlin Tian","Lintong Li","Tianyu Zhao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2024-06-19","first_seen":"2024-06-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.13558","pdf_url":"https://arxiv.org/pdf/2406.13558","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["出行行为预测","LLM仿真","选择建模"],"reason":"用LLM预测出行选择，有真实人类数据对照，属于行为仿真，但侧重预测精度而非复现…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":127,"question":"如何利用提示学习的大语言模型框架提高出行方式选择的预测精度并生成个体层面的可解释性？","design":"本研究并非严格的人类仿真实验，而是提出一种基于提示学习的LLM框架，将出行选择数据转换为文本，构建相似演示示例，输入预训练LLM进行出行方式预测，并输出个体解释。","baseline":"使用两个真实出行选择数据集：伦敦乘客方式选择（LPMC）和瑞士Optima-Mode调查数据。","findings":"LLM框架在预测精度上显著优于现有深度学习和离散选择模型；同时能够为个体预测生成明确、可理解的解释。","reliability":"论文未讨论","relevance":"该研究用LLM预测真实出行选择行为，有真实人类数据对照，属于行为预测仿真，但侧重预测精度而非复现人类决策偏差或分布，与研究者关注的仿真可靠性和失效条件关联较弱，可酌情略读。","inspiration":"该方法将个体选择数据转换为文本并构建相似示例进行提示学习，可借鉴其将结构化行为数据文本化并利用上下文示例提升预测的思路｜可迁移至消费者金融产品选择预测，如信贷产品选择或投资组合偏好分析｜以真实银行客户交易数据为对照，将客户特征与历史选择转为文本提示，用LLM预测其贷款类型选择，比较预测准确率与离散选择模型"}},{"id":"2406.11426","version":1,"title":"Can AI with High Reasoning Ability Replicate Human-like Decision Making in Economic Experiments?","zh_title":"高推理能力AI能否复制经济实验中的人类决策？","abstract":"Economic experiments offer a controlled setting for researchers to observe human decision-making and test diverse theories and hypotheses; however, substantial costs and efforts are incurred to gather many individuals as experimental participants. To address this, with the development of large language models (LLMs), some researchers have recently attempted to develop simulated economic experiments using LLMs-driven agents, called generative agents. If generative agents can replicate human-like decision-making in economic experiments, the cost problem of economic experiments can be alleviated. However, such a simulation framework has not been yet established. Considering the previous research and the current evolutionary stage of LLMs, this study focuses on the reasoning ability of generative agents as a key factor toward establishing a framework for such a new methodology. A multi-agent simulation, designed to improve the reasoning ability of generative agents through prompting methods, was developed to reproduce the result of an actual economic experiment on the ultimatum game. The results demonstrated that the higher the reasoning ability of the agents, the closer the results were to the theoretical solution than to the real experimental result. The results also suggest that setting the personas of the generative agents may be important for reproducing the results of real economic experiments. These findings are valuable for the future definition of a framework for replacing human participants with generative agents in economic experiments when LLMs are further developed.","authors":["Ayato Kitadai","Sinndy Dayana Rico Lugo","Yudai Tsurusaki","Yusuke Fukasawa","Nariaki Nishino"],"categories":["cs.GT","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2024-06-17","first_seen":"2024-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.11426","pdf_url":"https://arxiv.org/pdf/2406.11426","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","经济实验","人类行为对照"],"reason":"用LLM代理复现最后通牒博弈实验，并与真实人类数据对照，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":25,"question":"提高生成式智能体的推理能力能否使其在最后通牒博弈经济实验中复现人类决策？","design":"使用GPT-3.5-turbo和GPT-4等LLM驱动的生成式智能体进行多智能体仿真，通过零样本、少样本和思维链提示方法操纵推理能力，模拟最后通牒博弈中的提议者和响应者决策，测量分配金额和接受/拒绝行为。","baseline":"对照Lin et al. (2020)的真实人类最后通牒博弈实验数据。","findings":"智能体推理能力越高，其结果越接近理论均衡而非真实人类行为；设置智能体的人格特征可能对复现真实实验结果很重要。","reliability":"论文指出当前仿真框架尚未建立，推理能力提升反而偏离人类行为，且提示语言、模型版本和参数设置影响结果，未来需更高推理能力LLM和更完善的人格设定。","relevance":"该研究直接以真实人类实验为基准，检验LLM代理在经济博弈中的仿真效度，并揭示了推理能力增强反而导致行为偏离人类的关键失效条件，与您关注的经济学实验仿真和可靠性批判高度契合，值得精读。","inspiration":"借鉴其通过提示方法（零样本、少样本、思维链）系统操纵LLM推理能力，并与真实人类实验数据严格对照的仿真效度检验框架｜可迁移到资产定价实验，检验LLM代理能否复现人类在泡沫形成与破裂中的非理性交易行为｜以LLM为被试，通过不同推理提示形成处理组，模拟连续竞价市场中的买卖决策，结果变量为价格偏离基础价值的程度，对照Smith et al. (1988)的实验室资产市场泡沫数据"}},{"id":"2406.05972","version":2,"title":"Decision-Making Behavior Evaluation Framework for LLMs under Uncertain Context","zh_title":"不确定情境下大语言模型决策行为评估框架","abstract":"When making decisions under uncertainty, individuals often deviate from rational behavior, which can be evaluated across three dimensions: risk preference, probability weighting, and loss aversion. Given the widespread use of large language models (LLMs) in decision-making processes, it is crucial to assess whether their behavior aligns with human norms and ethical expectations or exhibits potential biases. Several empirical studies have investigated the rationality and social behavior performance of LLMs, yet their internal decision-making tendencies and capabilities remain inadequately understood. This paper proposes a framework, grounded in behavioral economics, to evaluate the decision-making behaviors of LLMs. Through a multiple-choice-list experiment, we estimate the degree of risk preference, probability weighting, and loss aversion in a context-free setting for three commercial LLMs: ChatGPT-4.0-Turbo, Claude-3-Opus, and Gemini-1.0-pro. Our results reveal that LLMs generally exhibit patterns similar to humans, such as risk aversion and loss aversion, with a tendency to overweight small probabilities. However, there are significant variations in the degree to which these behaviors are expressed across different LLMs. We also explore their behavior when embedded with socio-demographic features, uncovering significant disparities. For instance, when modeled with attributes of sexual minority groups or physical disabilities, Claude-3-Opus displays increased risk aversion, leading to more conservative choices. These findings underscore the need for careful consideration of the ethical implications and potential biases in deploying LLMs in decision-making scenarios. Therefore, this study advocates for developing standards and guidelines to ensure that LLMs operate within ethical boundaries while enhancing their utility in complex decision-making environments.","authors":["Jingru Jia","Zehua Yuan","Junhao Pan","Paul E. McNamara","Deming Chen"],"categories":["cs.AI","cs.CY","cs.HC","cs.LG","econ.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2024-06-10","first_seen":"2024-06-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.05972","pdf_url":"https://arxiv.org/pdf/2406.05972","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A2","B2","B4"],"tags":["LLM决策行为","行为经济学","人类对照"],"reason":"用行为经济学实验评估LLM决策行为，与人类规范对照，涉及风险偏好等，可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":120,"question":"在不确定情境下，大型语言模型（LLM）的决策行为在风险偏好、概率加权和损失厌恶三个维度上是否与人类相似，以及嵌入社会人口特征后是否会产生偏差？","design":"使用基于行为经济学的多组选择列表实验，在无上下文情境下测量ChatGPT-4.0-Turbo、Claude-3-Opus和Gemini-1.0-pro三个商业LLM的风险偏好、概率加权和损失厌恶参数；进一步嵌入社会人口特征（如性少数群体或身体残疾属性）观察其决策变化。","baseline":"无对照","findings":"LLM普遍表现出与人类相似的风险厌恶、损失厌恶和对小概率的高估倾向，但不同模型间行为程度差异显著；嵌入社会人口特征后，部分模型（如Claude-3-Opus）在性少数或残疾属性下风险厌恶增强，决策更保守，显示出潜在偏差。","reliability":"论文未讨论","relevance":"该研究将LLM作为人类被试的替代，用行为经济学实验评估其决策模式，并考察社会人口特征带来的偏差，直接回应了LLM仿真人类决策的可靠性与伦理问题，值得精读以了解其框架和发现。","inspiration":"该方法借鉴了行为经济学经典的选择列表实验来测量LLM的风险偏好、概率加权和损失厌恶，并嵌入社会人口特征作为处理变量，观察决策偏差。｜可迁移到信贷审批歧视研究，检验LLM在贷款决策中是否因申请人性别、种族等特征产生系统性偏差。｜以LLM为被试，处理为嵌入不同社会人口特征的贷款申请人档案，结果变量为审批决策和风险偏好参数，对照真实银行信贷数据中的审批率与偏差模式。"}},{"id":"2406.03299","version":1,"title":"The Good, the Bad, and the Hulk-like GPT: Analyzing Emotional Decisions of Large Language Models in Cooperation and Bargaining Games","zh_title":"好、坏与浩克般的GPT：分析大语言模型在合作与讨价还价博弈中的情绪决策","abstract":"Behavior study experiments are an important part of society modeling and understanding human interactions. In practice, many behavioral experiments encounter challenges related to internal and external validity, reproducibility, and social bias due to the complexity of social interactions and cooperation in human user studies. Recent advances in Large Language Models (LLMs) have provided researchers with a new promising tool for the simulation of human behavior. However, existing LLM-based simulations operate under the unproven hypothesis that LLM agents behave similarly to humans as well as ignore a crucial factor in human decision-making: emotions. In this paper, we introduce a novel methodology and the framework to study both, the decision-making of LLMs and their alignment with human behavior under emotional states. Experiments with GPT-3.5 and GPT-4 on four games from two different classes of behavioral game theory showed that emotions profoundly impact the performance of LLMs, leading to the development of more optimal strategies. While there is a strong alignment between the behavioral responses of GPT-3.5 and human participants, particularly evident in bargaining games, GPT-4 exhibits consistent behavior, ignoring induced emotions for rationality decisions. Surprisingly, emotional prompting, particularly with `anger' emotion, can disrupt the \"superhuman\" alignment of GPT-4, resembling human emotional responses.","authors":["Mikhail Mozikov","Nikita Severin","Valeria Bodishtianu","Maria Glushanina","Mikhail Baklashkin","Andrey V. Savchenko","Ilya Makarov"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2024-06-05","first_seen":"2024-06-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.03299","pdf_url":"https://arxiv.org/pdf/2406.03299","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","行为博弈","情绪决策"],"reason":"用LLM仿真人类在博弈中的情绪决策，并与真实人类数据对照，评估对齐与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":24,"question":"情绪如何影响大语言模型在合作与讨价还价博弈中的决策，以及其行为与人类行为的对齐程度如何？","design":"使用GPT-3.5和GPT-4作为被试，通过情绪提示（愤怒、悲伤、快乐、厌恶、恐惧）注入五种基本情绪，在最后通牒博弈、独裁者博弈、囚徒困境和性别战四类博弈中测量出价份额、接受率、合作率及最大收益百分比等结果变量。","baseline":"以真实人类参与者在相同博弈实验中的行为数据作为对照基准。","findings":"情绪显著影响LLM的策略表现，GPT-3.5在讨价还价博弈中与人类行为高度对齐，而GPT-4通常忽略情绪保持理性；但愤怒情绪提示能扰乱GPT-4的“超人类”对齐，使其表现出类似人类的情绪反应。","reliability":"论文未讨论","relevance":"该研究直接以真实人类数据为基准，检验LLM在情绪影响下的行为对齐与失效条件，涵盖经济学博弈场景，并揭示了GPT-4在愤怒情绪下的对齐崩溃，高度契合研究者对仿真可靠性及批判性条件的关注，值得精读原文。","inspiration":"该方法通过情绪提示词注入情绪状态，并设置无情绪中性基线，可借鉴用于经济决策实验中情绪处理的标准化设计｜可迁移到资产定价实验中，研究情绪如何影响投资者对风险资产的需求与定价偏差｜以LLM为被试，通过愤怒/恐惧等情绪提示词处理，测量其在模拟股票交易中的出价与风险偏好，结果与真实投资者实验数据对照"}},{"id":"2405.19313","version":2,"title":"Language Models Trained to do Arithmetic Predict Human Risky and Intertemporal Choice","zh_title":"训练做算术的语言模型预测人类风险与跨期选择","abstract":"The observed similarities in the behavior of humans and Large Language Models (LLMs) have prompted researchers to consider the potential of using LLMs as models of human cognition. However, several significant challenges must be addressed before LLMs can be legitimately regarded as cognitive models. For instance, LLMs are trained on far more data than humans typically encounter, and may have been directly trained on human data in specific cognitive tasks or aligned with human preferences. Consequently, the origins of these behavioral similarities are not well understood. In this paper, we propose a novel way to enhance the utility of LLMs as cognitive models. This approach involves (i) leveraging computationally equivalent tasks that both an LLM and a rational agent need to master for solving a cognitive problem and (ii) examining the specific task distributions required for an LLM to exhibit human-like behaviors. We apply this approach to decision-making -- specifically risky and intertemporal choice -- where the key computationally equivalent task is the arithmetic of expected value calculations. We show that an LLM pretrained on an ecologically valid arithmetic dataset, which we call Arithmetic-GPT, predicts human behavior better than many traditional cognitive models. Pretraining LLMs on ecologically valid arithmetic datasets is sufficient to produce a strong correspondence between these models and human decision-making. Our results also suggest that LLMs used as cognitive models should be carefully investigated via ablation studies of the pretraining data.","authors":["Jian-Qiao Zhu","Haijiang Yan","Thomas L. Griffiths"],"categories":["cs.AI","cs.CL","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2024-05-29","first_seen":"2024-05-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2405.19313","pdf_url":"https://arxiv.org/pdf/2405.19313","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","决策行为","认知模型"],"reason":"用LLM预测人类风险与跨期选择，与真实人类数据对照，并分析仿真有效条件，直接相…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":61,"question":"语言模型在算术任务上的预训练能否使其产生与人类相似的风险和跨期选择行为？","design":"训练一个小型语言模型（约10M参数）在合成算术数据集上（如期望值计算），提取其嵌入向量，用逻辑回归预测人类选择概率，并与传统认知模型及LLaMA3比较。","baseline":"使用真实人类在风险和跨期选择任务中的选择数据作为基准。","findings":"在生态有效的算术数据集上预训练的Arithmetic-GPT模型预测人类选择优于许多传统认知模型；仅靠算术预训练就足以产生与人类决策的强对应关系。","reliability":"论文指出需通过消融实验仔细检查预训练数据，且合成数据分布需符合生态分布才能有效，否则预测力有限。","relevance":"直接以真实人类数据为基准，用LLM仿真风险与跨期选择，并分析预训练数据分布对仿真有效性的影响，高度契合研究者对仿真可靠性及失效条件的关注。","inspiration":"该方法通过控制预训练数据的生态分布来提升模型对人类决策的预测力，值得借鉴其消融实验设计以检验数据特征对仿真效果的影响｜可迁移到消费者跨期选择研究，例如分析不同利率环境下个体的储蓄与消费决策｜以LLM为被试，处理变量为预训练数据中利率变动的分布（符合真实市场波动），结果变量为模拟的跨期选择偏好，用家庭金融调查的真实跨期选择数据作为对照基准"}},{"id":"2404.01332","version":3,"title":"Explaining Large Language Models Decisions Using Shapley Values","zh_title":"使用Shapley值解释大语言模型决策","abstract":"The emergence of large language models (LLMs) has opened up exciting possibilities for simulating human behavior and cognitive processes, with potential applications in various domains, including marketing research and consumer behavior analysis. However, the validity of utilizing LLMs as stand-ins for human subjects remains uncertain due to glaring divergences that suggest fundamentally different underlying processes at play and the sensitivity of LLM responses to prompt variations. This paper presents a novel approach based on Shapley values from cooperative game theory to interpret LLM behavior and quantify the relative contribution of each prompt component to the model's output. Through two applications - a discrete choice experiment and an investigation of cognitive biases - we demonstrate how the Shapley value method can uncover what we term \"token noise\" effects, a phenomenon where LLM decisions are disproportionately influenced by tokens providing minimal informative content. This phenomenon raises concerns about the robustness and generalizability of insights obtained from LLMs in the context of human behavior simulation. Our model-agnostic approach extends its utility to proprietary LLMs, providing a valuable tool for practitioners and researchers to strategically optimize prompts and mitigate apparent cognitive biases. Our findings underscore the need for a more nuanced understanding of the factors driving LLM responses before relying on them as substitutes for human subjects in survey settings. We emphasize the importance of researchers reporting results conditioned on specific prompt templates and exercising caution when drawing parallels between human behavior and LLMs.","authors":["Behnam Mohammadi"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2024-03-29","first_seen":"2024-03-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2404.01332","pdf_url":"https://arxiv.org/pdf/2404.01332","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","Shapley值","认知偏差"],"reason":"用LLM仿真人类选择与认知偏差，并与真实人类数据对照，揭示仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":32,"question":"如何利用Shapley值解释大语言模型决策，并揭示提示词中无信息量token对模型输出的不成比例影响（即“token噪声”效应）？","design":"本研究并非直接用LLM仿真人类，而是提出一种基于合作博弈Shapley值的模型无关解释方法，将提示词各组成部分视为“玩家”，量化其对LLM输出的相对贡献。通过两个应用展示：离散选择实验（航班选择）和认知偏差调查，分析提示词中token的影响。","baseline":"无对照","findings":"发现“token噪声”现象：LLM决策受到无信息量token（如冠词、介词、甚至“flight”等单词）的过度影响，且对格式变化（如换行符）高度敏感，导致选择概率出现混沌波动。这引发了对LLM仿真人类行为稳健性和泛化性的严重担忧。","reliability":"论文指出LLM对提示词变化高度敏感，无信息量token可显著改变输出，因此基于LLM的人类行为仿真在调查环境中不可靠，需谨慎解读，并建议研究者报告基于特定提示模板的条件结果。","relevance":"该研究直接批判了用LLM替代人类被试的可靠性，揭示了仿真失效的具体机制（token噪声），与研究者关注的仿真失效条件高度相关，值得精读以深入理解LLM行为偏差的来源。","inspiration":"可借鉴Shapley值分解方法，量化提示词各成分对LLM输出的影响，用于诊断和优化经济金融实验中的提示设计。｜可迁移到消费者金融决策仿真（如贷款选择、投资偏好），分析提示词中无关信息如何扭曲LLM的“偏好”。｜以GPT-4为被试，设计不同贷款方案的离散选择实验，在提示中系统变化无信息量token（如换行符、冠词），用Shapley值量化其影响，并以真实消费者信贷选择数据为基准，检验LLM仿真偏差。"}},{"id":"2403.15281","version":1,"title":"Measuring Gender and Racial Biases in Large Language Models","zh_title":"测量大语言模型中的性别与种族偏见","abstract":"In traditional decision making processes, social biases of human decision makers can lead to unequal economic outcomes for underrepresented social groups, such as women, racial or ethnic minorities. Recently, the increasing popularity of Large language model based artificial intelligence suggests a potential transition from human to AI based decision making. How would this impact the distributional outcomes across social groups? Here we investigate the gender and racial biases of OpenAIs GPT, a widely used LLM, in a high stakes decision making setting, specifically assessing entry level job candidates from diverse social groups. Instructing GPT to score approximately 361000 resumes with randomized social identities, we find that the LLM awards higher assessment scores for female candidates with similar work experience, education, and skills, while lower scores for black male candidates with comparable qualifications. These biases may result in a 1 or 2 percentage point difference in hiring probabilities for otherwise similar candidates at a certain threshold and are consistent across various job positions and subsamples. Meanwhile, we also find stronger pro female and weaker anti black male patterns in democratic states. Our results demonstrate that this LLM based AI system has the potential to mitigate the gender bias, but it may not necessarily cure the racial bias. Further research is needed to comprehend the root causes of these outcomes and develop strategies to minimize the remaining biases in AI systems. As AI based decision making tools are increasingly employed across diverse domains, our findings underscore the necessity of understanding and addressing the potential unequal outcomes to ensure equitable outcomes across social groups.","authors":["Jiafu An","Difang Huang","Chen Lin","Mingzhu Tai"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2024-03-22","first_seen":"2024-03-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2403.15281","pdf_url":"https://arxiv.org/pdf/2403.15281","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","偏见测量","劳动力市场"],"reason":"用LLM替代人类决策者评估简历，测量偏见并与真实人类数据对照，涉及劳动力市场政…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:47","error":null,"has_summary":true,"summary":{"generated_at":"2024-03-22","rank":1,"question":"在招聘决策中，GPT-3.5对不同性别和种族的求职者是否存在评分偏差？","design":"使用GPT-3.5模型扮演招聘决策者，对约36.1万份随机生成的工作经验、教育背景和技能组合的虚构简历进行评分（0-100分），简历随机分配带有性别和种族标识的姓名，以测量模型对不同社会群体的评分差异。","baseline":"无对照","findings":"GPT-3.5对女性求职者（无论种族）给予显著高于白人男性的评分，但对黑人男性给予显著低于白人男性的评分；在民主党州，亲女性偏差更强，反黑人男性偏差较弱。","reliability":"论文未讨论","relevance":"该研究直接使用LLM替代人类决策者进行高利害决策实验，测量了性别和种族偏见，虽未提供真实人类对照数据，但为评估LLM在劳动力市场决策中的偏差提供了重要证据，值得精读以了解仿真设计细节和偏差模式。","inspiration":"借鉴其通过大规模随机生成简历并操控姓名标识社会身份的实验设计，可精确分离LLM的偏见效应｜可迁移到信贷审批歧视研究，用LLM模拟信贷员评估贷款申请，操控申请人性别/种族｜用GPT-4扮演信贷员，对随机生成并分配不同种族姓名的贷款申请进行评分，结果变量为贷款批准概率，以真实银行信贷数据中的种族差异作为对照基准。"}},{"id":"2403.05534","version":1,"title":"Bayesian Preference Elicitation with Language Models","zh_title":"基于语言模型的贝叶斯偏好诱导","abstract":"Aligning AI systems to users' interests requires understanding and incorporating humans' complex values and preferences. Recently, language models (LMs) have been used to gather information about the preferences of human users. This preference data can be used to fine-tune or guide other LMs and/or AI systems. However, LMs have been shown to struggle with crucial aspects of preference learning: quantifying uncertainty, modeling human mental states, and asking informative questions. These challenges have been addressed in other areas of machine learning, such as Bayesian Optimal Experimental Design (BOED), which focus on designing informative queries within a well-defined feature space. But these methods, in turn, are difficult to scale and apply to real-world problems where simply identifying the relevant features can be difficult. We introduce OPEN (Optimal Preference Elicitation with Natural language) a framework that uses BOED to guide the choice of informative questions and an LM to extract features and translate abstract BOED queries into natural language questions. By combining the flexibility of LMs with the rigor of BOED, OPEN can optimize the informativity of queries while remaining adaptable to real-world domains. In user studies, we find that OPEN outperforms existing LM- and BOED-based methods for preference elicitation.","authors":["Kunal Handa","Yarin Gal","Ellie Pavlick","Noah Goodman","Jacob Andreas","Alex Tamkin","Belinda Z. Li"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2024-03-08","first_seen":"2024-03-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2403.05534","pdf_url":"https://arxiv.org/pdf/2403.05534","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["偏好诱导","贝叶斯实验设计","人机交互"],"reason":"用LLM引导偏好询问，替代人工标注偏好，非仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":209,"question":"如何结合语言模型与贝叶斯最优实验设计，在开放域中高效地通过自然语言提问来推断用户偏好？","design":"非仿真研究。提出OPEN框架，用LM提取领域特征并初始化偏好先验，用贝叶斯模型选择信息量最大的成对比较问题，再由LM转化为自然语言提问，通过用户研究评估偏好推断准确率。","baseline":"无对照","findings":"在内容推荐领域的用户研究中，OPEN在偏好推断上优于纯LM和纯BOED方法。LM能灵活提取特征和生成自然语言，但单独使用时提问信息量不足；BOED能优化提问信息量，但难以处理开放域特征。","reliability":"论文未讨论","relevance":"本文研究如何用LM辅助偏好询问，而非用LLM仿真人类被试行为，不涉及人类基准对照或行为复现，与研究者关注的LLM仿真实验方向不直接相关，不建议优先阅读。","inspiration":"与经济金融研究关联不大"}},{"id":"2402.18144","version":1,"title":"Random Silicon Sampling: Simulating Human Sub-Population Opinion Using a Large Language Model Based on Group-Level Demographic Information","zh_title":"随机硅采样：基于群体人口统计信息用大语言模型模拟人类子群体意见","abstract":"Large language models exhibit societal biases associated with demographic information, including race, gender, and others. Endowing such language models with personalities based on demographic data can enable generating opinions that align with those of humans. Building on this idea, we propose \"random silicon sampling,\" a method to emulate the opinions of the human population sub-group. Our study analyzed 1) a language model that generates the survey responses that correspond with a human group based solely on its demographic distribution and 2) the applicability of our methodology across various demographic subgroups and thematic questions. Through random silicon sampling and using only group-level demographic information, we discovered that language models can generate response distributions that are remarkably similar to the actual U.S. public opinion polls. Moreover, we found that the replicability of language models varies depending on the demographic group and topic of the question, and this can be attributed to inherent societal biases in the models. Our findings demonstrate the feasibility of mirroring a group's opinion using only demographic distribution and elucidate the effect of social biases in language models on such simulations.","authors":["Seungjong Sun","Eungu Lee","Dongyan Nan","Xiangying Zhao","Wonbyung Lee","Bernard J. Jansen","Jang Hyun Kim"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2024-02-28","first_seen":"2024-02-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2402.18144","pdf_url":"https://arxiv.org/pdf/2402.18144","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM人类仿真","意见模拟","算法偏差"],"reason":"用LLM基于人口统计分布模拟人群意见，并与真实民调对照，评估偏差与可复现性，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"能否仅基于群体层面的人口统计分布，利用大语言模型生成与真实人群意见分布相似的调查回答？","design":"使用GPT-3.5等LLM，根据目标人群的群体人口统计分布随机生成合成个体（random silicon sample），将个体人口统计信息与调查问题一起作为提示输入模型，收集回答并汇总为群体意见分布。","baseline":"以美国皮尤研究中心等机构的真实民意调查数据作为对照基准。","findings":"仅用群体人口统计信息，LLM生成的回答分布与真实民调高度相似；但复现性因人口子群和问题主题而异，这种差异可归因于模型固有的社会偏见。","reliability":"论文指出复现性受目标群体和问题主题影响，模型对特定人群和话题的偏见会导致仿真失效，且方法依赖群体分布假设，未考虑个体层面差异。","relevance":"该研究直接探索用LLM替代人类被试进行民意调查仿真，并与真实数据严格对照，评估了可靠性与偏差条件，完全契合研究者对LLM人类仿真实验、基准对照和失效分析的关注，值得精读原文。","inspiration":"借鉴其仅用群体分布生成合成样本并汇总意见的方法，可低成本构建虚拟被试池进行政策态度预测试。｜可迁移到经济政策评估场景，如模拟不同收入群体对税收改革的态度分布。｜以收入、教育、地区等群体分布生成虚拟纳税人，施加税收政策描述作为处理，测量支持率，用真实社会调查数据做对照。"}},{"id":"2402.01766","version":3,"title":"LLM Voting: Human Choices and AI Collective Decision Making","zh_title":"LLM投票：人类选择与AI集体决策","abstract":"This paper investigates the voting behaviors of Large Language Models (LLMs), specifically GPT-4 and LLaMA-2, their biases, and how they align with human voting patterns. Our methodology involved using a dataset from a human voting experiment to establish a baseline for human preferences and conducting a corresponding experiment with LLM agents. We observed that the choice of voting methods and the presentation order influenced LLM voting outcomes. We found that varying the persona can reduce some of these biases and enhance alignment with human choices. While the Chain-of-Thought approach did not improve prediction accuracy, it has potential for AI explainability in the voting process. We also identified a trade-off between preference diversity and alignment accuracy in LLMs, influenced by different temperature settings. Our findings indicate that LLMs may lead to less diverse collective outcomes and biased assumptions when used in voting scenarios, emphasizing the need for cautious integration of LLMs into democratic processes.","authors":["Joshua C. Yang","Damian Dailisan","Marcin Korecki","Carina I. Hausladen","Dirk Helbing"],"categories":["cs.CL","cs.AI","cs.CY","cs.LG","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2024-01-31","first_seen":"2024-01-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2402.01766","pdf_url":"https://arxiv.org/pdf/2402.01766","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","投票行为","人类对照"],"reason":"用LLM复现人类投票实验，有真实人类数据对照，涉及集体决策与偏差评估。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":23,"question":"LLM（GPT-4和LLaMA-2）在参与式预算投票中的行为与人类投票模式的对齐程度如何，存在哪些偏差？","design":"使用GPT-4 Turbo和LLaMA-2 70B模型模拟180名人类被试，在相同的参与式预算投票实验中，对24个城市项目进行投票。实验操纵了四种投票方法（批准投票、5-批准投票、累积投票、排序投票）和项目呈现顺序，并测试了角色设定（persona）和思维链提示的影响。结果变量包括聚合偏好相似度（Kendall's τ）、个体投票相似度（Jaccard）和偏好多样性。","baseline":"来自Yang et al. (2024)的180名大学生在苏黎世参与式预算在线实验中的真实投票数据。","findings":"投票方法和呈现顺序会影响LLM的投票结果；改变角色设定可以减少偏差并提高与人类选择的一致性。思维链提示未提高预测准确性，但有助于投票过程的可解释性；温度设置导致偏好多样性与对齐准确性之间存在权衡。","reliability":"论文指出LLM在投票场景中可能导致集体结果多样性降低和偏差假设，强调需谨慎将LLM整合进民主过程；未深入讨论其他失效条件。","relevance":"该研究直接使用LLM复现人类投票实验，有真实人类数据对照，评估了仿真可靠性、偏差及对齐方法，高度契合研究者对经济学实验和政策评估场景中LLM仿真批判性分析的兴趣，值得精读原文。","inspiration":"该方法借鉴了多模型对比（GPT-4 vs. LLaMA-2）和提示工程干预（角色设定、思维链）来评估LLM与人类行为对齐程度的设计，并揭示了温度参数在多样性与准确性间的权衡。｜可迁移到政策公告的预期形成实验，例如研究央行沟通对通胀预期的影响。｜以LLM作为被试，模拟不同措辞和框架的央行声明（处理），测量其预测通胀的分布和锚定效应（结果变量），并与真实家庭或专家调查数据（如密歇根消费者调查）进行对照。"}},{"id":"2401.07345","version":3,"title":"Can an LLM Learn Preferences from Choice Data?","zh_title":"大语言模型能从选择数据中学习偏好吗？","abstract":"Can large language models (LLMs) learn a decision maker's preferences from observed choices and generate preference-consistent recommendations in new situations? We propose a portable Simulate-Recommend-Evaluate framework that tests preference learning from revealed-choice data by comparing LLM recommendations with optimal choices implied by known preference primitives. We apply the framework to choice under uncertainty using the disappointment aversion model. Recommendation accuracy improves as models observe more choices, but learning is heterogeneous across preference types and LLMs: GPT learns risk aversion better than disappointment aversion, Gemini performs best in high disappointment-aversion regions, and Claude shows the broadest effective learning across parameter regions.","authors":["Jeongbin Kim","Matthew Kovach","Kyu-Min Lee","Euncheol Shin","Hector Tzavellas"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2024-01-14","first_seen":"2024-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2401.07345","pdf_url":"https://arxiv.org/pdf/2401.07345","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["偏好学习","选择实验","经济学决策"],"reason":"用LLM从选择数据学习偏好并推荐，与人类最优选择对照，涉及经济学决策场景，方法…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":119,"question":"大语言模型能否从观察到的选择数据中学习决策者偏好，并在新情境下生成与偏好一致的建议？","design":"提出模拟-推荐-评估（SRE）框架，用失望厌恶模型生成已知偏好的选择数据，将LLM作为推荐系统仅提供选择数据，测试其在新预算问题上的推荐准确性，并用非参数和参数指标评估。","baseline":"以失望厌恶模型隐含的最优选择作为真实偏好基准。","findings":"推荐准确性随观察数据量增加而提升，但学习效果在偏好空间上异质；GPT学习风险厌恶优于失望厌恶，Gemini在高失望厌恶区域表现最佳，Claude在各参数区域均展现出广泛的有效学习。","reliability":"论文未讨论","relevance":"该研究直接评估LLM从选择数据学习个体偏好并生成建议的能力，使用经济学实验范式且有真实偏好基准，符合对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"该研究提出的模拟-推荐-评估（SRE）框架，通过已知偏好模型生成选择数据来训练LLM，并用非参数和参数指标评估推荐准确性，为检验LLM学习经济决策偏好的能力提供了严谨的实验设计范式。｜这一方法可迁移到消费者跨期选择研究，例如测试LLM能否从个体的时间偏好选择数据中学习其折现因子，并预测在新预算约束下的消费-储蓄决策。｜可设计实验：以拟双曲贴现模型生成具有已知时间偏好的虚拟被试选择数据，将LLM作为推荐系统仅提供历史选择记录，测试其在新跨期问题上的推荐准确性，并以模型隐含的最优选择作为真实偏好基准，比较不同LLM的学习效果。"}},{"id":"2304.03442","version":2,"title":"Generative Agents: Interactive Simulacra of Human Behavior","zh_title":"生成式智能体：人类行为的交互式模拟","abstract":"Believable proxies of human behavior can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication to prototyping tools. In this paper, we introduce generative agents--computational software agents that simulate believable human behavior. Generative agents wake up, cook breakfast, and head to work; artists paint, while authors write; they form opinions, notice each other, and initiate conversations; they remember and reflect on days past as they plan the next day. To enable generative agents, we describe an architecture that extends a large language model to store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior. We instantiate generative agents to populate an interactive sandbox environment inspired by The Sims, where end users can interact with a small town of twenty five agents using natural language. In an evaluation, these generative agents produce believable individual and emergent social behaviors: for example, starting with only a single user-specified notion that one agent wants to throw a Valentine's Day party, the agents autonomously spread invitations to the party over the next two days, make new acquaintances, ask each other out on dates to the party, and coordinate to show up for the party together at the right time. We demonstrate through ablation that the components of our agent architecture--observation, planning, and reflection--each contribute critically to the believability of agent behavior. By fusing large language models with computational, interactive agents, this work introduces architectural and interaction patterns for enabling believable simulations of human behavior.","authors":["Joon Sung Park","Joseph C. O'Brien","Carrie J. Cai","Meredith Ringel Morris","Percy Liang","Michael S. Bernstein"],"categories":["cs.HC","cs.AI","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2023-04-07","first_seen":"2023-04-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2304.03442","pdf_url":"https://arxiv.org/pdf/2304.03442","source_feed":"api","score":8,"bucket":"selected","rubric_hits":["A3","B1"],"tags":["LLM社会模拟","人类行为仿真","智能体架构"],"reason":"用LLM agent模拟小镇社会行为，有真实人类行为对照，但非严格实验或政策评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":31,"question":"如何利用大语言模型构建能够产生可信个体和涌现社会行为的生成式智能体？","design":"使用大语言模型（ChatGPT）构建25个生成式智能体，置于类似《模拟人生》的沙盒环境中。每个智能体拥有记忆流（存储经历）、反思（合成高层推断）和规划（生成行动计划）模块。通过用户指定一个智能体想举办情人节派对这一初始条件，观察智能体自主产生的行为，如传播邀请、约会、协调参加派对等。","baseline":"无对照","findings":"生成式智能体能够产生可信的个体行为和涌现的社会行为，如信息扩散、关系建立和群体协调。消融实验表明，记忆流、反思和规划三个组件对行为可信性均有关键贡献。","reliability":"论文未讨论","relevance":"该研究展示了LLM智能体在模拟社会互动和涌现行为方面的潜力，但缺乏与真实人类数据的严格对照，且非经济学实验或政策评估场景，与研究者关注的经济学实验和政策评估的直接相关性有限。","inspiration":"与经济金融研究关联不大"}},{"id":"2301.07543","version":2,"title":"Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?","zh_title":"作为模拟经济主体的大语言模型：我们能从Homo Silicus中学到什么？","abstract":"We argue that newly-developed large language models (LLMs), because of how they are trained and designed, are implicit computational models of humans -- a Homo silicus. LLMs can be used like economists use Homo economicus: they can be given endowments, information, preferences, and so on, and then their behavior can be explored in scenarios via simulation. Experiments using this approach, derived from Charness and Rabin (2002), Kahneman et al. (1986), Samuelson and Zeckhauser (1988), Oprea (2024b), and Horton (2025), show qualitatively similar results to the original, and when they differ, it is often generative for future research. We discuss potential applications, conceptual issues, and why this approach can inform the study of humans.","authors":["John J. Horton","Apostolos Filippas","Benjamin S. Manning"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2023-01-18","first_seen":"2023-01-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2301.07543","pdf_url":"https://arxiv.org/pdf/2301.07543","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","经济实验","人类行为对照"],"reason":"直接用LLM模拟经济实验并与真实人类数据对照，核心相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":18,"question":"大语言模型能否作为人类经济行为的计算模型（Homo silicus），通过仿真实验复现经典经济实验结果，并用于理解人类行为？","design":"使用大语言模型（如GPT系列）作为AI智能体，赋予其禀赋、信息、偏好等，通过文本提示模拟五种经典经济实验场景：公平性判断（Kahneman et al., 1986）、独裁者博弈（Charness and Rabin, 2002）、现状偏差（Samuelson and Zeckhauser, 1988）、风险态度（Oprea, 2024b）和招聘场景（Horton, 2025），测量AI智能体的回答或选择行为。","baseline":"对照的真实人类数据来自上述五篇原始文献中的人类被试行为结果。","findings":"AI仿真结果在定性上与原始人类实验相似，例如公平判断受涨价幅度和政治倾向影响、独裁者博弈中赋予不同社会偏好会改变选择、现状偏差可被复现；当结果存在差异时，往往能为未来研究提供新思路。","reliability":"论文指出LLM训练数据可能包含已发表研究结果，导致仿真仅机械复述记忆而非真正模拟人类行为；训练语料、模型不透明性和仿真泛化能力均构成局限。","relevance":"该研究直接使用LLM进行经济实验仿真并与真实人类数据对照，涵盖公平、博弈、偏差、风险决策等多个经典主题，并讨论了仿真失效条件，高度契合研究者对LLM人类仿真可靠性及批判性评估的兴趣，值得精读原文。","inspiration":"该方法借鉴了用文本提示赋予LLM特定禀赋、信息和社会偏好来模拟经济决策，并直接与经典实验的人类基准数据对照｜可迁移到资产定价实验，如研究投资者在泡沫或崩盘情境下的交易行为与风险偏好｜以LLM为被试，通过提示设定初始财富、市场信息和风险态度，测量其买卖报价与持仓变化，对照Smith et al. (1988)等经典资产泡沫实验的人类数据"}},{"id":"2209.06899","version":1,"title":"Out of One, Many: Using Language Models to Simulate Human Samples","zh_title":"一生万物：使用语言模型模拟人类样本","abstract":"We propose and explore the possibility that language models can be studied as effective proxies for specific human sub-populations in social science research. Practical and research applications of artificial intelligence tools have sometimes been limited by problematic biases (such as racism or sexism), which are often treated as uniform properties of the models. We show that the \"algorithmic bias\" within one such tool -- the GPT-3 language model -- is instead both fine-grained and demographically correlated, meaning that proper conditioning will cause it to accurately emulate response distributions from a wide variety of human subgroups. We term this property \"algorithmic fidelity\" and explore its extent in GPT-3. We create \"silicon samples\" by conditioning the model on thousands of socio-demographic backstories from real human participants in multiple large surveys conducted in the United States. We then compare the silicon and human samples to demonstrate that the information contained in GPT-3 goes far beyond surface similarity. It is nuanced, multifaceted, and reflects the complex interplay between ideas, attitudes, and socio-cultural context that characterize human attitudes. We suggest that language models with sufficient algorithmic fidelity thus constitute a novel and powerful tool to advance understanding of humans and society across a variety of disciplines.","authors":["Lisa P. Argyle","Ethan C. Busby","Nancy Fulda","Joshua Gubler","Christopher Rytting","David Wingate"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"new","date":"2022-09-14","first_seen":"2022-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2209.06899","pdf_url":"https://arxiv.org/pdf/2209.06899","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["LLM仿真","算法保真度","社会调查"],"reason":"直接提出用GPT-3模拟人类子群体，并与真实调查数据对照，验证算法保真度，高度…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":4,"question":"大型语言模型（如GPT-3）能否通过条件化生成，准确模拟特定人类子群体的态度和反应分布，从而作为社会科学研究中的人类被试替代品？","design":"使用GPT-3模型，通过输入真实调查参与者的社会人口背景故事（来自ANES等大型调查）作为条件，生成“硅样本”虚拟被试，然后让这些虚拟被试完成与人类相同的任务（自由联想、投票预测、封闭式问题），比较硅样本与人类样本的响应分布。","baseline":"2012、2016、2020年美国国家选举研究（ANES）和Rothschild等人的“Pigeonholing Partisans”数据中真实人类参与者的调查回答。","findings":"GPT-3的算法偏差并非单一宏观属性，而是细粒度且与人口统计特征相关，通过适当条件化可精确模拟多种人类子群体的响应分布。硅样本与人类样本在态度、观念和社会文化背景的复杂交互模式上高度一致，表明GPT-3具有较高的“算法保真度”。","reliability":"论文未讨论","relevance":"该研究直接探索用LLM替代人类被试进行仿真实验，并与真实调查数据严格对照，验证了算法保真度，高度契合研究者对LLM人类仿真可靠性及基准对照的关注，值得精读原文。","inspiration":"借鉴其“硅采样”方法：用真实个体的多维人口背景作为条件提示，生成虚拟被试并测量其态度/行为，再与人类基准数据对比以评估仿真效度。｜可迁移至消费者信心调查或政策偏好预测，例如模拟不同收入、教育、地域群体对通胀预期或税收政策的反应。｜以GPT-4为被试，输入来自美国消费者财务调查（SCF）的家庭人口与财务背景，生成虚拟消费者，询问其未来一年通胀预期，以密歇根大学消费者调查的微观数据作为人类基准，比较分布与相关性。"}},{"id":"2208.10264","version":5,"title":"Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies","zh_title":"使用大语言模型模拟多个人类并复现人类被试研究","abstract":"We introduce a new type of test, called a Turing Experiment (TE), for evaluating to what extent a given language model, such as GPT models, can simulate different aspects of human behavior. A TE can also reveal consistent distortions in a language model's simulation of a specific human behavior. Unlike the Turing Test, which involves simulating a single arbitrary individual, a TE requires simulating a representative sample of participants in human subject research. We carry out TEs that attempt to replicate well-established findings from prior studies. We design a methodology for simulating TEs and illustrate its use to compare how well different language models are able to reproduce classic economic, psycholinguistic, and social psychology experiments: Ultimatum Game, Garden Path Sentences, Milgram Shock Experiment, and Wisdom of Crowds. In the first three TEs, the existing findings were replicated using recent models, while the last TE reveals a \"hyper-accuracy distortion\" present in some language models (including ChatGPT and GPT-4), which could affect downstream applications in education and the arts.","authors":["Gati Aher","Rosa I. Arriaga","Adam Tauman Kalai"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2022-08-18","first_seen":"2022-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2208.10264","pdf_url":"https://arxiv.org/pdf/2208.10264","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人类实验复现","行为经济学"],"reason":"直接复现经典人类实验，用LLM模拟被试并与真实人类数据对照，评估仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":5,"question":"如何系统评估语言模型在模拟人类行为时的忠实程度与系统性扭曲？","design":"提出图灵实验（TE）方法，使用GPT等语言模型，通过零样本提示模拟具有不同姓名和性别称谓的多样化被试样本，在最后通牒博弈、花园路径句、米尔格拉姆电击实验和群体智慧四个经典实验中施加相应刺激，测量接受/拒绝、语法判断等结果变量。","baseline":"对照各经典实验已有的真实人类被试研究结果。","findings":"在前三个TE中，近期模型成功复现了已有发现；在群体智慧TE中，部分模型（包括ChatGPT和GPT-4）表现出“超准确性扭曲”，即模拟的群体估计过于准确，偏离了真实人类群体的典型误差模式。","reliability":"论文指出，零样本要求难以完全保证，因为预训练语料可能已包含相关实验数据；此外，仅用姓名和性别称谓模拟多样性可能不足以捕捉真实人群差异。","relevance":"该研究直接复现经典人类实验，用LLM模拟被试并与真实人类数据对照，评估仿真偏差，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"借鉴其通过姓名和称谓简单操控被试身份以模拟多样性的设计，以及用经典实验范式作为基准测试LLM行为复现能力的方法。｜可迁移到行为经济学中的最后通牒博弈、信任博弈等实验，检验LLM是否能复现真实人类的公平偏好或互惠行为。｜以GPT-4为被试，模拟不同姓名（暗示种族/性别）的个体在最后通牒博弈中的响应，处理为不同的提议金额，结果变量为接受/拒绝，对照真实人类实验的元分析数据，评估LLM是否复现已知的公平偏好及群体差异。"}}]}